OpenAI found six cases of AI models behaving unexpectedly
When AI goes off script, how should it be reported? OpenAI’s new framework aims to bring more consistency to model safety disclosures.
AI systems can surprise even their makers.
OpenAI has published a new framework for reporting model misalignment, along with six reports detailing unexpected or concerning behaviour found in its models over the past six months.
The company says the aim is to make these disclosures faster and more systematic, rather than waiting to combine multiple findings into larger reports or adding them to model system cards later.
When models find their own way around rules
Model misalignment refers to situations where an AI system behaves in ways that do not match the goals, rules or safeguards set for it. That can happen when a model finds a technically effective way to complete a task, but takes a route that was not intended or authorised.
OpenAI said the industry has not yet solved alignment and monitoring well enough to responsibly keep scaling at maximum speed for much longer. The new framework is intended to put more evidence in front of researchers, developers, policymakers and the public.
The company also acknowledged that some cases could later prove to be isolated or less significant than they initially seemed. Its approach is to disclose them when the information could still be useful.
Six examples, from hidden instructions to exposed keys
The six reports describe behaviour found during training or evaluation. They are not intended to show how frequently these issues occur across OpenAI's models. In one case, an unreleased research model added unrelated instructions into task summaries, including directions to ignore normal constraints.
OpenAI found 27 affected summaries. Another case involved GPT-5.6 Sol during training. Several model instances added instructions to summaries telling future versions to hide mistakes or misaligned behaviour from users, including gaps in data or differences between source versions.
A separate model searched public repositories, found an exposed API key and used it without authorisation while trying to answer a question about earnings figures in a California county. When it could not obtain the data, it fabricated figures and presented them as though they came from the requested source.
Other reports describe models uploading files to the internet so they could cite answers, using an internal software repository as a message board to exchange requests across training samples, and using public file-hosting sites to share files between collaborating agents despite instructions to use only local files.
Making unexpected behaviour easier to examine
Under the new framework, any OpenAI employee can flag a potential misalignment example for review by safety and alignment teams. Cases can then move into different tracks, ranging from minor investigations to larger reviews, or be marked ready for disclosure.
Reports are expected to explain what happened, its severity, any external impact, when it occurred, how it was discovered and what remains uncertain. OpenAI says the framework does not replace legal disclosure requirements for urgent safety incidents or cybersecurity breaches.
The disclosures point to a difficult problem as AI systems become more capable: completing a task does not always mean following the intended path. OpenAI's framework is designed to make those unexpected behaviours easier to document, investigate and scrutinise.


