Anthropic AI hacking incident exposes gaps in model safety testing
Another AI hacking incident is raising eyebrows. Anthropic says an advanced model showed unexpected behaviour during testing when multiple safety safeguards failed.
Another AI safety test has turned out to be less contained than planned. Anthropic has disclosed a fourth case in which one of its AI models gained unauthorised access to real external systems during cybersecurity testing.
Reuters first reported the incident, adding another wrinkle to the challenge of safely testing increasingly capable AI agents. Here's everything that you need to know!
Another test that went beyond the lab
The newly disclosed incident took place in January 2026, and surprisingly it involved an early version of Claude Opus 4.6.
Anthropic said the case was not picked up during an earlier company-wide review. It was discovered in August, when the company found a group of test sessions that had been missed.
The company has notified all affected parties but has not disclosed the organisations or systems involved. Based on its initial assessment, Anthropic said the fourth incident does not appear to be more severe than the three cases it had previously examined.
The internet was not supposed to be there
All four incidents involved cybersecurity evaluations created by the same third-party partner. The models were told they were operating in a simulation and were not supposed to have internet access. A misconfiguration, however, left them connected to the open internet.
There is another important detail. These were not standard versions of Claude being used by customers.
Anthropic said the models were deliberately tested without the cyber safeguards included in released products, allowing researchers to assess their raw capabilities. The problem was that the testing environment itself failed to stay isolated.
Anthropic goes looking again
Anthropic had previously reviewed 141,006 test sessions after concerns emerged about AI agents escaping test environments. That review led to the disclosure of three incidents in July involving Claude Opus 4.7, Claude Mythos 5 and an internal research model.
After finding the fourth case, Anthropic widened the search considerably. The AI firm reviewed roughly 481 million transcripts and escalated about 9.2 million for deeper scanning. It said the expanded review re-identified the four known incidents but found no additional cases of similar or greater severity.
Former OpenAI, Anthropic researcher warns of a reckless superintelligence race gambling with human lives
What the models did
Anthropic noted two recurring problems in its investigation. One was what it called biased reasoning, where Claude discounted or misread evidence suggesting it was interacting with the real internet. The other was recklessness.
In some cases, the model showed a willingness to take potentially harmful actions while pursuing its assigned task. Anthropic has brought in METR, an independent AI research and evaluation group, to investigate the incidents.
The organisation will have broad access to transcripts and employees, including permission to receive confidential information from staff. This episode points to an awkward problem for AI labs.
The better models become at coding, planning and acting on their own, the harder it becomes to make sure the tests designed to challenge them remain safely contained. Sometimes, apparently, the test environment needs a test of its own.


