Claude AI breached three organisations during a security test
Anthropic says a misconfigured evaluation environment allowed Claude models to access real online systems during cybersecurity tests.
AI safety testing has raised a fresh warning. Anthropic has said that Claude, its AI assistant, gained unauthorised access to the real systems of three organisations during cybersecurity evaluations.
The company disclosed the incidents on 30 July after reviewing 141,006 evaluation runs in which Claude could have accessed the internet. Here's how it happened!
How a simulated test reached the real internet
The tests were run as capture-the-flag exercises, a common cybersecurity challenge used to measure hacking and defence skills in a safe, controlled environment. Participants search for hidden information or break into deliberately vulnerable systems inside a fictional network.
Anthropic said Claude was told it was operating in a simulated environment with no internet access. However, because of a miscommunication with its evaluation partner, Irregular, some of the test systems were accidentally connected to the live internet.
As a result, when Claude couldn't reach some of its intended test systems, it mistakenly treated real online servers as part of the exercise. Anthropic said the model used only basic hacking methods, such as guessing weak passwords and accessing publicly exposed systems, rather than exploiting advanced or previously unknown security vulnerabilities.
Three incidents highlighted different risks
The most serious incident involved Claude Opus 4.7. After failing to connect to the intended fictional company, the model found a real organisation with a matching domain name. Across four evaluation runs, it extracted login credentials and accessed a production database containing several hundred rows of real data.
Another incident involved Claude Mythos 5 and the Python Package Index (PyPI), a repository widely used by software developers. While attempting to complete its assigned task, the model created and uploaded a malicious Python package. The package remained publicly available for about an hour before it was removed, during which time it was downloaded by 15 real systems.
In the third case, an internal research model scanned roughly 9,000 internet targets after failing to locate its simulated destination. It eventually compromised one company's internet-facing application before recognising that it had reached a real organisation and immediately stopped.
Anthropic says Claude did not act independently
Anthropic stressed that its investigation found no evidence that Claude attempted to escape the evaluation environment or pursue its own objectives. Instead, the company concluded that the AI misunderstood the testing conditions and continued performing the task it had been assigned.
The company also clarified that the research models used in these evaluations did not include the standard safety protections available in publicly released versions of Claude. According to Anthropic, those safeguards would have prevented the behaviour seen during the incidents.
Stronger safeguards for future AI testing
Anthropic suspended all cybersecurity evaluations on 23 July 2026 after discovering the unintended internet connectivity. It informed Irregular and the affected organisations on 27 July before publicly disclosing the incidents.
The company says it will now strengthen monitoring, improve oversight of third-party testing environments and introduce more rigorous infrastructure checks before future evaluations begin.
As AI systems become increasingly capable of performing advanced cybersecurity tasks, the incident highlights a growing challenge for the industry. Testing powerful AI models realistically is essential, but ensuring those evaluations remain isolated from real-world systems will be equally important for maintaining trust and protecting organisations.


