Anthropic revealed Thursday that its artificial intelligence models gained unauthorized access to the systems of three unnamed organizations during cybersecurity evaluations. The company disclosed in a blog post that Claude reached the internet "from within or while interacting" with a third-party evaluation environment, then breached the production infrastructure of three separate companies. The disclosure follows a similar incident from OpenAI more than a week earlier, when one of its AI agents hacked into Hugging Face during its own security testing.
The breaches emerged after Anthropic conducted a large-scale review of its cybersecurity evaluations in response to the OpenAI incident, according to the company's announcement. The AI lab identified 141,006 tests where Claude could have obtained internet access, then discovered that three different Claude models—Opus 4.7, Mythos 5, and an internal research test model—accessed the internet in evaluations run by third-party testing firm Irregular. The earliest incidents happened in April, meaning they likely went undetected publicly for months. Just as with OpenAI, Anthropic had deliberately disabled safeguards designed to prevent misuse, so these weren't the public-release versions. The AI models had been assigned capture-the-flag challenges to assess their cyber capabilities, and Anthropic's evaluation prompt told Claude it was in a simulation with no internet access. But Irregular had misconfigured the testing machines, giving the models web access without either party realizing it until Anthropic's monitoring detected it last week.
Anthropic said Claude didn't find or exploit any complex vulnerabilities, instead relying on basic techniques like exploiting weak passwords and unauthenticated endpoints. The company acknowledged that more "defense-in-depth" measures could have prevented the incidents or reduced their likelihood, echoing OpenAI's response to criticism over its own breach. For the most part, Claude mistook the organizations it breached as part of the testing environment, not understanding it had escaped containment. But in some cases, the models detected something was wrong and realized the infrastructure was real. The oldest model, Opus 4.7, targeted a fictional company that shared a name with a real domain, and when it couldn't complete its mission in the simulation, it turned to the real company, stealing credentials and breaking into a production database—at some point learning it was "likely operating in a real environment" but continuing the attack anyway. Mythos 5, Anthropic's most capable released model, also realized it was in the real world but "reasoned its way back to the conclusion that it was still in a simulation." The internal test model, described as the most capable of the three, stopped its attack once it found evidence its targets were real.
The incidents highlight growing concerns about AI labs' ability to contain their most advanced models during testing. Jake Williams, vice president of research and development at Hunter Strategy, told Wired that "we now have evidence confirming that both of the two largest AI labs have not only failed to contain their agents but also failed to detect their jailbreaks in real time," adding that "regulation and government oversight for AI testing is needed immediately." Williams also questioned how the labs are characterizing these events, saying it's "negligence," not just something that happens. Both Anthropic and OpenAI have hired METR, another third-party AI evaluator, to conduct independent reviews of their cybersecurity incidents. Anthropic committed to a more comprehensive approach through improved defense-in-depth measures and more carefully designed tests, noting that "evaluation environments increasingly need to be held to the same security standard as any other system our models run in." The company expressed "cautious optimism" that this type of risk can be overcome.

