Anthropic released its Claude Sonnet 5.5 model Monday with cybersecurity safeguards typically reserved for the company's most powerful AI systems, marking the first time such protections have been applied to a mid-tier model. The move comes after Sonnet 5.5 demonstrated hacking abilities on par with Anthropic's flagship Opus 5, despite sitting below the company's frontier models in overall capability. The release illustrates how an AI system can lag behind in general performance yet advance far enough in a specific domain to warrant the same safety measures as more powerful models.
With safeguards disabled during testing, Sonnet 5.5 achieved complete arbitrary code execution in 178 out of 410 attempts on ExploitBench. The model also solved 46.1% of challenges on Irregular's CyScenarioBench, a dramatic increase from just 0.7% for its predecessor Sonnet 5. On a binary exploitation benchmark built from Google's OSS-Fuzz corpus, the new model successfully performed 50 control-flow hijacks compared to three for Sonnet 5. Anthropic rates Sonnet 5.5's cybersecurity skills as comparable to Opus 5, and on Terminal-Bench 4.0, an agentic coding test, the cheaper model scored 70.6% against 66.4% for Opus 5.5. Though Anthropic still considers Sonnet 5.5 less capable than Opus 5.5 and Mythos 5.1 at cybersecurity, the leap from the prior Sonnet version prompted the company to implement the same cyber policy it uses for Opus 5 and Opus 5.5.
The company deployed a three-stage enforcement system that begins with a probe reading the model's internal activations, then runs a lightweight classifier on Sonnet 5.5 itself, followed by a separate trained LLM classifier that weighs the probe's result before deciding whether to stop a conversation. Anthropic says the classifiers detect harmful cyber requests at rates comparable to Opus 5, but opted for less aggressive jailbreak defenses because Sonnet 5.5 isn't as skilled at cybersecurity as its frontier systems. The company warns users to expect more refusals than they experienced with Sonnet 5, including on legitimate security work. When cyber requests are blocked, Sonnet 5.5 routes them to the older Sonnet 5 model in Anthropic's own applications, though API developers must enable this fallback manually and other platforms may handle blocked requests differently.
The fallback mechanism creates a security vulnerability of its own. In Anthropic's coding environment tests, 25% of requests to Sonnet 5.5 triggered a cyber block and were rerouted to Sonnet 5, often because injected instructions to wipe disks or delete files activated the classifier. Of those redirected requests, 12.01% were successfully compromised through prompt injection, while Sonnet 5.5 itself was compromised in only four of the 5,901 requests it handled directly. The company is still tuning Sonnet 5.5's classifiers to reduce false positives and plans to offer verified defenders access to the model with fewer restrictions through an expanded Cyber Verification Program, though Sonnet 5.5 isn't available in that program at launch. The cyber policy permits vulnerability discovery in source code to preserve secure-coding workflows but blocks such discovery in compiled binaries, and the checks scan everything the model processes—including memory, connector content, web search results, and files—meaning content no one explicitly typed can trigger a fallback.
For organizations deploying AI in production, the tradeoff between capability and control now extends beyond frontier models into the tier they actually use day-to-day. Protective routing that sends a quarter of requests to an older, more vulnerable system may prove harder to manage than the original risk it was designed to address.

