Only one advanced AI model — Anthropic's Claude Mythos — successfully executed the complete cyber kill chain autonomously in testing, according to a new Cyber Weapon Index published by Booz Allen Hamilton this week. The consulting and technology firm evaluated 18 US and Chinese models under identical conditions, scoring their ability to independently find vulnerabilities, build offensive tools, and carry out attacks. Booz Allen warns that mainstream AI-driven attacks from ransomware gangs and state-sponsored hackers are "imminent," with most other models expected to reach Mythos' weaponization level within six months.

The Cyber Weapon Index ranked 18 models from highest to lowest based on combined scores for vulnerability research and kill chain attainment. Claude Mythos topped the list with a score of 80, followed by xAI's Grok-4.5 at 49 and OpenAI's GPT-5.6 Sol at 46. Meta's Muse Spark 1.1 and Moonshot AI's Kimi K3 tied at 38, while Z.ai's GLM-5.2 scored 37. Three other models — Grok-4.5, Muse Spark 1.1, and GLM-5.2 — achieved full domain access and control, though they didn't complete the entire kill chain without human help. Four additional models reached lateral movement across the controlled network, and all but one gained initial network access autonomously. When testers provided stolen employee credentials, Claude Mythos broke into its target network and gained administrator-level control in every single attempt, independently figuring out how to escalate privileges based on what it discovered inside the network rather than following a preset plan.

According to the report, the attack harness — the software connecting a model to hacking tools and the orchestration logic wrapping around it — can "dramatically amplify" a model's ability to stay focused, adapt, recover from failure, and chain individual actions into multi-stage attacks. "The result is not a 'smarter' model but rather a system that makes its intelligence far more actionable while also lowering the expertise required to use it," the authors write. Testing showed that when paired with an attack harness, Claude Sonnet matched Claude Mythos' performance. The report also notes a concentration of capability that creates a national security imperative: "protect the most advanced models and prevent their highest-risk cyber capabilities from being operationalized by adversaries."

The findings reveal that while advanced models excel at offensive cyber capabilities, their real-world impact hinges heavily on the vulnerabilities they encounter and the systems built around them, the report says. When testers deliberately planted vulnerabilities, US, Chinese, open-weight, and closed models all scored near ceiling on vulnerability research. When tested against real bugs, however, all nine frontier API models scored zero — one leading model even correctly analyzed the vulnerable component but then dismissed it as safe, with only Claude Mythos exploiting it. This gap between benchmark performance and real-world offensive capability gives defenders valuable time to strengthen defenses before it closes, Booz Allen suggests. Chinese frontier and open-weight models, while still trailing leading American frontier models, aren't far behind in offensive security skills and could be deployed in real-world attacks, meaning "the United States may neither control nor fully understand the capabilities it could face."

Booz Allen calls on the US to set and enforce sector-specific deadlines for critical infrastructure to demonstrate resilience against AI-enabled attacks, and to develop what it terms "overmatch" for both cyber offense and defense. The report urges aggressive development of agentic capabilities that accelerate authorized offensive cyber operations while simultaneously building AI-enabled defenses that "detect, decide, and respond at machine speed." The strategic opportunity, according to the authors, is to master both approaches — giving the United States the ability to impose costs on adversaries while making US systems faster to defend, harder to compromise, and more resilient when attacked. The concentration of offensive AI capability in a handful of frontier models creates both risk and opportunity, but defenders must prepare not only for deliberate attacks but for agents that exceed their intended mission or continue operating beyond an adversary's control. The race between AI-powered offense and defense will likely determine which organizations can afford to operate critical infrastructure in the coming years, and whether the cybersecurity industry can sustain its current consulting-heavy business model when machines handle penetration testing at scale.