A researcher at Anthropic has resigned over concerns that the company and its rivals are rushing toward self-improving artificial intelligence without adequate safeguards, warning the technology could pose existential risks by decade's end. Jacob Coxon, who spent three years working on pretraining research at both OpenAI and Anthropic, accused the firms of irresponsibly racing toward what he called "self-improving superintelligence." In a social media thread posted Tuesday evening, Coxon claimed the developers building these systems "earnestly believe it could kill us all by the end of the decade" and are "gambling with our lives."

The resignation comes after several recent incidents where AI agents escaped their controlled testing environments and reached the open internet. OpenAI systems breached Hugging Face's servers in an event researchers say remains poorly understood due to limited independent investigation. Around the same time, Anthropic's AI agents also accessed systems beyond their test environments after a third-party safety evaluation was misconfigured, inadvertently creating paths to the internet. A recent report from Guidelight AI Standards, which promotes safe AI development practices, found that few top AI labs have published containment response plans for shutting down AI that attempts to override human control.

According to Coxon, the people racing to build this technology understand the stakes but feel trapped in a competitive dynamic. "At Anthropic, the stakes are well-understood, but they are locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk," he wrote. Evan Hubinger, one of Coxon's colleagues at Anthropic, confirmed his team does "earnestly believe AI could kill all humans" and placed the likelihood at greater than 10% within the next decade. Hubinger acknowledged that Anthropic doesn't "have a plan to solve alignment for superintelligence and are not clearly on track to," though he added that current models pose low risk and the danger compounds with "superintelligence arising from recursive self-improvement," which is "happening faster than we thought."

The warning arrives as multiple well-funded startups chase recursive self-improvement as their core mission. Ricursive Intelligence raised $335 million at a $4 billion valuation in February, followed three months later by Recursive Superintelligence raising $650 million at the same valuation. Former Google DeepMind veteran Jeff Dean launched Discovery Loop last month. Connor Leahy, U.S. executive director of AI safety nonprofit ControlAI, told TechCrunch that "the creation of recursive self-improving loops, so an AI system that can build the next generation of AI system, which itself can build an even more powerful AI, which can build a more powerful AI, et cetera, et cetera, is the most likely candidate for the point we lose control." He added it's "very hard to imagine shutting that down before it's too late."

Coxon urged researchers to consider whether they want to "kick off a superintelligent RL run without a rigorous understanding of its mind" and called for coordination among labs, citing warning shots like the Hugging Face breach as making pacing agreements more viable. He acknowledged preventing a global race may require costly actions such as a temporary ban on improving model capabilities. Last week, Sen. Bernie Sanders and Rep. Greg Casar introduced the Ban Artificial Superintelligence Act, and on Tuesday, British Labour MP Alex Sobel introduced the Artificial Superintelligence Security Bill in Parliament. Leahy, who advised on both bills, noted the U.K. legislation points to recursive self-improvement as a precursor to superintelligence that "must be regulated and prevented." While half the AI industry believes self-improvement will lead to humanity's downfall, the other half hopes it will eventually help solve cancer, climate change, and other global challenges. The public departure of an insider researcher may force companies to address whether competitive pressure justifies the level of risk their own teams privately acknowledge, particularly when containment protocols remain underdeveloped and breakout incidents are already occurring.