A startup that deliberately crashes software systems to test their strength is now using artificial intelligence to speed up the process. Gremlin announced Foresight AI, a new add-on that autonomously breaks software infrastructure to uncover hidden bugs before they appear in production, according to a report published by The Register on October 8, 2026. The technology automates chaos engineering, a practice popularized by Netflix over a decade ago in which companies intentionally introduce controlled failures into working systems to test their resilience.

The AI tool handles many preparatory and post-testing tasks that used to be done manually, significantly accelerating the testing cycle, according to Gremlin founder Kolton Andrus, who previously performed chaos engineering at Netflix. Gremlin's approach relies on its Failure Atlas, a repository containing millions of chaos engineering experiments the company ran over the past decade on tens of thousands of systems. The company developed a harness that connects large language models to the Failure Atlas, enabling the AI to generate more technically grounded answers about what failed rather than drawing on random information from the internet. Gremlin uses a variety of closed and open LLMs depending on which performs best at any given time, with the Atlas keeping the models focused and minimizing hallucinations.

Once Foresight AI induces a system failure, it determines the root cause, generates code or configuration changes it believes will fix the problem, and then either applies the solution or prepares a report for a site reliability engineer to review, the report states. The system operates as an agentic loop of test-and-replace, keeping humans involved at critical junctures, and runs until the failure no longer occurs. Gremlin's current customers tend to be larger compute-heavy enterprises, many in financial services, retail, and enterprise SaaS, who use the platform to test disaster recovery plans, ensure correct operation under less-than-ideal conditions, and verify that Kubernetes scales properly. "A lot of AI solutions are, 'Hey, we took a guess. Here you go. Good luck,'" Andrus said, contrasting Gremlin's deliberate approach with diagnostic tools that only react after systems have already crashed.

The shift toward using AI to write applications and deploy infrastructure creates its own complications, Andrus explained. AI operates at such velocity that human reviewers can't keep pace, and AI can make careless errors anywhere in the process. Complex distributed systems prove particularly difficult to diagnose when they fail because some bugs can't be caught with routine integration and unit testing—they only surface when faults like network latency, memory exhaustion, or other conditions expose weaknesses in distributed architectures. Any service that promises to break systems for the greater good will need sign-off at the highest corporate levels, noted Intellyx analyst Jason English, adding that "we need to change our mindset about risk" for the technology to work effectively. For companies betting their uptime on AI-driven fault injection, the pitch is simple: find out what breaks before customers do. Organizations willing to embrace controlled chaos may gain resilience that manual testing can't match, though success hinges on whether executive leadership trusts machines to tip over production systems in search of hidden flaws.