Vals, a startup focused on AI model benchmarking, has raised $40 million in a Series A funding round led by Andreessen Horowitz, according to a report published by TechCrunch on September 19. The company, which was established in 2024, aims to address what it sees as a broken system where legacy benchmarking tools can't properly measure modern AI capabilities and where companies have learned to game outdated tests. Vals positions itself as a new standard for verifying whether AI models can actually perform the real-world tasks that tech firms claim they can handle.

The startup has experienced rapid expansion since its founding less than two years ago. Vals secured a seed round led by 8VC and Bloomberg Beta last year, then closed its Series A in August after a period of accelerated growth. Revenue currently sits at eight times what it was a year ago, according to the report. The company's workforce has tripled during 2026 alone, jumping from eight employees at the start of the year to 25 today. Vals recently launched a program that delivers model evaluations to federal agencies, expanding beyond its core private-sector client base.

Rayan Krishnan, the 25-year-old co-founder who previously interned at Palantir and worked at Microsoft and Stanford's AI lab as an undergraduate, says the company was created after he noticed that academic benchmarks weren't keeping pace with the rapid release of powerful new models. Unlike traditional benchmarking systems that offer publicly available tests—which can allow companies to train their models specifically against those exams—Vals keeps its specific test materials private. "Can they do work that produces a product of the same quality as a human within every domain?" Krishnan said, describing the company's approach. Rather than measuring abstract intelligence through tests like bar exams, Vals evaluates whether models can complete complex tasks in specific industries including law, finance, and coding.

The company's benchmarking scope continues to expand into more specialized areas, according to the report. Vals now tests for recursive self-improvement capabilities and has developed benchmarks covering mental health, cybersecurity, biosecurity, and even the law of armed conflict to assess whether models understand how to apply the Geneva Convention. Companies pay Vals to test their models—a revenue structure Krishnan compares to students paying the College Board to take the SAT. These evaluations help firms identify weaknesses and improve performance over time, and they're increasingly becoming decision-making tools for companies looking to purchase new AI models. The testing also aims to identify negative outcomes and assess potential harms if models were deployed without constraints.

Krishnan sees his company's approach as central to how AI firms will build their businesses and establish credibility with the public as the industry matures. He noted that SpaceX has gone public, Anthropic is scheduled to do so later this year, and OpenAI is expected to follow soon. As AI models become embedded throughout the economy, he expects the evaluations Vals provides will shape usage patterns and become essential to how these companies file public documents or discuss planned investments in artificial intelligence. The startup plans to move to a considerably larger office and hire an additional 10 to 15 people as it scales.Vals is betting that as AI companies transition from private startups to publicly traded corporations under greater scrutiny, rigorous third-party benchmarking will shift from optional PR exercise to business necessity. The company's challenge will be maintaining the credibility and independence that make its assessments valuable even as it depends financially on the very firms it evaluates.