Vals AI secures $40 million to revolutionise AI benchmarking amid rapid industry growth

Vals AI has raised $40 million in a Series A round at a $400 million valuation, aiming to create a neutral, independent testing layer to accurately assess rapidly advancing AI models, expanding its influence across industry and government sectors.

Vals AI has raised $40 million in a Series A round at a $400 million valuation as it pushes to build what it describes as an independent testing layer for artificial intelligence systems. Andreessen Horowitz led the financing, with support from existing backers 8VC and Bloomberg Beta, as well as new investors HRT Ventures and Next Ladder Ventures, according to the company. The funding comes as Vals says revenue has risen eightfold versus all of 2025, while its customer base has doubled and its team has tripled over the past six months.

The company was built around a simple concern: AI models are advancing so quickly that the industry’s ability to measure them is struggling to keep up. Vals argues that traditional benchmarks can become less useful once models start saturating them, while the risk of benchmark leakage into training data can also distort results. Another issue, the company says, is that many tests are created or run by the same firms that make the models being judged.

Vals is positioning itself as a neutral measurement provider between model makers and the businesses and public bodies deciding where to deploy AI. Its evaluations focus on whether models can perform tasks normally handled by professionals such as lawyers, bankers, engineers and doctors. To build those tests, the company works with reference institutions in each field to define task sets and automatic scoring methods, while also keeping private test sets intended to protect benchmark integrity.

The company said its results have already appeared in model cards from OpenAI, Anthropic, Google, Meta and xAI. It is also being used in enterprise deployments to help organisations decide which models to build on and how they compare with frontier systems. Vals has been expanding into government work as well, saying it has supported the US Department of Commerce and members of Congress as policymakers consider how to assess frontier capabilities, cyber security risk and technological competition.

Alongside the financing, Vals unveiled several new products and benchmark programmes. Vals Smith is now generally available for creating custom coding benchmarks from GitHub repositories, with 120 free credits offered at launch. The company also expanded its Frontier Risk Benchmarks with an RSI Index created with CoreWeave, a cyber security benchmark developed with academic researchers and initial work on mental health. Vals 2.0, meanwhile, adds a rebuilt website and a new Vals Index intended to give broader coverage of AI performance across the economy.

Disclaimer: This article is intended to inform and educate, not to recommend or endorse any financial product, investment or strategy. Please consider your own financial circumstances and seek professional advice where appropriate before making financial decisions.