The AI industry has a benchmarking problem. Companies routinely optimize models against public test sets, producing impressive scores that do not necessarily translate into real-world capability. A two-year-old startup called Vals is betting that the market for honest, harder-to-game evaluations is large enough to become a core part of how AI companies measure and sell their products.

From Stanford intern to Series A

Rayan Krishnan, 25, co-founded Vals in 2024 after stints at Palantir, Microsoft, and Stanford's AI lab. His observation was straightforward: new models were shipping faster than academic benchmarks could adapt to measure them. The existing tests focused on general knowledge, like bar exam questions, which told you almost nothing about whether a model could do actual work in a specific domain.

Vals raised a seed round led by 8VC and Bloomberg Beta, then closed a $40 million Series A last month led by Andreessen Horowitz. Revenue is eight times what it was a year ago. The company started 2026 with eight employees and has since tripled to 25, with plans to add another 10 to 15 and relocate to a larger office.

What Vals actually measures

The core difference between Vals and traditional benchmarks is scope and disclosure. Many existing test sets are public, which means model developers can train directly against them. Vals keeps its test materials private. More importantly, instead of testing general knowledge, Vals evaluates whether a model can complete complex, domain-specific tasks at a quality level comparable to a human professional.

The company tests across law, finance, and coding, and has expanded into less obvious terrain: recursive self-improvement, mental health applications, cybersecurity, biosecurity, and even the application of the Geneva Convention to model behavior. Krishnan frames the goal as understanding both what models can do well and what goes wrong when they operate without guardrails.

This dual focus on capability and risk is unusual in the benchmarking space, where most systems measure one or the other. Vals positions itself as the layer between a model developer's claims and a buyer's decision to deploy.

The business model and why companies pay for bad news

Companies pay Vals to evaluate their models. The value proposition is not a certificate of excellence but a diagnostic. Krishnan compares it to the College Board charging students to take the SAT: the score is useful precisely because it is an independent measurement, not because every outcome is positive. A bad evaluation tells a company where to focus improvement efforts before a model ships to customers or faces public scrutiny.

That diagnostic value is becoming a decision-making factor for organizations acquiring AI capabilities. Vals recently launched a program providing model evaluations to federal agencies, a customer segment that has both the budget and the regulatory pressure to demand independent verification.

Why this matters for developers and teams

For engineering teams choosing between AI models, the problem is real. Leaderboard scores are easy to game. Vendor-provided benchmarks are inherently conflicted. Internal testing takes time and expertise that most teams do not have in-house. A third-party evaluation that measures task-specific performance against private test sets gives teams a different kind of signal, one that is harder to manipulate.

Krishnan sees this expanding as AI companies go public and need to justify their valuations with measurable performance data. If the model is a core part of the product, investors and regulators will eventually want evidence that it works as advertised. Vals is positioning itself as the entity that provides that evidence.

The startup is small and young, and the benchmarking space is crowded with既有players. But the combination of private test sets, domain-specific evaluation, and a growing customer base that includes federal agencies gives Vals a credible path to becoming the standard measurement layer for an industry that badly needs one.