Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 19, 2026 3 min read
Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking

AI benchmarking has become the primary way companies prove their models work and stand out from rivals. When the numbers look good, it usually results in better public relations. However, legacy systems struggle to measure modern capabilities because they were built for older technology.

Vals, a startup founded in 2024, aims to fix this flawed process. The company raised $40 million in a series A round led by Andreessen Horowitz last month. That funding followed a seed round from 8VC and Bloomberg Beta secured last year.

Rayan Krishnan, the 25-year-old co-founder, interned at Palantir and worked for Microsoft and Stanford’s artificial intelligence lab while studying there. He started the company because he felt evaluation methods were falling behind industry advances.

“We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks were not keeping up with that frontier advance,” Krishnan said. He believes benchmarks need to verify if models can actually do what companies claim.

Last week, Krishnan showed me his two-floor office on San Francisco’s Folsom Street. The old brick building once housed a brewery but now hosts several startups shipping future technology.

“Historically, I think evaluation has been done to evaluate intelligence in a very abstract way. Like, do models know enough information to be able to take a bar exam type test?” Krishnan asked.

Vals differs from other systems that release test data publicly. Public availability allows companies to train their models against those tests, which can be seen as cheating. Vals does not disclose its specific test materials. The company also evaluates models on complex tasks in specific industries like law, finance, and coding.

“What we’re doing is actually looking at what are the real impacts of the models,” Krishnan said. “Can they do work that produces a product of the same quality as a human within every domain?”

The evaluation checks for positive outcomes as well as negative ones. The goal is to understand the implications if these models operated without restriction.

The capabilities being measured are expanding. Beyond traditional sectors, the startup is testing recursive self improvement, mental health, cybersecurity, biosecurity, and the law of armed conflict to see how models apply the Geneva Convention.

Companies pay Vals to test their models. It is an unusual arrangement for a business to pay to find out its product is underperforming. Krishnan compares the revenue model to a student paying the College Board to take the SAT. Effective measurement helps companies troubleshoot and improve over time.

These evaluations are becoming key factors for companies acquiring new AI models. The startup reported revenue eight times higher than last year. Staff numbers have tripled from eight to 25 employees since the beginning of the year.

Krishnan plans to move to a larger office and hire 10 to 15 more people as the company grows. The business also launched a program providing model evaluations to federal agencies.

Krishnan views this benchmarking system as the future for how AI companies establish public trust and grow their businesses.

“AI companies are starting to go public. SpaceX went public. Anthropic is slated for later this year. I suspect OpenAI will be public soon. I think as AI models become a core part of the economy and are diffused more broadly, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings or talk about the prospective investments they’re going to make in AI,” he said.

What it means

For people making things, this shift means the era of simple knowledge tests is ending. Companies will need to prove their tools can actually perform specific, high-stakes work in fields like law or medicine without failing. The cost of building a model will increasingly include the cost of independent verification, similar to how regulated industries already operate. Trust will depend on passing these stricter, private evaluations rather than public leaderboards.

Scroll to Top