Powered by Smartsupp

Vals Launches to Revolutionize AI Benchmarking With Next-Gen Validation for Modern Models



By admin | Sep 21, 2026 | 4 min read


Vals Launches to Revolutionize AI Benchmarking With Next-Gen Validation for Modern Models

Benchmarking has turned into the standard method AI companies use to demonstrate what their models can do—and when the numbers look good, it becomes a powerful marketing tool. In practice, strong benchmark results almost always translate into favorable publicity. The problem is that many companies have learned how to game older benchmarking systems, which were designed for a previous generation of AI and aren't equipped to assess today's models. Vals, a startup founded in 2024, aims to overhaul this flawed setup. In under two years, the company has become a recognizable name in tech, securing a seed round led by 8VC and Bloomberg Beta last year. Then, following a stretch of rapid expansion, it closed a $40 million Series A last month, led by Andreessen Horowitz.

Rayan Krishnan, the company's 25-year-old co-founder, previously interned at Palantir and, while an undergraduate at Stanford, worked for Microsoft and the university's highly regarded AI lab. Krishnan says Vals grew out of his own observations about how benchmarking was lagging behind the very industry it was meant to measure. "We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance," Krishnan explains. As AI becomes embedded in every facet of society, benchmarks ought to exist to confirm that models can actually do what their makers claim, Krishnan said.

Last week, the young founder gave me a tour of his company's two-story office on Folsom Street in San Francisco—an old brick building that housed a large brewery a century ago. Rather than churning out beer, the historic structure now shelters a variety of startups aiming to build the future of tech. "Historically, I think evaluation has been done to evaluate intelligence in a very abstract way," Krishnan tells me. "Like, do models know enough information to be able to take a bar exam type test."

This is where Vals tries to set itself apart. Many benchmarking systems rely on tests that are publicly available—which means a company could train its model against those tests, essentially cheating on the exam. Vals, by contrast, does not disclose its specific test materials. And rather than gauging an AI model's general knowledge, Vals assesses how well models handle complex tasks tied to particular fields such as law, finance, and coding. "What we're doing is actually looking at what are the real impacts of the models," said Krishnan. "Can they do work that produces a product of the same quality as a human within every domain."

The goal, he says, is to check for negative outcomes as well as positive ones. The idea is to examine how, "if these models ran wild in the world, what the negative implications would be."

The capabilities Vals measures keep expanding. Beyond more conventional industries, the startup continues to move into more unusual territory. "We have a benchmark on recursive self improvement. We're doing some work in mental health, cybersecurity, biosecurity, and even law of armed conflict to models to understand how to apply the Geneva Convention," Krishnan shares.

Companies pay Vals to test their models, which can seem like a strange arrangement at first glance. Why would a company pay to find out its model isn't performing well? But having an effective measurement helps companies diagnose problems and improve over time. Krishnan likens their revenue model to how a student might pay the College Board to take the SAT. These evaluations, in turn, are becoming crucial decision-making factors for companies looking to acquire new AI models.

The startup recently disclosed that its revenue is currently eight times what it was last year. Its headcount is growing too. Vals, which began the year with just eight people, has already tripled to a team of 25. Krishnan said that as the startup expands, the plan is to move to a considerably larger office and to add another 10 to 15 people. The company also recently launched a program focused on providing model evaluations to federal agencies.

Krishnan sees his company's benchmarking system as the future of how AI companies will think about growing their businesses and building public trust. "AI companies are starting to go public. SpaceX went public. Anthropic is slated for later this year. I suspect OpenAI will be public soon. I think as AI models become a core part of the economy and are diffused more broadly, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings or talk about the prospective investments they're going to make in AI," he said.




RELATED AI TOOLS CATEGORIES AND TAGS

Comments

Please log in to leave a comment.

No comments yet. Be the first to comment!