Announcing Our Investment in Vals

Posts
News
Aug 13, 2026
Share

Posts
News
Aug 13, 2026
Share

Intelligence is becoming abundant, and deciding what intelligence to use is not getting easier. Two years ago, choosing a model meant picking between a handful of chat endpoints. Today the option space spans dozens of frontier and open models, reasoning modes, agent harnesses, and domain applications, each updated monthly, each arriving with its own performance claims. Every enterprise now has to construct its own ontology for intelligence: which model, on which task, in which harness, at what cost. That mapping is the core axis of decision making for this wave of technology, and it can only be built on measurement.

The measurement most of the industry runs on was never designed for these decisions. Academic benchmarks test contrived tasks far from real work. Open test sets leak into training corpora, directly or through synthetic data, and results quietly inflate. Public leaderboards get optimized against. And when a lab reports its own numbers, even in good faith, it is reporting on an exam it wrote and studied for. None of this is scandal; it is the natural state of a field that built its capability faster than its instruments. But it leaves the most important question in technology (what can these systems actually do?) without a trusted answer.

Every foundational technology eventually gets its instruments, and the instrument maker is independent by construction. Credit got ratings agencies. Electrical equipment got UL. We believe AI adoption at scale requires the same: an unbiased, research-grade third party that communicates clearly, to enterprises and to the world, what the intelligence in your pocket is capable of.

That is Vals. 

The Anatomy of Vals

Vals builds benchmarks the way the work is actually done. Its evaluations are authored alongside practitioners (lawyers, accountants, finance professionals) on real professional tasks: legal research through a consortium of major law firms, expert-written finance agent workflows, tax, healthcare, and coding. Test sets are held private so they cannot leak into training data, with public validation sets published for transparency. The output is rare in this market: numbers that mean something, produced by a party with no model to sell.

The public leaderboards and the Vals Index are what is visible. Underneath is evaluation infrastructure that both sides of the market run on. Labs use Vals to understand where their models stand on real work. Enterprises use Vals to decide what to deploy and to keep measuring it after they do. As models, harnesses, and applications multiply, every serious adoption decision routes through the same question, and Vals is building the trusted answer.

Team 

Rayan and Langston started Vals out of Stanford with a view that was contrarian at the time: that as capability compounded, independent measurement would become more valuable, and that nobody was building it with real research rigor. They pair that rigor with an instinct for working alongside practitioners, recruiting the experts whose work the benchmarks are meant to reflect. Their results are now cited across major publications and relied upon in regulated industries where the cost of a wrong answer is highest, and are a core part of model cards. 

Vals x 8VC

We led Vals' seed round. Today, Vals announced a  Series A led by our friends at a16z. We are proud to keep partnering with Rayan, Langston, and the team as they build the measurement layer for the most consequential technology of our time.

To learn more, visit vals.ai