📰 Key Takeaways

Vals is a startup founded in 2024 that’s trying to fix the AI industry’s outdated model benchmarking system. In under two years, the company has grown fast: last year 8VC and Bloomberg Beta led its seed round, and last month it closed a $40M Series A led by Andreessen Horowitz.

Co-founder Rayan Krishnan, now 25, interned at Palantir and worked at Microsoft and the Stanford AI Lab while studying at Stanford. He says the idea for Vals came from watching increasingly capable new models get released at a breakneck pace, while academic benchmarks couldn’t keep up with the frontier. In his view, the whole point of a benchmark is to verify whether a model can actually do what vendors claim it can.

The biggest difference between Vals and traditional benchmarks: many existing test questions are public, so model makers can effectively “cheat” by training directly on them. Vals keeps its test content private. It also moves away from abstract “intelligence” tests (like having a model take the bar exam or other knowledge-based tests) and instead evaluates how well models complete complex, real-world tasks in specific industries — law, finance, coding — checking whether the output quality actually matches what a human worker would produce. Krishnan also notes that evaluation isn’t just about measuring upside; it also looks at what could go wrong if a model operates “out of control.” The original article didn’t go into detail on exactly which capabilities Vals tests — check the source link for more.


💬 JudyAI Lab’s Take

Something worth flagging today: Vals, less than two years old, has already raised a $40M Series A led by Andreessen Horowitz across two rounds. That’s a strong signal that the market has real appetite for rethinking how AI gets evaluated.

This points to a broader shift: when model capability moves faster than academic benchmarks can update, “leaderboard scores” start to lose meaning — worse, they become numbers vendors can game by training against the test. Two things Vals is doing are worth studying if you’re building with AI: keeping test content private to prevent targeted training, and shifting the focus from abstract knowledge tests to real completion quality on concrete industry tasks in law, finance, and coding. The angle on evaluating “what happens when a model goes off the rails” is also a good reminder — it’s not just about whether a model can do something, but how costly it is when it gets it wrong. This shift toward judging models by real task outcomes instead of standardized test questions may well be where evaluation methodology is headed.

Next time you’re picking a model, it’s worth asking: was that score earned on a public test bank, or on a real task where the model never saw the answer key?


📅 Source Info


🔗 Further Reading