Andreessen Horowitz-Backed Startup Vals Eyes AI Benchmarking Crown
Vals, a startup founded in 2024, is revolutionizing AI benchmarking by offering secure, rigorous, and industry-specific evaluations for modern models. Co-founder Rayan Krishnan aims to fix outdated systems, ensuring AI accountability and building public trust as the company experiences rapid growth and expands its influence.
In the rapidly evolving landscape of artificial intelligence, benchmarking has long served as the standard for validating model capabilities, allowing companies to showcase their prowess and gain a competitive edge. However, this system, often reliant on older, easily gamed methodologies, has struggled to keep pace with the swift advancements of modern AI. Vals, a startup established in 2024, has emerged with a mission to overhaul this imperfect system, positioning itself as a critical player in the tech industry.
Vals quickly garnered attention, securing a seed round led by prominent investors 8VC and Bloomberg Beta, followed by a significant $40 million Series A funding round led by Andreessen Horowitz. At the helm is 25-year-old co-founder Rayan Krishnan, whose past experiences at Palantir and contributions to Microsoft and Stanford’s renowned AI lab as an undergraduate fueled his insights into the inadequacies of existing benchmarking practices. Krishnan observed a clear disconnect: new, highly capable models were rapidly entering the market, yet the academic benchmarks designed to measure their progress were lagging behind the frontier of AI innovation.
Vals differentiates itself from traditional benchmarking systems by adopting a more rigorous and secure approach. Unlike many systems that offer publicly available tests, which can inadvertently enable companies to 'train to the test' and artificially inflate their scores, Vals maintains strict confidentiality over its specific test materials. Furthermore, instead of merely assessing general AI knowledge, Vals delves into evaluating models' capacity to perform intricate, industry-specific tasks across diverse sectors such as law, finance, and coding. Krishnan emphasizes that their objective is to gauge the 'real impacts' of these models, determining if they can produce work of human-level quality within specialized domains.
This advanced evaluation process extends beyond just positive outcomes; Vals also meticulously analyzes potential negative implications should these models operate unsupervised. The startup's capabilities are continuously expanding, venturing into unique and critical areas. This includes benchmarks for recursive self-improvement, and specialized work in mental health, cybersecurity, biosecurity, and even the law of armed conflict, assessing how models might apply principles like the Geneva Convention.
The business model of Vals involves companies paying for its rigorous model evaluations, a concept Krishnan likens to students paying the College Board for the SAT. These evaluations are proving invaluable for companies, aiding in troubleshooting and continuous improvement of their AI models. Consequently, Vals' assessments are becoming pivotal decision-making factors for organizations seeking to acquire new AI technologies. The company’s robust growth is evident in its financial performance, with revenue currently eight times greater than the previous year. Its team has also rapidly expanded, tripling from eight to 25 employees within the current year, with plans for further expansion and a relocation to a larger office space.
Looking ahead, Vals recently initiated a program to provide model evaluations to federal agencies, further cementing its role in the broader AI ecosystem. Krishnan envisions Vals’ benchmarking system as instrumental in shaping how AI companies approach growth and build public trust, especially as more AI entities like Anthropic and potentially OpenAI prepare to go public. He believes that as AI models become an integral part of the global economy, the sophisticated benchmarks and evaluations offered by Vals will be central to driving their adoption, influencing public filings, and guiding prospective investments in artificial intelligence.