According to monitoring by Dongcha Beating, AI model evaluation company Vals AI has completed a $40 million Series A funding round, achieving a valuation of $400 million, led by a16z. Existing investors such as 8VC, Pear VC, and Bloomberg Beta also participated. Vals specializes in third-party evaluations of large models. Currently, many models report their performance metrics based on self-assessments by the manufacturers. Vals, however, conducts unified testing of different models using real tasks from programming, finance, law, and healthcare. The company claims its results have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI, with this year's revenue already reaching eight times that of the entire year of 2025. a16z is particularly interested in the value of this 'third-party evaluation.' Public benchmarks are increasingly susceptible to score manipulation, inclusion of training data, and even targeted optimizations by manufacturers, meaning high scores on leaderboards do not necessarily equate to real-world usability. Additionally, Vals has fully opened Vals Smith, allowing users to create coding benchmarks directly from any GitHub repository. The system extracts real development tasks from historical pull requests and uses hidden tests to check whether models can complete them. This enables companies to directly test which model is best suited for their codebase, rather than relying solely on public leaderboards.
All Comments