IT之家 AIUpdated

In just 10 months, nearly doubling, AI model evaluation platform Arena's valuation rises to 3.1 billion USD

IT之家, October 9 news: On the 8th local time, AI model evaluation platform Arena announced it had completed a US$200 million Series B funding round (IT之家 note: approximately RMB 1.343 billion at current exchange rates)…

Image source · IT之家 AI

IT之家 October 9 news: On the 8th local time, AI model evaluation platform Arena announced the completion of a 200 million USD (IT之家 note: approximately 1.343 billion RMB at current exchange rates) Series B funding round,reaching a valuation of 3.1 billion USD(approximately 20.817 billion RMB at current exchange rates). Arena was originally a 2023 research project at the University of California, Berkeley, ranking AI models through public voting.

Arena previously disclosed that its annualized revenue reached 100 million USD (approximately 672 million RMB at current exchange rates) in June this year.

This funding round was co-led by Lightspeed Venture Partners and Khosla Ventures, with participation from Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis, and other institutions.

In January this year, Arena completed a 150 million USD (approximately 1.007 billion RMB at current exchange rates) Series A funding round,at a post-money valuation of 1.7 billion USD(approximately 11.416 billion RMB at current exchange rates), with annualized revenue at that time of 30 million USD (approximately 201 million RMB at current exchange rates). In just about 10 months, Arena's valuation has nearly doubled.

Arena's evaluation platformis free for individual users. Users can enter prompts, or have AI write programs and develop projects according to their requirements, then compare the results of different models and score them. Arena says the platform attracts tens of millions of visitors each month.

In September (the 9th month) last year, Arena launched AI Evaluations, a commercial service for AI model research institutions and enterprises, which uses feedback data from community users to provide detailed model performance analysis.

This year, AI research institutions found that their models could achieve high scores bygaming the rules of benchmark tests, yet may not possess the corresponding real-world capabilities. Enterprises are also no longer satisfied with standardized test scores and want to know which model better suits their business needs.

In its funding announcement, Arena pointed out: "The pace of AI development has already outstripped our ability to evaluate it. Once a model detects it is being tested, fixed benchmarks become ineffective. The world needs a neutral third party to test whether AI is safe in actual use and whether its behavior matches human expectations. Arena has begun to take on this role."

To this end, Arena hasadded alignment capability evaluations to its leaderboard, focusing on whether models perform actions users did not request without authorization, whether they misattribute statements or facts to the wrong sources, and whether they falsely claim to have completed tasks. Arena calls the last behavior "deceptive completion."

On the initially published alignment capability leaderboard, several OpenAI models ranked at the top, Claude Opus 5.5 ranked sixth, and Claude Fable ranked ninth.

Original source

IT之家 AI

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original