Mistral Large 4 by @MistralAI is now in the Arena!
Head to Agent Arena to test it out, and your votes will shape its evaluation. Scores coming soon.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
Mistral Large 4 is also available in Code Arena: WebDev, Text, and Vision.

