Arena · XOriginal · English

We evaluated Jev Router by @typesafeai on Agent Arena.

Image source · Arena · X

We evaluated Jev Router by @typesafeai on Agent Arena.

Our tests spanning more than 4,700 real-world agentic sessions show:

1. Jev Router does not improve on the current Pareto frontier. For similar performance as DeepSeek V4.1 Flash (Max), the solution costs 38% more and its median model request latency is 1.7x higher.

2. However, it does mostly route to Pareto efficient models. The most LLM it picks is DeepSeek V4.1 Flash, with GPT-6.1 Sol and GPT-6 Luna also being frequent choices.

3. Jev Router’s key strength is steerability. Its score of +10% almost matches Claude Opus 5.5 (High) (+10.48). This highlights the benefit of Jev effectively routing to stronger LLM in response to user feedback.

More insights in the thread below.

Original source

Arena · X

Content notes

Original publication and rights belong to the source.