The DecoderOriginal · English

AI agent teams waste massive tokens for barely measurable quality gains, research finds

Teams of AI agents barely outperform solo agents but cost up to 5.1x more, according to Vals AI. Only one out of four tests with GPT-6 Sol and Claude Opus 5.5 showed a measurable gain. Anthropic's own data backs this up…

Matthias Bastian
Image source · The Decoder

AI agent teams waste massive tokens for barely measurable quality gains, research finds

Matthias Bastian Matthias Bastian View the LinkedIn Profile of Matthias Bastian Oct 11, 2026 Image description

AI agent teams deliver almost no better results than single agents, research finds.

Evals company Vals AI tested GPT-6 Sol and Claude Opus 5.5 on the "Vibe Code Bench," both solo and as teams, at two reasoning levels: medium and maximum reasoning effort. The teams cost between 1.8x and 5.1x more than single agents.

Out of four comparisons between teams and solo agents, only one showed a statistically significant improvement: GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher. At maximum reasoning, the team setup gave neither Sol nor Opus 5.5 any real advantage. The results suggest that the extra cost of agent teams isn't worth it in most cases, especially when models are already running at full compute.

Vals AI's benchmark shows cost per app versus score on Vibe Code Bench. Arrows point from each model's single agent to its team at the same reasoning effort. Teams cost far more but barely improve scores. | Image: Vals AI[
Anthropic saw quality gains shrink as it added more agents in two of its own tests with Opus 5.5. Larger teams reached a given performance level faster, but going from ten to 100 agents only nudged scores up slightly after 24 hours. In separate ProgramBench tests, speed gains came with higher token usage.

Task with Opus 5.5 1 agent 10 agents 30 agents 100 agents
Knowledge base 0.53 0.70 0.71 0.74
Lean Theorem Proving 0.39 0.66 0.66 0.68

Fable 5.1 showed stronger quality gains on the Lean theorem proving task above ten agents, but still scored below Opus 5.5 across all tests. On the knowledge base task, Fable's score actually dipped slightly when scaling from 30 to 100 agents.

More agents buy speed but not better output

OpenAI researcher Noam Brown confirmed in the Dwarkesh Podcast that multi-agent systems mainly buy speed, not better quality. Four agents solved tasks twice as fast but also cost twice as much. At 16 agents, the pattern held but grew slightly less efficient.

The effect depends heavily on the task. Web research and math parallelize well, but writing a novel doesn't, he said. Throwing 10,000 agents at a novel would be just as pointless as throwing 10,000 people at it. Brown acknowledged that scaling to very large numbers of agents remains largely unexplored because the costs are simply too high.

OpenAI developer Eric Provencher recently warned against using agent swarms for exactly this reason. They're most likely wasted money, he argued, because coordination between agents breaks down. He called it the coordination tax.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Read on for the full picture.
Subscribe for hype-free coverage.

  • Full access to every article on THE DECODER
  • No ads
  • Join the comments and community discussions
  • A weekly AI news recap via mail
  • 6x/year: "AI Radar" — deep dives on the AI topics that matter most
  • Daily AI news, always up to date
  • Our full ten-year archive
  • Covered by a team with 10+ years in AI
Subscribe to The Decoder
Original source

The Decoder

Content notes

Original publication and rights belong to the source.