THE CRUNCH

Multi-agent AI setups, where several models or instances work on a task as a team, cost between 1.8x and 5.1x more than a single agent while delivering almost no measurable quality improvement, according to testing by the evals company Vals AI. The firm ran GPT-6 Sol and Claude Opus 5.5 solo and in teams on its Vibe Code Bench at two reasoning levels, medium and maximum effort.

Of four team-versus-solo comparisons, only one showed a statistically significant gain: GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher. At maximum reasoning, the team setup gave neither model any real advantage, suggesting the extra spend buys little once models are already running at full compute.

Other findings point the same way. Anthropic's own tests with Opus 5.5 saw quality gains shrink as more agents were added: larger teams reached a given performance level faster, but going from ten to 100 agents only nudged scores up slightly after 24 hours. In separate ProgramBench tests, the speed gains came with higher token usage, and Fable 5.1's score on a knowledge base task actually dipped slightly when scaling from 30 to 100 agents.

OpenAI researcher Noam Brown put the trade-off plainly on the Dwarkesh Podcast: multi-agent systems mainly buy speed, not better quality. Four agents solved tasks twice as fast but cost twice as much, and the pattern held, slightly less efficiently, at 16 agents. He noted the effect depends heavily on the task, with web research and math parallelising well while writing a novel does not. OpenAI developer Eric Provencher has separately warned that agent swarms are most likely wasted money because coordination between agents breaks down, a problem he calls the coordination tax.