Back to News
InnovationAI Understanding briefing

Study finds multi‑agent AI teams consume far more tokens for modest or no quality gains

Vals.ai benchmarked OpenAI’s GPT‑6 Sol and Anthropic’s Claude Opus 5.5 as single agents versus coordinated teams on a full‑stack web‑app task, revealing steep cost increases and only limited performance improvements.

5 min readRead the linked source
Source-provided image accompanying Study finds multi‑agent AI teams consume far more tokens for modest or no quality gains
Source referenceSource recorded
Publisher
vals.ai
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Key terms

API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Feature
An input variable used by a model to make predictions.
Test yourselfAI Agents Quiz
Source video from vals.ai · shown with attribution.

What happened

Vals.ai ran its Vibe Code Bench, a that builds complete web applications from product specifications, using two leading multi‑agent‑capable models—OpenAI’s GPT‑6 Sol and Anthropic’s Claude Opus 5.5. Each model was evaluated in four configurations: single‑agent at medium effort, single‑agent at max effort, team‑based at medium effort, and team‑based at max effort. Teams consisted of a lead agent that delegated sub‑tasks to up to five subagents. Across 50 apps, the study measured UI‑test pass rates (score), runtime, and token‑based cost. Key findings include: - Only GPT‑6 Sol’s medium‑effort team showed a statistically significant score increase of 7.3 points (p = 0.005). All other configurations showed non‑significant differences. - Team runs cost 1.8‑5.1× more than single‑agent runs. For Claude Opus, max‑effort teams consumed a median of 224 million cached tokens per app versus 17 million for the lead and 55 million for a single agent. - Runtime savings were modest for Sol (team finished slightly faster) but substantially longer for Opus (up to 2.3× the single‑agent time). - Higher reasoning effort (max) improved Sol’s single‑agent score by 11.4 points but did not yield a comparable gain for Opus. - The study notes that benefits depended on how well the work could be split into independent parts; Sol’s medium‑effort teams helped most on harder apps, while Opus’s results were mixed. The authors caution that results stem from a single benchmark and may not generalize to other domains.

Vals.ai used its Vibe Code Bench to evaluate 50 full‑stack web‑app projects. Each app required database, API, UI, and deployment components, with correctness measured by automated UI tests. The runs were performed in a sandbox that logged token usage, runtime, and cost. Both OpenAI’s GPT‑6 Sol and Anthropic’s Claude Opus 5.5 were executed in two modes: a single agent handling the entire build, and a team mode where a lead agent split the work among up to five subagents. The lead received a brief instruction to divide the task, delegate, and later integrate and verify the results. All other prompts, sandbox environments, and grading criteria were identical across configurations. Statistical analysis (paired t‑tests) showed only one significant improvement: Sol’s medium‑effort team outperformed its single‑agent counterpart by 7.3 points. The other three comparisons—Sol max‑effort, Opus medium‑effort, and Opus max‑effort—did not achieve statistical significance. Cost analysis revealed that team configurations consumed far more tokens, primarily due to each subagent maintaining its own context that was repeatedly sent on every model call. For Opus max‑effort teams, cached token usage reached a median of 224 million per app, inflating the estimated cost to about $189 for a single run. Runtime effects varied: Sol’s teams were slightly faster than single agents, while Opus’s teams were markedly slower, often running in sequential waves that waited for the slowest subagent. The extra time did not translate into higher scores, indicating inefficiencies in coordination. The authors note that the study’s scope is limited to this benchmark and that results may differ for other types of tasks.

The study also examined how the lead agents divided work. Sol’s medium‑effort lead split tasks along architectural boundaries (database, API, UI, deployment) and created subagents quickly. At max effort, it added more subagents and split the UI into finer areas. Opus’s lead, in contrast, authored a shared contract document outlining schema and ownership before spawning subagents in waves—first builders, then feature agents, then testing agents. Opus’s max‑effort teams employed more subagents (median 6.8 per app) and issued many more tool calls, yet the additional testing did not yield proportionally higher scores. Overall, the research suggests that multi‑agent orchestration can be beneficial when tasks are naturally parallelizable and when the coordination overhead is low. Otherwise, the token and time costs may outweigh any marginal quality gains.

Source details: vals.ai ↗

Why it matters

The findings challenge the assumption that adding subagents automatically yields better outcomes. For developers and enterprises considering multi‑agent APIs, the study highlights a trade‑off: substantial token‑cost inflation and longer runtimes for only marginal, sometimes statistically insignificant, quality improvements. This has direct implications for budgeting cloud‑based AI usage, especially given per‑token pricing models. Moreover, the research underscores the importance of task decomposition—only when work can be cleanly parallelized do teams provide measurable gains. As OpenAI and Anthropic promote native multi‑agent support, stakeholders need realistic expectations about performance versus cost, and may prioritize smarter delegation strategies or improved coordination mechanisms before widespread adoption.

Cost Efficiency: Token‑based pricing means that the 2‑5× cost increase for team runs can quickly become prohibitive for large‑scale deployments, especially for enterprises budgeting AI usage. Performance vs. Parallelism: The modest score improvements—significant only for one configuration—indicate that naïve parallelism does not guarantee better outcomes. Effective delegation requires careful task analysis. Product Roadmaps: Both OpenAI and Anthropic are marketing multi‑agent capabilities as a core . This study provides early, data‑driven feedback that could influence future API designs, such as shared context caches or smarter subagent scheduling. Strategic Decision‑Making: Companies evaluating whether to adopt multi‑agent APIs now have empirical evidence to weigh against their specific workloads, potentially opting for single‑agent approaches when tasks are tightly coupled. Research Direction: The findings highlight a gap in current multi‑agent research—efficient coordination and context sharing—guiding future academic and industry investigations.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interactive Concept Check+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

What to watch next

Future work will likely explore: 1. More efficient coordination protocols that reduce cached‑token overhead. 2. Benchmarks across diverse domains (e.g., data analysis, research synthesis) to test generality. 3. Pricing model adjustments by providers to reflect multi‑agent token usage. 4. Tooling that surfaces when a task is amenable to parallel sub‑agents versus when a single agent suffices. 5. Potential model or API updates from OpenAI and Anthropic aimed at lowering the cost of multi‑agent orchestration.

API Enhancements: Look for announcements from OpenAI or Anthropic about reduced token duplication or shared context mechanisms for subagents. Alternative Benchmarks: New studies applying multi‑agent setups to domains like data cleaning, scientific literature review, or business workflow automation could validate or contradict these results. Pricing Adjustments: Providers may introduce tiered pricing for multi‑agent usage to reflect the higher token consumption. Tooling Improvements: Development of orchestration frameworks that automatically identify independent sub‑tasks could lower coordination overhead and improve cost‑benefit ratios. Policy Implications: As multi‑agent systems become more prevalent, regulators may consider guidelines for transparent reporting of token usage and cost in AI services.

Related guides & quizzes

Found this useful?