Eric Provencher, developer at OpenAI Codex, has issued a direct warning about the scalability economics of multi-agent AI systems. Running more than two parallel sub-agents almost always wastes tokens without improving output quality, according to his analysis. The culprit: agents don't inherently trust each other and spend computational resources re-verifying work across the swarm.

Provencher calls this overhead the "coordination tax." The problem surfaces when agents operate in parallel without shared context or trust mechanisms. Each agent re-validates decisions made by others, creating redundant processing loops that multiply costs exponentially while output quality plateaus or degrades.

His evidence is stark. A single project deployed 1,393 agents to refactor Python code. The final bill: $20,000 in token consumption. The same task completed by a single Astra agent cost a fraction of that amount with comparable or better results. The 1,393-agent swarm burned resources without delivering commensurate gains in code quality or execution speed.

This finding challenges the current enthusiasm around multi-agent architectures. The AI development community has embraced agent swarms as a solution to complex problems, betting that distributed reasoning and parallel task execution would beat single-agent approaches. Provencher's work suggests that assumption breaks down at scale without proper architectural safeguards.

The coordination tax problem manifests differently than most engineers expect. It isn't about communication overhead in the traditional sense. Rather, it reflects a fundamental trust deficit in distributed systems. When one agent completes a subtask, downstream agents cannot assume correctness. They re-run validation, re-check logic, and re-execute portions of work already completed. This redundancy multiplies with each additional agent in the swarm.

Two agents appear to be a practical threshold. With a pair, basic division of labor becomes feasible and one-level-deep verification remains affordable. Beyond that, the trust costs outweigh benefits. A swarm of ten agents doesn't just cost five times more than two. It costs substantially more because each agent verifies multiple others' work simultaneously.

The implications reshape how AI development teams should architect production systems. Swarm approaches marketed as solutions to hard problems may actually represent engineering waste at current token prices. Teams building with Claude or GPT-4 should reconsider the default assumption that "more agents equals better results."

This doesn't eliminate multi-agent value entirely. Hierarchical designs with clear authority structures, or tightly scoped agent teams with built-in trust mechanisms (like shared memory or verified hand-offs), might avoid coordination tax. But loosely coupled swarms executing in parallel without explicit trust protocols appear fundamentally uneconomical.

The lesson applies beyond OpenAI's products. Any organization running multi-agent systems on commercial API pricing needs to audit actual token consumption versus output quality. If Provencher's findings hold across different models and domains, the AI infrastructure industry may need to rebuild agent coordination from first principles before swarm-based systems become practical at scale.