Anthropic released Claude Opus 5, its new flagship model, claiming it matches the performance of competing systems while cutting token costs in half. The model scores at the top tier for coding and knowledge work tasks, undercutting rivals on price.
On ARC-AGI-3, a benchmark measuring performance on novel problem-solving tasks, Opus 5 achieved 30.2 percent accuracy, nearly four times higher than GPT-5.6 Sol. This benchmark tests a model's ability to solve unfamiliar problems without domain-specific training, a key metric for general intelligence.
The cost advantage matters more than the raw performance numbers. Anthropic positioned Opus 5 to deliver near-parity with Fable 5, another high-performing model, while reducing the per-token expense by 50 percent. For enterprises running AI inference at scale, token costs drive operational budgets. Halving expenses while maintaining top-tier capability creates direct financial pressure on competitors.
Anthropic's pricing strategy reflects the broader shift in AI economics. Raw capability benchmarks have plateaued among leading models. Companies now compete on efficiency, cost, and reliability. Opus 5 targets organizations already committed to Claude but looking to reduce spending, and price-sensitive customers currently using alternatives.
The ARC-AGI-3 results suggest Anthropic improved reasoning and problem-solving. This aligns with Anthropic's published research on scaling laws and constitutional AI, which focuses on making models more reliable for complex tasks rather than chasing raw scale.
The release intensifies competition in the frontier model space. OpenAI, Google, and other developers track similar benchmarks and pricing. An established player undercutting on cost while matching performance typically triggers market share shifts among enterprise customers, where switching costs are real but not insurmountable when savings are substantial.
Anthrop