AI · Opinion

Your startup is paying double for two Elo points of Claude

Sonnet 5.5 lands within two Elo points of Opus 5.5 at half the price. For most production traffic, the flagship now buys benchmark points you will never ship.

Two Elo points separate Claude Sonnet 5.5 from Claude Opus 5.5 on GDPval-AA v2.1, Artificial Analysis’s test of real work across 44 occupations: 1,844 against 1,846. Sonnet 5 scored 1,449. Opus 5.5 costs $4 per million input tokens and $20 per million output tokens. Sonnet 5.5 costs $2 and $10. If your startup pointed every agent call at the flagship when Opus 5.5 shipped on 22 September, you spent the days until 28 September paying double for a model that, on this measure, cannot be told apart from the one Anthropic released on the 28th.

I expected the usual launch choreography, where the mid-tier model gets close to the flagship on a few flattering charts and trails it on everything hard. What surprised me is that Anthropic’s own table puts Sonnet 5.5 ahead on Terminal-Bench 4.0, with 70.6% at Max effort against 66.4% for Opus 5.5 at Xhigh. Kingy AI points out that Anthropic reports standard errors of about 2.5 points on this 66-task benchmark, so the lead sits inside the noise. I am happy to call it a tie, and at half the list price I would take the tie.

Token counts matter more than list price, because you pay per token and a model that reaches the answer in fewer of them is cheaper than its price sheet implies. Balyasny Asset Management ran 2,441 finance tasks and measured about 121K tokens per answer on Sonnet 5.5, where Sonnet 5 used 497K, roughly a quarter of the spend for a better score. Base44 averaged 3.6 iterations per app build across 118 builds, against 7.7 for Opus 5, with the apps scoring level. GitHub made the model generally available in Copilot on launch day. Anthropic claims that on several benchmarks, running the new model at its Low or Medium setting tops the old model’s best result at roughly a tenth of the per-task cost.

This has now happened in consecutive generations. In June, Sonnet 5 beat Opus 4.8 on Terminal-Bench 2.1, 80.4 to 74.6, and edged it on GDPval-AA v2, 1,618 to 1,615, while Opus 4.8 cost 67% more per input token. Two releases running, a Sonnet has tied or beaten an Opus on Anthropic’s headline agentic and knowledge-work tests, and OpenAI prices its mid-tier GPT-6 Sol at exactly half of Opus 5.5 for uncached tokens, so the pressure is coming from both sides.

The case for Opus is real, and Anthropic makes it itself. On benchmarks built from actual coding work, Opus 5.5 still leads: 57.8% to 55.5% on CursorBench 4.0, and 54.4% to 52.1% on FrontierCode 1.1 with Sonnet at Xhigh. Anthropic writes that Opus 5.5 “remains clearly stronger at complex, open-ended work requiring sustained judgment.” It also concedes that at higher effort Sonnet 5.5 costs about the same per task as Opus, so the saving lives at low and medium effort. None of the benchmark figures have been independently reproduced, and Artificial Analysis ran GDPval-AA on a pre-release build with a structured-outputs bug.

I accept every one of those points, and I think they argue for routing rather than for a flagship default. Both coding gaps are small, and Anthropic itself describes the CursorBench one as within about two points. Most production traffic at the startups I cover is well-scoped work: ticket triage, extraction, bug fixes, slide decks, code review. Zendesk’s 20% faster ticket processing came from exactly that kind of task. The saving living at low and medium effort suits me fine, because routine traffic belongs at those settings, and Claude Code already defaults to Medium. Genuinely open-ended architecture work deserves Opus, and one tester quoted in the launch post described the split neatly: let Opus set a game’s architecture and hand implementation to Sonnet. Flagship spend is also moving the wrong way at the top end. Emergent says independent testing found Opus 5.5 at max effort costs 21% more than Opus 5 on long coding-agent runs, which undercuts Anthropic’s 40% cheaper headline, a figure that compares Opus 5.5 at medium with Opus 5 at high.

I cannot tell you what share of startup API spend goes to Opus today. Nobody has published market-wide numbers by model tier, and I looked. You can find your own share in an afternoon, though. Export a week of Opus 5.5 calls from your logs, replay them on claude-sonnet-5-5 at Medium effort, and have someone who does not know which model wrote which output grade the pairs blind. Move every task type where Sonnet wins or ties, keep Opus on the rest, and each task type you move will cost half as much per token from the day you switch it.

Prompted by Introducing Claude Sonnet 5.5, anthropic.com.