Claude Sonnet 5.5: Anthropic says 30% cheaper per task, independent measurement says 50% dearer
Anthropic released Claude Sonnet 5.5 on 28 September 2026, the second model of the 5.5 family after Claude Opus 5.5 on 22 September. The price list stays exactly where it has been since Sonnet 5 launched on 30 June 2026. Anthropic advertises up to 30 percent lower cost per task. Artificial Analysis measured $7.60 per task at maximum effort the same day, roughly 50 percent above Sonnet 5. Both numbers manage to be right, and the difference sits behind a dial most calls never touch. The third model is called Claude Haiku 5.5 and had not shipped as of 29 September 2026.

What is on offer
Sonnet 5.5 (model ID claude-sonnet-5-5) succeeds Sonnet 5 from late June: one million tokens of context, 128,000 output tokens via the API, 300,000 in the batch preview, text and image input, a knowledge cutoff of June 2026. Available across the Claude API, AWS, Google Cloud, Azure and the rest of the list, including arrangements that retain no input text (documentation, as of 29 September 2026). Prices: $2 per million input tokens, $10 per million output, $0.20 cache reads, cache writes at $2.50 for five minutes and $4.00 for one hour, batch API at half price. The shortest cacheable prompt drops from 1,024 to 512 tokens. Two sentences from the documentation deserve to be read before any cost calculation. The tokenizer is the same as Sonnet 5's, so the same text yields the same token counts. And the prices stand literally at the Sonnet 5 level, cache and batch included. The 30 percent is therefore a pure efficiency claim, not a price move. The 30 percent speed plus, per the wording, refers to the output generation rate; no wall-clock methodology is given. The documentation lists "high" as the standard effort level.
The vendor benchmarks and their footnotes
The numbers from the announcement of 28 September 2026, all measured by the vendor itself: Terminal-Bench 4.0 jumps from Sonnet 5's 10.3 percent to 70.6, CursorBench 4.0 from 34.1 to 55.5, HLE with tools from 54.9 to 64.5 percent, OSWorld 2.1 to 80.1 partial points at 57.0 percent of tasks completed. On GDPval-AA v2.1 it scores 1,844 points, Opus 5.5 scores 1,846, a tie at half the price on the Sonnet side. Plus the line everyone is repeating: the first Sonnet model to beat Pokémon Red working only from screenshots. Customer trials quoted on the page: 121,000 tokens per answer instead of 497,000 across 2,441 finance tasks, 2.4x speed at 12 percent fewer tokens, 14 percent fewer output tokens, 20 percent faster ticket handling.
The footnotes belong in any quotation. GDPval and AA-Briefcase ran on a pre-release deployment with a bug that could degrade responses, so those values may sit too low. On the Opus 5.5 page the same GDPval figure appears as 1,735 in a table and 1,846 in the running text, with no effort level labeled. And where comparison columns show GPT-6 Sol, some cells carry GPT-5.6 Sol values instead, because the numbers for the newer line are not public.
The independent measurement runs the other way
The measurement shop Artificial Analysis had Sonnet 5.5 through its suite on release day: intelligence index 56 at maximum effort, Opus 5.5 holds 58, Sonnet 5 holds 38. The output rate of 138.7 tokens per second is the fastest Anthropic value on record there, Sonnet 5 manages about 75. Cost per task: $7.60, roughly 50 percent above Sonnet 5, at an average of 193,000 output tokens per task, the highest consumption the site has ever measured, about seven times that of GPT-6 Astra. On the intelligence-versus-cost-per-task chart, Sonnet 5.5 sits off the efficiency frontier (article of 28 September 2026). Their Terminal-Bench re-measurement of 64 percent instead of the advertised 70.6, they themselves attribute to a pre-release build with structured outputs, a re-run is promised, which is more honesty than many vendor dashboards carry.
The Hacker News thread reached 772 points and 517 comments on release day. The most visible objection lands exactly where the documentation does too: every thinking level above medium approaches the cost of Opus 5.5 and sometimes exceeds it, while the advertised best scores run at maximum (comment of 29 September 2026). A community bakeoff measured four times the cost for the same programming task with a worse result than GPT-6 Sol, a single case without methodological depth, pointing the same direction.
The dial decides which number is true
The two measurements only contradict each other on the cover. Anthropic's savings evidence describes well-scoped everyday tasks on the default levels, and there it shows effect: 121,000 tokens against 497,000 is a real example of less thinking ballast for the same work. Artificial Analysis measured at maximum effort, where the model burns 193,000 output tokens per task on its index tasks, approaches Opus cost and in parts exceeds it. The heise summary of 28 September 2026 puts it without fanfare: on several tests Sonnet 5.5 performs at Opus 5.5 level given maximum compute, then also at similar cost. The German write-up at the-decoder notes the curiosity that Sonnet 5.5 scores worse with Max than with Xhigh on one coding probe. The unit "cost per task" needs two qualifiers from now on: which effort, which task. Whoever quotes vendor charts quotes the dial along with them.
The family and the missing model
| Model | Released | Price per million in/out | Status |
|---|---|---|---|
| Fable 5.1 and Mythos 5.1 | 1 September 2026 | $10 / $50 | Fable open, Mythos only for vetted partners |
| Opus 5.5 | 22 September 2026 | $4 / $20 | available, claimed at Fable 5.1 level |
| Sonnet 5.5 | 28 September 2026 | $2 / $10 | available, priced as Sonnet 5 |
| Haiku 5.5 | announced 28 September 2026 | no figure given | not released |
Everything known about Haiku 5.5 fits in one sentence of 28 September 2026: "Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks." A price, a spec or a date is nowhere stated. the-decoder frames the family against OpenAI: with Fable, Opus and Sonnet, counterparts to GPT-6 Astra, Sol and Luna are standing, and the low end is missing. In the Hacker News debate the same observation lands shorter: Anthropic does not seem to care about competing on the low end. Both are directional claims, not measurements. The OSWorld 2.0 leaderboard carries no entry for any 5.5 model, but the board stands at 17 September 2026, so the gap is expected and not a negative signal.
The system card in brief
The Sonnet 5.5 system card (148 pages, 28 September 2026) reports: no new threshold of the Responsible Scaling Framework crossed; the most prompt-injection-robust Sonnet Anthropic has built so far; clearly above Sonnet 5 on cyber tasks (quote: develops sophisticated exploits much more capably), yet below Opus 5.5; Sonnet 5 serves as the fallback for cyber refusals. Two sentences belong in the record because no marketing chart carries them: harmlessness testing shows regressions against Sonnet 5 in multi-turn tracking and surveillance scenarios, and the model's thinking reads less legibly than many of its predecessors.
What we do with it
Claude Code runs on most of our Linux hosts as a daily remote-maintenance tool; we measured the installed version 2.1.199 on two hosts on 29 September 2026, and the billing of Claude subscriptions hangs on exactly this family whose prices and effort levels this post sorts out. None of the 17 model entries on our inference gateway points to Anthropic (gateway configuration file, as of 29 September 2026); documents and traffic data are produced on our own GPUs and stay there. There is no migratable estate, we have not tested Sonnet 5.5, and this post changes nothing about that. Before any deployment we do this: run 20 real tasks from our mail and document processes across the medium and high levels, note the token counter and wall time per task, redraw the efficiency frontier locally, before a number from a vendor chart meets the house bill. For volume applications the eye stays on Haiku 5.5, the only family member that could push costs below Sonnet level; so far, no date is promised for it.
Further reading
- Anthropic: Introducing Claude Sonnet 5.5, 28 Sep 2026
- Claude documentation: Sonnet 5.5 overview and pricing
- System Card Claude Sonnet 5.5, PDF, 148 pages, 28 Sep 2026
- Anthropic: Introducing Claude Opus 5.5, 22 Sep 2026
- Anthropic: Introducing Claude Sonnet 5, 30 Jun 2026 (baseline)
- Artificial Analysis: Claude Sonnet 5.5, independent measurement of 28 Sep 2026
- Discussion on Hacker News, 28/29 Sep 2026
- heise online on Sonnet 5.5, 28 Sep 2026
- the-decoder: Claude Sonnet 5.5, 28 Sep 2026
- The predecessor reviewed: Claude, 5th generation
Is Sonnet 5.5 cheaper than Sonnet 5?+
The price list is unchanged: $2 per million input tokens, $10 per million output, $0.20 for cache reads, in place since 30 June 2026. Whether a task gets cheaper depends on the effort level. Anthropic cites up to 30 percent less for well-scoped tasks; Artificial Analysis measured around 50 percent more at maximum effort.
What is actually confirmed about Haiku 5.5?+
Only the announcement sentence of 28 September 2026: built for high-volume and cost-sensitive applications, joining in the coming weeks. A price, a specification or a date is absent. While the price stays open, the model is useless for any cost plan.
Does Opus 5.5 replace the top model class?+
Anthropic claims Opus 5.5 performs at the level of Fable 5.1 on most work at 40 percent lower cost than Opus 5. No independent measurement of that claim existed as of 29 September 2026. Fable 5.1 remains the class with the largest headroom for multi-day agentic runs.
senn-tech