Warning: Trying to access array offset on false in /www/wwwroot/speedinet.co.za/wp-content/themes/Divi/includes/builder/functions.php on line 2182
Why Claude Sonnet 5.5 Is Cheaper Without a Price Cut, and When It Isn’t | Speedinet Wireless and Fiber Internet Service Provider

Need Reliable Internet?

Business Fibre • Wireless • Hosting • VoIP

Get Quote

When Anthropic released Claude Sonnet 5.5 on 28 September, it left the price card exactly where Sonnet 5 had it, at $2 per million input tokens, $10 per million output tokens and $0.20 per million cached reads, yet it still told customers the new model would cost up to 30% less per task for most work. Both statements can be true at once only because the bill for an AI model is the number of tokens it spends multiplied by the price of each token, and Anthropic changed the first number while leaving the second one alone.

Need Reliable Internet?

Business Fibre • Wireless • Hosting • VoIP

Get Quote

That distinction explains where the saving comes from, and it also explains the part of the launch that got less attention, which is that the saving belongs to the lower and middle settings of the model rather than to the model as a whole. Once Sonnet 5.5 is pushed to its highest effort level, independent testing shows it spending so many tokens that it costs more per task than Claude Opus 5.5, a model whose tokens are priced at twice the rate.

The Price Card Did Not Move, So the Saving Comes From Token Count

Claude Sonnet 5.5 is cheaper than Sonnet 5 because it usually finishes the same job with fewer tokens and fewer tool calls, while the per-token price stays at $2 input and $10 output. Anthropic says as much on its launch page, where it describes the model as priced the same as its predecessor but typically needing far fewer tokens to do the same work, and it qualifies the 30% figure as the result of its own testing.

The baseline matters here, because Sonnet 5 was not a lean model to begin with. It launched in June with a new tokenizer and an introductory 2/10 price that Anthropic framed as a way to keep migration roughly cost-neutral, a point Memeburn flagged at the time when it noted that Sonnet 5 could use up to 35% more tokens per request than the model it replaced. Anthropic later made the introductory rate permanent on 10 August and cancelled the planned September rise to 3/15, which is why “same price as Sonnet 5” now means 2/10 rather than the older Sonnet rate. So when Sonnet 5.5 claims a 30% saving, it is claiming it against a predecessor that customers already knew as token-hungry, and one of the testers quoted by Anthropic, CodeRabbit, said plainly that the older model’s heavy token consumption and its tendency to run unnecessary web searches had disappeared in the new one.

Three Habits That Cut the Bill

The customer reports that Anthropic published alongside the launch describe three distinct behaviours, each of which removes tokens from a job in a different way:

  • Sonnet 5.5 writes and reasons more briefly, since Slack measured about 14% fewer output tokens across its Slackbot evaluations without changing any prompts, Box reported 12% fewer total tokens, and Balyasny Asset Management found the model using about 121,000 tokens per answer on 2,441 finance tasks where Sonnet 5 had used 497,000, which is roughly a quarter of the previous amount.
  • It groups tool calls together instead of making them one at a time, which Anthropic says early testers noticed in head-to-head runs, and Lovable put numbers on it by reporting a third fewer tool calls and about half as many shell executions to finish a coding task.
  • It takes fewer detours, as Base44 found across 118 real app builds, where Sonnet 5.5 needed 3.6 iterations per build on average against 7.7 for Opus 5, failed fewer tool calls than every other model in the company’s comparison and rarely stopped mid-build to wait for a user’s answer.

The second and third points matter more than they first appear to, because in an agent loop every tool call ends one turn and starts another, and each new turn sends the growing conversation back to the model. Caching makes those repeated reads cheap at $0.20 per million tokens, but a job that finishes in half as many turns still re-reads its context half as often, and it also generates less reasoning between those turns, which is the expensive part of the bill since output tokens cost five times as much as uncached input. The spread in those customer numbers, from 12% at Box to roughly 75% at Balyasny, is also the reason Anthropic wrote “up to” and “for most work” rather than promising a fixed discount.

Speed is a separate benefit that often gets folded into the same claim. Anthropic puts the output speed gain over Sonnet 5 at more than 30%, and Zendesk reported tickets being processed 20% faster, but faster output lowers waiting time rather than the invoice, since the bill follows the token count and not the clock.

Need Reliable Internet?

Business Fibre • Wireless • Hosting • VoIP

Get Quote
Claude Sonnet 5.5 token-saving habits infographic

Need Reliable Internet?

Business Fibre • Wireless • Hosting • VoIP

Get Quote

The Effort Dial Decides Whether Sonnet 5.5 Is Actually Cheaper

Effort is the setting that controls how long Claude reasons before answering, and it moves the cost of a Sonnet 5.5 task by a factor of more than 18 from the lowest level to the highest. Anthropic sets the default to Medium in the Claude apps and Claude Code and to High on the Claude Platform, and its launch page makes the strongest cost claims at those lower settings, where it says that on a number of evaluations the new model running at Low or Medium already outscores the best result Sonnet 5 ever posted, while spending roughly 10% as much on each task.

Independent numbers from Artificial Analysis support that claim at the bottom of the ladder, since Sonnet 5.5 at Low effort scores 36 on its Intelligence Index for $0.41 per task, while Sonnet 5 at Xhigh scored 34 for $2.87. The same data becomes more complicated once Opus 5.5 is placed beside it, because the two models share Anthropic’s effort labels but not the same cost curve.

Effort level Sonnet 5.5 score Sonnet 5.5 cost per task Opus 5.5 score Opus 5.5 cost per task
Low 36 $0.41 42 $0.55
Medium 41 $0.59 51 $1.34
High 47 $1.08 54 $1.82
Xhigh 52 $2.74 56 $3.46
Max 56 $7.60 58 $5.98

Reading across a single row makes Sonnet 5.5 look cheaper at every level except Max, but reading diagonally tells a more useful story. Opus 5.5 at High scores 54 for $1.82, which beats Sonnet 5.5 at Xhigh on both score and cost, and Opus 5.5 at Medium reaches 51 for $1.34, close to Sonnet’s Xhigh result at about half the price. In practice that puts the crossover around High effort, where Artificial Analysis also says Sonnet 5.5 is at its most cost-competitive, landing a point behind GPT-6 Sol’s best score while costing almost exactly the same per task. Below that point Sonnet 5.5 is the cheap option in Anthropic’s lineup, and above it a team is paying Opus-level money for a result that Opus can deliver at a lower setting. Anthropic’s own launch page concedes a softer version of this, saying Sonnet 5.5 complements Opus 5.5 best at lower effort and performs comparably at a similar cost at higher settings.

Claude Sonnet 5.5 benchmark comparison table

At Max Effort, Sonnet 5.5 Costs More Than Opus 5.5

At Max effort, Claude Sonnet 5.5 costs about $7.60 per task on the Artificial Analysis Intelligence Index, compared with $5.98 for Opus 5.5, even though Opus charges twice as much per token. Artificial Analysis reported that Sonnet 5.5 used around 193,000 output tokens per index task at that setting, the highest figure it has measured on any model, and that the resulting cost per task was roughly 50% higher than Sonnet 5’s, which is the opposite direction from the launch headline.

Anthropic’s own benchmark table carries hints of the same pattern. The 70.6% Terminal-Bench 4.0 score that put Sonnet 5.5 ahead of Opus 5.5’s 66.4% compares Sonnet at Max with Opus at Xhigh, and a reading of Anthropic’s cost chart by Kingy AI puts Sonnet at 61.5% when both models run at Xhigh, with Sonnet’s Max attempt costing $12.54 against $7.35 for Opus at Xhigh. A footnote on the launch page goes further on FrontierCode, where Sonnet 5.5 scored 46.2% at Max but 52.1% at Xhigh, because at the higher setting it more often ran a code-review skill that split work across many subagents, which in the cases Cognition examined led to timeouts or edits outside the task’s scope. Max, in other words, buys more work, and on at least one test that extra work made the result worse.

This is the same caveat that surrounded Opus 5.5 a week earlier, when Anthropic’s 40% saving turned out to rest on default settings and shrank as effort rose, a gap Memeburn covered in its look at how far Opus 5.5 subscriptions actually stretch. It is also why the comparison with GPT-6 Sol, which lists at the same 2/10 as Sonnet 5.5, depends so heavily on settings, as our Opus 5.5 and GPT-6 Sol breakdown found with the larger models.

Two caveats apply to the independent figures. Artificial Analysis ran part of its suite on a pre-release deployment that Anthropic later found had a bug affecting structured outputs, which Anthropic expects to have understated Sonnet 5.5’s scores slightly rather than inflated them, and Artificial Analysis plans to re-run the affected tests. The index is also one fixed mix of ten evaluations, so the dollar figures describe that workload and will not match any particular company’s bill.

Cheaper Tasks Tend to Mean More Tasks

Anthropic’s product page pitches the lower task cost as a reason to run Sonnet 5.5 repeatedly, and one of its testers described letting Opus 5.5 set a game’s architecture while Sonnet 5.5 carried out the implementation, a split that turns one expensive job into many cheaper ones. That framing points to the limit of any per-task saving, because Gartner described the same dynamic in August as the inference paradox, predicting that inference costs per agentic workflow will rise more than fivefold through 2028 as falling token prices subsidise longer and more complex workflows, with analyst Will Sommer warning that product leaders “cannot rely on more efficient token economics” to justify their AI spending.

Anthropic itself says Sonnet 5.5 does not advance the frontier of its models’ capabilities, which makes it the kind of release that fits comfortably inside the slowdown Dario Amodei called for this month, since it competes on efficiency rather than on new capability, a debate Memeburn has followed in its analysis of what pacing the frontier actually binds. For teams deciding whether to switch, the practical reading is that Sonnet 5.5 is genuinely cheaper than Sonnet 5 at the settings most people will use, that the saving comes from how the model works rather than from what Anthropic charges, and that anyone raising effort above High to chase benchmark-level results should compare the cost against Opus 5.5 at a lower setting before assuming the smaller model is the budget choice.

FAQs

Is Claude Sonnet 5.5 cheaper than Sonnet 5?

At the default Medium and High settings, yes, because Claude Sonnet 5.5 uses fewer tokens and fewer tool calls to finish the same work, and Anthropic says that cuts the cost per task by up to 30%. At Max effort, Artificial Analysis measured the opposite, with Sonnet 5.5 costing about 50% more per task than Sonnet 5.

How much does Claude Sonnet 5.5 cost?

Claude Sonnet 5.5 costs $2 per million input tokens, $10 per million output tokens, $0.20 per million cached reads and $2.50 per million cache writes on Anthropic’s API, which is the same as Sonnet 5 and half the per-token price of Opus 5.5.

Why does Claude Sonnet 5.5 use fewer tokens?

Early testers report that Claude Sonnet 5.5 reasons more briefly, batches tool calls instead of making them one by one, and avoids detours such as unnecessary web searches or stopping to ask questions mid-task, so each job takes fewer turns and produces less output.

Is Claude Sonnet 5.5 cheaper than Opus 5.5?

At the same effort label Claude Sonnet 5.5 is cheaper everywhere except Max, where Artificial Analysis measured $7.60 per task against $5.98 for Opus 5.5. Above High effort, however, Opus 5.5 at a lower setting often scores higher for less money.

Which effort level should I use with Claude Sonnet 5.5?

For routine coding and document work, Medium or High gives the best balance, and Artificial Analysis rates High as Claude Sonnet 5.5’s most cost-competitive setting. Tasks that seem to need Xhigh or Max are usually cheaper to route to Opus 5.5 at Medium or High.

The post Why Claude Sonnet 5.5 Is Cheaper Without a Price Cut, and When It Isn’t appeared first on Memeburn.

Need Reliable Internet?

Business Fibre • Wireless • Hosting • VoIP

Get Quote