News · Model catalogue
Why Claude Sonnet 5.5 can cost less at the same price
Claude Sonnet 5.5 is reportedly priced the same as Sonnet 5 while claiming lower task cost. The distinction is usage, not a headline discount.

Simon Willison writes in his report on Claude Sonnet 5.5 that Anthropic’s new Sonnet model is priced the same as Sonnet 5. He also relays Anthropic’s claim that it runs “30%+ faster” and costs up to 30% less for most work.
That sounds contradictory only if a model’s listed price and the cost of finishing work are treated as the same thing. They are not necessarily the same. CNBC and TechCrunch independently reported the release: CNBC reports that Anthropic said the model does not advance the frontier, while TechCrunch reports faster response times and less token burn.
The important phrase is not “same price”. It is “less token burn”. That is the bridge between a stable price card and a lower charge for a completed task—and it is also where the claim needs the most scrutiny.
A listed price is a rate, a task bill is a total
A useful accounting model is simple: task charge equals the quoted unit price multiplied by the usage counted for that task. Hold the rate steady and reduce counted usage, and the total can fall. Raise the usage enough and the total can rise, even when the published rate has not moved.
That is the distinction Willison’s account puts into view. “Priced the same” describes the comparison with Sonnet 5; “costs up to 30% less for most work” describes a claimed result after some work has been done. The latter cannot be read as a new universal price for the model.
The published summaries do not provide a price card, a billing formula or a breakdown of the tokens involved. So the equation is a way to parse the claim, not a reconstruction of an Anthropic invoice. For lower token burn to become lower spend, the usage described as token burn must be usage for which the buyer is charged.
TechCrunch’s summary supplies the proposed mechanism—less token burn—but not the method. It does not say which tokens declined, how they were counted, or whether the saving came from a changed model behaviour, a changed default setting or a particular definition of the work being measured. Those missing details decide whether a claim transfers to a given workload.
A task needs a boundary before it has a cost
“Cost per task” sounds more concrete than it is. A task could be a narrowly defined response, a coding job assessed against a stated requirement, or a wider piece of work whose boundary includes several model interactions. The supplied reporting does not define what Anthropic includes in “most work”.
That omission is not a minor footnote. A result is only comparable when the input, the required output, the acceptance standard, the thinking-effort setting, the token count and the response-time measurement are kept visible. Without that, a lower total could reflect a genuinely more economical completion, a different amount of work, or a different rule for deciding that work is finished.
This is why a team should not turn “up to 30% less” into a budget forecast by itself. The phrase has a ceiling, a qualifier—“for most work”—and no disclosed workload definition in the material here. It is a vendor claim reported by Willison, not a measurement protocol someone else can reproduce from the release summaries.
There is a second boundary to keep in mind. The reports support a claim about the model’s task cost, not an established claim about the total cost of a business process. Nothing in the supplied material says how Anthropic defines that broader boundary, if it does at all.
Lower usage is the lever when the rate stays put
The commercial logic behind the release is straightforward in the abstract. If a vendor holds the advertised model price steady but promises a lower charge for a completed task, the claim must rest on less counted usage, a different billing arrangement, or both. TechCrunch’s reference to less token burn gives readers a reason to focus on usage rather than assuming the headline means a price cut.
What is not settled is how Sonnet 5.5 is meant to achieve that reduction. The supplied reporting does not describe an implementation, a task suite, or a before-and-after token count. It would be overclaiming to infer a particular technical mechanism from the cost claim alone.
The trade-off is also something to test, rather than something these summaries establish. A lower-usage completion is valuable only if it still meets the relevant acceptance standard; a faster response is valuable only if it is measured in a way that matters to the user. CNBC reports Anthropic’s separate claim that Sonnet 5.5 is better at coding and knowledge work, but the available summary does not set out the evaluation behind that statement.
That separation matters. A performance claim, a response-time claim and a cost claim can all be true under different conditions. None automatically proves the others.
Thinking effort can turn an average claim into an edge case
Willison’s report contains a more operationally useful warning than the release slogan. He writes that, at “max” thinking effort, Sonnet 5.5 thought for 128,000 tokens, describing the behaviour as the same bug seen with Opus 5.5.
The supplied material does not say whether those tokens were billable, what caused the bug, whether it was corrected, or how often it occurred. It should not be used to make a wider claim about normal Sonnet 5.5 behaviour. But it does show why the setting used in a cost comparison belongs beside the result.
An effort setting can be more important to a task bill than a broad statement about the model. If a maximum setting permits a very large amount of thinking, the task-level cost can become sensitive to that choice. A planning model that records only the model name and its listed price has missed a parameter that the report itself identifies as consequential.
This is also the failure mode in the phrase “up to”. A claim framed around typical or favourable work does not describe a ceiling on consumption. The max-thinking incident does not prove that all expensive tasks behave that way; it shows why an operator needs to know what happens at the edges, not only at the headline result.
The release changes what belongs on the evaluation sheet
The immediate operational lesson is to compare a model release at the level of completed work, not only at the level of its published rate. For a fixed set of jobs, record the model, the thinking-effort setting, token consumption, response time, acceptance condition and resulting charge. That makes it possible to distinguish a lower unit price from lower usage at the same unit price.
For Sonnet 5.5, that means treating the release as a claim to evaluate: same listed price as Sonnet 5, claimed lower cost for most work, reported lower token burn, and a reported edge case at maximum thinking effort. It does not mean every workload will cost 30% less. It does not mean a lower task cost establishes a frontier advance, either.
The phrase “token burn” is therefore not marketing texture. It is the accounting question behind the headline. Until the task boundary and settings are explicit, it is impossible to tell whether the apparent saving belongs to the model, the workload, or the way the result was counted.
The cost-control guide covers budgets, allowed models, right-sizing a model, token caps and offline batching.
- anthropic
- claude-sonnet-5-5
- ai-pricing
- token-usage
- thinking-effort
Questions
Is Claude Sonnet 5.5 cheaper than Sonnet 5?
Not in the price comparison Simon Willison describes: he writes that Sonnet 5.5 is priced the same as Sonnet 5. He also relays Anthropic’s claim that it costs up to 30% less for most work.
What does less token burn mean for task cost?
It can lower a task bill if the lower token burn is billable usage and the listed rate is unchanged. TechCrunch reports less token burn, but its summary does not state the billing formula or define a task boundary.
Does Sonnet 5.5 advance the frontier?
No, according to Anthropic as reported by CNBC. CNBC reports that the company said Sonnet 5.5 does not advance the frontier, while describing it as better at coding and knowledge-work tasks than its predecessor.
What was the Sonnet 5.5 thinking-effort bug?
Simon Willison writes that, at “max” thinking effort, Sonnet 5.5 thought for 128,000 tokens. He describes it as the same bug he observed with Opus 5.5; the supplied reporting does not explain the cause or outcome.
Sources
Every page this piece was written from.
About the author
odnoga Team
The odnoga team writes about artificial intelligence for the people who build with it: what shipped, what the research actually found, and what it means for the week ahead. Every piece names its sources.
More from this section
- GPT-6 Sol and Luna make routing a cost test · 25 September 2026
- Model IDs retired from the catalogue on 2026-09-14 · 21 September 2026
- Nine more models, and a prober that checks the rest · 20 September 2026
- AMD’s World Labs deal would bring a world-model lab in-house · 30 September 2026