Sonnet 5.5 vs Opus 5.5: Half the Price, Not Half the Bill
Sonnet 5.5 costs half as much per token as Opus 5.5. Then Artificial Analysis measured what a task actually costs, and the rate card stopped mattering.
- Per token, Sonnet 5.5 is half the price of Opus 5.5: $2 in and $10 out per million tokens against $4 and $20. Per task, it depends on the effort dial.
- At max effort, Artificial Analysis measured Sonnet 5.5 at about 193k output tokens per task against 119k for Opus 5.5. That puts a task at $7.60 on Sonnet and $5.98 on Opus.
- Matched by score, Opus 5.5 is cheaper. Both reach an index score of 56, and Opus gets there at xhigh for $3.46 a task while Sonnet needs max at $7.60.
- Sonnet 5.5 wins on speed (139 tokens a second at max vs 94) and at low to high effort on well-scoped work. Anthropic's own staff say not to run it at max.
Sonnet 5.5 is not more token-efficient than Opus 5.5 at the top end. At max effort, Artificial Analysis measured about 193k output tokens per task for Sonnet 5.5 against 119k for Opus 5.5, so a task costs $7.60 on Sonnet and $5.98 on Opus despite Sonnet's half-price tokens. At low to high effort on well-scoped work, Sonnet 5.5 is cheaper and faster.
I posted this the night Sonnet 5.5 dropped:
“I honestly don’t know where I fit this model into my portfolio… Subagent for Fable/Opus 5.5? I mean Opus is so good so nah… This is a weird one here.”
Then the numbers came in, and weird turned out to be the right word.
Half the price is not half the bill
The rate card says Sonnet 5.5 is a steal. Per Anthropic’s pricing docs, it’s $2 in and $10 out per million tokens. Opus 5.5 is $4 and $20. Same 1M context on both, and cache reads cost $0.20 on both.
So half price. Case closed… except nobody pays per token. You pay per task.
The number that broke the story
Artificial Analysis ran both models at max effort:
| Sonnet 5.5 | Opus 5.5 | Where it loses | |
|---|---|---|---|
| Intelligence Index | 56 | 58 | Sonnet, by 2 points |
| Output tokens per task | ~193k | ~119k | Sonnet, by 60% |
| Cost per task | $7.60 | $5.98 | Sonnet, by 27% |
| Output speed | 139 tokens/s | 94 tokens/s | Opus, by a lot |
| Price per million tokens (in / out) | $2 / $10 | $4 / $20 | Opus, by 2x |
Sonnet writes 193k tokens a task. Opus writes 119k. Artificial Analysis also puts Sonnet 5.5 at about 7x the tokens of GPT-6 Astra at max.
Caveat, straight from the source: those Sonnet numbers came from a pre-release build with a structured-outputs bug, and Artificial Analysis says it will re-run them. Anthropic expects “minimal change or slightly understated performance.” I’ll update when they do.
Wait, Anthropic said fewer tokens
They did. The launch post says Sonnet 5.5 “typically needs far fewer tokens” and costs up to 30% less per task. Read what it’s compared to. It’s Sonnet 5, not Opus 5.5.
Customer numbers follow the same pattern. Balyasny went from about 497k tokens an answer to 121k. Slack saw 14% fewer output tokens. Box saw 12% fewer total tokens. All versus Sonnet 5.
@mreflow said what a lot of people were thinking: “very confused by Anthropic’s claims” next to the Artificial Analysis data. A follow-up post nailed the gap. The benchmarks run at max effort. The savings are measured at the default.
The effort dial is the price list
Same model, five settings. From Artificial Analysis’s Sonnet 5.5 page:
| Effort | Index score | Cost per task |
|---|---|---|
| Low | 36 | $0.41 |
| Medium | 41 | $0.59 |
| High | 47 | $1.08 |
| Xhigh | 52 | $2.74 |
| Max | 56 | $7.60 |
Going from high to max buys 9 points and costs 7x more. And it isn’t only Artificial Analysis. In a ComputingForGeeks test of three infrastructure tasks, run three times each, high used 11,868 output tokens for $0.12. Max used 135,413 for $1.36. Every lint passed both times.
Simon Willison hit the wall harder. The pelican test at max burned the full 128,000 tokens for $1.28 and never returned an answer. At xhigh, it finished in 41 seconds for 5.74 cents.
An Anthropic staffer, @edwinarbus, put it plainly:
“do not use Sonnet with max effort! at that point, you should probably be using Opus.”
Match the score, not the label
This is where the comparison flips. Line the models up by score, not by effort name, using the Artificial Analysis tables and beri.net’s matched-score breakdown:
Opus 5.5 at low scores 42 for $0.55. Sonnet 5.5 at medium scores 41 for $0.59. Opus is cheaper and a hair better.
At a score of 56, Sonnet needs max at $7.60. Opus gets there at xhigh for $3.46. That’s 2.2x the bill for the same result.
Sonnet’s cheaper at the same effort label. That’s the trap. The label isn’t the product.
Where each model actually wins
Anthropic’s own launch table, Sonnet 5.5 first:
| Benchmark | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% |
| FrontierCode | 46.2% | 54.4% |
| CursorBench 4.0 | 55.5% | 57.8% |
| Humanity’s Last Exam (tools) | 64.5% | 67.7% |
| OSWorld 2.1 | 80.1% | 81.8% |
| GDPval-AA | 1844 | 1846 |
Opus wins most rows. Sonnet takes Terminal-Bench, though Vals had it the other way: Opus at 61.62%, Sonnet at 53.03%, and almost the same cost per task ($19.07 vs $19.33). Vals notes 30 of Opus’s 198 attempts were served by older models through provider-side fallback, so read that gap with care. Anthropic’s footnote also says Sonnet scores lower at max than at xhigh on FrontierCode because of code-review skill timeouts.
On speed, Sonnet wins clean, and that’s real money when a human is waiting on the answer. One Claude Code user ran both and measured Sonnet about 1.25x faster. The same user found cache hits, which cost the same on both models, made a week of usage only 19% cheaper on Sonnet, not 50%. One account, one week, so treat that as a data point and not a law.
Then there’s the part no benchmark captures. One developer, after a day of real use, wrote that Sonnet is not “same as opus just 50% cheaper.” That only holds for problems that don’t need much wisdom. Opus asks what the real goal is. Sonnet asks how to get the task done.
What I’d run
Sonnet 5.5 at medium or high, for scoped work where the spec is clear and speed matters. Opus 5.5 for anything with an open question in it. Opus at xhigh where you’d have cranked Sonnet to max.
If you’re asking whether Sonnet is the cheap option for agents that run all day… only if you keep the dial low. Set the effort in the request and cap the output, or the meter runs while you sleep.
Here’s the prompt I’d give any agent before it touches a paid API:
Before you run this job, interview me one question at a time about what it needs to be right.
Then tell me the lowest effort setting likely to pass and what you would check to prove it.
Start at that setting. Only move up one level if the check fails, and tell me roughly what each step will cost before you take it.
Price per token was never the price. The dial is. Set it like you’re paying for it… because you are.
I’ll be watching for Artificial Analysis’s re-run. If the pre-release bug moves the Sonnet numbers, this piece gets updated with the new ones. For the last time cost per task changed a ranking, see Opus 5.5 vs Grok 4.7 vs MiMo V2.6, and for the budget tier, Muse Spark 1.3 vs Gemini 3.8 Flash.
#TheAIMogul
Bottom lineRun Sonnet 5.5 at medium or high for well-scoped work and Opus 5.5 for anything that needs judgment. Never pay for Sonnet at max, because Opus does the same job for less. The effort setting is the real price list.