Comparison Head to Head

Sonnet 5.5 vs Opus 5.5: Half the Price, Not Half the Bill

Sonnet 5.5 costs half as much per token as Opus 5.5. Then Artificial Analysis measured what a task actually costs, and the rate card stopped mattering.

Two identical vintage gas pumps in a dark garage, the left priced at $2 with a long paper receipt coiling across the floor and the right priced at $4 with only a short receipt.
Illustration generated for Run the Eval
The receipts
  • Per token, Sonnet 5.5 is half the price of Opus 5.5: $2 in and $10 out per million tokens against $4 and $20. Per task, it depends on the effort dial.
  • At max effort, Artificial Analysis measured Sonnet 5.5 at about 193k output tokens per task against 119k for Opus 5.5. That puts a task at $7.60 on Sonnet and $5.98 on Opus.
  • Matched by score, Opus 5.5 is cheaper. Both reach an index score of 56, and Opus gets there at xhigh for $3.46 a task while Sonnet needs max at $7.60.
  • Sonnet 5.5 wins on speed (139 tokens a second at max vs 94) and at low to high effort on well-scoped work. Anthropic's own staff say not to run it at max.
Short answer

Sonnet 5.5 is not more token-efficient than Opus 5.5 at the top end. At max effort, Artificial Analysis measured about 193k output tokens per task for Sonnet 5.5 against 119k for Opus 5.5, so a task costs $7.60 on Sonnet and $5.98 on Opus despite Sonnet's half-price tokens. At low to high effort on well-scoped work, Sonnet 5.5 is cheaper and faster.

I posted this the night Sonnet 5.5 dropped:

“I honestly don’t know where I fit this model into my portfolio… Subagent for Fable/Opus 5.5? I mean Opus is so good so nah… This is a weird one here.”

@MicahBerkley, September 28

Then the numbers came in, and weird turned out to be the right word.

Half the price is not half the bill

The rate card says Sonnet 5.5 is a steal. Per Anthropic’s pricing docs, it’s $2 in and $10 out per million tokens. Opus 5.5 is $4 and $20. Same 1M context on both, and cache reads cost $0.20 on both.

So half price. Case closed… except nobody pays per token. You pay per task.

The number that broke the story

Artificial Analysis ran both models at max effort:

Sonnet 5.5Opus 5.5Where it loses
Intelligence Index5658Sonnet, by 2 points
Output tokens per task~193k~119kSonnet, by 60%
Cost per task$7.60$5.98Sonnet, by 27%
Output speed139 tokens/s94 tokens/sOpus, by a lot
Price per million tokens (in / out)$2 / $10$4 / $20Opus, by 2x

Sonnet writes 193k tokens a task. Opus writes 119k. Artificial Analysis also puts Sonnet 5.5 at about 7x the tokens of GPT-6 Astra at max.

Output tokens per task at max effort: Sonnet 5.5 about 193k, Opus 5.5 about 119k.
Same effort label, very different appetite.

Caveat, straight from the source: those Sonnet numbers came from a pre-release build with a structured-outputs bug, and Artificial Analysis says it will re-run them. Anthropic expects “minimal change or slightly understated performance.” I’ll update when they do.

Wait, Anthropic said fewer tokens

They did. The launch post says Sonnet 5.5 “typically needs far fewer tokens” and costs up to 30% less per task. Read what it’s compared to. It’s Sonnet 5, not Opus 5.5.

Customer numbers follow the same pattern. Balyasny went from about 497k tokens an answer to 121k. Slack saw 14% fewer output tokens. Box saw 12% fewer total tokens. All versus Sonnet 5.

@mreflow said what a lot of people were thinking: “very confused by Anthropic’s claims” next to the Artificial Analysis data. A follow-up post nailed the gap. The benchmarks run at max effort. The savings are measured at the default.

The effort dial is the price list

Same model, five settings. From Artificial Analysis’s Sonnet 5.5 page:

EffortIndex scoreCost per task
Low36$0.41
Medium41$0.59
High47$1.08
Xhigh52$2.74
Max56$7.60

Going from high to max buys 9 points and costs 7x more. And it isn’t only Artificial Analysis. In a ComputingForGeeks test of three infrastructure tasks, run three times each, high used 11,868 output tokens for $0.12. Max used 135,413 for $1.36. Every lint passed both times.

Sonnet 5.5 used 11,868 output tokens at high effort and 135,413 at max effort on the same nine tasks, with all nine lints passing both times.
Max bought nothing on this job but a bigger bill.

Simon Willison hit the wall harder. The pelican test at max burned the full 128,000 tokens for $1.28 and never returned an answer. At xhigh, it finished in 41 seconds for 5.74 cents.

An Anthropic staffer, @edwinarbus, put it plainly:

“do not use Sonnet with max effort! at that point, you should probably be using Opus.”

@edwinarbus, September 28

Match the score, not the label

This is where the comparison flips. Line the models up by score, not by effort name, using the Artificial Analysis tables and beri.net’s matched-score breakdown:

Matched by score, Opus 5.5 costs less per task than Sonnet 5.5 at every point: 0.55 vs 0.59 dollars near a score of 42, 1.82 vs 2.74 near 53, and 3.46 vs 7.60 at 56.
Half the price per token. Not half the price per task.

Opus 5.5 at low scores 42 for $0.55. Sonnet 5.5 at medium scores 41 for $0.59. Opus is cheaper and a hair better.

At a score of 56, Sonnet needs max at $7.60. Opus gets there at xhigh for $3.46. That’s 2.2x the bill for the same result.

To reach an Intelligence Index score of 56, Sonnet 5.5 needs max effort at $7.60 per task, while Opus 5.5 gets there at xhigh for $3.46.
The rate card says half. The task bill says double.

Sonnet’s cheaper at the same effort label. That’s the trap. The label isn’t the product.

Where each model actually wins

Anthropic’s own launch table, Sonnet 5.5 first:

BenchmarkSonnet 5.5Opus 5.5
Terminal-Bench 4.070.6%66.4%
FrontierCode46.2%54.4%
CursorBench 4.055.5%57.8%
Humanity’s Last Exam (tools)64.5%67.7%
OSWorld 2.180.1%81.8%
GDPval-AA18441846

Opus wins most rows. Sonnet takes Terminal-Bench, though Vals had it the other way: Opus at 61.62%, Sonnet at 53.03%, and almost the same cost per task ($19.07 vs $19.33). Vals notes 30 of Opus’s 198 attempts were served by older models through provider-side fallback, so read that gap with care. Anthropic’s footnote also says Sonnet scores lower at max than at xhigh on FrontierCode because of code-review skill timeouts.

On speed, Sonnet wins clean, and that’s real money when a human is waiting on the answer. One Claude Code user ran both and measured Sonnet about 1.25x faster. The same user found cache hits, which cost the same on both models, made a week of usage only 19% cheaper on Sonnet, not 50%. One account, one week, so treat that as a data point and not a law.

Then there’s the part no benchmark captures. One developer, after a day of real use, wrote that Sonnet is not “same as opus just 50% cheaper.” That only holds for problems that don’t need much wisdom. Opus asks what the real goal is. Sonnet asks how to get the task done.

What I’d run

Sonnet 5.5 at low to high for well-scoped work, Opus 5.5 at low to medium for ambiguous hard work, Opus 5.5 at xhigh instead of Sonnet at max, and skip Sonnet 5.5 at max.
Pick the dial before you pick the model.

Sonnet 5.5 at medium or high, for scoped work where the spec is clear and speed matters. Opus 5.5 for anything with an open question in it. Opus at xhigh where you’d have cranked Sonnet to max.

If you’re asking whether Sonnet is the cheap option for agents that run all day… only if you keep the dial low. Set the effort in the request and cap the output, or the meter runs while you sleep.

Here’s the prompt I’d give any agent before it touches a paid API:

Before you run this job, interview me one question at a time about what it needs to be right.
Then tell me the lowest effort setting likely to pass and what you would check to prove it.
Start at that setting. Only move up one level if the check fails, and tell me roughly what each step will cost before you take it.

Price per token was never the price. The dial is. Set it like you’re paying for it… because you are.

I’ll be watching for Artificial Analysis’s re-run. If the pre-release bug moves the Sonnet numbers, this piece gets updated with the new ones. For the last time cost per task changed a ranking, see Opus 5.5 vs Grok 4.7 vs MiMo V2.6, and for the budget tier, Muse Spark 1.3 vs Gemini 3.8 Flash.

#TheAIMogul

Bottom lineRun Sonnet 5.5 at medium or high for well-scoped work and Opus 5.5 for anything that needs judgment. Never pay for Sonnet at max, because Opus does the same job for less. The effort setting is the real price list.

Frequently asked

Is Sonnet 5.5 cheaper than Opus 5.5?
Per token, yes: Sonnet 5.5 costs $2 per million input tokens and $10 per million output, against $4 and $20 for Opus 5.5. Per task, it depends on effort. At max effort Artificial Analysis measured $7.60 a task for Sonnet 5.5 and $5.98 for Opus 5.5, because Sonnet writes about 60% more output tokens at that setting.
Which model uses fewer tokens, Sonnet 5.5 or Opus 5.5?
At max effort, Opus 5.5 uses fewer: about 119k output tokens per task against about 193k for Sonnet 5.5, per Artificial Analysis. Anthropic's claim that Sonnet 5.5 needs far fewer tokens compares it with Sonnet 5, not with Opus 5.5.
What effort level should I use for Sonnet 5.5?
Medium or high. Anthropic's default is high on the API and medium in Claude Code and the apps. A hands-on test found max used about 11 times the tokens of high on the same nine tasks with identical lint results, and an Anthropic staffer publicly recommends not using Sonnet at max effort.
Is Sonnet 5.5 faster than Opus 5.5?
Yes. Artificial Analysis measured 139 output tokens a second for Sonnet 5.5 at max against 94 for Opus 5.5. Anthropic says Sonnet 5.5 is 30% or more faster than Sonnet 5.
When should I pick Opus 5.5 over Sonnet 5.5?
For ambiguous plans, hard judgment calls and long agentic work. Anthropic's benchmark table has Opus 5.5 ahead on FrontierCode (54.4% vs 46.2%) and Humanity's Last Exam with tools (67.7% vs 64.5%), and matched by score Opus is the cheaper way to get there.