It’s only been a couple of days since the amazing Claude Opus 5.5 was released. People have used it to build really cool projects because of how efficient the model is with tokens and how good it is at following instructions. In my review, I claimed that Claude is officially back.
Today, Anthropic released Claude Sonnet 5.5.
What’s new? How does it compare to Sonnet 5? Should you upgrade? And what’s in the fine print that most people won’t read?
Let’s answer all these questions below.
Sonnet 5.5 Is Faster and Cheaper
Here’s how Anthropic describes it in the official announcement:
Sonnet 5.5 is priced the same as Sonnet 5, but it typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task than its predecessor. It also generates output more than 30% faster, making it our fastest Sonnet model to date.
The price is still $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 per million, and you get 50% off if you use the Batch API.
Here’s a side-by-side comparison of the two similar models:

Sonnet 5.5 vs Sonnet 5. Image by Jim Clyde Monge
So how is it cheaper if the price is the same?
It’s not because of a new tokenizer. According to the migration docs, Sonnet 5.5 uses the same tokenizer as Sonnet 5, so the same text gives you the same token count. The model just writes less to finish the same task.
There’s also a small change that helps with cost. The minimum prompt size you can cache went down from 1,024 tokens to 512, so shorter system prompts can now be cached too.
Here’s a quick look at the specs:
- Context window: 1M tokens
- Max output: 128K tokens (300K on the Batch API beta)
- Knowledge cutoff: June 2026
- Default effort: Medium in Claude Code and Claude.ai, High on the API

Sonnet 5.5 vs Sonnet 5 vs Opus 5.5 vs Opus 5. Image by Jim Clyde Monge
On Terminal-Bench 4.0, Sonnet 5.5 scored 70.6%, which is higher than Opus 5.5 (66.4%). On GDPval-AA, a test built around real office work, it got 1844 Elo. Opus 5.5 got 1846. That’s pretty much a tie.
Now, there are two numbers here I’d take with a grain of salt.
First, look at Sonnet 5’s score on Terminal-Bench 4.0. It’s 10.3%. That’s the same model Anthropic called its most agentic Sonnet yet three months ago, as Jake Handy pointed out on his Substack post.
A jump from 10% to 70% looks great on a chart, but I think some of it comes from how poorly the older model handles this version of the test.
Second, Artificial Analysis ran Terminal-Bench on its own setup and got 64% for Sonnet 5.5, not 70.6%. It still came out ahead of Opus 5.5 and GPT-6 Astra (both at 60%), so Sonnet 5.5 is still the leader. The gap is just smaller than what Anthropic showed.
Sonnet 5.5 Is More Expensive Than Opus 5.5 on Artificial Analysis
Let’s look at the technical summary from Artificial Analysis:

Artificial Analysis Intelligence Index. Image by Jim Clyde Monge
Sonnet 5.5 scored 56 on the Intelligence Index. That puts it at #2, only two points behind Opus 5.5, and 18 points ahead of Sonnet 5.
Now, what’s interesting about the results is in the Cost per Intelligence Index Task section. I can see that it’s the most expensive model after Fable 5.1.

Artificial Analysis Cost per Intelligence Index Task. Image by Jim Clyde Monge
It’s even more expensive than the recently released Claude Opus 5.5.
Why is that?
It comes down to how much the model thinks. At max effort, Sonnet 5.5 used about 193,000 output tokens per task. According to Kingy AI’s breakdown, around 142,000 of those were reasoning tokens.

Sonnet 5.5 vs Sonnet 5 vs Opus 5.5 vs GPT 6 Sol pricing. Image by Jim Clyde Monge
For comparison, Opus 5.5 used about 119,000, and Sonnet 5 used about 118,000.
Artificial Analysis says this is the highest number of output tokens per task they’ve ever recorded. It’s around 60% more than Opus 5.5 and about 7x more than GPT-6 Astra.
So even though each token is cheaper, Sonnet 5.5 uses so many of them that the total goes up. At max effort, it costs about $7.60 per task. That’s roughly 50% more than Sonnet 5, which is the opposite of what the announcement says.
Effort Setting = How Much You Pay
So is Anthropic wrong about the 30% savings? Not really. There’s one line in the docs that explains it:

An effort level doesn’t produce the same amount of thinking as it did on Claude Sonnet 5. Re-run your effort sweep rather than carrying a setting over.
Basically, “medium” on Sonnet 5.5 doesn’t mean the same thing as “medium” on Sonnet 5. You save money when you run Sonnet 5.5 at a lower effort, and it still does as well as Sonnet 5 did at a higher one. If you keep your old settings and leave it on max, you’ll probably end up paying more.
The difference between settings is big. On low effort, Artificial Analysis has Sonnet 5.5 at 36 on the index, and it only used 23 million output tokens for the whole test. On max, it scores 56. Those extra 20 points cost a lot of thinking.
And max effort doesn’t always give you better results. Here are two examples I found:
- Simon Willison’s pelican test. On max, Sonnet 5.5 used 128,000 tokens ($1.28) and still failed to draw a pelican riding a bicycle. On xhigh, it got it right in 41 seconds for about 6 cents.
- FrontierCode 1.1. Sonnet 5.5 scored 52.1% on xhigh and only 46.2% on max. Vellum’s breakdown says the model kept sending work to subagents on max, which led to timeouts and edits outside the task.
So if you’re going to use Sonnet 5.5, I’d start with high or xhigh and only go up to max if you really need to.
What Devs Should Know
To my fellow devs out there who build on Anthropic’s API, take note that some of these changes in Sonnet 5.5 will break your code.
You can find all of them in the migration docs.

- You can’t turn thinking off anymore. Sending
thinking: {"type": "disabled"}returns a 400 error. The lowest option isbetween_tools, which skips thinking at the start. - Forced tool use was removed. Setting
tool_choicetoanyor to a specific tool gives you an error. Onlyautoandnonework now. If you force a tool call to get structured output, you'll have to change that. - Progress updates can disappear. Longer notes between tool calls now come back as thinking blocks, and with the default display setting, they’re empty. Your app won’t crash, but it’ll stop showing updates.
- Thinking blocks are tied to your account. They only work in the account that created them, and you can’t reuse them after editing earlier messages.
That last item is part of something new. Sonnet 5.5 is the first Sonnet model with classifiers that block “reasoning extraction.” That’s when someone tries to pull out the model’s chain of thought, usually to train another model with it.
The API even has a refusal type called frontier_llm, for requests that "could assist development of competing AI models." I found that pretty interesting to see written down so directly.
A few more things from the system card are worth mentioning. Sonnet 5.5 hallucinates more than Opus 5.5, even though it’s more honest under pressure.
Artificial Analysis measured a 47% hallucination rate on its Omniscience test. So if you use it for research or anything heavy on facts, I’d double-check its answers.
What Cool Things Can You Do with Sonnet 5.5?
Here’s Claude Code creator Boris Cherny sharing a side-by-side comparison of Sonnet 5 vs Sonnet 5.5 fixing a bug.
I shared in my previous post about the . Most of them were motion graphics and animation made with JavaScript.
Sources
- Claude Opus 5.5generativeai.pub
- really cool projectsgenerativeai.pub
- official announcementanthropic.com
- Jim Clyde Mongemedium.com
- migration docsplatform.claude.com
- Claude.aiclaude.ai
- Jake Handy pointed outhandyai.substack.com







