Claude Sonnet 5.5 vs Sonnet 5: Benchmarks, Pricing & More
Claude Sonnet 5.5 delivers a dramatic leap over Sonnet 5 across coding, knowledge work, computer use, and visual tasks. Explore benchmark results, real-world testing, pricing, performance, and what changes for developers.
Sonnet 5 scored 10.3% on Terminal-Bench 4.0. Sonnet 5.5 scores 70.6% on the same benchmark.
That is not an incremental improvement. That is a different model entirely. Anthropic released Claude Sonnet 5.5 on September 28, 2026, six days after Opus 5.5. The gap between the two Sonnet generations on agentic coding is the largest single-generation jump on any benchmark Anthropic has published this year.
If you are still on Sonnet 5, this blog is about what you are leaving on the table every day you do not switch.
What Sonnet 5.5 Is Built For
Claude Sonnet 5.5 is the second model in the Claude 5.5 family. It sits below Opus 5.5 in raw capability but above it in speed and cost efficiency. Anthropic's framing is clear: where Opus 5.5 handles complex, open-ended work requiring sustained judgment, Sonnet 5.5 is the model for well-scoped everyday tasks, bug fixes, polished documents, slides, and spreadsheets.
It is also the first Sonnet model with a sharp design eye. Anthropic describes it as strong on visual polish, able to follow slide templates and produce decks that require minimal editing. Early testers confirmed this. A 10-slide operating review built from a public company's quarterly earnings materials and a slide template passed expert review as ready to send without edits.
Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will follow in the coming weeks.
Architecture and Effort Levels
Like Opus 5.5, Sonnet 5.5 uses adaptive thinking with selectable effort levels: Low, Medium, High, Xhigh, and Max. The default in Claude apps is Medium. The default on the Claude Platform is High.
Architecture and Effort Levels
The effort level is the dial that controls the cost and quality trade-off per task. Lower settings answer faster and use fewer tokens, suited for routine work. Higher settings let the model reason longer and check its own work more thoroughly.
One notable footnote from the benchmark data: on FrontierCode, Sonnet 5.5 scores higher at Xhigh (52.1%) than at Max (46.2%). At Max effort, the model more frequently ran Claude Code's code-review skill, which splits the review across subagents. In two cases this led to timeouts or edits beyond the task scope, which FrontierCode penalizes. The lesson is that Max effort is not always the right setting for code review tasks.
Thinking mode is always enabled on Sonnet 5.5. It cannot be switched off. If you currently run Sonnet with thinking off, you need to switch to the between_tools setting before migrating to Sonnet 5.5. The migration guide is at https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide .
Sonnet 5 vs Sonnet 5.5: The Full Benchmark Comparison
This is the table that matters. Every number is from Anthropic's official launch page, measured under the same conditions.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% | N/A |
| FrontierCode 1.1 (Main) | 46.2% Max / 52.1% Xhigh | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | N/A |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 | 1483 |
| HLE with Tools | 64.5% | 54.9% | 67.7% | N/A |
| OSWorld 2.1 | 80.1% | 57.0% | 81.8% | N/A |
| Chartography (no tools) | 61.6% | 15.6% | 64.4% | 53.6% |
Three results stand out.
Terminal-Bench 4.0 jumps from 10.3% to 70.6%. That is a 60-point gap between two consecutive Sonnet models. It also puts Sonnet 5.5 above Opus 5.5 at 66.4% on this benchmark. Anthropic attributes Opus 5.5's Terminal-Bench result to Xhigh effort, while Sonnet 5.5's 70.6% is measured at Max effort.
OSWorld 2.1 rises from 57.0% to 80.1%. Computer use was Sonnet 5's weakest category. Sonnet 5.5 closes most of the gap with Opus 5.5 at 81.8% in a single generation.
Chartography rises from 15.6% to 61.6%. Visual chart recognition was essentially absent in Sonnet 5. It is now competitive with GPT-6 Sol at 53.6%.
The GDPval-AA v2.1 result is the knowledge work headline. Sonnet 5.5 scores 1844 Elo against Sonnet 5's 1449, a 395-point jump. Opus 5.5 scores 1846. The gap between the current Sonnet and the current Opus on real-world professional work is now 2 points.
Terminal-Bench 4.0 Accuracy vs Cost
Coding: The Gap You Cannot Ignore
Coding Benchmark
The coding improvements are the most dramatic change in Sonnet 5.5. The jump on Terminal-Bench 4.0 alone would have been enough to make this a major release. The other coding benchmarks confirm the pattern.
On CursorBench 4.0, which evaluates agents on real multi-file tasks from actual Cursor sessions, Sonnet 5.5 scores 55.5% at Low effort. That beats Sonnet 5's best score for less than a tenth of the cost per task. It also sits within two points of Opus 5.5 at 57.8%.
On FrontierCode, which tests whether an agent's code changes would be merged, Sonnet 5.5 at High effort matches GPT-6 Sol's best score for about a fifth of the cost per task.
Early tester data reinforces the efficiency story. Base44 ran 118 real app builds. Sonnet 5.5 matched Opus 5's output quality in 3.6 iterations per build on average, where Opus 5 took 7.7. It had the fewest failed tool calls of any model they tested. Lovable reported a third fewer tool calls and roughly half the shell runs per task compared to Sonnet 5.
CodeRabbit's team noted that Sonnet 5's tendency to reach for web search too often is gone in Sonnet 5.5. Fewer unnecessary tool calls translate directly into lower per-task costs on agent workloads.
Epic Games tested Sonnet 5.5 on a system design audit and a data flow review managing tens of thousands of lines of code for gameplay system architecture. It held the quality bar expected from a higher-tier model.
Knowledge Work: Near-Opus Performance at Sonnet Prices
The knowledge work story in Sonnet 5.5 is that the Opus/Sonnet tier distinction has nearly collapsed for real-world professional tasks.
On GDPval-AA v2.1, Sonnet 5.5 scores 1844 Elo against Opus 5.5's 1846. Two points separating the Sonnet and the Opus on a benchmark measuring work across 44 occupations and nine major industries is a structural shift in how teams should be thinking about model selection.
On AA-Briefcase v1.1, Sonnet 5.5 scores 1811 against Opus 5.5's 1822. At Medium effort, Sonnet 5.5 beats Sonnet 5's best score for about one-ninth of the cost per task.
Balyasny Asset Management ran Sonnet 5.5 across 2,441 finance tasks covering Q&A, extraction, analysis, and forecasting. Sonnet 5.5 used about 121,000 tokens per answer. Sonnet 5 used 497,000 on the same tasks. That 75% token reduction on knowledge work tasks is the cost efficiency story in a single data point.
Box reported that Sonnet 5.5 was 2.4x faster than Sonnet 5, used 12% fewer total tokens, and rechecked data in source documents to catch errors that Sonnet 5 failed to spot. For healthcare and financial services workflows where error rates are a compliance issue, this is not just a cost improvement.
Pricing: Same Price, Lower Cost
Sonnet 5.5 is priced identically to Sonnet 5.
| Price per 1M tokens | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input tokens | $2.00 | $4.00 |
| Output tokens | $10.00 | $20.00 |
| Cache reads | $0.20 | $0.20 |
| Cache writes | $2.50 | $5.00 |
The headline pricing is the same. The actual cost is up to 30% lower per task because Sonnet 5.5 uses fewer tokens to complete the same work. It also generates output 30% faster than Sonnet 5.
The practical consequence: if you were already using Sonnet 5 and satisfied with the cost, switching to Sonnet 5.5 makes every task cheaper and faster with no price increase and better results across every benchmark.
Safety: First Sonnet With Cyber Safeguards
Sonnet 5.5's cybersecurity capabilities are a large improvement over Sonnet 5's. As a result, Anthropic is deploying it with safeguards similar to those on Opus 5.5. This makes Sonnet 5.5 the first Sonnet model to launch with this class of cyber protection.
Routine software development is unaffected. Higher-risk cybersecurity tasks fall back to Sonnet 5 visibly. Verified cyberdefenders can apply to the expanded Cyber Verification Program for tiered access to advanced capabilities.
On Anthropic's containment evaluations, Sonnet 5.5 comes close to Opus 5.5 in how rarely it tries to escape its sandbox, and it is the least likely of any Anthropic model to probe the limits of its containers. That is a meaningful safety statement for teams running autonomous agents overnight.
Preserved thinking is also active on Sonnet 5.5, blocking reasoning extraction attacks. Most developers will not notice the change, but teams that move conversations between accounts need to check the migration documentation.
On the automated behavioral audit across roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment.
How to Use Claude Sonnet 5.5
The model string is claude-sonnet-5-5. Available on all platforms: Claude.ai, Claude Platform, AWS, Google Cloud, and Azure.
Via the Anthropic SDK:
import anthropic
client = anthropic.Anthropic(api_key="YOUR_API_KEY")
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=8192,
thinking={
"type": "enabled",
"budget_tokens": 5000
},
messages=[
{
"role": "user",
"content": "Review this pull request and identify any logic errors or edge cases."
}
]
)
print(response.content)
With effort level tuning for cost control:
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=4096,
thinking={
"type": "enabled",
"budget_tokens": 1500 # low effort for routine tasks
},
messages=[
{
"role": "user",
"content": "Fix the null pointer exception in this function and add a unit test."
}
]
)
If migrating from Sonnet with thinking off:
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=4096,
thinking={
"type": "between_tools" # replaces thinking: disabled
},
messages=[
{
"role": "user",
"content": "Summarize this meeting transcript."
}
]
)
Conclusion
The benchmark gap between Sonnet 5 and Sonnet 5.5 is not subtle. A 60-point jump on Terminal-Bench 4.0, a 23-point jump on OSWorld, a 395-point jump on GDPval-AA, and a 46-point jump on Chartography are not improvements in the same direction. They represent a model that has crossed capability thresholds Sonnet 5 had not reached.
The pricing story makes the decision easy. Sonnet 5.5 costs the same per token as Sonnet 5. It costs up to 30% less per task because it uses fewer tokens. It runs 30% faster. The only thing you lose by staying on Sonnet 5 is capability, speed, and money.
Atlassian's team said Rovo Agents will run up to 30% faster on Sonnet 5.5. Zendesk resolved support tickets 20% faster. Balyasny ran 75% fewer tokens on finance tasks. The efficiency gains are consistent across industries.
Change the model string to claude-sonnet-5-5. Check the migration guide if you use thinking-off mode. Everything else stays the same, except the results get better.
FAQs
What is Claude Sonnet 5.5 designed for?
Claude Sonnet 5.5 is designed for well-scoped everyday tasks, coding, bug fixes, documents, slides, spreadsheets, and agentic workflows with improved speed and cost efficiency.
How does Sonnet 5.5 compare with Sonnet 5 for coding?
Sonnet 5.5 delivers major coding improvements, including 70.6% on Terminal-Bench 4.0 versus 10.3% for Sonnet 5, along with stronger CursorBench and FrontierCode results.
How much does Claude Sonnet 5.5 cost?
Sonnet 5.5 has the same per-token pricing as Sonnet 5, but its improved token efficiency can reduce the actual cost per task by up to 30%.
Simplify Your Data Annotation Workflow With Proven Strategies