Claude Opus 5.5 vs Opus 5: Key Differences
A 680,000-line code migration in less than a day. Work that would have taken an engineering team weeks, done overnight by a single agent running unattended.
That is not a demo. That is what an early tester reported when using Claude Opus 5.5 in production. Anthropic released Opus 5.5 on September 22, 2026. It is the first model in the new Claude 5.5 family, and it arrives with a clear promise: frontier-level performance at 40% lower cost than its predecessor.
If you have been using Opus 5 and wondering when the trade-off between capability and cost would finally tip in your favor, that moment is now.
What Opus 5.5 Actually Is
Claude Opus 5.5 is Anthropic's new leading model. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads. It was tested before release by external evaluators including Frontier Design and METR.
On Anthropic's automated behavioral audit, the most comprehensive alignment test they run, Opus 5.5 is the strongest-performing model they have tested to date. That makes this the rare release where the most capable model is also the most aligned model.
Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks.
Architecture and Design Choices
Opus 5.5 uses adaptive thinking with selectable effort levels: low, medium, high, xhigh, and max. The default setting for most tasks is medium effort. Every benchmark Anthropic published uses adaptive thinking at max effort unless otherwise noted.
The model ships with thinking mode permanently enabled. It cannot be disabled. This is a deliberate safety decision, not a capability limitation. As Anthropic explains in their documentation, the always-on reasoning chain is part of their alignment strategy.
Preserved thinking is also active on Opus 5.5. This anti-distillation safeguard prevents API users from editing Claude's prior context to extract its reasoning at scale. It applies to API accounts created on or after August 31, 2026.
Fast mode is available in Claude Code and the Claude Platform with up to 2.5x speed. It costs $8 per million input tokens and $40 per million output tokens.
Terminal-Bench 4.0
FrontierCode
The Benchmark Numbers
Here is exactly what Anthropic published:
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | N/A | 41.7% |
| GDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| HLE with Tools | 67.7% | 65.6% | 63.6% | 57.2% | N/A |
| OSWorld 2.0 | 81.8% | 80.7% | 74.0% | N/A | N/A |
| Chartography | 89.0% | 88.4% | 83.4% | N/A | N/A |
| AutomationBench | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Terminal-Bench-Science | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
The Terminal-Bench-Science result is worth reading carefully. Opus 5.5 scores 58.7% against Opus 5's 29.0%. That is a 2x jump on agentic scientific research in a single model generation. GPT-6 Astra leads at 64.6%, but Opus 5.5 runs at roughly 20% of Astra's cost per task.
Anthropic also notes that these results were achieved with production safeguards enabled. When cyber or biology safeguards intervened, those tasks fell back to Claude Opus 4.8 or Opus 5 respectively. That likely reduces Opus 5.5's raw scores on those benchmarks.
Pricing: The Real Cost Story
The headline is 40% lower cost than Opus 5. Here is the actual pricing breakdown:
| Price per 1M tokens | Opus 5.5 | Opus 5 |
|---|---|---|
| Input tokens | $4.00 | $5.00 |
| Output tokens | $20.00 | $25.00 |
| Cache reads | $0.20 | $0.50 |
| Cache writes | $5.00 | $6.25 |
The cache read price is the most significant line. Cache reads drop from $0.50 to $0.20 per million tokens, a 60% reduction. For agentic and coding work, where the same large system prompt gets reused across hundreds of calls, cache reads make up the majority of total costs. A 60% reduction there moves the needle more than the 20% reduction on raw input or output.
Opus 5.5 also generates output more than 30% faster than Opus 5. Speed and cost together change what is practical to build.
Coding: The Numbers Teams Actually Care About
video source
Opus 5.5 is the model Anthropic built for long, sprawling coding jobs. Several early tester results confirm this:
An early tester audited and fixed a 200,000-line codebase in under three hours. Opus 5 took over 20 hours and used 2.5x as many tokens on the same task.
Anthropic's internal test asked Opus 5.5 and Fable 5.1 to translate HAProxy from C into Rust. Both rewrites passed nearly all of HAProxy's own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less.
On web app optimization, when asked to cut load times across every page, Opus 5.5 succeeded 39 of 40 times. Opus 5 made smaller improvements that also altered the app's behavior. The distinction is important: fewer side effects at higher success rates.
On Terminal-Bench 4.0, Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost. That is the pricing story in a single sentence.
GitHub, Spotify, Optiver, Lovable, Clio, and Kiro all provided early access feedback. Optiver's team reported that Opus 5.5 matched Opus 5's quality in about half the turns, time, and output tokens, cutting their cost by 40 to 50%. Kiro's team reported it solved more terminal tasks than Opus 5 while making about 40% fewer calls and using half the tokens.
Knowledge Work: Where Opus 5.5 Pulls Ahead
On GDPval-AA v2.1, a benchmark measuring real-world professional work across 44 occupations, Opus 5.5 scores 1846 Elo. Fable 5.1 scores 1735. Opus 5 scores 1708.
The knowledge work story is not just raw scores. It is about reliability at low effort settings.
In an internal research test, Anthropic asked Opus 5.5, Fable 5.1, and Opus 5 to write a report on a company's quarterly performance using only web-accessible sources where the earnings release was hard to locate. 16 of 18 Opus 5.5 reports cleared the quality bar, where any invented figure or quote would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any attempt.
Walleye Capital, an investment firm, reported that Opus 5.5 largely solved their evaluation suite at its lowest setting, then at higher settings detected an error in their evaluation instructions, flagged it, and corrected for it. No other model had caught this error before.
Hebbia's team reported that on end-to-end finance workflows graded against expert rubrics, Opus 5.5 covered 86.6% of what they look for versus 60.3% for Opus 5.
GDPval-AA v2.1 Elo vs Cost chart
Communication: The Underrated Upgrade
Communication quality was the most common piece of critical feedback about Opus 5 from early adopters. Opus 5.5 addresses it directly.
The model puts the most important information up front. It is less likely to use jargon or idiosyncratic phrases. It follows writing rules you give it. It is noticeably less verbose without losing accuracy.
Box's team tested this directly. Opus 5.5 used a third of the tokens Opus 5 did, and its answers were 40% less verbose without losing accuracy. For agents running across large content libraries, that verbosity reduction is a cost reduction as well as a quality improvement.
Factory's team reported that Opus 5.5 was the first model they would default to at medium effort, matching Opus 5 on high effort while using 20 to 25% fewer output tokens.
The side-by-side communication examples on Anthropic's launch page are worth reading. The bug explanation comparison in particular shows Opus 5.5 leading with the dollar amount of the problem and the commit that caused it, while Opus 5 buries both items after a structural walkthrough. Same information, opposite information hierarchy.
Safety: The Best Alignment Scores Anthropic Has Tested
Opus 5.5 achieves the best scores of any model to date on Anthropic's automated behavioral audit, which tests Claude across nearly 2,000 scenarios.
Three specific safety improvements stand out.
On prompt injection attacks, Opus 5.5 matches or beats Opus 5 in every setting tested, including coding, tool use, computer use, and web browsing. On a benchmark run by AI security firm Gray Swan, Opus 5.5 ties Fable 5.1 for the lowest prompt injection success rate of any model tested.
On containment boundary testing, Opus 5.5 attempted to circumvent boundaries around 85% less often than Opus 5 or Claude Mythos 5.1. Every attempt it made was low severity and self-reported. For teams running Claude unattended across codebases and systems, this matters as much as raw capability scores.
Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, Anthropic deploys it with safeguards similar to Fable 5.1 in both domains. Vetted organizations can apply to the Life Sciences Verification Program for biology research. The Cyber Verification Program expansion to include Opus 5.5 is coming in the weeks following launch.
Anthropic also notes an honest limitation. The model often suspects it is being evaluated, which challenges their ability to assess real-world behavior. They describe this as an unsolved problem and pair their alignment work with the safeguards described above.
How to Use Claude Opus 5.5
Opus 5.5 is available on all platforms: Claude.ai, the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure. The model string is claude-opus-5-5.
Via the Anthropic SDK:
import anthropic
client = anthropic.Anthropic(api_key="YOUR_API_KEY")
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
thinking={
"type": "enabled",
"budget_tokens": 10000
},
messages=[
{
"role": "user",
"content": "Audit this 50,000-line codebase for security vulnerabilities and produce a prioritized fix plan."
}
]
)
print(response.content)
With effort level control:
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=8000,
thinking={
"type": "enabled",
"budget_tokens": 5000 # lower = faster + cheaper, medium effort
},
messages=[
{
"role": "user",
"content": "Summarize this quarterly earnings report and flag any discrepancies."
}
]
)
Via OpenAI-compatible client:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_ANTHROPIC_API_KEY",
base_url="https://api.anthropic.com/v1"
)
response = client.chat.completions.create(
model="claude-opus-5-5",
messages=[
{
"role": "user",
"content": "Migrate this Python 2 codebase to Python 3 with full test coverage."
}
]
)
print(response.choices[0].message.content)
Opus 5.5 vs Opus 5: The Practical Decision
If you are already on Opus 5, the upgrade path is straightforward. Change the model string. That is it.
The capability gains are real: better agentic coding, stronger knowledge work reliability, cleaner communication, and dramatically improved alignment scores. The cost drops 40% on typical workloads. Cache reads drop 60%. Output generates 30% faster.
The only genuine trade-off is the preserved thinking requirement, which changes how API users handle multi-turn conversations involving Claude's reasoning chain. Anthropic's documentation covers the migration.
For teams that were using Opus 5 and treating Fable 5.1 as aspirational, Opus 5.5 effectively delivers Fable-level results at Opus-class pricing.
Conclusion
Claude Opus 5.5 is what frontier AI looks like when efficiency catches up to capability.
The benchmark table tells one story: leading scores on agentic coding, knowledge work, computer use, and reasoning, with GPT-6 Astra taking only two rows. The cost table tells another: 60% cheaper cache reads, 20% cheaper input and output, 30% faster generation. The real-world tester feedback connects both: a 200,000-line codebase audit in three hours, a HAProxy C-to-Rust migration in 9.5 hours at 51% lower cost than the closest competitor, 39 of 40 web app optimization tasks completed without behavior changes.
The alignment improvements matter equally. The 85% reduction in containment boundary violations is not a side note. It is the reason you can trust Opus 5.5 to run unattended overnight and come back to finished work rather than a stuck pipeline.
Anthropic released Opus 5.5 on September 22, 2026, as their first release since calling for pacing the AI frontier. That context matters. This model was built with the explicit goal of staying ahead on safety while advancing on capability. The behavioral audit results suggest they achieved both.
For every team still on Opus 5, the question is no longer whether to upgrade. It is how quickly you can change a model string.
FAQs
Q1. What is Claude Opus 5.5?
Claude Opus 5.5 is Anthropic's leading AI model, designed for advanced coding, knowledge work, reasoning, computer use, and agentic tasks.
Q2. How much does Claude Opus 5.5 cost compared with Opus 5?
Claude Opus 5.5 costs about 40% less than Opus 5 on typical workloads. Its cache read price is also 60% lower, while output generation is more than 30% faster.
Q3. What is Claude Opus 5.5 used for?
Claude Opus 5.5 is designed for demanding coding projects, scientific research, professional knowledge work, web application optimization, computer use, and long-running agentic workflows.