Anthropic has released Claude Sonnet 5.5. The model scores 70.6% on Terminal-Bench 4.0, a massive jump from the 10.3% achieved by its predecessor, Sonnet 5. It is the second release in the Claude 5.5 family, positioned as a faster, cheaper alternative to the more expensive Opus 5.5.
In this article
The model is live on the Claude Platform, AWS, Google Cloud, and Microsoft Azure. Developers cannot host it themselves as it uses closed weights. It targets specific tasks like bug fixing and creating polished documents, slides, and spreadsheets.
What changed versus Sonnet 5
Anthropic lists four main improvements over the previous version:
- Speed: Output generation is 30% faster, making it the quickest Sonnet model yet.
- Cost per task: Expenses can drop by up to 30% due to reduced token and tool call requirements.
- Writing: Early testers describe the prose as clearer and the model as a better collaboration partner.
- Vision and long-horizon work: It is the first Sonnet model to beat Pokémon Red using only screenshots.
Technical specifications include a 1M-token context window and a 128K maximum output. The knowledge cutoff is June 2026. Adaptive thinking is enabled by default. Users can select effort levels ranging from low to max.
Benchmarks
The following scores are vendor-reported in the launch post. Detailed methodology is available in the Sonnet 5.5 System Card.
- Terminal-Bench 4.0: 70.6%, compared to 10.3% for Sonnet 5 and 66.4% for Opus 5.5 at Xhigh.
- CursorBench 4.0: 55.5%, roughly two points below Opus 5.5 (57.8%).
- FrontierCode 1.1: 52.1% at Xhigh and 46.2% at Max. GPT-6 Sol scored 49.3%.
- GDPval-AA v2.1: 1844, versus 1846 for Opus 5.5 and 1449 for Sonnet 5.
- OSWorld 2.1: 80.1% on computer use, close to Opus 5.5 (81.8%).
- Humanity’s Last Exam: 64.5% with tools, up from 54.9%.
Results at Max level are lower than at Xhigh. At Max, the model frequently runs multi-agent code reviews. This process sometimes caused timeouts or edits that fell outside the scope, which FrontierCode penalises. Anthropic notes that Opus 5.5 remains stronger for complex, open-ended work.
Pricing and efficiency
Standard pricing remains unchanged from Sonnet 5: $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 and cache writes cost $2.50 per million. This is half the price of Opus 5.5, which is $4 and $20 respectively. Savings come from token efficiency rather than a price cut.
Customer data supports these claims. Balyasny Asset Management measured about 121K tokens per answer versus 497K on Sonnet 5. Base44 reported 3.6 iterations per app build, where Opus 5 took 7.7. Zendesk processed tickets 20% faster.
Default effort levels vary by interface. Claude Code and the Claude apps default to Medium. The Claude Platform defaults to High.
How it compares
| Feature | Claude Sonnet 5.5 | GPT-6 Sol | Gemini 3.1 Pro Preview | Claude Opus 5.5 |
|---|---|---|---|---|
| Developer | Anthropic | OpenAI | Anthropic | |
| Price per 1M tokens (input/output) | $2 / $10 | $2 / $10 | $2 / $12 | $4 / $20 |
| Context window | 1M tokens | 1,050,000 tokens | 1,048,576 tokens | 1M tokens |
| Max output | 128K tokens | 128K tokens | 65,536 tokens | 128K tokens |
| Knowledge cutoff | June 2026 | April 20, 2026 | Not listed | June 2026 |
| Reasoning control | Adaptive thinking, 5 effort levels | 6 effort levels (none to max) | Thinking supported | Adaptive thinking (always on) |
| Inputs | Text, image | Text, image | Text, image, video, audio, PDF | Text, image |
| FrontierCode 1.1 | 52.1% (Xhigh) | 49.3% | Not reported | 54.4% |
| GDPval-AA v2.1 | 1844 | 1487 | Not reported | 1846 |
| Chartography (no tools) | 61.6% | 53.6% | Not reported | 64.4% |
| Release stage | Generally available | Generally available | Preview | Generally available |
| Open weights | No | No | No | No |
Benchmark scores are vendor-reported by Anthropic; GDPval-AA runs by Artificial Analysis. Prices are standard API list rates, verified September 28, 2026.
What it means
The shift to a 70.6% score on Terminal-Bench 4.0 is significant for developers. The previous version struggled at just 10.3%. This new benchmark measures the ability to write and debug code directly in the terminal. A score of 70.6% suggests the model can now handle complex code generation and debugging tasks that were previously unreliable. For teams relying on automated code generation, this moves the model from a novelty to a viable tool for production environments.




