Claude Sonnet 5.5: When Anthropic's Fastest Model Started Beating Itself
Anthropic's new Sonnet 5.5 runs 30% faster, costs 30% less, and scores 70.6% on Terminal-Bench — up from 10.3%. The model that was supposed to be the budget option is now dangerously close to Opus territory.
Anthropic just released Claude Sonnet 5.5, and the numbers are genuinely startling. This is the second model in the Claude 5.5 family, positioned as the faster, cheaper complement to Opus 5.5. But the benchmarks tell a more interesting story: Sonnet 5.5 isn't just a budget alternative anymore. It's pressing uncomfortably close to Opus on several fronts, and on one agentic coding benchmark, it actually beats its bigger sibling.
The Terminal-Bench Jump Nobody Expected
Here's the number that stops you cold: Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation. Sonnet 5 scored 10.3%. That's not an incremental improvement — that's a completely different model wearing the same name badge. For context, Opus 5.5 scores 66.4% on the same benchmark. The cheaper model outscored the flagship.
On FrontierCode 1.1, Sonnet 5.5 hits 46.2% at Max effort, compared to Opus 5.5's 54.4%. On CursorBench 4.0, it scores 55.5% versus Opus's 57.8%. The gap is narrowing with every release, and at this pace, the distinction between Sonnet and Opus is becoming less about capability and more about sustained judgment on open-ended problems.
30% Faster, 30% Cheaper — But the Real Savings Are Bigger
Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it Anthropic's fastest Sonnet model to date. Pricing stays the same as Sonnet 5: $2 per million input tokens, $10 per million output tokens, and $0.20 per million for cache reads. But the real cost story is in token efficiency.
Because the model needs fewer tokens to accomplish the same task, Anthropic reports it costs up to 30% less per task than its predecessor — even at identical per-token prices. This is the quiet revolution in AI pricing: models getting cheaper not because tokens cost less, but because models waste fewer of them.
First Sonnet to Beat Pokémon Red
Among the benchmark highlights, one stands out for its sheer audacity: Sonnet 5.5 is the first Sonnet model to beat Pokémon Red working only from screenshots. This isn't just a party trick — it demonstrates strong image understanding and long-horizon planning, the ability to maintain context across hundreds of game states and make coherent sequential decisions. It's the kind of task that requires both perception and persistence, and it signals that Sonnet 5.5 has moved beyond narrow task execution into something resembling sustained agentic behavior.
Safety: The First Sonnet With Cyber Safeguards
Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards and fallbacks previously reserved for Anthropic's most capable models. Its cybersecurity capabilities are comparable to Opus 5's, which triggered the added guardrails. Biology safeguards remain the same as Sonnet 5's. Both target a narrow set of high-risk requests — routine software development and most life sciences work are unaffected.
On automated behavioral audits, Sonnet 5.5 improves on or matches Sonnet 5 on most alignment measures. This matters because it suggests Anthropic isn't trading safety for speed — the model gets better and safer simultaneously.
The Shrinking Gap Between Sonnet and Opus
Here's what the broader picture looks like across key benchmarks:
- Terminal-Bench 4.0: Sonnet 5.5 at 70.6% vs Opus 5.5 at 66.4% — Sonnet wins
- FrontierCode 1.1: Sonnet 5.5 at 46.2% vs Opus 5.5 at 54.4% — Opus leads by 8 points
- CursorBench 4.0: Sonnet 5.5 at 55.5% vs Opus 5.5 at 57.8% — Opus leads by 2.3 points
- GDPval-AA: Sonnet 5.5 at 1844 vs Opus 5.5 at 1846 — effectively tied
- OSWorld 2.0: Sonnet 5.5 at 80.1% vs Opus 5.5 at 81.8% — Opus leads by 1.7 points
- Humanity's Last Exam: Sonnet 5.5 at 64.5% vs Opus 5.5 at 67.7% — Opus leads by 3.2 points
On GDPval-AA, a test of real-world work across occupations, the two models are separated by two points. On OSWorld, a computer-use benchmark, it's under two percentage points. The message from Anthropic's own data is clear: for most everyday work, Sonnet 5.5 is close enough to Opus that the price difference becomes hard to justify.
What This Means for Developers and Teams
The practical implications are significant:
- Agentic coding workflows just got dramatically cheaper. If Sonnet 5.5 can handle 70%+ of Terminal-Bench tasks, most CI/CD integrations and automated bug-fixing pipelines can switch from Opus to Sonnet without meaningful quality loss.
- Speed-sensitive applications — chatbots, real-time code completion, interactive agents — benefit from the 30%+ speed increase. Latency matters as much as quality for user-facing tools.
- The token efficiency gain means long-running agentic tasks (where token costs compound) see compounding savings. A 30% reduction in tokens per task on a 10-step agent workflow is a 30% reduction in total cost, not just per-call cost.
- Teams running mixed model strategies can now use Sonnet 5.5 as the default and reserve Opus 5.5 for genuinely complex, open-ended problems — getting the best of both worlds at a fraction of the cost.
The Bigger Picture: Model Tiering Is Getting Harder
Anthropic has maintained a clear three-tier strategy: Opus for complex work, Sonnet for everyday tasks, and Haiku for high-volume cost-sensitive applications. Haiku 5.5 is coming in the next few weeks. But as Sonnet climbs into Opus territory on benchmark after benchmark, the tiers are blurring.
This is a good problem for consumers. It means Anthropic is pushing capability down the stack rather than hoarding it at the top. The Sonnet 5.5 release suggests that the gap between tiers will continue to shrink — and that the next Opus will need to reach significantly higher to justify its price premium.
Anthropic notes that in their own testing and that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. That's probably true. But 'clearly stronger' is doing a lot of heavy lifting when Sonnet is faster, cheaper, and within a few percentage points on most benchmarks. For the vast majority of real-world use cases, Sonnet 5.5 isn't just good enough — it's the better choice.
The model that was supposed to be the budget option just became the default. That's not a pricing strategy — that's a paradigm shift.
Related Posts
Strata: When a 125B AI Model Ran on a Gaming PC at 100 Tokens Per Second
A new open-source tool called Strata lets you run Qwen 3.8 Flash Next — a 125-billion-parameter model — on an ordinary gaming PC with an RTX 4090. Nothing leaves your machine, and it's faster than you can read.
When AI Agents Spend Your Money While You Sleep: Why Hard Budget Caps Are Becoming Non-Negotiable
AWS and Google Cloud finally launched hard spending limits in the same month. It's not a coincidence — it's a response to AI agents that can rack up thousands of dollars before you wake up.
When Utah Banned VPNs: How a Court Stopped a Law That Demanded the Technically Impossible
A federal judge just blocked Utah's unprecedented anti-VPN law, ruling that lawmakers cannot mandate perfect geolocation — a technical impossibility. The case reveals a deeper problem: when legislation outruns engineering.