Claude Sonnet 5.5 Release September 2026 - All the Info | Lexicon

Redaktion ·

What Claude Sonnet 5.5 is

Claude Sonnet 5.5 is Anthropic’s new mid tier — the faster, cheaper companion to Opus 5.5. Released on September 28, 2026, barely a week after Opus 5.5 (September 22). The pattern is routine for Anthropic by now: the expensive flagship first, the workhorse tier below it shortly after, meant to carry the bulk of daily volume. So Sonnet 5.5 is not the smartest model in the house but the one that brings together quality and affordable volume — exactly the role the Sonnet line has always played.

The model ID is claude-sonnet-5-5 (Claude API, Google Cloud, Microsoft Foundry, AWS; on Bedrock anthropic.claude-sonnet-5-5, via OpenRouter and Vercel as anthropic/claude-sonnet-5.5). In Claude Code it runs under the alias sonnet. Its knowledge cutoff is June 2026, the context window holds 1M tokens, output goes up to 128K (up to 300K in batch with a beta header). One detail for caching workflows: the cache minimum drops from 1,024 to 512 tokens.

Suitability profile in comparison

Suitability profile: Claude Sonnet 5.5
CodingReasoningTextVisionSpeedKosten-Eff.
  • Coding 5 / 5 · Feld-Spitze
  • Reasoning 4.5 / 5 · Claude Opus 5.5 u. a.
  • Text 4.5 / 5 · Claude Opus 5.5
  • Vision 4 / 4.5 · Claude Opus 5.5
  • Speed 4 / 4 · Feld-Spitze
  • Kosten-Eff. 4 / 4.5 · GPT-6.1 Sol

Eignung 0–5 · redaktionelle Einordnung, kein Benchmark · gestrichelt = Feld-Bestwert je Achse

The axes (Coding, Reasoning, Text, Vision, Speed, Cost efficiency) are an editorial read, not a benchmark. They show the Sonnet 5.5 pattern: strong coding and reasoning close to the Opus top, plus a clear edge over Opus 5.5 on speed and cost efficiency. Not a model that dominates one axis — one that doesn’t sag on any. The hard numbers are in the benchmark bars below.

Benchmarks: just behind Opus 5.5, clearly ahead of Sonnet 5

With the numbers, it pays to separate them by origin. Some come from Anthropic’s own launch test suite, some from vendor-independent measurements by Artificial Analysis.

From Anthropic’s own test suite (each in “max” effort):

| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | |---|---|---|---| | Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% | | FrontierCode 1.1 (xhigh) | 52.1% | 42.4% | 54.4% | | CursorBench 4.0 | 55.5% | 34.1% | 57.8% | | GDPval-AA v2.1 (Elo) | 1844 | 1449 | 1846 | | AA-Briefcase v1.1 (Elo) | 1811 | 1359 | 1822 | | Humanity’s Last Exam (tools) | 64.5% | 54.9% | 67.7% | | OSWorld 2.1 | 80.1% | 57.0% | 81.8% | | Chartography (no tools) | 61.6% | 15.6% | 64.4% |

Two things stand out. First, the gap to the predecessor: against Sonnet 5 the jump is clear in every row, dramatic on Terminal-Bench and Chartography. Second, the closeness to Opus 5.5: the big model leads in seven of eight benchmarks, but mostly by a thin margin. Only on Terminal-Bench 4.0 does Sonnet 5.5 flip it, landing ahead of Opus 5.5 at 70.6%. Important: these are vendor numbers, not a neutral comparison.

The vendor-independent Artificial Analysis Intelligence Index confirms the picture from a second direction:

Intelligence Index — Sonnet 5.5 against Opus 5.5, Astra and Sol 6.1

Artificial Analysis Intelligence Index, Stand Sep 2026 · höher = besser

58 ClaudeOpus 5.5 Anthropic 56 ClaudeSonnet 5.5 Anthropic 53 GPT-6Astra OpenAI 52 GPT-6.1Sol OpenAI

Werte: Artificial Analysis · artificialanalysis.ai

Sonnet 5.5 scores 56 points — behind Opus 5.5 (58, currently rank 1 in the index), but ahead of GPT-6 Astra (53) and GPT-6.1 Sol (52). For a mid-tier model that’s notable: Sonnet 5.5 beats OpenAI’s expensive Astra flagship in the index. So general intelligence is no compromise — the compromise sits elsewhere, in the effort.

The actual core: it all hangs on effort

This is the point the headline numbers hide. Sonnet 5.5 has five reasoning tiers — low, medium, high, xhigh, max — and performance scales hard with the tier. The impressive 70.6% on Terminal-Bench only holds in the most expensive “max” mode. The same task looks completely different depending on effort:

| Effort | Terminal-Bench 4.0 | Cost per attempt | Role | |---|---|---|---| | max | 70.6% | $12.54 | headline figure | | xhigh | 61.5% | $5.30 | — | | high | 43.0% | $1.94 | API default | | medium | 28.8% | $0.83 | Claude Code default | | low | 20.0% | $0.76 | — |

Read the two marked rows twice. Anyone calling Sonnet 5.5 over the API with the default setting gets high — and therefore 43.0% instead of the advertised 70.6%. Anyone using it in Claude Code gets medium — so 28.8%. The record value from the press release costs 15x the default per attempt and simply isn’t active in normal use.

On top of that: thinking is on by default and can only be reduced via {"type":"between_tools"}, and only up to high. From xhigh on, the model always thinks. The former fixed thinking budgets are gone.

Price and cost per task

Sonnet 5.5 costs $2 per 1M input and $10 per 1M output tokens — exactly Sonnet 5’s list price. Cache read is $0.20, cache write (5 min) $2.50, batch halves to $1/$5. There’s no surcharge inside the 1M window. On blended price, Sonnet 5.5 sits level with GPT-6.1 Sol and clearly below the flagships:

Price per 1M tokens (blended, list price, as of Sep 2026)

Ø aus Input- und Output-Listenpreis · niedriger = besser

$6.00 ClaudeSonnet 5.5 Anthropic $6.00 GPT-6.1Sol OpenAI $12.00 ClaudeOpus 5.5 Anthropic $15.00 ClaudeOpus 4.8 Anthropic $30.00 GPT-6Astra OpenAI $30.00 ClaudeFable 5.1 Anthropic

Werte: Anbieter-Preislisten, Stand Sep 2026 ·

But the token price is only half the math. Anthropic cites up to 30% lower cost per task for Sonnet 5.5 versus Sonnet 5 — at the same list price, because the model reaches the same result with fewer tokens (the range runs from 12 to 75% depending on workload). It also works more than 30% faster than Sonnet 5.

The interesting comparison is with its direct price rival GPT-6.1 Sol: same blended price, different profiles. Sonnet 5.5 leads the Intelligence Index (56 vs. 52) and, at around 142 tokens/s, is more than twice as fast as Sol 6.1 (~67). The flip side shows in the vendor-independent task cost: in “max” mode an index task runs about $7.60 on Sonnet 5.5, but only $0.72 on GPT-6.1 Sol. For long, high-volume agent runs where task cost decides, Sol 6.1 stays the cheaper pick — Sonnet 5.5 wins where speed and general intelligence matter. Again: task cost hangs directly on the chosen effort.

Positioning

The positioning map places Sonnet 5.5 against Opus 5.5 and the frontier field — x-axis speed and cost efficiency, y-axis capability and reasoning:

Positioning: Claude Sonnet 5.5 in the frontier field
Frontier Geschwindigkeit / Kosten-Effizienz → Fähigkeit / Reasoning ↑ 1 1 Claude Opus 5.5 2 2 GPT-6 Astra 3 3 Claude Fable 5.1 4 4 Claude Sonnet 5.5 5 5 GPT-6.1 Sol 6 6 Claude Opus 4.8

Redaktionelle Einordnung, kein Benchmark · Ausschnitt, Achsen gezoomt

The axes are an editorial read, not a benchmark. The pattern is clear regardless: Sonnet 5.5 sits top-right — high capability at high cost efficiency — exactly the zone where a workhorse belongs. Opus 5.5 sits just above it on capability but well to the left on price; the expensive frontier models Astra and Fable 5.1 cost more without leading the index correspondingly.

Breaking changes versus Sonnet 5

Sonnet 5.5 is not a pure drop-in replacement. Five API changes cause 400 errors if existing Sonnet 5 code stays untouched:

  1. You can no longer turn thinking off. {"type":"disabled"} is replaced by {"type":"between_tools"} (and that doesn’t apply at xhigh/max). The fixed thinking budgets are gone.
  2. No forced tool choice. tool_choice with any or tool is replaced by auto plus strict: true.
  3. Sampling parameters. temperature, top_p or top_k with non-default values combined with assistant prefill trigger a 400.
  4. Computer Use only runs with the computer_toolset_20260801 toolset.
  5. The advisor tool only works with Opus 5/5.5, Sonnet 5.5, Fable 5/5.1 and Mythos 5/5.1 — no longer with Sonnet 5 or Opus 4.8.

One more trap for multi-model setups: thinking blocks are bound to the conversation and account and cannot be ported between models.

Safety: the first Sonnet with cyber fallbacks

Sonnet 5.5 is the first Sonnet with server-side cyber fallbacks. Riskier security requests get routed internally to Sonnet 5 instead of being left to the stronger model. Add new refusal categories (cyber, bio, frontier_llm, reasoning_extraction) and protection against distillation. The fallback rate is low, around 1.5% — on Opus 5.5 it’s about 10%, so Sonnet 5.5 reaches for the fallback less often.

Three known limitations come with it: on Bedrock, structured outputs are missing. On long agent tasks, the model sometimes pauses for clarification instead of running through. And part of the knowledge-work benchmarks ran, per Anthropic, on a faulty pre-release deployment — read those values with care.

When Sonnet 5.5, when Opus 5.5, when GPT-6.1 Sol

The choice rarely follows “which model is best” but “which fits the task and the budget”:

  • Claude Sonnet 5.5 is the new default for the broad everyday: coding agents, RAG with long context, tool-use chains and high-volume pipelines where you want near-Opus quality without paying Opus prices on every run. Dial the effort up on purpose for heavy tasks, leave it on the default for volume.
  • Claude Opus 5.5 is the pick when maximum quality across many steps counts — it leads the Intelligence Index and is worth the premium on heavy, error-intolerant runs (architecture, complex analysis). The gap to Sonnet 5.5 is small in most benchmarks, though.
  • GPT-6.1 Sol is the direct price rival at the same blended price. It wins where cost per task in long agent runs decides (around $0.72 versus $7.60 per index task). Sonnet 5.5 counters with a higher index score and more than double the speed. Details in the GPT-6.1 Sol lexicon entry.

And because this market turns over fast — a week lay between Opus 5.5 and Sonnet 5.5 — model choice belongs in one central, configurable place in your system, not hard-wired into individual calls.

Availability, pricing and specifications

  • Release: September 28, 2026 (Anthropic), barely a week after Opus 5.5.
  • Access: Claude API, Google Cloud, Microsoft Foundry, AWS Bedrock (anthropic.claude-sonnet-5-5), OpenRouter/Vercel (anthropic/claude-sonnet-5.5), GitHub Copilot (Pro, Pro+, Max, Business, Enterprise — rolling out). In Claude Code as the alias sonnet.
  • Pricing: $2 input / $10 output per 1M tokens, cache read $0.20, cache write (5 min) $2.50, batch $1/$5. No surcharge inside the 1M window.
  • Context: 1M input tokens, up to 128K output (up to 300K in batch with a beta header). Cache minimum 512 tokens.
  • Reasoning effort: low, medium (Claude Code default), high (API default), xhigh, max. Thinking on by default.
  • Knowledge cutoff: June 2026.

FAQ

FAQ

Is Claude Sonnet 5.5 better than Opus 5.5?
In most benchmarks Opus 5.5 leads narrowly — it sits ahead of Sonnet 5.5 in the Artificial Analysis Intelligence Index (58 vs. 56) and wins seven of eight Anthropic benchmarks. Sonnet 5.5 beats Opus 5.5 only on Terminal-Bench 4.0 (70.6% vs. 66.4%), while being clearly cheaper and faster.
Why doesn't Sonnet 5.5 hit the advertised benchmark values for me?
Because the record numbers only hold in the most expensive 'max' effort. The API defaults to 'high', Claude Code defaults to 'medium'. On Terminal-Bench that means 43.0% (high) or 28.8% (medium) instead of the 70.6% from 'max'. For the top values you have to dial the effort up on purpose — which costs up to 15x per attempt.
What does Claude Sonnet 5.5 cost?
$2 per 1M input and $10 per 1M output tokens — the same list price as Sonnet 5. Anthropic cites up to 30% lower cost per task because the model reaches the goal with fewer tokens. Cache read costs $0.20, batch halves the prices.
Sonnet 5.5 or GPT-6.1 Sol?
Both have the same blended token price. Sonnet 5.5 has the higher Intelligence Index (56 vs. 52) and, at around 142 tokens/s, is more than twice as fast as Sol 6.1 (67). GPT-6.1 Sol wins on task cost in long agent runs (around $0.72 versus $7.60 per index task).
Can I move Sonnet 5 code to Sonnet 5.5 unchanged?
Not quite. Five changes cause 400 errors: thinking can no longer be turned off via 'disabled', forced tool choice is gone, certain sampling parameters with assistant prefill are rejected, Computer Use needs the new toolset, and the advisor tool no longer supports Sonnet 5.
See everything in one place:Sonnet