Anthropic has released Claude Sonnet 5.5 — the second model in the 5.5 lineup after Opus 5.5, designed for clearly defined everyday tasks: bug fixing, working with documents, presentations, and spreadsheets. The company positions the new model as a faster and cheaper addition to the flagship, while the pricing remains at the Sonnet 5 level. In an independent measurement of work tasks, the model nearly matched Opus 5.5, and in agentic command-line programming it surpassed it. You can try Sonnet 5.5 right now — in Claude apps, via API, and with major cloud providers.

What happened
Anthropic released a mid-tier model and accompanied the release with its own measurements and an independent evaluation. According to the company, Sonnet 5.5 generates a response more than 30% faster than Sonnet 5 and is up to 30% cheaper per task due to lower token consumption. In Terminal-Bench 4.0, which tests agentic command-line programming, the model scored 70.6% versus 10.3% for Sonnet 5 and 66.4% for Opus 5.5. In OSWorld 2.1, which evaluates computer use, it showed 80.1% versus 57% for Sonnet 5 and 81.8% for Opus 5.5. In GDPval-AA — a measurement of work tasks across 44 professions conducted by Artificial Analysis — Sonnet 5.5 scored 1844 points versus 1846 for Opus 5.5 and 1449 for Sonnet 5. The model operates under the identifier claude-sonnet-5-5 and is available in Claude apps, via API, and also in AWS, Google Cloud, and Microsoft Azure at previous prices: $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads.
Context
The release completes the 5.5 lineup, which Anthropic started with the flagship Opus 5.5, and the promised Haiku 5.5 in the coming weeks should close the cheap segment for high-load scenarios. The concept of the series is to distribute models by roles: Haiku for mass inexpensive tasks, Sonnet as the main work model, Opus for the most complex cases. The second layer of context is the steady convergence of the mid-tier with flagships, due to which model choice is increasingly determined not by access to the best model, but by price and task type. The structure of the evidence base is also indicative: the key figure on near-complete parity in work tasks was provided not by Anthropic itself, but by the independent Artificial Analysis, which makes it more weighty than usual vendor assurances.
Why this matters for the industry
Sonnet 5.5 effectively closes the gap between the mid-tier and top-tier: nearly the level of Opus 5.5 in work tasks and higher in Terminal-Bench 4.0 at twice the lower input token price ($2 versus $4 per million). For product builders, this is a rare case where the mid-tier catches up to the flagship without changing vendors and without price increases: clearly defined steps in agentic pipelines — CLI coding, bug fixes, documents, presentations, spreadsheets — can be moved to claude-sonnet-5-5 with almost no loss of quality, leaving Opus 5.5 only for heavy steps. This price-performance ratio puts pressure on competitors' pricing policies and solidifies multi-level model routing as a typical inference architecture. A separate precedent concerns security: according to the System Card, this is the first Sonnet with Opus 5-level cyber protection, redirection of elevated-risk requests to Sonnet 5, and classifiers against distillation, meaning protective mechanisms have for the first time been scaled from the flagship to a mass model. The release of Haiku 5.5 in the coming weeks will complete the three-tier pricing ladder and will require reconfiguring routing in agentic products.
Why this matters for users
You can use the model right now: in Claude apps, via API, and in AWS, Google Cloud, and Azure, with the pricing remaining the same and the response becoming noticeably faster. It makes sense to move routine tasks — bug fixing, preparing documents, presentations, and spreadsheets — to Sonnet 5.5 from Sonnet 5 or partially from Opus 5.5 to reduce costs on typical operations with almost no loss of quality. Those building services on cheap models should watch the promised Haiku 5.5: it is aimed at high-load scenarios and may become an even cheaper option for mass requests. Of interest: Sonnet 5.5 was the first Sonnet model to complete Pokémon Red, relying only on screen screenshots, which clearly shows how far computer use by image has advanced.
What is still unknown / limitations
It is more accurate to speak not of proven parity with Opus 5.5, but of a two-point gap in GDPval-AA: confidence intervals and result variance have not been published. The measurement itself was conducted on a pre-release version with a bug in structured output, so final figures for the release build may differ. The Terminal-Bench 4.0 result is methodologically the most debatable: a sevenfold jump within the lineup in one generation is atypical, the sources mention "best effort level" for Opus 5.5, but the full conditions — version of the agent harness, number of runs, success criteria — have not been disclosed. Key figures currently rely on the vendor's report and a single independent measurement, so conclusions about parity should be rechecked against independent tests accumulated over the coming months. Finally, Haiku 5.5 is only promised, and its characteristics and timing will remain undefined until the official announcement.
Sources
- Introducing Claude Sonnet 5.5 — Anthropic
- Claude Sonnet 5.5 System Card (PDF) — Anthropic
- The Decoder — review of the Claude Sonnet 5.5 release
Author
Look at AI, editorial team
