The official Claude Platform documentation has released the guide Prompting Claude Opus 5.5 — a set of prompting patterns for the claude-opus-5-5 model. The key difference from Opus 5: adaptive thinking is always enabled, and reasoning depth is set by a single effort parameter with a default of medium. According to the documentation, at this level the model matches or exceeds Opus 5 at high on coding and knowledge benchmarks, generating tokens more than 30% faster and using fewer tokens. At the same time, four breaking API changes have been documented that break configurations migrated from Opus 5 without modifications.

image
image

What Happened

The prompt engineering section of the Claude Platform documentation has been expanded with a page, Prompting Claude Opus 5.5, featuring patterns for working with the claude-opus-5-5 model. It lists the levels of the effort parameter — low, medium, high, xhigh, and max — and documents four breaking changes. First: requests with thinking set to disabled or with budget_tokens return a 400 invalid_request_error. Second: forced tool selection via tool_choice with values any and tool is no longer supported. Third: thinking blocks are tied to the model and a specific conversation, and prefix modifications are blocked by default for accounts from August 31, 2026. Fourth: the computer_20251124 tool has been replaced by computer_toolset_20260801 in the Claude API and Google Cloud. The documentation also includes a related page, What's new in Claude Opus 5.5.

Context

Previously, reasoning depth in flagship Claude models was a manual setting: the reasoning budget was set in tokens via budget_tokens, reasoning could be completely disabled for simple tasks, and guaranteed tool invocation was ensured via tool_choice. The new guide establishes a paradigm shift: reasoning becomes an inseparable part of inference, and instead of manual budgets and forced invocations, only the single effort parameter remains, determining how deeply the model thinks. In essence, this is not a list of new features but a change to the API contract: teams no longer decide whether the model thinks at all, but only choose the degree of depth. At the same time, the document remains vendor documentation, not a research publication, so its quantitative claims should be read as manufacturer statements.

Why This Matters for the Industry

For teams maintaining production agents, the effect is both migrational and economic. Existing pipelines with old settings break when migrated to claude-opus-5-5, and new deployments are possible immediately, but Opus 5 configurations can only be transferred manually: the first engineering task today is to audit current calls for thinking disabling, forced tool_choice, and the outdated computer tool. A separate operational risk: changing effort between requests resets the prompt cache, which directly impacts the cost and latency of agent cycles that vary reasoning depth on the fly. The economic signal is that Opus 5 high-level quality is now, according to the documentation, achievable at medium cheaper, faster, and with fewer output tokens; if parity is confirmed by independent measurements, the typical profile of agent pipelines will shift to medium, the ecosystem of harnesses, cost calculators, and task routers will begin to adapt to the semantics of effort, and comparing models by benchmarks without specifying effort will lose methodological meaning.

Why This Matters for Users

If you are testing claude-opus-5-5, the documentation recommends starting with effort medium and max_tokens 128000 for agentic coding: thinking consumes the max_tokens limit, even when reasoning blocks are not returned, so a low limit silently cuts quality. Opus 5 settings should not be transferred — the model's starting profile is different, and response blocks are better selected by the type field rather than by position in the array. Reasoning depth is proposed to be reduced by the effort parameter, not by prompting instructions: according to the documentation, this path provides Opus 5 high-level quality cheaper and faster.

What Is Still Unknown / Limitations

The main claim — that claude-opus-5-5 at effort medium matches or exceeds Opus 5 at high on coding and knowledge benchmarks — is not accompanied in the sources by a list of benchmarks, a description of the harness, sample sizes, and variance, so it requires independent verification on real workloads. The methodology for measuring the 30% token generation speedup is also not disclosed. The documentation remains vendor material: long-term scenarios such as the establishment of always-on thinking as an industry standard and the open question of whether effort will remain a controllable parameter or collapse into an automatic mode are interpretations, not guarantees. Before migrating, details such as the behavior of prefix modification blocking for thinking blocks for accounts marked with the date August 31, 2026, should be rechecked against the primary source.

Sources

Author

Look at AI, editorial team