Anthropic’s Opus 5.5 Guide Flags Agent Migration Traps
Always-on thinking, cache-sensitive effort settings and ambiguous progress updates make this prompting guide a practical checklist for production agents.
Listen
AI narration
13:30
0:00 / 13:30
AI SummaryGenerated from this article
Anthropic's Opus 5.5 migration guide warns that valid model responses can break production agents due to always-on thinking, cache-sensitive effort settings, and ambiguous progress signals. Developers must remove disabled thinking requests, rebudget max_tokens for thinking overhead, use per-message effort to preserve prompt cache, and track task completion with explicit checklists rather than relying on end_turn signals. Progress text may appear in omitted thinking blocks, and applications should inspect response blocks by type. Teams migrating Opus 5 integrations should test harness logic, tool declarations, cache behavior, and completion checks against actual workflows, since lower effort might generate fewer tokens yet lose cache reuse benefits.
An existing Opus 5 prompt may work on Claude Opus 5.5 while the application around it fails. A request that disables thinking returns an error. A token limit sized for a visible answer can cut that answer short. An agent loop can mistake a progress report for completed work.
Anthropic’s Opus 5.5 prompting guide describes how the model behaves in long-running tasks and what developers should change in their requests, interfaces and agent harnesses. For teams migrating an Opus 5 integration, the question is less “Should we rewrite our prompts?” than “Which assumptions in our application no longer hold?”
Remove Disabled Thinking and Rebudget Output
Claude Opus 5 could accept thinking: {"type":"disabled"} at high effort or below. Opus 5.5 cannot. Its thinking is always adaptive, and requests that disable it return a 400 invalid_request_error. Anthropic’s Opus 5.5 API change notes say to omit the thinking field or set its type to adaptive. Requests specifying a manual thinking budget with thinking: {"type":"enabled", "budget_tokens": N} also fail.
For an application that previously disabled thinking to keep simple turns fast, Anthropic recommends starting Opus 5.5 at low effort and measuring latency and answer quality on its own traffic. That advice applies specifically to the thinking-disabled migration path. More generally, the new model defaults to medium effort, whereas Opus 5 defaulted to high. Copying an old effort setting without an evaluation can change both behavior and cost.
The output budget needs attention too. Thinking counts toward max_tokens even when its contents are not returned to the application. A limit chosen for an Opus 5 request with thinking disabled may leave too little room for both Opus 5.5’s thinking and its visible reply. Anthropic says a max_tokens value of 128,000, the model’s maximum, has worked well in its testing of long agentic coding turns. That is guidance for workloads that need the room, not a reason to give every short chat response the maximum allowance.
Developers should inspect response blocks by type instead of assuming the first block contains displayable text. A response can begin with a thinking block; under the default omitted display setting, its thinking field is empty. Prompts that previously asked the model to write out reasoning as a substitute for disabled thinking deserve review too. Anthropic advises removing those instructions and using summarized thinking blocks where appropriate instead of trying to force full reasoning into the visible answer.
Work with Zeniteq
Let’s work together
We’re open to thoughtful collaborations with teams building in AI. Explore the ways we can work together.
Anthropic reports that Opus 5.5 at medium effort matches or exceeds Opus 5 at high effort on its coding and knowledge-work evaluations. That is a vendor-reported comparison, not evidence that medium is the best setting for every production task. An effort sweep against representative requests remains the safer way to choose.
Keep Effort Experiments From Breaking Cache Reuse
Effort is now a primary control for thinking depth, latency and cost. Where an application changes it also matters: Anthropic warns that changing top-level effort between requests invalidates the prompt cache. A router that alternates effort settings across turns of the same conversation may pay for repeated prompt processing even if its system instructions and task history appear unchanged.
The guide points to per-message effort, currently in beta, for turns that need a different level while preserving the cache. In an agent workflow, a routine status turn and a difficult code-review turn need not receive the same thinking allowance. Changing the top-level setting to accommodate each one, though, has a caching consequence.
Check cache behavior as well as answer quality during migration. A lower-effort request might generate fewer tokens yet save less than expected if it forfeits reuse of a large shared prefix. Anthropic also cautions that effort names do not represent the same amount of thinking across Opus 5 and 5.5; keeping “high” because an older deployment used “high” is not a meaningful like-for-like comparison.
An end_turn May Finish a Message, Not the Job
For unattended work, the guide’s harness warning is consequential. Opus 5.5 can report progress on a multipart task in a text-only message whose stop_reason is end_turn. An agent loop that interprets every such response as task completion can stop after an update, leaving the work unfinished.
Anthropic recommends tracking task parts in a checklist, such as a to-do tool or file, and checking that state when a turn ends. If items remain open and the model has not identified a blocker, the harness can send a short continuation naming the outstanding work. Another option is to define the completion condition in advance and use a separate, smaller model to check the conversation against it.
Neither option calls for an unlimited “continue” loop. Anthropic advises stopping after two or three automatic continuations on the same task so a genuinely stuck run can be reviewed. A background command or subagent that is still running also needs its result returned to the model before the harness declares the overall task complete.
end_turn says the model ended its turn. The application must decide whether the user’s task is complete, blocked, awaiting a running tool, or ready for another turn. A migration test should include a long task that produces an interim update before its final deliverable.
Check Where Progress Appears in the Response
Completion logic is only half the progress-update problem. Anthropic says some text that appeared between tool calls on Opus 5 comes back in thinking blocks on Opus 5.5. With the default display: "omitted" setting, the text in those blocks is empty in the returned response. An application that previously streamed that material as user-facing progress can appear to go quiet between tool calls without receiving an API error.
Review which block types the interface consumes, whether the chosen display setting returns the information it expects, and what users see while a long-running task continues. A silent interface is not proof that the model is idle. A visible update should not be presented as a finished result, either.
Give Multiagent Teams a Time Budget
Anthropic’s guide separately addresses time signals for multiagent harnesses. A team of agents coordinating a long task benefits from an explicit time budget instead of an open-ended instruction to keep investigating. A harness can make the deadline and remaining time available as task context, then reserve time for combining results and producing an answer instead of allowing every subagent to continue expanding its own work.
This is a coordination technique, not a substitute for application-level limits. Developers still need to decide when to stop spawning work, when to collect outstanding results and what to report if the deadline arrives before every branch finishes. Time guidance in a prompt cannot itself enforce a timeout.
Mark Pasted Material Without Treating Prompts as a Firewall
The guide also calls out instructions arriving inside text a user has pasted. When an agent handles logs, documents, emails or copied web content, the user may want the material analyzed. Commands inside that material are not automatically part of the user’s request.
A harness should make the boundary legible by clearly marking pasted content and telling the model what task to perform with it. For example, a request to summarize a pasted document should not silently promote a sentence inside the document into an instruction to change the agent’s tools or goals. Anthropic presents this as prompting guidance, not a guarantee against prompt injection. Where tools can modify files, send messages or access sensitive data, tool permissions and application checks must still carry the security burden.
Run the API Migration Checks Alongside Prompt Tests
The prompting guide is not the entire migration specification. Anthropic’s API notes identify other changes that can fail requests outright: Opus 5.5 does not support forced tool choices using tool_choice type any or a named tool; on the Claude API and Google Cloud, it also rejects the older computer_20251124 computer-use tool. Anthropic says that older tool continues to work on Amazon Bedrock.
Conversation history needs care as well. Thinking blocks are tied to the model and conversation, and changes to earlier content can cause a replayed block to be rejected under the documented binding rules. Applications that switch models mid-conversation should check whether the destination model can read the earlier thinking blocks; an unsupported block can be dropped without the request failing.
A release test needs to cover more than a sample prompt. Validate request settings, tool declarations, replayed histories, cache behavior, token limits, UI progress and completion checks against the workflows the application actually runs.
Final Thoughts
A migration can return valid responses while losing cache reuse, hiding progress from users or stopping an agent before its checklist is complete. Those failures are harder to notice than a 400 error and can make a capable model look unreliable.
For an Opus 5 integration, I’d test the harness before spending much time polishing prompt wording. The model’s reported performance gains matter only if the surrounding application gives it enough tokens, preserves the intended conversation and knows when the job is truly done.
Frequently Asked Questions
4 questions
1
Can Claude Opus 5.5 Run With Thinking Disabled?
No. Opus 5.5 rejects requests that set thinking to disabled, as well as requests that specify a manual thinking budget using enabled and budget_tokens. Anthropic says to omit the field or use adaptive thinking. If an Opus 5 application disabled thinking for speed, start testing Opus 5.5 at low effort and measure the effect on latency and quality.
2
Does Changing Effort Affect the Prompt Cache?
Yes. Anthropic warns that changing top-level effort between requests invalidates the prompt cache. For a turn that needs a different effort level, the guide points to per-message effort, a beta feature that preserves the cache. Developers should measure cache reuse during effort experiments instead of judging an effort setting only by its output tokens or latency.
3
Does end_turn Mean an Opus 5.5 Agent Has Finished?
No. A text-only progress update can end with stop_reason: "end_turn" while parts of the user’s task remain unfinished. Anthropic recommends checking an explicit task list or completion condition before stopping an unattended run. If work remains, the harness can request a bounded continuation; it should also wait for running commands or subagents to return their results.
4
Why Might an Opus 5.5 Agent Appear Silent Between Tool Calls?
Some progress text that previously appeared between tool calls can arrive in thinking blocks on Opus 5.5. Under the default omitted display setting, those blocks contain no displayable thinking text, so an interface built around the older response shape may show no update. Developers should inspect blocks by type and test what their chosen display setting actually returns.