Anthropic’s Boris Cherny, Head of Claude Code, has publicly acknowledged what a vocal group of users has been telling the company: Claude Opus 5 has real behavioral problems. In a recent exchange of replies on X, Cherny wrote:
“We know Opus is not perfect, and it is a big priority for the team to fix it.”
This wasn’t a formal Anthropic announcement. It was a direct response inside an X discussion, which makes the context important. The comment followed criticism that Opus 5 can feel lazy, sloppy, overly verbose, and inconsistent. A day earlier, Thariq Shihipar from Anthropic had similarly described Opus 5 as a “spiky” model and said improving its consistency was a major priority.
The X Exchange Shifted From Defense to Acknowledgment
The discussion began after Cherny pointed to tasks where Opus excels, including long-running optimization and coding work. That is a fair defense of a genuine strength, but it sounded to critics like the familiar argument that users simply hadn’t discovered the correct way to use the model.
Cherny’s follow-up changed the tone. He said he hadn’t intended to be dismissive, acknowledged verbosity and other reported issues, and offered claude /config outputStyle=concise as a short-term workaround. More importantly, he separated what Opus does well from the question of whether its current behavior is good enough.
That distinction is the real story. No AI model is perfect, so the sentence is unremarkable in isolation. In context, however, two Anthropic employees have confirmed that the Opus 5 discussion reflects issues the company wants to address. User criticism has clearly reached the people responsible for Claude.
My Problem With Opus 5 Goes Beyond Verbosity
From my experience, Opus 5 has been good at coding and disappointing almost everywhere else. It can navigate codebases, sustain engineering tasks, and occasionally identify implementation paths that other models miss. That aligns with Anthropic’s launch positioning, which emphasizes software engineering and long-running agentic work.
Outside coding, I’ve struggled to justify using it. Its writing is often inflated, awkward, and less controlled than I expect from a flagship model. It also fails to follow explicit instructions often enough that I have to repeat constraints or repair the answer. On top of that, it consumes tokens incredibly quickly in my workflows.
Those aren’t cosmetic complaints. When a model writes too much, ignores constraints, and burns through usage while doing it, the user pays three times: editing time, retry time, and token cost. A concise output setting may hide some verbosity, but it doesn’t automatically fix instruction adherence or poor judgment.
Anthropic argues that newer Claude models benefit from less prescriptive prompting. Its Claude 5 context-engineering guidance says the Claude Code team removed more than 80% of its system prompt without a measurable loss on coding evaluations. That may help users carrying complicated instructions from older coding setups. It doesn’t explain away weak writing or ignored constraints during ordinary use.
Opus 4.6 Set a Better Usability Standard
Opus 4.6 remains the comparison point because it felt more balanced. Anthropic introduced Opus 4.6 on February 5, 2026, as a model for both coding and everyday professional work. In my use, it was easier to direct and produced cleaner prose. Opus 5 may have stronger capabilities on certain long-horizon tasks, but raw capability isn’t the same as usability.
What I miss isn’t an old benchmark score. It is the confidence that a clear request will produce a controlled answer. When that confidence disappears, a newer model can feel like a downgrade even if it performs better on selected internal evaluations.
Anthropic Has Heard the Discussion. Now It Must Deliver
Two Anthropic figures responding in related X discussions is meaningful, but it isn’t a fix. The next step should be visible behavioral improvement across normal workflows, not another chart demonstrating Opus 5’s coding strength.
Constructive criticism will be most useful when it includes reproducible evidence: the prompt, model and effort setting, the instruction that was missed, the unwanted output, and a session or feedback ID when available. Tagging Cherny and Shihipar can keep the discussion visible, but detailed reports give Anthropic something its engineers can actually test.
I’d watch four areas: instruction following, default writing quality, verbosity, and token efficiency. If Anthropic improves those without weakening Opus 5’s coding ability, it could recover the broader appeal that Opus 4.6 had.
Why I Switched to GPT-5.6 Sol
For now, I use GPT-5.6 Sol. OpenAI says its updated ChatGPT version is tuned to deliver more focused answers, avoid unnecessary formatting, and behave more consistently across different reasoning levels. That matches what I currently value.
This isn’t a claim that Sol wins every task. It is a workflow decision. I need an AI model that can code, write clearly, respect detailed constraints, and avoid turning every request into a token-heavy negotiation. Opus 5 still has a place for difficult coding work, but it is no longer my default for everything else.
Final Thoughts
The encouraging part isn’t simply that Cherny used the phrase “not perfect.” It is that both he and Shihipar acknowledged specific behavioral concerns and described improving Opus 5 as a priority. Anthropic now has an opportunity to turn a tense X discussion into useful product feedback.
I hope the next Opus update feels less like a specialist coding engine and more like the dependable general model Opus 4.6 was for me. Until that happens, constructive criticism should continue, and my default will remain GPT-5.6 Sol.
Frequently Asked Questions
4 questions
1What did Boris Cherny say about Opus 5?
Boris Cherny said Anthropic knows Opus 5 isn’t perfect and that fixing it is a major priority for the team. The statement came during an exchange of replies on X rather than through a formal company announcement. He specifically acknowledged verbosity and other concerns while offering Claude Code’s concise output setting as a temporary workaround.








