Anthropic’s Boris Cherny, Head of Claude Code, has publicly acknowledged what a vocal group of users has been telling the company: Claude Opus 5 has real behavioral problems. In a recent exchange of replies on X, Cherny wrote:
“We know Opus is not perfect, and it is a big priority for the team to fix it.”
This wasn’t a formal Anthropic announcement. It was a direct response inside an X discussion, which makes the context important. The comment followed criticism that Opus 5 can feel lazy, sloppy, overly verbose, and inconsistent. A day earlier, Thariq Shihipar from Anthropic had similarly described Opus 5 as a “spiky” model and said improving its consistency was a major priority.
The X Exchange Shifted From Defense to Acknowledgment
The discussion began after Cherny pointed to tasks where Opus excels, including long-running optimization and coding work. That is a fair defense of a genuine strength, but it sounded to critics like the familiar argument that users simply hadn’t discovered the correct way to use the model.
Cherny’s follow-up changed the tone. He said he hadn’t intended to be dismissive, acknowledged verbosity and other reported issues, and offered claude /config outputStyle=concise as a short-term workaround. More importantly, he separated what Opus does well from the question of whether its current behavior is good enough.
That distinction is the real story. No AI model is perfect, so the sentence is unremarkable in isolation. In context, however, two Anthropic employees have confirmed that the Opus 5 discussion reflects issues the company wants to address. User criticism has clearly reached the people responsible for Claude.
My Problem With Opus 5 Goes Beyond Verbosity
From my experience, Opus 5 has been good at coding and disappointing almost everywhere else. It can navigate codebases, sustain engineering tasks, and occasionally identify implementation paths that other models miss. That aligns with Anthropic’s launch positioning, which emphasizes software engineering and long-running agentic work.
Outside coding, I’ve struggled to justify using it. Its writing is often inflated, awkward, and less controlled than I expect from a flagship model. It also fails to follow explicit instructions often enough that I have to repeat constraints or repair the answer. On top of that, it consumes tokens incredibly quickly in my workflows.








