Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, a 90% reduction from Haiku 4.5’s published rates. Those prices apply only to requests with prompts of 100,000 tokens or fewer. Longer prompts move to a five-times-higher schedule.
Anthropic announced Claude Haiku 5.5 on October 7, 2026, with immediate availability through the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure. The model also introduces adjustable effort to the Haiku tier, giving developers another control over the cost and capability of individual tasks.
The release positions Haiku against OpenAI’s GPT-6 Luna for high-volume work. Anthropic reports stronger results on selected computer-use and coding benchmarks, while also halving Sonnet 5.5 cache-read pricing and introducing monthly API credits for Max and Team subscribers.
How much cheaper a completed task becomes depends on prompt length, the new tokenizer, reasoning effort and retries.
The 100,000-Token Threshold Changes the Entire Price Schedule
Haiku 5.5 has two published pricing tiers. The threshold affects output and caching rates as well as ordinary input charges.
| Price per million tokens | Haiku 5.5: prompts up to 100,000 tokens | Haiku 5.5: prompts over 100,000 tokens | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache reads | $0.01 | $0.05 | $0.10 |
| Cache writes | $0.125 | $0.625 | $1.25 |
The distinction applies at the request level. Anthropic publishes one schedule for prompts at or below the threshold and another for prompts above it; the higher rate does not apply only to the portion beyond 100,000 tokens.
For an illustrative request containing 90,000 uncached input tokens and 2,000 output tokens, the published rates produce a $0.010 bill. A request containing 110,000 input tokens and the same 2,000 output tokens costs $0.060. These examples exclude cache operations, additional reasoning tokens and other charges.
That is six times the cost for about 22% more input. Developers should monitor prompt size explicitly, particularly in agents whose conversation history grows over successive turns.
Anthropic says prompts within the cheaper tier accounted for roughly 90% of requests to its previous Haiku model. That helps explain the pricing strategy, though request share alone does not describe spending. A smaller number of long, output-heavy requests could still contribute substantially to an application’s bill.
The New Tokenizer Can Push Existing Prompts Into the Higher Tier
The Haiku 5.5 migration guide says the same input text produces approximately 30% more tokens than on Haiku 4.5, with the exact increase depending on content. Developers must recount prompts using the new model instead of reusing old token measurements.
An illustrative prompt previously measured at 80,000 tokens could become approximately 104,000 tokens. At Haiku 4.5’s input rate, the original prompt cost $0.08. At Haiku 5.5’s higher input tier, the retokenized version would cost approximately $0.052, excluding output and caching.
The reduction is still approximately 35%, well short of 90%.
Anthropic separately estimates an average per-task cost reduction of around 75% after tokenizer effects. This company-reported average does not guarantee the same saving for every application. Actual usage determines the result.
Adjustable Effort Makes Haiku’s Cost a Configuration Choice
According to Anthropic, Haiku 5.5 is the first Haiku-class model with adjustable effort. Its announcement shows performance and cost across settings ranging from low to max.
The migration documentation describes adaptive thinking, with effort controlled through output_config.effort. Lower effort can reduce thinking, and the model can skip it entirely on simpler requests.
A tightly scoped classification request may not need the same reasoning allocation as a browser task that requires several decisions and tool calls. Developers need to evaluate each setting against their application’s own acceptance criteria.
Adaptive thinking also creates a migration risk: it is on by default, and thinking tokens count toward max_tokens. A limit previously sufficient for Haiku 4.5 can leave too little room for Haiku 5.5’s final answer. Applications that assume the first response block contains answer text must instead select blocks by type.
One community observation illustrates how widely effort settings can differ. In the Hacker News discussion, a commenter posting as simonw reported that a pelican-on-a-bicycle SVG task took seven seconds at low effort and five minutes, nine seconds at max, with substantially different costs. A single reported task is not a general latency benchmark, but it cautions against treating every effort level as the same product experience.
Anthropic’s GPT-6 Luna Comparison Favors Haiku, With Caveats
Anthropic’s benchmark table presents Haiku 5.5 as a substantial improvement over Haiku 4.5 and a competitive small-model alternative to GPT-6 Luna. Two of the clearest comparisons concern computer use and terminal-based agent work:
| Company-reported benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| OSWorld 2.1, offline subset | 72.4% | 15.7% | 48.9% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
Anthropic describes OSWorld 2.1 as evaluating agents operating a computer through long, multi-step tasks. Its announcement labels the displayed computer-use evaluation as an offline subset, and the accompanying chart uses a partial-credit score. The 72.4% figure should not be recast as a universal task-completion rate.
Terminal-Bench 4.0 covers complex professional tasks performed through a command-line interface. Haiku’s reported improvement is large, while Sonnet’s considerably higher score preserves an important distinction within Anthropic’s range.
All figures here come from Anthropic’s announcement, not independent reproductions. A production configuration at low effort should not be assumed to reproduce a headline benchmark result. Establishing a lower total bill than GPT-6 Luna also requires a matched workload with equivalent task conditions, effort settings, tool usage, retry behavior and pricing.
For more complex agentic coding, Anthropic itself recommends Sonnet 5.5 and Opus 5.5. It positions Haiku around narrower work such as summarization, compaction and subagents. The benchmark gains are relevant to model routing, but they do not establish Haiku as a replacement for every larger-model task.
Multi-Cloud Availability Does Not Mean a Drop-In Upgrade
Anthropic says Haiku 5.5 is available now across its own platform and the three major cloud providers. On the Claude Platform, developers use claude-haiku-5-5.
The migration guide identifies platform-specific model IDs, including anthropic.claude-haiku-5-5 for Amazon Bedrock. Integration requirements and feature support are not necessarily identical across providers.
Sources
- Claude Haiku 5.5anthropic.com
- Haiku 5.5 migration guideplatform.claude.com
- Hacker News discussionnews.ycombinator.com





