Ramp’s latest spending data puts a number on a question that AI labs usually answer with benchmarks: how much will businesses pay for the best available model?
In the August 2026 Ramp AI Index, Anthropic’s Claude Fable 5 accounted for 11.4% of dollars spent on Anthropic models but only 6% of the company’s tokens. OpenAI’s GPT-5.6 Sol had a much broader footprint, representing 23% of OpenAI model spending and 25% of tokens. In absolute terms, Ramp estimates that Fable produced only about 75% as much model-attributed spending as Sol during July.
That would be notable on its own. The stronger signal comes from what happened next. Anthropic’s cheaper Opus 5 quickly overtook Fable in Ramp’s spending data, while OpenAI cut Sol’s price on August 21, 2026. Together, these developments suggest that the enterprise AI market is becoming less tolerant of large premiums for relatively narrow capability gains.
Ramp’s Data Reveals a Price Ceiling for Frontier AI

Image: August 2026 Ramp AI Index: Cracks in the AI thesis.
Fable’s share of Anthropic spending is almost 1.9 times its share of tokens. That gap reflects the model’s premium price, but it also shows that the higher price has not translated into comparable workload volume. Sol’s spending and token shares, by contrast, are close to proportional.
The percentages use different vendor-level denominators, so 6% of Anthropic tokens cannot be compared directly with 25% of OpenAI tokens as if they were portions of one market. Ramp’s absolute spending estimate addresses part of that problem: despite its higher price, Fable generated less model-attributed revenue than Sol in July.
There is also an important sample distinction. Ramp serves more than 70,000 businesses, and its economics dataset covers more than 50,000 U.S. businesses. However, the model-level Fable analysis comes specifically from Ramp’s token spend management product, not every company in the broader AI Index. Ramp says that subset is more technology-heavy than its normal sample, but it does not publish the exact number of firms included in this chart.
Ramp’s conclusion is still direct: Fable may reveal an upper bound on what businesses currently pay for frontier intelligence. Better performance has value, but it does not automatically justify a much higher cost across production-scale workloads.
Fable 5 Is an Escalation Model, Not a Default
Anthropic launched Fable 5 on June 9, suspended access on June 12, and restored it globally on July 1. Ramp’s July data therefore covers its first full month of stable availability, although the disrupted introduction could still have slowed initial integration and testing.
The model was always positioned for unusually difficult work. Anthropic describes Fable as a system for long-running coding, research, scientific, and knowledge tasks that lesser models cannot reliably sustain. Its standard API price is $10 per million input tokens and $50 per million output tokens.
That combination naturally limits volume. A business may use Fable to investigate a critical production failure, complete a complex software migration, or analyze a high-value research problem. It is less likely to route routine classification, extraction, summarization, and support tasks through the same model.
Low token share does not prove that Fable lacks value. A small number of successful runs could produce considerable economic benefit. It does show that businesses are treating the model as a scarce resource rather than making it their general-purpose AI layer.
Opus 5 Gives Anthropic Customers a Cheaper Substitute
Anthropic complicated Fable’s commercial position when it released Claude Opus 5 on July 24. The company says Opus approaches Fable’s frontier intelligence at half the price, charging $5 per million input tokens and $25 per million output tokens. On some of Anthropic’s evaluations, Opus came close to Fable’s peak results while completing tasks at a substantially lower cost.
Ramp’s chart shows Opus 5 overtaking Fable in model-attributed enterprise spending shortly after launch. That is an unusually clear example of internal substitution: Anthropic’s own near-frontier model appears to be absorbing demand that might otherwise have gone to its most capable product.
For most businesses, the relevant metric is not intelligence in isolation. It is the cost of reaching an acceptable result. If Opus completes a task reliably for half the model price, Fable must either solve something Opus cannot or reduce enough retries, human review, and execution time to recover the difference.
This is good product segmentation for customers. It is more difficult for the economics of a premium frontier tier, because every improvement to Opus narrows the set of workloads capable of supporting Fable’s price.
OpenAI Cut Sol’s Price After It Had Already Gained Traction
Ramp published its findings on August 12. Nine days later, OpenAI reduced GPT-5.6 Sol’s API pricing, strengthening the price-performance contrast identified in the report.
As of August 25, 2026, standard prices per million tokens are:
| Model | Input tokens | Output tokens |
|---|---|---|
| Claude Fable 5 | $10 | $50 |
| Claude Opus 5 | $5 | $25 |
| GPT-5.6 Sol | $4 | $20 |
OpenAI reduced Sol from $5 to $4 per million input tokens and from $30 to $20 per million output tokens. The promotional pricing is scheduled to remain available through at least November 21, 2026. At current standard rates, Fable costs 2.5 times as much as Sol for both input and output, while Opus costs 25% more. Cached tokens, long-context requests, batch processing, and other service tiers can alter the effective comparison.
The timing matters. The price cut cannot explain Sol’s July usage because it happened afterward. Sol had already reached 25% of OpenAI token volume and 23% of spending at its previous price. The reduction may now expand the number of workloads for which companies consider Sol economical.
OpenAI is effectively betting that greater volume, stronger customer retention, and broader production use are more valuable than maintaining the original frontier-model margin. That puts additional pressure on competitors to justify premium prices through measurable differences in completed work, not benchmark leadership alone.
The Best AI Model Is Becoming a Routing Decision
Enterprise AI systems rarely need one model for every request. Production teams can route work according to complexity, risk, latency, and expected value:
- Use an inexpensive model for repetitive, high-volume tasks.
- Escalate uncertain or difficult cases to Opus, Sol, or another frontier model.
- Reserve Fable-level capability for tasks where failure or human delay is particularly expensive.
- Measure the total cost of a successful outcome, including retries, tool calls, latency, and human review.
This structure makes the “best model” less important as a universal product. It becomes one component in a portfolio. The premium model handles the tail of difficult requests, while cheaper models process most of the token volume.
Model providers can still earn significant revenue from that tail, especially as AI agents take on longer and more valuable projects. The challenge is that near-frontier alternatives keep improving. A model priced at twice the level of a close competitor needs to produce a sufficiently higher task-completion rate, not merely a slightly better benchmark score.
Ramp’s findings also suggest that token prices are becoming visible to finance departments. Its AI Token Spend Management product tracks usage by provider, model, user, and project. Once businesses can identify which model generated a bill, routing expensive routine work through a flagship becomes harder to defend.
The Upper Bound Is Real, but It Is Not Permanent
Ramp’s evidence has limits. The model-level sample skews toward technology-focused businesses, covers a short adoption window, and does not reveal what companies accomplished with their tokens. The data cannot tell us whether a Fable task produced more business value than ten cheaper Opus or Sol runs.
The early availability disruption also makes Fable’s first month unusual. Enterprises need time to run evaluations, obtain approvals, update model routers, and confirm that a new system behaves reliably in production. Its share could rise as more applications are designed around long-running autonomous work.
Ramp’s “upper bound” should therefore be understood as a competitive price ceiling, not a fixed law. Businesses may pay much more when a model provides a unique capability tied to a valuable outcome. The problem for Fable is that Opus 5 and GPT-5.6 Sol already cover much of the surrounding territory at lower prices.
Final Thoughts
Fable 5 does not need to dominate token volume to be a useful or successful model. It can occupy the escalation layer where a difficult task is worth substantially more than the inference bill.
The harder commercial question is whether frontier capability can retain a broad pricing premium when near-frontier models arrive weeks later at half the cost. Ramp’s data, Opus 5’s rapid uptake, and OpenAI’s Sol price cut point in the same direction: enterprise AI spending will favor the cheapest model that reliably completes the work. Frontier intelligence still matters, but the market is forcing it to prove its value one workload at a time.
Frequently Asked Questions
4 questions
1What did Ramp find about Claude Fable 5 usage?
Ramp found that Claude Fable 5 represented 6% of Anthropic tokens and 11.4% of spending on Anthropic models during the measured period. GPT-5.6 Sol accounted for 25% of OpenAI tokens and 23% of OpenAI model spending. Ramp also estimated that Fable generated about 75% as much absolute model-attributed spending as Sol in July.







