The price-per-token lie: why cheaper models don't mean cheaper bills
Opus 5.5 changes input, output and cache economics. Price the accepted task, including retries and review, before choosing a cheaper model.
Jump to summaryWritten by Florian Bruniaux
AI Founding Engineer at Méthode Aristote, 13 years scaling engineering teams from developer to CTO. Builds open-source developer tools, see what else I've shipped.
TL;DR
| Decision | What to measure |
|---|---|
| Current reference | Opus 5.5, Anthropic’s recommended default for most workloads, with a dated price sheet and explicit effort; Fable 5.1 ($10 / $50) is the higher tier |
| Token accounting | Separate fresh input, cache creation, cache reads and output |
| Workflow accounting | Include failed attempts, retries and human corrections |
| Adoption decision | Cost per accepted task at the required quality |
Opus 5.5 changes the economics worth testing now. A model upgrade can reduce the price of tokens and the number of turns needed to finish. That is a reason to replace an old cost baseline, including the assumption that sending every easy task to a smaller model must save money.
Price an Opus 5.5 task from its actual usage
The standard Claude API prices, checked on 30 September 2026, give the following Opus 5.5 rates in USD per million tokens. These are standard rates, excluding Fast Mode, batch discounts, regional adjustments and tool charges.
| Token category | USD per million tokens |
|---|---|
| Uncached input | 4.00 |
| Five-minute cache creation | 5.00 |
| One-hour cache creation | 8.00 |
| Cache read | 0.20 |
| Output | 20.00 |
For illustration, a task with 100,000 uncached input tokens, 50,000 five-minute cache-creation tokens, 900,000 cache-read tokens and 20,000 output tokens costs $1.23 in model tokens at those rates. The categories are disjoint in this constructed example. It is not a measured review run. Retries and another review pass add their own usage.
Token counts do not carry over between model generations either. Anthropic’s pricing page states that the Claude 4.7+ tokenizer produces about 30% more tokens for the same text than earlier Claude models, so a per-token price comparison across generations has to be rerun on the same text.
Addy Osmani’s task-cost walkthrough on claude.dev, dated September 25, 2026, connects token prices to turns, cache reuse and completed work. Applied to code review, the next calculation is total campaign spend divided by accepted review tasks, with failed and partial attempts retained in the numerator. Report human triage separately instead of hiding it inside an attractive API number.
A subscription adds another measure: quota consumed to finish the same task. An API-equivalent estimate is useful for comparison, but it is not the subscription invoice. Part nine keeps those meters separate.
Earlier cost observations explain the mechanisms
The conference accounts below explain why token price alone was already insufficient. Their rates and model names belong to the period described. Opus 5.5, which Anthropic recommends for most workloads, is the evaluation candidate here, with Fable 5.1 reserved for demanding reasoning; these accounts supply mechanisms to inspect, not a model shortlist.
The catalog price hides three multipliers
The first multiplier is the input-output split. Cursor’s own usage data, which Martignole cites without fully vouching for the theory behind it, puts 90% of the tokens a coding agent consumes on the input side: the files it reads, the conversation history it resends on every single turn, the tool descriptions it carries along. Only 10% is the code actually generated. Providers price output several times higher than input per token, so a naive read of the price sheet suggests output should dominate the bill. In practice, the side priced cheaper per unit is the side that shows up nine times out of ten, and that’s what actually drives the total.
The second multiplier is caching. In Martignole’s July 2026 account, a cache read cost roughly a tenth of fresh input; the Opus 5.5 rates above make that ratio one twentieth, after Anthropic cut the Opus 5.5 cache-read price to $0.20 from its predecessor’s $0.50 on September 22, 2026. Neither ratio applies to the whole task bill. A model has no memory between turns: every message resends the entire conversation, system prompt, tool definitions, and codebase context included. Without caching, that repetition gets billed at full price every single time. With caching enabled, writing a new turn to the cache costs about 1.25 times the input price, but reading a cached prefix on the next turn costs a fraction of it (a tenth in Martignole’s account, a twentieth on Opus 5.5). On Anthropic’s raw API, caching is opt-in through a cache_control field; Claude Code enables it automatically, and other agents built on the API may or may not. Martignole learned this the expensive way on a personal project, a role-playing game he built on top of Claude’s API: without caching switched on, single sessions ran him thirty to forty dollars.
The third multiplier is the work needed to finish. Martignole’s July account described a lower-priced model consuming enough extra tokens to cost more, and a different coding harness changing the spend again. Those examples motivate a current test: hold the task and acceptance criteria fixed, then measure Opus 5.5 with the actual tools and effort you plan to deploy. A shorter successful run and an early incomplete stop can both look cheap until the outcome is checked.

Volume tells you nothing about spend
Vercel AI Gateway’s routing data for the week of July 9, 2026, gives Martignole his cleanest illustration of why the sticker price misleads. DeepSeek V4 Flash accounted for 39% of the tokens routed through the gateway that week, but only 3% of the actual dollars billed. Anthropic’s models, at roughly 20% of the same volume, accounted for 55% of the spend. Nearly twice Anthropic’s traffic on DeepSeek translated into 3% of the invoice, while Anthropic’s fifth of the traffic accounted for most of the budget. The price sheet has moved since: DeepSeek ended its off-peak discounts on September 5, 2025, and V4-Pro brought off-peak half price back on August 16, 2026, after the gateway week above, so a share measured in one week carries that week’s terms. Reading volume as a proxy for cost, as a leaderboard-driven adoption push would, gets the picture backwards.

Inference is the cost that never stops arriving
Percy Liang, Professor of Computer Science at Stanford, teaches that same distortion from the infrastructure side, in his CS336 course, “Language Modeling from Scratch.” Lecture 10 covers inference, and Liang frames the economics in one sentence that Martignole’s talk arrives at independently, from a completely different angle: training is a one-time cost, however large; inference is a cost you incur every single day. To make the scale concrete, Liang points to OpenAI’s API, estimated at roughly 8.6 trillion tokens generated per day (a figure that matches a16z’s own tracking of OpenAI’s API volume for October 2025). Against a GPT-4-sized training run, which Liang puts at around 32 trillion tokens, that means OpenAI regenerates a comparable volume of tokens in under four days of ordinary serving. For anyone trying to control a bill, Liang’s conclusion is that the number worth tracking sits downstream of the catalog price, in the realized cost per token actually generated in production, after caching, retries, and verbosity have all had their say. Martignole’s talk reconstructs that same number from Vercel and OpenRouter routing data instead of Stanford lecture slides.
The frontier models still get most of the money
Tuhin Srivastava, CEO and co-founder of the inference platform Baseten, made the corresponding point from the supply side in a June 2026 guest lecture for Stanford’s MS&E435, “Economics of the AI Supercycle”: at that point, about 90% or 95% of inference spend still went to frontier models, even as open-weight alternatives kept closing the capability gap. Asked where that spend is headed, Srivastava pointed to more agentic applications and bigger models, both of which mean more inference and more compute. That growth lands on the recurring side of the ledger, the side priced per token and billed every day.
What that leaves you with
Put Martignole’s receipts next to Liang’s lecture and Srivastava’s numbers, and the same gap appears in each: the unit price on a vendor’s page does not describe what a workload costs, and most spend sits in the recurring inference that the unit price hides. Part four of this series already named the fix organizations reach for once they notice: a FinOps function, a role with a name, someone accountable for the real number instead of the catalog one. But a role only works if it has something to measure against, and part two showed how often that visibility simply isn’t there, on a flat-rate plan with no per-call invoice to check the math against.
For Opus 5.5, use the published rates and the measured token categories, then inspect failed attempts and human rework. Reconsider the routing policy if the stronger baseline finishes enough tasks with fewer passes. Keep vendor-price changes in the budget scenarios, since a price sheet can move in either direction: the Sonnet 5 increase to $3 / $15 planned for September 1, 2026 was cancelled, while the Opus 5.5 cache-read price fell.
YSNK
(You should now know)
- Opus 5.5 needs a fresh cost baseline separating uncached input, cache creation, cache reads, output and retries
- Vercel’s gateway data makes the disconnect concrete: DeepSeek V4 Flash carried 39% of routed volume but only 3% of spend, while Anthropic’s 20% of volume accounted for 55%
- Stanford’s Percy Liang frames the same distortion from the infrastructure side: training is a one-time cost, inference is incurred every day, and he estimates OpenAI’s API regenerates a GPT-4-sized training run’s worth of tokens in under four days
- Baseten’s Tuhin Srivastava said in June 2026 that about 90-95% of inference spend still went to frontier models, with usage shifting toward agentic, compute-heavy inference
- A budget built on one vendor’s price sheet is fragile by construction. Part six of this series covers single-vendor dependence
Go Further in the Claude Code Guide
Practical resources selected to help you take the next step.
Related articles
Why AI ROI stays invisible, even to the people measuring it
Claims that nobody can measure AI ROI and that 95% of pilots return zero are weaker than their usual retellings, once their sources are checked.
A $214 AI coding agent rewrite
A verified $214 bill for a production rewrite, checked against the source transcript, sits beside a corpus of about 2,900 tech talks on coding-agent costs.
FinOps applied to tokens: who owns the AI bill
Cloud FinOps put a cost study first and a named owner for the number second. Three practitioner cases show that sequence applied to token bills.