FinOps applied to tokens: who owns the AI bill
Cloud FinOps put a cost study first and a named owner for the number second. Three practitioner cases show that sequence applied to token bills.
Jump to summaryWritten by Florian Bruniaux
AI Founding Engineer at Méthode Aristote, 13 years scaling engineering teams from developer to CTO. Builds open-source developer tools, see what else I've shipped.
TL;DR
| Decision | Current test |
|---|---|
| Review baseline | Measure Opus 5.5 on accepted changes before preserving the old routing policy |
| Small-model filter | Count savings and the risky changes it prevents the reviewer from seeing |
| Additional reviewer | Require a verified contribution after overlap and false alarms |
| Cost owner | Track all-attempt spend, quota and human triage separately |
Opus 5.5 gives FinOps a new decision to make: whether Anthropic’s recommended default for most workloads (Fable 5.1 sits above it at $10 / $50 per million tokens, for demanding reasoning) can complete work cheaply enough to simplify the routing system. Part five works through the current rates. A small model in front of every review is a candidate design whose savings and omissions need to be measured again. The AI FinOps overview on the Claude Code guide site maps the cost regimes and levers behind these routing choices.
Retest the filter against an Opus 5.5 baseline
Run a reserved set of changes through the bounded Opus 5.5 reviewer, then compare the proposed filtered route on the same task set. Keep independent reference defects hidden from both paths. Count the full filter cost, the reviews it admits, the verified defects it screens out and the human time required to recover them.
The filter earns its place when the resulting workflow meets the required quality at a worthwhile total cost. If it mainly adds an inference call and another failure mode, simplify the route. The same rule applies to an extra general reviewer. A review protocol that defines the effort comparison and adjudication has been drafted; those runs remain to be executed.
The earlier practitioner cases below explain how those routing patterns arose. Their model choices describe those teams at the time, while the adoption decision here starts from Opus 5.5.
In June 2025, a roundtable at the techready conference put three cloud practitioners on stage to talk about FinOps in practice: measuring and controlling cloud spend before it turns into a surprise on the finance team’s desk. Moderator Seifeddin Mansri asked the panel a pointed question. Did FinOps show up bottom-up, pushed by engineering teams already anxious about their own AWS or GCP bill? Or top-down, mandated by decision-makers who wanted costs under control before signing off on a migration? The panel’s histories split between the two paths, and the top-down one came with a sequence that AI spend now echoes.
Martin, who runs FinOps at SNCF Connect and Tech, described the top-down version from the inside: a full AWS migration across 2020 and 2021, with a mandatory step attached, a total cost of ownership study, run before the migration started. “The company decided to migrate fully to Amazon in 2020 and 2021, and there was a study done beforehand to evaluate the TCO, okay, we’re going to the cloud, how much is that going to cost us, partly to defend the case that it would cost less on the cloud. It’s always a very complicated calculation” (“l’entreprise a décidé de migrer intégralement sur Amazon sur 2020 et 2021, et il y a eu une étude qui a été faite avant pour évaluer le TCO, ça allait nous coûter combien, pour en fait aussi défendre le fait que ça allait coûter moins cher sur le cloud, c’est toujours très compliqué comme calcul”), he said. Martin joined the FinOps team afterward, recruited internally from a DevOps role, to keep answering that cost question once the migration was live.
That sequence, a cost study first and a named owner for the number second, has a counterpart in AI spend now that the cloud version of the discipline has settled into a job title. Three cases in this corpus show what it looks like moving from infrastructure bills to token bills: a small model standing guard in front of an expensive one, a local model tested against both a dollar figure and a carbon footprint, and one developer applying the identical logic to a side project with no budget committee in sight.
A cheap model reads the diff before the expensive one does
Sneha Tuli, Principal Product Manager at Microsoft, described the mechanism at aidevcon’s DevCon Fall 2025, in a talk published November 21, 2025. Her team had rolled out AI-assisted code review across the company: more than 90% of pull requests, touching more than 95% of developers. Getting there ran through a trust problem first. “We had to earn trust,” she said. “Initially we were kind of spamming the developers with a lot of comments, so we created a small language model based on past reviews. Any time a PR was created, the diffs were passed through this SLM, which was hosted by us. Not only did it save cost, but only important, risky diffs were now getting sent to the LLM. This reduced the noise significantly.” The same rollout produced a 10 to 20% cut in PR completion time and a 50 to 60% usefulness rate on the suggestions developers saw, per Tuli’s own numbers.
The filter does two jobs at once. Tuli reports that it saved cost, since a small model screening every diff costs less per diff than a frontier LLM. She also reports less noise for developers, because only the diffs the SLM flags as risky reach the model that writes the review comments. She gives no savings figure.

Doctolib tries the same fix for cost and carbon
Thomas Bentkowski, Product Manager at Doctolib, named cost and environmental impact together as one of five friction points his team hit while scaling AI tooling from 30 to 600 engineers, in a talk published September 24, 2025. On cost, running enough volume through an expensive frontier model billed per million tokens “could indeed create a huge bill at the end,” he said. On environmental impact, he stayed careful about the limits of what’s actually documented. “There’s no real big papers available yet on the carbon footprint emitted by those LLMs,” he said, “but we are trying to contextualize the usage of those LLMs inside the global scheme of everything.”
In September 2025, Doctolib was testing the same kind of fix. “Recently OpenAI released gpt-oss with 20 billion parameters, which is great to actually work in local,” Bentkowski said, “so we’re trying to leverage those on specific use cases and limiting the LLM usage running into data centers.” The experiment was intended to address both cost and environmental impact by handling narrow use cases locally. Bentkowski did not report a measured saving or footprint reduction in the talk.
Apply the same decision to a personal project
Nicolas Martignole’s December 2025 account of a role-playing game separated frequent narration from less frequent coding work and gave them different model budgets. The transferable decision is to identify the repeated call, its required quality and its cost. The historical model pair does not determine today’s route.
For an Opus 5.5 setup, compare a simple route against any proposed split on the tasks the application actually repeats. Include retries and the cost of maintaining another model adapter. A role with low per-call cost can dominate the total if it runs on every interaction.
Where the individual version runs out of signal
Martignole’s game engine still runs on a metered API key, the same billing model behind his $214 receipt from part one, where every model choice shows up as a line item he can inspect. A Claude Code user on a subscription doesn’t have that. Part two of this series covered the same blind spot from the spending side: a Claude Max or Pro account has no per-call invoice to read, so the arithmetic Martignole did on camera has no inputs to run on. A flat-rate plan has no invoice line to prove a cheaper model saved anything, because the flat monthly price hides the token math a metered account exposes by default. Usage limits still bind: Anthropic announced a separate credit for claude -p and Agent SDK usage in June 2026, then paused it, so that usage still draws from the subscription’s limits (Claude Help Center).
Mapping the token-reduction toolbox covers that version of the problem directly: what a token-reduction discipline looks like on Claude Max or Pro, where the monthly price is flat and the main lever left is the context window. The organizational FinOps this article describes, a dedicated role, a TCO study, a small model screening diffs, assumes an invoice exists to optimize against. The individual version has to build its own signal from usage patterns instead, with no bill to check the math against.
What FinOps still gets wrong
The earlier cases made token prices a visible trigger for cost control. The Opus 5.5 evaluation adds a more useful decision unit: the accepted task, including failed attempts and triage. Part three of this series showed the gap that number is actually supposed to be managing: a measured 15-20% productivity gain sitting next to marketed multipliers running past 10x, and a real budget has to be built against one of those figures rather than the other. That’s the gap a FinOps role exists to close, by pricing what a token costs and routing around the expensive ones.
FinOps for AI, as practiced above, assumes the price per token is the unit that measures the spend. It mostly isn’t. Some tokens cost far more than others on the very same model: the ones a verbose agent burns to reach an answer a leaner one reaches faster, and the ones it resends untouched on every turn because nothing cached them. The next part in this series takes that assumption apart.
YSNK
(You should now know)
- FinOps for AI is starting to echo the sequence SNCF Connect followed for its cloud migration: a cost study first, then a named owner for the number
- Microsoft’s AI review covers over 90% of pull requests, and a small model sends only the risky diffs to the LLM, which Sneha Tuli reports cut cost and review noise
- In September 2025, Doctolib was testing a local 20B-parameter model (gpt-oss) to address both token cost and carbon footprint, without a measured saving reported in the talk
- An Opus 5.5 baseline lets you test whether a filter or additional reviewer still earns its execution and maintenance cost
- A flat-rate Claude Max or Pro plan has no per-call invoice line to show whether any of this saved money. The organizational version of FinOps assumes a bill exists to optimize against, and the individual version doesn’t have one
Go Further in the Claude Code Guide
Practical resources selected to help you take the next step.
Related articles
Doctolib's cross-verified AI adoption data
Four conference talks and outside sources document Doctolib's AI adoption from 30 engineers to 600, without reducing it to one figure.
Measuring your own AI usage: three meters, three different answers
Twenty days of my own Claude Code and Codex usage, measured three ways. The tools disagree by a factor of three, and dollars are not what binds.
The price-per-token lie: why cheaper models don't mean cheaper bills
Opus 5.5 changes input, output and cache economics. Price the accepted task, including retries and review, before choosing a cheaper model.
Go deeper
Step-by-step guides that put this into practice.