Instrumented field report
Leaving direct Anthropic access
what Back Market measured, and what the other sources say
Two field reports from Nicolas M on leaving the direct Anthropic API, checked against the other public videos that discuss OpenRouter and against its documentation.
Documented report · Dated sources and measurements in each chapter
Key takeaways
- The fact. Between May and August 2026, the monthly spend on coding agents at Back Market falls from 87 k$ with Anthropic direct to 23 k$ with OpenRouter, for the same population of about 245 developers.
- The mechanism. The saving does not come from a negotiated discount but from a change in which models are consumed, coupled with a cap of 200 $ per developer per month introduced in the same window. The 2 levers arrive together, and no account allows crediting them separately.
- What the aggregator does not do. A single contract with OpenRouter simplifies purchasing. Routing to 10 providers leaves 10 data policies to check, and OpenRouter does not answer for what happens at Fireworks, Google Vertex or Amazon Bedrock.
- What the other sources say. Stripe announced the acquisition of OpenRouter, for more than 7 billion dollars according to Bloomberg; a team of 200 engineers left it for Vertex AI on security grounds; Chinese models pass 60 % of usage there in July 2026. Part IV lists the videos concerned.
- Where to dig deeper. The Claude Code Ultimate Guide carries these lessons into 4 pages, each tied to its question, listed in part V.
Terms with a dotted underline show their definition on hover, keyboard focus or tap.
The Back Market case
The Anthropic bill goes from 57 k$ to 87 k$ in one month
The spend nearly doubles in a month, before any technical problem.
Nicolas M opens on what he calls « là où ça fait mal », where it hurts. In May 2026, about 230 to 240 developers consume 57 k$ of Anthropic API, through Claude Code alone. In June, the same perimeter reaches 87 k$.
He speaks as the person in charge of the rollout: « Imaginez ceci, 250 développeurs pendant 1 an qui peuvent utiliser Claude Code sans aucune limite de budget avec une facturation à l'usage. » Picture 250 developers for a year, using Claude Code with no budget limit and usage-based billing. His assigned role is to deploy the tool and train people, while the spend climbs.
The decision is taken to leave direct Anthropic access. OpenRouter usage starts on 18/06/2026, and the rollout becomes general at the end of July.
Monthly spend drops from 87 k$ to 23 k$
The 2 billing columns read month by month, with no service interruption for users.
Nicolas M shows the 2 consoles side by side. The Anthropic column goes dark while the OpenRouter column rises, over the same user population.
| Month 2026 | Anthropic direct | OpenRouter | Comment |
|---|---|---|---|
| May | 57 k$ | not used yet | 230 to 240 Claude Code users |
| June | 87 k$ | start on 18/06 | The peak that triggers the decision |
| July | 25 k$ | 21 k$ | The OpenRouter amount covers 18/06 to 31/07 |
| August | 574 $ | 23 k$ | Reference month for every analysis in the video |
| September | 17 $ | not disclosed | Read around 09/09 |
Asked whether AI use had stopped, he answers: « Est-ce que vous avez arrêté d'utiliser l'IA ? Non, tous ces gens-là ont été transférés sur OpenRouter. » Did you stop using AI? No, all those people were moved to OpenRouter. The Anthropic column falls because the usage moved to OpenRouter.
DIAGRAM · Before and after the gateway
A single contract with the aggregator does not transfer liability
The video of 22/07/2026, earlier than the review, sets the legal constraint the figures do not show.
September's review is not the first episode. On 22/07/2026, Nicolas M had already explained the mechanism, and above all what it costs in legal work. That is the part spend tables never show.
His commercial argument is plain: « vous en tant qu'entreprise vous allez avoir qu'un seul contrat avec Open Router et vous n'aurez pas forcément besoin d'avoir ensuite un contrat avec tous les fournisseurs qu'il y a derrière ». As a company you will have a single contract with OpenRouter, and you will not necessarily need a contract with every provider behind it.
He draws a rule from it: « si vous routez vers 10 providers, vous avez 10 politiques de données à vérifier ». Route to 10 providers and you have 10 data policies to check. He says out loud that this concerns providers and not models, correcting a slip on his own slide.
| Item | Value | Reliability |
|---|---|---|
| Providers available at OpenRouter | « à peu près 70 aujourd'hui », about 70 today | Audible in the video |
| Providers shortlisted and authorised by Back Market | not readable | The subtitle renders the number as « H » |
| Providers named in July | Google Vertex, Fireworks, Amazon Bedrock | Named explicitly |
| Providers named in September | + Microsoft Azure, Mistral | Named explicitly |
| Providers in use in September | « une dizaine », about 10 | Audible in the video |
The same episode defines, in passing, the vocabulary used from here on: a frontier model is one « parce qu'en fait c'est un modèle très puissant et qui ne peut pas tourner aujourd'hui dans votre infrastructure ou sur un provider, et c'est surtout un modèle qui est privé », a very powerful model that cannot run today in your own infrastructure or at a provider, and above all a private model.
Anthropic was then the only host of the Claude models. The move to OpenRouter does not change that. It does not open Claude to other hosts, it opens usage to other models.
Why the bill goes down
GLM 5.2 absorbs 53 % of the tokens for 35 % of the bill
Back Market pays less because it consumes models other than Anthropic's.
In August, Nicolas M crosses 2 views of the same console, the split by amount spent and the split by token volume. A model that weighs more in tokens than in spend costs less per token than the average of the mix.
| Model | Share of tokens | Share of the bill | Amount cited |
|---|---|---|---|
| GLM 5.2 (open weight) | 53 % | 35 % | more than 8 k$ |
| Sonnet 5 | 8 % | not disclosed | barely 2 k$ |
| Opus | not disclosed | not disclosed | « a pas mal disparu du radar », largely off the radar |
| Gemini, GPT | « loin derrière », far behind | not disclosed | not disclosed |
GLM 5.2 is both the most used model and, in his words, « le modèle aussi qui va coûter le plus cher », the one that will cost the most, among those carrying the volume. Yet when the sort runs on money rather than tokens, Sonnet and Opus climb back despite their low volume.
Hence his conclusion: « les modèles d'Anthropic sont vraiment vraiment chers, même quand vous passez par OpenRouter », Anthropic's models are really expensive, even when you go through OpenRouter.
He takes care not to overread the disappearance of Opus: « le coût a été réduit pas parce qu'on s'en sert moins mais parce que simplement on a accès à plus de providers », the cost went down not because we use it less but simply because we have access to more providers. He also answers the productivity objection: « les utilisateurs n'ont pas perdu soudainement leur productivité ou leur usage », users did not suddenly lose their productivity or their usage.
Speed: faster on simple tasks, weaker on long ones
Speed, for him, is an underrated criterion: « les modèles d'OpenAI sont plus rapides. Le nombre de tokens par seconde qui sont générés par ces modèles sont plus rapides et une session de travail va prendre moins de temps parce que vous allez moins attendre. » OpenAI's models are faster, they generate more tokens per second, and a working session takes less time because you wait less. Same impression on GLM 5.2.
On complex, long tasks he observes that these models « tirent un peu la patte », drag their feet, or « font un peu n'importe quoi », go off the rails, and « restent quand même moins bons en terme d'intelligence que les modèles Anthropic », remain less capable than Anthropic's models. On simple tasks, however, « ça marche super bien », it works very well.
Back Market therefore segments usage by task difficulty.
A team whose workload is mostly complex should not expect the same factor. The AI Unit Economics guide details this routing by complexity and its risks, including cache loss when the model changes mid-task.
Public pricing confirms a 5.7 ratio between GLM 5.2 and Opus 4.8
The OpenRouter catalog as of 16/09/2026 checks the orders of magnitude cited.
No access to Back Market's console is needed. The OpenRouter catalog is public and requires no authentication, and his claim about the price of Anthropic models can be checked against it.
| Identifier | Input | Output | Output ratio vs GLM 5.2 |
|---|---|---|---|
z-ai/glm-5.3-flash | 0.09 $ | 0.30 $ | 0.07 |
z-ai/glm-5.2 | 1.40 $ | 4.40 $ | reference |
anthropic/claude-haiku-4.5 | 1.00 $ | 5.00 $ | 1.1 |
anthropic/claude-sonnet-5 | 2.00 $ | 10.00 $ | 2.3 |
anthropic/claude-opus-4.8 | 5.00 $ | 25.00 $ | 5.7 |
z-ai/glm-5.3-flash. Between 2 open weight models of the same family, the price gap exceeds the one separating GLM 5.2 from Sonnet 5.About ten providers and automatic allocation by token price
What the gateway does in place of the team, and what it does not do.
Back Market works with about ten providers, including Microsoft Azure, Amazon Bedrock, Google, Fireworks and Mistral. Nicolas M notes that Mistral also hosts GLM 5.2: « vous pouvez aller taper sur l'infrastructure Mistral », you can go hit Mistral's infrastructure, to consume a Chinese model from European infrastructure.
His definition, for anyone who confuses the 2 levels: « un provider c'est comme un hébergeur internet qui aurait plein de sites internet, et chaque site internet en fait ce sont les fameux modèles », a provider is like a web host that has lots of websites, and each website is actually one of the famous models.
DIAGRAM · What the gateway arbitrates
He describes the arbitration: « cette ventilation, elle est faite automatiquement par OpenRouter selon aussi le prix du token. Parfois c'est plus intéressant d'aller chercher chez Microsoft un accès à OpenAI que de l'avoir en direct, et parfois c'est l'inverse. », this allocation is done automatically by OpenRouter, also based on token price. Sometimes it is more advantageous to get access to OpenAI through Microsoft than to have it directly, and sometimes it is the other way around. The benefit he values is operational more than financial: « moi, en tant qu'utilisateur, j'ai pas besoin de m'ennuyer », as a user, I don't need to bother with it.
On OpenRouter's side, these constraints go through documented routing parameters. zdr: true restricts a request to endpoints with no data retention, data_collection: "deny" excludes providers that retain data. The only list restricts routing to named providers, within those allowed at the account level. This is the technical equivalent of Back Market's short list described in part I. By default, OpenRouter distributes load favoring the cheapest providers, weighted by the inverse square of their price. Source: provider routing documentation.
Making a workstation independent of a model provider takes more than a gateway URL. The article "Portability becomes a Scale concern" describes what survives a controlled migration from one provider to another.
Blended cost per million tokens falls 13 % in a month
The single indicator Nicolas M uses to steer, and what it actually proves.
His preferred metric sits in the top right of the OpenRouter console, the blended dollar per million. Between July and August, it falls 13 %.
He reads it this way: « mes utilisateurs pour un million de tokens ont payé moins cher que le mois d'avant, et je suppose que du coup ils ont utilisé des modèles moins chers. Pour mesurer l'efficience d'une entreprise et de 245-250 utilisateurs, c'est plutôt un bon indicateur. », for a million tokens, my users paid less than the previous month, and I assume they therefore used cheaper models. To measure the efficiency of a company and of 245 to 250 users, it is a fairly good indicator.
The same blind spot exists at the task level. Inferya compares 2 models that each cost 0.07 $ per attempt on SWE-bench Verified, with the mini-SWE-agent harness: one solves 75.8 % of tasks, the other 9 %. Per task solved, the first works out to 0.092 $ per task and the second to 0.778 $, that is 8.4 times more.
A similar blended cost can therefore hide very different costs per task solved.
The FinOps Foundation formalizes these indicators, from cost per token to cost per inference and per API call, in its "FinOps for AI" guide, updated on 17/02/2026. The article "Mapping the token-reduction toolbox" applies the same caution to token-reduction tools. A shorter output is not enough to prove that a completed task costs less.
Nicolas M presents this drop as an inference, not as a measurement. The verb he uses is « je suppose », I assume.
Budget control
200 $ per developer per month, with 3 outcomes when it runs out
The lever the gateway does not provide, which Nicolas M presents as co-responsible for the result.
The budget resets to zero at the end of the month. Asked whether 200 $ is enough, he answers: « Oui, les gens le prouvent aujourd'hui. », yes, people prove it today.
DIAGRAM · The path of a developer who exhausts their allowance
- Self-served 20 $ extension
Requestable from Slack at any time, « parce que c'est pas un humain qui va approuver votre demande », because it's not a human who will approve your request. It requires filling in a form with figures taken from the OpenRouter console: number of sessions, most expensive session, breakdown by model. It is granted only once: « si j'éclate ces 20 dollars, j'aurai pas le droit à 240 dollars, ça s'arrêtera là », if I blow through those 20 dollars, I won't get another 240 dollars, it stops there.
- Project exception arbitrated by committee
The manager fills in a Confluence form with an estimate of the amount needed. A committee, combining budget holders and product profiles, checks that the request « va générer du business », will generate business. Stated turnaround: about half a day. Result: a dedicated OpenRouter workspace, outside the individual allowance.
- Dedicated keys for automated processing
Non-human agents, pull request review, CI tasks, Slack bots, product features, have their own keys with an imposed model, and do not consume a developer's allowance.
The intent of the form is educational, not punitive: « c'est pour que les gens puissent comprendre un peu où est parti l'argent et qu'ils puissent se poser les bonnes questions, et que nous en fait on n'ait pas besoin de leur expliquer à chaque fois qu'est-ce que c'est qu'un modèle, qu'est-ce que c'est qu'une session. », it's so people can understand a bit where the money went and ask themselves the right questions, and so that we don't need to explain every time what a model is, what a session is. He adds that it is « super bien accepté par nos collaborateurs », very well accepted by our colleagues.
His own 190 $ session serves as an example
Rather than expose a colleague, Nicolas M names himself. He shows a session from 13 to 14 August at 190 $ on a single run. « Vous êtes motivé et vous faites ça. », you're motivated and you do that.
| Model | Cost |
|---|---|
| Opus 4.8 | 139 $ |
| Opus 5 | 8 $ |
| Sonnet 5 | 5 $ |
| Remainder | negligible |
A single model accounts for 73 % of the session. The form asks for the breakdown by model, which makes this gap visible: « je prends conscience qu'il y a des modèles qui font mal » I become aware that some models really hurt.
The guide AI Unit Economics distinguishes budget policies by who is spending. An agent, a CI task or an unsupervised service call for a strict cap, since there is nobody to read an alert. A developer in a session can receive an alert, then go through an approval.
Exceptions tie each expense to a named project
What Nicolas M values most in the setup, and what none of the previous figures show.
He cites amounts tied to projects: « on sait qu'on a mis 1100 € sur ce petit projet là qui est important, et que ce hackathon là finalement a coûté que 38 dollars et pas 500 dollars », we know we put 1,100 € into that small but important project, and that hackathon actually cost only 38 dollars, not 500 dollars.
The same reasoning applies to automated agents, with an additional constraint: « si jamais un système commence à faire n'importe quoi et à dépenser plein d'argent, grâce à OpenRouter on peut très facilement contrôler ça, mettre de l'alerting dessus », if a system ever starts acting erratically and spending a lot of money, thanks to OpenRouter we can very easily control that, put alerting on it.
His conclusion deliberately widens the subject: « je sais que je parle beaucoup d'OpenRouter, mais au-delà d'OpenRouter, ce que je voudrais vous montrer c'est comment on fait pour que la valeur que génère l'intelligence artificielle, on s'assure que cette valeur est là » I know I talk a lot about OpenRouter, but beyond OpenRouter, what I want to show you is how we make sure the value generated by artificial intelligence is actually there. He adds a non-financial argument: « quand on voit des tokens et la finance, faut réfléchir que ça a un certain coût en énergie », when we look at tokens and finance, we need to remember that it has a certain energy cost.
What the other sources say
75 public videos mention OpenRouter, 12 bring a new fact
Usage reports, market reading and practical implementation, everything that complements, qualifies or contradicts the Back Market case.
Between 2025 and September 2026, 75 public videos mention OpenRouter, across French- and English-speaking technical channels. Most cite it in passing; 12 bring a fact, a figure or a firsthand account.
For an overview of the market, OpenRouter and a16z published on 04/12/2025 a « State of AI » report based on more than 100,000 billion tokens routed in one year. It finds that the adoption of open weight models has accelerated and that they now serve as building blocks for production applications.
The 12 selected videos
| Angle | Date | Channel | Video | What it contributes |
|---|---|---|---|---|
| Usage reports | 14/09/2026 | Nicolas Martignole | Bilan de 2 mois avec OpenRouter, pour 245 utilisateurs : les vrais chiffres | Assessment of the switch: 87 k$ per month at Anthropic in June, 23 k$ at OpenRouter in August, for about 245 developers, with the associated budget mechanism. |
| Usage reports | 22/07/2026 | Nicolas Martignole | Claude Max (abo prix fixe) ou API (pay-as-you-go): ce n'est pas une question de prix | About 70 providers available, a short list authorised by Back Market, and a liability that stays with the client. |
| Usage reports | 11/07/2026 | Nicolas Martignole | Le vrai coût de l'IA : pourquoi le prix au token est faux | According to a presentation by OpenRouter's COO, Chinese models there exceed American models in volume; Nicolas M concludes from this that price per token is misleading. |
| Usage reports | 30/06/2026 | Nicolas Martignole | Premiers benchmarks avec Sonnet 5, et comparo avec GLM 5.2 | Evaluation of OpenRouter to get a single contract, and providers tested: Google Vertex, Amazon Bedrock, Fireworks. |
| Usage reports | 29/05/2026 | AI Native Dev | How We Built an AI Code Reviewer for 200 Engineers | A team of 200 engineers starts on OpenRouter (a single API, same prices as the model publishers) then moves to Vertex AI: some models did not meet its security rules, and Vertex gives it more control over logs and budgets. |
| Market | 17/08/2026 | Bloomberg | Anthropic's Revenue Jump, The Wealthy Bet on SpaceX · Bloomberg Tech | Stripe would be acquiring OpenRouter for more than 7 billion dollars, according to Bloomberg's sources. |
| Market | 22/08/2026 | Bloomberg | Chinese AI Models Gain Ground on US AI in Price and Use | OpenRouter's usage data show more than 60 % Chinese models in July 2026. Explicit caveat: OpenRouter is only a platform, and direct usage of the Anthropic and OpenAI APIs is not visible there. |
| Market | 06/08/2026 | AI Engineer | The State of Model Routing — NVIDIA, Cognition, OpenRouter | Panel discussion with a representative from OpenRouter, NVIDIA and Cognition on routing: sending each task to the smallest model that suffices, and orchestrating several models together. |
| Market | 23/04/2026 | The Product Crew | Tech @ France : Comment l’État français déploie sa propre stack IA | The French State presents OpenGateLLM, its open source inference gateway, as an alternative to LiteLLM and OpenRouter. |
| Practical implementation | 04/06/2026 | Cole Medin | Claude Plans, Gemini Designs: The Workflow to Build BEAUTIFUL Frontends | Claude plans, Gemini 3.5 Flash designs the interface: access to Gemini goes through OpenRouter from the Pi agent. |
| Practical implementation | 06/04/2026 | The Next New Thing | Best & cheapest AI for OpenClaw | Configuring an agent with cheaper models via OpenRouter: a few dollars of credit, then comparing models down to the choice of GLM 5.1. |
| Practical implementation | 15/06/2025 | Alex so yes | Tutoriel Complet : Coder une app entière avec Cline (OpenSource) | OpenRouter as a single interface that hides API format differences between model publishers, with its ranking of the most used models. |
The 75 videos that mention OpenRouter
What the Claude Code guide keeps
4 pages of the Claude Code Ultimate Guide incorporate these lessons
Each lesson joins the guide page that already addresses its question.
The Claude Code Ultimate Guide already covered the question of gateways, but scattered across pages, and with a contradiction between 2 of them. One advised against pointing Claude Code at another backend in production, the other recommended exactly this mechanism.
Each lesson from this document is incorporated page by page, without repeating the same figure in 2 places.
| Question | Guide page | What it takes from the Back Market case |
|---|---|---|
| How to cap and attribute developer spend? | API Gateway, Model Allowlists | The 200 $ per month cap with self-served extension; about 70 providers available and around 10 used; the rule of 10 providers, 10 data policies; the liability that stays with the client |
| Named-user licenses, service identities, or both? | Subscription Strategy, section 6 | The move from 87 k$ to 23 k$ per month, with the caveat that 2 levers acted together |
| Is pointing Claude Code elsewhere legitimate? | Pointing Claude Code at Another Backend | The distinction between a contracted commercial gateway and a reverse-engineered proxy, the warning remaining intact for the second |
| What hardware, and how much does a served token cost? | Local vs Cloud Inference | A reference table pointing to the page that addresses each question |
2 other pages of the guide extend the topic without revisiting the Back Market case: AI Unit Economics, on cost per accepted task, and Team Metrics, on the metrics of teams working with agents.
Reproducing the measurement
Comparing 2 models on real tasks, via the OpenRouter API
A minimal protocol tests model substitution before touching a team's configuration.
Model substitution is the first thing to test. It needs neither a team gateway nor a spending cap. The test runs on a small, fixed batch of real tasks, 5 to 10 requests representative of everyday work, sent identically to 2 models.
The most direct test connects Claude Code itself to OpenRouter. The official integration guide fits in 3 environment variables, with no local proxy, to place in the shell profile or in the project's .claude/settings.local.json file.
Terminal Connecting Claude Code to OpenRouter 4 lines
export OPENROUTER_API_KEY="sk-or-..."
export ANTHROPIC_BASE_URL="https://openrouter.ai/api"
export ANTHROPIC_AUTH_TOKEN="$OPENROUTER_API_KEY"
export ANTHROPIC_API_KEY="" # doit être vide, et non absenteIf Claude Code was already connected to an Anthropic account, run /logout once and restart. The /status command should then display https://openrouter.ai/api on the base URL line. This URL takes no /v1 suffix, reserved for the OpenAI-compatible API, which triggers model-not-found errors.
For a controlled comparison outside Claude Code, the first step is to list the available identifiers rather than copy a name seen in a video. The catalogue is public and requires no authentication.
Terminal Listing models and their pricing, sorted by output price 10 lines
curl -s https://openrouter.ai/api/v1/models -o or-models.json
python3 - or-models.json <<'PY'
import json, sys
rows = json.load(open(sys.argv[1]))["data"]
for m in sorted(rows, key=lambda x: float(x["pricing"]["completion"])):
p = m["pricing"]
print("%-44s in=%7.2f $/M out=%8.2f $/M" % (
m["id"], float(p["prompt"]) * 1e6, float(p["completion"]) * 1e6))
PYThe response contains data, total_count and links; each entry carries id, pricing.prompt, pricing.completion and context_length. As of 16/09/2026, the catalogue holds 443 entries. Identifiers prefixed with a ~ are aliases that follow a family's latest version: convenient for exploring, best avoided for comparisons, since their target can change between 2 passes.
Second step, send the same request to both models. The following command reads the request from prompt.txt and saves each response to a separate file, from any working directory.
Terminal Same request, 2 models (requires curl and jq) 10 lines
export OPENROUTER_API_KEY='sk-or-...'
for model in anthropic/claude-sonnet-5 z-ai/glm-5.2; do
jq -n --arg m "$model" --rawfile p prompt.txt \
'{model: $m, messages: [{role: "user", content: $p}]}' \
| curl -s https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d @- > "reponse-${model//\//_}.json"
done/api/v1/chat/completions path come from OpenRouter's documentation, consulted on 16/09/2026. The completion call itself was not executed for this document, for lack of a funded key. Verify the structure of a first response before automating the reading of subsequent ones.Third step, compare on verifiable criteria rather than impression: is the response usable without rework, does the code compile, do the tests pass. Then relate each pass's token consumption to the price noted in the first step.
Without this count, there is no way to know whether the cheaper model stays cheaper once rework is counted. The AI Unit Economics guide explains how to build this pairwise comparison, and what a small sample allows one to conclude.
Reproducing the per-person cap with a credit-limited key
The 200 $ mechanism is a documented OpenRouter feature, testable alone.
Back Market's individual allocation rests on a documented capability of the management API. A key can carry its own credit limit.
| Need | Element | Verified |
|---|---|---|
| Entry point | https://openrouter.ai/api/v1 | yes, Quickstart |
| Authentication | Authorization: Bearer <key> header | yes, Quickstart |
| Completion call | POST /api/v1/chat/completions | yes, Quickstart |
| Catalogue and pricing | GET /api/v1/models | yes, called live |
| Per-key cap | POST /api/v1/keys, limit field | yes, Provisioning doc |
| Cap reset | Daily, weekly or monthly, per key | yes, Enterprise page |
On its Enterprise page, OpenRouter documents per-key credit caps with automatic reset, daily, weekly or monthly. One key per developer with a monthly reset therefore reproduces Back Market's 200 $ per month allocation, without in-house accounting.
POST /api/v1/keys was not called for this document. This operation creates a billable resource and requires a provisioning key. The existence of the limit field (« Optional credit limit ») comes from the documentation, not from an execution. The exact behavior when the cap is reached, error code and message returned to the client, remains to be observed before building a process on it.For an individual test, reproducing the 3-question form depends on no API. After a week of use, answer for yourself: how many sessions, which session cost the most, and what is the split by model. Nicolas M credits it with a pedagogical effect, and it can be tested without a budget.
Minimal protocol over 2 weeks
- Week 1, measure without changing anything
Switch everyday usage to a dedicated OpenRouter key, without a cap and without changing model. The goal is to get a baseline and a starting blended cost.
- End of week 1, answer the form
The 3 questions, on one's own console. Write down the answer before looking at the following week. That is the measure of the pedagogical effect, and it disappears if reconstructed after the fact.
- Week 2, segment by difficulty
Explicitly route simple tasks to a cheap model and keep the high-end model for complex tasks, following the cutoff Nicolas M describes. This week tests segmentation.
- Compare the 2 blended costs
The gap between the 2 weeks is the individual equivalent of the 13 % observed at Back Market. Even at very different volumes from one week to the next, the indicator stays comparable since it is expressed per million tokens, but the total bill is not.
What this experience report does not demonstrate
Applying the observed factor elsewhere takes 4 precautions.
The population figures vary within the video itself: the title announces 245 users, the introduction speaks of 250, the first Anthropic screen of 230 to 240. Nicolas M uses them interchangeably to refer to the same population. This document keeps his wording rather than harmonizing a figure he never settled on.
Glossary
Terms used in this report.
- Blended cost
- Average price paid per million tokens over a period, across all models. OpenRouter displays it on its console. At constant volume, it falls as users shift toward cheaper models.
- Open weight
- Model whose weights are published, so it can be hosted by any provider. Distinct from a proprietary model, accessible only at its publisher.
- Provider
- Host that serves a model. The same open-weight model can be served by several providers, at different prices.
- Zero data retention
- Commitment by a provider to retain no request data. Selection criterion cited to rule out otherwise well-ranked models. Abbreviated ZDR.
Other glossary terms
- BYOK
- Bring Your Own Key: using one's own keys or credits at a model provider through a gateway. OpenRouter then applies a separate fee schedule.
- Claude Code
- Command-line development agent published by Anthropic. It consumes either the Anthropic API directly or a compatible gateway.
- GLM
- Open-weight model family published by Z.ai. GLM 5.2 is the most consumed model in the experience report described here.
- Frontier model
- Proprietary model, too heavy to run at a third party or on internal infrastructure, served only by its publisher. Claude models belong to this category.
- OpenRouter
- Commercial gateway that exposes several hundred models behind a single API and a single bill, and splits requests across providers.
- Gateway
- Intermediary service between a client and several model providers. It centralizes authentication, billing and routing. Known in French as passerelle.
- Token
- Unit of text splitting billed by providers. Prices are expressed per million tokens, separately for input and output.
- Workspace
- Isolated billing space within OpenRouter. Used here to carve a project budget out of a developer's individual allocation.
Conclusion, sources and updates
Report scope and method
This document starts from 2 videos by Nicolas M, published on 22/07/2026 and 14/09/2026, which describe the migration of about 245 developers from direct Anthropic access to the OpenRouter gateway, then the budget mechanism that came with it. It checks them against other public videos that discuss OpenRouter, against OpenRouter's documentation and pricing, and points to the pages of the Claude Code Ultimate Guide that cover each question. Quotations come from the subtitles, mostly automatic, which garble proper nouns. Model names were restored when they were certain, and flagged when they were not. OpenRouter prices were read from the public API on 16/09/2026. Quotations stay in French, their source language, with their meaning in English right after, so every figure remains checkable against the video. This document measures neither the quality of the code produced nor the time developers spent. It covers a spend, and the traceability of that spend.
Conclusion
Moving to a gateway and capping spend per person are 2 separate decisions, which Nicolas M presents as joint. The gateway opens access to other models and breaks the spend down by model; the cap limits what each developer commits. The video does not allow saying which share of the drop belongs to each.
The cheapest transferable mechanism is the 3-question form required before the 20 $ extension. It forces the developer to read their own console before spending more. It needs neither a gateway nor a budget, and can be tested alone, on one's own usage.
Dated sources
- Bilan de 2 mois avec OpenRouter, pour 245 utilisateurs : les vrais chiffres, Nicolas Martignole · 14 September 2026 · Source documentaire
- OpenRouter, Quickstart: base URL, authentication header and endpoints · 16 September 2026 · Observation vérifiée
- OpenRouter, Provisioning API Keys: creating keys with a credit cap · 16 September 2026 · Observation vérifiée
- OpenRouter, public catalogue of models and pricing (443 models at the time of reading) · 16 September 2026 · Observation vérifiée
- Claude Max (abo prix fixe) ou API (pay-as-you-go), ce n'est pas une question de prix, Nicolas Martignole · 22 July 2026 · Source documentaire
- The State of Model Routing, NVIDIA, Cognition, OpenRouter, AI Engineer · 6 August 2026 · Source documentaire
- How We Built an AI Code Reviewer for 200 Engineers, AI Native Dev · 29 May 2026 · Source documentaire
- Anthropic's Revenue Jump, The Wealthy Bet on SpaceX · Bloomberg Tech, Bloomberg · 17 August 2026 · Source documentaire
- Chinese AI Models Gain Ground on US AI in Price and Use, Bloomberg · 22 August 2026 · Source documentaire
- Tech @ France : Comment l’État français déploie sa propre stack IA, The Product Crew · 23 April 2026 · Source documentaire
- Le vrai coût de l'IA : pourquoi le prix au token est faux, Nicolas Martignole · 11 July 2026 · Source documentaire
- Claude Code Ultimate Guide, API Gateway · 16 September 2026 · Source documentaire
- Stripe, announcement of the agreement to acquire OpenRouter · 19 August 2026 · Observation vérifiée
- OpenRouter, provider routing: zdr, data_collection, only and ignore · 17 September 2026 · Observation vérifiée
- OpenRouter, Enterprise offering: per-key caps, ZDR, regional lock-in · 17 September 2026 · Observation vérifiée
- OpenRouter, pricing: platform fees and BYOK · 17 September 2026 · Observation vérifiée
- OpenRouter, Claude Code integration guide · 17 September 2026 · Observation vérifiée
- OpenRouter and a16z, The 2025 State of AI Report · 4 December 2025 · Observation vérifiée
- Superconductor, Kimi K3 benchmark on an internal SWE-bench · 3 August 2026 · Observation vérifiée
- Inferya, cost per solved task of agent benchmarks · 21 August 2026 · Observation vérifiée
- FinOps Foundation, FinOps for AI Overview · 17 February 2026 · Observation vérifiée
“Accessed on” gives the date the source was read. A source may have been published earlier and later updated.
Update history
| Date | Version | Changes |
|---|---|---|
| 18 September 2026 | 2.8.0 | Rhythm and repetition pass on both languages: 25 announce colons turned into sentences, 7 paragraphs opening on « On » given back their subject, 5 chapter decks returned to a finite verb, 5 paragraphs split to create short breaks, and the English glosses that follow a French quotation set off by a comma throughout. Repetitions removed: the monthly comparison amounts replayed in a callout, 2 quotations repeated word for word from one section to the next, the part II trailer for a session detailed in part III, and glossary definitions repeated in the body. |
| 17 September 2026 | 2.7.0 | English edition. Quotations stay in French, their source language, with their meaning in English right after, so every figure remains checkable against the video. Video and source titles keep their published wording. |
| 17 September 2026 | 2.6.0 | Additions verified against primary sources: official announcement of the Stripe acquisition, OpenRouter routing parameters and per-key caps, platform fees, connecting Claude Code to OpenRouter and its official reservation, Superconductor's and Inferya's quality and cost-per-task measurements, the FinOps for AI guide. Links to the guide's AI Unit Economics and Team Metrics pages and to 2 published blog articles. Dates in DD/MM/YYYY format. |
| 17 September 2026 | 2.5.0 | Amounts in k$ notation, links to the companies and products cited at their first mention, link to the speaker's profile, highlighted quotations, copy button on commands and an "Other resources" menu. |
| 16 September 2026 | 2.4.0 | Syntax highlighting for commands and links to the portfolio, the Claude Code Ultimate Guide and GitHub in the menu. |
| 16 September 2026 | 2.3.0 | Full editorial pass: removal of repeated opposition phrasings, signposting subheadings and sentencious closing lines; titles rewritten to state the fact; infographic captions replaced with their source and date. 4 claims that overreached the source were corrected, including the idea that neither of the 2 levers would suffice alone, which the video does not establish. |
| 16 September 2026 | 2.2.0 | Addition of 5 infographics: contractual liability, tokens-versus-bill gap, drop in the monthly bill, scope of the blended cost, 3-question form. Each image rests on a fact already established in the document; its text and proportions were checked one by one. Wording corrections: a leftover imperative and spoken remarks presented as written. |
| 16 September 2026 | 2.1.0 | Document made readable by any reader: full list of videos mentioning OpenRouter and an annotated selection of 12 of them, cross-references to the guide's public pages, test protocol rewritten around the OpenRouter API alone. Removal of elements specific to the author's own environment. |
| 16 September 2026 | 2.0.0 | Scope widened: the video of 22/07/2026 and the liability constraint it raises, comparison with other public videos, cross-references to the Claude Code Ultimate Guide. The count of authorized providers remains an open reservation, not an established figure. |