Guide contents
Other resources

Instrumented field report

Leaving direct Anthropic access
what Back Market measured, and what the other sources say

Two field reports from Nicolas M on leaving the direct Anthropic API, checked against the other public videos that discuss OpenRouter and against its documentation.

Documented report · Dated sources and measurements in each chapter

Key takeaways

Open the glossary

Terms with a dotted underline show their definition on hover, keyboard focus or tap.

I

The Back Market case

01

The Anthropic bill goes from 57 k$ to 87 k$ in one month

The spend nearly doubles in a month, before any technical problem.

Nicolas M opens on what he calls « là où ça fait mal », where it hurts. In May 2026, about 230 to 240 developers consume 57 k$ of Anthropic API, through Claude Code alone. In June, the same perimeter reaches 87 k$.

He speaks as the person in charge of the rollout: « Imaginez ceci, 250 développeurs pendant 1 an qui peuvent utiliser Claude Code sans aucune limite de budget avec une facturation à l'usage. » Picture 250 developers for a year, using Claude Code with no budget limit and usage-based billing. His assigned role is to deploy the tool and train people, while the spend climbs.

More tokens does not mean more value « On se rend compte rapidement que le fait de dépenser beaucoup de tokens et donc beaucoup d'argent n'a pas forcément un rapport direct avec la valeur que l'on génère. » Spending many tokens, and so much money, bears no direct relation to the value generated. Everything described next tries to tie the spend back to a result.

The decision is taken to leave direct Anthropic access. OpenRouter usage starts on 18/06/2026, and the rollout becomes general at the end of July.

02

Monthly spend drops from 87 k$ to 23 k$

The 2 billing columns read month by month, with no service interruption for users.

Nicolas M shows the 2 consoles side by side. The Anthropic column goes dark while the OpenRouter column rises, over the same user population.

Figures read on screen in Nicolas M's video of 14/09/2026.
Monthly spend per channel, as read on screen in the video
Month 2026Anthropic directOpenRouterComment
May57 k$not used yet230 to 240 Claude Code users
June87 k$start on 18/06The peak that triggers the decision
July25 k$21 k$The OpenRouter amount covers 18/06 to 31/07
August574 $23 k$Reference month for every analysis in the video
September17 $not disclosedRead around 09/09
A comparison to read with care The 21 k$ figure covers 6 weeks, from 18/06 to 31/07, not a calendar month. The only clean month against month comparison sets June with Anthropic against August with OpenRouter, and that is the one Nicolas M keeps.

Asked whether AI use had stopped, he answers: « Est-ce que vous avez arrêté d'utiliser l'IA ? Non, tous ces gens-là ont été transférés sur OpenRouter. » Did you stop using AI? No, all those people were moved to OpenRouter. The Anthropic column falls because the usage moved to OpenRouter.

DIAGRAM · Before and after the gateway
The same workstation, 2 supply paths. The gateway does not add a model, it opens a catalogue.
03

A single contract with the aggregator does not transfer liability

The video of 22/07/2026, earlier than the review, sets the legal constraint the figures do not show.

September's review is not the first episode. On 22/07/2026, Nicolas M had already explained the mechanism, and above all what it costs in legal work. That is the part spend tables never show.

His commercial argument is plain: « vous en tant qu'entreprise vous allez avoir qu'un seul contrat avec Open Router et vous n'aurez pas forcément besoin d'avoir ensuite un contrat avec tous les fournisseurs qu'il y a derrière ». As a company you will have a single contract with OpenRouter, and you will not necessarily need a contract with every provider behind it.

OpenRouter does not answer for what the providers do « Il faut faire très attention juridiquement. En fait, il y a une chaîne de responsabilité et Open Router ne va pas porter pour vous ce qui se passe chez Fireworks, chez Google Vertex ou chez Amazon Bedrock. Il y a une histoire que chacun reste chez soi. » Be very careful legally. There is a liability chain, and OpenRouter will not carry for you what happens at Fireworks, Google Vertex or Amazon Bedrock. Each provider stays responsible for its own processing.
Diagram built from Nicolas M's video of 22/07/2026.

He draws a rule from it: « si vous routez vers 10 providers, vous avez 10 politiques de données à vérifier ». Route to 10 providers and you have 10 data policies to check. He says out loud that this concerns providers and not models, correcting a slip on his own slide.

What the video of 22/07/2026 establishes about the catalogue
ItemValueReliability
Providers available at OpenRouter« à peu près 70 aujourd'hui », about 70 todayAudible in the video
Providers shortlisted and authorised by Back Marketnot readableThe subtitle renders the number as « H »
Providers named in JulyGoogle Vertex, Fireworks, Amazon BedrockNamed explicitly
Providers named in September+ Microsoft Azure, MistralNamed explicitly
Providers in use in September« une dizaine », about 10Audible in the video
Why the exact count stays open July speaks of providers « short listés et autorisés », shortlisted and authorised; September speaks of those « avec lesquels on travaille », the ones they work with. These are not the same question. The authorised list is a superset of the list in use. The September figure narrows the order of magnitude to about 10 out of 70, it does not give the size of the authorisation.

The same episode defines, in passing, the vocabulary used from here on: a frontier model is one « parce qu'en fait c'est un modèle très puissant et qui ne peut pas tourner aujourd'hui dans votre infrastructure ou sur un provider, et c'est surtout un modèle qui est privé », a very powerful model that cannot run today in your own infrastructure or at a provider, and above all a private model.

Anthropic was then the only host of the Claude models. The move to OpenRouter does not change that. It does not open Claude to other hosts, it opens usage to other models.

II

Why the bill goes down

04

GLM 5.2 absorbs 53 % of the tokens for 35 % of the bill

Back Market pays less because it consumes models other than Anthropic's.

The saving comes from the models consumed « On paye pas moins cher parce qu'on a un prix plus intéressant chez Anthropic. On paye moins cher parce qu'en fait on utilise plus uniquement les modèles Anthropic, on utilise d'autres modèles. » We do not pay less because we got a better price at Anthropic. We pay less because we no longer use only Anthropic models, we use other models. On his account, the price of Claude models does not drop by going through OpenRouter.

In August, Nicolas M crosses 2 views of the same console, the split by amount spent and the split by token volume. A model that weighs more in tokens than in spend costs less per token than the average of the mix.

August 2026, split by model as announced in the video
ModelShare of tokensShare of the billAmount cited
GLM 5.2 (open weight)53 %35 %more than 8 k$
Sonnet 58 %not disclosedbarely 2 k$
Opusnot disclosednot disclosed« a pas mal disparu du radar », largely off the radar
Gemini, GPT« loin derrière », far behindnot disclosednot disclosed
Split announced by Nicolas M for August 2026, video of 14/09/2026.

GLM 5.2 is both the most used model and, in his words, « le modèle aussi qui va coûter le plus cher », the one that will cost the most, among those carrying the volume. Yet when the sort runs on money rather than tokens, Sonnet and Opus climb back despite their low volume.

Hence his conclusion: « les modèles d'Anthropic sont vraiment vraiment chers, même quand vous passez par OpenRouter », Anthropic's models are really expensive, even when you go through OpenRouter.

He takes care not to overread the disappearance of Opus: « le coût a été réduit pas parce qu'on s'en sert moins mais parce que simplement on a accès à plus de providers », the cost went down not because we use it less but simply because we have access to more providers. He also answers the productivity objection: « les utilisateurs n'ont pas perdu soudainement leur productivité ou leur usage », users did not suddenly lose their productivity or their usage.

Speed: faster on simple tasks, weaker on long ones

Speed, for him, is an underrated criterion: « les modèles d'OpenAI sont plus rapides. Le nombre de tokens par seconde qui sont générés par ces modèles sont plus rapides et une session de travail va prendre moins de temps parce que vous allez moins attendre. » OpenAI's models are faster, they generate more tokens per second, and a working session takes less time because you wait less. Same impression on GLM 5.2.

On complex, long tasks he observes that these models « tirent un peu la patte », drag their feet, or « font un peu n'importe quoi », go off the rails, and « restent quand même moins bons en terme d'intelligence que les modèles Anthropic », remain less capable than Anthropic's models. On simple tasks, however, « ça marche super bien », it works very well.

Back Market therefore segments usage by task difficulty.

A team whose workload is mostly complex should not expect the same factor. The AI Unit Economics guide details this routing by complexity and its risks, including cache loss when the model changes mid-task.

05

Public pricing confirms a 5.7 ratio between GLM 5.2 and Opus 4.8

The OpenRouter catalog as of 16/09/2026 checks the orders of magnitude cited.

No access to Back Market's console is needed. The OpenRouter catalog is public and requires no authentication, and his claim about the price of Anthropic models can be checked against it.

Pricing recorded at https://openrouter.ai/api/v1/models on 16/09/2026, in dollars per million tokens
IdentifierInputOutputOutput ratio vs GLM 5.2
z-ai/glm-5.3-flash0.09 $0.30 $0.07
z-ai/glm-5.21.40 $4.40 $reference
anthropic/claude-haiku-4.51.00 $5.00 $1.1
anthropic/claude-sonnet-52.00 $10.00 $2.3
anthropic/claude-opus-4.85.00 $25.00 $5.7
What the 5.7 ratio changes An Opus 4.8 output token costs 5.7 times a GLM 5.2 output token. Shifting volume from one to the other brings the bill down, with no discount and no renegotiation.
GLM 5.2 costs 15 times the price of GLM 5.3 Flash A million output tokens is billed at 4.40 $, against 0.30 $ on z-ai/glm-5.3-flash. Between 2 open weight models of the same family, the price gap exceeds the one separating GLM 5.2 from Sonnet 5.
The catalog price does not include platform fees OpenRouter charges 5.5 % on pay-as-you-go and 8 % on the Business plan; the Enterprise plan opens up discounts. With your own provider keys (BYOK), the first 25 k$ of monthly consumption at catalog price is fee-free (200 k$ on Enterprise), then 5 % beyond that. The video does not specify which plan Back Market subscribes to, so not the rate applied to its 23 k$ monthly spend. Source: OpenRouter pricing, recorded on 17/09/2026.
06

About ten providers and automatic allocation by token price

What the gateway does in place of the team, and what it does not do.

Back Market works with about ten providers, including Microsoft Azure, Amazon Bedrock, Google, Fireworks and Mistral. Nicolas M notes that Mistral also hosts GLM 5.2: « vous pouvez aller taper sur l'infrastructure Mistral », you can go hit Mistral's infrastructure, to consume a Chinese model from European infrastructure.

His definition, for anyone who confuses the 2 levels: « un provider c'est comme un hébergeur internet qui aurait plein de sites internet, et chaque site internet en fait ce sont les fameux modèles », a provider is like a web host that has lots of websites, and each website is actually one of the famous models.

DIAGRAM · What the gateway arbitrates
The developer requests a model. OpenRouter chooses who runs it. The only allocation criterion cited in the video is token price.

He describes the arbitration: « cette ventilation, elle est faite automatiquement par OpenRouter selon aussi le prix du token. Parfois c'est plus intéressant d'aller chercher chez Microsoft un accès à OpenAI que de l'avoir en direct, et parfois c'est l'inverse. », this allocation is done automatically by OpenRouter, also based on token price. Sometimes it is more advantageous to get access to OpenAI through Microsoft than to have it directly, and sometimes it is the other way around. The benefit he values is operational more than financial: « moi, en tant qu'utilisateur, j'ai pas besoin de m'ennuyer », as a user, I don't need to bother with it.

The criterion that rules out well-ranked models Tencent models are climbing OpenRouter's rankings but remain unusable at Back Market: they « ne sont pas encore disponibles sur des providers en Europe en zero data retention », are not yet available on European providers with zero data retention. The ZDR constraint filters the catalog before price does. A team without a ZDR requirement has access to more models.

On OpenRouter's side, these constraints go through documented routing parameters. zdr: true restricts a request to endpoints with no data retention, data_collection: "deny" excludes providers that retain data. The only list restricts routing to named providers, within those allowed at the account level. This is the technical equivalent of Back Market's short list described in part I. By default, OpenRouter distributes load favoring the cheapest providers, weighted by the inverse square of their price. Source: provider routing documentation.

Making a workstation independent of a model provider takes more than a gateway URL. The article "Portability becomes a Scale concern" describes what survives a controlled migration from one provider to another.

07

Blended cost per million tokens falls 13 % in a month

The single indicator Nicolas M uses to steer, and what it actually proves.

His preferred metric sits in the top right of the OpenRouter console, the blended dollar per million. Between July and August, it falls 13 %.

Indicator displayed by the OpenRouter console, cited by Nicolas M in the video from 14/09/2026.

He reads it this way: « mes utilisateurs pour un million de tokens ont payé moins cher que le mois d'avant, et je suppose que du coup ils ont utilisé des modèles moins chers. Pour mesurer l'efficience d'une entreprise et de 245-250 utilisateurs, c'est plutôt un bon indicateur. », for a million tokens, my users paid less than the previous month, and I assume they therefore used cheaper models. To measure the efficiency of a company and of 245 to 250 users, it is a fairly good indicator.

Blended cost measures neither volume, nor the bill, nor quality It measures a shift in the model mix, at a given token volume. The blended cost can therefore fall 13 % while spend rises, if volume grows faster. The quality obtained enters nowhere in that calculation.

The same blind spot exists at the task level. Inferya compares 2 models that each cost 0.07 $ per attempt on SWE-bench Verified, with the mini-SWE-agent harness: one solves 75.8 % of tasks, the other 9 %. Per task solved, the first works out to 0.092 $ per task and the second to 0.778 $, that is 8.4 times more.

A similar blended cost can therefore hide very different costs per task solved.

The FinOps Foundation formalizes these indicators, from cost per token to cost per inference and per API call, in its "FinOps for AI" guide, updated on 17/02/2026. The article "Mapping the token-reduction toolbox" applies the same caution to token-reduction tools. A shorter output is not enough to prove that a completed task costs less.

Nicolas M presents this drop as an inference, not as a measurement. The verb he uses is « je suppose », I assume.

III

Budget control

08

200 $ per developer per month, with 3 outcomes when it runs out

The lever the gateway does not provide, which Nicolas M presents as co-responsible for the result.

Nicolas M attributes the drop to 2 levers « Vous avez compris, on est passé de 87 000 dollars à 22-25 000 dollars par mois. Et la question, c'est comment on a fait ? Oui, on a basculé OpenRouter. Oui, je vous ai montré que les gens utilisent des modèles open weight comme GLM 5.2. Mais il y a aussi une chose qui est importante, c'est qu'aujourd'hui on a mis un budget de 200 dollars par développeur et par mois. » You understood, we went from 87,000 dollars to 22 to 25,000 dollars per month. And the question is, how did we do it? Yes, we switched to OpenRouter. Yes, I showed you that people use open weight models like GLM 5.2. But there is also something important, which is that today we have set a budget of 200 dollars per developer per month.

The budget resets to zero at the end of the month. Asked whether 200 $ is enough, he answers: « Oui, les gens le prouvent aujourd'hui. », yes, people prove it today.

DIAGRAM · The path of a developer who exhausts their allowance
Three outcomes, only one of which is automatic. The form is not an approval filter: there is no human approver on the 20 $.
  1. Self-served 20 $ extension

    Requestable from Slack at any time, « parce que c'est pas un humain qui va approuver votre demande », because it's not a human who will approve your request. It requires filling in a form with figures taken from the OpenRouter console: number of sessions, most expensive session, breakdown by model. It is granted only once: « si j'éclate ces 20 dollars, j'aurai pas le droit à 240 dollars, ça s'arrêtera là », if I blow through those 20 dollars, I won't get another 240 dollars, it stops there.

  2. Project exception arbitrated by committee

    The manager fills in a Confluence form with an estimate of the amount needed. A committee, combining budget holders and product profiles, checks that the request « va générer du business », will generate business. Stated turnaround: about half a day. Result: a dedicated OpenRouter workspace, outside the individual allowance.

  3. Dedicated keys for automated processing

    Non-human agents, pull request review, CI tasks, Slack bots, product features, have their own keys with an imposed model, and do not consume a developer's allowance.

The intent of the form is educational, not punitive: « c'est pour que les gens puissent comprendre un peu où est parti l'argent et qu'ils puissent se poser les bonnes questions, et que nous en fait on n'ait pas besoin de leur expliquer à chaque fois qu'est-ce que c'est qu'un modèle, qu'est-ce que c'est qu'une session. », it's so people can understand a bit where the money went and ask themselves the right questions, and so that we don't need to explain every time what a model is, what a session is. He adds that it is « super bien accepté par nos collaborateurs », very well accepted by our colleagues.

His own 190 $ session serves as an example

Rather than expose a colleague, Nicolas M names himself. He shows a session from 13 to 14 August at 190 $ on a single run. « Vous êtes motivé et vous faites ça. », you're motivated and you do that.

Breakdown of the 190 $ session, shown on screen
ModelCost
Opus 4.8139 $
Opus 58 $
Sonnet 55 $
Remaindernegligible

A single model accounts for 73 % of the session. The form asks for the breakdown by model, which makes this gap visible: « je prends conscience qu'il y a des modèles qui font mal » I become aware that some models really hurt.

About ten exceptions granted since early August On the exception process: « personne n'est arrivé en disant : moi, je veux 30 000 dollars pour tous mes développeurs. C'est pas comme ça que ça fonctionne. », nobody showed up saying, I want 30,000 dollars for all my developers. That's not how it works. About ten exceptions granted since early August, ranging from a 3 to 4 day hackathon to the on-call week for the monolith release.

The guide AI Unit Economics distinguishes budget policies by who is spending. An agent, a CI task or an unsupervised service call for a strict cap, since there is nobody to read an alert. A developer in a session can receive an alert, then go through an approval.

09

Exceptions tie each expense to a named project

What Nicolas M values most in the setup, and what none of the previous figures show.

He cites amounts tied to projects: « on sait qu'on a mis 1100 € sur ce petit projet là qui est important, et que ce hackathon là finalement a coûté que 38 dollars et pas 500 dollars », we know we put 1,100 € into that small but important project, and that hackathon actually cost only 38 dollars, not 500 dollars.

Tying each expense to a project « On arrive enfin à rattacher le volume financier qu'on engage à des vraies histoires, avec des vrais projets, avec des choses où les gens nous racontent pourquoi ils veulent utiliser l'IA pour coder. », we finally manage to tie the financial volume we commit to real stories, real projects, things where people tell us why they want to use AI to code.

The same reasoning applies to automated agents, with an additional constraint: « si jamais un système commence à faire n'importe quoi et à dépenser plein d'argent, grâce à OpenRouter on peut très facilement contrôler ça, mettre de l'alerting dessus », if a system ever starts acting erratically and spending a lot of money, thanks to OpenRouter we can very easily control that, put alerting on it.

His conclusion deliberately widens the subject: « je sais que je parle beaucoup d'OpenRouter, mais au-delà d'OpenRouter, ce que je voudrais vous montrer c'est comment on fait pour que la valeur que génère l'intelligence artificielle, on s'assure que cette valeur est là » I know I talk a lot about OpenRouter, but beyond OpenRouter, what I want to show you is how we make sure the value generated by artificial intelligence is actually there. He adds a non-financial argument: « quand on voit des tokens et la finance, faut réfléchir que ça a un certain coût en énergie », when we look at tokens and finance, we need to remember that it has a certain energy cost.

IV

What the other sources say

10

75 public videos mention OpenRouter, 12 bring a new fact

Usage reports, market reading and practical implementation, everything that complements, qualifies or contradicts the Back Market case.

Between 2025 and September 2026, 75 public videos mention OpenRouter, across French- and English-speaking technical channels. Most cite it in passing; 12 bring a fact, a figure or a firsthand account.

3 facts published between May and August 2026 An official acquisition: Stripe announced on 19/08/2026 an agreement to acquire OpenRouter, without disclosing the price; Bloomberg put it at over 7 billion dollars. A documented departure: a team of 200 engineers left OpenRouter for Vertex AI, for security and control reasons. A market shift: on OpenRouter, Chinese models exceed 60 % of usage in July 2026, with the caveat that this platform does not see direct usage of the US APIs.

For an overview of the market, OpenRouter and a16z published on 04/12/2025 a « State of AI » report based on more than 100,000 billion tokens routed in one year. It finds that the adoption of open weight models has accelerated and that they now serve as building blocks for production applications.

The 12 selected videos
Selected videos, sorted by angle
AngleDateChannelVideoWhat it contributes
Usage reports14/09/2026Nicolas MartignoleBilan de 2 mois avec OpenRouter, pour 245 utilisateurs : les vrais chiffresAssessment of the switch: 87 k$ per month at Anthropic in June, 23 k$ at OpenRouter in August, for about 245 developers, with the associated budget mechanism.
Usage reports22/07/2026Nicolas MartignoleClaude Max (abo prix fixe) ou API (pay-as-you-go): ce n'est pas une question de prixAbout 70 providers available, a short list authorised by Back Market, and a liability that stays with the client.
Usage reports11/07/2026Nicolas MartignoleLe vrai coût de l'IA : pourquoi le prix au token est fauxAccording to a presentation by OpenRouter's COO, Chinese models there exceed American models in volume; Nicolas M concludes from this that price per token is misleading.
Usage reports30/06/2026Nicolas MartignolePremiers benchmarks avec Sonnet 5, et comparo avec GLM 5.2Evaluation of OpenRouter to get a single contract, and providers tested: Google Vertex, Amazon Bedrock, Fireworks.
Usage reports29/05/2026AI Native DevHow We Built an AI Code Reviewer for 200 EngineersA team of 200 engineers starts on OpenRouter (a single API, same prices as the model publishers) then moves to Vertex AI: some models did not meet its security rules, and Vertex gives it more control over logs and budgets.
Market17/08/2026BloombergAnthropic's Revenue Jump, The Wealthy Bet on SpaceX · Bloomberg TechStripe would be acquiring OpenRouter for more than 7 billion dollars, according to Bloomberg's sources.
Market22/08/2026BloombergChinese AI Models Gain Ground on US AI in Price and UseOpenRouter's usage data show more than 60 % Chinese models in July 2026. Explicit caveat: OpenRouter is only a platform, and direct usage of the Anthropic and OpenAI APIs is not visible there.
Market06/08/2026AI EngineerThe State of Model Routing — NVIDIA, Cognition, OpenRouterPanel discussion with a representative from OpenRouter, NVIDIA and Cognition on routing: sending each task to the smallest model that suffices, and orchestrating several models together.
Market23/04/2026The Product CrewTech @ France : Comment l’État français déploie sa propre stack IAThe French State presents OpenGateLLM, its open source inference gateway, as an alternative to LiteLLM and OpenRouter.
Practical implementation04/06/2026Cole MedinClaude Plans, Gemini Designs: The Workflow to Build BEAUTIFUL FrontendsClaude plans, Gemini 3.5 Flash designs the interface: access to Gemini goes through OpenRouter from the Pi agent.
Practical implementation06/04/2026The Next New ThingBest & cheapest AI for OpenClawConfiguring an agent with cheaper models via OpenRouter: a few dollars of credit, then comparing models down to the choice of GLM 5.1.
Practical implementation15/06/2025Alex so yesTutoriel Complet : Coder une app entière avec Cline (OpenSource)OpenRouter as a single interface that hides API format differences between model publishers, with its ranking of the most used models.
How the selection was made Each selected video was reviewed on its passages devoted to OpenRouter. The mention count served to spot candidates, not to rank them. A single video can carry more weight through one fact than 10 passing mentions. The summaries rely on the subtitles, automatic for most of them.
The 75 videos that mention OpenRouter
From most recent to oldest. Mentions: number of occurrences of OpenRouter in the subtitles
DateChannelVideoMentions
14/09/2026Nicolas MartignoleBilan de 2 mois avec OpenRouter, pour 245 utilisateurs : les vrais chiffres29
02/09/2026Latent SpaceThe Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO1
28/08/2026Stanford OnlineWhen AI Stops Being a Project: Turning Technology into Real Value for Patients and Providers2
27/08/2026Cole MedinWatch This If Your Coding Agent is Ignoring Your Rules (You Need Hooks)1
24/08/2026IndyDevDanIntelligence EXPLOSION: Harness Engineering with Pi Agent, Deepseek, and Gemini2
22/08/2026BloombergChinese AI Models Gain Ground on US AI in Price and Use3
22/08/2026BloombergBloomberg This Weekend · Trade War Escalates, Pushback on Beef Plan, Bond Market Intervention2
21/08/2026The Next New ThingTop 10 GitHub: AI videos, gorgeous diagrams, token savings and more1
20/08/2026Cole MedinDeepSeek Just Built the Next Generation of Coding Agents1
17/08/2026BloombergAnthropic's Revenue Jump, The Wealthy Bet on SpaceX · Bloomberg Tech6
17/08/2026AI EngineerSecurity Firewall for Agents — Ryan Dahl, Deno1
10/08/2026IndyDevDanEngineers… Your Software Factory NEEDS Agent Sandboxes to SCALE (exe.dev)3
06/08/2026AI EngineerThe State of Model Routing — NVIDIA, Cognition, OpenRouter6
03/08/2026Latent SpaceNext 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten1
01/08/2026The Next New ThingFree: Agent creation, chat, task-assignment and other tools.1
30/07/2026LangChainThe misaligned incentives behind AI coding agents2
27/07/2026IndyDevDanIs Anthropic STEALING Your Data? (While You PAY FOR IT)2
24/07/2026Cole MedinIs Kimi K3 Really That Good?! (Don't Just Believe The Hype)1
22/07/2026Nicolas MartignoleClaude Max (abo prix fixe) ou API (pay-as-you-go): ce n'est pas une question de prix14
20/07/2026The Next New ThingTutorial: how non-developers build apps in Claude Code1
20/07/2026IndyDevDanEngineers... STOP Picking GPT-5.6 Sol OR Claude Fable 5… FUSE THEM1
17/07/2026AI EngineerSpecial Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth4
16/07/2026LangChainThe best AI agents cost less than you think2
11/07/2026Nicolas MartignoleLe vrai coût de l'IA : pourquoi le prix au token est faux31
10/07/2026The Next New ThingFree transcription app, YouTube scraper, and GitHub’s top repos!1
09/07/2026Nicolas MartignoleRAISE Summit 2026 - Paris - Visite rapide et impressions4
04/07/2026Alex so yesJ'ai testé Mistral Vibe : la meilleure alternative open source à Claude Code ?1
02/07/2026AI Native DevLars Trieloff - Building AI agents in the browser, for the browser, of the browser - AI Native DevCo1
02/07/2026Nicolas MartignoleSonnet 5.0, Fable et le benchmark le plus scientifique du monde...10
01/07/2026LangChainIntroducing OpenWiki, an open source agent for repo documentation1
30/06/2026Nicolas MartignolePremiers benchmarks avec Sonnet 5, et comparo avec GLM 5.218
25/06/2026The Next New ThingNew AI Drops: AI does your work, Claude in Slack, cheap AI & more10
23/06/2026Stanford OnlineStanford MS&E435 Economics of the AI Supercycle · Spring 2026 · Applications, Coding AI1
21/06/2026BloombergWeekend Listen: Anthropic's Co-Founder and Top Economist on Doing Research at the AI Frontier ·...1
19/06/2026BloombergAnthropic's Co-Founder and Top Economist on Doing Research at the AI Frontier · Odd Lots1
17/06/2026AI Native DevLiran Tal - Your AI Agent Installed Malware Because a SKILL.md Told It To - AI DevCon June 20261
10/06/2026AI EngineerSovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind2
04/06/2026Cole MedinClaude Plans, Gemini Designs: The Workflow to Build BEAUTIFUL Frontends4
29/05/2026AI Native DevHow We Built an AI Code Reviewer for 200 Engineers2
29/05/2026LangChainIntroducing Managed Deep Agents · Interrupt 261
22/05/2026AI EngineerLobster Trap: OpenClaw in Containers from Local to K8s and Back — Sally Ann O'Malley, Red Hat1
21/05/2026The Next New ThingFounders react: real Hermes use1
19/05/2026Stanford OnlineStanford CS336 Language Modeling from Scratch · Spring 2026 · Lecture 12: Evaluation3
19/05/2026The Next New ThingAnthropic beats OpenAI? + 5 AI Stories You Missed3
19/05/2026AI EngineerDon't Build Slop (4 Levels of AI Agent Maturity) - Ara Khan, Cline1
15/05/2026The Next New ThingFree Claude Code + 9 other apps1
13/05/2026AI EngineerBuilding a Chess Coach — Anant Dole and Asbjorn Steinskog, Take Take Take2
09/05/2026AI EngineerVoice AI: when is the "Her" moment? — Neil Zeghidour, CEO, Gradium AI1
02/05/2026AI EngineerHuman-in-the-Loop Automation with n8n — Liam McGarrigle5
28/04/2026NVIDIA DeveloperIntroducing NVIDIA Nemotron 3 Nano Omni1
23/04/2026The Product CrewTech @ France : Comment l’État français déploie sa propre stack IA2
16/04/2026Cole MedinI'm Building an AI Dark Factory That Ships Its Own Code (Public Experiment)2
15/04/2026AI EngineerPaperclip: Open Source Human Control Plane for AI Labor — Dotta Bippa3
08/04/2026AI EngineerLet LLMs Wander: Engineering RL Environments — Stefano Fiorucci1
06/04/2026The Next New ThingBest & cheapest AI for OpenClaw9
06/04/2026Cole MedinI Built Self-Evolving Claude Code Memory w/ Karpathy's LLM Knowledge Bases1
31/03/2026The Next New ThingFounder of Paperclip shows how1
19/03/2026Cole MedinThe Subagent Era Is Officially Here - Learn this Now1
17/03/2026NVIDIA DeveloperGet Started with Unsloth Studio: Generate Data & Fine-Tune LLMs Locally on any NVIDIA GPU1
13/03/2026Dev With AIAlphonse Terrier - dataaaaa! : Comment l'IA a rendu possible un projet que je n'osais pas commencer2
08/02/2026AI Native DevGitHub's Former CEO On What AI Competition Means for IDEs1
29/01/2026Cole MedinClaude Skills Aren't Just for Claude - Here's How to Build Them for ANY Agent1
27/01/2026AI Native DevWhat GitHub's Ex-CEO Says About the Future of Developer Skills1
22/01/2026Cole MedinI Was Wrong About Ralph Wiggum4
08/01/2026AI EngineerDSPy: The End of Prompt Engineering - Kevin Madura, AlixPartners2
22/12/2025IndyDevDanTOP 2% Engineering: /PLAN 20263
01/12/2025IndyDevDanClaude Opus 4.5: The Engineers' Model1
27/11/2025AI Native DevAdam Larson - Evaluating An LLM's Ability To Code · DevCon Fall 20253
07/10/2025DevoxxRoboCoders: Judgment Day – AI IDEs Face Off By Viktor Gamov, Baruch Sadogursky1
15/06/2025Alex so yesTutoriel Complet : Coder une app entière avec Cline (OpenSource)14
04/06/2025Stanford OnlineStanford CS336 Language Modeling from Scratch · Spring 2025 · Lecture 12: Evaluation1
30/05/2025Alex so yesOpenAI Codex CLI : Coder avec une IA dans un terminal ?1
14/05/2025AI Native DevAgentic workflow powered by Dagger with Kambui Nurse7
07/04/2025IndyDevDanCRACKED Aider AND Claude Code Combo. Zuck ships Llama4. Gemini 2.5 Pro is SOTA?9
19/02/2025AI Native DevBuilding the Ultimate AI-Powered Development Environment with Farhath Razzaque4
V

What the Claude Code guide keeps

11

4 pages of the Claude Code Ultimate Guide incorporate these lessons

Each lesson joins the guide page that already addresses its question.

The Claude Code Ultimate Guide already covered the question of gateways, but scattered across pages, and with a contradiction between 2 of them. One advised against pointing Claude Code at another backend in production, the other recommended exactly this mechanism.

Each lesson from this document is incorporated page by page, without repeating the same figure in 2 places.

Each question has its page
QuestionGuide pageWhat it takes from the Back Market case
How to cap and attribute developer spend?API Gateway, Model AllowlistsThe 200 $ per month cap with self-served extension; about 70 providers available and around 10 used; the rule of 10 providers, 10 data policies; the liability that stays with the client
Named-user licenses, service identities, or both?Subscription Strategy, section 6The move from 87 k$ to 23 k$ per month, with the caveat that 2 levers acted together
Is pointing Claude Code elsewhere legitimate?Pointing Claude Code at Another BackendThe distinction between a contracted commercial gateway and a reverse-engineered proxy, the warning remaining intact for the second
What hardware, and how much does a served token cost?Local vs Cloud InferenceA reference table pointing to the page that addresses each question

2 other pages of the guide extend the topic without revisiting the Back Market case: AI Unit Economics, on cost per accepted task, and Team Metrics, on the metrics of teams working with agents.

Why not a dedicated page Each of these 4 pages answers a different question. A page devoted to « consommer autre chose que Claude », consuming something other than Claude, would have repeated the hardware, the spend and the governance already covered elsewhere. The guide prefers to tie each lesson to the page that owns its question and to link the pages together.
VI

Reproducing the measurement

12

Comparing 2 models on real tasks, via the OpenRouter API

A minimal protocol tests model substitution before touching a team's configuration.

Model substitution is the first thing to test. It needs neither a team gateway nor a spending cap. The test runs on a small, fixed batch of real tasks, 5 to 10 requests representative of everyday work, sent identically to 2 models.

Prerequisites, and what skews the measurement A funded OpenRouter key is required. Every call consumes real tokens, on both models. Start small. Changing a request between the 2 passes invalidates the comparison; keeping them in text files avoids rewriting them.

The most direct test connects Claude Code itself to OpenRouter. The official integration guide fits in 3 environment variables, with no local proxy, to place in the shell profile or in the project's .claude/settings.local.json file.

Terminal Connecting Claude Code to OpenRouter 4 lines
export OPENROUTER_API_KEY="sk-or-..."
export ANTHROPIC_BASE_URL="https://openrouter.ai/api"
export ANTHROPIC_AUTH_TOKEN="$OPENROUTER_API_KEY"
export ANTHROPIC_API_KEY=""   # doit être vide, et non absente

If Claude Code was already connected to an Anthropic account, run /logout once and restart. The /status command should then display https://openrouter.ai/api on the base URL line. This URL takes no /v1 suffix, reserved for the OpenAI-compatible API, which triggers model-not-found errors.

OpenRouter only guarantees Claude Code with Anthropic as provider The integration guide states it in plain terms: « Claude Code with OpenRouter is only guaranteed to work with the Anthropic first-party provider. » OpenRouter recommends setting Anthropic as the priority provider. Running Claude Code on GLM 5.2 or another open-weight model therefore falls outside the guaranteed scope. This case remains to be tested on one's own tasks.

For a controlled comparison outside Claude Code, the first step is to list the available identifiers rather than copy a name seen in a video. The catalogue is public and requires no authentication.

Terminal Listing models and their pricing, sorted by output price 10 lines
curl -s https://openrouter.ai/api/v1/models -o or-models.json

python3 - or-models.json <<'PY'
import json, sys
rows = json.load(open(sys.argv[1]))["data"]
for m in sorted(rows, key=lambda x: float(x["pricing"]["completion"])):
    p = m["pricing"]
    print("%-44s in=%7.2f $/M  out=%8.2f $/M" % (
        m["id"], float(p["prompt"]) * 1e6, float(p["completion"]) * 1e6))
PY

The response contains data, total_count and links; each entry carries id, pricing.prompt, pricing.completion and context_length. As of 16/09/2026, the catalogue holds 443 entries. Identifiers prefixed with a ~ are aliases that follow a family's latest version: convenient for exploring, best avoided for comparisons, since their target can change between 2 passes.

Second step, send the same request to both models. The following command reads the request from prompt.txt and saves each response to a separate file, from any working directory.

Terminal Same request, 2 models (requires curl and jq) 10 lines
export OPENROUTER_API_KEY='sk-or-...'

for model in anthropic/claude-sonnet-5 z-ai/glm-5.2; do
  jq -n --arg m "$model" --rawfile p prompt.txt \
    '{model: $m, messages: [{role: "user", content: $p}]}' \
  | curl -s https://openrouter.ai/api/v1/chat/completions \
      -H "Authorization: Bearer $OPENROUTER_API_KEY" \
      -H "Content-Type: application/json" \
      -d @- > "reponse-${model//\//_}.json"
done
Scope of verification The base URL, the authentication header and the /api/v1/chat/completions path come from OpenRouter's documentation, consulted on 16/09/2026. The completion call itself was not executed for this document, for lack of a funded key. Verify the structure of a first response before automating the reading of subsequent ones.

Third step, compare on verifiable criteria rather than impression: is the response usable without rework, does the code compile, do the tests pass. Then relate each pass's token consumption to the price noted in the first step.

Without this count, there is no way to know whether the cheaper model stays cheaper once rework is counted. The AI Unit Economics guide explains how to build this pairwise comparison, and what a small sample allows one to conclude.

13

Reproducing the per-person cap with a credit-limited key

The 200 $ mechanism is a documented OpenRouter feature, testable alone.

Back Market's individual allocation rests on a documented capability of the management API. A key can carry its own credit limit.

API elements verified in OpenRouter's documentation on 16/09/2026
NeedElementVerified
Entry pointhttps://openrouter.ai/api/v1yes, Quickstart
AuthenticationAuthorization: Bearer <key> headeryes, Quickstart
Completion callPOST /api/v1/chat/completionsyes, Quickstart
Catalogue and pricingGET /api/v1/modelsyes, called live
Per-key capPOST /api/v1/keys, limit fieldyes, Provisioning doc
Cap resetDaily, weekly or monthly, per keyyes, Enterprise page

On its Enterprise page, OpenRouter documents per-key credit caps with automatic reset, daily, weekly or monthly. One key per developer with a monthly reset therefore reproduces Back Market's 200 $ per month allocation, without in-house accounting.

What was not verified by execution POST /api/v1/keys was not called for this document. This operation creates a billable resource and requires a provisioning key. The existence of the limit field (« Optional credit limit ») comes from the documentation, not from an execution. The exact behavior when the cap is reached, error code and message returned to the client, remains to be observed before building a process on it.

For an individual test, reproducing the 3-question form depends on no API. After a week of use, answer for yourself: how many sessions, which session cost the most, and what is the split by model. Nicolas M credits it with a pedagogical effect, and it can be tested without a budget.

Form described in Nicolas M's video of 14/09/2026.
Minimal protocol over 2 weeks
  1. Week 1, measure without changing anything

    Switch everyday usage to a dedicated OpenRouter key, without a cap and without changing model. The goal is to get a baseline and a starting blended cost.

  2. End of week 1, answer the form

    The 3 questions, on one's own console. Write down the answer before looking at the following week. That is the measure of the pedagogical effect, and it disappears if reconstructed after the fact.

  3. Week 2, segment by difficulty

    Explicitly route simple tasks to a cheap model and keep the high-end model for complex tasks, following the cutoff Nicolas M describes. This week tests segmentation.

  4. Compare the 2 blended costs

    The gap between the 2 weeks is the individual equivalent of the 13 % observed at Back Market. Even at very different volumes from one week to the next, the indicator stays comparable since it is expressed per million tokens, but the total bill is not.

14

What this experience report does not demonstrate

Applying the observed factor elsewhere takes 4 precautions.

What the video does not count It contains no measurement of the quality of the code produced, of the rework rate, of time spent, or of the number of tasks completed. The claim, « les utilisateurs n'ont pas perdu leur productivité », users did not lose their productivity, is a qualitative observation, presented as such. Without a quality measurement, the 3.8 factor on spend is not enough to establish a saving. The agentic metrics in the Team Metrics guide offer a framework for measuring what the video does not.
A public benchmark gives a first order of magnitude Superconductor compared several agents on 03/08/2026 on tickets drawn from its own merged Rails pull requests, scored by separate evaluator models on correctness, completeness and code quality. Kimi K3 reaches about 80 % quality there, on par with Opus 4.8, for roughly a quarter of its average cost per ticket; GLM 5.2 lands about 5 points behind. Kimi K3 is also the slowest, at about 44 minutes per ticket, roughly double Opus 4.8. These results cover series of 10 tickets and a single codebase. They give an order of magnitude to confirm elsewhere.
The gateway and the cap arrived in the same window The video does not allow attributing a share of the drop to each. Nicolas M presents them as joint. Attributing the factor to the gateway alone would go beyond what he states.
Automatic captions distort names The transcript renders « Anthropic » as « entropic », « Claude Code » as « clô code », « Tencent » as « 10 cent ». The most recent model names shown on screen could not be reliably reconstructed from audio alone. The exact identifiers used in this document come from the OpenRouter catalogue, not from the transcript.
245 developers, a result obtained at company scale 245 developers allow a multi-provider split, an arbitration committee and volume negotiation. Solo use or a 5-person team gets neither the same mix effect nor the same governance need. At small scale, segmenting by task difficulty and regularly reading the blended cost remain applicable.

The population figures vary within the video itself: the title announces 245 users, the introduction speaks of 250, the first Anthropic screen of 230 to 240. Nicolas M uses them interchangeably to refer to the same population. This document keeps his wording rather than harmonizing a figure he never settled on.

A question about Back Market and OpenRouter?

Have a question, a test result or a correction to share? Message me on LinkedIn.

Contact me on LinkedIn
15

Glossary

Terms used in this report.

Blended cost
Average price paid per million tokens over a period, across all models. OpenRouter displays it on its console. At constant volume, it falls as users shift toward cheaper models.
Open weight
Model whose weights are published, so it can be hosted by any provider. Distinct from a proprietary model, accessible only at its publisher.
Provider
Host that serves a model. The same open-weight model can be served by several providers, at different prices.
Zero data retention
Commitment by a provider to retain no request data. Selection criterion cited to rule out otherwise well-ranked models. Abbreviated ZDR.
Other glossary terms
BYOK
Bring Your Own Key: using one's own keys or credits at a model provider through a gateway. OpenRouter then applies a separate fee schedule.
Claude Code
Command-line development agent published by Anthropic. It consumes either the Anthropic API directly or a compatible gateway.
GLM
Open-weight model family published by Z.ai. GLM 5.2 is the most consumed model in the experience report described here.
Frontier model
Proprietary model, too heavy to run at a third party or on internal infrastructure, served only by its publisher. Claude models belong to this category.
OpenRouter
Commercial gateway that exposes several hundred models behind a single API and a single bill, and splits requests across providers.
Gateway
Intermediary service between a client and several model providers. It centralizes authentication, billing and routing. Known in French as passerelle.
Token
Unit of text splitting billed by providers. Prices are expressed per million tokens, separately for input and output.
Workspace
Isolated billing space within OpenRouter. Used here to carve a project budget out of a developer's individual allocation.
16

Conclusion, sources and updates

Report scope and method

This document starts from 2 videos by Nicolas M, published on 22/07/2026 and 14/09/2026, which describe the migration of about 245 developers from direct Anthropic access to the OpenRouter gateway, then the budget mechanism that came with it. It checks them against other public videos that discuss OpenRouter, against OpenRouter's documentation and pricing, and points to the pages of the Claude Code Ultimate Guide that cover each question. Quotations come from the subtitles, mostly automatic, which garble proper nouns. Model names were restored when they were certain, and flagged when they were not. OpenRouter prices were read from the public API on 16/09/2026. Quotations stay in French, their source language, with their meaning in English right after, so every figure remains checkable against the video. This document measures neither the quality of the code produced nor the time developers spent. It covers a spend, and the traceability of that spend.

Conclusion

Moving to a gateway and capping spend per person are 2 separate decisions, which Nicolas M presents as joint. The gateway opens access to other models and breaks the spend down by model; the cap limits what each developer commits. The video does not allow saying which share of the drop belongs to each.

The cheapest transferable mechanism is the 3-question form required before the 20 $ extension. It forces the developer to read their own console before spending more. It needs neither a gateway nor a budget, and can be tested alone, on one's own usage.

Dated sources

“Accessed on” gives the date the source was read. A source may have been published earlier and later updated.

Update history

DateVersionChanges
18 September 20262.8.0Rhythm and repetition pass on both languages: 25 announce colons turned into sentences, 7 paragraphs opening on « On » given back their subject, 5 chapter decks returned to a finite verb, 5 paragraphs split to create short breaks, and the English glosses that follow a French quotation set off by a comma throughout. Repetitions removed: the monthly comparison amounts replayed in a callout, 2 quotations repeated word for word from one section to the next, the part II trailer for a session detailed in part III, and glossary definitions repeated in the body.
17 September 20262.7.0English edition. Quotations stay in French, their source language, with their meaning in English right after, so every figure remains checkable against the video. Video and source titles keep their published wording.
17 September 20262.6.0Additions verified against primary sources: official announcement of the Stripe acquisition, OpenRouter routing parameters and per-key caps, platform fees, connecting Claude Code to OpenRouter and its official reservation, Superconductor's and Inferya's quality and cost-per-task measurements, the FinOps for AI guide. Links to the guide's AI Unit Economics and Team Metrics pages and to 2 published blog articles. Dates in DD/MM/YYYY format.
17 September 20262.5.0Amounts in k$ notation, links to the companies and products cited at their first mention, link to the speaker's profile, highlighted quotations, copy button on commands and an "Other resources" menu.
16 September 20262.4.0Syntax highlighting for commands and links to the portfolio, the Claude Code Ultimate Guide and GitHub in the menu.
16 September 20262.3.0Full editorial pass: removal of repeated opposition phrasings, signposting subheadings and sentencious closing lines; titles rewritten to state the fact; infographic captions replaced with their source and date. 4 claims that overreached the source were corrected, including the idea that neither of the 2 levers would suffice alone, which the video does not establish.
16 September 20262.2.0Addition of 5 infographics: contractual liability, tokens-versus-bill gap, drop in the monthly bill, scope of the blended cost, 3-question form. Each image rests on a fact already established in the document; its text and proportions were checked one by one. Wording corrections: a leftover imperative and spoken remarks presented as written.
16 September 20262.1.0Document made readable by any reader: full list of videos mentioning OpenRouter and an annotated selection of 12 of them, cross-references to the guide's public pages, test protocol rewritten around the OpenRouter API alone. Removal of elements specific to the author's own environment.
16 September 20262.0.0Scope widened: the video of 22/07/2026 and the liability constraint it raises, comparison with other public videos, cross-references to the Claude Code Ultimate Guide. The count of authorized providers remains an open reservation, not an established figure.

Copy the report as JSON

Automatic copying is unavailable. The text is selected: press ⌘C or Ctrl+C, then paste it into your assistant.

Share this report

Download the HTML report