Measuring your own AI usage: three meters, three different answers
Twenty days of my own Claude Code and Codex usage, measured three ways. The tools disagree by a factor of three, and dollars are not what binds.
Jump to summaryWritten by Florian Bruniaux
AI Founding Engineer at Méthode Aristote, 13 years scaling engineering teams from developer to CTO. Builds open-source developer tools, see what else I've shipped.
TL;DR
| What | Number | Status |
|---|---|---|
| Volume across both tools, 20 days, one machine | 38.97 billion tokens | ccusage, cross-checked against my own recount within 1% |
| Share of that volume that is cache read | 94.8% on Claude Code, 96.1% on Codex | Same source |
| Tokens that represent new work (input + cache write + output) | 1.75 billion, a 22:1 ratio against the headline number | Same source |
| API-equivalent price of those 20 days | $24,351 | Public list prices applied by ccusage, not an invoice. I pay four flat-rate subscriptions |
| Codex 7-day quota ceiling | Hit 100% on 6 days out of 20, 90% or more on 11 | Read from the used_percent field Codex writes into its own session files |
| A raw scan of the same transcripts, without deduplication | Overcounts by 87% | Measured against ccusage on the same window |
| Output filtering by rtk on shell commands | 63.9% cut, on 7.4% of Claude’s new tokens | rtk’s own SQLite history, same window |
Every article in this series so far has quoted someone else’s number. This one quotes mine, because I wanted to know what happens when you try to measure your own usage carefully, and the answer turned out to be more interesting than the total.
I run Claude Code and Codex side by side on one Mac, on four flat-rate subscriptions. Three tools claim to tell me what that consumes. They disagree by a factor of three. Working out which one to trust took longer than building the export pipeline that now feeds this article, and the reasons they disagree are the lesson.
Three meters, three definitions
The first meter is rtk, a proxy I use to compress command output before it reaches the model. It reports a cut of about 64% and it is telling the truth, on its own territory: 66,342 filtered commands over twenty days, 66.7 million tokens of shell output reduced to 24.1 million. That territory is narrow. Those 66.7 million tokens are 7.4% of the new tokens Claude Code consumed in the same window, and the 42.7 million saved are about a tenth of one percent of total volume. A filter on shell output is a good idea but not a lever on your bill.
The second meter is ccboard, a dashboard that reads Claude Code’s own statistics cache. Its headline figures had been frozen since August 19, because that cache is written by Claude Code and Claude Code had stopped updating it. Its analytics tab, which scans the transcript files directly, counted every subagent transcript as a separate session, which inflated the session count roughly twofold. Neither behavior is a bug in the arithmetic. Both are a mismatch between what the tool counts and what a reader assumes it counts.
The third meter is ccusage, which reads the same transcript files for both tools and applies published per-token prices. It is the one I kept, for one reason: it deduplicates. Claude Code writes the same assistant message into its transcript several times, once per streamed revision. Summing those lines naively overcounts by 87% against the same window. The fix is to deduplicate on the pair formed by the message id and the request id, which is exactly what ccusage does and what a hand-rolled jq one-liner does not.
I did not take that on trust. I wrote an independent recount straight from the raw transcripts for both tools, and it lands within 0.11% of ccusage on Claude and 0.49% on Codex over the calibration window, with an exact match on 16 days out of 21. Three things had to be right before those numbers agreed, and each one was a silent error first:
- Codex keeps a second directory,
archived_sessions, larger than the live one. Reading only the live directory loses two thirds of the volume. - Codex re-emits a usage event when its rate-limit window refreshes, with no new consumption attached. Counting those adds 2%.
- Group dates in the same timezone on both sides. Splitting days on UTC instead of local time moved individual days by up to 48%.
None of those are exotic. They are the normal state of a log format nobody promised would be stable, and they are the reason I would not publish a figure that came out of a single unverified script, mine included.
Opus 5.5 needs its own measurement window
The export below is a dated calibration window. A measurement on Opus 5.5 needs a separate window that records its resolved model ID, effort, provider, cache categories and task outcomes with the counting-tool version. Use the current cost calculation for API-equivalent spend, and capture subscription quota independently. Opus 5.5 uses the Claude 4.7+ tokenizer, which produces about 30% more tokens for the same text than earlier Claude models, according to Anthropic’s pricing page (read September 30, 2026), so raw token counts across the switch will not compare like for like. Since September 22, 2026, Claude Pro and Team Standard plans default to Opus instead of Sonnet (Claude Code v2.1.280).
The comparison should answer whether the same class of accepted tasks consumes less budget and operator time. A higher token count after adopting more work is not a regression by itself. A smaller count after abandoning difficult tasks is not a saving. Retain failed attempts and the task mix beside the total.
The number that is not a bill
The API-equivalent amounts below belong to the original calibration export. They are not repriced at today’s rates and do not predict the bill for replacement models. Reproducing them requires the per-model token breakdown, resolved model IDs, pricing snapshot, cache categories and counting-tool version. This article does not publish those inputs, and the aggregate totals alone cannot recover them.
Twenty days of my usage carries an API-equivalent price of $24,351: $7,855 on the Claude side, $16,496 on the Codex side. That figure is what the same tokens would have cost at published per-call prices. It is not what I paid. I pay four flat-rate subscriptions. For scale, published list prices are $100 (5x) or $200 (20x) a month for Claude Max and $100, $200 or $500 a month for ChatGPT Pro.
I keep both numbers visible because the ratio is the story. Assume a generous $200 a month per licence, prorated to twenty days across four licences (about $530), and my usage still prices out at roughly forty-five times the subscription cost it sits under. That gap is not a personal anomaly. It is part of the commercial proposition of flat-rate agent subscriptions, and it is also the reason the invoice tells you nothing about whether your usage is reasonable. Under metered billing, waste shows up as money; under a subscription it stays invisible until the quota runs out.
The number that binds
The wall is the quota window, not dollars.
Codex writes its own rate-limit state into every session file it saves: a used_percent value for a rolling 10,080-minute window, which is seven days. Over the calibration window my ceiling hit 100% on six days out of twenty, and 90% or more on eleven. That is the constraint I actually feel while working, and it is invisible in every dollar figure on this page. Usage through claude -p and the Agent SDK draws from the same subscription limits: Anthropic announced a separate credit for it in June 2026, then paused it (Claude Help Center).
Codex logs that percentage where you can read it back, whereas on the Claude side I found no equivalent in the transcript files, so the only way to document it is to screenshot /usage by hand, every day, which is what I now do. Claude Code v2.1.80 added a rate_limits field to statusline scripts (5-hour and 7-day windows); this measurement does not report whether it can replace the manual capture. If you take one operational habit from this article, take that one: under a subscription, capture the quota curve, because nothing else reconstructs it after the fact.
What the volume is made of
The headline 38.97 billion tokens is a bad unit. Strip the cache reads and 1.75 billion tokens remain, the ones that represent content genuinely produced or ingested. Everything else is the same context being read back, turn after turn. Two details behind that ratio matter more than the ratio itself:
- Context size. My median Claude Code turn carries 261,000 tokens of context in a main session, and 74% of main-session turns run above 150,000. Subagents sit lower, at a 149,000 median, and account for 40% of total volume. A long-running session is a large context read back repeatedly, not a small one.
- Cache write TTL. Claude Code wrote 58.7% of its cache entries with the one-hour time to live, billed at twice the input rate, against 1.25 times for the five-minute one. That accounting detail alone can shift the equivalent price of a working day.
Codex reports no cache writes at all. That is a difference in accounting, not in behavior, and it is a good reminder that cross-tool comparison of any single column is a trap. The price-per-token lie covers the same failure mode one level up, at the model price list.
What I am not claiming
One machine, twenty days, one person’s habits. Licences used elsewhere are not in these numbers. The dollar figures are list-price equivalents, recomputed from a pricing table archived with each export so they stay reproducible. The Claude transcripts only go back as far as the retention sweep allows, which is thirty days by default and which I had to raise before this measurement window could exist at all. The Codex quota percentages belong to whichever account was active at the time, which I log separately. The Next Web reported on September 29, 2026 that the included Work and Codex usage on ChatGPT Pro $200 drops from 20x to 10x Plus on October 30, 2026, at the same price. That is a secondary report, not an OpenAI page. On that plan a quota ceiling measured before that date will not carry over. Anthropic’s subscription limits were also set to change on September 14, 2026, according to a message attributed to Anthropic and cited in a t3dotgg video dated August 31, 2026, so a measurement window that spans that date does not compare two identical periods.
These figures are a dated measurement of one twenty-day window. The series does not report a later window, and the numbers above are not updated for the plan and price changes in this section.
The method is the transferable part. Deduplicate before you count, verify your counter against a second implementation, keep the pricing table you used, and measure the constraint that binds you rather than the one your billing page happens to display. FinOps for AI is where that discipline goes once more than one person is spending.
YSNK
(You should now know)
- Three tools measuring the same usage disagreed by a factor of three, and every gap traced to a definition mismatch rather than to arithmetic
- Claude Code writes each assistant message into its transcript several times, so any naive scan overcounts. On my data, by 87%
- Cache reads are 94.8% of Claude volume and 96.1% of Codex volume, so the headline token count says almost nothing about work done. New tokens are the honest unit, at a 22:1 ratio here
- Under a flat-rate subscription the dollar figure is an API-equivalent, not a bill. Mine runs about forty-five times the subscription it sits under, which means the invoice cannot tell you whether your usage is sane
- The binding constraint is the quota window, not money: my Codex seven-day ceiling hit 100% on six days out of twenty. Codex records that percentage in its session files, Claude Code writes none to its transcript files, so capture it by hand, or test the statusline
rate_limitsfield (v2.1.80 and later), while the window is live
Go Further in the Claude Code Guide
Practical resources selected to help you take the next step.
Related articles
Portable agent configuration is a release system, not a shared folder
A release model for instructions, skills, Output Styles, hooks, MCP definitions and BM25 routing without confusing installed files with working behavior.
A $214 AI coding agent rewrite
A verified $214 bill for a production rewrite, checked against the source transcript, sits beside a corpus of about 2,900 tech talks on coding-agent costs.
1/2 · Why I combined three Claude Code skills
Flow Lean fused three response skills. Its historical eval beat each comparator and caught a candidate offering a migration against an inferred target.
Go deeper
Step-by-step guides that put this into practice.
Une seule source configure Claude Code et Codex
Une configuration d'agents inspectée, partagée par Claude Code et Codex : releases immuables, 4 couches d'exécution, routage BM25 des skills et écarts relevés. Rapport complet à lire en ligne.
One source configures Claude Code and Codex
An inspected agent configuration shared by Claude Code and Codex: immutable releases, 4 execution layers, BM25 skill routing and the gaps found. Full report to read online.
Back Market a sorti 245 développeurs d'Anthropic en direct
Ce que Back Market a mesuré après 2 mois sur OpenRouter : budget par personne, choix des fournisseurs, chaîne de responsabilité. Rapport complet à lire en ligne.