Why AI ROI stays invisible, even to the people measuring it
Claims that nobody can measure AI ROI and that 95% of pilots return zero are weaker than their usual retellings, once their sources are checked.
Jump to summaryWritten by Florian Bruniaux
AI Founding Engineer at Méthode Aristote, 13 years scaling engineering teams from developer to CTO. Builds open-source developer tools, see what else I've shipped.
TL;DR
| Claim | What actually holds up | Status |
|---|---|---|
| ”Nobody can measure AI task cost or ROI” | Ed Zitron, Bloomberg Technology, June 2, 2026 | A media commentator’s argument about market incentives rather than a measured figure |
| ”Only 5% of organizations see real value from AI” | Repeated on conference stages, including at aidevcon | The report itself switches denominators: its executive summary says organizations, while its implementation funnel measures custom enterprise AI tools |
| MIT NANDA report behind that 5% | Aditya Challapally et al., MIT Media Lab, July 2025 | Preliminary finding with limits: conference-sourced sample, six-month window, and an undisclosed institutional interest in agent infrastructure |
| ”5 to 10% needle movement” on business results | One attendee, DevCon London, aidevcon, June 23, 2026 | Verified against the raw transcript; the AI-generated summary was excluded |
| ”A token is intelligence… turning tokens into industrial ROI” | Clément Dietschy, CEO of Ask For The Moon, tpc, June 24, 2026 | Verified against transcript |
| $214, Devoxx France ticketing rewrite | Nicolas Martignole, covered in part 1 of this series | The rare case in this corpus with both a complete project and a checkable cost |
Ed Zitron, CEO of the tech PR firm EZPR and host of the Better Offline podcast, told Bloomberg Technology on June 2, 2026 that nobody can measure what an AI task costs or what it returns, a gap he argued lets companies like Anthropic and OpenAI chase IPO valuations without proof that their spending pays for itself. That’s a media commentator’s argument about market incentives, and it supplies no measured number of its own. A second account comes closer to the ground. At DevCon London, on the aidevcon stage in June 2026, one attendee described the same problem in blunter terms. “Everyone is spending so much money,” they said. “So much effort on this, but not actually seeing any needle movement. The needle movement is very, very small, around the 5 to 10%. And that’s the main challenge.”
That invisibility isn’t only a macro problem investors squint at from a distance. It shows up at the tooling layer too. Mapping the token-reduction toolbox covers a specific blind spot. A team on a Claude Max or Pro subscription has no per-call invoice to read, because the metered billing that let Nicolas Martignole add up his Devoxx France spend doesn’t exist for them. There are no line items and no arithmetic to redo. If the spend itself never surfaces per call, measuring the return against it gets harder, whether the missing number sits behind a multi-billion-dollar IPO pitch or a single engineer’s monthly subscription.
The 5% everyone mis-cites
The number that gets repeated most often in this corpus is a variant of “only 5% of organizations get real value from AI.” Richmond Alake, introduced in the recording as Oracle’s Director of AI Developer Experience, used it at an aidevcon talk on memory engineering published November 21, 2025, to argue for a specific fix: “The 95% that didn’t realize value and the 5% that did, the main difference was that the 5% were building agents with memory.” The extracted summary of that talk frames the 5% as a share of “organizations.” That’s the version that keeps circulating from stage to stage. The underlying report uses that framing too, then switches denominators in its implementation funnel.
The number traces back to a report titled The GenAI Divide: State of AI in Business 2025, written by Aditya Challapally and colleagues at the MIT Media Lab’s NANDA initiative and published in July 2025. Its executive summary says 95% of organizations see zero return, while a later implementation funnel says roughly 5% of custom enterprise AI tools reach successful implementation. Those are different denominators, and the report does not publish enough underlying data to reconcile them. Its sample also leans heavily on conference attendees instead of a representative cross-section of businesses, and its measurement window runs six months, short enough to miss value that compounds over a year or more. Project NANDA develops decentralized agent infrastructure aligned with the report’s recommendations on memory and integration. The report does not disclose that institutional interest as a conflict, and the available text does not establish that the authors sell a competing commercial stack. None of this makes the 5% false. It makes the number preliminary, internally inconsistent, and unsafe to repeat without its denominator problem.

Measuring the wrong thing
Sourcegraph’s product team made a related point from a different angle, in a talk framing enterprise AI adoption back in January 2025: “We can no longer hide behind output metrics as developers, we need to report business impact like every other department.” Stripped of the vendor framing, the argument is methodological. A completion rate, a lines-of-code count, an acceptance percentage on an autocomplete suggestion, none of it says whether the business is better off. It says whether the tool got used. Those are different questions, and most of the numbers that circulate in this corpus answer the easy one because it’s the one an IDE can log automatically, while the hard one requires someone to sit down and connect a deployment to a dollar figure months later.
The pattern isn’t new to AI tooling specifically. AI velocity is bidirectional documented the identical blind spot on the delivery side: velocity and technical debt climb together, and no ledger anywhere nets one against the other while there’s still time to act. ROI measurement fails the same way, one layer up. The win gets counted because someone built a dashboard for it. The cost that offsets it usually doesn’t, because nobody built the matching dashboard on that side.
A job description, not a receipt
Clément Dietschy, CEO and co-founder of Ask For The Moon, put the same idea into a line built to be quoted, on the French podcast tpc on June 24, 2026: “A token is intelligence and enormous value. My job is turning tokens into industrial ROI” (“Un token, c’est de l’intelligence et énormément de valeur. Mon boulot, c’est de transformer du token en ROI industriel”). It’s a clean formula, and it’s honest about what the job consists of once the platform pitch gets stripped away. The job converts a unit that costs money into an outcome someone downstream can point to and defend.
By its own terms, it describes intent rather than a documented result. Dietschy states what his job is. He doesn’t name a specific project his platform delivered with a number attached, and that gap runs through every claim collected so far in this piece: a market analyst without a measured number, a statistic whose denominator got swapped in transmission, a methodological correction about which metric to trust, and now a mission statement. Four different registers share one gap: none names a project with a number attached.
The one number with a receipt
Put these four claims side by side and a pattern falls out that none of them states directly. Zitron’s argument is about the absence of any measurement at all. The MIT NANDA report switches between organization-level language and a tool-level implementation funnel. Sourcegraph’s point is that the metric usually being measured is the wrong one. Dietschy’s line is aspirational. None of them, on its own, hands you a project with both a complete scope and a number that can be checked against a source.
The spend part 1 of this series documented does. $214, an entire ticketing system rebuilt end to end by one engineer, with a total that Martignole added up by hand from his billing dashboard on camera and that was checked word for word against the transcript. That pairing, a finished project plus a checkable cost, is what makes Nicolas Martignole’s number usable as evidence. Most of the claims gathered for this piece never clear that bar. The 5% figure comes from a report whose headline and implementation funnel use different units without naming the successful deployments. The 5-to-10% needle movement describes a trend without a specific project behind it. Even Dietschy’s formula, sharp as it is, describes a job. It isn’t a record of a job already done.
$214 is a rare figure in this corpus with that pairing, a complete project and a checkable cost. It is a cost, not a measured return, so it shows what a checkable number looks like without settling the ROI question. Until more of the industry publishes numbers with both a complete scope and a source, the evidence collected here cannot support a specific industry-wide ROI percentage.
YSNK
(You should now know)
- None of the four ROI claims collected in this piece, Zitron’s, MIT NANDA’s, Sourcegraph’s, Dietschy’s, pairs a complete project with a checkable outcome
- The widely-cited “only 5% of organizations see value” figure comes from MIT NANDA’s July 2025 report, which switches denominators between its executive summary (organizations) and its implementation funnel (custom enterprise tools)
- A Claude Max or Pro subscription has no per-call invoice, so the spend that would let a team measure ROI does not surface at call level
- Martignole’s $214 Devoxx France rewrite, covered in part 1, remains the rare pairing in this corpus of a finished project and a number checkable against a source
- Until more of the industry publishes numbers with that same pairing, this corpus cannot support a specific industry-wide AI ROI percentage
Go Further in the Claude Code Guide
Practical resources selected to help you take the next step.
Related articles
The price-per-token lie: why cheaper models don't mean cheaper bills
Opus 5.5 changes input, output and cache economics. Price the accepted task, including retries and review, before choosing a cheaper model.
A $214 AI coding agent rewrite
A verified $214 bill for a production rewrite, checked against the source transcript, sits beside a corpus of about 2,900 tech talks on coding-agent costs.
FinOps applied to tokens: who owns the AI bill
Cloud FinOps put a cost study first and a named owner for the number second. Three practitioner cases show that sequence applied to token bills.