Skip to main content
FB.
Share on LinkedIn

Lean and AI: follow the work until it reaches the user

AI can speed up one step while the request waits elsewhere. A practical way to map the full service, protect quality and measure accepted user outcomes.

• 10 min read • Updated Oct 4, 2026
Florian Bruniaux

Written by Florian Bruniaux

AI Founding Engineer at Méthode Aristote, 13 years scaling engineering teams from developer to CTO. Builds open-source developer tools, see what else I've shipped.

AI makes it easier to produce a draft, a specification, a pull request or a customer reply. That is useful only if the work survives the next decision and reaches the person it was meant to help. A faster step can feed a longer queue.

This is the question I would ask before scaling an AI tool: which request will reach a user sooner, with the expected quality, because this step is faster? If the answer names only generated volume, the value stream has not been measured yet.

Start with the request, not the tool

Lean starts from value as the customer experiences it and from the sequence of work needed to deliver that value. The Toyota Production System describes two useful principles: Just-in-Time, producing what is needed when it is needed, and jidoka, detecting a problem and stopping to prevent defective work from passing downstream. These principles come from manufacturing. Applying them to an AI-assisted service is an analogy and a design choice, not evidence that a software team will reproduce Toyota’s results.

For a software change, the sequence may run from customer problem to product decision, implementation, review, testing, release, adoption and support. For customer service, it may run from a request to classification, evidence gathering, a proposed answer, approval, delivery and resolution. In both cases, generating the middle artifact faster is a local improvement. The user experiences the end of the sequence.

Andrew Ng has described product, legal and marketing work as possible constraints once engineering becomes faster. That is a practitioner’s observation, not a measured law of bottlenecks. It is a useful prompt to look beyond code before buying more coding capacity.

Draw the path of one real request

Pick one type of demand and follow a small sample of completed requests. Record when each request entered and left every stage. Include rejected work and corrections, not just successful releases. Use the same definition and observation window before and after changing the workflow.

StageQuestion to answerObservable signal
IntakeIs the problem worth solving?Time waiting for a decision; requests declined or clarified
CreationWhat artifact does AI help produce?Active work time and artifacts submitted
VerificationCan a reviewer establish that it is correct and appropriate?Time waiting for review; reviewer effort; rejected or revised artifacts
DeliveryDid it reach the intended user?Time to release or response; failures and corrective work
OutcomeDid the user get the intended result?Adoption, resolution or another outcome defined for this request

Do not add these rows into a magic productivity score. Their value is diagnostic. If creation falls while verification rises, the constraint may have moved. If delivery is faster but outcomes do not change, the faster path may be carrying work users do not need.

The unit matters. A count of pull requests is not a count of useful changes, just as a count of support replies is not a count of resolved cases. Choose an outcome and keep its denominator visible: the target users, eligible requests or accepted tasks.

Add a stop signal where mistakes become expensive

Jidoka is often flattened into “add more automated tests.” The stronger question is where the process detects a defect, who can stop the flow, and how the cause is removed before the same defect arrives again. In an AI workflow, a failed evaluation, an unsupported claim or a reviewer unable to understand a change should be visible before publication or release.

That does not require identical scrutiny for every item. Low-risk, reversible work can take a short path. A change with a large blast radius, sensitive data or an irreversible user effect needs a stronger check. Define those paths before the queue fills, and record how often the stronger path catches something material. A review gate that nobody has capacity to operate merely hides work in a new waiting room.

DORA’s qualitative analysis of 1,110 Google engineers’ responses reports that AI can move developer work toward auditing and verification. The responses describe a shift in experience; they do not establish that every team will see the same effect. They do justify measuring review effort alongside generation speed.

Prepare the repository before adding agent capacity

In a FlowCon talk on Lean standards for AI, Theodo’s Marek Kalnik shows an agent producing inconsistent translation keys from a repository that already mixed naming conventions and languages. He describes shared working standards and criteria for a project that agents can use reliably. The example names a plausible source of rework; it does not measure how much time or money those standards saved.

Factory’s Agent Readiness model checks repository conditions such as build commands, tests, instructions, observability and security. In Eno Reyes’s talk, the suggested order is to make one task verifiable before assigning many tasks in parallel. A targeted test that catches the bad translation key is a software analogue of poka-yoke, or error-proofing. An engineer who knows the repository must choose which mistake matters and verify that the check catches it. A configuration file or readiness score alone cannot show that the mistake was prevented.

I would start with one recurring failure in a real task, change the check or instruction that allowed it, and repeat comparable work with the same model and acceptance criteria. I would record retries, token use and reviewer corrections, then follow accepted work through release and the intended user result. Factory’s five-level rubric can help locate a gap, but its score is not a measure of customer value or a controlled estimate of agent productivity.

Test a mistake the harness failed to catch

A legacy migration gives this a concrete test. Suppose an agent cites a real source line but misreads an operator or misses a side effect in a library. The citation check passes, yet the behavior claim is false. I would keep the citation check, ask a separate reviewer to follow the call path, and run characterization cases against the existing application before changing it. The first migrated unit should reach the normal review and release boundary before parallel agents take on more units. A source-only map is a candidate for investigation, not permission to migrate unobserved behavior.

The control has its own cost. Track false blocks, repeated CI runs, manual takeovers, reviewer effort and maintenance alongside accepted changes. For a fragile file format, test that the actual edit path preserves bytes and line endings, including a write through another tool. A rule installed in a repository is not proof that the running agent obeyed it. The guide develops this as a repository-harness countermeasure loop, with context selection and method choice treated as separate decisions. This is a proposed test design, not a measured migration gain.

Keep the evidence in its lane

Two field studies show why an output metric needs a downstream check. In Microsoft’s early-2026 CLI-agent rollout, adopters merged about 24% more pull requests in the four-month observation window. A study of one AI-forward company found per-capita merged-PR throughput at 2.09 times its pre-mandate level while reviewer load roughly doubled. Both count merged PRs. Neither measures customer value, and the second study cannot isolate AI as the sole cause of the change.

A six-week ICSE-SEIP case study at Globo provides a useful counterexample to the idea that review must absorb every gain. Tasks the team tagged as using GenAI had an average Jira development cycle time 23% below the company’s historical figure. Participation was voluntary, tasks were tagged by the team, and the comparison was not randomized. The paper does not establish whether AI caused the difference, whether the task mix was equivalent, or whether release quality and user outcomes improved. It does establish that a longer downstream queue is a hypothesis to test, not a result to assume.

The studies are relevant because they show that creation and acceptance capacity can grow at different rates. They do not tell another organization what its own gain will be. A Doctolib product manager’s warning makes the same point as practitioner testimony: more AI requests could also mean more bugs. The warning is a reason to check defects, not a claim that bugs actually doubled there.

What the company cases actually measure

Qonto provides a concrete example of Lean applied to an AI-assisted engineering task. In January 2025, a small team started a kaizen to migrate an Ember application toward React. Their engineering account describes a CLI that runs an AI conversion, a deterministic codemod, a second AI pass and human review. One engineer, with limited help, migrated 8,632 lines in a two-week evaluation starting February 10. Qonto’s roughly 50 lines per engineer-day without the tool was an estimate, not an observed control group. The reported peak of 20 times manual output is therefore a comparison against that estimate, and the unit is migrated code. The account does not give the full review cost, defect rate after integration or time until customers benefited from the migration.

The kaizen matters more than the headline multiplier. Qonto first tried a web chatbot, then moved into the development workflow, and later reconsidered the CLI because engineers preferred an IDE extension. Its earlier engineering talk on quality discusses right-first-time work and limiting work in progress. That 2023 talk predates the migration experiment; it establishes an existing Lean vocabulary, not the effect of AI. In a 2026 vision statement, Qonto’s president invokes jidoka and kaizen for AI in financial workflows. A stated direction is not a measured product result.

Theodo reports a different unit. Its ComPaRe EMA account says fifteen features reached production in three weeks for AP-HP. The two-month comparison is the team’s estimate of delivery without agents. The case describes a human reviewing the technical strategy, tests and code, and stopping the agent to improve instructions or checks when it fails. Those are observable control choices; the article does not publish subsequent adoption or a comparable defect series. In a separate conference talk, Fabrice Bernhard describes a team facing pull requests of about 4,000 lines across 53 files. The team measured merge lead time and experimented with reusable code blueprints to focus review. The talk reports that merge lead time fell, without a published series or a user outcome. Do not combine these two stories into one measured project.

A five-year practitioner study at Siemens Electronics Works Amberg examines a different question: how AI decisions fit a Lean quality culture. Van Giffen and coauthors identify tensions between human and AI analysis, transparent and opaque reasoning, and specification-led and data-led processes. The study proposes ways to manage those tensions. It does not quantify an AI-driven speed or quality gain. For software teams, its practical lesson is to keep the person responsible for the process able to inspect a recommendation and decide what stops when it is wrong. That is an application to software, not a result measured at Siemens.

The reverse direction, AI supporting an existing Lean practice, is visible at the Lean Enterprise Institute. Art Smalley describes RootCoach, a tool that critiques root-cause reasoning. His claim that it beats 90% of coaching he has seen is personal judgment, without a published blind comparison. Matthew Savas names the risk: frequent AI feedback can make the tool an authority and displace the dialogue through which a manager develops a problem solver. A useful evaluation would score reasoning quality and the person’s later independence, not count generated A3 documents.

Open-source projects make some of these ideas inspectable. Andon records recurring agent defects and proposes hooks as countermeasures; its own README says a completion check can accept an agent-written verification entry that is false. RaiSE and AI-Kaizen describe Lean or PDCA workflows. Their repositories show proposed controls, not independently measured improvements in customer outcomes.

Run one bounded improvement

Before adding AI to a stage, write down the request type, intended user result and current waiting points. Then change one step and compare equivalent requests over equivalent windows. Keep four observations together: elapsed time to accepted outcome, time spent waiting, correction or failure rate, and human effort across the whole path. DORA’s delivery metrics help with the release portion; they do not replace a product or service outcome.

If a new tool reduces drafting time but increases review or support effort, adjust the work admitted to the queue, the verification method or the task boundary. If the end-to-end result improves without degrading quality or team health, expand carefully to the next request type. The point is to learn from the system you actually run, not to import a productivity percentage from someone else’s study.

For software teams, the guide’s team-metrics chapter develops the measurement layer. My AI productivity review examines the field-study denominators in more detail. This article’s job is narrower: keep the user request in view when the middle of the workflow becomes very fast.

Contact