UVAL: the protocol I built to stop accepting code I don't understand
Jeremy Twei coined it. Addy Osmani popularized it. Margaret-Anne Storey extended it to teams. Here's what I built to fight all three.
Jump to summaryWritten by Florian Bruniaux
AI Founding Engineer at Méthode Aristote, 13 years scaling engineering teams from developer to CTO. Builds open-source developer tools, see what else I've shipped.
TL;DR
| What | Details |
|---|---|
| The problem | Comprehension debt: code you shipped, code that works, code you can’t explain three weeks later |
| The protocol | UVAL (Understand, Verify, Apply, Learn) puts understanding before shipping, not after |
| What it covers | Individual comprehension across four concrete steps, with honest limits on team-level cognitive debt |
| What it doesn’t fix | Team shared-model loss (cognitive debt) needs additional layers: decision logs, architecture records, PR reviews focused on the “why” |
| Who it’s for | Anyone writing production code with AI assistance who wants to stay the author, not just the approver |
The moment I remember most clearly isn’t dramatic. It was 11pm on a Tuesday. A session management bug in the platform I’m building at Méthode Aristote was blocking a demo the next morning, so I asked Claude Code to fix it. It generated a 40-line solution touching three files, restructuring how tokens were refreshed. I read it quickly, ran the tests, they passed, I shipped it.
Three weeks later, a colleague asked me why we refreshed tokens on every request instead of on expiry. I opened the file, read it again, and could not explain the reasoning, trade-offs, or alternative I had apparently rejected. I could point at the code and say “that’s what it does,” without explaining why we made this choice.
The code worked. I couldn’t explain why it was structured the way it was.
The problem has a name now
For a while I thought this was a discipline issue. A me problem. Then two researchers gave it actual names.
Jeremy Twei coined the term comprehension debt: code that works, code you’ve shipped, code that lives in your codebase, but code you don’t understand. Addy Osmani (Google Chrome team) picked it up in January 2026 and built it into a framework for thinking about AI-assisted development. The framing is precise: comprehension debt is distinct from technical debt, where you know the code is suboptimal and have a plan. Comprehension debt is invisible. You don’t know you have it until someone asks you to explain your own system, and you can’t.
Margaret-Anne Storey (Canada Research Chair, University of Victoria) expanded the idea to teams with cognitive debt: the loss of shared mental models across a codebase and the people maintaining it. Her framing is precise: “Cognitive debt tends not to announce itself through failing builds or subtle bugs after deployment, but rather shows up through a silent loss of shared theory.”
Both describe the same failure mode at different scales. Osmani covers what happens individually when AI generates code faster than you can absorb it. Storey covers what happens to a team. The signature is difficult to catch because nothing breaks visibly. Builds pass, tests stay green, and nobody flags the drift until someone asks a question nobody can answer, usually in a PR review where everyone is too embarrassed to admit they have lost the thread.
I recognized exactly what had happened that Tuesday night. By the time I saw the pattern clearly across multiple sessions, I’d already started building something to address it.
UVAL: Understand, Verify, Apply, Learn
I formalized UVAL in the Claude Code Ultimate Guide, in the learning-with-ai section. The name is an acronym, and the logic is simple: understanding comes before using and shipping.
Understand First
For months, my first move when hitting a problem was to open a new chat and describe it. The results were correct but opaque. I’d paste the solution, it would work, I’d move on. What I wasn’t doing was arriving at the conversation with context that would make the AI’s response useful to me as someone with a specific codebase, specific constraints, and specific things I already understood.
Now I do four things before asking. First, I write the problem in one sentence. If I can’t, I don’t understand the problem yet. The difference between “the auth doesn’t work” and “the login form doesn’t display inline validation errors when the email field is empty on the first submission” matters. The second version produces a useful response. Then I force myself to name three possible approaches, even rough ones, even wrong ones. Having to name them makes me realize how much of the problem space I already understand. After that I identify what specifically I don’t know, not “I don’t know how to do this” but “I know Zod handles schema validation, I don’t know how to surface field-level errors inline in React without triggering a full re-render.” Then I ask.
Those 15 minutes force collaboration on a problem rather than delegation. The hardest part is the first sentence because it reveals whether I understand the problem. When I can’t write it, more often than I’d like to admit, I’m trying to solve something I haven’t fully defined.
Verify (Explain It Back)
When I first tried talking the code out loud (the Rubber Duck approach from Hunt and Thomas’s Pragmatic Programmer), I assumed I’d breeze through it. I’d read the AI-generated code, it made sense, I understood what it was doing. Then I tried to explain not what each line did, but why it was written that way. Why this middleware level and not the route level. Why reduce instead of forEach. Why the retry logic was here and not in the caller.
I got stuck on the third line.
That’s where the real work happens. The places where I can describe what the code does but not why it’s structured that way are the gaps, and finding them is the whole point. Once I’ve identified them, I ask specifically about those lines. “Why is the token refresh triggered here rather than in the middleware?” produces a concrete answer. I remember it because it’s anchored to something I almost understood but didn’t. It’s the difference between passive reading and active verification. The code makes sense when I read it. That’s not the same as understanding it.
If I can’t explain the code to a colleague, I haven’t understood it. That’s the test. Simple to state, and it keeps catching things I thought I’d understood.
Apply (Predict, Test, Adapt)
Apply the solution to the actual requirement. If the generated code already fits, leave it intact. Renaming a variable or changing the iteration style does not by itself demonstrate understanding and can create unnecessary review work.
Before running a check, predict one boundary case and explain the invariant it exercises. Then test that prediction. When a correction is needed, connect it to a requirement or observed failure. A useful next exercise changes an assumption and asks whether you can diagnose the result without repeating the assistant’s explanation.
These are proposed learning checks. This article has no longitudinal result showing that a small edit produces retention six weeks later. Record what you can explain independently, what required help, and what transfers to a new task.
Learn (Capture the Insight)
I have started four learning journals in my career. None survived past week two. Life moves fast, the insight that seemed worth documenting feels obvious by Saturday, and by Monday the context has evaporated along with the motivation to write it down. I don’t think the problem is discipline. I think the problem is friction.
Instead, I use a Claude Code hook on the Stop event. After each session, Claude asks “What’s one thing you learned today?” The answer goes automatically to a log file. Ten seconds, no friction, it happens before I close the terminal. The Stop event is part of the hook catalog covered in Claude Code Under the Hood if you want to understand how to set this up. The insights are rough and sometimes useless, but occasionally there’s one I’d have completely forgotten by the next day, and it’s there in the log.
The other mechanism is CLAUDE.md used as a decision log. Not “we use Zod for validation” but “we use Zod because manual validation missed a nested object edge case in production, and we needed the schema as the single source of truth for field rules.” Without a decision log, you know what the code does but not why it’s structured that way. The code can be read. The reasoning behind it can only be recovered if someone wrote it down at the time. Three months later, that someone is often you, and you don’t remember.
What UVAL covers (and what it doesn’t)
UVAL is the individual layer. It addresses your comprehension: your ability to explain your own code three weeks later, your capacity to onboard someone onto logic you built months ago. The four steps are checkpoints that force engagement precisely at the moments when it’s easiest to skip.
For Storey’s cognitive debt (the team version), UVAL is a starting point, not a solution. A team where everyone uses it individually still needs shared decision logs, architecture records, PRs with “why” sections. That’s the team layer. Neither replaces the other.
Before shipping, I need to explain why the code is structured that way. I do not ship it if I cannot. Reading what code does is easy. Explaining the why is the real work.
Match the learning exercise to the review policy
Explaining code line by line is a learning exercise I use to expose gaps. It is not a requirement that a human deeply reread every generated line in every change. A team can delegate routine review within an established risk policy while retaining ownership of behavior, architecture and recovery. Sensitive paths, uncertain requirements and interacting changes still need scrutiny proportionate to their consequences.
UVAL asks whether I can explain and maintain the decisions I accept. The guide’s allocation of judgment addresses how to divide that responsibility between humans and agents. Passing a test or receiving an agent approval does not by itself answer the comprehension question.
Where UVAL stops
UVAL slows you down on purpose.
If you’re building a prototype you plan to throw away, or exploring a solution space with no intention of shipping, the protocol doesn’t apply. Exploratory work, throwaway experiments, the first hour of a feature when you’re not sure what the feature even is: vibe coding belongs there. Skip the steps.
The protocol is for code that goes to production. For code that others will maintain, or that you’ll debug in six months with no recollection of the decisions that shaped it. For that code, time spent on Understand and Verify is an investment in later diagnosis and explanation. This article does not measure how much debugging time it saves.
It also does not fully solve the collective version of the problem. A team of five each using UVAL individually still needs explicit shared processes: decision logs, architecture records, and PR reviews focused on why rather than only “does it work.” UVAL covers one layer.
It requires discipline. No tooling forces you through the steps. The habit will break when you are tired, under pressure, or in a hurry. The goal is a raised baseline with fewer Tuesday nights where code ships without understanding.
The Verify step is the exception, the one place the habit can become a guardrail the tooling enforces. Test-driven development with Claude turns “explain it before you accept it” into a red test the implementation has to satisfy, and the discipline of never letting the model weaken a test to reach green is exactly the engagement UVAL is trying to protect.
Test what the learner can do after the explanation
In IFTTD episode 362, Yacine Hmito distinguishes formalizing quality criteria in a skill from teaching another person to judge the output. A shared instruction can improve the agent’s context without establishing the operator’s understanding.
After an explanation, change one assumption and ask the learner to predict the effect, reproduce a defect and identify when the task exceeds their authority. Record which steps required help. Use a fresh task for the next assessment: repeating a known solution does not show transfer. This extends UVAL’s verification step; it does not establish that learning through review develops the same judgment as learning through writing.

Check the explanation against behavior
Necessary or Sufficient? tests model explanations through input interventions on two synthetic decision tasks. Its results motivate checking explanations against behavior, but do not validate UVAL or measure human learning. For a code exercise, predict a boundary case before running it, then compare the observed behavior with the explanation.
Make supervision possible before asking for approval
Mitchell, Ghosh and Passi argue that agent interfaces and work organization can undermine the human oversight they rely on. Their paper is a position supported by prior research, not a trial validating UVAL.
For a selected exercise, show the learner the requirement and evidence before the agent’s recommendation. Preserve their first judgment, then record whether the recommendation changes it and why. Give them a real way to challenge the evidence and resume the work. Measure decision accuracy, ability to resume, effort and satisfaction separately. An approval click is not a comprehension test, and this exercise does not require a second human review of every routine PR.

Where to go from here
If you want to try this without committing to the full protocol, start with V. Next time AI generates code for you, before you paste it, explain it back to yourself line by line. Not what each line does, but why it’s there. Notice where you get stuck. Ask about those places. That single change adds five minutes and always produces something useful: either you understood it, or you found the gap you were missing.
The full UVAL implementation, including the Claude Code Stop hook configuration, the CLAUDE.md decision log template, and the complete protocol, is in the learning-with-ai section of the guide. The raw source is also on GitHub if you want to go deeper.
People skip Understand first whenever I walk them through this. I get impatient at Verify. I am curious which one breaks down for you first.
Owning code after it ships
Comprehension debt stays invisible until it costs you something: a derailed sprint review, a colleague’s question you cannot answer, or a debugging session that goes sideways because you left no trace of the reasoning three months ago. UVAL names that failure mode and adds four steps at the moments where it is easiest to skip engagement. The result should be code you can still own when something breaks on a Tuesday night six months from now.
References: Jeremy Twei, original coinage of “comprehension debt”. Addy Osmani, “The 80% Problem in Agentic Coding” (January 2026). Margaret-Anne Storey, “Cognitive Debt” and “Cognitive Debt Revisited” (February 2026).
YSNK
(You should now know)
- Comprehension debt (individual, Osmani) and cognitive debt (team, Storey) are the same failure mode at different scales: nothing breaks visibly, builds pass, tests stay green, until someone asks a question nobody can answer
- The Understand step has four concrete sub-actions before you ask anything: state the problem in one sentence, name three possible approaches even if rough, identify specifically what you don’t know, then ask
- The Verify test asks whether you can explain why code is structured that way to a colleague. Getting stuck on the third line signals a gap
- Apply means predicting behavior, testing a requirement or diagnosing a defect. Leave correct code unchanged; assess retention and transfer rather than inferring them from a cosmetic edit
- The protocol doesn’t apply to throwaway prototypes or the first hour of a feature when you’re still exploring. It’s for code that ships to production and that someone, often you in three months, will need to explain
Go Further in the Claude Code Guide
Practical resources selected to help you take the next step.
Open-source galaxy
Projects used in this path
Related appearances
Live appearances and podcast episodes about this article.
Related articles
From afterthought to infrastructure: how AI config evolves in a real project
Nine months of AI configuration in one production project: evolving responsibilities, profile-based generation and a measured reduction in always-on lines.
Claude is my second contributor: what real Git stats show
6 contributors in our git history. One is an AI. What the commit patterns actually look like after months of Claude Code, beyond the marketing claims.
AI velocity is bidirectional
Everyone talks about shipping 10x faster with AI. Nobody talks about accumulating debt 10x faster. 7 months of production data from a real EdTech platform.
Go deeper
Step-by-step guides that put this into practice.
Claude Code security: the attack surface nobody audits
Hooks are shell scripts with your user permissions. MCP servers are third-party code with access to your credentials. Their timing and access depend on the configured events and server.
Context engineering: the L0-to-L5 playbook
Choose context controls from L0 to L5 according to the failure you observe, from project documentation to scoped rules, behavior checks and shared configuration.
MCP servers: what they actually cost and when to use them
Compare eager and deferred MCP tool loading, distinguish context use from billed tokens, and choose servers for a concrete workflow need.