TDD with Claude Code
Separate test creation from implementation, verify the expected failure, and give automated TDD retries an explicit exit path.
Jump to summaryWritten by Florian Bruniaux
AI Founding Engineer at Méthode Aristote, 13 years scaling engineering teams from developer to CTO. Builds open-source developer tools, see what else I've shipped.
What you'll set up
- ✓ Copy-paste RED/GREEN/REFACTOR prompts for pytest (Django) and Vitest (React)
- ✓ A PostToolUse hook that runs the matching test file after an edit
Prerequisites
- → Claude Code installed and configured (at minimum Level 1 of the setup guide)
TL;DR
I recommend asking for the failing test and the implementation in two separate requests, and reading the failure reason before you accept the green result. Keep the tests out of the model’s reach once they exist, with a hook when habit does not hold, and give any automatic retry loop an explicit exit, because mine ran for months without one. I expect the cycle to move from prompt discipline to hooks that enforce it, since a habit does not hold the line on its own. I have no measured sample behind this guide, only repeated use of my own setup.
In my use, an ambiguous request to write a test often produced a passing test alongside implementation. That looks like TDD without demonstrating that the test detects the missing behavior. Separate the phases and inspect the failure before relying on the green result.
The structural problem
The requests “write a test for X” and “write a FAILING test for X; do not implement yet” express different task boundaries. The latter asks for the contract first. It still needs a check of the output.
The RED phase must fail for the expected reason. A missing interface can justify an ImportError for a new feature; an unrelated missing dependency cannot establish a regression. If the test passes already, investigate what it actually exercises before moving to GREEN.
A failing test acts as a specification only when it fails for the expected reason. A passing test verifies an implementation against that specification. Keeping those roles separate matters because code generated in under a minute can make an interface feel settled before anyone has stated its inputs, outputs, errors, and boundary behavior.
Writing the test first forces that interface decision into the open. If the expected behavior cannot be expressed as an assertion, the implementation request is not ready. This is the useful design pressure in TDD; coverage produced after the implementation does not provide the same signal.
The same pattern shows up wherever people push an LLM on real work: a model will not constrain itself. The constraint has to be engineered in. The guide’s TDD workflow page covers the same discipline as a general Claude Code workflow, RED/GREEN/REFACTOR gates wired into a project from day one rather than applied prompt by prompt.
Two words to state the phase boundary
The first word is FAILING, in uppercase. In practice, writing it uppercase cut down the cases where Claude slipped an implementation past the RED phase, though that is an observation from repeated use, not a result measured across a large sample. “Write a failing test” is ambiguous enough that Claude might still produce a test that technically fails but only because of a contrived assertion. FAILING as a standalone emphasis cuts through that.
The second is Stop. at the end of the prompt, or the equivalent Do NOT implement yet. Without it, “write a test for the subscription endpoint” is syntactically the first half of “write a test for the subscription endpoint, then implement it.” Claude’s training optimizes for task completion. In my use it often continues beyond the phase you intended. Stop. is the explicit boundary.
Use both words in the same request:
Write a FAILING test for [feature]. Do NOT implement yet. Stop.

Three prompts per phase
Copy these verbatim. Adjust the feature description and file paths, leave everything else intact.
RED phase
Write a FAILING test for [feature description].
File: [test file path]
Do NOT implement the feature. Do NOT create any implementation files. Stop.
Run the test and confirm that it fails because the requested behavior or interface is missing.
Do not accept a broken dependency, fixture, or test setup as RED evidence.
Report the actual failure reason; the exception class alone is not sufficient.
For Django/pytest variant (tests/api/test_subscriptions.py, DRF APITestCase):
Write a FAILING test for POST /api/v1/subscriptions/ creating a new subscription.
File: tests/api/test_subscriptions.py
Use APITestCase. Assert response.status_code == 201 and response.data["id"] is not None.
Do NOT create views.py or urls.py. Do NOT implement anything.
Run the test and identify the expected failure from the missing endpoint or behavior.
Exclude unrelated dependency, fixture, and test-setup failures. Stop.
For React/Vitest variant:
Write a FAILING test for SubscriptionCard displaying the plan name and price.
File: src/components/SubscriptionCard.test.tsx
Use @testing-library/react. Assert the plan name and formatted price are in the document.
Do NOT create SubscriptionCard.tsx.
Run the test and identify the expected failure from the missing component or behavior.
Exclude unrelated dependency and test-setup failures. Stop.
GREEN phase
The test at [test file path] is failing with [paste the actual error output].
Write the minimum implementation to make it pass. Nothing else.
File: [implementation file path]
“Minimum” is load-bearing here. Without it, in my use Claude often implements the feature, handles edge cases, adds error handling, and returns a full service layer. That’s all work that belongs in later cycles.
REFACTOR phase
The tests at [test file path] are passing.
Refactor [implementation file path] for readability and structure.
Do not change behavior. Run the tests to confirm they still pass after each change.
Verifying the cycle works
Each phase has a specific success signature. If you see a different outcome, the cycle broke.
RED: inspect the actual failure and confirm it comes from the missing behavior or intended new interface. ImportError, NameError and AssertionError alone do not establish that; exclude broken dependencies, fixtures and test setup. If it passes already, stop and investigate before moving on.
GREEN: all tests pass, and Claude has created exactly the files you asked for, nothing more. If Claude added a serializer you didn’t ask for, it extrapolated. That work belongs in the next cycle, written with its own test first.
REFACTOR: all tests still pass after the refactor. If any test breaks, the refactor changed behavior. Roll back and start the phase again with a more constrained prompt.
The PostToolUse auto-run hook
After a matching Edit or Write on a test file, this hook runs pytest or Vitest against that file and returns the captured output. It does not run the full suite or cover every source-file edit.
Create .claude/hooks/run-tests.sh:
#!/bin/bash
INPUT=$(cat)
TOOL=$(echo "$INPUT" | jq -r '.tool_name')
if [[ "$TOOL" != "Edit" && "$TOOL" != "Write" ]]; then exit 0; fi
FILE=$(echo "$INPUT" | jq -r '.tool_input.file_path // ""')
if [[ -z "$FILE" ]]; then exit 0; fi
# Only fire on test files
if ! echo "$FILE" | grep -qE "(test_|\.test\.|\.spec\.)"; then exit 0; fi
ROOT=$(git rev-parse --show-toplevel 2>/dev/null || pwd)
cd "$ROOT"
if echo "$FILE" | grep -qE "\.py$"; then
OUTPUT=$(pytest "$FILE" -x -q 2>&1 | tail -20)
elif echo "$FILE" | grep -qE "\.(ts|tsx|js|jsx)$"; then
OUTPUT=$(npx vitest run "$FILE" --reporter=verbose 2>&1 | tail -20)
else
exit 0
fi
# A PostToolUse hook that just prints and exits 0 sends its output to the debug
# log, not to the model. additionalContext is what reaches Claude.
jq -n --arg ctx "$OUTPUT" \
'{hookSpecificOutput: {hookEventName: "PostToolUse", additionalContext: $ctx}}'
exit 0
Register it in .claude/settings.json:
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{
"type": "command",
"command": "$CLAUDE_PROJECT_DIR/.claude/hooks/run-tests.sh",
"timeout": 30
}
]
}
]
}
}
chmod +x .claude/hooks/run-tests.sh
The jq line at the end makes the result reach Claude. A PostToolUse hook that prints to stdout and exits 0 sends its output to the debug log, where the model never reads it. These common output forms have different effects:
| Hook output | What Claude receives |
|---|---|
exit 0 + plain stdout | nothing (debug log only) |
exit 0 + JSON additionalContext | the text, injected as context |
exit 2 + stderr | feedback; the tool has already run |
This hook uses additionalContext to provide test output after matching test-file edits. It returns exit 0 and does not block; other configured hooks may have their own decisions. PostToolUse cannot undo or prevent that edit, including on exit 2. Use a PreToolUse decision to deny a later action or a tested Stop policy to control retries. See the official event-specific behavior.
Five anti-patterns specific to Claude
Combining RED and GREEN in one prompt. “Write a failing test and then implement it” asks for a sequence in a single unit. In my use Claude often writes both and optimizes them against each other. The test will be exactly as hard to fail as the implementation is easy to write.
Not checking the failure reason. An ImportError can be expected for a new interface, or it can indicate a broken dependency. An AssertionError reports a mismatch; a value close to the expected result does not establish that implementation was added during RED. Inspect the diff and the traceback to confirm that the test fails for the intended reason.
Tests that assert nothing useful. assert result passes as long as result is truthy. assert result.status_code == 201 catches regressions. Claude defaults to the former when given a vague RED prompt. Specify the assertion shape explicitly in the prompt.
Multiple features per cycle. “Write tests for user registration, login, and password reset” leads to one RED covering three features, one GREEN implementing all three at once, and a REFACTOR touching too much to reason about safely. Keep one feature per cycle to avoid the debugging time that follows.
Skipping REFACTOR. Claude writes minimum implementations that pass tests, and minimum is not the same as readable. The REFACTOR step turns “technically correct” into “maintainable.” Skipping it means the next cycle starts with code that’s already hard to modify.

All five reduce to one rule, stated by Quentin Adam on IFTTD episode 341 and confirmed independently at Devoxx Belgium 2025 (Chris Simon, TDD & DDD From the Ground Up): lock the tests before touching the implementation, and never let the LLM modify a test to make it pass. Adam’s phrasing is the one that sticks.
Tell the model the tests have to pass and it behaves like Ultron deciding the problem with the planet is humanity, so it deletes your tests and reports green. A test weakened to get green protects nothing. When a habit does not hold that line, a hook can.
nizos/tdd-guard (2.3k stars, MIT, actively maintained) packages exactly that: a PreToolUse hook on Write|Edit|MultiEdit that inspects each change before it lands and blocks the ones that skip a test or over-implement. Its decision mechanism has a limit. A PreToolUse denial is deterministic, while tdd-guard asks an LLM (Claude Sonnet by default) whether a change breaks the cycle. That judgment produces its occasional false positives. This is a hard gate around a soft judgment. Claude Code Under the Hood covers where that layer sits in the loop.
When the loop won’t go green
A TDD loop that retries on its own needs a floor, and mine ran without one for months. The Stop hook that re-ran the suite pushed Claude back into the cycle on every red, capped at ten iterations, and when it reached that cap it gave up quietly. Nothing surfaced. The tokens were already spent.
In my installation, a second Stop hook provides the escape condition. For this event, the exit-code meanings are:
| Exit code | Effect on a Stop hook |
|---|---|
exit 0 | let Claude stop |
exit 2 | block the stop, Claude keeps going |
An andon cord, the rope that halts the assembly line, therefore cannot work by blocking. It has to exit 0. Mine stops the line by taking the fuel away instead: it deletes the flag files that the TDD loop hook reads before deciding to exit 2, sends me a notification, and prints the reason to stderr. In that controller, three turns in a row ending on a failing gate command (tsc, vitest, eslint) trigger the escape condition. This is my configuration, not a built-in TDD guarantee; another blocking Stop hook can still prevent stopping.
Two details that earn their keep. The threshold and the kill switch are environment variables (CLAUDE_ANDON_THRESHOLD, default 3, and CLAUDE_ANDON_SKIP=1), because the first thing you want while debugging the hook is a way to turn it off. And the counter resets on any passing gate, so a flaky test can’t walk you to the threshold over an afternoon.
I shipped this on 20 July 2026, eight days after an audit of my own pipeline found that ten-iteration cap terminating in silence. It had been running that way for a long time and I never noticed, because nothing was wired to tell me.

Keep acceptance criteria outside the repair loop
The IFTTD discussion with Yacine Hmito treats reproduction and correction as separate artifacts. Apply that boundary when an agent repairs a failing test: retain the original expected behavior and inspect changes to both implementation and tests.
For a regression fix, verify that the new test fails on the unchanged implementation for the expected behavioral reason, then passes with the fix. A missing dependency is not a valid red phase. If the requirement itself was wrong, have its owner approve the correction and record why; silently weakening an assertion proves nothing about the original defect. This complements the protected-test pattern already described here.
Mutation testing
Once you have a passing suite, ask Claude to verify the tests actually catch regressions:
In [implementation file path], make one mutation: change a > to >=,
remove a condition, or flip a boolean. Run the tests. Report whether they
catch it. Then revert and try a different mutation.
If the tests pass through a meaningful change, they’re not covering what you think they are. Claude can run several mutations in sequence and report which ones slip through, faster than reading the test file line by line. A suite that survives > becoming >= on a boundary condition is a suite that will miss the same regression in production.
This works because the test’s pass/fail state after a mutation is a binary, deterministic signal. Test execution judges the result; Claude runs the loop and reads the outcome.
The manual prompt fits one function at a time, when you want to sanity-check a suite you just wrote. For a whole codebase, Trail of Bits ships a mutation-testing skill in its Claude Code marketplace (trailofbits/skills, 6.1k stars, the skill itself around 2.4k installs) that configures mewt or muton campaigns, scopes the targets, and tunes the timeouts that otherwise make full-suite mutation runs unusable. Its companion skill genotoxic goes further. It triages the mutants that survive against a call graph, separating dead code from genuinely missing tests, and runs necessist to find assertions weak enough that deleting them changes nothing. That last check automates the “tests that assert nothing useful” anti-pattern from earlier in this guide.
YSNK
(You should now know)
- For PostToolUse, plain stdout on exit 0 usually goes to debug output;
additionalContextreaches Claude, while exit 2 supplies feedback after the tool has already run. It does not prevent or reverse the edit - tdd-guard’s block is deterministic, a PreToolUse hook that denies actually denies, but the ruling that triggers it comes from an LLM judging the change. A hard gate wrapped around a soft judgment, which is why it has occasional false positives
- For the Stop event,
exit 0lets Claude stop,exit 2blocks the stop and keeps it going. An automated retry loop that needs a kill switch can’t block its own way out, it has to remove the fuel instead (deleting the flag file the loop reads) and exit 0 - A ten-iteration retry cap that gives up silently on failure is worse than no cap: the tokens are already spent and nothing surfaces unless something is wired to say so
- Mutation testing catches suites that don’t test what they claim: flip a
>to>=on a boundary condition, rerun, and if the tests still pass, they’re not covering that boundary at all
Go Further in the Claude Code Guide
Practical resources selected to help you take the next step.
Open-source galaxy
Projects used in this path
Why this matters
The research and reasoning behind this playbook.
2/6 · Diagnose and repair context drift
Use L0 to L5 to diagnose context drift, then maintain adherence through observation, repair, and replay instead of treating setup as finished.
UVAL: the protocol I built to stop accepting code I don't understand
Jeremy Twei coined it. Addy Osmani popularized it. Margaret-Anne Storey extended it to teams. Here's what I built to fight all three.
AI velocity is bidirectional
Everyone talks about shipping 10x faster with AI. Nobody talks about accumulating debt 10x faster. 7 months of production data from a real EdTech platform.
Related guides
Claude Code security: the attack surface few teams audit
Hooks are shell scripts with your user permissions. MCP servers are third-party code with access to your credentials. Their timing and access depend on the configured events and server.
Claude Code setup, level by level
Three configuration layers for project context, daily tools and persistent memory, with checks for what loads and how it behaves.
Context engineering: the L0-to-L5 playbook
Choose context controls from L0 to L5 according to the failure you observe, from project documentation to scoped rules, behavior checks and shared configuration.