Plan-execute: running 20+ agents on a migration
An illustrative 243-file migration split across bounded agents, with a shared brief, local checks and sequential merge validation.
Jump to summaryWritten by Florian Bruniaux
AI Founding Engineer at Méthode Aristote, 13 years scaling engineering teams from developer to CTO. Builds open-source developer tools, see what else I've shipped.
What you'll set up
- ✓ The two-phase model configured and ready to run
- ✓ A decomposition checklist to know if a task can be parallelized
- ✓ The /batch command with quality gates for your stack
- ✓ A review workflow for exceptions that need human judgment
Prerequisites
- → Claude Code installed and configured
- → Git worktrees available (git 2.5+)
- → A migration task with at least 20+ files
TL;DR
I recommend plan-execute for a migration only when you can write a quality gate that returns pass or fail with no human interpretation. If the files are not independent, do not parallelize; if you cannot write the gate yet, write it first. The shape is one read-only planner that writes a brief, bounded agents in separate worktrees, and a sequential merge that reruns the gate after every merge. I expect review to move to the plan and the gate rather than the individual diffs. The 243-file example below is illustrative, and nothing in this summary depends on it.
A JSX to TSX migration can span more files than one session can handle comfortably. Interruptions, shared dependencies and incomplete checks can leave partially converted branches behind. The plan-execute pattern gives each unit a bounded assignment and makes integration a separate step.
In the illustrative example, one Opus agent plans a 243-file migration, then roughly 24 Sonnet agents execute bounded assignments in separate branches. Three files require human review while 240 pass their local checks. This is an illustrative example, not a run I can document: I do not have a log, a measured exception rate or a calibrated duration behind these numbers, and the guide’s rules do not depend on them.
This guide covers how to set it up, why it fails when people skip the decomposition step, and what to do with the exceptions.
Why sequential falls apart at scale
A long sequential session accumulates tool output and decisions. Whether earlier conventions remain available depends on the model, context management and what is retained during compaction. File count alone does not tell you when a convention will disappear or stop being followed.
Each unit still requires dependency inspection, transformation and validation. A mistaken convention in a shared brief can affect many files, whether execution is sequential or parallel. Detect it with explicit checks and review before expanding the run.
Bounded assignments reduce each agent’s scope, while separate branches keep normal checkout edits apart. Neither prevents a shared mistake or an integration failure. Local gates catch the properties they test; the coordinator must run integration checks again after each merge.
The two-phase model
Phase 1 runs a single Opus agent in read-only mode. It reads the entire codebase, identifies patterns, finds edge cases, and flags files that will need special treatment. It touches no files. Its only output is a brief.md that covers: the target conventions, per-directory exceptions, known edge cases from the actual code, and the exact quality gate command each agent must pass before its branch is considered ready. This brief is the load-bearing artifact. Everything that follows depends on its quality. It’s context engineering applied to a one-shot task; the durable, always-on version of the same discipline is what the context engineering guide covers.
Phase 2 spawns N Sonnet agents, each receiving the brief and a file subset. Each agent works in its own isolated worktree (a separate working copy with its own branch), applies the migration, and runs the quality gate. If the gate passes, it marks its branch ready. If it doesn’t, it writes an exception report and stops. A coordinator agent collects the ready branches, merges them sequentially and reruns the quality gate after each merge. Exceptions get routed to a human reviewer with the exact context of what blocked each one.
Phase 1 (Opus)
Read codebase → identify patterns → identify edge cases
Output: brief.md
↓
Phase 2 (Sonnet × N)
Agent 01: files 1-10 → quality gate → branch ready ──┐
Agent 02: files 11-20 → quality gate → branch ready ──┤
Agent 03: files 21-30 → quality gate ✗ → exception ─┤
... ↓
Agent 24: files 231-243 → quality gate → branch ready Coordinator merges
ready branches
Exception queue → human

The key constraint on Phase 1: Opus does not write code. If it touches files, two things go wrong. The read-only invariant breaks, which means Phase 2 agents might encounter modified files that differ from what the brief describes. More subtly, Opus starts making micro-decisions that should be explicit conventions in the brief, and Phase 2 agents receive an incomplete specification. The brief must be complete enough that any Sonnet agent can execute correctly without seeing what any other agent is doing.
Decomposition: the step everyone skips
Poor decomposition can leave agents with overlapping responsibilities and incompatible outputs. Resolve that split before spawning them; this guide has no measured failure rate for this step.
The rule is this: if you can’t write an automatable quality gate, a command that returns pass or fail with no human interpretation required, you cannot parallelize safely. Vague migration instructions multiplied across 24 agents produce 24 different implementations of the vague part. And since each agent can’t see what the others are doing, you don’t find out until you try to merge.
This is the rule I recommend, and it is not mine alone. Four practitioners at Dev With AI, Sciara, Grégoire, Negouai and Drode, Laforge, plus an independent voice on the IFTTD podcast (episode 311), all describing completely different projects, landed on the same rule for automating any kind of review without citing each other: binary, deterministic criteria only, typed code, function under 30 lines, config present. Use pass/fail gates for checks that can be expressed as binary criteria, and keep subjective judgments in human review. If your migration’s correctness can’t be reduced to a binary check, you’re not ready to parallelize it, you’re ready to write the check.
Before spawning anything, work through this checklist against your specific migration:
Are the files independent within the scope of the migration? A JSX to TSX migration is usually safe here because each file converts in isolation. A migration that renames a shared type used across 50 files is not, because the change to the type definition needs to land before any of the 50 files can convert. When independence isn’t obvious, dep-scope maps symbol-level TS/JS dependencies so you can see which files actually share a type before you split the work.
Is the quality gate a command, not a judgment call? tsc --noEmit returning zero is a quality gate. “Looks correct” is not. If you can’t encode the criterion as a shell command, go back to the brief and make the conventions more specific until you can.
Does the brief cover the edge cases Phase 1 identified? If Phase 1 found that files in src/legacy/ use a different import pattern, the brief must say what to do about it. An agent encountering an undocumented edge case will either make a wrong call or raise an exception. The first is worse.
Do you have characterization tests for the behavior this migration could change? This matters for behavioral migrations more than mechanical ones. For a mechanical JSX to TSX pass, tsc checks type compatibility; it does not prove runtime behavior. For a migration that touches runtime behavior, you need a baseline to compare against.

Worktree isolation
Without worktrees, parallel agents working on the same filesystem collide. Git stash conflicts, concurrent writes to the same file, and inconsistent index state are all real failure modes when two agents happen to touch overlapping paths.
Worktrees give each agent a completely separate working copy of the repository in its own directory, with its own branch. Agent 01 works in ../migration-agent-01 on branch feat/migration-jsx-01. Agent 02 works in ../migration-agent-02 on feat/migration-jsx-02. They share the same Git history and avoid editing the same checkout by default. Process permissions can still allow access to other worktrees.

You can create them manually:
git worktree add ../migration-agent-01 -b feat/migration-jsx-01
git worktree add ../migration-agent-02 -b feat/migration-jsx-02
# ... repeat for each agent
Or let Claude Code handle it automatically with isolation: worktree in the agent frontmatter:
---
name: migration-agent
model: claude-sonnet-5
isolation: worktree
allowed-tools:
- Read
- Edit
- Write
- Bash
---
You will receive: brief.md (conventions), a list of files, and a quality gate command.
Apply the migration to your file list following brief.md exactly.
Run the quality gate command. If it passes, exit cleanly.
If it fails, write exception-report.md explaining what blocked you and exit.
Do not touch files outside your assigned list.
The isolation: worktree option creates a fresh worktree per agent invocation and cleans it up on exit if no changes were made. If the agent makes changes, the worktree path and branch name come back in the result so the coordinator knows where to find the work.
Worktrees separate normal checkout operations. They leave network calls, credentials in scope, and files outside the migration target reachable from inside the worktree. For a JSX to TSX pass on already-versioned code, that’s usually fine. For a migration touching anything closer to secrets, deploy scripts, or external calls, add the isolation layer that shows up independently across three separate practitioner communities: an ephemeral container, read-only mount outside the target files, a network allowlist, and a hard time quota. Worktrees keep git clean. The additional layer keeps the blast radius small.
The /batch command
/batch here is a skill you author yourself, not a built-in Claude Code command; it orchestrates Phase 2. It doesn’t take flags: you give it a plain-language instruction, it proposes a decomposition into independent units (typically 5 to 30), and nothing runs until you validate the plan. Three things belong in that instruction.
The parallelism you want. For a 243-file migration, around 24 agents at 10 files each is a reasonable split. The right number depends on the independence of your files and the time each agent takes: if your quality gate is fast (under 30 seconds), you can push toward 30 agents, and if it’s slow (tsc on a large project can take 2-3 minutes), fewer agents means less wasted wall time when one fails early.
The quality gate, as the exact command each unit must pass. For the JSX to TSX migration: tsc --noEmit --project tsconfig.json.
And the brief produced in Phase 1, so every agent reads the conventions before touching its assigned files.
A complete invocation for the JSX to TSX case:
/batch migrate every *.jsx file under src/ to TSX, following the conventions
in .claude/migration/brief.md. Split the work across roughly 24 parallel
agents. Each unit must pass "tsc --noEmit --project tsconfig.json" before
it counts as done.
What /batch doesn’t handle automatically: detecting cross-file dependencies that should have blocked decomposition, resolving conflicts between ready branches that touched the same shared file, and reviewing the exceptions. Those three points stay human. The ready branches merge cleanly precisely because the decomposition step ensured they’re independent. If something slips through and two branches conflict, that’s a signal the decomposition had a gap.
The merge itself should stay boring: merge the ready branches sequentially into the integration branch and re-run the quality gate after each merge, stopping at the first failure. A branch that passed its gate in isolation can still break the build once its neighbors have landed; the serialized re-run is what catches those interaction effects. Resist the urge to octopus-merge 24 branches in one shot, because when that build fails you’re back to bisecting by hand.
Characterization tests
For behavioral migrations, characterization tests make Phase 2 safe. They compare the output after migration with the output before it, rather than verifying correctness. If test-first workflows with Claude are new territory, the TDD guide covers the discipline this builds on.
The prompt to generate them:
For [filename], capture its current output for [these inputs] as tests.
Do not assess whether the output is correct. Do not optimize anything.
Capture only what the function currently returns for the given inputs,
so we can verify the migration produces the same results.
For Python numerical migrations (Pandas, NumPy), pd.testing.assert_frame_equal with rtol=1e-5 handles floating-point tolerance without false negatives. For Django view migrations, assertContains on the response body catches behavior drift without parsing HTML.
The test does not prove the original code was right, or that every behavior is equivalent. It compares the outputs for the chosen inputs and tolerances. Keep review and checks for properties those cases do not exercise.
Exceptions in the illustrative example
In the illustrative 243-file example, three files trigger exception reports. These are illustrative outcomes, not an observed exception rate. What distinguished them from the 240 that passed isn’t interesting in general terms, but the pattern holds across migrations: exceptions are files with dependencies outside the agent’s assigned scope, patterns the brief didn’t cover, or circular imports the quality gate can’t resolve in isolation.
The coordinator agent produces an exception queue: a list of branch names with their exception-report.md content. Each report contains the file that failed, the exact quality gate output, and a short description of what the agent couldn’t resolve. A human reviewer picks up the queue and handles each one with that context already loaded.
The three illustrative JSX to TSX exceptions have different blockers. One had a circular import that tsc flagged only when combined with its dependent. One used a third-party type that hadn’t been updated to support JSX transforms correctly. One depended on a utility in src/legacy/ that the brief hadn’t covered because Phase 1 hadn’t indexed that directory.
Resolving each exception requires investigating its specific cause; no measured resolution duration is attached to this illustrative example. The alternative was having all three problems scattered invisibly through a sequential session, discovered at code review or after merge.
What makes Phase 1 worth the time
Planning is useful when it exposes dependencies and exceptions before execution starts. Its duration depends on repository scope, tooling and the checks required; this guide provides no calibrated estimate.
Phase 1 should inspect directories such as src/legacy/, document import patterns, and flag dependencies that make the split unsafe. Review the brief for missed cases; choosing a planning model does not guarantee that it finds every exception.
The brief is the specification for the bounded assignments. Check its file coverage, conventions and gate commands before dispatch, then use the coordinator to handle integration and exceptions.
Measure input, output, cache use and retries on a representative unit before setting a run budget. Multiply only with explicit assumptions about the remaining units and verify current provider prices. This illustrative example supplies neither token telemetry nor a dollar estimate.
The pattern generalizes past migrations. Stefan Negouai and Sébastien Drode’s team, presenting at Dev With AI, moved their review discipline up a level for the same reason: review the plan, not the twenty-four diffs it produces. At the industrial end of the spectrum, Jean-Louis Rigau and Emmanuel Sciara’s Gastown orchestrator and Hubert Grégoire’s Agentic Factory are the same two-phase idea with permanent infrastructure around it: a durable git-based work registry, specialized agents, and an LLM judge at the merge gate. The guide’s agent teams workflow covers the broader orchestration patterns this two-phase model is one instance of.
Write the brief before touching any files. When it cannot drive an automatic quality gate, finish the brief before parallelizing the migration.
YSNK
(You should now know)
- Phase 1’s Opus planner must stay strictly read-only. The moment it touches files, it starts making micro-decisions that should be explicit brief conventions, and Phase 2 agents inherit an incomplete specification
isolation: worktreein the agent frontmatter creates a fresh worktree per invocation and cleans it up automatically if the agent made no changes. If it did, the path and branch name come back in the result- Worktrees separate normal checkout edits; they don’t sandbox what an agent can reach once running: network calls, credentials in scope, files outside the target. Migrations touching secrets or deploy scripts need an added layer: ephemeral container, read-only mount, network allowlist, time quota
- Merge ready branches sequentially and re-run the quality gate after each one, not as one octopus merge. A branch that passed its gate in isolation can still break the build once its neighbors have landed
- A characterization test compares selected pre- and post-migration behavior. It does not prove whole-program equivalence or remove the need to verify integration
Go Further in the Claude Code Guide
Practical resources selected to help you take the next step.
Open-source galaxy
Projects used in this path
Why this matters
The research and reasoning behind this playbook.
AI velocity is bidirectional
Everyone talks about shipping 10x faster with AI. Nobody talks about accumulating debt 10x faster. 7 months of production data from a real EdTech platform.
Claude is my second contributor: what real Git stats show
6 human commit authors plus Claude in our git history (May 2026 snapshot). What the commit patterns actually look like after months of Claude Code, beyond the marketing claims.
From afterthought to infrastructure: how AI config evolves in a real project
Nine months of AI configuration in one production project: evolving responsibilities, profile-based generation and a measured reduction in always-on lines.
Related guides
Claude Code security: the attack surface few teams audit
Hooks are shell scripts with your user permissions. MCP servers are third-party code with access to your credentials. Their timing and access depend on the configured events and server.
Claude Code setup, level by level
Three configuration layers for project context, daily tools and persistent memory, with checks for what loads and how it behaves.
Context engineering: the L0-to-L5 playbook
Choose context controls from L0 to L5 according to the failure you observe, from project documentation to scoped rules, behavior checks and shared configuration.