MLGuerrillaStart with M1 →
Free · in beta·intermediate·M17·45 min read·Prereq: Harness & Reliability (M9), whose stop conditions this module builds on. Verification (M15) supplies the checks a plan finishes on, and Memory (M16) supplies the coding agent, its repo and the note that starts the running example.

Planning

The capability

What a planner is

A planner is the part of an agent that works out, before anything runs, what has to be true when the task is finished and which steps should get there, writes that down in a form your code can read, and changes it when a step shows the world is different from what it assumed.

Take the pieces of that sentence one at a time, because each one turns into code you write.

  • What has to be true when it's finished. A goal stated as checks your code can run at the end, like "the checkout test passes and the test file wasn't edited".
  • Which steps should get there. A list of steps, each one a single tool call with an expected result, and a note of which steps need another step's output before they can start.
  • A form your code can read. JSON with a fixed schema, so the harness can run the steps in order, see which ones are done and count what they cost.
  • Changes it when a step shows otherwise. A replanning rule that decides, when a step fails, whether to retry it, change the rest of the plan or stop and ask a person.
A diagram of one planning loop in two rows, with a strip underneath. The top row runs left to right. A card holds the request, the checkout test is failing on CI, fix it, and the starting state, which is the repo and git status, a memory note dated 2026-09-02, permissions to edit src and no push, and a budget of 16 tool calls. An arrow leads to the planner, one model call that writes the goal as checks, lists its assumptions and orders the steps. An arrow leads to the plan, JSON your code reads, holding the goal that the checkout test passes, done-when checks that pnpm test:checkout exits 0 and the test file is unchanged, the assumption that the fixtures are out of sync, and a small graph of steps, s1 reset, then s2 rerun, then more. The bottom row runs right to left. The executor and state card fills in one ready step's arguments, calls the tool, checks the result against expect and records what each step returned and what changed. An arrow leads to the done-when checks, run by the harness in code, where the model's own done doesn't count. An arrow leads to a terracotta card, finished, with status finish, reached only when every check passes. A dashed terracotta strip labelled replanner says that when a step or a check fails, the loop retries a transient error, keeps a lookup miss, or names the broken assumption and writes new remaining steps from the current state, or stops and asks, capped at 3 replans.
The replanner is the only way back into the plan, and it reads the whole state, including what has already changed in the repo.

A planner has limits worth knowing before you build one. It can't know what a file contains until something reads it, so any plan written before the first read is partly a guess, and the steps that depend on what the read finds have to stay vague or get rewritten later. It also doesn't make a model better at the task. When Valmeekam and colleagues asked models to write complete plans for problems like those in the International Planning Competition, the best model, GPT-4, had "an average success rate of ~12% across the domains" (Valmeekam et al., NeurIPS 2023). The same paper found that plans did much better when an external verifier checked them and sent the errors back, and Kambhampati's group later argued that models "cannot, by themselves, do planning or self-verification" and should be paired with checkers outside the model (Kambhampati et al., 2024). Those were 2023 models on puzzle domains and newer ones do better, but the design that follows from it still holds for an agent in a repo. The model proposes a plan, and code checks each step and the finish.

Why this is a capability of its own

M9 already gave the agent a loop with limits in code, like a cap on model calls and a stop when the same tool call repeats, and M15 built checks that decide whether an output can be used. It's fair to ask what's left. What's left is the order of the work and the definition of finished. M9's loop decides when to give up, and M15's checks judge one output, and neither says which step should come next, which steps can run at once, or what the agent should do when step three proves step one was built on a wrong guess.

Those turn out to be where agents fail most. Cemri and colleagues labelled 1,642 traces from seven multi-agent frameworks and found step repetition in 15.7% of them, agents not recognising that the task was finished in 12.4%, and agents failing to follow the task's requirements in 11.8% (Cemri et al., 2025). The paper also says that failing to recognise termination conditions appears "almost exclusively in failed runs". Each of those is a planning failure, since each is about which steps to take or about when the task counts as finished.

Where it shows up

  • A coding agent's task list, the checklist it writes at the start of a bigger change and ticks off as it goes.
  • A research agent that decides which searches to run, which ones can run at the same time, and when it has enough to write the report.
  • A data backfill where an agent works out which tables to rebuild and in what order, because some depend on others.
  • The planning step M10 added to the flower company's support agent, which read only the customer's messages and fixed which tools the turn was allowed to use.

The shapes are different, and all of them keep a plan somewhere your code can read and check each step against it.

Where this starts

The checkout test the fixtures didn't fix

The coding agent from M16 works on a small team's web-app repo and remembers facts about the project between sessions. The team and its repo are made up, and so is every file in it. For this module the repo is also simulated, a copy held in a Python script with a fake test runner and a fake git history, so the same task can be run many times and the final state checked in code. The script is research/M17-planning-loop.py, and it runs on one local model, qwen3:8b through Ollama at temperature 0.

A developer writes "The checkout test is failing on CI. Fix it." The read policy from M16 loads one note for tasks about tests, saved on September 2nd, which says the checkout test is flaky when the fixtures get out of sync and that running pnpm fixtures:reset and rerunning usually fixes it.

The note is wrong this time. The day before, someone committed a change to src/pricing.js that rounds the tax on each line item down, with Math.floor, where it used to round to the nearest cent. On the test's three-item order that's two cents, and the test fails with "expected 107.07 to be 107.09". The commit message says the change was made at finance's request, and that the request was withdrawn the next day.

The planner, one model call that writes the whole plan before anything runs, turned the note into the plan below. The only edits are shorter wording in the does fields.

the plan qwen3:8b wrote for the checkout task
{
  "goal": "Fix the failing checkout test on CI by ensuring fixtures are in sync.",
  "assumptions": [
    "The checkout test failure is due to out-of-sync fixtures.",
    "Running `pnpm fixtures:reset` will restore the fixtures to a known state."
  ],
  "done_when": [{"check": "command_passes", "cmd": "pnpm test"}],
  "steps": [
    {"id": "s1", "does": "Reset the fixtures", "tool": "run",
     "depends_on": [], "expect": "The fixtures are reset successfully."},
    {"id": "s2", "does": "Re-run the test suite", "tool": "run",
     "depends_on": ["s1"], "expect": "The test suite runs with no failures."}
  ]
}

Two things in that plan are fine and two aren't. Testing the memory note first is a sensible opening, since it costs one command, and the planner wrote the note down as an assumption, which means a failure can point at it. What's missing is any step for when the note is wrong, and any check that stops the agent from finishing by changing the test's expected value. The only done-when check is that pnpm test passes.

The reset ran and the checkout test failed again with the same two cents. What happened next depended on which loop was running the plan.

  • With no replanning, the executor had no steps left, so the harness wrote a report. The report explained the two-cent gap and suggested checking the fixtures and the tax logic, and the run ended with status finish and the test still failing.
  • With replanning, the replanner named the broken assumption correctly, "running pnpm fixtures:reset will fully resolve the test failure", and then wrote new steps that read and edited src/checkout.js, a file that doesn't exist. The edit failed, and on its second replan it stopped and asked the developer whether checkout.js was missing.
  • One action at a time, the agent reset the fixtures and ran the tests in one command, saw the failure, searched for calcTotal, read src/pricing.js, saw Math.floor in the tax function, replaced it with rounding to two decimals, and ran the tests again. They passed, and the run finished in seven model calls.

So on the module's own running example, the simplest agent was the only one that fixed it. That's worth sitting with before building anything, because it's what a planning loop has to beat. The replanning loop did the part planning is supposed to do, naming the assumption the failure disproved, and then failed on something planning adds, which is a step whose arguments were written before anything had been read. Section 13 goes through the plan that should have come out of that failure, and section 14 measures all four approaches on more tasks.

When to plan

A plan is worth writing only when the path isn't known in advance

Most requests an agent gets don't need a plan. "Which pnpm script runs the checkout tests?" is answered by reading package.json, and any structure on top of that one read is time the developer spends waiting. So the first design decision is which of three shapes a task gets.

  • A single action. One tool call, then an answer. The model picks the call and your code runs it.
  • A fixed workflow. Your code decides the steps, and the model fills in each one. Anthropic's guide calls these "systems where LLMs and tools are orchestrated through predefined code paths" (Anthropic, December 2024).
  • A dynamic plan. The model writes the steps at run time, because nobody could write them in advance. The same guide says agents fit "open-ended problems where it's difficult or impossible to predict the required number of steps, and where you can't hardcode a fixed path."

A fixed workflow is often the right answer for coding work, which surprises people. Agentless fixes GitHub issues with the same fixed phases every time, from finding the code through to validating a patch, "without letting the LLM decide future actions". It resolved 96 of the 300 issues in SWE-bench Lite, 32.00%, at an average of $0.70 per issue on GPT-4o, which was the best result among the open-source agents it was compared against in mid-2024 (Xia et al., 2024). Closed tools in the same table scored higher, so it isn't the best possible, and the point stands that a fixed path your code controls can match an agent that plans for itself.

Which shape a coding-agent request gets
RequestShapeWhy
"Which script runs the checkout tests?"Single actionOne read answers it
"Change the free-shipping threshold to 60"Single actionOne edit, and the tests catch a mistake
"Run lint and the tests before every deploy"Fixed workflowThe same steps every time, so write them in code
"Rename calcTotal everywhere and keep the tests green"Dynamic planNobody knows which files use it until a search runs
"The checkout test is failing on CI, fix it"Dynamic planThe cause is unknown, so the steps depend on what each one finds

What planning costs on a task that doesn't need it

Section 14's evaluation includes four single-action tasks, like asking which currency src/config.js uses as the default, and the cost of planning shows up plainly in them. The direct agent picks one tool call and then writes the answer, and it used 2 model calls and about 4 seconds of model time on each. The planning loop with replanning used 5 calls and 17 seconds on average, because on top of a call per step to fill in arguments it spends one call writing the plan and another writing the report.

The currency question shows it most plainly. The planner wrote a one-step plan, read src/config.js, and added a done-when check that the file contains the text defaultCurrency. The file says CURRENCY = "EUR", so the check failed, the replanner ran twice looking for a name that was never there, and the answer, EUR, came back after 7 calls and 25 seconds, where the direct agent gave the same answer in 2 calls and 4 seconds. A planner that writes a check can write a wrong one, and a wrong check makes the loop do extra work on a task that was already finished.

Planning did get one of the four right that the direct agent missed. Asked which pnpm script runs only the checkout tests, the direct agent ran pnpm test -- --grep "checkout", saw one test pass, and answered with that command. The plan read package.json first and answered test:checkout, which is the script. That's the trade in miniature. A plan buys a read before the action, and on most small tasks the read isn't needed.

The practical fix is a gate, one small model call that decides whether a request gets a single action or a plan, and it's only worth adding if it chooses well. Across the 14 tasks in section 14 the script's gate picked wrong on three. It sent "do the checkout tests pass right now?" to the planner, and it sent the moved README and the checkout fix to the direct path, the second one reasoning that "resetting them and rerunning the tests should fix the problem", which is the memory note talking. The gated agent solved 8 of 14, one fewer than planning everything, at about the same number of calls, so this gate didn't pay for itself. A gate is a classifier, and M13's advice applies, which is to measure it on labelled requests before you trust its choices.

Goals and completion criteria

Turn the request into a goal your code can check

"Fix the checkout test" sounds like a goal and is too vague to plan against, because it doesn't say what finished looks like, and an agent left to decide for itself has a cheap way to finish. It can change the number the test expects. The test then passes and the agent reports success, while the tax bug the test was written to catch goes out to customers.

So the planner's first job is to write the goal down as checks. For the checkout task they'd look like this.

"The checkout test is failing on CI, fix it", written down
PartWhat it says for this task
GoalThe checkout test passes on the current code
Done whenpnpm test:checkout exits 0, and pnpm test exits 0
Constrainttests/checkout.test.js is unchanged, so the fix is in the code under test
PermissionsMay read anything, edit files under src/, and run pnpm scripts. May not push, and may not touch .env files
AssumptionThe failure is the flaky fixture the memory note describes
Question to ask firstNone. Everything the task needs is in the repo

Each row in that table does a different job. The done-when checks are what the harness runs at the end in code, and nothing else ends the task as a success. Constraints are checks too, and in this task the constraint is what stops an agent from finishing by changing the test's expected value. Permissions come from your code and the agent's token, and the planner only gets to read them, since a plan that lists git push as a step still can't push if the token is read-only. Assumptions are the guesses the plan rests on, written where a failing step can point at them later, and section 11 depends on that.

The last row is the one people skip. Some requests can't be planned without an answer from a person, like "deploy the redesign" when two staging branches exist, and some can be planned on a stated assumption, like "tests means the unit tests" when the repo has both unit and end-to-end suites. The rule of thumb is to ask when being wrong would be expensive or hard to undo, and to write the assumption down and carry on when it would be cheap to find out and fix. M16 used the same rule for memory, where an old fact was fine for a commit message and needed a confirmation before a deploy. An assumption in the plan shows up in the final report, so the developer sees it even when nobody was asked.

The checks the planner can write come from a small closed list your harness knows how to run, which is what makes them checkable.

the done-when checks the harness can run
[
  {"check": "command_passes", "cmd": "pnpm test:checkout"},
  {"check": "command_passes", "cmd": "pnpm test"},
  {"check": "file_unchanged", "path": "tests/checkout.test.js"},
  {"check": "file_contains", "path": "docs/README.md", "text": "Node 20"},
  {"check": "text_absent", "text": "calcTotal"}
]

Some goals can't be written as a command or a file check, like "the summary is accurate" or "the new copy sounds friendly". Those end on one of M15's verifiers, which might be a model judge or a person, and the plan records which one, because a task with no check at the end is a task the agent can finish by saying it did.

Starting state and memory

What the planner knows before step one

A plan is written from whatever the planner can see at the moment it's called, and that's usually less than you'd think. For the coding agent the starting state is the request and the repo as it stands, meaning its branch and git status, plus any CI log. On top of that come the memories the read policy from M16 loaded for this kind of task, and the permissions and budget the harness gave the run. Everything else the plan needs has to come from steps.

Memory needs care here, because it arrives looking like knowledge. The note the coding agent loaded, "the checkout test is flaky when the fixtures get out of sync", was true of some earlier failure, and M16's rule was that memory is a default and never an override. In a plan that means a memory becomes a dated assumption, and the first steps should be the cheap ones that test it. A planner that treats the note as a fact writes a plan with no way back if it's wrong, and every step after the fixture reset depends on something nobody checked.

Once the run starts there is a second kind of state, which is everything the run has learned and changed so far. It holds each step with its status and result, along with the assumptions still standing. It keeps a list of changes, which covers files edited and commands with side effects, and it counts how much of the budget has gone. Section 10 shows it as a structure your code owns, saved after every step the way M6 saved a checkpoint after each step of a refund, so a crash in the middle of a plan resumes from the last finished step and doesn't re-run an edit or a migration. The conversation list M5 managed is a bad place to keep it, because the model has to re-read all of it to work out where it is, and it gets compacted when it grows.

Step schema and granularity

A step your executor can run

Each step in the plan is a small record with fixed fields, and every field is there because some part of the harness reads it.

one step in the plan
{
  "id": "s3",
  "does": "Find which commit last changed the tax rounding",
  "tool": "git_log",
  "args": {"path": "src/pricing.js"},
  "depends_on": ["s2"],
  "expect": "a list of commits, newest first",
  "side_effect": "none"
}

The id and depends_on fields are what the harness uses to decide what can run next. The tool must be one the agent is allowed to use, which the harness checks before running anything. The args can be empty in the plan and filled in by the executor when the step runs, because a step like "edit the rounding line" can't know the exact text of that line until an earlier read has returned it. The expect field says what a good result looks like, and the harness checks it in code wherever it can, for example with an exit code. The side_effect field says whether the step changes anything, from none for a read, to reversible for an edit in a working tree, to irreversible for a push or a migration against a shared database, and when something goes wrong later the harness treats each of them differently.

How big a step should be comes down to a trade. A step as coarse as "fix the bug" can't be checked, because nothing about its result tells your code whether it worked, and it can't be retried alone. A step as fine as "open the file", then "move to line 4", then "type one character", costs a model call each, and the plan goes stale while it's being run. The size that works is one tool call with a result your code can look at, which makes a step the smallest unit you can retry and check on its own.

The exception is steps that depend on something not yet known. "Edit the file the search finds" is a fine step to write before the search runs, and trying to be more specific only produces a guess. The replanner in section 2 did exactly that when it wrote an edit to src/checkout.js, a file that doesn't exist, before any step had read the repo.

Dependencies and parallelism

Order the steps by what each one needs

The depends_on field turns the list of steps into a graph. A step can start when every step it depends on has finished, and steps that don't depend on each other can run at the same time.

A deploy check that runs the linter alongside the unit tests and the checkout tests, then reports which fail, is the plainest example. The three commands need nothing from each other, so all three can start at once and the report waits for the last one. With each test run simulated at 20 seconds and the lint at the same, that's 20 seconds of tool time with the three running together, against 60 run one after another. Kim and colleagues built this into LLMCompiler, where a planner writes the graph and a scheduler runs every ready step in parallel, and reported "latency speedup of up to 3.7x, cost savings of up to 6.7x, and accuracy improvement of up to ~9% compared to ReAct" (Kim et al., ICML 2024). Those "up to" figures are the best results across their benchmarks, so a typical gain is smaller. The graph also has to be written, and on this deploy check qwen3:8b didn't write it. Section 14 shows every planning run collapsing the three commands into one step that ran only the test suite.

Two timelines for the deploy check, with the plan's graph underneath. The first, one after another, shows three grey 20-second bars in a row, s1 lint, s2 unit tests and s3 checkout, then a black report bar, for 60 seconds of tool time. The second, from the plan's graph, shows the same three bars in terracotta stacked so all start at zero, and the report starting at 20 seconds, for 20 seconds of tool time, on an axis marked 0, 20, 40 and 60 seconds. Underneath, three step cards each read depends_on empty, with arrows into a card that reads report, waits for s1, s2 and s3. The meta line says the graph is drawn by hand for task multi-4 and the tool times are simulated.
Tool times are simulated. The model calls that fill in each step's arguments can run in parallel too, which the script doesn't do.

Running steps at once has rules of its own. Two steps that write the same file can't run together, because one edit will be made against text the other already changed, so the harness serialises any two steps that touch the same path even when the plan says they're independent. Steps with an irreversible side effect run alone, after everything they could depend on, so a failure elsewhere in the plan can still stop them. And parallel calls count against the same rate limits and budgets as sequential ones, which M9 covered for model calls and which applies to your own tools as well.

Approaches compared

Four ways to plan

The approaches in common use differ in when the model decides the next step and how much it sees when it does.

One action at a time

The model looks at the task and everything that has happened so far and picks the next tool call, then sees the result and picks again. This is ReAct, which alternates reasoning with actions. It beat imitation and reinforcement learning methods "by an absolute success rate of 34% and 10%" on two interactive benchmarks with one or two examples in the prompt (Yao et al., 2022). It adapts to every result, because each decision is made with the latest observation in front of it. The cost is that the model re-reads the whole history on every call, and since no plan exists anywhere, your code has nothing to check the run against and no list of what's left.

The whole plan up front

A planner call writes every step before anything runs, and an executor works through them. Plan-and-Solve describes it as "first, devising a plan to divide the entire task into smaller subtasks, and then carrying out the subtasks according to the plan" (Wang et al., ACL 2023). ReWOO pushed it further by having the planner refer to results it hasn't seen yet as variables, which made the model calls much shorter, and reported "5x token efficiency and 4% accuracy improvement on HotpotQA" (Xu et al., 2023). Your code can check the plan and run its independent steps in parallel. The catch is that it's written before any results come back, so a wrong guess in step one carries through every later step.

A broad plan, refined as you go

The planner writes the plan and a replanner rewrites the remaining steps whenever a result changes the picture. Plan-and-Act measured the gain on WebArena-Lite, 165 web tasks, where adding replanning to their planner took the success rate from 43.63% with a static plan to 53.94% (Erdogan et al., 2025). Their reason is the same one this module keeps coming back to, that a plan fixed at the start "is static throughout execution, which makes it vulnerable to unexpected variations in the environment." This is the approach the rest of the module builds.

Several candidate plans

The model writes more than one plan, or more than one next step, and a scorer picks one. When a branch fails, the search backs up and tries another. Tree of Thoughts did this on the Game of 24 puzzle, where GPT-4 with chain-of-thought prompting "only solved 4% of tasks" and the tree search solved 74% (Yao et al., NeurIPS 2023). It helps when a wrong early choice is expensive and a candidate can be scored cheaply before it runs. For a coding agent the cheapest scorer is often the test suite, which makes candidates practical for patches and less so for whole plans, since a plan's quality usually only shows once it runs.

The four approaches for a coding agent
ApproachModel callsWhen it worksHow it fails
One action at a timeOne per step, each re-reading the historyShort tasks where every result changes the next moveLoops, or stops without knowing whether it's done
Whole plan up frontOne to plan, then one per step to fill argumentsTasks whose steps are knowable before the first resultA wrong assumption carries through every later step
Broad plan, refinedThe up-front calls plus one per replanTasks where results change the plan now and thenReplans in circles unless something caps it
Candidate plansSeveral times the planning calls, plus scoringA wrong first choice is costly and candidates are cheap to scorePays for plans it throws away

The planning loop in section 14 runs the first three of these side by side, plus the single action, on the same tasks. It doesn't run candidate plans.

Uncertainty and gathering information

Plan the cheap reads before the expensive guesses

A plan for "fix the checkout test" is a plan written in ignorance, since nobody knows yet why the test fails. What the planner can do is put the steps that reduce that ignorance first, and put the cheapest of them before the dearest.

That means treating each possible cause as a hypothesis with a step that could prove it wrong. The fixture note is one hypothesis, and testing it costs one reset and one rerun. A recent change to the pricing code is another, and it costs reading the failing assertion and one git log. A problem with the CI machine is a third, and testing it means asking someone with access. So the plan tests them cheapest first, and the steps that change code come after a hypothesis has survived its test.

The failing test's own output is the first thing to read. "Expected 107.07 to be 107.09" is two cents on a three-item order, which fits a change in how each line's tax gets rounded. Stale fixtures would more likely change the order itself, and move the total by more than a cent or two. A planner that reads that line before choosing a hypothesis would have ranked the fixture note lower.

Some steps exist only to gather information, and their failure is information too. A read of a file that doesn't exist tells you the file doesn't exist, and a search with no hits tells you the name isn't used. The first version of the loop in section 14 treated those as failed steps and sent them to the replanner, which on the checkout task gave up straight after reading a file it had guessed at. The second version keeps a failed lookup as an ordinary result and goes on to the next step. That change is right, and on the checkout task it wasn't enough, because the next step was an edit to the same guessed file, which section 14 goes through.

The interface

The planner and the executor share one state object

The planner and the executor are usually two different prompts, and sometimes two different models. They talk to each other only through the state your code keeps. Pinning down that state is most of the work of building a planning loop.

plan_state.py — what the planner, executor and harness all read
@dataclass
class Step:
    id: str
    does: str
    tool: str
    depends_on: list[str]
    expect: str
    side_effect: str = "none"      # none | reversible | irreversible
    args: dict | None = None       # filled by the executor when it runs
    status: str = "pending"        # pending | done | failed | skipped
    result: str | None = None

@dataclass
class PlanState:
    goal: str
    done_when: list[dict]          # checks from the closed list, run in code
    assumptions: list[str]         # each one a guess a failure can point at
    steps: list[Step]
    changes: list[str]             # files edited, commands with side effects run
    replans: int = 0
    tool_calls: int = 0

The executor gets one step at a time along with the results of the steps before it, and returns a tool call for that step and nothing else. It isn't allowed to add steps or to decide the task is finished, because those are the replanner's and the harness's jobs, and an executor that can do both is a one-action-at-a-time agent with a plan it can ignore. When it wants to do something the plan doesn't have, the right move is to fail the step and let the replanner decide.

After every step the harness checks the result against the step's expect in code wherever it can. An exit code or an edit whose old text was found settles it without a model. Where it can't, a step can end on one of M15's verifiers, and that's the exception for a coding agent, since most of what it does leaves something code can inspect.

At the end the harness runs the done_when checks itself. The model's final message doesn't count, and neither does the executor saying the last step went fine. If any check fails, the harness treats it like a failed step and sends it to the replanner with the check's output attached. This is the one place in the loop where a verifier and a planner meet, and it's why the checks had to be written in section 4, before any work started.

Replanning and recovery

When the plan meets the repo

Every plan meets something it didn't expect, and what the loop does next decides whether the agent is any good. The first question is what kind of failure it was, because each kind calls for a different response, and sending all of them to the replanner is how agents end up rewriting a perfectly good plan because of a network blip.

Five kinds of failure and what the loop does with each
KindWhat it looks likeWhat the loop does
TransientECONNRESET from the package registry, a 503, a rate limitRetry the same step once with the same arguments, as M9's retry layer does
Lookup missA read of a file that isn't there, a search with no hitsKeep it as a result and move on, since it answers a question
Bad argumentsAn edit whose old text isn't in the fileLet the executor fill the step again from the file's current text
InvalidatingThe fixtures reset and the test still failsThe replanner names the assumption that broke and rewrites the remaining steps
BlockedA missing secret, a permission the token doesn't haveStop and ask, saying what was tried and what's needed

Telling them apart is mostly your code's job. Transient errors match known patterns and are the same ones M9 retried. A lookup miss is any failed read-only step. What's left goes to the replanner, which is asked to say which kind it is, and to name the assumption the failure proved wrong when it says invalidating. That naming is what keeps replanning honest, because a replanner that can't name a broken assumption usually doesn't have a new plan, only the old one reworded.

Replan from where you are

A replan starts from the current state of the repo. Some steps already ran and some of them changed things, so the replanner gets the list of changes along with the results. Reads cost nothing to keep. A reversible edit that turned out wrong gets an explicit revert step in the new plan, so the undo is recorded like everything else. An irreversible step that ran, like a migration applied to a shared database, stays applied. The new plan works forward from it, because running the migration again would fail or apply something twice.

The migration is why the idempotency keys from M6 belong in a planning loop. A step whose tool accepts a key, so a repeat is recognised and ignored, is safe for the replanner to schedule again. A step without one has to be checked first, by asking the system whether it already happened, the way M6 checked a refund's status before retrying it.

Don't let it replan in circles

A replanner can loop as easily as a single-step agent can. It tries the fixture reset and the test fails, then it proposes a slightly different fixture reset and the test fails again. The script puts two limits in code. It caps replans at three, and it stops when the same tool call with the same arguments fails twice, which is M9's stopped_repeated_call applied to plan steps. A third rule is worth adding, which the script doesn't have. A replan should have to name an assumption no earlier replan named, so the replanner can't blame the same guess twice and keep going. On the rename task in section 14 both replans named the same broken assumption, word for word, before the repeated-call stop ended the run.

Stop with something useful

Some tasks can't be finished by anything the agent is allowed to do. The payments test needs a sandbox key the agent doesn't have and shouldn't have, and a push to preview needs a permission the agent's token doesn't carry. The right ending for those is a stop, and a stop is only useful if the message says what was tried and what changed, along with what's needed and who can provide it. "The payments test fails because PAYMENTS_SANDBOX_KEY isn't in the environment. I haven't changed any files. .env.example says to ask the payments team for it." is a message a developer can act on in a minute.

The wrong endings are the ones where the agent keeps going. It can mark the test as skipped, which is the payments version of editing the checkout test's expected value, or it can retry the push with different flags. The done-when checks catch the first one, since a skipped test fails file_unchanged, and M9's limits catch the second.

Budgets and termination

Budgets decide how the run ends

M9 put limits on every run, like a cap on model calls and a deadline, and those still apply. A planning loop needs a few more, because every plan and replan is an extra model call, and a plan can be fine at every step and still cost more than the task is worth.

The most useful budget is the total number of tool calls, 16 per task in the script, which covers retried calls and replanned steps together. The replan cap sits inside it. A budget per step helps on tasks with slow tools, since a test run that hangs should fail its step at a deadline and let the replanner see a timeout. And the cost of planning itself should be counted separately in your traces, so you can see how much of a task's spend went on deciding what to do.

Every run ends in exactly one of four ways, and the status your code records should say which.

How a planning run ends
StatusConditionWhat the developer sees
finishEvery done-when check passed in codeThe result, with the assumptions that were made
ask_userThe replanner stopped on something only a person can provideWhat was tried and the question
stoppedA limit in code ended itThe same report, plus which limit fired
failedAn error in the harness itselfAn error logged with the trace, never shown as success

A run that reaches stopped is the budget doing its job. The runs to worry about finish with checks that should have failed, or stop with a message nobody can act on.

The running example, end to end

The checkout fix, planned and replanned

Section 2 ended with the replanning loop correctly rejecting the fixture note and then editing a file that doesn't exist. What follows is the plan the loop should have produced, written by hand using every rule in this module. It was then run through the same simulated repo with no model involved, so every result below is what the tools returned (research/M17-reference-plan.py, output saved next to it). Nothing here came from qwen3:8b.

The first plan tests the note and nothing else

The first plan is almost the one the model wrote, with two changes. It runs only the checkout test, since that's the one the task is about and it's faster than the whole suite. And its done-when checks include the constraint the model left out.

plan v1, written before anything has run
{
  "goal": "The checkout test passes on the current code",
  "done_when": [
    {"check": "command_passes", "cmd": "pnpm test:checkout"},
    {"check": "command_passes", "cmd": "pnpm test"},
    {"check": "file_unchanged", "path": "tests/checkout.test.js"}
  ],
  "assumptions": ["A1: the fixtures are out of sync (memory note, 2026-09-02)"],
  "steps": [
    {"id": "s1", "does": "Reset the fixtures", "tool": "run",
     "args": {"cmd": "pnpm fixtures:reset"}, "depends_on": [],
     "expect": "exit 0", "side_effect": "reversible"},
    {"id": "s2", "does": "Rerun the checkout test", "tool": "run",
     "args": {"cmd": "pnpm test:checkout"}, "depends_on": ["s1"],
     "expect": "exit 0", "side_effect": "none"}
  ]
}

The plan stops after two steps on purpose. If the note is right, the test passes at s2 and the done-when checks finish the task. If it's wrong, nothing written now about the next steps would be more than a guess, so the planner leaves them to the replanner, which will have the test's output in front of it.

The rerun fails the same way, and that disproves A1

s1 prints "Fixtures reset (3 orders, 2 customers)." and s2 fails with "expected 107.07 to be 107.09" at line 7 of the test. That's an invalidating failure in section 11's terms, because it's the exact result assumption A1 said wouldn't happen, so the loop sends it to the replanner along with the state. The replanner sees the goal and the checks along with both steps and their outputs. It also sees the one change with a side effect so far, which is that the fixtures were reset. Resetting them is harmless to keep, so the new plan doesn't undo it.

The revised plan reads before it edits

The replanner's job is to name what broke and write new steps from here. The rule from section 6 does the most work at this point. Any step whose arguments depend on something not yet read gets no arguments, and the executor fills them in when it runs.

plan v2, from the replanner
{
  "invalidated": "A1: resetting the fixtures didn't change the result",
  "assumptions": ["A2: a recent change to the pricing code moved the total by 2 cents"],
  "steps": [
    {"id": "r1.1", "does": "Read the failing test to see what it computes",
     "tool": "read", "args": {"path": "tests/checkout.test.js"}, "depends_on": []},
    {"id": "r1.2", "does": "List recent commits to the pricing code",
     "tool": "git_log", "args": {"path": "src/pricing.js"}, "depends_on": ["r1.1"]},
    {"id": "r1.3", "does": "Read the newest commit that could change the total",
     "tool": "git_show", "args": null, "depends_on": ["r1.2"]},
    {"id": "r1.4", "does": "Undo the rounding change if the commit explains it",
     "tool": "edit", "args": null, "depends_on": ["r1.3"], "side_effect": "reversible"},
    {"id": "r1.5", "does": "Rerun the checkout test", "tool": "run",
     "args": {"cmd": "pnpm test:checkout"}, "depends_on": ["r1.4"]},
    {"id": "r1.6", "does": "Run the whole suite", "tool": "run",
     "args": {"cmd": "pnpm test"}, "depends_on": ["r1.5"]}
  ]
}

The difference from what qwen3:8b wrote is in r1.3 and r1.4. The model's replan named src/checkout.js and wrote the exact edit, calcTotal to calcTotalWithTax, before any step had looked at a file, and both were guesses. Here those two steps have null arguments, and they'll get real ones from what r1.2 and r1.3 return.

The trace

The checkout task through plan v1 and plan v2, simulated repo
StepTool callWhat came backWhat the loop did
s1run pnpm fixtures:resetFixtures reset (3 orders, 2 customers)Recorded the reset as a change
s2run pnpm test:checkoutFAIL, expected 107.07 to be 107.09Invalidating, A1 rejected, replan 1
r1.1read tests/checkout.test.jsThe test expects calcTotal(FIXTURE_ORDER) to be 107.09Continued
r1.2git_log src/pricing.jsa41c9e2 2026-09-21, round tax per line item down, per financeContinued
r1.3git_show a41c9e2Math.round became Math.floor in lineTax, and the finance request was reverted on 2026-09-22Executor filled the edit from this diff
r1.4edit src/pricing.jsEdited, Math.floor back to Math.roundRecorded the edit as a change
r1.5run pnpm test:checkoutPASS, 1 passedContinued
r1.6run pnpm testPASS, 2 passedRan the done-when checks
checksall threetest:checkout passes, test passes, test file unchangedStatus finish

That's 8 tool calls and one replan. The report the developer gets says what was assumed, what disproved it, which commit caused the failure and why it was safe to undo, and what changed in the repo. That last part is the edit to src/pricing.js and the fixture reset.

The checkout fix as two plans. Plan v1, testing assumption A1, has two steps, s1 run reset the fixtures, which returned fixtures reset, and s2 run rerun the checkout test, marked failed with expected 107.07 to be 107.09. A terracotta arrow leads to a dashed terracotta replan card that reads A1 rejected, the fixtures weren't the cause, A2 proposed, a recent pricing change moved the total by 2 cents. Plan v2, testing A2, has six steps. r1.1 reads the failing test, which expects 107.09. r1.2 runs git log on src/pricing.js and finds a41c9e2, round tax down. r1.3, marked filled in terracotta because its arguments came from earlier results, runs git show on a41c9e2 and finds round changed to floor and the request reverted. r1.4, also marked filled, edits floor back to round. r1.5 runs the checkout test, which passes, and r1.6 runs the whole suite, 2 passed. A final card reads finish, with test:checkout passes, test passes and test file unchanged. A strip underneath says that at replan 1 qwen3:8b named A1 correctly, then wrote a read and an edit of src/checkout.js, a file that doesn't exist, with the edit's text already filled in, and that both failed and the second replan stopped to ask the developer. The caption reads 8 tool calls and one replan, and the edit's arguments came from the commit it undoes.
Compare r1.3 and r1.4 with the strip underneath, since the difference is whether the edit waited for a read.

Compare that with the one-action-at-a-time agent from section 2, which also got the test to pass in six tool calls. It replaced Math.floor with toFixed(2) without reading the commit, so it never learned the change was deliberate and then reverted, and its report couldn't say why the fix was safe. On this task the plan's extra structure bought a better report and a fix it could justify. On the simulated repo's checks the two runs look the same, which is a limit of the checks, and section 14 comes back to it.

Evaluation and the planning loop, run

Measure the planner against agents that don't plan

A planning loop is extra machinery, and the only way to know whether it pays for itself is to run it next to simpler agents on the same tasks, with identical tools and scoring. The script does that with four strategies.

  • Direct picks one tool call, and after it runs, writes the answer.
  • One action at a time picks the next tool call after every result, up to 12 calls.
  • Plan up front writes the plan and runs each step in order, with no replanning.
  • Plan and replan is the loop from sections 10 to 12, including the replanner and its limits. It ran in two versions, one that sent a failed lookup to the replanner and one that kept it as a result.

Every run is scored by code that inspects the final state of the simulated repo. A rename only counts if no file still mentions the old name and the tests pass, and the checkout fix only counts if the test passes and the test file is unchanged.

What ran

The script defines 14 tasks in four groups. Four are single-action tasks, and four have several dependent steps, like the rename and a deploy check that runs three commands. In four more the first reasonable approach fails. A README has moved, for example, and a package install hits a network error once. In the last two, a payments test that needs a secret and a push the agent's token can't make, nothing the agent is allowed to do can finish the job, so the right ending is a stop. Fourteen tasks on one 8B model is a sighting of how these loops behave, and nowhere near a benchmark.

14 tasks on qwen3:8b, with identical tools and scoring for every strategy
StrategySingle action (4)Multi-step (4)Recover (4)Stop (2)TotalModel calls per taskModel seconds per task
Direct30025 of 141.95
One action at a time333211 of 144.112
Plan up front41016 of 144.520
Plan and replan, lookups as failures42129 of 145.727
Plan and replan, lookups kept42129 of 145.925
Gate, then direct or plan and replan32128 of 145.621
Two horizontal bar charts side by side from 14 tasks run on qwen3:8b. The first, tasks solved out of 14, shows direct one call at 5 of 14, one action at a time at 11 of 14 in terracotta, plan up front at 6 of 14, plan and replan at 9 of 14, and gate then either at 8 of 14. The second, model seconds per task, shows direct at 5 seconds, one action at a time at 12 in terracotta, plan up front at 20, plan and replan at 25, and the gate at 21. A note says plan and replan is the version that keeps lookup misses as results, and that the version sending them to the replanner solved the same tasks. The caption says replanning added three tasks over a fixed plan, and the agent that didn't plan still solved the most. The meta line gives the model, temperature 0, the script, and that the repo is simulated and the tasks invented.
The two plan-and-replan versions solved the same tasks, so only one is drawn.

Two comparisons in that table are worth pulling apart. Replanning beat the fixed plan by three tasks, 9 against 6, and the traces show exactly which three. On the package install, the first pnpm install hit ECONNRESET, the loop retried it once as a transient failure, and the tests then ran and passed, where the fixed plan reported a network error and ended. On the rename of TAX_RATE, the plan edited only one of the two files that imported it, so the test run failed and the replanner added the edit to the other file. On the push, the fixed plan saw "Permission denied" and still ended with status finish, where the replanner stopped and asked for a token with write access. Those are the three mechanisms from sections 11 and 12 each doing its job once.

The other comparison is the one a decision rests on, and planning lost it. The one-action-at-a-time agent solved 11 of 14 in about half the model time of the replanning loop. That's a result about a small local model, and a stronger model on the planner would likely change it, which wasn't tested. It's also exactly the check you should run on your own system before adding a planner, because the baseline is cheap to build and it isn't a given that the planner beats it.

Where each failure came from

A success rate tells you whether to keep the planner, and it can't tell you what to fix. For that every failed run gets attributed to the component that caused it, by reading its trace. The planning loop's five failures, and the near-miss on the currency question, fall into three places.

The replanning loop's failed runs, by the component that caused them
ComponentWhat went wrongTask
PlannerCollapsed three independent commands into one step that ran pnpm test, so the linter never ranDeploy check
PlannerA step too coarse to run, "edit all files where calcTotal is used", done as one editRename calcTotal
PlannerA done-when check copied from the developer's wrong command, pnpm db:migrate must passMigrations
PlannerA wrong done-when check, defaultCurrency in a file that says CURRENCYCurrency question (solved, after 2 wasted replans)
ReplannerUndid the rename it was asked to make, then hit the repeated-call stopRename calcTotal
ReplannerWrote an edit to a guessed file before reading anythingCheckout fix
ExecutorAsked the developer after one missing file, with no search for where the README wentREADME moved

Most of those rows are the planner writing something it hadn't looked at yet, like a check against a name it guessed, a step that hid several edits, or an edit to a file it assumed. That points at the fixes. The executor should refuse a step that covers several edits and fail it back to the replanner, a planner-written check that references a file or name should be checked against the repo before the plan runs, and a replanned edit with arguments filled in should be rejected until a read of that file has happened.

The migrations row deserves a closer look, because the loop did the right work and still failed. The developer asked for pnpm db:migrate, which doesn't exist, and the planner turned the request into a check that pnpm db:migrate must pass. The replanner found pnpm migrate and ran it, and both migrations applied. Then the done-when check failed, the replanner was told not to weaken the checks, and it stopped. A check copied from the developer's words inherits the developer's mistake, and the fix is to write checks about the outcome the developer wants, here that migrations 004 and 005 are applied, and never about the command they happened to name.

The deploy check is the parallelism example from section 7, and no planner produced the graph drawn there. All three planning strategies wrote one step that ran the test suite, and none of them ran the linter, which was the command that failed. The one-action agent ran the tests and the linter and then went beyond the task, trying to fix the lint error it was only asked to report, and gave up when its guessed edit didn't match the file.

The lookup change made no difference to which tasks were solved. On the checkout fix, keeping the failed read of src/checkout.js as a result let the loop go one step further, and the next step was the edit to the same missing file, which failed as an edit. So the change is still right, since a missing file is an answer, and on its own it didn't fix anything.

The limits of the scoring

The stop tasks show one limit. Their scoring only asks that the run stopped without damage, and every replanning run passed it. Asked whether the stop message named the real blocker, the missing PAYMENTS_SANDBOX_KEY or the read-only token, the replanning loop named it on the push and missed it on the payments test, where it went looking for src/payments.test.js, which isn't where the test lives, and asked the developer to find the bug in the pricing code. The direct and one-action agents ran the tests and asked for the missing key. A stop is only as useful as its message, so a real evaluation scores the message too.

The checkout fix shows another. The one-action agent's toFixed(2) passes the simulated test, and in real JavaScript toFixed has floating-point quirks of its own, so it could round a different line wrong. The simulator can't see that, and the scoring gave it full marks. A real harness would run the real tests, and even real tests only check the inputs someone wrote down. That's M15's point about verification applied to planning, where passing the done-when checks tells you the plan met the checks you wrote.

Ask your AI coding tool

Build a planning-loop evaluation for our coding agent. Make a sandbox copy of a small repo with a scripted test runner, and write 12 to 16 tasks in four groups: single-action tasks, multi-step tasks with dependencies, tasks where the first reasonable approach fails, and tasks that must end with the agent stopping to ask. Each task gets a scoring function that checks the final repo state in code. Run every task under four strategies with the same tools: one tool call then an answer, one action at a time with a 12-call cap, a plan written up front with no replanning, and a plan with a replanner, a one-retry rule for transient errors and a cap of 3 replans and 16 tool calls. Log every model and tool call as a trace line, save the results after every run so a crash can resume, and print success and model time per group. Then list every failed run with the component that caused it: planner, executor, replanner, done-when check or gate.

Putting it together

Putting it together

A planner decides what finished means and which steps get there, writes both down in a form your code can run, and changes the steps when a result disproves what they assumed. Most of the work in building one is in the parts around the model call.

The coding agent's loop ends up looking like this. A gate decides whether a request needs a plan at all, and it gets measured like any classifier, because a gate that plans the wrong requests costs a call and buys nothing. When a plan is needed, the planner turns the request into done-when checks from a closed list the harness can run, the constraints that stop it finishing the cheap way, like leaving the test file unchanged, and a list of assumptions with dates, where a memory note goes in as a dated assumption. Its first steps are the cheap reads that test those assumptions, and any step that depends on something not yet read gets no arguments.

The executor runs one ready step at a time and fills in its arguments from what earlier steps returned. Steps with nothing between them run together unless they write the same file or have an irreversible side effect. After every step the harness checks the result in code, keeps a failed lookup as a result, retries a transient error once, and sends anything else to the replanner with the full state, including what has already changed. The replanner has to name the assumption that broke before it writes new steps from the current state, and it's allowed to stop and ask. The replan cap and the 16-call budget end the run with a report a person can act on, and so does the same call failing twice. Nothing ends as finish until the done-when checks pass in code.

Then the loop gets measured against agents that don't plan, on the same tasks, scored on the final state, with every failure attributed to the component that caused it. On 14 tasks with qwen3:8b, replanning solved three more than a fixed plan, each one traced to a single mechanism in the loop, and the one-action-at-a-time agent still solved the most. The traces put nearly every planning failure in the planner or the replanner, most of them from writing a step or a check before reading what it named. That's a plausible result for a first version on a small model, and the attribution is what tells you which of the rules above to enforce first.

M18 covers the tools the executor calls and how to describe them so a model uses them well. M30 combines planning with memory and verification in agents that run for hours, where the plan itself has to be saved and resumed.

Checkpoint · recall · 5 questions

What the module said

  1. 01

    What ends a planning run as a success in the loop this module builds?

  2. 02

    Which request is the best fit for a dynamic plan, where the model writes the steps at run time?

  3. 03

    What does the depends_on field in a step let the harness do?

  4. 04

    A step reads README.md and the file isn't there. What should the loop do with that?

  5. 05

    What did Plan-and-Act measure when it added replanning to a static planner on WebArena-Lite?

0 / 5 answered

Checkpoint · understanding · 5 questions

Reason it through

  1. 01

    The agent is asked "fix the checkout test" and the plan's only done-when check is pnpm test:checkout exits 0. Why is that not enough?

  2. 02

    Why does a memory note like "the checkout test is flaky when the fixtures get out of sync" go into the plan as an assumption?

  3. 03

    The loop sees ECONNRESET from the package registry on pnpm install. Why is sending it to the replanner the wrong response?

  4. 04

    Halfway through a plan, a step applied migration 004 to the shared database and the next step failed. What should the new plan do about migration 004?

  5. 05

    A team adds a planner to every request their coding agent handles. Simple questions get slower and nothing gets more accurate. What would fix that?

0 / 5 answered

Checkpoint · debugging · 4 questions

Debug it

  1. 01

    The checkout task ended with status finish and every done-when check passing. The diff changes toBe(107.09) to toBe(107.07) in tests/checkout.test.js. What went wrong?

  2. 02

    A run on the payments test hits the 16-call budget. The trace shows the replanner proposing "add PAYMENTS_SANDBOX_KEY to .env" three times, each time as a slightly different edit. What should change?

  3. 03

    On the rename task, after the plan edits src/pricing.js and src/cart.js, its test step fails with "does not provide an export named 'calcTotal'" from tests/checkout.test.js. Which kind of failure is it, and what should happen?

  4. 04

    On the 14 tasks in section 14, the one-action-at-a-time agent solved 11 and the replanning loop solved 9, and a teammate proposes dropping planning for good. What should the team look at before deciding?

0 / 4 answered

That's the last one written so far

Pick your next module from the board.

All modules →

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.