I named this series before I’d counted anything. “Loops all the way down” was a hunch. When I went through my factory’s code and data, there were more loops than I’d thought, and they were doing more of the work than I’d thought.
Here’s the surprise up front: over the last month, the loop that adds features merged 529 pull requests. The loop that simplifies existing code merged 617. Technical debt from a project that grew organically. Now 617 commits easier to read, easier to follow, and smaller to build and deploy.
The one rule
Every loop in Leaf follows one rule from its README:
Every change enters through one of four loops. One adds; three correct. A loop runs unattended only when a machine can tell that it is done.
That last sentence is the design test for everything. If a machine can tell the loop is done, the loop can run without a person. If it can’t, a person has to be in it. So the job isn’t to remove humans. It’s to make “done” something a machine can check, wherever you can.
Loops inside a plan
Start small. A single plan is already a stack of loops, each one closing on evidence a machine can read:
| Loop | Cycle | Done when |
|---|---|---|
| Worker | Edit, run the node’s check, fix | The check passes and the commit is pushed |
| Foreman | Review each diff, merge it, start the next wave | Every node has merged into the plan branch |
| Suite | Run the end-to-end tests, fix what fails, run again | The suite passes. A report with no tests in it counts as a failure |
| PR | Open it, wait for CI, merge | GitHub says it merged. Nothing else counts |
note: not every repo runs a full E2E suite
The whole Development path, from Leaf’s flow docs:
save → approve → start → foreman → spawn → pod → done → review → merge → suite → fix → PR → finish → teardown
And the suite loop nested inside it:
last merge → suite → fail → fix → merge → suite → … → pass
The foreman’s instruction for that inner loop is three words long: fix until green.
If the suite fails, spawn a fix agent and brief it: which tests failed, their errors, the trace and run-log links, and what you think is wrong. It fixes the code or the test, whichever is wrong. Review and merge its work like any node, then run the suite again. Repeat until it’s green.
No human gets paged when a test fails. A test failing is just the loop doing its job.
Here’s what that stack looks like on a real plan: a Sadie PRD that slimmed the app down to what my household actually uses. Two planner passes, then the foreman, then 17 workers on three different models across three machines, running in waves. Each row is one agent. The arrows are the foreman sending a message straight to a worker. 38 minutes from start to merged PR.

Loops that recover
Things die. Containers run out of memory. Hosts drop off the network. A foreman crashes halfway through a plan.
Leaf treats all of that as normal. A dead foreman gets replaced, and the new one reads the room log and the branch history before it does anything. A worker that disappears gets resumed with its saved conversation and any messages it missed.
The principle behind it:
Recovery should use the same paths as normal operation; a special recovery state machine is a second implementation that fails precisely when it is least exercised.
Recovery is a loop too. It’s just one you hope runs rarely.
The four big loops
Zoom out and the factory has four loops that decide what work happens at all.

Only one of them needs a person. The other three file their own issues, plan them, build them and merge them. 669 of the 1,198 merges this month had no human in them at all.
They also share the machines in a fixed order. Quality goes first, because nothing else moves while the factory is broken. Development comes next. Improvement gets whatever’s left. The entire system is also quota-aware for all of my AI subscriptions. I’m a boy on a budget.
Improvement: entropy, handled
The Improvement loop is fully automated. It’s designed to find existing code that can be simplified. It walks the codebase one directory at a time, and an examining agent gives each area a verdict:
Improvement counters entropy mechanically. It walks the code for duplication and dead weight and lands reductions unattended, because a check can prove that a reduction is smaller and that behaviour held.
The key phrase is “a check can prove.” Behaviour holding is easy: the tests have to keep passing. “Smaller” turned out to be harder. Leaf used to enforce it with a line-count gate (a reduction had to delete more lines than it added). That gate came out on September 29. Today a reduction has to survive the suite and the repository’s own stated intent. The README still describes the old gate.
Over three weeks it examined 506 areas:

The trend is the interesting part. In its first week, the sweep found something worth changing in 38 areas. The second week, 27. This week, 4, across 341 areas, a faster pace than ever. When the examiner says “nothing,” its reasons often point at a cleanup that had already merged. As far as the sweep can tell, the code is getting simpler and staying that way.
What it did to a real codebase
The best place to see this is my product monorepo. It’s the oldest code I have. It grew organically over a couple of years, through a lot of different harnesses and a lot of different processes, mine included. That’s exactly the kind of codebase that piles up dead weight.

That’s 18% less product code and 586 fewer files in 19 days, while the tests grew 12%. In the four months before Leaf, that code grew about 1%. My biggest manual cleanup weeks back then removed 12K to 14K lines. Improvement removed 31K to 37K lines a week, every week, and I didn’t schedule any of it.
Across every repo, Improvement PRs deleted 139K lines and added 24K. For every line the Development loop added, Improvement took about one away. Features shipped. The codebase didn’t grow.
Did the deletions break things? I traced every bug fix since May back to the commits that wrote the lines it changed. Improvement PRs were implicated in a later fix about 1% of the time. (There’s a blind spot: that method can’t see a bug caused by deleting something that was needed, because deleted lines can’t be traced. The Quality loop is the cross-check there. It traced two broken builds to merges all month, and neither was a reduction.)
What it doesn’t do yet
Look at that verdict chart again. Most of what Improvement found was unreachable code: 49 of the 69 areas where it acted. Dead code is worth deleting. It’s 105K fewer lines for the next agent to read and get led astray by, and fewer to build and ship. But dead code doesn’t throw bugs at runtime, because it never runs.
The harder target is code that’s over-engineered but works: three layers of abstraction around something that needed one. That’s the cleanup that has always fallen through the gaps, and agentic development is historically bad at it. Agents add. They rarely take away.
Leaf’s examiner can’t catch it as designed. Its overbuilt verdict means “a capability no caller and no test asks for.” That’s unused code. Code that’s used, works, and is far too complicated gets the verdict nothing.
That isn’t an oversight. It’s the one rule at work. A loop can only run unattended when a machine can tell it’s done, and “simpler” is hard for a machine to check. Fewer lines isn’t enough; I already tried a line-count gate. To go after over-engineered working code, the loop needs a new verdict and a real measure of done: the suite stays green and a complexity metric goes down. I think it’s possible. But I’m not there yet.
What it costs
It isn’t free. Improvement used about 23% of the factory’s tokens this month. But a merged reduction costs a median of 303K input-equivalent tokens, well under half of a typical feature. I’ll come back in Part 6 to whether that spend pays for itself.
Trajectory: the factory watching itself
This is the loop I’m proudest of, and the one that’s hardest to explain.
Every run Leaf does leaves a full record: every message, every tool call, every error. A reader agent goes through those records and notes friction: a command that failed, a check that couldn’t run, a pod that ran out of disk. Those notes get grouped into problems. When a problem shows up in at least three moments across at least two runs in a week, Factory files an issue about itself and works it like a bug.
Over the month, it tracked 256 problems. 84 stopped recurring. It merged 46 fixes, mostly to its own setup and its prompts. A couple of the real titles:
- “Sadie’s baseline dependency install fills the host disk with CUDA torch wheels”
- “Examination prompts exceed the codex input limit”
That second one is the factory noticing that one of its own loops (Improvement) was choking on its own history. Nobody told it to look.
Quality: the loudest thing Factory says
Quality keeps the machine moving. … It closes when the machine is observed moving again, not when a fix is claimed. It calls a person only when it cannot find the cause, and that call is the loudest thing Factory says.
“Observed moving again, not when a fix is claimed.” That line does a lot of work. A loop that closes on a claim will happily report success while the build stays red. A loop that closes on observation can’t.
Quality opened 15 fault rooms this month and merged 6 fixes. Two of its diagnoses traced a broken build back to a Development merge. The rest were flaky infrastructure, things that weren’t defects, or cases it couldn’t settle.
The nuance
A loop is only as honest as its “done.” The Improvement loop’s trend looks great, and I believe it. But “the examiner found nothing” and “there’s nothing to find” aren’t the same claim. A weaker examiner would produce the same chart.
The watcher can under-watch. The reader agent that feeds Trajectory is told that “nothing” is the usual answer. That keeps noise down. It also means real, quiet problems might never reach the three-moment threshold.
Quality only sees what breaks the build. A bug that ships and gets found by a user next week never enters the Quality loop. That’s a real gap, and I cover it in Part 5.
What the data says to change
- Record examiner failures as failures. When the examining agent ends without answering, Leaf records the area as “nothing.” That happened five times this month. It’s small, but it’s exactly the kind of quiet lie a loop shouldn’t tell.
- Calibrate the reader. Spot-check a sample of runs where the reader found nothing. If real friction is in there, its prior is too strong.
- Update the README’s Improvement claim. “A check can prove that a reduction is smaller” stopped being true when the line-count gate went. Either find a better machine check for “smaller” or say plainly that the examiner’s judgement and the suite are the gate.
- Feed user-found bugs into Quality. A bug that names the PR it regressed should open a fault, the same way a red build does.
- Add an
overcomplexverdict. Behaviour that’s used, with a simpler implementation available. Its done: the suite stays green and a measured complexity score (cyclomatic or cognitive complexity, call depth, layers of abstraction) goes down. That’s the cleanup agents have never been good at, and the one a loop could finally do.
So what?
When you design an agentic workflow, don’t start with “which steps can I automate?” Start with “how will a machine know this is done?”
Where you have a good answer, build a loop and let it run. Where you don’t, that’s where a person belongs, and that’s where you should spend your effort writing the contract.
In my factory, only one loop out of four needed me. The other three found work, fixed it and cleaned up after everyone, including me.
Loops All the Way Down is a six-part series on the contract between humans and agents. New to Leaf? Start with Introducing Leaf.
- Two Kinds of Work
- The Contract
- One Adds, Three Correct (you’re here)
- Where the Tokens Actually Go
- Quality Without the Ceremony
- Output, Not Craft