I like well-made things, including the parts nobody will ever see. So it’s a little uncomfortable to spend six posts arguing that, in software, nobody should be judged by how the joints look.

But that’s where the data took me. The people who use software see output. Does it work? Is it fast? Does it break? The craft underneath matters exactly as much as it changes those answers, and no more.

This last part is about what that means, what my data could and couldn’t prove, and what we give up.

Fix the harness, not the artifact

If you judge by output, a strange rule follows. When an agent produces bad work, you don’t fix the work. You fix whatever produced it.

Leaf’s AGENTS.md says it in one line:

When an agent misbehaves, fix its prompt.

Code only refuses what a prompt can’t prevent: a webhook with no auth, a host out of memory, a test report with no tests in it read as green. Everything else is behaviour, and behaviour lives in prompts I can read and change.

Kief Morris calls this a flywheel: humans improve the harness, and the harness improves the output. In his words, “The flywheel becomes more powerful as we feed it richer signals.” It sounds tidy in an essay. Here’s what it looked like in one month of mine:

  • 158 commits to Leaf’s prompts and skills.
  • First-try landing went from 83.5% to 97.2%.
  • The median cost of a merged PR fell from 2.1M to 337K input-equivalent tokens (the mix of work shifted too, so that one comes with an asterisk).

I didn’t hand-fix 1,198 pull requests. I fixed the thing that wrote them, about five times a day.

And the factory does some of this itself. In Part 3, I described the Trajectory loop: Leaf reads its own run records, notices problems that keep coming back, and files issues about itself. It tracked 256 problems this month. 84 stopped recurring. That’s the flywheel with nobody pushing it.

Does clean code pay for itself?

Here’s the claim I most wanted to prove: if internal quality only matters when it changes output, then cleaning up code should show up in the output. It did. Just not where I first looked.

Where I looked first: cheaper agents

My first guess was that agents would feel messy code as cost, so cleanup would make the next change cheaper: fewer tokens, fewer retries. Leaf’s Improvement loop simplified code 617 times this month. So I checked whether later feature work in simplified areas got cheaper.

I couldn’t find it. Week by week, work in simplified areas cost about the same as work elsewhere (a median of 105K vs 101K tokens per node one week, 253K vs 245K another), with about the same first-try rate. Three weeks is short, my measure was crude (same directory, not same code), and by the last week three-quarters of all feature work landed in simplified areas, which left almost nothing to compare against. But I said I’d publish it either way: I can’t show that cleanup makes each agent run cheaper.

Where it showed up: the codebase itself

Measuring the code instead of the agents tells a different story. In my product monorepo, the oldest and most organically grown codebase I have:

  • Product source shrank 18% in under three weeks: about 105K lines and 586 files gone, while tests grew.
  • Those cleanups almost never broke anything. Traced back from later bug fixes, Improvement PRs were implicated about 1% of the time.
  • Bugs per line of new code held steady, while merged changes per week went up 3.8×.

Not every benefit of less code shows up as a cheaper agent run. Some of it is 105K fewer lines for the next agent to read and get led astray by. Some is smaller bundles and builds. And some is the oldest rule in the book: less code, fewer bugs.

What it hasn’t touched yet

Most of what the loop removed was dead code. The harder target, and the one agentic development has always been worst at, is code that’s over-engineered but works. Leaf’s examiner has no verdict for that today, because “simpler” is hard for a machine to check, and a loop only runs unattended when a machine can tell it’s done. Giving it a real measure of done (the suite stays green and a complexity score goes down) is the next step. That’s the design rule from Part 1, deciding what gets built next.

What humans give up

This model takes things away from people. I’d rather name them than pretend it doesn’t.

Reading the code. I read plans, not diffs. That’s a real loss. Reading code is how a lot of us learned, and how we keep a mental model of a system we’re responsible for. My bet is that the plan, the contract and the run records carry enough understanding. I’m not sure that holds for a team of fifteen.

The feeling of control. For most teams, clicking merge feels like a decision. In Part 5, I showed it wasn’t one. Who pressed the button made no measurable difference. But it felt like control, and losing that feeling is uncomfortable even when the control was never real.

The middle of the work. The human loop is the start (judgement, contracts) and the edges (the harness). The middle belongs to the machine. If you love the middle, this is a sad trade. A lot of us got into this work for the middle. The output is better for the trade anyway.

Where this breaks

A few places where I know it doesn’t hold yet:

  • Choosing the work. Leaf runs contracts. It doesn’t decide which contracts matter. That’s still all judgement, and it’s still all me.
  • The contract is the ceiling. A complete contract for the wrong fix gets built quickly and well. The factory doesn’t make my judgement better. It makes it faster and more visible.
  • Quality past the merge is young. SZZ says the factory introduces bugs at the same rate per line as before. That’s two to four weeks of data and a heuristic. The bug-to-PR link should be recorded, not reconstructed.
  • Teams. I’m the only human in every loop that has one. A team has to decide who writes contracts, who approves plans, and who owns the harness. The model holds. The roles get harder.
  • Regulated work. If an auditor needs a second person to approve every change, plan approval and an evidence trail might satisfy them. That’s a conversation to have, not a given.
  • The human loop costs real tokens. In Part 4, my own contract-writing sessions often cost more than the factory’s build. That’s judgement, and it’s worth paying for. But I’m paying for a lot of it inefficiently.

Everything the data told me to change

This series promised that the data gets to push back. It did. Here’s the full list, across all six parts:

  1. Record why a node blocks (contract, scope, check, infrastructure, harness). Without it, I can’t test whether better contracts make better outcomes.
  2. Fix class 3 routing. It lands as reliably as class 4 at about 5× the tokens.
  3. Link user-found bugs to the PR that caused them natively (not just by SZZ after the fact), and let them open a Quality fault like a red build does.
  4. Test smaller plans against big graphs. An 8+ node plan costs 17× a single-node plan.
  5. Scope my own sessions. Fresh context per problem. My human sessions are 96% cache reads.
  6. Record examiner failures as failures, not as “nothing found.”
  7. Calibrate the reader that feeds the Trajectory loop. Its “nothing is usual” prior may hide quiet problems.
  8. Rewrite three README claims that stopped being true: who decides “done,” how a reduction proves it’s smaller, and the empty models column in the class table.
  9. Watch the planner’s one-turn limit. It has to dry-run every check in a single turn. Big graphs will collide with that.
  10. Teach Improvement to find over-engineered code that works, with a done a machine can check: the suite stays green and a complexity score goes down.

Ten changes, from one month of data, on a system I designed. That’s the strongest argument I have for measuring the harness: it’s wrong in ways you can’t see until you count.

Where this leaves Morris’s loops

I started this series from Kief Morris’s picture of humans on the loop. Here’s how mine maps onto it:

MorrisThis series
The why loop and the how loopThe human loop (judgement) and the implementation loop (mechanics)
Nested how loopsWorker, foreman, suite and PR loops, plus three loops that correct the code and the factory itself
Humans on the loop, building the harnessYes, plus a designed contract as the harness’s input
Shift quality checks leftQuality in layers, most of it before code exists, and proof over review
The flywheel”Fix the prompt,” and a Trajectory loop that files issues about its own runs
(Not covered)Wide tools for the human loop, lean ones for the machine. The merge as a mechanical step. Token efficiency and output quality as the two measures

So what?

Here’s the whole series in one line:

Humans decide what done means. Machines decide whether it happened.

Spend your judgement where it counts: finding what’s wrong, deciding what matters, writing it down well enough that a machine can hold you to it. Give that judgement the widest tools you have. Then hand it to a loop that’s lean, checks its own work, and doesn’t need you watching.

And measure all of it. Not lines of code, not PR counts, not how good the diff looks. Measure the tokens, the outcomes, and the harness that turns one into the other.

Measure the harness, not the code.


Loops All the Way Down is a six-part series on the contract between humans and agents. New to Leaf? Start with Introducing Leaf.

  1. Two Kinds of Work
  2. The Contract
  3. One Adds, Three Correct
  4. Where the Tokens Actually Go
  5. Quality Without the Ceremony
  6. Output, Not Craft (you’re here)