For most of the short history of agentic coding, we’ve had a name for the thing we do while the agent writes code. We call it watching TV.

You describe the change. The agent starts working. Text scrolls by. You watch. Sometimes you catch it heading the wrong way and steer. Mostly you just watch, because stepping away feels risky and staying feels pointless.

I did a lot of that. Then I started building a factory so I’d stop, and along the way I figured out why watching TV felt so wrong. We’re running two different kinds of work through the same chat window.

Two loops, not one

The first loop is the human loop. It’s where work gets found and shaped. A bug report. A screenshot. “This feels slow, why?” An idea I want to pressure-test. These sessions start like:

issue: we're seeing the following in sadie.
<paste or screenshot>
goal: determine a root cause and open a properly-scoped issue.

or:

analyze how we do <some thing>. Does it work with this the way I think it does?

note: sadie is my home AI agent who keeps calendars and lights working smoothly for our family.

This loop is wide open on purpose. I want the browser, other codebases, the web, logs, data, and every skill and plugin I’ve got. I want to pick the model, because its habits shape what questions get asked. You can never outrun the training. This is the playground, and it typically runs on my own machine.

The second loop is the implementation loop. It’s where work gets done. Change these files. Make this test pass. Merge it. For most of our short history, the first loop has done this work too. That’s the TV part.

They aren’t the same work. We shouldn’t treat them the same.

Judgement and mechanics

The cleanest way I’ve found to say it: the human loop is judgement, and the implementation loop is mechanics.

Judgement is deciding what’s wrong, what matters, what “fixed” looks like, and what’s out of scope. Mechanics is everything after that decision: editing files, running checks, fixing what fails, merging.

My factory, Leaf, has this split written into its README:

Automate mechanics, not judgement

Predictable transitions should be performed by code: they are cheaper, replayable, and easier to test than another model turn. Decisions that depend on meaning stay with the agent that owns the graph.

And a test for where the line goes:

encode a rule only when the available facts determine one correct action. If context can make either action right, expose the facts and leave the choice with the responsible agent or person.

But the rule that does the most work is shorter:

A loop runs unattended only when a machine can tell that it is done.

That’s the whole design test. If a machine can tell the work is done, the work doesn’t need a person watching it. If it can’t, it does. So the job isn’t to remove humans from software development. It’s to make “done” something a machine can check, and to do that before the work starts.

Humans decide what done means. Machines decide whether it happened.

Focusing on the handoff

If there are two loops, something has to pass between them. In most setups that’s a prompt typed into the same session that’s about to do the work. The thinking and the building blur together, and the agent fills gaps with guesses or dogmatic prose.

I think that handoff is the real design problem in most agentic development theories right now. Not which model. Not which agent. The contract between what a human does and what the automated workflow does.

Leaf is my current answer. Work enters as a GitHub issue written to a standard defined in a skill: what’s happening, what should happen, what’s in scope, and how to check it’s done. An independent agent checks the issue before anything else happens. A planner turns it into a graph of small jobs, each with a check a machine can run. Then the factory takes it from there, all the way to a merged pull request. Rinse. Repeat. Track and improve.

Standing on someone else’s loop

I’m not the first person to draw this picture. Kief Morris wrote a piece on Martin Fowler’s site this spring, Humans and Agents in Software Engineering Loops, that splits development into a why loop (turning ideas into outcomes) and a how loop (building the software). He argues that humans shouldn’t sit outside the loop and let agents run wild, or in it reviewing every line. They should be on it. As he puts it, “The challenge when we insist on being too closely involved in the process is that we become a bottleneck.” Instead, “The ‘on the loop’ way is to change the harness that produced the artefact so it produces the results we want.”

I agree with all of that. This series is my attempt to push on three things his piece leaves open:

  1. The interface. If humans are on the loop, what exactly do they hand it? I think the answer is a contract, and it can be designed.
  2. The tools. The human loop and the implementation loop shouldn’t load the same tools. One should be wide. The other should be brutally lean.
  3. The measures. If we’re going to argue about this, we should measure it. I use two: token efficiency and quality of output.

Output, not craft

One more frame, and then the data.

I care about craft. I’m a woodworker. I like a clean joint. But the people who use my software never see the joints. They see whether it works, whether it’s fast, and whether it breaks. That’s output. Most people don’t judge a chair on its build quality. They judge it on whether or not it’s comofrtable.

This series judges everything by output. Clean code matters when it changes output: when it makes the next change cheaper, faster or safer. If it doesn’t, it’s taste. I test that idea against my own data in Part 6. Some of it held, and some of it didn’t.

What the data looks like

Between September 9 and October 10, Leaf ran 1,206 plans and 1,198 of them ended in a merged pull request. Nodes landed on the first try 97% of the time by the end of the month. The median issue went from filed to merged in about an hour. Most of the non-merged changes were test runs for major workflow changes that I didn’t delete.

The best test is my product monorepo. It’s my oldest codebase, built over time with a pile of different harnesses and processes. Comparing the month with Leaf against the four months before it:

  • Merged changes per week went up 3.8×. Feature and fix work alone went up about 1.8×. The rest is a cleanup loop that didn’t exist before.
  • Bugs introduced per line of code stayed about the same. No better, no worse, at several times the pace.
  • Product source shrank 18% (about 105K lines and 586 files) in under three weeks, while the tests grew.
  • For 75% of the hours the factory was working, I wasn’t in an interactive session on this machine at all.

Four headline numbers from a month of the factory: 3.8 times more merged changes per week, about the same bugs per line of code (9.9 to 9.1 per 10K lines), 18% less product source code, and 75% of factory hours with me not in a session

That last one is the number I care about most. Cheaper comes in a lot of forms. Tokens are one. My attention is the expensive one. Leaf runs on a small collection of small systems in my office closet. It runs whether or not my laptop lid is open.

It still cost real tokens, on both sides. Here’s a number I didn’t expect: for a typical piece of work, my own interactive session writing the contract cost about three-quarters as many tokens as the factory spent building it. For gnarly bugs, my side cost far more than the fix. Finding out what’s wrong costs more than fixing it. That turns out to be the best argument for the split, not against it. Simple bug fixes were just that; simple. An error snippet or a screen shot with an obvious problem and fix is fewer tokens to diagnose than a full plan, implementation, and full e2e suite to validate it.

A caveat that applies to everything in this series: it’s one person’s factory. My repos, my habits, my month. It’s a practitioner’s data, not a benchmark.

A second caveat: for the large monorepo app, I run a full e2e suite once all the Leaf work is done. Before a PR is created. This hasn’t been for the whole month. But it’s a great quality gate IMO.

The series

  1. Two kinds of work (this post). Judgement and mechanics, and the handoff between them.
  2. The contract. What a good issue looks like when a machine has to act on it. Humans define done; machines check it.
  3. One adds, three correct. The loops inside my factory, and why only one of four needs a person.
  4. Where the tokens actually go. Counting honestly, lean context, and what judgement really costs.
  5. Quality without the ceremony. Where quality comes from when no human reads the diff, and why the merge button doesn’t matter.
  6. Output, not craft. What humans give up, what the data couldn’t prove, and where this breaks.

Each post ends with a section called “What the data says to change.” I built Leaf, and the data still found plenty I got wrong. That’s the point. This is about improvement, not about defending a design.

So what?

If you’re spending your days watching an agent scroll, ask which loop you’re in.

If you’re still figuring out what’s wrong or what to build, stay. That’s judgement, and it deserves your attention and your best tools.

If you already know, write it down well enough that a machine can tell when it’s done. Then go do something else.


Loops All the Way Down is a six-part series on the contract between humans and agents. New to Leaf? Start with Introducing Leaf.

  1. Two Kinds of Work (you’re here)
  2. The Contract
  3. One Adds, Three Correct
  4. Where the Tokens Actually Go
  5. Quality Without the Ceremony
  6. Output, Not Craft