We’ve turned the pull request into a ritual. Request reviewers. Wait. Address nits. Re-request. Wait again. Somewhere in there, a person clicks a green button, and we call the code “approved.”
My factory merged 1,198 pull requests between September 9 and October 10. I clicked merge on 27 of them. Automation merged the rest.
That isn’t because I stopped caring about quality. It’s because I moved quality somewhere it does more good.
Quality is a moving target
I’ll start with the framing, because it matters more than the numbers.
Quality isn’t a gate you pass once. It’s what your loops keep pushing toward. The code changes, the models change, the work changes. A check that caught everything last week misses something this week. So the question isn’t “did a human approve this?” It’s “how many independent chances did we have to catch the problem, and are those chances getting better?”
Current quality gates don’t go away in my world. Humans just stop staring at them.
Where quality actually enters
In Leaf, quality shows up in layers. Most of them act before a line of code exists.
- The contract. Every piece of work starts as an issue with a written definition of done. A question nobody has answered yet becomes a
[NEEDS CLARIFICATION]marker, and that marker blocks the issue from passing its check. - The skills. Each node carries the repo’s own skills, so “what good looks like here” comes along with it.
- The plan’s detail. Every node names the files it may touch, a
done_when, and a runnablecheck. The planner dry-runs those checks before it saves the plan. - The worker’s proof. The worker has to run its check before it can say it’s done.
- The foreman’s review. The foreman reads every diff against the node’s intent before it merges it into the plan branch.
- The end-to-end suite. For repos that declare one, it runs on the plan branch before the PR exists, with a fix loop until it passes.
- CI. The same gate it always was.
The worker’s rule is short and blunt:
Make the change on the branch that’s already checked out. Prove it with the node’s check, or with the unit tests for the files you changed: those tests only, never a whole suite or a build. Reading the code, a diff or a grep is never proof; before the done line, after the final edit, run the node’s check.
And the foreman’s review is scoped on purpose. It judges each diff against the node’s intent and done_when only:
style, wording or improvements the node did not ask for are not grounds to send it back.
That second rule matters. Review that wanders into taste is where ceremony comes from. Standards are defined. But I’ve found no reason to quibble over them. Agents are the audience. Not humans. Take that PEP-8!
What the layers catch
Each layer catches things before a human would ever see them:
- When Leaf had a separate validator pod, it failed 11% of the work it reviewed.
- The first end-to-end run failed on 19% of the plans that had a suite. The suite loop then fixed them, usually in one more run.
- Overall, 99.3% of plan runs ended in a merge GitHub confirmed.
And the result: nodes landing on the first try went from 83.5% to 97.2% between mid-September and this week. The median time from a filed issue to a merged Development PR is about an hour.
What the changes taught me
I changed Leaf a lot this month. 158 commits touched its prompts and skills. A few of those changes were structural, and the data around them says more than the totals do.

Three things jump out.
Removing a quality layer made quality better. On September 30, I retired the separate validator pods and made the foreman own diff review. Blocked nodes went from 14.5% to almost nothing. First-run e2e failures halved. It also got cheaper. A second reviewer that doesn’t own the outcome isn’t a stronger gate. It’s another handoff. Gates matter less than who owns the judgement.
Proof beat review. The biggest single jump came from one prompt rule: prove your change before you say it’s done. First-try went from about 90% to about 97%. First-run e2e failures fell from 34% to 7%. Making the worker run its own check did more than any reviewer I added.
The merge button changed nothing. On October 6, the foreman started merging its own PRs instead of relying on GitHub’s auto-merge setting. Quality didn’t move. That’s the point. Who presses the button doesn’t matter once the gates before it have done their work. The merge is a mechanical step after a quality goal is met. It isn’t a ceremony.
What review was really for
If the merge doesn’t need a human, what was code review doing all this time?
Pull it apart and it does at least four jobs:
| What review does | Where it moved |
|---|---|
| Finding defects | Left, into the plan’s checks, the worker’s proof, the foreman’s review, the suite and CI |
| Sharing understanding | Up, into the issue and the plan. I read plans, not diffs |
| Approval and accountability | Into the contract. I approve the plan, or I set the repo to run on its own, and I own the definition of done |
| Consistency and style | Into the harness: repo skills, AGENTS.md, linters, and an Improvement loop that walks the code looking for things to simplify |
This isn’t a new idea. In 2013, Alberto Bacchelli and Christian Bird studied code review at Microsoft (Expectations, Outcomes, and Challenges of Modern Code Review). Finding defects was the top reason developers gave for reviewing. But reviews turned out to be less about defects than people expected. What they delivered was knowledge transfer, team awareness and alternative solutions. If that’s right, the ceremony has been guarding the wrong thing. Defects are better caught by checks that run every time. Understanding is better shared in the plan, where the decisions actually get made.
Did quality hold after the merge?
Everything above stops at the merge. The real question is whether the code that merged was any good. Did bugs show up later?
A pull request lists the bugs it closes. Nothing records the bugs it causes. When a bug gets filed next week, the issue doesn’t say “this regressed in PR #4521.” So I reconstructed that link the way researchers do, with a technique called SZZ: take every bug fix, look at the lines it had to change, and trace those lines back to the commit that wrote them.
I ran it on my product monorepo, the oldest and messiest codebase I have, built over years with a mix of harnesses and processes. I traced 972 bug fixes since May. A commit counts as “implicated” if a fix within two weeks had to change a line it wrote.

Per line of code, the factory introduces bugs at about the same rate as the process it replaced. Not better, not worse. It’s slightly lower per commit, but factory commits are smaller, so that’s mostly size. A one-week window says the same thing.
That might sound underwhelming. It isn’t. “The same” is at several times the pace: merged changes per week in that repo went up 3.8× (feature and fix work alone about 1.8×), mostly while I was doing something else. And the baseline wasn’t me hand-writing code. It was the mix of harnesses and AI-assisted processes that built the repo in the first place.
Three quality signals I didn’t have before
The factory also made quality easier to see.
- Bugs are easier to identify. Every bug comes in through the same contract with the same label. My old process was ad hoc. Funny side effect: better identification makes the factory look worse compared with history, because it finds bugs the old process never labelled. To keep the comparison fair, I only counted fixes the old way.
- Bugs are easier to track. Leaf’s database and git history together tell me what was planned, what ran, what merged and what broke. The analysis above took about a minute of compute. Before, it wasn’t possible at all.
- Less code means fewer bugs. That’s been true for decades: more code, more defects. Leaf’s Improvement loop cut that repo’s product source by 18% in under three weeks, and its own PRs were almost never implicated in a later bug. (Part 3 has the details, and the blind spot.)
The nuance
I’m not going to pretend this is finished.
The quality data is young. The factory-era numbers cover two to four weeks. SZZ is a heuristic: a fix that only adds lines blames nothing, and a refactor inside a fix can blame the wrong commit. I’d call the result “no worse, at several times the pace” and not one word stronger. Your mileage, and my next month, may differ.
The link should be native, not reconstructed. I had to rebuild the bug-to-PR link after the fact. A bug that names the PR it regressed would make this a recorded fact instead of an inference.
First-run e2e failures are creeping back up. They went 7% → 11% → 18% over the last week as more work hit areas the suite covers. That’s what a moving target looks like. It’s also the next thing to fix.
Some teams need a human at merge. Compliance rules often say a second person has to approve a change. I think plan approval plus a complete evidence trail (the issue, the plan, every check, every diff review, the room log) is a stronger record than a thumbs-up on a diff. But that’s an argument you’d have to win with your auditor, not with me.
This is one person’s factory. I’m the human in every loop that has one. A team would need to decide who writes contracts and who approves plans. The model holds; the roles get harder.
What the data says to change
- Link post-merge bugs to the PR that caused them. A
type:bugshould name the change it regressed. SZZ got me most of the way after the fact. A native link would make output quality a recorded fact, and that’s the quality that matters. - Record why a node blocks. Leaf can’t tell “the contract was ambiguous” from “the disk filled up.” Those are different failures with different owners.
- Fix the README. It says that in the Development loop, “a person decides when it is done.” That’s misleading. A person defines done when they approve the plan. The machine checks it.
So what?
If your team’s quality strategy is “a senior engineer reads every diff,” you have a single point of failure that gets slower as your agents get faster.
Move the gates left. Make every unit of work prove itself. Give review to whoever owns the outcome, and scope it to the contract. Then let the merge be what it always should have been: the last mechanical step, not the moment of truth.
Loops All the Way Down is a six-part series on the contract between humans and agents. New to Leaf? Start with Introducing Leaf.
- Two Kinds of Work
- The Contract
- One Adds, Three Correct
- Where the Tokens Actually Go
- Quality Without the Ceremony (you’re here)
- Output, Not Craft