Most agent workflows start with a prompt. Mine start with a contract.
That sounds like a distinction without a difference. It isn’t. A prompt is the start of a conversation. A contract is the end of one. It’s what’s left when the thinking is done and the work can be handed to something that doesn’t have a direct path to me to ask follow-up questions.
In Part 1, I argued that agentic development is two kinds of work: judgement and mechanics. Humans should spend judgement. Machines should run mechanics. This post is about the thing that sits between them, because that’s where too many setups quietly fall apart.
Two contracts, not one
There are really two contracts in play.
The first is standing. It’s true for every piece of work: how each role behaves, how this codebase works, what good looks like here. Leaf’s README splits it into layers:
| Layer | What it holds | Where it lives |
|---|---|---|
| Factory-wide behaviour | How a worker, foreman or planner acts | Role prompts |
| Repository knowledge | How this codebase works | The repo’s AGENTS.md and skills |
| Requirements | What this job is | The issue and the plan |
The second is per job. That’s the issue. It’s the one I want to talk about, because it’s the one a person writes every time.
Anatomy of a real one
Here’s a real contract from a personal project of mine called Sadie, a household assistant. The household had stopped wanting a weekly “favorites on sale” email. I worked out what was going on in an interactive session and filed this:
## Observed
Every Sunday at 08:00 America/New_York, every user account with an email
gets a separate "This week's favorites on sale" email. ...
## Expected
No weekly specials email is sent, and nothing in the tree schedules or
builds one. Favorites and the per-list deals check work exactly as they
do today.
## Scope
Delete: (five files) ... Reword comments that name the digest and would
be left false ... Leave `sadie/mail/` alone.
## Host step (not doable from a pod)
sudo systemctl disable --now groceries-digest.timer ...
## Done when
- `git grep -n groceries.digest` and `git grep -n groceries-digest` return nothing.
- `scripts/test.sh open_webui/sadie/groceries` passes, with the favorites
and deals tests unchanged.
- On `apps`, `systemctl list-timers --all` no longer lists the timer.
A few things about it.
It’s complete without being long. It says what’s happening, what should happen, what’s in scope, what’s explicitly out of scope, and how to tell when it’s done. Nobody reading it has to guess.
It doesn’t say how to write the code. It names the files and the boundaries. The implementation is the factory’s problem.
It’s honest about what a machine can’t do. One step has to happen on the host, by a person. The contract says so, in its own section. That matters, and I’ll come back to it.
Done, made checkable
The planner turned that issue into a one-node plan. Here’s the part of the node that matters (trimmed):
{
"id": "groceries-1",
"files": ["deploy/groceries-digest.timer", "...", "groceries/tools.py"],
"done_when": "the five files are gone; `git grep ...` returns nothing; neither service.py nor tools.py mentions a digest; `scripts/test.sh open_webui/sadie/groceries` passes with the favorites and deals tests unchanged; no file under sadie/mail/ is modified",
"check": "scripts/test.sh open_webui/sadie/groceries && ! git grep -n -e groceries.digest -e groceries-digest -e 'digest.py' -- deploy backend/open_webui/sadie",
"class": 2,
"effort": "low"
}
Look at done_when next to check.
done_when is the human’s definition of done, sharpened. check is one shell line that proves it. The planner’s prompt is strict about this:
check: one shell line the worker runs to prove its change: the unit tests for its files, or a lint or type-check of them where no test applies. … Before you save the graph, run every check in the tree: it must run and reach what it proves.
The planner runs every check before it’s allowed to save the plan. The worker runs it again before it’s allowed to say it’s done. The proof is tested before the work starts.
And notice what isn’t in the check: the host step. The planner left the systemctl line out of the machine’s definition of done, because no pod can verify it. It went into the commit message as a note for me. The contract kept the human’s obligation and the machine’s obligation separate.
We’re also not filling the agent doing the work with dogma and restrictions. Modern models make solid judgements at these levels. We scope the complexity, assign the right level of model by capability vs. complexity and let it do its job with a clear purpose and a clear test of when it’s done.
That’s the whole idea in one example:
Humans decide what done means. Machines decide whether it happened.
What it cost
The Sadie contract ran like this:
- I filed it, and it waited for my approval.
- Once approved, the plan ran and merged in 8 minutes.
- Haiku did the work. Opus ran the foreman.
- Total cost: about 155K input-equivalent tokens, with planning at about a third of that.
Cheap model, small context, one shot. That only works because the contract left nothing to figure out.
The terms
If you write it out like a real contract, it looks like this:
| The human owes | The factory owes | |
|---|---|---|
| Before | A complete issue, including what done means | An independent check: Ready. or Not ready. and what’s missing |
| During | Nothing, unless asked | A plan (graph, files, checks, a class for each node), then the work |
| When stuck | Answer needs detail, or join the planner’s room | Stop and ask rather than guess |
| After | Nothing. The merge is mechanical | A merged PR that meets the stated done, then cleanup |
| When it goes wrong | Fix the harness, not the artifact | The evidence: room log, trajectories, telemetry |
Two rules hold the human side together.
Unresolved questions are marked, not guessed. If I can’t answer something yet, it becomes a [NEEDS CLARIFICATION] marker in the issue. That marker blocks the issue from passing its check. It records a decision nobody has made, which is exactly what you want to find before eight agents start building on it.
The check is independent. The agent that judges whether an issue is ready isn’t the one I wrote it with. It reads the issue cold, against a written standard, the way a stranger would.
On the factory side, Leaf has a rule I’ve come to love:
Settled: prompts carry the rules, code carries the physics
Code refuses only what a prompt can’t prevent: an unauthenticated webhook, a host out of memory, a suite report with no tests in it read as green. Everything else is judgement, and judgement lives in the prompts, where I can read it and change it. I don’t over-burden the system with complex rule sets in code that should be a continual prompt improvement. It causes pain long term.
Do contracts actually arrive complete?
Over the last month, Leaf ran 1,225 planning rooms. Only 26 issues came back Not ready. That’s about 2%.
That’s either a sign the contracts are good or a sign the check is soft. I think it’s mostly the first, and here’s why: in Part 4 I measured my own sessions, and the ones that ended in a filed issue cost about three-quarters of what the factory spent carrying the work out. Sometimes much more. The contracts arrive complete because the human side already paid for them.
The nuance
A complete contract can still be wrong. If I write a clear, checkable issue for the wrong fix, the factory will build the wrong fix quickly and well. The contract moves risk. It doesn’t remove it. Judgement is still judgement.
I can’t yet prove that better contracts make better outcomes. I wanted to. The obvious signal is a worker getting stuck and reporting blocked. Plans where a worker blocked cost 1.5 to 2× more at every plan size. But when I read the block reasons, most were infrastructure: a failed turn, a full disk, a check that refused three times. Only a few pointed back at the contract itself. Leaf doesn’t record why a node blocks in a way I can count. That’s a feature to file.
Some work doesn’t fit a contract yet. If I can’t say what done looks like, it isn’t ready for the factory. It’s still exploration, and it belongs in the human loop. That’s not a failure of the model. That’s the model working.
What the data says to change
- Record why a node blocks. Contract, scope, check, infrastructure or harness. Without that, “better contracts make better outcomes” is a belief, not a finding.
- Fix the README’s wording. It says that in the Development loop, “a person decides when it is done.” It should say a person defines done, by approving the plan, and the machine verifies it.
- Watch the planner’s one-turn limit. The planner has exactly one turn, and it has to dry-run every check in that turn. On a large graph with slow tests, those two rules will collide.
So what?
If you hand work to agents with a prompt, you’re handing them a conversation they can’t have. They’ll fill the gaps with guesses, and you’ll pay for the guesses in tokens, retries, review, and rework.
Write the contract instead. Say what’s true now, what should be true after, what’s out of scope, and how to check it. Mark what you don’t know. Keep the human’s obligations and the machine’s obligations separate.
Then let the machine hold you to it.
Loops All the Way Down is a six-part series on the contract between humans and agents. New to Leaf? Start with Introducing Leaf.
- Two Kinds of Work
- The Contract (you’re here)
- One Adds, Three Correct
- Where the Tokens Actually Go
- Quality Without the Ceremony
- Output, Not Craft