Last month my software factory spent 7.5 billion tokens. That sounds like a lot. It’s also close to meaningless.

94% of those tokens were cache reads. A cache read is the model re-reading context it has already seen. It’s cheap and fast, and on most plans it barely counts against your limits. If you add cache reads to everything else and call the total “usage,” you get a big scary number that tells you almost nothing about what the work cost.

So before I talk about efficiency, I have to talk about how to count.

Counting tokens honestly

I use one unit for everything in this series. I call it input-equivalent tokens, or IET. It weights each kind of token by what it costs relative to a plain input token:

IET = input + 1.25 × cache_write + 0.10 × cache_read + 5 × output

The weights come from published API price ratios. I apply them to every provider so the numbers stay comparable. This isn’t a dollar figure. Most of my usage runs on subscriptions, not metered billing. It’s a consistent way to say “how much work did this context and this output represent.”

By that measure, the factory spent 1.44 billion IET across 1,206 plan runs between September 9 and October 10. 1,198 of those runs ended in a pull request that GitHub confirmed as merged.

That’s the number I’ll work from.

Where the tokens go

Leaf, my factory, has four roles that spend tokens on a plan. Here’s the split:

Share of input-equivalent tokens by role: workers 60% (make one change and prove it), planner 17% (check the issue, build the graph), foreman 16% (run the graph, review each diff, merge), validator 7% (a separate reviewer pod, retired Sept 30)

The planner number is the one I keep coming back to. On a typical plan, planning is 28% of the total. That number barely moves across loops. Bug fixes, features, cleanups, it’s always about a quarter.

That’s the price of a principle I wrote into Leaf’s README early on:

The plan is the highest-leverage review point. A requirement clarified there benefits every worker, validator, and dependent node; ambiguity discovered during execution consumes scarce agent turns and can propagate through the graph.

You pay a quarter of the plan up front so the other three quarters don’t wander. I think that’s a good trade. But I didn’t know the price until I measured it.

The human side isn’t free

Here’s the part that surprised me.

I went back through a week of my own interactive sessions. I looked only at sessions that ended with a Leaf issue being filed: the conversations where I worked out a bug or a feature and handed it to the factory. There were 29 of them, filing 37 contracts (bugs and PRDs) plus 19 user stories.

For a typical contract, my side of the conversation cost about three-quarters of what the factory spent carrying it out. The median human session was 1.7M IET. The median factory plan it fed was 1.3M.

The average tells a different story. Across the week, the human side cost 1.5× the factory side. Three sessions drove most of that. All three were root-cause hunts: a bug I didn’t understand, with hundreds of turns and up to eight subagents digging through code, logs and the browser. Each cost 11 to 25 million IET. The factory then fixed what they found for 1.4 to 1.9 million each.

Finding out what’s wrong costs more than fixing it.

Scatter plot of 23 interactive sessions, my tokens against the factory's tokens for the work each session filed. Most sit near the equal-cost line; the median human session cost about three-quarters of the factory side. Three root-cause hunts sit far above it at 11M to 25M tokens against 1.4M to 1.9M for the fix.

That’s not a problem with the split. It’s the reason for it. The expensive, wide-open sessions are where judgement happens. They need the browser, other repos, research, a person in the conversation. Once that judgement is written down as a contract, the narrow, cheap work follows. If those investigation sessions also had to write the code, they’d cost even more, and I’d be watching them do it. The context windows would be huge. They’d have irrelevant data in them. All of the usual issues with that sort of development.

Something I have said for years has held up: Once you have all the data, the answer is almost always trivial. I’m glad that truism is holding up.

One more caveat on my numbers: my interactive setup is unusually lean (more on that below). An average developer’s human-loop sessions would likely cost more, not less.

The other half of cheaper: attention

Tokens are the half of “cheaper” that’s easy to count. The other half is me.

A person can only hold a few things in their head at once, and every context switch has a cost. When I’m in an interactive session watching an agent implement something, I’m not doing anything else (or I’m trying not to). I’m on the hook for every turn.

The research on that cost is old and consistent. Gloria Mark and her colleagues at UC Irvine studied what interruptions do to work (The Cost of Interrupted Work: More Speed and Stress, CHI 2008). The surprise was that people finished interrupted tasks just as well, and even faster. But in their words, “this comes at a price: experiencing more stress, higher frustration, time pressure and effort.” Babysitting an agent is a steady stream of small interruptions. You keep up. You just pay for it.

DORA has been finding the team-scale version of this for years. Its 2024 report puts it bluntly: “Unstable organizational priorities cause meaningful decreases in productivity and substantial increases in burnout” (DORA 2024). A developer flipping between three projects and an agent that needs a nudge every few minutes is running unstable priorities at the scale of one person.

DORA also found something about AI that stuck with me. Developers who use gen AI heavily report more flow, higher job satisfaction and less burnout. They also report spending less time on valuable work, while the time spent on toil doesn’t change at all (How gen AI affects the value of development work, December 2024). My read: AI sped up the parts we already liked and left the grind exactly where it was. Watching an agent work is part of that grind. It feels productive. It isn’t the valuable part.

That’s the case for getting the human out of the middle. Not that I’m too busy to watch, but that watching is the toil, and every minute of it is a context switch I’m paying for in stress, not just time.

So I checked the overlap. Across the month, the factory was actively running plans during 282 hours, often in up to five repos at once. I was in an interactive session on this machine for 71 of those hours, and for 38 of them I was working on a different project.

For 75% of the time the factory was working, I wasn’t in an interactive session on this machine at all.

A grid of every hour from Sept 9 to Oct 10. Most hours where the factory was working are amber, meaning I wasn't in an interactive session; far fewer are blue, meaning I was. Factory work runs through nights and whole days.

That number only sees one laptop, so it can’t tell me what I was doing instead. A lot of it was other projects. What it does show is that the work didn’t need me there. That’s a cost no token count captures, and it’s the one I value most.

How does this compare?

I looked for industry baselines. They’re thin.

  • Anthropic’s own docs put the average Claude Code cost at about $13 per developer per active day, with 90% of users under $30 (Claude Code costs). That’s a daily spend, not a per-change cost, so it doesn’t line up neatly with anything here.
  • One team published a trace of their agent work: an average of about 234 million context tokens per single-ticket PR across 18 PRs (beri.net). Leaf’s average is about 6.3 million raw tokens per merged PR, planning included.

I wouldn’t hang anything on that second comparison. It’s one team’s blog against one person’s factory. But it points the same way as everything else here: lean context and small, well-defined units of work are where the savings live.

I couldn’t find a peer-reviewed tokens-per-PR baseline. If you know of one, I’d love to see it.

Lean context is the whole game

Every token a model reads is a token it can get distracted by. My interactive sessions are already leaner than most: a custom system prompt and a small set of plugins and skills. But they still carry what exploring needs: the browser, research, other repos, a long conversation. That’s right for exploring. It’s wrong for building.

Leaf’s workers get almost nothing. A worker’s whole system prompt is about 1,400 tokens. The foreman’s is about 3,300. The /leaf skill I use to write contracts is about 7,400 on its own, bigger than any prompt inside the factory.

The README states the rule:

Factory-wide behaviour belongs in role prompts. Repository-specific knowledge belongs in that repository’s AGENTS.md. Requirements belong in the plan. Mixing those layers makes instructions drift, forces every agent to pay for irrelevant context, and leaves no clear owner when they disagree.

A worker gets its role, the repo’s own guidance, the skills the planner named for that node, and one node. That’s it. No research tools. No other repos. No chat history. It isn’t asked to be curious. It’s asked to finish.

The result shows up as waste. Only 5.6% of the factory’s IET went to seats that didn’t finish their work: pods that failed, went silent, or got stuck. I expected that number to be much higher.

Classes, not models

The planner never picks a model. It picks a class, a rough measure of how much reasoning a node needs:

ClassThe work
2The pattern is in the tree and the brief says where
3The brief has decided the work, or the area is small enough to read whole
4The node has to find the pattern itself (the default)
5The design inside the node is genuinely open

The foreman picks a model within the class when the node runs, based on what’s available and how much quota is left. Complexity is a judgement. Model choice is mechanics.

Did it work? For quality, yes. Nodes land on the first try 96 to 97% of the time in classes 2, 3 and 4. The planner isn’t under-classing hard work, or those numbers would sag at the bottom.

For tokens, it’s messier. Class 3 costs about 5× more per finished node than class 4 (median 534K IET vs 111K). That’s backwards. The reason is routing, not classification: almost all class 3 work ran on one model through one harness, and that model reads a lot. The most efficient models in the data finished nodes at a fraction of that: a local Qwen model at 34K, Haiku 5.5 at 59K, Codex’s lighter model at 63K. This is an optmization to make in how Leaf works and in what I subscribe to.

Size is the biggest lever

Nothing moves cost like the shape of the plan:

Median input-equivalent tokens per merged PR by plan size: 1 node 322K, 2 to 3 nodes 865K, 4 to 7 nodes 2.0M, 8 or more nodes 5.6M, 17 times a single-node plan

An 8+ node plan costs 17× a single-node plan. Some of that is just more work. Some of it is coordination: more merges, more context passed between nodes, more chances to block. Leaf’s README says file ownership is the unit of safe parallelism, and that splitting one coherent change across agents “buys coordination rather than throughput.” The data agrees, loudly.

The nuance

A few things I can’t claim yet.

I can’t say a prompt change cost more tokens. On October 4 I added a rule that workers must prove their change before reporting done. Quality jumped (more on that in Part 5). Cost per small Development PR went up at the same time. But that same week, the work shifted away from the most token-efficient models toward heavier ones. Prompt effect and routing effect are tangled together. I’m not going to pretend I can untangle them.

Tokens aren’t money. Most of this ran on subscriptions with weekly quotas. IET is a fair way to compare work. It isn’t a bill.

This is one person’s factory. Every number here comes from my repos, my habits and my month. It’s a practitioner’s data, not a benchmark.

What the data says to change

I said at the start of this series that the data gets to push back. Here’s where it does:

  1. Fix class 3 routing. It lands as reliably as class 4 and costs 5× the tokens. Either that model’s quota is effectively free and I should say so, or class 3 should prefer a lighter model.
  2. Test smaller plans against big graphs. If a PRD with eight stories runs as eight small plans instead of one big graph, does the cost per merged story drop? I suspect yes. Now I can measure it.
  3. Scope investigation sessions. My human sessions run 96% cache reads. Long sessions re-read their whole history every turn. Starting a fresh session per problem is a cheap win I’m not taking.

So what?

If you’re running agents, stop measuring raw tokens. Weight them, or you’ll optimize the wrong thing. And count the cost that doesn’t show up in any token report: the hours a person spends watching.

Then look at where judgement is spent. In my factory it’s about a quarter of every plan, and in my own sessions it’s often more than the build itself. That isn’t waste. That’s the job. The trick is making sure judgement gets spent once, by the right loop, with the right tools, and written down so nothing downstream has to buy it again.

Measure the harness, not the code.


Loops All the Way Down is a six-part series on the contract between humans and agents. New to Leaf? Start with Introducing Leaf.

  1. Two Kinds of Work
  2. The Contract
  3. One Adds, Three Correct
  4. Where the Tokens Actually Go (you’re here)
  5. Quality Without the Ceremony
  6. Output, Not Craft