Your Agents Outwrite Your Readers: Budget Review Like a Resource

A developer I know has not written a line of code at work in ten months. He is not hiding it. His pull requests merge, his features ship, his team is happy. He mentioned it to me the way you'd mention a new commute.
My first question was not how. It was: how much of it have you read?
Because that is the number that decides whether his ten months are a career milestone or a slowly compounding liability. He could not answer it. Neither could most of us.
A loop is a feedback loop, not a mailing list
"Human-in-the-loop" has become a phrase people argue about, mostly because they mean two different things by it. The common reading is the office reading: the person who gets CC'd on the email. Informed, not involved. There's a well-worn analogy for this — the parent who texts their kid before a party. They know where the kid is going. The kid goes anyway.
That is not a loop. In engineering, a loop is a three-step chain: input goes in, something acts on it, the result gets checked, and the check decides what happens next. A thermostat doesn't warn the furnace. It turns it on or off. If you put a human in that loop, the work stops until they do their part. That's the whole point.
A developer I'll call Marcus put it well in a piece on dev.to about being a pragmatic AI developer: a human-in-the-loop is someone responsible for reading, evaluating, and exercising the code. Every one of those people can stop the work. Nothing gets approved because a Slack message arrived. An approval tapped on a phone between meetings is not a link in the chain. It's a notification.
Fine. So why does this keep being a debate instead of a policy? Because we never asked the harder question underneath it: if nothing moves until a human reads it, what happens when the machines produce ten times more than they used to?
Generation scales sideways. Reading doesn't.
Here is the asymmetry nobody budgets for.
Generating code is now cheap and parallel. Spin up five agents in five git worktrees and they will all work at once. One engineer at MetalBear, an infrastructure company that surveyed its own team about how they actually use AI, described trying to save money by having a strong model write a detailed implementation guide and a cheaper model write the code. It failed for a boring reason: the cheap model's code wasn't good enough, so someone had to redo it. But underneath that failure is a habit worth noticing — the team keeps trying to add more writers to a problem that is really about readers.
Their write-up is unusually honest about this. In January, their engineers said AI was useful around the edges of the work. Eight months later, most of the people who replied barely write code by hand anymore. One of them said: "I'm not really writing any code, I just review the slop and try to tame the beast." The part of the job that grew the most in that company was not writing, and not testing. It was review.
Now put a number on it. A widely-read analysis of one month of Claude Code usage — 455 sessions, 63,000 requests, deduplicated — found that subagents consumed 48.1% of all tokens and wrote 64% of all output. Every subagent pays an entry price before it does anything: its opening request already carries a median of about 47,000 tokens of instructions, tool lists, and skills. Launching work has never been cheaper. Reviewing it has never been more expensive.
This is the sentence I'd put on a wall: you can add writers, but you cannot add readers. Five agents can produce five diffs in the time it takes you to read one. That queue — not model quality, not token cost — is the constraint in an AI-heavy team.
Set a review budget, even a crude one
You almost certainly do not track this today. Start with three numbers you can pull out of git and your pull-request tooling in an afternoon. No new platform required.
- Change entering review per engineer per week. Lines added plus removed, not commits. This is your arrival rate.
- Median time to first review. If it's climbing, you are producing faster than you can absorb.
- Re-open rate. How often a merged change gets a follow-up fix within two weeks. This is your best cheap proxy for "nobody really read this."
Then anchor against a rough human limit. A careful read of unfamiliar code — the kind where you follow the call sites and check the edge cases — runs on the order of a couple hundred lines an hour. A diff you can genuinely hold in your head is a few hundred lines. A 1,400-line diff approved from a phone is a signature, not a review. Those numbers are rules of thumb, not research; calibrate them against your own team's re-open rate. The point is that you now have a ratio: lines arriving, divided by lines anyone could realistically read. If that ratio is above one, you are not running a review process. You are running a queue with a rubber stamp at the end.
Tier the change, not the reviewer
Reading everything does not scale, and pretending it does produces a worse outcome than tiering honestly. So tier deliberately.
Tier 1 — read it, run it, test it yourself. Authentication and authorization, billing, data deletion, migrations, anything that crosses a customer boundary. Here you read the diff, check out the branch, run it locally, and write at least one test the agent did not write. The agent's tests tend to agree with the agent's code; that is their defining weakness.
Tier 2 — read the diff, run the tests, skim the plumbing. Internal tooling, build scripts, CI tweaks, admin endpoints behind a login. You want to know the shape of the change and roughly how the data flows. You do not need to reconstruct every line.
Tier 3 — read to understand it, never merge it. One-off scripts, scratch analysis, data poking, a repro someone needed at 4 p.m. Let an agent write these all day. Then delete them. The graveyard of every codebase is full of scripts that were never meant to be permanent and quietly became load-bearing.
Be honest about how this fails. Tiering requires judgment, and judgment is the thing under pressure. The failure mode has a signature: a change lands in a sensitive area, someone labels it "internal tooling" because the ticket was vague, and three weeks later it turns out the internal endpoint reads customer data. Write your tier-1 list down and treat it as a policy, not a mood.
Spend the cheap output on making reading cheap
If generating text is nearly free, spend some of it on the reader instead of the writer. This is the most underused move available to you right now. Ask for the change the way you would want to read it:
- Split it into small commits, each with a one-line reason for existing.
- Name the two or three riskiest lines and say why they are risky.
- Include the test that would fail if the change were wrong — not a test that passes because the code says so.
- List the files by risk, so you read in that order and can stop when your hour is up.
One caveat, and it comes straight out of the MetalBear survey: agents write confident pull-request descriptions that do not always match the code. Several of their engineers write every commit message and PR description by hand for exactly this reason. So treat the description as a claim to verify, not a summary to trust. If the description is doing your reviewing for you, you have replaced the loop with a nicer-looking distribution list.
Keep a ledger of what you didn't read
This is the part that changes how you feel about the whole thing. If you merged code without reading it, record it. Not as a confession — as inventory.
Three columns in a file in the repo, or an issue with a label, is enough: what changed, who touched it, and what you skipped. Over a quarter you get something genuinely useful: a list of code that nobody has read and nobody can explain, distinct from code that nobody has read but one person understands deeply. Those two categories fail very differently. The first one is archaeology at 2 a.m. The second one is a phone call.
Unread code is not automatically bad code. Most of it is fine, and some of it you were right to skip. But the skipping is a decision, and decisions deserve records. A ledger converts hidden debt into visible debt, and visible debt is the only kind you can pay down on purpose.
The three questions that tell you if you're actually in the loop
Before approving anything an agent wrote, ask yourself:
- Could I explain this change to someone else without opening the pull-request description?
- Do I know what happens when the input is empty, null, duplicated, or wrong?
- If this breaks in production tonight, do I know where to look first?
If the answer to any of them is no, you are adjacent to the loop, not in it. That is fine occasionally. It is not fine as a policy.
The trade-off you have to choose on purpose
There is no clever trick here. If you enforce real reads on everything, you cap how much agent output your team can consume — which means you deliberately leave speed on the table. If you run a dozen agents and read nothing, you get velocity now and a codebase nobody can confidently change later. Both are choices. The problem is teams making the second choice while believing they made the first.
So go back to my friend who hasn't typed code in ten months. He isn't doing anything wrong. Writing was never the hard part of the job — it was just the most visible part. The job moved, and what it moved into is harder than it looks: reading code that is usually clean-looking whether or not it is correct, and deciding what deserves to exist. Ten months of merged pull requests is a lot of code with your name on it. The question is how much of it you could explain from memory.
Key Takeaways
- The bottleneck moved. Agents made generation cheap and parallel; human reading is still serial and expensive. That queue is your real constraint.
- A loop stops work. A reviewer who can halt the change is in the loop. Someone who gets a notification, or taps approve from a phone, is not.
- Track three numbers. Change entering review per engineer per week, median time to first review, and two-week re-open rate. Crude beats nothing.
- Tier by blast radius, not by vibe. Auth, money, deletion, and migrations get a full read and a test you wrote yourself. Throwaway scripts get deleted, not merged.
- Spend free output on the reader. Small commits, a ranked risk list, and a failing test cost the agent almost nothing and save you an hour.
- Keep a ledger of unread code. Skipping a read is a decision; recording it turns hidden debt into something you can actually pay down.