2026-10-04 · 8 min read

A Soft Cap Is Just a Receipt: Building Budget Limits That Actually Hold

Flat isometric diagram of several indigo process nodes routing glowing request lines into one central budget gate, with a single green path passing through to success and other lines blocked.

It is 3:40 a.m. and the only thing watching your product is an email. “You have used 90% of your monthly budget.” You read it at breakfast. By then you are over by $9,400, because the thing that ran while you slept never checks its inbox.

This is the gap Simon Willison wrote about in his case for default hard budget caps on pretty much everything: “after $X/month, cut this thing off and return errors.” His argument is that soft caps — the ones that send a warning email and keep going — will not cut it. He is right, and it is worth restating in engineering terms rather than financial ones. An alert is a measurement. A cap is an admission decision. They are different mechanisms, and only one of them can stop a request from happening.

The industry is starting to agree. AWS shipped a monthly spend limit that pauses a project when it is reached, and Google Cloud shipped Spend Caps in July. Both are good news, and both carry a footnote: AWS’s version was initially rolled out to “a limited number of customers.” So if you run your own product on pay-per-use services, or run an agent that calls them, you probably cannot wait for your vendor. You have to build the cap yourself. Here is what that actually takes.

Why “check the budget first” does not work

The naive version looks like this: before you call the expensive API, read how much you have spent this month, compare it to the limit, and refuse if you are over. It reads like a solution. It is not one, for two reasons.

First, you learn the price after you pay it. For a model call, the cost depends on how many tokens came back. For storage, on how many bytes got written. You cannot decide whether you can afford something whose price is only known on the receipt. Any real cap has to reserve an estimate up front and reconcile it afterwards — the same reason a hotel puts a hold on your card at check-in.

Second, reading and then acting is a race. Your agent fans out twenty parallel calls. Each one reads “$0.40 left in the budget.” All twenty pass the check. All twenty spend. You finish the minute at $8 against a $0.40 ceiling. This is the oldest bug in distributed systems — check-then-act — and agents are unusually good at triggering it, because fan-out is exactly what they are built to do.

The five pieces a cap needs

A hard cap is a small reservation system. If you have ever implemented a semaphore, a rate limiter, or an inventory hold, you already know the shape. It needs five things:

  1. An estimate before the call. Reserve a generous upper bound: maximum output tokens times price, or the p95 cost of that job type with a safety margin. Over-estimating costs you nothing, as long as reconciliation exists.
  2. An atomic reserve. One counter, one writer, one transaction. In SQL, that means an update that adds a reservation only when the new committed total stays under the limit, and returns a row when it succeeds. In Redis, a small Lua script does the same job. Two processes must never read the same headroom.
  3. A reconciliation. When the call finishes, subtract the actual cost and return the difference to the pool. Reserve an upper bound, settle at the real number.
  4. A lease with a time limit. If the process dies between reserving and reconciling — and it will — the reservation has to expire on its own. Set that expiry at roughly four times your p99 job duration. Without it, a crash-looping worker leaks budget forever and your ceiling quietly shrinks every day.
  5. A declared failure mode. What happens when a reservation is refused? Return an error, fall back to a cheap model, queue the job for tomorrow, or serve a cached answer. Decide before your first incident, because the lazy answer — a generic 500 — turns a cost control into a customer-facing outage.

One chokepoint, not four call sites

A cap only holds if every path to money goes through it. Most teams have three or four helpers that call the vendor: one in the API service, one in the worker, one in a script somebody wrote during a hack day. Enforcement lives in three of them. The fix is boring and effective — wrap the vendor SDK behind a single client module that nothing else imports, and make a reservation a required argument to it. Then prove it: set a limit of $0.01 in staging and confirm the request is refused before the socket opens. If you cannot make your cap refuse on purpose, you have not verified it.

If you run agents in containers you do not fully trust, apply the same rule one level down and enforce at the egress proxy rather than inside the agent’s process. There is a limit to how much a cap inside the thing being capped can protect you from.

What a cap costs you

None of this is free, and the costs are not hypothetical.

  • Legitimate work gets cut off. A monthly number is an absurd unit for a batch job that costs $40 and runs for twenty minutes. Use two layers instead: a small per-job cap that protects you from a runaway loop, and a much larger global cap that protects the company. Set the per-job cap near three times your p99 unit cost, and the global one near four times your seven-day median spend.
  • You create a new failure mode. The worst moment to hit a monthly ceiling is the last hour of the last day of the month, which is exactly when a busy month ends. Pausing a project is honest, but it is still an outage. Say so in your status copy and your customer email, not just in the runbook.
  • Somebody will need to raise it. Build the raise path as a human action with an audit trail, and treat it as a change to production configuration rather than a retry.

Measure whether the cap is binding

A cap you never hit is a cap you cannot trust. Three numbers belong on the same dashboard as your latency:

  • Refusal rate. How many reservations were refused this week? If it is zero for months, either your limit sits far above your usage or the enforcement path is dead. Both are worth knowing.
  • Reservation leak. Reserved minus reconciled, aged past the lease. Anything non-zero means your accounting is drifting away from reality.
  • Cost per completed unit of work. Dollars per finished task, not per token. It is the only number that tells you whether a change made the product cheaper or merely made the invoice look different.

Why agents make this urgent

For most of software history, spending money required a human to write code and deploy it. That friction was not a feature, but it was a brake. Agents remove it: the same property that lets one person ship a useful service lets that service ship a bill. A loop told to keep trying until it works will cheerfully retry a failing paid call thousands of times, and every retry is a fresh reservation against your ceiling.

The same agents can work in your favour. Willison suggests agents should bias toward providers that offer hard caps and warn newer builders away from uncapped services. That is a good default — your tooling can enforce a policy your team never wrote down.

Fifteen years of observability taught us to alert on everything. Alerts are for the things that are still going to happen. A cap is for the thing you have decided will not. If the only control between your product and a five-figure invoice is an email at 90%, you do not have a budget. You have a receipt.

Key Takeaways

  • An alert measures spending after the fact; a cap decides whether a request may happen at all. Only one of them scales.
  • You cannot cap a price you only learn after the call. Reserve an estimate, then reconcile the actual.
  • Naive check-then-call logic loses to concurrency — your own agent’s fan-out will spend past the limit twenty times over.
  • Give reservations a lease with a time limit, or a crashed worker leaks budget permanently.
  • Enforce at one chokepoint — an SDK wrapper or an egress proxy — and prove it by making the cap refuse a request in staging.
  • Use two layers: a per-job cap near 3x p99 unit cost and a global cap near 4x your seven-day median.
  • Decide the refusal behaviour in advance: error, degrade, or queue. A generic 500 turns cost control into an outage.
  • Track refusal rate, reservation leak, and cost per completed unit of work. A cap you never hit is a cap you cannot trust.

[ Call to action ]

Ready to scale your engineering team?

Tell us the roles you need to fill and we'll get back within 24 hours.