The Alarm Was the Feature: Load Shedding Lessons From Apollo's 1202

On July 20, 1969, roughly six minutes into the Eagle's descent toward the Sea of Tranquility, a display inside the lunar module flashed a code the crew had never seen on a real mission: 1202. A minute later it happened again, then a 1201. Neil Armstrong asked what it meant. Mission Control had seconds to decide whether to abort the first moon landing.
The short version you have probably heard is that the guidance computer was overloaded. The more useful version is this: the overload was handled by design. A radar that had been left on was flooding the computer with data it did not need during descent. The computer's executive — the part of the software that decides which job runs next — ran tasks in strict priority order. As the flood of low-value work arrived, the software delayed or dropped it, kept the critical descent calculations running, and told the crew it was doing so. The 1202 alarm was not a crash. It was a status report: we are behind, but we are still flying.
Margaret Hamilton, who led the MIT Instrumentation Laboratory team that wrote that software and who died on September 30 at 90, spent her career arguing that this was not luck. She popularized the term “software engineering” precisely because she believed this kind of work deserved the same rigor as building a bridge: priorities, budgets, recovery paths, and tests. Jack Garman, an engineer in Mission Control, had already run the alarm scenarios in simulation. Steve Bales, the guidance officer, heard the code and said “go.” The landing continued with about 20 seconds of propellant left.
Here is the uncomfortable part for anyone running software today: you also have a priority system. You just did not write it down.
Your Priority Order Already Exists. You Just Did Not Write It.
When your database connection pool fills up, something decides which requests wait and which fail. When your LLM provider rate-limits you, something decides which calls get retried and which get dropped. When a node runs out of memory, something decides which containers get evicted. In every case, a policy is being applied.
The question is whether it is the policy you would have chosen.
Most teams cannot answer a simple question in under 30 seconds: name the three user-visible outcomes that must survive at any cost. If the answer is “everything,” you have no priority order, and your system will pick one for you under stress — usually “whoever got there first” or “whoever retried last.”
Modern agent systems make this worse rather than better. A retry storm is the opposite of load shedding: every retry adds load to a system that was already struggling. An agent that loops “try again, try a different approach, escalate to a bigger model” when a dependency is slow will happily turn a partial outage into a total one. Retrying is a decision with a cost, not a default behavior.
Degradation Is a Ladder, Not a Switch
The most common failure in production is treating availability as binary: up or down. Real systems have rungs, and each rung has a cost and a customer experience attached to it. A workable ladder usually looks something like this:
- Full service. Everything runs.
- Reduced. Personalization, recommendations, or enrichment is skipped; the core answer stays correct.
- Stale. You serve a cached result and label it as cached, with a timestamp the user can see.
- Deferred. You accept the request, queue it, and tell the user when it will be done. This only works if the work is idempotent, meaning it can safely run twice without duplicating its effect.
- Honest failure. You return a clear error saying what is unavailable and what to do next.
Apollo 11's ladder went straight from “full” to “we are deferring the radar work.” That was the right call because rendezvous radar data was not needed while the crew was descending. Your ladder has to make the same kind of judgment: what does this user actually need in the next ten seconds?
A useful test: if you could serve only 10% of your usual capacity, which 10% of features would you keep? Write that down. That list is your priority order, and it should be decided by someone who owns the customer relationship — not reverse-engineered from a dashboard at 2 a.m.
Make the Drop Loud
Silent degradation is the most expensive kind. If your system quietly serves stale data or drops a background job, two things happen. First, nobody knows you are in a degraded state, so nobody fixes the underlying cause. Second, when that state lasts three days, you have renamed an outage to “normal.”
Emit a distinct signal every time the system enters a lower rung, separate from your error rate. Errors mean “this failed.” Degradation means “this worked, but at lower fidelity.” Mixing them guarantees your on-call rotation learns to ignore both.
Track three things about each rung: how often you enter it, how long you stay there, and whether you ever came back. That last one is the one teams forget. A degraded mode that never recovers is not graceful; it is a permanent capacity problem you have learned to live with.
Run a 1202 Drill
Apollo's crew and flight controllers had rehearsed the alarms. You should rehearse yours, because an untested degraded path is a hypothesis, not a capability. The drill takes an afternoon:
- Pick your peak-traffic window. Write down what normal looks like.
- Make a dependency slow rather than dead. Slow is the harder case — a dead dependency fails fast and cleanly, while a slow one absorbs your threads, your connections, and your patience. Most real incidents are slowness, not death.
- Push traffic to roughly twice your normal peak, or throttle a critical resource to a fraction of its usual capacity. On a cloud provider this is usually a small configuration change.
- Watch whether the ladder fires in the order you declared. If browse traffic keeps serving while checkout dies, your priority order is wrong — or the shedding mechanism does not know about it.
Do it quarterly, and change which dependency you slow down. The point is not to prove the ladder works once. It is to keep the people who would have to operate it familiar with what it looks like.
The Four Decisions You Owe Your Team
- The critical path. Three user-visible outcomes, written in plain language, approved by whoever owns the customer promise. Not thirty.
- The priority list. Every significant capability gets a rung order and a named owner. Capabilities without owners get shed first, and they should — that is the honest answer.
- The trigger. What metric crosses what threshold, and who can pull the lever manually. Manual overrides need an audit trail; otherwise you will spend a week figuring out why checkout was disabled on a Tuesday.
- The exit. What has to be true to return to full service. If the answer is “when it feels better,” you will never leave the degraded state.
What You Are Giving Up
Load shedding is not free, and pretending otherwise is how these projects fail.
- You maintain two paths. Every rung above the bottom is code you keep alive but rarely run. That is real maintenance cost, and rarely-run code rots quietly.
- Partial results can be subtly wrong. A cached price or an unpersonalized answer looks fine and may mislead. Label it. Users forgive “this might be out of date” far more readily than being quietly misled.
- You can shed revenue by accident. If the priority list is built from infrastructure metrics instead of business outcomes, you will protect whatever is easiest to measure and drop whatever pays for the servers.
- Deferred work needs idempotency. A queue that runs a job twice will double-charge someone eventually. This is not optional plumbing.
The trade-off is real, but compare it to the alternative: when the next slow dependency arrives, you either decide what to sacrifice, or you watch the whole thing fall over and explain it afterward.
Hamilton's team made that decision years in advance, in a machine with roughly 72 kilobytes of memory, and it bought them a moon landing. You have far more memory and far more warning. Spend an afternoon writing the list down.
Key Takeaways
- Apollo 11's 1202 alarm was not a bug; it was a priority scheduler reporting that it had dropped low-value work and kept flying. The overload behavior was designed, tested, and rehearsed.
- You already have a priority order — it is baked into timeouts, connection pools, retry loops, and eviction rules. Either you choose it or your system chooses it for you.
- Degradation is a ladder, not a switch. Define your rungs (reduced, stale, deferred, honest failure) and the owner of each capability.
- Emit a separate signal for “degraded” versus “failed,” and track how often you enter a rung, how long you stay, and whether you ever come back.
- Test the shedding path on purpose, quarterly, with slowness rather than death — and change the dependency each time.
- Write down four things: the critical path (three outcomes), the priority list, the trigger, and the exit.
- Retries are load, not recovery. An agent that escalates on every failure does not degrade gracefully; it multiplies the problem.