Your Oldest System Is One Retirement Away From an Outage

In 1993, a brand-new Stratus fault-tolerant server was booted at a steel-coating plant in Dearborn, Michigan. Twenty-four years later it was still running: no unplanned shutdowns, roughly 80% of its original components, an operating system untouched since the early 2000s, no vendor maintenance contract, and spare parts sourced from a third-party reseller. The machine was so dependable that its owner had already won Stratus’s own longevity contest in 2010. Computerworld reported the whole story in 2017.
Then the plant changed hands, a modernization plan appeared, and the server finally got a retirement date.
Here is what almost every retelling of that story skips. The impressive number was not 24. It was one: the number of people who could say, honestly and with evidence, that the system would not fail on its own, and who knew exactly what to do on the rare occasions when it did. When that server is switched off for the last time, the hardware loses nothing. Its owner loses a career’s worth of context that was never written down anywhere.
If you run anything old — a mainframe batch job, a 2011 Rails app, a data pipeline held together by cron and hope — the same arithmetic applies to you. The migration is not the risk. The knowledge transfer is. And unlike hardware, people don’t announce their failure six months early with a blinking amber light.
The hardware was never the fragile part
Look at what fault tolerance actually bought: redundant disks, redundant power supplies, redundant everything. Components were swapped over the years, but close to 80% of the system was original. That is a hardware success story, and it is the kind of thing vendors love to put in a case study.
Now look at what kept it running in practice: one person who had been paying attention for two decades. He knew which errors were noise and which ones were the prelude to a very bad night. He knew which third-party vendor still stocked parts, and which parts to hoard. He knew the sequence that looked harmless but wasn’t, and the sequence that looked alarming but happened every Tuesday at 4 a.m.
That knowledge is not documentation. Documentation tells you what happens. Experience tells you what happens when it doesn’t. A runbook cannot contain a sentence like “the users like the reliability of it, and the screens are actually pretty simple” — which is a real observation from that site, and a real requirement that never made it into a spec. If you modernize a plain character interface into a polished one with three extra clicks, you can make people slower and less accurate while shipping something objectively nicer. You won’t find that in the code. You’ll find it in the person who watched them use it every day for twenty years.
Three kinds of knowledge, one of which is flammable
Before you plan anything, sort what lives inside that system into three buckets. Their recovery costs are wildly different.
- What it does. Inputs, outputs, schedules, job chains, downstream consumers. This is mostly recoverable from logs, database schemas, and file drops. Tedious, but a competent newcomer with three weeks and read access can rebuild the map.
- Why it does it that way. The workarounds, the regulatory quirk from 1998, the reason job 42 must finish before job 43 even though nothing obvious depends on it. This lives in one or two heads. It is not recoverable from artifacts, because the artifacts describe the solution, not the problem.
- Who depends on the weird part. The person in accounting who exports a file every month and does something with it that nobody has asked about in a decade. This knowledge is usually held by different people than the system owner — and neither group knows the other’s half.
The first bucket shows up in your migration plan. The second shows up in month two as a “surprise requirement.” The third shows up after cutover as a phone call.
Your deadline isn’t technical. It’s biological.
Notice how the server got its upgrade date: ownership changed, and the new owner had a plan. The story is clear that upgrades had been eyed for years and that business cycle changes kept derailing them. Nothing technical blocked it. Priority did.
People work the same way. An expert doesn’t leave because the system failed. They retire, take another job, get reorganized onto a team where the old system isn’t their problem, or simply get tired of being the only person who can be paged at 2 a.m. When that happens, the knowledge doesn’t degrade gracefully. It goes to zero in a single afternoon.
So here is a rule you can apply this quarter: treat a system as P1, not negotiable, when the number of people who understand its failure modes is one — regardless of how healthy the hardware looks. Two is fragile but survivable. Three is a real system. And if that one person is within five years of retirement, or has already asked to be rotated off, the clock is not theoretical.
A 90-day knowledge cutover plan
You already know how to migrate a workload. Here is what to add around it, with rough timings you can stretch or compress.
- Days 0–10: inventory behavior, not code. Pull twelve months of logs and list every scheduled job, every file that leaves the system, every inbound feed, and every human who touches it. Aim for a list of consumers by name, not by team. If you can’t name them, you don’t have an inventory.
- Days 10–30: classify each behavior. Mark it load-bearing (it runs and something depends on the output), ritual (it runs, nobody can explain why, no consumer found), or zombie (it hasn’t succeeded in a year but nobody dared delete it). Zombies are free wins. Rituals are where you find the regulatory quirk, or the thing that will silently break at cutover.
- Days 30–50: write the failure playbook, not the feature spec. The last five incidents: symptom, diagnosis, fix, time to fix. Which alerts are noise. Which errors are benign. Which vendor still stocks parts. What the manual workaround is when everything else has failed. One page per critical behavior beats forty pages of architecture diagrams.
- Days 50–70: run shadow sessions in both directions. First the newcomer performs the task while the expert narrates. Then the expert performs it while the newcomer narrates. The second direction is the one that exposes false confidence, because you find out what the newcomer thinks is happening — and it is usually wrong in an interesting way.
- Days 70–90: parallel run, then close the window. Run old and new side by side on real data for at least two full cycles, including one month-end, because that is when the ritual jobs earn their name. Then set an explicit date after which questions are no longer free, and back it with a paid consulting retainer.
Budget the human, not the binary
Migration estimates usually cover infrastructure, licenses, engineering weeks, and testing. They almost never cover overlap time for the one person who understands the system. That is an accounting error, not an optimism error.
Write two things into the plan. First, real overlap: two person-weeks of the expert’s time per critical behavior, scheduled and defended, not “available if needed.” Second, a paid exit deliverable — the failure playbook, the consumer list, and a recorded walkthrough — as an explicit scope item with a date. If knowledge transfer isn’t in someone’s performance goals, it doesn’t happen.
There’s a hiring angle too. You cannot recruit a specialist for a platform that shipped before many of your candidates were born. What you can do is hire for pattern recognition and pay for transfer time. So ask about the undocumented system they kept alive: what did they have to figure out alone, what broke first, and how did they find out. A developer who spent three years as the only person who understood a pile of cron jobs has debugging instincts no take-home exercise will surface. Treat that experience as a senior signal, not a résumé gap.
Sometimes the right answer is to leave it alone
Fair counterargument: not every old system deserves a migration. If it serves one department, runs three stable jobs, has contained vendor risk, a parts supply, and a documented failure playbook, then “keep it, buy spares, revisit in two years” can be the cheapest correct decision on the table. A big-bang rewrite nobody asked for is a way to convert working software into an outage.
But that decision is only defensible after the capture work. Without it, “leave it alone” isn’t a strategy. It’s a bet that the one person who understands the system will never leave — a bet with a payoff you’ll never see and a downside you’ll meet at 2 a.m.
When that Dearborn server is finally powered down, nothing will break, because a plan exists and the people are ready. That is the whole point. The goal was never to keep a 1993 machine alive. It was to make the shutdown boring, and to make sure the knowledge inside it survived the power cut.
Key Takeaways
- Long uptime on old hardware is a distraction. The fragile asset is the small number of people who understand why the system behaves the way it does.
- Sort knowledge into three buckets: what it does (recoverable from artifacts), why it does it that way (one head), and who depends on the odd parts (often a different head entirely).
- Treat a system as P1 when only one person understands its failure modes, especially if that person is near retirement or asking to rotate off.
- Run a 90-day knowledge cutover: inventory consumers, classify behaviors, write the failure playbook, shadow in both directions, parallel run, then close the question window with a retainer.
- Budget overlap time and a paid exit deliverable. Knowledge transfer that isn’t scoped, dated, and owned does not happen.
- Keeping an old system is legitimate — but only after the knowledge is captured. Otherwise it’s a single-person bet.