Your Agent Is a Demo Until It Survives a Restart

Three teams independently shipped the same word this week, and it was not the name of a model. Earendil shipped Pi 1.0 alongside an experimental package called Pi Durable. DeepSeek opened the public preview of DeepSeek Harness. Different codebases, different philosophies, same noun.
That noun is harness: the storage and machinery that runs conversations with a language model, gives those conversations tools, and records what happened. The model is now the easily replaced part. You can swap providers with a config string. The harness is what decides whether your agent is a product or a demo, and the difference almost never shows up in a benchmark.
What a harness actually is
Strip away the branding and a harness is four things:
- A transcript. The record of an interaction between a person and an agent.
- An agent definition. The model, its thinking budget, and the tools it is allowed to call.
- Tools and an execution environment. Where the work actually happens: your laptop, a remote VM, an in-memory sandbox.
- Storage. Where the transcript lives between runs.
Pi Durable's documentation is unusually blunt about the last one. A harness opens over a storage backend, and Pi ships three: memory, SQLite, and JSONL. One process owns a storage at a time; other clients attach to that process. Tools get their environment from the conversation's working directory, which means conversation A can run on your laptop while conversation B runs somewhere else entirely.
The other half of the design space is composition. DeepSeek Harness describes itself as everything is a plugin, with plugins that add tools, skills, and interface. Pi took the opposite route: the Pi 1.0 release note is mostly a list of things it refused to adopt, on the grounds that agentic tooling changes every week and most changes do not last. Both positions are defensible. What matters for you is that neither of them is a model question.
The one test that separates a demo from a product
Here is a test you can run in a meeting, on a vendor, or on your own codebase, and it takes about ninety seconds.
What happens to the run if the process dies right now?
Not does it error gracefully. Not does it retry. What happens to the twenty minutes of work, the four files it edited, and the two external APIs it already called. If the honest answer is we look at the logs and start over, you have a coding assistant, not a durable agent. That is fine. A terminal coding agent driven by one person is a completely legitimate product, and it is what Pi 1.0 set out to be. It is not fine if six people are steering the same agent over three days.
Durability has three observable properties, and you can check each one:
- The unit of state is the transcript, not the process. A restart re-opens the same root conversation instead of creating a new one. The conversation has an identity that outlives any particular machine.
- The working set is bounded. Pi Durable keeps only active transcripts, live tasks, and pending submissions in memory. Everything else stays on disk until it is needed. Active transcripts stay small because compaction summarizes older messages before they overflow the context window, so a conversation with tens of thousands of messages still fits in memory.
- Tools are re-derivable. If a tool call needs a shell or a filesystem, the harness rebuilds that environment on demand from the conversation's metadata. Nothing important lives only on a process stack.
Why just retry it does not survive contact with production
Retrying is the obvious fix and the wrong default, because a retried agent is a retried side effect. If the process died after the tool call that sent the email and before the transcript recorded it, your retry sends the email twice. If it died after the transcript recorded the call but before the call succeeded, your retry skips work that never happened. Neither failure is visible in a log line that says task failed.
This is the same trap as any distributed write, wearing an LLM costume. You need at-least-once execution plus idempotent tools, or exactly-once bookkeeping plus a recovery path that can tell done from maybe done. Durable harnesses are popular right now precisely because that bookkeeping is tedious, and it is the part every team rebuilds badly.
The trade-off is real, though, and worth saying out loud. Durability means you now own a database. You own a schema, a migration story, and a backup story. You own the case where two clients attach and disagree. For a single-user terminal tool, all of that is overhead with no payoff, which is presumably why Pi's authors kept it in a separate experimental package rather than bolting it onto Pi 1.0. A durable backend is not an upgrade for an agent you run in a terminal. It is a different product.
Durable state is not durable truth
Once you can survive a crash, a new failure mode appears: you can faithfully resume a conversation built on a fact that stopped being true three weeks ago.
A neat illustration comes from an agent built to answer which AWS limit is actually current. It reads typed facts from a structured knowledge base and cross-checks them against the live AWS API. On the EC2 on-demand vCPU quota it found three numbers: 32 in an official user guide, 5 in the reconciled record, and 16 in the live account. On the gp3 EBS IOPS limit it found 80,000 current, 16,000 superseded, and the stale 16,000 ranked highest by keyword search because the old doc used the exact phrase a person would type.
That is not a durability bug. It is the honest answer, and it is a design requirement. A durable agent needs to record when a fact was verified, not just what it was. If your transcript stores the vCPU limit is 5 with no timestamp and no source, your crash recovery is now reliably reproducing a lie. Store provenance next to the value, and re-verify anything that decays.
What this means for how you hire and how you buy
Interview questions about agents have gotten lazy. Explain how RAG works tells you whether someone read a blog post. Try these instead:
- Your agent edits a config file, then the pod is evicted. Walk me through what state exists when it comes back, and which tool calls are safe to replay.
- Two users want to steer the same conversation. Where does the lock live, and what happens when the second one sends a message before the first one's tool call returns?
- Your transcript is 40,000 messages. What is in memory right now, and what is on disk?
Anyone who has actually run an agent in production will have an opinion about the third one. It is also the question that reveals whether someone has only ever used an agent inside a chat window.
The same shift applies to buying. Vendor comparisons today are mostly model comparisons: which model, which context window, which benchmark. Those numbers will be obsolete in a quarter. Ask instead about storage backends, conformance tests, resume semantics, and what happens when a long-running conversation hits a schema change. Pi Durable's own note is a useful benchmark for how much code you are really adopting: the entire source, minus tests, is about 15,000 lines, and the storage backends alone are about 3,000 lines that an agent can usually skip. That is small enough to read, which is the actual point. A harness you cannot audit is a harness you cannot debug at 2am.
Start with the failure, not the framework
The practical version of all this is boring and cheap. Write down the five ways your agent's process can die: laptop sleeps, container redeploys, host runs out of memory, network partition, user closes the tab. For each one, write down the exact state you would need to resume without duplicating work. If that list is short, you do not need a durable harness. If it is long, you have just written your storage schema, and you can adopt one instead of inventing it.
Models will keep changing. Restart semantics will not. That is why the harness is the product.
Key Takeaways
- A harness is storage plus the machinery to run conversations in parallel: transcript, agent config, tools, execution environment. The model is the swappable part.
- The core test is what happens to the run, and to already-executed side effects, if the process dies mid-task. Start over is only an acceptable answer for single-user tooling.
- Durability has three checkable properties: state lives in the transcript, the in-memory working set is bounded by the context window, and tool environments are re-derivable.
- Retrying a durable task means retrying side effects. You need idempotent tools, or recoverable bookkeeping that distinguishes done from maybe done.
- Durable state can preserve stale facts. Store provenance and verification timestamps next to values, and re-check anything that decays.
- When comparing agent platforms, compare backends, resume semantics, and auditability, not model benchmarks that expire in a quarter.