2026-10-01 · 7 min read

Bit-Exactness Is a Spec, Not a Hardware Property

Flat isometric illustration of two parallel stacks of identical indigo and violet block plates rising from a chip and a cube, linked by thin glowing indigo lines with one green node marking a matching checkpoint.

Two implementations of the same neural network. The first runs on an NVIDIA GPU using FP8 arithmetic on tensor cores, the specialized matrix-multiply hardware that makes modern AI fast. The second runs in a browser on WebGPU, with no tensor cores, no FP8, and no fusion between layers. Different execution model, different precision format, different engineering effort entirely.

They produce the same bytes. Not similar images. The same bytes.

The project is OpenDLSS-NR, a Vulkan reimplementation of a neural rendering network. Seventy-one transformer blocks across six pooling levels, 141 MiB of weights, FP8 (E4M3) activations with FP16 accumulation. Its README claims the intermediates match the original, not just the final image: all 75 block boundaries, byte for byte. And a second, independent port puts the same network in a browser and matches those same captures.

Most teams would call that impossible, or pointless, or both. It is neither. It is a decision about where you put your correctness contract, and it is one of the highest-leverage choices in any system that runs in more than one place.

"Bit-exact" is not the same as "looks the same"

Here is the trap. If you define correctness at the output, the image, the answer, the ranked list, you have a weak test. Dozens of different internal states can produce output that passes. Floating-point arithmetic does not associate the way math class taught you: (a+b)+c and a+(b+c) can differ in the last bit, and once one layer drifts, everything downstream inherits the drift.

So an output-only check tells you that nothing exploded. It does not tell you where anything changed.

OpenDLSS-NR publishes the whole interior instead. Every intermediate tensor at every block boundary has to match a reference capture. Not the final frame, the parts in between. With 75 boundaries checked, an equivalence claim stops being a vibe and becomes a diff you can read.

That is expensive. It is also why the project ships the tooling you would want if you took this seriously: a CPU reference implementation of the arithmetic, captured fixtures to compare against, a mode that reports per-dispatch timings, and a bisect mode that walks the very first block kernel by kernel. If the first block diverges, you never look at the seventy-first.

The contract gets cheaper as it gets coarser, and you pay later

The trade-off is direct, and it is worth naming plainly.

  • Coarse contract, output only: maximum freedom inside. You can swap kernels, reorder accumulation, fuse layers, change precision. Minimum proof. When accuracy moves 0.4% after an upgrade, you have no idea which of 241 compute steps did it.
  • Fine contract, stage boundaries: less freedom, far more proof. You can localize a regression to a single stage in one run.
  • Full bit-exact: almost no freedom inside, and portability you cannot get any other way.

The fine print matters here. FP8 with FP16 accumulation is itself a specification. Rounding mode, accumulation order, where you permit a fused multiply-add: these are not hardware facts. They are choices you can write down. Once written down, they can be reproduced on hardware that has none of those features.

The browser port is the proof. No tensor cores, no FP8, no fusion between blocks, and still bit-exact against the same fixtures. As the project's own README puts it, "the exactness is in the specification, not in the hardware."

Notice what stays free even under that constraint. The fast path on the GPU swaps cooperative-matrix GEMMs for mma.sync instructions, adds asynchronous copy rings, and chains kernels without barriers. All of that is performance engineering with zero effect on observable state, because the contract lives at the boundaries and nothing above them is allowed to notice.

What it buys you, and what it costs

What you gain: you can replace a shipping implementation and prove nothing behavioral changed. You can run the same model on three backends and know they agree. Debugging becomes bisection instead of archaeology. And you can accept an aggressive optimization without wondering whether you just quietly degraded the product.

What you pay: you give up fused shortcuts that change rounding. You need fixture discipline, meaning captures generated from a known-good build, dated, versioned, and regenerable. You may leave speed on the table. And you take on a maintenance burden that only makes sense if something downstream actually depends on the guarantee.

So decide by situation, not by taste:

  • Go bit-exact when you are rewriting something that already ships, when the same model runs on more than one backend, when debugging costs more than compute, or when an external party will check your numbers.
  • Skip it when the output is consumed by human judgment. For a summary, a recommendation, a chat reply, exactness buys you almost nothing. There, a tolerance band plus a distribution check catches the failures that exactness would never surface.

That second case is the bigger one in practice, and pretending otherwise is how teams end up with brittle golden-file tests that break on every dependency bump and teach everyone to ignore them.

Three habits worth stealing even if you never go bit-exact

  1. Put the contract at a boundary you can observe. Set thresholds per stage, not only at the end: logits within 1e-3, embedding cosine similarity above 0.999, final result set exactly equal. Then report which stage broke, not merely that something did.
  2. Treat fixtures as build artifacts with provenance. A golden capture should carry a date, a source build hash, and a regeneration command, and regenerating it should require a review. An unversioned fixture is a rumor with a test runner attached.
  3. Bisect from the front. Check the first stage first. Most teams debug backward from the symptom, which means walking seventy blocks to find a bug sitting in block zero.

The measurement discipline is part of the spec

Two details in that README deserve a highlight, because they are the kind of honesty that is rare in benchmarks.

First: 241 dispatches at every resolution. Constant. That means when you compare 768x768 to 4K, you are comparing per-kernel time, not launch overhead. If the dispatch count moved with resolution, the numbers would be telling you a story about the CPU instead of the GPU.

Second: the GPU alternates between two clock states under sustained load, so medians run a few percent higher than the minimum. The advice is to compare minima. That instinct generalizes. When your instrument drifts, change the statistic, not the story. A benchmark whose median sits 4% above the floor is not broken. It is measuring a different thing, and if you do not say so, your readers will find out for you.

Then there is the number the project volunteers: 72 ms at 512x512 in the browser against 2.7 ms on the native path. Roughly 27 times slower. That is fine, because the browser port was never trying to be fast. It was trying to test whether the specification was real. It turned out to be the cheapest, and one of the most convincing, tests available.

One question that separates engineers

If you are hiring, ask this: "You have ported a model to a new backend. How do you know it is correct?"

The answers divide people quickly. "I ran the tests" is output-level thinking. "I compared intermediate tensors stage by stage and bisected on the first divergence" is someone who has been burned and knows where the bodies are. Neither answer is about a framework or a GPU vendor, which is exactly why the question is worth asking. It measures whether the candidate thinks in contracts or in vibes.

Key Takeaways

  • Output-level equality is a weak claim. Many internal states produce the same picture, and none of them tell you where a regression started.
  • Numerical behavior is a specification you write, not a property of the hardware you happen to run on. FP8 rounding and accumulation order are choices, and choices can be reproduced anywhere.
  • Push your correctness contract to the coarsest boundary that still lets you localize a failure. For OpenDLSS-NR that was all 75 block boundaries, and it made a browser port with no tensor cores possible.
  • Bit-exactness is worth it for replacements, multi-backend deployments, and externally audited numbers. It is close to worthless for anything a human judges case by case.
  • Version your fixtures, bisect from the first stage, and when your measurement instrument drifts, change the statistic rather than the story.

[ Call to action ]

Ready to scale your engineering team?

Tell us the roles you need to fill and we'll get back within 24 hours.