2026-10-06 · 8 min read

A Watermark Is a Maybe, Not a Verdict: Building Policy Around AI-Content Detection

Flat isometric illustration of translucent document sheets passing through violet filter rings, an indigo wave pattern fading into fragments, and one sheet keeping an intact pattern ending at a single green node.

A watermark that vanishes when you swap 10% of the words for synonyms is not a compliance strategy. It's a signal — and signals need thresholds, policies, and a human on the other end. Most teams shipping AI features into Europe are about to find that out the hard way.

On Monday, OpenAI said it will start adding an invisible watermark to text generated by ChatGPT and Codex inside the European Union, to comply with the EU AI Act's transparency rules that took effect on August 2. The method, called textGrain, doesn't stamp a symbol into the text. It subtly shapes the model's word choices with a secret key, so a detector holding the same key can spot the pattern. The mark rolls out to EU users on all plans; it's off by default for API developers worldwide; and detector access is limited to approved researchers and expert organizations. Anthropic went further and applied its text watermark worldwide.

You are probably not building the watermark. You're on the receiving end — piping model text into a product, a document pipeline, a CMS, a hiring process. And that's where the trap sits. The moment a probabilistic mark becomes a checkbox in your spec, you have turned a statistical hint into an accusation, with no error bars attached.

A watermark is a detector, not a proof

OpenAI published numbers that make this argument better than I can. In one test, replacing 10% of words with synonyms dropped detection from about 92% to 66%. Short passages, math answers and translated text are harder to detect. Then the company said the quiet part plainly: a missing watermark 'does not prove human authorship.'

That sentence is the whole article. A watermark can suggest that an OpenAI system generated or processed part of a passage. It cannot tell you how much human judgment went into it, whether another vendor's model wrote it, or whether a person edited the output heavily enough to erase the pattern.

Now the part most write-ups skip: you cannot read a detector alarm without two rates. The true-positive rate tells you how often AI-written text gets flagged — that's the 92%. The false-positive rate tells you how often human writing gets flagged, and it isn't published. Without it you're holding a likelihood, not a verdict. Run the numbers on a concrete pile:

  • 1,000 documents, 100 of them substantially AI-written. A detector with 90% true positives and 2% false positives flags 90 real ones and 18 humans. That's 108 flags, 83% of them correct. Fine for a review queue. Catastrophic as an auto-reject.
  • Same 1,000 documents, but only 10 AI-written. The same detector flags 9 real ones and 20 humans. Now roughly 69% of your flags are wrong.

Same detector. Same error rates. Wildly different meaning, because the meaning depends on the mix, not the model. This is exactly the math that turned antivirus alerts and spam filters into review queues instead of delete buttons twenty years ago. We already learned this lesson once. It just came back wearing a new hat.

Trap 1: the binary gate

If your product says 'AI-generated: yes or no', you've built a machine that emits certainty it doesn't have. Offer three states instead — flagged, clean, unknown — and route only flagged items to a human. Make unknown the default for anything short, translated, or heavily edited, which in production is most of it.

Trap 2: you can't verify your own output

The mark depends on a secret key and a detector OpenAI is handing to approved researchers only. That's a defensible safety call, and it also means you cannot run the check on your own text as a regular customer. If your roadmap promises 'watermark-verified provenance,' you are promising something you can't independently test. Write that sentence into the design doc before someone writes it onto a marketing page.

Trap 3: your own pipeline erases the mark

Watermarks live in word choices, so they travel with copy-paste and die in every well-meaning rewrite. Summarizing, translating, SEO rewriting, tone adjustment, markdown-to-CMS conversion, a human editor tightening the prose — each is a paraphrase-laundering pass, whether anyone intends it or not. If your compliance story is 'the mark will be there when someone checks,' your compliance story is a coin flip.

There's a one-day experiment that fixes this ignorance: take a sample of your AI-generated output, run it through your real production pipeline end to end, and measure how often the mark survives. Put that number on a dashboard. If it falls from 90% to 40%, at least you know. Right now, most teams know nothing at all.

Record provenance where you create the artifact

Detection is archaeology: inferring what happened by inspecting the remains. Provenance is bookkeeping: writing down what happened while it happens. Bookkeeping is cheaper, more accurate, and entirely under your control. For every artifact your system produces, log:

  • run id, model name, model version, provider;
  • a hash of the prompt and of the inputs that mattered;
  • which tools ran, and the identifiers of what they returned;
  • a timestamp and the account that kicked it off;
  • every later human edit, with who made it.

You don't need a cryptographic protocol to get most of this value. You need rows in a table nobody deletes. When a customer, an auditor, or a journalist asks whether a model wrote something, you answer from your own records instead of guessing from a word pattern.

For artifacts you control end to end — signed reports, images, PDFs, code commits — you can go further with cryptographic content credentials such as C2PA manifests, which bind provenance claims to the file itself and can be verified offline by anyone. Watermarks and signatures aren't competitors. One is a weak statistical signal for text you're about to lose control of; the other is a strong claim on artifacts you keep.

Write the response policy before the first flag

Someone will paste a detector result into a ticket and ask you to act on it. Decide now what happens:

  1. No single signal decides anything. A flag opens a review, not a verdict.
  2. Reviewers see evidence, not a score — the matched span, its source, and the error-rate context.
  3. Every decision gets logged with its reason, so clusters of false positives become visible.
  4. The affected person can respond, especially when they're a student, a candidate, or a contractor whose income depends on the outcome.
  5. Track the queue's precision over time, the way a fraud team tracks chargebacks.

The urge to skip this is strong, because the tool feels objective. It isn't. It's a classifier with an unpublished error profile, and treating it as an oracle is how you end up in a headline about wrongly accusing someone — which costs far more than building a review path ever would have.

What this means in practice

If you publish model output: turn watermarking on where it's available, log provenance at generation time, and never claim in your docs that the absence of a mark means a human wrote something.

If you build content tools: expose an unknown state in the UI. Your users will get flagged for writing they did themselves, and they need somewhere to say so.

If you hire or evaluate: don't put a detector in the filter. Use it as one input next to the things that actually distinguish people — a portfolio walk-through, a live discussion of their own code, questions about decisions they made six months ago.

One honest trade-off: watermarking can nudge a model's word choices, and some users argue that they supplied the instructions, the context, and the decisions, so labeling the text as machine-made misrepresents their work. That's a real objection, not a technicality. The EU rule asks for machine-readable marking; it doesn't ask you to make a claim about authorship. Keep those two jobs separate in your product copy and in your policy.

Key Takeaways

  • A watermark is a probabilistic detector: 92% detection drops to roughly 66% after a 10% synonym swap, and the false-positive rate isn't published.
  • Without a known false-positive rate, a flag is a likelihood, not a verdict — and on a mostly-human pile, most flags can be wrong.
  • You can't verify the mark on your own output; the key and detector are vendor-held and access-limited.
  • Your own pipeline — translation, summarization, editing, CMS import — erases the mark, so 'no watermark' never means 'human' (OpenAI says so itself).
  • Log provenance at generation time: run id, model version, prompt hash, tool calls, human edits. Bookkeeping beats archaeology.
  • Use cryptographic manifests for artifacts you control; treat statistical watermarks as one weak signal among better evidence.
  • Write the review policy before the first flag, and never let a single classifier be the decision.

[ Call to action ]

Ready to scale your engineering team?

Tell us the roles you need to fill and we'll get back within 24 hours.