watts.it.com // daily AI micro-learning
The landscape provenanceevidenceverificationdisclosurejudgment 2026·08·27 · 4 min · dated

A flag is not a verdict

// listen · 2 ai hosts · audio edition

AI-generated audio discussion of this module — same content, spoken.

Overview

Why now. As of August 2026, text coming out of one of the frontier assistants arrives with an invisible watermark in it — applied worldwide, and carried through the big cloud platforms into whatever your employer has built on top, without your employer choosing it.

The sentence worth your attention is in the maker’s own help documentation. A detected mark means the content “may have been processed by” the model — not that the model wrote it. The mark attaches to whichever words the model chose, so a paragraph you accepted from a suggestion carries it exactly as one drafted from nothing does.

That distinction is about to land on desks that are not ready for it, because two different technologies are both called “AI detection” and their failures are nothing alike. So: when you are shown a result saying something was AI-written, ask which kind of check produced it — a watermark from the model’s own maker, or a third-party detector’s guess — before you treat the answer as settled.

The content

A watermark is put there as the text is written. The maker holds a cryptographic key and uses it to settle the choice between words that would have been equally good either way. Nothing looks different to you, but the pattern is there and the maker can check it. That makes it strong in one direction: when a watermark says yes, it says so with real force.

Its weakness runs the other way: it goes quiet. Marks thin out on very short passages, where there is too little text to carry a signal, and on factual writing, where the wording is constrained. They are sparse in code. And they may vanish once someone has heavily edited, paraphrased, translated or remixed the text. Google describes the same limits in its own scheme, which suggests the behaviour is inherent to the technique.

So a clean result tells you very little — less still just now, with the rollout to older models still running and no detection tool released yet, though the maker says one is coming.

You would assume, then, that the watermark is the strict instrument and the detector the loose one. Half right.

A detector holds no key and saw nothing. It is a statistical guess made afterwards by a stranger to the text, and it can fail in both directions — including the one that lands on a person. That was the alarm in 2023, when a study of seven early detectors reported heavy bias against writing by non-native English speakers, and it is still the version most people carry.

A 2026 study revisits that story and corrects it in both directions at once. Researchers tested 160 long papers across four tools, using human papers written before 2019 by non-native-English graduate students — precisely the group said to be penalised. Three of the four produced zero false positives. The fairness problem that dominated the coverage did not reappear, and that is worth saying plainly rather than letting a frightening old number stand.

It also found these tools miss far more than they invent. Against text put through a humanising step, one tool’s detection rate fell to 22.5% and another’s to 2.5% — the commoner error in both studies, and it points the opposite way to the fear. The authors’ conclusion is the sentence to carry out of all this: detection tools “should not be used as sole evidence in high-stakes decision-making.”

Two instruments. Both informative, neither proof, and they miss on different days.

So when a result lands in front of you: if it is a watermark, a hit is evidence a model touched the text and none at all about who wrote it — an accepted suggestion leaves the same trace as a wholesale draft. An absence is not a clearance. If it is a detector, a flag is a reason to go and look at the work: the drafts, the sources, the person. A prompt to investigate, not the finding.

The reverse case matters more, because you are likelier to be the subject of one than the person holding it. What settles it is almost never the tool — it is whether you can show your working.

Try it

Take a document you produced with AI help this month.

  1. Rehearse the reading. Ask your assistant: “Two things get called AI detection: a watermark placed by the model’s maker, and a third-party statistical detector. For each, tell me what a positive result does and does not establish, and what a negative result establishes.” Check whether it keeps them apart. If it blurs them, you have just watched the confusion happen.

  2. Build the record instead. In one or two lines in your own notes, write what the assistant actually contributed — drafted, restructured, proofread, translated, checked — and where the source material came from. Thirty seconds, and the only artefact that reliably answers the question later.

Make it a habit where you can. If your workspace lets you save instructions or a reusable prompt, add a line asking the assistant to summarise its contribution at the end of a drafting session. Check what your workspace has enabled; where saved instructions aren’t available, keep it as a copy-paste.

Where this breaks. None of this helps if the result arrives as an accusation rather than a question. Then the move is not to argue about accuracy but to ask what it is evidence of, and to offer the record.

Additional reading

Editor’s note

Look, I hate this. I am not a supporter. There is an ethical ignorance to implementing a feature that, when a model is presented with my own work for consideration, will push for that work to be marked in a manner that implies the work was done originally by the model. But the decision is made, and the parameters of its implementation are settled, so this module exists as an explanation of how it works. Good luck.

signed-off-by: Luke Topfer <editor> · 2026·08·27
06 Self-check

// three assertions against what you just read · results stay in this browser

assert 1/3

The module says two different technologies are both called "AI detection". What is the difference the module asks you to hold on to?

assert 2/3

A colleague forwards you a third-party detector's report flagging a supplier's proposal as AI-written, and asks whether that settles it. Following the module, what is the right next step?

assert 3/3

According to the module, what does a clean result — no watermark detected — tell you right now?