watts.it.com // daily AI micro-learning
Judgment & limits reviewverificationcheckingjudgmentdrafting 2026·09·07 · 4 min · dated

What's missing doesn't look wrong

// listen · 2 ai hosts · audio edition

AI-generated audio discussion of this module — same content, spoken.

Overview

Why now. In January 2026, researchers recruited a hundred people online and showed them charts with something quietly left out of the data. Asked what they noticed, one group flagged the missing piece in 30.7% of their answers. Then the researchers asked one more question — what’s missing here? — and the same people, on the same charts, found it 88.0% of the time.

Nothing changed except the question. The gap between a third and nine-tenths sat entirely in whether anyone had thought to look for absence.

That is the shape of every review you do of an AI draft, and why they feel so quick. You read what is on the page and ask whether any of it is wrong. Wrong is easy — a wrong figure sits there being wrong at you. Missing does nothing. There is no sentence to trip over, no claim to check, no seam.

So here is the habit. List from the source what the draft should contain, then make yourself answer that list against the draft, item by item, present or absent.

The content

The researchers have a name for the underlying tendency, and it is worth borrowing because it makes the problem portable: presence bias, “the systematic tendency to detect, learn, and reason from things that are present more readily than from things that are absent.”

Machines have it too, and worse. A benchmark published on 31 August 2026 built 500 note pairs — 298 where a specific fact had certainly been dropped, 202 where something had instead been added or altered — and put eight different AI reviewer designs at the benchmark. On the added-or-altered material, the reviewers scored 0.79 to 0.94. On the omissions, 0.50 to 0.63. Half is a coin toss.

The authors then tried the things you would try. Rewording. Voting across eight runs. Their finding is the useful part: every one of those moves shifted how much the reviewer flagged without giving it usable detection of what was gone. You can make a reviewer flag more, or flag less. You cannot make it see a thing that isn’t there by asking more firmly.

Which matters, because asking more firmly is exactly what most of us do. It also matters for a habit that is spreading — running an AI draft past a second AI to check it. That second pass catches the invented figure. On the caveat that never got written, it is barely better than not asking.

The obvious fix does not work either. If absence is the problem, hand the reviewer a checklist — that is the instinct, and it has been tested properly. Two randomised trials, reported in JAMA Network Open in 2023, emailed peer reviewers a short list of the items most often reported badly, and asked them to check whether the manuscript addressed each one and to press the authors for anything it didn’t. Across 243 published papers in one trial, completeness moved by 2.7 percentage points (p=.31); in the other it moved less. The authors’ own summary of what that means is blunt: reminding reviewers of specific items “is not useful in increasing reporting completeness in published articles.”

Two well-run trials, and nothing measurable happened.

So what does work? The same benchmark ran the control that separates the two ingredients, and the split is the useful part.

Give an ordinary AI reviewer the list of facts the source establishes, but let it answer with one overall judgement, and its score on omissions climbs, but stops inside the range it started in. Now make it return a verdict against each item on that list — present, or absent, one line at a time — and it reaches 0.786, clear of that range altogether. The authors apportion it: of the gap between an ordinary reviewer and the full procedure, having the list accounts for roughly a third, and the closed verdicts for the rest.

So the list is not the fix. The list is what makes the fix possible. The work is done by being made to say absent about a named thing, which a general read never asks of you — and never asked of those peer reviewers either. They were sent the list. Nobody made them answer it, line by line, on the record.

The distinction is fussier than use a checklist. It is also the difference between two trials that moved nothing and a procedure that works.

Try it

Take the next AI-generated summary, recap or brief you’re handed — something built from a source you also have. Run it as two turns.

Turn one, source only.

From this transcript/document alone, list the facts, decisions and open questions it establishes. Numbered list, one line each, maximum ten items. Do not summarise.

Read that list yourself and add anything you know should be on it.

Turn two, draft against list.

Here is the write-up. For each numbered item above, answer only PRESENT or ABSENT. If present, quote the words that carry it. Do not comment on quality.

Then read only the ABSENT rows. Those are the ones you would not have found.

Two turns rather than one is belt and braces: it stops the draft seeding the list. What the evidence actually requires is the list and the per-item verdict, so a single prompt carrying both will also work. If your workspace’s assistant can reach the source through a connector your admin has enabled, point it there instead of pasting. If it can’t, paste it — nothing here needs a feature.

Where it breaks: when the source is itself incomplete. This finds what the draft dropped, not what the meeting never said.

Additional reading

Editor’s note

This is hard to remember, and it is worth the effort. In my experience larger tasks carry the greater risk of critical omissions and also the longer list of things to check, so the discipline gets more valuable and more expensive at the same moment. There are lots of ways to manage that, including asking the model to break the list into sections you can work through. But the list itself is still yours to write. Nothing in the procedure knows what should have been in the document.

signed-off-by: Luke Topfer <editor> · 2026·09·07
06 Self-check

// three assertions against what you just read · results stay in this browser

assert 1/3

The module says reviewing an AI draft feels quick for a reason. What is it?

assert 2/3

You are handed an AI-written recap of a meeting whose transcript you also have. Following the module, what do you actually do?

assert 3/3

What does the module say about using a checklist to catch what's missing?