What's missing doesn't look wrong
AI-generated audio discussion of this module — same content, spoken.
Overview
Why now. In January 2026, researchers recruited a hundred people online and showed them charts with something quietly left out of the data. Asked what they noticed, one group flagged the missing piece in 30.7% of their answers. Then the researchers asked one more question — what’s missing here? — and the same people, on the same charts, found it 88.0% of the time.
Nothing changed except the question. The gap between a third and nine-tenths sat entirely in whether anyone had thought to look for absence.
That is the shape of every review you do of an AI draft, and why they feel so quick. You read what is on the page and ask whether any of it is wrong. Wrong is easy — a wrong figure sits there being wrong at you. Missing does nothing. There is no sentence to trip over, no claim to check, no seam.
So here is the habit. List from the source what the draft should contain, then make yourself answer that list against the draft, item by item, present or absent.
The content
The researchers have a name for the underlying tendency, and it is worth borrowing because it makes the problem portable: presence bias, “the systematic tendency to detect, learn, and reason from things that are present more readily than from things that are absent.”
Machines have it too, and worse. A benchmark published on 31 August 2026 built 500 note pairs — 298 where a specific fact had certainly been dropped, 202 where something had instead been added or altered — and put eight different AI reviewer designs at the benchmark. On the added-or-altered material, the reviewers scored 0.79 to 0.94. On the omissions, 0.50 to 0.63. Half is a coin toss.
The authors then tried the things you would try. Rewording. Voting across eight runs. Their finding is the useful part: every one of those moves shifted how much the reviewer flagged without giving it usable detection of what was gone. You can make a reviewer flag more, or flag less. You cannot make it see a thing that isn’t there by asking more firmly.
Which matters, because asking more firmly is exactly what most of us do. It also matters for a habit that is spreading — running an AI draft past a second AI to check it. That second pass catches the invented figure. On the caveat that never got written, it is barely better than not asking.
The obvious fix does not work either. If absence is the problem, hand the reviewer a checklist — that is the instinct, and it has been tested properly. Two randomised trials, reported in JAMA Network Open in 2023, emailed peer reviewers a short list of the items most often reported badly, and asked them to check whether the manuscript addressed each one and to press the authors for anything it didn’t. Across 243 published papers in one trial, completeness moved by 2.7 percentage points (p=.31); in the other it moved less. The authors’ own summary of what that means is blunt: reminding reviewers of specific items “is not useful in increasing reporting completeness in published articles.”
Two well-run trials, and nothing measurable happened.
So what does work? The same benchmark ran the control that separates the two ingredients, and the split is the useful part.
Give an ordinary AI reviewer the list of facts the source establishes, but let it answer with one overall judgement, and its score on omissions climbs, but stops inside the range it started in. Now make it return a verdict against each item on that list — present, or absent, one line at a time — and it reaches 0.786, clear of that range altogether. The authors apportion it: of the gap between an ordinary reviewer and the full procedure, having the list accounts for roughly a third, and the closed verdicts for the rest.
So the list is not the fix. The list is what makes the fix possible. The work is done by being made to say absent about a named thing, which a general read never asks of you — and never asked of those peer reviewers either. They were sent the list. Nobody made them answer it, line by line, on the record.
The distinction is fussier than use a checklist. It is also the difference between two trials that moved nothing and a procedure that works.
Try it
Take the next AI-generated summary, recap or brief you’re handed — something built from a source you also have. Run it as two turns.
Turn one, source only.
From this transcript/document alone, list the facts, decisions and open questions it establishes. Numbered list, one line each, maximum ten items. Do not summarise.
Read that list yourself and add anything you know should be on it.
Turn two, draft against list.
Here is the write-up. For each numbered item above, answer only PRESENT or ABSENT. If present, quote the words that carry it. Do not comment on quality.
Then read only the ABSENT rows. Those are the ones you would not have found.
Two turns rather than one is belt and braces: it stops the draft seeding the list. What the evidence actually requires is the list and the per-item verdict, so a single prompt carrying both will also work. If your workspace’s assistant can reach the source through a connector your admin has enabled, point it there instead of pasting. If it can’t, paste it — nothing here needs a feature.
Where it breaks: when the source is itself incomplete. This finds what the draft dropped, not what the meeting never said.
Additional reading
- Making Absence Visible: The Roles of Reference and Prompting in Recognizing Missing Information — Ben Shoshan, Lanir, Goldstein & Mokryn, January 2026. The hundred-participant study, including the per-domain detection rates and the authors’ note that they did not test a matching prompt for surplus information.
- LLM Judges Verify Presence, Not Absence — Fox, Markham, Lail & Karotsieris, 31 August 2026. A preprint, and the setting is clinical notes; the mechanism is not. Read it for what rewording and voting did and did not achieve.
- Reminding Peer Reviewers of Reporting Guideline Items — Speich et al., JAMA Network Open, June 2023. The null result, in full, with both trials reported.
Editor’s note
This is hard to remember, and it is worth the effort. In my experience larger tasks carry the greater risk of critical omissions and also the longer list of things to check, so the discipline gets more valuable and more expensive at the same moment. There are lots of ways to manage that, including asking the model to break the list into sections you can work through. But the list itself is still yours to write. Nothing in the procedure knows what should have been in the document.
// three assertions against what you just read · results stay in this browser
The module says reviewing an AI draft feels quick for a reason. What is it?
You are handed an AI-written recap of a meeting whose transcript you also have. Following the module, what do you actually do?
What does the module say about using a checklist to catch what's missing?
Was this useful for your daily work?