watts.it.com // daily AI micro-learning
Judgment & limits summarisationjudgmentefficiencyverification 2026·06·24 · 4 min · evergreen

The summary that ate the caveat: what AI quietly drops when it condenses your documents

Overview

This is about the failure mode hiding inside the most useful thing AI does for you: condensing a long document. Not hallucination, not getting the news wrong — omission. The model keeps the conclusion and quietly drops the caveat, the dissent, the dollar figure, the condition that the whole document hung on.

Why now: a 2025 Royal Society Open Science study found AI summaries of scientific papers were roughly five times more likely than expert-written summaries to strip out the qualifiers that limit a finding’s scope — and, counterintuitively, newer models did this more, not less. And it isn’t a 2025 artefact: independent testing in 2026 on current frontier models found the same omission pattern — in one case every summary dropped the source text’s only actively participating female character.

What you’ll be able to do: summarise a real document so the relevant exceptions survive, and check whether they did.

The content

The obvious read is that summarisation is the safe, boring use of AI — low stakes, hard to get wrong. Overturn that. A summary’s job is to decide what is relevant, and that judgement is exactly where it fails you. The model isn’t lying; it is editing. It treats the carefully hedged sentence — the one the author agonised over — as a low-value detail and cuts it. You’re left with conclusions without conditions, a recommendation without the “unless”, a number without the asterisk.

Call it caveat collapse. The clause that constrains a claim is usually the longest, dullest, most qualified sentence in the document, so it reads as noise to a system optimising for a clean, confident précis. Peters and Chin-Yee (Royal Society Open Science, 2025) measured this on research papers: most models overgeneralised, producing broader claims than the source, and frontier models were worse than older ones. The unsettling part for anyone who thinks they’ve solved this with prompting: asking the model to be accurate and faithful roughly doubled the overgeneralisation. The instruction to “be precise” made it lop off the very nuance that precision lives in.

This isn’t confined to journal abstracts. The FABLES study (Kim et al., 2024) found that across book-length summaries, between a third and two-thirds omitted key events, and up to nearly 40 per cent dropped significant details — a consistent typology of omission, not random noise. And the cost compounds: a 2025 PNAS Nexus study of over 10,000 people (Melumad & Yun) found those who learned from AI summaries ended up with shallower, less factual knowledge than people who read the sources, partly because a summary hides what it left out. You can’t notice the caveat you were never shown.

So the risk isn’t a wrong fact you might catch. It’s a true summary that is materially misleading because of what’s missing — and missing things are invisible by construction.

Try it

Don’t ask for a shorter version. Ask the model to surface what a summary would normally bury, so you can decide what survives. Run this on a real document you’re about to act on — a contract, a board paper, a research report, a long email thread.

You are summarising the attached document for me to make a decision on.

Before you write any summary, list separately:
1. Every caveat, exception, condition, or "unless/except/provided that" clause.
2. Every dissent, minority view, or hedged/uncertain statement.
3. Every specific number, date, threshold, or dollar figure, with the
   condition attached to it.

Quote the exact source sentence for each. Do not paraphrase these.

THEN write a 5-sentence summary — and explicitly mark which items from
the lists above you carried into the summary and which you dropped,
and why you judged the dropped ones safe to omit.

The “mark what you dropped, and why” step is the point: it converts an invisible omission into a visible decision you can overrule. Where this won’t save you: if you can’t or won’t read the quoted clauses, you’re still trusting the model’s relevance judgement — the protocol surfaces the caveats, but only you can decide which ones are load-bearing for your call. And on a document you haven’t opened at all, no prompt makes the summary safe.

Additional reading

Editor’s note

The thing that should worry you is not only that the model could get something wrong, but that the summary is correct and still misleads, because the clause that mattered is simply gone. I see this most with contracts and guidance documents: the model returns the position cleanly and drops the “subject to” that the whole position depended on. The counterintuitive finding here (that telling the model to “be accurate” made omission worse) is the part to ruminate on, because it means your instinct to prompt harder is wrong. What you need to do instead is make it show you what it cut, then read those lines yourself.

signed-off-by: Luke Topfer <editor> · 2026·06·24
06 Self-check

// three assertions against what you just read · results stay in this browser

assert 1/3

When an AI model summarises a long document, what failure mode is most likely to catch you out?

assert 2/3

You're about to make a decision on a 40-page supplier contract and want an AI summary you can rely on. What does this module say to do?

assert 3/3

You've run the caveat-listing prompt on a board paper and the model has quoted every exception. Where can this still fail you?