The summary that ate the caveat: what AI quietly drops when it condenses your documents
Overview
This is about the failure mode hiding inside the most useful thing AI does for you: condensing a long document. Not hallucination, not getting the news wrong — omission. The model keeps the conclusion and quietly drops the caveat, the dissent, the dollar figure, the condition that the whole document hung on.
Why now: a 2025 Royal Society Open Science study found AI summaries of scientific papers were roughly five times more likely than expert-written summaries to strip out the qualifiers that limit a finding’s scope — and, counterintuitively, newer models did this more, not less. And it isn’t a 2025 artefact: independent testing in 2026 on current frontier models found the same omission pattern — in one case every summary dropped the source text’s only actively participating female character.
What you’ll be able to do: summarise a real document so the relevant exceptions survive, and check whether they did.
The content
The obvious read is that summarisation is the safe, boring use of AI — low stakes, hard to get wrong. Overturn that. A summary’s job is to decide what is relevant, and that judgement is exactly where it fails you. The model isn’t lying; it is editing. It treats the carefully hedged sentence — the one the author agonised over — as a low-value detail and cuts it. You’re left with conclusions without conditions, a recommendation without the “unless”, a number without the asterisk.
Call it caveat collapse. The clause that constrains a claim is usually the longest, dullest, most qualified sentence in the document, so it reads as noise to a system optimising for a clean, confident précis. Peters and Chin-Yee (Royal Society Open Science, 2025) measured this on research papers: most models overgeneralised, producing broader claims than the source, and frontier models were worse than older ones. The unsettling part for anyone who thinks they’ve solved this with prompting: asking the model to be accurate and faithful roughly doubled the overgeneralisation. The instruction to “be precise” made it lop off the very nuance that precision lives in.
This isn’t confined to journal abstracts. The FABLES study (Kim et al., 2024) found that across book-length summaries, between a third and two-thirds omitted key events, and up to nearly 40 per cent dropped significant details — a consistent typology of omission, not random noise. And the cost compounds: a 2025 PNAS Nexus study of over 10,000 people (Melumad & Yun) found those who learned from AI summaries ended up with shallower, less factual knowledge than people who read the sources, partly because a summary hides what it left out. You can’t notice the caveat you were never shown.
So the risk isn’t a wrong fact you might catch. It’s a true summary that is materially misleading because of what’s missing — and missing things are invisible by construction.
Try it
Don’t ask for a shorter version. Ask the model to surface what a summary would normally bury, so you can decide what survives. Run this on a real document you’re about to act on — a contract, a board paper, a research report, a long email thread.
You are summarising the attached document for me to make a decision on.
Before you write any summary, list separately:
1. Every caveat, exception, condition, or "unless/except/provided that" clause.
2. Every dissent, minority view, or hedged/uncertain statement.
3. Every specific number, date, threshold, or dollar figure, with the
condition attached to it.
Quote the exact source sentence for each. Do not paraphrase these.
THEN write a 5-sentence summary — and explicitly mark which items from
the lists above you carried into the summary and which you dropped,
and why you judged the dropped ones safe to omit.
The “mark what you dropped, and why” step is the point: it converts an invisible omission into a visible decision you can overrule. Where this won’t save you: if you can’t or won’t read the quoted clauses, you’re still trusting the model’s relevance judgement — the protocol surfaces the caveats, but only you can decide which ones are load-bearing for your call. And on a document you haven’t opened at all, no prompt makes the summary safe.
Additional reading
- Generalization bias in large language model summarization of scientific research (Peters & Chin-Yee, Royal Society Open Science, 2025) — measures how models strip scope-limiting qualifiers; accuracy prompts made it worse.
- FABLES: Evaluating faithfulness and content selection in book-length summarization (Kim et al., 2024) — a typology of what long-document summaries omit, and how often.
- Experimental evidence of the effects of large language models versus web search on depth of learning (Melumad & Yun, PNAS Nexus, 2025) — 10,000+ participants; why learning from summaries leaves knowledge shallower.
- Loss by Omission: GenAI Summarization Tools and History Knowledge (Ross, The Digital Orientalist, March 2026) — independent field testing (not peer-reviewed) of current frontier models (ChatGPT 5.1, Claude 4.5 Sonnet, Microsoft 365 Copilot) summarising a primary text; the same omission pattern as the 2025 studies, including dropping “almost all narrative substance” and the source’s only actively participating female character.
- A Large-Scale Multi-Dimensional Empirical Study of LLMs for Conversation Summarization (Zhou et al., OmniCSEval, June 2026) — a 2026 study of 28 models across 1,800 conversations that measures completeness — whether a summary keeps the essential facts — as a first-class metric alongside faithfulness.
Editor’s note
The thing that should worry you is not only that the model could get something wrong, but that the summary is correct and still misleads, because the clause that mattered is simply gone. I see this most with contracts and guidance documents: the model returns the position cleanly and drops the “subject to” that the whole position depended on. The counterintuitive finding here (that telling the model to “be accurate” made omission worse) is the part to ruminate on, because it means your instinct to prompt harder is wrong. What you need to do instead is make it show you what it cut, then read those lines yourself.
// three assertions against what you just read · results stay in this browser
When an AI model summarises a long document, what failure mode is most likely to catch you out?
You're about to make a decision on a 40-page supplier contract and want an AI summary you can rely on. What does this module say to do?
You've run the caveat-listing prompt on a board paper and the model has quoted every exception. Where can this still fail you?
Was this useful for your daily work?