Citation theatre: the more a research agent cites, the more links it fabricates
Overview
A deep research agent runs its own web searches and hands back a long report stacked with citations. The stack of footnotes is the thing that makes it feel authoritative. This module is about which of those links to click, and what a dead one is actually telling you.
Why now. In April 2026 three University of Pennsylvania researchers — Delip Rao, Eric Wong and Chris Callison-Burch — measured whether AI citation URLs go anywhere. Across more than 200,000 citation links — ten models and agents on one benchmark, three on a separate 32-field dataset — they found 3–13% were hallucinated (no record in the Internet Archive; in their words, they “likely never existed”) and 5–18% did not resolve at all. The twist that matters for how you work: the deep research agents, the tools built to be exhaustive, generated more citations per query than ordinary search-augmented chat — and hallucinated them at higher rates.
By the end you’ll have a thirty-second habit for which links to click on any AI-assisted report, and what a broken one means.
The content
The obvious read is that turning on web search fixes citations — the model isn’t guessing any more, it’s retrieving, so the links must be real. That gets it half right and half backwards. Retrieval reduces invention; it doesn’t remove it. And the sheer volume of footnotes a research agent produces does its own quiet work on you: twenty sources look like twenty checks. Call it citation theatre — the density of citations performs a rigour the links don’t always have.
The Penn study is worth reading for one decomposition in particular. When they sorted the broken links, some models’ non-resolving URLs were almost entirely link-rot — real pages that have since died — while other models fabricated every non-resolving URL outright. So a link that fails to load is ambiguous: it might be a genuine source that moved, or a sentence the model invented whole and dressed with a plausible address. Either way the citation is now doing nothing for you, and the claim sitting on top of it is unsupported.
Here’s the limit, stated plainly so you don’t over-trust the habit: this study measured whether URLs resolve, not whether the page — if it’s live — actually says what the report claims. A working link is necessary, not sufficient. Clicking weeds out the fabricated and the dead; it does not catch a real page being misread or stretched. Retrieval is not verification.
Try it
Use this on your most recent AI-assisted report with citations — the one closest to going somewhere with your name on it, not a throwaway query.
- Pick the three or four claims you would actually lean on — the load-bearing ones, not every line.
- Click each of their links. If one 404s, redirects to a bare homepage, or lands on a page that doesn’t mention the claim, that citation is empty — cut the claim or find a real source for it.
- For a broken link on a claim you really need, paste the URL into the Wayback Machine (web.archive.org). If the Internet Archive has never captured it, that is the study’s own test for “never existed” — treat the surrounding sentence as fabricated, not merely stale.
To build the list fast, you can have the agent lay it out for you — though remember it cannot confirm its own links, so you do the clicking:
List every factual claim in the report above that rests on a citation.
For each, give me one row: the claim as written, the single URL it cites,
and the exact sentence on that page it relies on.
Do not reassure me the links are valid — you can't. Just build the list so
I can open each one myself.
Where it breaks: a link that loads is not proof. The page can resolve and still not support what the report built on it — clicking catches fabrication and rot, not misreading. The heaviest-cited report is the one to check hardest, not the one to trust most.
Additional reading
- Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents — Rao, Wong & Callison-Burch (arXiv, 3 Apr 2026) — the source for the 3–13% hallucinated / 5–18% non-resolving figures and the finding that deep research agents cite more and fabricate more.
- Full text (HTML) — the per-model and per-domain breakdown, including which models fabricate every broken link versus which suffer genuine link-rot.
- Wayback Machine — Internet Archive — paste a suspect URL here; no capture, ever, is the practical tell that a link was invented rather than merely dead.
Editor’s note
A citation is a promise that a reader can follow it back to the source. A research agent makes that promise dozens of times in a single report, and some of the trails lead nowhere. The reflex to trust the well-referenced document is the wrong one here — the volume of footnotes is doing the persuading, not the sources behind them. You don’t need to open every link. Open the few the decision rests on, and treat the ones you didn’t as unsourced, not as checked.
// three assertions against what you just read · results stay in this browser
Deep research agents run their own web searches, so their citations should be solid. What did the study behind this module actually find about them?
A deep research agent hands you a report backed by twenty citations, and it's about to go to your director with your name on it. What does this module say to do first?
A working link is necessary but not sufficient. Which failure does clicking through a report's citations NOT catch?
Was this useful for your daily work?