watts.it.com // daily AI micro-learning
The landscape ai-transitionverificationreliabilityjudgementtrust 2026·07·10 · 7 min · dated

Vibe citing: when generation is free, being able to vouch for it is the job

Overview

This weekly starts with a report that no longer exists. In October 2025 KPMG published a flagship piece on agentic AI — Total Experience: Redefining Excellence in the Age of Agentic AI. In June 2026 the firm quietly pulled it from its websites, because a forensic review by the AI-detection firm GPTZero found that of the report’s 45 citations, only five correctly pointed to the source they claimed. Forty of the forty-five titles were fabricated. Around half the claims those citations were meant to support turned out to be fake or misattributed — including case studies about UBS, NHS Greater Manchester, Swiss Federal Railways and Transport for London, each of which disputed the report’s account of their own AI use, as reported by the Financial Times.

A report about excellence in AI, undone by the AI that helped write it. It’s an easy story to enjoy and a lazy one to learn from. The comfortable read — a model hallucinated, a check got skipped, better tools will fix it — misses what actually changed. By the end of this edition you’ll have the real lesson: what becomes scarce, and therefore valuable, when generating a plausible page costs nothing.

The content

Start with the mechanism, because GPTZero gave it a name worth keeping: vibe citing. It’s the citation cousin of vibe coding — a model asked to support an argument stitches together fragments of real sources and invents the rest, producing references that look impeccable and don’t resolve. The tells were consistent throughout the reference list: fabricated titles, real authors welded to papers they never wrote, two genuine sources fused into one that doesn’t exist. GPTZero’s read was that the authors had likely used an AI referencing tool that over-complied — asked to find examples of agentic AI in the wild, it simply produced them, whether or not they were real.

The obvious response is that this is a discipline problem at one firm — someone should have checked. That’s true, and it’s also the least useful thing to take away, because it lets everyone who isn’t KPMG off the hook. The more honest reading is that this is the first clear shape of a failure mode that the current tools make easy for everyone, not a lapse peculiar to one team under deadline. When producing a fully-formed, confident, well-formatted artefact becomes almost free, the thing that used to be bundled inside “I wrote this” — and I can show you where every claim comes from — quietly comes unbundled. The output looks the same. The backing behind it does not.

Here’s why that gap is structural rather than teething, and this is the part the “better models will fix it” read gets wrong. METR measures how long a task a model can complete, and it reports two numbers that matter here. There’s the length a model can do with 50% reliability — a coin flip — and the length it can do with 80% reliability. The 80% figure is roughly five times shorter. In plain terms: the tasks a model can pull off often enough to impress you in a demo are several times longer than the tasks it can be depended on to get right. “It can do this” and “it reliably does this” are not the same claim, and the distance between them doesn’t close just because the headline number keeps climbing — METR’s latest measurements, updated in May 2026, show both frontiers rising together. Capability improving is not the same as the reliability gap shrinking. A tool can get genuinely, impressively better and still be exactly as unsafe to publish from unchecked.

So the reframe. When anyone can generate a credible-looking deck, memo or report in minutes, generation stops being the scarce thing. What becomes scarce — and what quietly separates work you can build a reputation on from work that detonates under scrutiny — is provenance: the unbroken line from each claim back to a source a human actually opened. The signature on a piece of work used to mean I made this. It is quietly coming to mean something harder and more valuable: I stand behind this, and I can show you why. Those are now two different acts, and only the second one is scarce.

Read the KPMG episode again through that lens and it stops being a story about one firm’s embarrassment. It’s a preview of the sorting that’s coming for everyone who publishes under an institution’s name. The organisations that will be trusted in this period are not the ones generating the most — everyone can do that now — but the ones that can vouch for what they ship. That capability doesn’t arrive with a better model. It’s built: in habits, in who-checks-what, in a culture where “where did this number come from” is a normal question and not an accusation. The leaders who treat that as the actual work of the AI transition — rather than assuming the tools have handled it — are building the thing that will still be worth something when the novelty of cheap generation has wholly worn off.

For you, at your desk, the move is smaller and completely within reach. Separate the two acts the tools have fused. Let AI generate freely — drafts, structure, first passes, candidate sources. Then, before your name goes anywhere near it, attest to it deliberately: every claim the argument leans on traces to something you personally opened and read, not to a citation that merely looks right. A fabricated reference is designed to survive a glance; it fails the click. The single most reliable tell that you’ve drifted into vibe citing is a reference list you assembled but never actually followed. Following it is now part of the writing, not a courtesy after it.

Two honest qualifications. First, this isn’t an argument against using AI to write — KPMG’s mistake wasn’t using the tool, it was shipping its output as though generation and verification were one step. Used with the attestation kept separate and human, these tools are extraordinary. Second, provenance is not the same as truth: a real, correctly-cited source can still be wrong, and checking that a citation resolves is the floor, not the ceiling. But it is a floor that the KPMG report — and a great deal of work going out right now under real names — never reached.

The promise of cheap generation was that the work would get easier. What it actually did was split the work in two and automate only the first half. Producing the artefact is nearly free now. Being able to stand behind it is not — and in a period where everyone can produce, that second half is quietly becoming the whole job.

Additional reading

Editor’s note

This website faces inward and outward for me. I want to improve my own knowledge and abilities, and I’d love for you to join me on the journey. As I commit this one for publication, you had better believe that I verified the sources.

signed-off-by: Luke Topfer <editor> · 2026·07·10
05 Self-check

// three assertions against what you just read · results stay in this browser

assert 1/3

When a credible-looking report, deck or memo can be generated in minutes, what becomes the scarce and valuable thing?

assert 2/3

An AI assistant has drafted your client briefing, complete with a tidy reference list it assembled. The argument leans on several cited claims. Before your name goes anywhere near it, what does this module say to do?

assert 3/3

Every reference in your report now resolves to a real source you have personally read. What limit does the module still warn about?