watts.it.com // daily AI micro-learning
Judgment & limits judgmentsecond opinionverificationreviewingsources 2026·09·17 · 4 min · evergreen

Leave out who said it

// listen · 2 ai hosts · audio edition

AI-generated audio discussion of this module — same content, spoken.

Overview

Why now. You are checking something a colleague told you, so you give the assistant the whole picture: our compliance lead says this clause is unenforceable — is that right? What comes back reads like an independent second opinion.

That sentence carried two things. A claim, and a name for who made it. On 8 September 2026, two researchers published an audit of what the second one does. They took questions a model had already answered correctly on its own, then asked again with one wrong option named — attributed either to “a subject expert” or to “most people”. Same wrong option, same verb, same instruction. Only the source changed.

The expert version pulled models off their own correct answers far more often than the majority version did. Across four models and roughly 220,000 answers, the pooled rates were 41.1% against 12.5%. The spread underneath is the part worth carrying: one hosted model moved 17.0% against 9.1%, another 37.0% against 12.9% — level with an open-weight model. The ordering held everywhere. The size did not.

The habit is one line. When you bring someone else’s claim to an assistant, describe the evidence and leave out who said it — naming an expert as the source pulls the answer toward that claim whether the claim is right or wrong.

The content

The obvious read is that models are gullible about experts, and that is the wrong shape for this. In the same study, an expert attribution pointing at the correct answer lifted accuracy to 84.9%, against 49.9% for a content-free filler cue. It was the experiment’s biggest mover, both ways.

So this is not a defect to be frightened of. It is a lever you are pulling without meaning to. Call the thing you hand over the credibility label — not the claim, but the standing of whoever made it. A separate study across three datasets put it plainly: “the dominant determinant is not group structure per se but the credibility label attached to the group.” Framing a source as a human expert moved the model reliably. Framing it as a friend did not.

The most direct evidence that the label does the work is a mechanistic study from July 2026. It held the wrong hint fixed and varied only who was giving it, across four levels of medical seniority. Against a baseline of roughly 60%, the physician hint dropped one model to 15%, dampening step by step as seniority fell. Same wrong answer, same wording. Only the job title changed.

Now the limits, because they decide how far to carry this.

The audit had reasoning turned off and forced a single letter as output, which is not a conversation. A benchmark published the same day covers that gap: seven models, open-ended dialogue of up to 25 turns, a persistent and mistaken user. The same ordering appeared — appeals to false authority moved models far more than appeals to what everybody thinks.

Resistance does not track model tier the way you would hope. It varies model to model — in the dialogue benchmark the most fragile system was a frontier one — which is why the spread matters more than the average. Nor is authority the strongest pressure: emotional framings ran higher there.

And here is the boundary that matters most: all of this is about relaying someone else’s claim. Nobody has cleanly tested the first-person version, where you announce your own credentials before asking. Do not rewrite how you introduce yourself on the strength of this.

Why has none of this simply been switched off? An August 2026 study of the neighbouring problem — models folding to a wrong majority — tested six mitigations, and five landed on one line: whatever made a model harder to talk out of a right answer made it harder to talk into one. The exception is the useful one: making the model reason the question through first raised both, where it could work the answer out itself. Which leaves the judgement with you, where it was anyway, and one finding to keep. Among models that show their reasoning, when they fold they mostly have not changed their minds: “collapse typically occurs while the correct position remains represented rather than after it disappears from the reasoning trace.”

Try it

Take a real one — a claim from a colleague, a consultant or an internal document that you need to check.

Ask it twice, in separate chats.

  1. Clean. State the question and the evidence, with no source attached. “A contract says X. Under Australian law, is that enforceable? Walk me through why.”
  2. Attributed. Same question, same words, with the label restored. “Our compliance lead says this clause is unenforceable — is that right?”

Then read the gap. That gap is what the name bought you, and it is the only part of this you can measure on your own work. Treat the clean answer as the assistant’s read. If your workspace lets you save prompts, keep the clean framing as your default.

Where it breaks: three ways. If the attribution genuinely matters — the compliance lead has seen the contract and you have not — stripping it out throws away real information; ask about the reasoning, not the verdict. If you have already argued with the assistant in that thread, start fresh. And a clean answer is not a correct answer: you have removed one thing pushing it around, not everything wrong with it.

Additional reading

Editor’s note

I see this often. I argue with lawyers about contracts. When I engage with an AI assistant to address a contract issue, the output I get is different depending on the delivery of the context. If weight is put on the fact that I’m a lawyer, then that can be influential. If weight is put on the fact that my counterparty is a lawyer, same effect but opposite direction. It’s one of the ways in which the prompt engineering skills we were all told were outdated after 2025 still feel relevant today.

signed-off-by: Luke Topfer <editor> · 2026·09·17
06 Self-check

// three assertions against what you just read · results stay in this browser

assert 1/3

The module rejects the obvious reading of the research — that models are simply gullible about experts. What does it put in its place?

assert 2/3

A colleague tells you a contract clause is unenforceable, and you want the assistant's own read. Following the module, what do you do?

assert 3/3

Which limit on this research does the module actually state?