The more it retrieved, the less it asked
AI-generated audio discussion of this module — same content, spoken.
Overview
Why now. In May 2026 two Cornell researchers put a thousand questions to ten assistants, more than half of them genuinely ambiguous — several correct answers each, depending on what you meant. Then they counted how often an assistant stopped to ask which one you meant. Seven of the ten managed it on under one per cent of the ambiguous questions.
Then they ran it again with retrieved passages in the window — the arrangement your work assistant uses when it reads your document store before answering. Eight of the ten asked less than before. The other two did not move. Not one asked more.
You know the situation from the other side. You typed a short question into the assistant wired to your files and got back something fluent, sourced and confident. What you cannot see is that the question had a second reading, and it picked one.
So the habit is small, and it comes after the answer rather than before it: ask it to list the other ways your question could have been read. It will not volunteer them, and the more it retrieved, the less likely it was to.
The content
Put a different question to those same models — is this question ambiguous? — and they often say yes. The recognition is there. It just never reaches the answer box.
The flattering version of that — it knew and hid it from you — is not quite what the numbers show. The same models also label plainly clear questions as ambiguous a good deal of the time. Some of what looks like insight is a tendency to say yes when asked. The recognition runs hot, and none of it surfaces where you are actually working.
What should change how you use a connected assistant is the direction of that effect. An independent team at Google found it from another angle. They measured abstention — how often a model declines to answer at all — and watched it fall once retrieval was on: for one model, from 84.1% of questions down to 52%. Their proposed explanation is the line worth keeping. It may arise, they suggest, from the model’s “increased confidence in the presence of any contextual information.”
The Cornell team reach for almost the same words. Once supporting passages are present, they suggest, the model “tends to treat the query as effectively unambiguous.”
Call it the licence to answer. The documents in the window do not have to be the right ones, or to address the reading you meant. They only have to be there. And the better your organisation’s connectors get, the more reliably that licence is granted.
Which inverts the usual pitch. Connect everything and the answers improve; the measured side effect is that its last impulse to check with you goes quiet.
This is not a quirk of one generation, either. The Cornell set sits two generations behind the current frontier, but researchers reported the same behaviour in December 2022 — rarely asking, answering wrongly instead — and a 2025 abstention benchmark, much of it built from underspecified questions, found models still answering definitively rather than flagging the gap.
One boundary decides where the habit applies. Hand a model a tool for asking you something, warn it the brief may be missing details, set it a long task, and the picture changes: in one May 2026 study one frontier model asked in about half its sessions, a second only selectively, and a third never asked at all, even when told to. What is described here is the answer box — the short question, the one confident reply. Which is where most people meet this at work.
And the obvious fix is the wrong one. Asking “was my question ambiguous?” invites exactly that yes-bias, and you will manufacture doubt about questions that were fine. Ask for the readings themselves. You want a list you can look at, not a verdict you have to trust.
Try it
Take a question you put to your work assistant this week — one where the answer mattered and the question was short.
Reopen the thread and ask:
List the other ways my question could have been read. For each reading, say in one line how the answer would have been different. Do not re-answer the original question.
Read what comes back as candidates, not findings. Most will be noise, and that is expected. You are looking for one thing: a reading you did not intend but a reasonable colleague might have assumed — the wrong time period, the wrong entity with a similar name, the wrong scope of “our team”.
If one is closer to what you meant, ask that version as a fresh question. Do not ask it to revise the first answer — you want the retrieval to run again on the sharper question.
If your workspace lets you save instructions or reusable prompts, keep this one there. Check what yours has enabled; if not, that is a fair thing to ask your admin for.
Where it breaks: two ways. It surfaces alternatives worth checking; it is not a reliable way of recovering what you meant. In a direct test of prompted disambiguation, on Chinese-language ambiguity across open-weight models, the complete set of readings came back in none of the conditions tried. And it works on the question, not the sources: a perfectly unambiguous question answered from the wrong document is a different failure, and this will not catch it.
Additional reading
- Knowing but Not Showing: LLMs Recognize Ambiguity but Rarely Ask Clarifying Questions — Su & Cardie, Cornell, 24 May 2026. The clarification rates are Table 6, with the retrieval change in brackets beside each. Table 4 is the yes-bias.
- Sufficient Context: A New Lens on Retrieval Augmented Generation Systems — Joren et al., Google, ICLR 2025. The “Models Abstain Less with RAG” paragraph inside Section 4.2 runs four sentences and carries the whole finding.
- CLAM: Selective Clarification for Ambiguous Questions — Kuhn, Gal & Farquhar, December 2022. The same behaviour, three and a half years and several model generations earlier. Its own contribution is the other half: prompting a model to detect, then ask, improved accuracy.
- Ask Early, Ask Late, Ask Right — May 2026. The boundary case: asking behaviour once a model has a tool for it and a long task to do.
- Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity — July 2025. Table 4 is that direct test; read the table itself, not the prose around it.
Editor’s note
I’ve encountered this issue in the particularly irritating context of a post-mortem in a bad session, where a session just couldn’t execute on what I felt was a reasonably clear set of instructions. On interrogating what went wrong, I realised that there was an unseen ambiguity in my instructions. Catching it beforehand is a level of prompt discipline that I honestly expect to be fairly rare, but maybe if this is in the back of your mind, you’ll avoid the pitfall every now and again.
// three assertions against what you just read · results stay in this browser
The Cornell researchers ran their whole set of questions twice — once plainly, once with retrieved passages in the window. What did adding the retrieved passages do to how often the assistants asked which reading you meant?
You asked your work assistant a short question this morning, and it came back with a fluent, sourced, confident answer that you are about to act on. Following the module, what do you actually do next?
The module draws a boundary around where its finding applies. What is that boundary?
Was this useful for your daily work?