"How did I do?" is not feedback
AI-generated audio discussion of this module — same content, spoken.
Overview
Why now. Rehearsing a hard conversation against an AI playing the other person is one of the few genuinely new things a plain chat window handed you. The salary ask, the underperformance conversation, the sceptical room — run it as often as you like, nobody is embarrassed.
Almost everyone runs it the same way: do the role-play, then ask how it went.
That last step has now been measured on its own, and it did nothing. A preregistered negotiation experiment published at EMNLP 2024 put 374 people through two negotiations with an objective outcome: the price they settled on as buyers. Feedback anchored in a specific, expert-built scheme of negotiation errors improved them between the first and the second (d=0.38, p=0.003). Feedback generated by a general-purpose model, with no such scheme behind it, did not improve them at all (p=0.82), and was no better than the arm given no feedback whatsoever (F(1,250)=0.57, p=0.45).
The rehearsal was not the intervention. The scoring was.
So: when you rehearse a hard conversation with an AI, write down first the three or four specific behaviours you are trying to use, and end the session by making it score you on each one separately. Model feedback with no scheme behind it has been measured against no feedback at all, and came out the same.
The content
Call the thing you write before you start the named list: three or four observable behaviours, not qualities. Not “be more empathetic” but “acknowledge their point before I counter it”. Not “be confident” but “state my number once, then stop talking”. The test is whether a stranger reading the transcript could mark each line yes or no, not knowing you or the outcome.
The same split turns up across three domains and thirteen years. A CHI ‘26 study ran 94 novice counsellors through identical AI role-plays, varying only whether a structured feedback layer was present. A 2013 job-interview study — pre-LLM, scored by career counsellors blind to condition — found practice-with-feedback beat practice-alone, not merely the control, though only among the women in the sample. A 2024 medical-education trial made the scoring step its only manipulated variable, and it was the difference, on 21 students at p=.049. Different tasks, eras and raters. Same result.
Why it does not feel like nothing. In that practice-only arm, one measured skill went backwards — strong uses of empathy fell 9.6% (d=-0.52) — while the same people’s self-assessment of their exploration skills moved from 11.6 percentile underconfidence to 5.7 percentile overconfidence (p<0.001, d=0.58), the only calibration measure that shifted. The honest reading is narrower: one arm, one skill of four, scored by a classifier, inside the simulator. It is not evidence that rehearsing makes you worse. It is evidence for something quieter — unscored rehearsal buys confidence whether or not anything improved.
Why more information beats a verdict. A 2020 meta-analysis pooling 994 effect sizes sorted feedback by information content. Pure reinforcement or punishment — good job, that was weak — came in at d=0.24. Corrective feedback at d=0.46. Corrective content plus information on regulating your own learning reached d=0.99. The ladder runs from verdict to instruction.
Kluger and DeNisi’s older meta-analysis supplies the mechanism, and it matters which way you read it: they were studying how often feedback harms performance, and over a third of their effects were negative. What predicted the damage was attention moving off the task and onto the person. “How did I do?” is a question about you. “Did I name the constraint before I made the ask, yes or no?” is a question about the task.
Two limits to carry. None of this has been shown to transfer to a real workplace conversation. A deployment of exactly this use case to 40,000-plus managers over six months says so itself: “we have not established that simulated practice improves real-world conversations.” The negotiation experiment is the closest thing here to workplace work, and it is a priced, closed-form task. And a narrow list is not free: in a 2023 surgical training study, AI feedback on four named metrics improved those metrics and degraded ones nobody was scoring. Rotate the list rather than carrying the same four.
Try it
Take a conversation you have coming up.
Before you open the chat, write the named list: three or four observable behaviours in the yes/no form above. Writing it afterwards defeats the point — you will name whatever you did.
Then set up the rehearsal. Give the model the other person’s position, incentives and likely objections, tell it to stay in role and not be agreeable, and run it long enough to get past the opening.
Then score it. Paste the named list back and ask for each behaviour separately: did I do this, yes or no, and quote the line where I did or did not. The quoted line is what makes this work — it gives you something to disagree with, which a rating out of ten does not.
If your workspace offers saved prompts or a custom assistant, the scoring instruction is the piece worth saving: the named list changes every time, the scoring step never does. Check what yours has enabled; failing that, the same three messages work in a plain chat window.
Where it breaks. The model is a lenient scorer and will hand out yeses for near misses — which is why you ask for the quoted line, not the verdict. You are checking its evidence, not its judgement. And a behaviour you cannot mark yes or no is not on the list yet.
Additional reading
- ACE: A LLM-based Negotiation Coaching System — Shea et al., EMNLP 2024. Preregistered, three arms, objective outcome. §7.2.2 carries the comparison, and notes that feedback which improves model negotiators did not help people.
- Can LLM-Simulated Practice and Feedback Upskill Human Counselors? — Louie et al., CHI ‘26. Calibration is in §5.2; Figure 3 shows what reached significance. Skills were scored by classifiers.
- The Power of Feedback Revisited — Wisniewski, Zierer & Hattie, 2020. The information ladder; the top rung rests on 42 effect sizes, and 17% of measured effects were negative.
- Feedback Interventions: The Effects on Performance — Kluger & DeNisi, 1996. Read it for the moderator table, not as an argument that feedback is good for you — a third of the effects went the other way.
- Conversation Coach — Amazon, August 2026. The honest ceiling: an organisation-wide deployment of AI conversation rehearsal that declines to claim transfer.
Editor’s note
I’ll confess that I’m not certain there’s actually a good way to design this type of exercise. In almost every case, there’s a personality element that can’t be captured. With that said, being able to practise your own role is rarely wasted effort, and if you’re going to do it, you might as well get as close as possible. I do think this is as close as it gets, but we have to acknowledge that it’s still your guesswork.
// three assertions against what you just read · results stay in this browser
The module opens on a three-arm negotiation experiment. What separated the arm that improved from the two that did not?
You have a salary conversation next week and want to rehearse it with an assistant. Following the module, what do you actually do?
Which limits does the module actually state about its own evidence and its own advice?
Was this useful for your daily work?