watts.it.com // daily AI micro-learning
Judgment & limits oversightverificationover-reliancefeedbackjudgment 2026·08·24 · 4 min · evergreen

The harsh verdict gets a pass

// listen · 2 ai hosts · audio edition

AI-generated audio discussion of this module — same content, spoken.

Overview

Why now. In June 2026 a preregistered experiment with 1,339 practising teachers tested the same wrong grade, attributed either to a colleague or to an algorithm. When the wrong grade was too harsh, teachers corrected it less often under the AI label. When the same wrong grade was too generous, the label made no measurable difference. So: when an assistant’s verdict on a person or a piece of work comes back negative, check it against the source before you accept it — that is the direction you are least likely to challenge.

The content

The standard warning assumes you will be too quick to dismiss the machine. That is a well-documented pattern — people drop a model fast after watching it get one thing wrong — and it shapes most of the guardrails organisations write.

This study points at the opposite failure. And the failure has a direction.

Teachers were shown a piece of student work alongside a grading recommendation. In one arm the work objectively merited 8/10 and the recommendation said 5 — harsh. In the other, it merited 2/10 and the recommendation still said 5 — lenient. Same recommended number both times; only the direction of the error changed. Half the teachers were told the recommendation came from a colleague, half from an algorithm. What was measured was how far each teacher’s own submitted grade landed from the objective benchmark.

In the harsh arm, the algorithm label pulled graders closer to the wrong recommendation. The paper’s controlled estimate puts that gap 22% wider than under human advice. In the lenient arm: nothing measurable.

Read that null precisely, because it is tempting to improve on it. It does not mean people pushed back against a soft AI. It means the label stopped mattering. Whatever makes an algorithmic verdict harder to argue with did not show up when the verdict ran in someone’s favour.

The proposed mechanism is more uncomfortable than “people trust computers”. These teachers did not rate the algorithm as more capable than a colleague — they rated it lower on ability in both conditions. Harshness did not make it look clever. It made it look less incompetent: the authors describe teachers perceiving “harsh algorithmic recommendations as less indicative of low ability than comparable human recommendations”. Severity partly cancels the discount you were already applying. Sounding tough reads as having looked closely, even when you think little of the thing doing the judging.

Which brings the part that should change what you do about it. Asked afterwards, these teachers were not favourable about AI grading — on a −5 to +5 scale they averaged −1.03 on willingness to let AI grade and −1.33 on whether it was ethically acceptable. Not an enthusiastic room, and it deferred anyway.

Take that as an aggregate, not a law about individuals: the study reports those attitudes descriptively and never tests whether the sceptics among them resisted more. What it shows is narrower, and still worth having. Disapproval, on its own, did not function as a check.

That fits what happens when people try to train the problem away. In one trial of a short AI-literacy session, “the educational intervention did not significantly reduce over-reliance” — what it did instead was make students ignore correct recommendations more often. General suspicion moves the dial in both directions at once, so it costs you accuracy without buying vigilance. A trigger — this verdict is negative — does not have that problem, because it fires on a specific occasion rather than on your whole relationship with the tool.

Where the evidence stops. This is one study, on Greek teachers, grading. The authors name hiring, healthcare and finance as places the pattern may extend; that testing has not happened. Separately, none of this describes what it is like to be on the receiving end of an AI decision about you — that is a different literature, and it runs in both directions. This module is about the seat you are in: reviewing a judgement about someone else, with the power to change it.

Try it

Pick the most recent time an assistant assessed something critically — a critique of your draft, a risk it flagged in a contract, a “this section doesn’t hold up”.

  1. Before you act on it, make it show its working. Paste the criticism back with the original and ask: “Quote the exact text from the source you’re basing this criticism on. If the wording isn’t there, say so plainly instead of paraphrasing.”
  2. Check the quote against the document yourself. Not the assistant’s summary of it — the document. Half the value is finding the objection rests on something that was never there.
  3. Then decide. You may well agree with the criticism. The habit isn’t distrust; it’s making the negative verdict earn the same scrutiny you would have given a colleague who said it.

Make it a standing rule. If your tool lets you save instructions or a reusable prompt, add one line to it: when your assessment of a person or a piece of work is negative, quote the specific evidence for it. That works in most enterprise assistants — check what your workspace has enabled; if saved instructions aren’t available, keep step 1 as a copy-paste.

Where it breaks. Do this on everything and you will spend your week checking. Scope it to verdicts with a consequence attached — the ones that get acted on, and the ones about people. And note the limit honestly: the study measured whether teachers corrected a wrong grade, not whether a check like this one fixes it. It is a sensible handle on a documented blind spot, not a proven remedy.

Additional reading

Editor’s note

I fell for this one myself, and I defended an overly harsh assessment proposed to me by an AI assistant. It was only when a colleague challenged my presentation (which I honestly attributed to AI as I presented it) that I committed the lesson from this module to my standard operating procedure. It is so easy to treat validation of AI output as applying only to facts, sources, and flattery, but it applies equally to criticisms. Not all are valid, and equal weight should be given to your assessment of them.

signed-off-by: Luke Topfer <editor> · 2026·08·24
06 Self-check

// three assertions against what you just read · results stay in this browser

assert 1/3

In the experiment the module describes, the same wrong recommendation was tested in two directions. What did the AI label change?

assert 2/3

An assistant reviews a contract you drafted and flags a clause as a serious risk. Following the module, what do you do first?

assert 3/3

The module reports that the teachers were not favourable about AI grading and deferred anyway. What does it conclude from that, and what does it NOT claim?