watts.it.com // daily AI micro-learning
Tools & connectors settingstools & connectorssummarisingverificationenterprise 2026·08·19 · 4 min · dated

Effort is a trade, not a dial: when to turn the thinking down

// listen · 2 ai hosts · audio edition

AI-generated audio discussion of this module — same content, spoken.

Overview

You have found the control by now. It might be called Think deeper, or Extended, or a reasoning level, or it might just be a slider with more at one end. And you have probably drawn the conclusion almost everyone draws: more is better, it only costs time.

That is the wrong model of it. Turning it up buys planning depth and analytical reach — and pays for them with precision and fidelity to what you actually gave it. The setting should follow the job.

So: turn the reasoning setting down whenever faithfulness to the source or to your instructions is the job — summarising, pulling fields out of a document, rewriting to a house format, holding a word count — and turn it up only when the work is genuine analysis.

Call it the fidelity trade.

The content

A widely-read practitioner recently wrote up an open-weights model shipping on its highest reasoning setting by default. Asked to draw a circle, it deliberated for several minutes — its visible reasoning has the line that this was a simple request, but it wanted the result carefully crafted — then produced an elaborate animated one. His advice: ignore the default, run it low.

A funny story about a bad default, not evidence. The evidence is duller.

A team at the University of Hawaii ran a large comparison of reasoning strategies on summarising — ordinary compression work, not maths puzzles — and in one experiment varied a frontier model’s think-ability setting from minimal to high, on three datasets. Their finding, in their words: factual faithfulness consistently declines as think ability increases. Their verdict on the study is that more reasoning is not always better, and that effective reasoning should preserve faithful compression rather than induce over-elaboration.

The reframe is worth stating plainly. The extra thinking is not spent on getting your document right. It is spent on doing more — elaborating, improving on what you handed over. On analysis that is what you want. On a summary, the thing it improves on is your source.

A second study makes the same point from the other direction, and its design closes the obvious objection: the author used the very same model and flipped only the internal signal that turns thinking on and off. Across a standard instruction-following benchmark the aggregate pass rate barely moved. But between 10% and 20% of individual prompts flipped between pass and fail — thinking reshuffled which instructions got obeyed rather than moving the score much. And the reshuffle had a direction: constraints about planning, structure and coordination improved, while constraints about exact local form got consistently worse. Word limits. Formats. “No commas.” The things you asked for precisely.

The damage is invisible. In that summarisation experiment, the metrics that look like quality — the ones scoring a summary against a reference — stayed flat while factual grounding fell. The output does not get worse to read. If anything it reads better: more considered, more confident. So checking your work by re-reading it will not catch this. Re-reading is the test it passes.

The vendors quietly say so already. Google’s developer documentation tells you to use minimal or low thinking for fact retrieval or classification, reserving maximum thinking for advanced coding, maths and multi-step planning. And in this month’s launch guidance, OpenAI reports its newest model at low reasoning beating the previous generation at high on an agentic benchmark, with everything else about the test held constant — a strange thing to publish if the dial were just a quality slider.

Three limits, because this cuts both ways. Higher effort helps on hard analytical work — the trade has a direction, not a winner, and a rule that just says “turn it down” is wrong. Nobody has measured this dial inside the assistant on your desk; the controlled results come from research settings, so take the direction, not a percentage. And the maker of that overthinking model warns that on long agent work, lower effort can mean thinner analysis and more retries.

Try it

  1. Pick a fidelity task you actually have. Something whose job is to be faithful rather than clever: condense a long document, pull the dates out of a contract, rewrite a draft into your team’s format.
  2. Find the control first. Look for a reasoning, thinking or effort setting near the model picker or in the chat’s options. If there isn’t one, that is the finding — an administrator set it and removed the choice, and it is a reasonable thing to ask about.
  3. Run the same task twice — highest setting available to you, then lowest — everything else constant.
  4. Check it the right way. Don’t compare the two outputs against each other; both read well and the elaborate one reads better. Compare each against the source, and count two things: claims that aren’t in your document, and instructions of yours that quietly went missing.

Where this breaks. If the task was genuinely analytical — options with trade-offs, a plan with dependencies — the high setting may win outright. Good. That is the trade working, and it tells you which half of your work belongs where.

Additional reading

Editor’s note

No one is an expert at this. The worry that you’ve missed something the model would have caught at a higher setting never quite goes away. But the effect is real, some models show it more than others, and experimenting with the setting is worth your time. There will be times you get better results on the lower one.

signed-off-by: Luke Topfer <editor> · 2026·08·19
06 Self-check

// three assertions against what you just read · results stay in this browser

assert 1/3

The module calls the reasoning setting "a trade, not a dial". What is being traded for what?

assert 2/3

You ask an assistant to condense a 40-page report to one page, and you have the reasoning setting turned up because the report matters. On the module's evidence, why will re-reading the summary not tell you whether that was a good idea?

assert 3/3

A colleague reads the module and proposes a team rule: "set everything to the lowest reasoning level." Which of the module's own limits most directly contradicts that?