The supervision paradox: why more capable AI means more babysitting, not less
Overview
This weekly is about a contradiction you may have felt in your own week: the better AI gets at working on its own, the more of your time seems to go on watching it — not less. If that runs against everything the tools were sold on, this is a map for it.
Why now: on 10 June 2026, Glean’s Work AI Institute published its Work AI Index, a survey of 6,000 full-time digital workers — 1,500 of them Australian. It found that the average worker now spends 6.4 hours a week — most of a working day — on what it calls botsitting: feeding AI the context it’s missing, checking its output, and cleaning up the confident mistakes it leaves behind. AI is handing people back around eleven hours a week and quietly taking a chunk of it straight back. And here is the finding that should stop you: the people getting the most out of AI botsit more than everyone else, not less.
By the end you’ll have the paradox, the reason it grows as the models get better rather than fading, the 1983 paper that predicted it, and the discipline that keeps supervision from curdling into something worse.
The content
Start with the promise, because it is real. 87% of digital workers now use AI at work — nearly nine in ten — most of them say it makes them more productive, and the Glean survey puts the time saved at roughly eleven hours a week per person. The natural expectation from there is a smooth one: as the tools get more capable and more autonomous, they need less hand-holding, so the watching-over shrinks towards zero. You hand over the task and walk away.
That is not what the data shows, and it is not what practitioners report. The same survey finds 6.4 hours a week going the other way — into botsitting. Break down all the time people spend with AI and the shape is stark: 37% is botsitting, 36% is actually using it to get work done, and 27% is learning and building. More than a third of your AI time is spent tending the AI. The obvious read is that this is transitional friction — teething trouble that better models will iron out. That read is wrong, and why it’s wrong is the whole point of this edition.
Here is the overturn. Supervision doesn’t shrink as the AI improves; it scales with it — because the thing that rises with capability isn’t your freedom, it’s your appetite. The moment you trust a task enough to let it run unwatched is the same moment you feel free to start three more. Capability buys you parallelism, and parallelism is just supervision wearing a more flattering outfit. Call it the supervision paradox: every increase in what the AI can do on its own increases the number of things you are now, at once, responsible for overseeing. You didn’t automate your way out of the loop. You got promoted — from doing the work to supervising a fleet of it.
The Glean data has this hiding in plain sight. The high achievers — the people who report gains in both productivity and quality — spend more of their AI time botsitting than everyone else: 40% against 33% for low achievers. They are not the ones who found the trick to stop watching. They are the ones who watch best. As the report puts it, they “don’t just prompt and pray”; they use their judgement. Getting more out of AI and spending more effort supervising it turn out to be the same skill.
None of this is new, which is oddly the reassuring part. In 1983 the cognitive psychologist Lisanne Bainbridge wrote a short paper on industrial automation, “Ironies of Automation”, and its central line reads as though it were written for this month: “By taking away the easy parts of the task, automation can make the difficult parts of the human operator’s task more difficult.” Her point was that automating a process doesn’t remove the human — it leaves them the residue the machine can’t do, which is usually the hardest part: monitoring, judging, catching the edge case, stepping in when it goes wrong. Forty-three years later, that residue has a new name and a far bigger surface area. The easy part — a competent draft, a first-pass analysis, a working block of code — is exactly what got automated. What’s left for you is the hard part: deciding whether it’s right.
And that hard part is genuinely hard. Sonar’s State of Code survey in January 2026 found that among developers — the group furthest down the agentic road — 38% say reviewing AI-generated work takes more effort than reviewing a human colleague’s, and only 48% always check AI output before they ship it. The bottleneck moved. It used to be that producing the work was the slow, expensive step; now producing it is cheap, and checking it is the slow, expensive step — “a new bottleneck at the verification stage”, in the report’s words. And unlike production, checking hasn’t been automated, for the obvious reason that the thing you’d automate it with is the thing you don’t fully trust.
So far this is just more work. The sharper danger is what parallelism does to the quality of your supervision. In a 2010 review of decades of human-factors research, Parasuraman and Manzey found that automation complacency — the drift into under-checking a system you’ve come to rely on — “occurs under conditions of multiple-task load, when manual tasks compete with the automated task for the operator’s attention.” Read that against your own bank of open agents. Running many streams at once is not a neutral way to get more done; it is the precise condition under which human oversight is known to decay. They found it in experts as readily as novices, and that it “cannot be overcome with simple practice.” You do not out-discipline this by trying harder.
This is where Glean’s darker coinage earns its place. Botsitting is the honest grunt work of supervising AI. Botshitting — their word — is what happens when that supervision quietly lapses: you ship the first output that looks good enough instead of one you could stand behind, and when it turns out wrong, you blame the tool. The survey found 69% of AI users admit to it, and that when AI-assisted work fails, 40% blame the AI while only 29% own the mistake. The slide from the first to the second isn’t a moral failing or a single reckless click. It is the ordinary result of running more parallel work than your attention can actually cover — which is exactly what the paradox tempts you into.
Two honest qualifications, because the strong version can mislead. First, be wary of the very feeling that got you here. METR’s controlled study in 2025 found experienced developers were 19% slower using AI on their own projects — while believing they had been faster. The sensation of “I’ve got room to run more in parallel” is exactly the kind of self-report that result should make you hold loosely. The confidence is real; whether the throughput is real is a separate question you have to actually check. Second, no one has yet measured how many agents one person can properly supervise — there is no published number for your span of control. The paradox is well-triangulated and it has a clear mechanism, but treat it as a limit to go looking for in your own work, not a figure to assume.
What the high achievers show is that the answer is not to botsit less. It is to botsit well, and to stop your fleet at the edge of what you can actually govern. Three moves fall out of that.
Size the fleet to your attention, not the tool’s capacity. This generation of tools will happily spawn a dozen agents at once; whether you can direct a dozen is a different question, and the honest answer is usually a much smaller number. Your binding constraint is not compute or licence seats — it is how many parallel streams one mind can hold a real model of. Run to that limit, not the software’s.
Decide the checkpoints before you start, not after it breaks. Supervision that’s ambient — a vague intention to “keep an eye on things” — is the supervision that lapses first under load. Settle in advance what you’ll verify on each stream and when you’ll look, so checking is a scheduled act rather than a guilty afterthought. Judging work you didn’t watch get made is its own discipline, and worth building deliberately.
Watch for the tells that sitting is turning into shitting. Two are reliable: you notice yourself shipping something because it looks right rather than because you confirmed it is, and you notice yourself reaching for “the AI got it wrong” when a mistake surfaces. Both mean your span of control is already exceeded. The fix isn’t more willpower; it’s fewer streams.
The promise of autonomous AI was that it would take the work off your hands. The quieter truth is that it changes what the work is — from producing the output to being accountable for output you increasingly didn’t produce. That job doesn’t get smaller as the models get better. It gets larger, and more consequential, and more clearly the part that was always yours.
Additional reading
- The Work AI Index 2026 — Work AI Institute / Glean (10 Jun 2026) — the source of the botsitting numbers: 6.4 hours a week, the 37/36/27 split, the high-achiever finding (40% vs 33%), and the “botshitting” tail.
- Ironies of Automation — Lisanne Bainbridge, Automatica (1983) — the paradox stated 43 years early; the verbatim line is reproduced and discussed in this readable Human Factors 101 summary.
- Complacency and Bias in Human Use of Automation — Parasuraman & Manzey, Human Factors (2010) — why over-trust degrades oversight specifically under multiple-task load, and why practice alone won’t fix it.
- Sonar State of Code Developer Survey 2026 (Jan 2026) — verification, not generation, as the new bottleneck; 38% find reviewing AI work harder than reviewing a colleague’s.
- Measuring the Impact of Early-2025 AI on Experienced Developers — METR (Jul 2025) — developers 19% slower with AI while feeling faster; the caution against trusting the confidence.
- Time Horizon 1.1 — METR (Jan 2026) — the length of task frontier agents can run unattended, still climbing; the precondition that makes parallel supervision the new normal.
- How we built our multi-agent research system — Anthropic (Jun 2025) — parallel agents multiply coordination cost, and human evaluation “catches what automation misses.”
Editor’s note
Earlier this week I ran a block of work using six primary agents running in parallel, and the amount of botsitting that I had to do was a lot. This isn’t an unusual occurrence for me, but what I find curious is that the amount of botsitting that is needed is not necessarily linked only to the number of agents you’re running at one time, but also the type of work that they’re doing. In that session, two of the agents that I ran had the explicit instruction to run their job to completion without pausing to ask me any questions — if they needed anything, they were to make a decision and tell me at the end of their run what the decision was. The work of those agents suited that instruction, but most knowledge work — where you’re working towards output that is functionally deliverable to someone else — does not. As you get better at orchestrating a fleet of agents, the temptation will often be to do more and oversee less; recognise that as you move that way, your judgement around whether the work is good is itself a new skill, one that requires discipline, sharpening and maintenance.
// three assertions against what you just read · results stay in this browser
The Glean survey found the highest achievers spend more of their AI time botsitting than everyone else, not less. What mechanism explains supervision growing as models improve?
Your AI tool can spawn a dozen agents at once and you have a heavy delivery week ahead. Following this module's advice, how should you set the work up?
Mid-week you catch yourself shipping an agent's output because it looks right, and reaching for 'the AI got it wrong' when a mistake surfaces. What do these two tells indicate?
Was this useful for your daily work?