watts.it.com // daily AI micro-learning
Tools & connectors visionmultimodalscreenshotsefficiency 2026·07·14 · 4 min · dated

Show, don't tell: paste the screenshot instead of describing it

Overview

On 16 April 2026, Anthropic shipped an upgrade that drew less notice than the benchmark numbers but changes more day to day: Claude’s image resolution roughly tripled. The announcement said the model could now accept images “up to 2,576 pixels on the long edge (~3.75 megapixels), more than three times as many as prior Claude models” — an increase aimed, in Anthropic’s own words, at “reading dense screenshots” and “data extractions from complex diagrams.” That higher-resolution reading carries forward to its current models. Reading a packed dashboard or a small-print form used to be a coin-flip. Now the fine detail comes through.

The practical version: the thing on your screen you were about to describe in words, you can just show.

This module gives you one habit — and a clear sense of where it still slips.

The content

The obvious read is that image input is a novelty, handy for asking what breed the dog is. The more useful read is about who does the typing. Every time you retype the figures off a dashboard, summarise a chart trapped in a PDF, or describe a baffling settings screen, you are working as a slow, error-prone layer of transcription between your screen and the model. Call it the retyping tax: the minutes you burn turning what you can already see into words the AI could have read for itself.

The reframe is that the gain isn’t the model describing a picture back to you. It’s that you stop paying the tax. Screenshot the dashboard and ask which region is trending down and by how much. Photograph the whiteboard after the workshop and ask for the actions as a list. Drop in a scanned table and ask for it back as rows for your spreadsheet. You skip the retyping — and the transcription errors you’d have added along the way. It runs the other way, too: when the AI hands back a visual that misses — a chart, a slide, a diagram — screenshot its own output and tell it to look at what it made, instead of describing the problem in words.

That is why the April resolution jump shifted the calculation rather than the spec sheet. Dense material — a figure with ten series, a form in small print, a spreadsheet at full zoom — was the exact case where earlier vision guessed. Anthropic now lists its high-resolution tier for current models such as Opus 4.8 and Sonnet 5 explicitly for “computer use, screenshot understanding, and dense documents,” and the same paste-an-image move is standard across today’s assistants — Gemini, ChatGPT, Copilot. The capability crossed from sometimes to usually.

Usually is not always, and the honest move is to know where it slips before you trust a number pulled off an image. Anthropic’s own guidance says the model “might hallucinate or make mistakes when interpreting low-quality, rotated, or very small images,” gives only “approximate” counts of objects, and — plainly — “cannot determine whether an image is AI-generated,” so it is no help deciding whether a picture is real. There is a mechanical catch as well: the model reads an image in small patches and shrinks anything oversized to fit, so a full 4K screenshot can lose the very small text you needed. Send the whole screen, and the label you were after can quietly disappear.

Try it

No prompt to memorise. Change one reflex this week.

The next time you catch yourself about to type out something that is
already on your screen — a table, a chart, a dashboard, a form, a
slide, a photo of a whiteboard — stop. Screenshot it, paste it in,
and ask your question about the image directly:

  "Pull the figures in this table into rows I can paste into a sheet."
  "Which line is falling here, and roughly by how much?"
  "List the actions written on this whiteboard."

Then verify: check one figure it read back against the original
before you rely on it.

Where it breaks, deliberately: take a dense dashboard at full 4K, send the whole thing, and ask for a small number buried in a corner. The model may misread it — not because it can’t see, but because your screenshot was downsized to fit and the small text blurred on the way in. The fix is a habit, not a cleverer prompt: crop to the panel you care about and enlarge it before you paste. Show it less, and it sees more. And whatever comes back, check the one figure that matters against the source — vision is at its most confident on exact values from a chart exactly where it is weakest.

Additional reading

Editor’s note

This skill is one that I developed out of pure frustration. I can’t tell you how many times I hated a visual output and directed an agent to “Look at what you’ve done”. I am consistently surprised, even today, at how well it works to use images. The risk is fairly low; the reward can be fairly high. Keep it in your tool box.

signed-off-by: Luke Topfer <editor> · 2026·07·14
06 Self-check

// three assertions against what you just read · results stay in this browser

assert 1/3

The module names a cost it calls "the retyping tax". What is it?

assert 2/3

You ask an AI assistant to build a chart for a client deck and the result is off — wrong emphasis, cluttered labels. Following this module's advice, what do you do next?

assert 3/3

You paste a full 4K screenshot of a dense dashboard and ask for a small figure buried in one corner. The number comes back wrong. What is the most likely cause?