watts.it.com // daily AI micro-learning
The landscape leadershipstrategyroimeasurementai-adoption 2026·07·24 · 6 min · dated

Measurement debt: when belief in AI outruns the proof

// listen · 2 ai hosts · audio edition

AI-generated audio discussion of this module — same content, spoken.

Overview

In March 2026, researchers surveyed 528 in-house legal leaders across six countries, most of them at companies with more than US$1 billion in revenue. Two findings came back from the same respondents. Every single team using AI plans to raise its AI budget next cycle — 100 per cent. And 83 per cent cannot demonstrate whether last year’s AI spending delivered results. The same people, in the same survey, are certain enough to spend more and unable to show what the spending bought. Both answers are honest.

The report — Axiom’s 2026 In-House Legal AI Report, published in late June — is the sharpest recent snapshot of a pattern that runs far beyond legal. This weekly is about what that gap actually is, why the obvious reading of it is wrong, and the small set of instruments — including one published by OpenAI’s CFO on 17 July — that leaders are starting to use to close it.

The content

Start with the obvious reading, because a budget committee somewhere is reaching for it right now: if nobody can prove the spending worked, the spending is probably hype, and the prudent move is to cut. The evidence says otherwise — and not because the proof exists somewhere else. It doesn’t. What exists is conviction of a very particular kind: the kind that comes from doing the work. In a small study published the same fortnight — thirty senior in-house leaders across nine countries, all customers of the legal AI firm Legora, so read it as a temperature check of committed users rather than a census — 87 per cent described AI as “essential” to their daily work, and said removing it would significantly disrupt their operations. Axiom’s chief legal AI and talent officer, Sara Morgan, put the sentiment plainly: “Legal AI is no longer a choice for in-house teams; it’s become the baseline.”

Essential, unproven, and funded for growth. Those three facts feel contradictory only if you assume measurement happens by itself. It doesn’t — and that’s the reframe this piece turns on. What the 83 per cent are carrying isn’t failure. It’s measurement debt: the gap that accrues, quietly and by default, every quarter a tool is deployed without anyone deciding what evidence of its value would look like. Like technical debt, nobody takes it on deliberately. It builds while you ship. And like technical debt, it compounds until the day it’s called in — usually by whoever owns the budget spreadsheet, usually at renewal time. The distinction matters because the two readings demand opposite responses. “The AI isn’t working” says cut. “We never instrumented it” says instrument. Getting that diagnosis wrong, in either direction, is expensive.

This is not a legal-sector quirk. When the consulting firm RGP surveyed 200 US finance chiefs in late 2025, only 14 per cent reported seeing clear, measurable impact from their AI investments — and 48 per cent said they are ultimately responsible for ensuring AI delivers measurable value. The bill for measurement debt, in other words, has an address. What makes the legal snapshot striking is the irony, and it’s worth being precise about: a profession whose entire craft is evidence — burden of proof, the documented record, the chain of reasoning that survives challenge — has been running its own tooling on trust. Not through carelessness. Adoption ran on a conviction clock, at the speed of daily experienced usefulness; measurement runs on a design clock, because someone has to decide up front what counts as working. Almost everyone let the first clock set the pace.

Here is the hopeful part, and it’s better than a platitude because it comes from the people who paid for the lesson. Axiom’s report asked teams what they would do differently, and the regrets rank like this: failing to establish measurement practices from day one (37 per cent), neglecting internal expertise before deployment (36 per cent), skipping rigorous tool evaluation (35 per cent), launching too many use cases at once (34 per cent), underinvesting in training (34 per cent), and lacking executive sponsorship (33 per cent). Morgan’s summary of the pattern deserves quoting: “The top disappointments teams report are not about AI performance. They are about the work of adoption: tool selection, implementation time, training burden, and ongoing maintenance.” Her colleague Chris Frickland made the same point about where the work lives: “The technology is not the hard part. The layer around it is where value either shows up or it doesn’t.” Read the regret list once and it’s a confession. Read it again and it’s a sequenced playbook — with measurement, the debt this piece is about, ranked first by the people who skipped it. One honest note: Axiom sells legal talent and AI enablement, so a report finding that teams need help adopting AI is a report that flatters its author’s business. The regret data survives that discount — it is respondents grading their own past decisions, not the vendor grading them.

Which brings us to the instrument. On 17 July, OpenAI’s CFO Sarah Friar published what she called a scorecard for the AI age, built on a metric she names “useful intelligence per dollar” and four questions: Is AI completing work that matters? What does each successful task cost? Can people depend on the result? Does each AI dollar produce more value as usage grows? Her framing of the underlying economics is genuinely clarifying: “The basic economic question facing CFOs and other business leaders is whether the value of the work AI completes grows faster than the cost of producing it.” The most practical piece is the smallest: classify AI output into three buckets — ready to use, needs correction, needs escalation — and track the full cost of a successful task, not the cost of a token. You could run that on one workflow, with a spreadsheet, starting Monday.

And you should also notice who’s offering it. Friar runs the finances of the company selling the intelligence, and her central move — measure the full cost of a successful outcome rather than the price of tokens — is, among other things, an argument for capable, expensive models over cheap ones. That doesn’t make the questions wrong; they’re good questions. It makes them a vendor’s questions. Treat the scorecard the way an experienced negotiator treats a counterparty’s first draft: a useful starting text whose defaults you check against your own position before you sign. The metric that finally matters will be denominated in your outcomes, not theirs.

The teams in Axiom’s 7 per cent — the ones that scaled — aren’t the ones that spent the most; the report’s through-line is that the differentiator was operating habits, the unglamorous layer around the tool. Measurement debt is unusual among the debts an organisation can carry: it can be repaid in fortnight-sized instalments, one workflow at a time, starting with the work you already believe in most. The point is not to slow conviction down. It is to let evidence catch up — because “100 per cent plan to spend more” and “83 per cent can’t show what it bought” cannot share many more budget cycles. The teams that instrument first get to keep their conviction, and prove it.

Additional reading

Editor’s note

My opinion (unsupported by research or collaboration) is that the measurement gap hides an exciting and unrecognised feature: that often value measurement is made even more difficult because people are still changing fundamental things about the way they work, and what their “work” actually is. It’s like building a highway — it takes a while, but when it’s done, the journey from A to B is cut in half. Enterprise work takes a long time to change, even at the most aggressively change-focused organisations. Issues with measurement may not reflect a lack of value. I think the highway is just unfinished.

signed-off-by: Luke Topfer <editor> · 2026·07·24
05 Self-check

// three assertions against what you just read · results stay in this browser

assert 1/3

In Axiom's 2026 survey, 83% of legal teams can't demonstrate whether last year's AI spending delivered results — yet 100% plan to raise their AI budgets. What does the module say that gap actually is?

assert 2/3

Your renewal committee sees strong usage of an AI tool but no proof of value, and is leaning toward cutting it. Following the module, what is the move that resolves the standoff?

assert 3/3

The module calls Friar's 'useful intelligence per dollar' scorecard genuinely useful AND says to treat it like a counterparty's first draft. Why both?