Every AI time-saving number you've read is a survey answer. The one time anyone used a stopwatch, it was off by 40 points.
You've seen the stat, in some form, about forty times this year. Businesses using AI save twenty hours a month. Or eight hours a week. Or thirty percent. It's on the pricing page of every tool you've been pitched, and it's usually printed in a big friendly number next to a photo of someone smiling at a laptop.
Here's the thing almost nobody says about those numbers, and it isn't that they're lies. It's that virtually every one of them was produced the same way: somebody was asked how much time AI saved them, and they answered. That's it. That's the methodology. Which raises a question worth about twenty hours a month to you — how good are people at answering that question?
The one time somebody actually held a stopwatch
In 2025 a research outfit called METR did the experiment nobody does, because it's expensive and slow and nowhere near as fun as a survey. They took 16 experienced open-source developers, gave them 246 real tasks in codebases they already knew well, randomly allowed AI on some tasks and not others, and recorded their screens. Actual measurement, not recollection.
Beforehand, the developers expected AI to make them about 24% faster. Afterwards — having done the work — they reckoned it had made them about 20% faster. The recordings said they were 19% slower. Both estimates missed the same way, and the gap between what these people believed about their own morning and what the tape showed was around 40 percentage points.
Now, the honest caveats, because they matter and because METR themselves lead with them. Sixteen developers is a small group. They were working in mature repos they knew intimately, which is the setting where an outside suggestion helps least. Most had only modest hours on the AI tooling. And the authors have since labelled the whole result historical — their explicit position is that it no longer reflects what current models do. If you want to read that study as "AI makes you slower," you're reading it against the wishes of the people who ran it, and against a year and a half of model releases.
So drop that reading. It isn't the interesting part anyway. The finding that survives every one of those caveats is the one about the humans: a room full of skilled professionals, working on their own familiar code, watching their own hands, could not tell whether the tool had sped them up or slowed them down. Not a bit off. Wrong in direction. That result has nothing to do with which model was current in 2025 and doesn't expire when the next one ships.
Then the same researchers went and asked people anyway
Here's what makes this genuinely useful rather than just cynical. In 2026, METR ran a survey — 349 engineers, researchers and founders, asked in the first few months of the year — knowing full well what their own trial had said about surveys. Their write-up says the quiet part plainly: survey results are "not necessarily grounded in reality." Publishing a survey with that sentence attached is a rare kind of honesty and worth borrowing.
But they did one clever thing. They asked two different questions instead of one. First, how much faster does AI make you — the usual question. Median answer: three times faster. Then they asked how much more you actually contribute because of it. Median answer to that one: somewhere between 1.4 and 2 times. Same people. Same tools. Same week.
That split is the most practical thing in any of this research. Speed and value are not the same measurement, and the gap between them is where the promised hours quietly go. A quote drafted in ten seconds instead of ten minutes is a genuine three-times speedup on that step — and worth nothing at all if it then sits in your drafts folder for two days, or gets rewritten because it priced a job you'd never take. The step got faster. The business didn't.
The view from the top of the market is stranger still
Sitting alongside all the enthusiastic survey numbers is a National Bureau of Economic Research study from early 2026 that surveyed something like 6,000 executives across the US, UK, Germany and Australia. Close to 90% of those firms said AI had made no measurable difference to productivity or employment across the previous three years.
Resist the obvious conclusion. That is not proof the technology doesn't work, and the same study contains the reason why: two-thirds of those executives reported using AI, but their average usage came to something like an hour and a half a week. Roughly a quarter weren't using it at all. Ninety percent seeing no productivity change is at least as likely to be a story about tourism as a story about the tools. You don't restructure a business around ninety minutes a week, and you don't show up in the productivity statistics for it either.
Which lands us somewhere more useful than either side of the argument wants. The people using AI heavily are reporting big gains they can't reliably verify. The people barely using it are reporting nothing, which is exactly what you'd predict. Almost nobody in either group is measuring anything.
Where the saved time actually goes
One more finding, from BCG's 2026 workplace research across roughly 12,000 frontline employees. Forty-two percent said regular AI use saved them about eight hours a week. Two-thirds said they'd had little or no guidance on what to do with the time.
Read those two together and you've got the whole productivity paradox in miniature. Time saved but not reassigned isn't saved. It's absorbed — into the inbox, into a slightly longer lunch, into the work expanding gently to fill the day, the way work always has. In a big company that's a management failure. In a two-van operation it's just Tuesday, because nobody is going to hand the owner guidance on what to do with a recovered forty minutes.
This is the actual leak in most small-business AI spending. Not that the tool didn't work — it probably did, on the narrow task, roughly as advertised. It's that the forty minutes never got pointed at anything. And because nothing downstream changed, the only evidence you'll ever have that the subscription is earning its keep is a feeling. Which we've now established is the one instrument known to be broken.
How to actually know, in about ten minutes of effort
The fix here isn't sophisticated and it isn't a dashboard. It's the thing the entire industry skipped on the way to quoting itself.
- Pick one task, not five. The quote you write most often, or the first reply to an enquiry. One task you do repeatedly, that you'd recognise on your own calendar.
- Time it five times before you change anything. A phone stopwatch is fine. Five is enough to see the spread, which matters more than the average — a task that swings between four minutes and twenty is a different problem to one that's reliably nine.
- Write the number somewhere that isn't your head. This is the whole discipline. Your memory of what invoicing used to take will helpfully revise itself the moment you start paying for something that promises to fix invoicing.
- Decide in advance where the time goes. "This buys me the Thursday afternoon I currently spend on quotes" is a plan. "This will free me up" is how eight hours a week evaporates without trace.
- Re-time it after a month of genuine use, not on day one. Then ask the second question — not just whether the step got faster, but whether more work actually got out the door. If speed went up and output didn't, you've found the gap, and it isn't in the AI.
Ten minutes of stopwatch beats any vendor's case study, including ours, because it's measured on your jobs, your customers and your Tuesday. And the perception gap isn't a failure of intelligence — it's how memory works. You remember the reply the AI nailed in four seconds. You don't remember the three you quietly rewrote, because rewriting felt like working, and working doesn't file itself as evidence.
Where we land on it
We should be straight about our position here: we sell AI features, so we have precisely the same incentive as everyone else to wave a big self-reported number at you. We'd rather tell you what the AI in Dispatch actually does, which is narrow and checkable — it reads your enquiries and drafts the replies, and quotes the standard jobs off your own price list. The send button stays behind you. The AI parts are priced as add-ons with a stated allowance and metered usage past it, rather than folded into an "unlimited" badge, because we know what they cost to run.
Every one of those claims is a thing you can point a stopwatch at, which is the only reason we'd make them in that form. Time your first replies for a fortnight, then time them again. If the number doesn't move, we haven't earned the subscription and no survey we could commission would change that.
That's really the takeaway, and it applies well beyond us. The AI is probably fine. The models are extraordinary and getting better on a schedule that's hard to keep up with. The broken instrument in this whole story is the one every business is currently using to evaluate them — asking yourself how it went, and believing the answer. Sixteen developers watching their own screens got that question wrong by forty points. You will not do better by feel. But you'll beat all of them with a phone stopwatch and a note on the fridge.
Software for service businesses — built by an operator.
Job management, books, and AI agents that actually know your business.