There is a number most AI programs can produce on demand, and a number most of them cannot.

The first is time saved per task. Half a day down to twenty minutes. These numbers are usually real, and the people reporting them are not exaggerating.

The second number is the one that shows up in the business. McKinsey's global survey has 88% of organizations using AI in at least one function, and 39% reporting any EBIT impact at all. RAND put roughly 80% of enterprise AI projects in the failed-to-deliver-promised-value column. Gartner expects more than 40% of agentic projects to be cancelled by the end of 2027.

Three explanations get offered for that gap.

  1. The model isn't good enough yet.

  2. People won't adopt it.

  3. The data isn't ready.

Each of those is true somewhere. There is a fourth reason that gets skipped, and I think it explains more of the gap than the other three combined — mostly because it doesn't look like an AI problem at all.

Most enterprise work is not slow. It is waiting.

Touch time and elapsed time

Don Reinertsen spent a career measuring this in product development, and the truth of it stretches far beyond the field it came from. Take any piece of work that crosses three or more functions and measure the time it spends in two states - touch time and queue time. Touch time is when somebody is actually working on it. Queue time is everything else.

A competitive brief takes four hours of work and three weeks to arrive. The four hours are touch time. The rest is the brief sitting in an inbox, sitting in a review, waiting on legal, waiting on a decision nobody has been asked for yet, waiting for Thursday because that is when the group meets.

My estimate is that touch time is under 10% of elapsed time in most enterprise workflows that cross three or more functions. Anyone holding real cycle-time data can check that and tell me I am wrong.

Here is why it matters. AI is extraordinarily good at touch time. By default, it does nothing at all to queue time.

So the four hours become forty minutes. And the three weeks become two weeks and four days.

That is the whole gap between the usage dashboard and the P&L, and it is not a technology failure. The technology did exactly what it promised.

The part that is genuinely counterintuitive

A team running at 95% utilization does not have queues that are marginally longer than a team at 70%. It has queues that are longer by roughly an order of magnitude. Wait time does not rise in a straight line as a resource fills up. It rises off the end of the chart.

Which means 100% busy is not a sign of a well-run system. It is usually the reason the system is slow.

Now put a productivity tool into that. Everyone gets 30% faster at their own step. Almost nobody responds by going home early. They take on more work, because there is always more work. Work in progress goes up. The queues get longer.

Every individual is genuinely faster and the system is measurably slower. Both things are true at once, and the org chart is arranged so that no single person can see both.

That is also why each function reports a different reality. IT is looking at usage and it is up. Finance is looking at the bill and it is up. Operations is looking at cycle times that have not moved. Three true numbers, three different objects, one meeting where they talk past each other.

The pricing makes this sharper rather than softer. Token prices for equivalent capability have fallen something like 98% in two to three years, and at least one industry analysis found enterprise AI bills roughly tripled over the same period. Cheap capability did not reduce cost. It raised the stakes on whether the usage is pointed at anything that matters.

What I would measure instead

Not task time. Measure cycle time, end to end, for one workflow that crosses functions.

Then the ratio between the two — touch time over elapsed time. That single ratio is the diagnosis. If it is 5%, no amount of speed at the touching part will save you, and you can stop arguing about which model to buy.

Then find the longest wait in the workflow and name who owns that step. Not who owns the tool. Who owns the wait.

And price a week of delay in that workflow. If nobody in the room can, the delay will never be prioritized against anything, because everything else on the list has a number attached and this does not.

One thing worth listening for. "We're seeing significant time savings across the team" is not an answer to "did anything downstream get faster." Ask the second question. A real answer names a specific step and says how long work waits there now. A hedge goes back to talking about usage.

Where this is settled, and where it isn't

The settled part: pointing capability at a step that already had slack does not improve the system. Goldratt established that in 1984 and Reinertsen put a price on the waiting. It is not controversial among people who have measured it.

What is genuinely still moving is agents. An agent that carries a piece of work across a handoff is a different animal from one that accelerates a station. The first removes queue. The second shortens touch time. Almost everything shipped so far is the second kind, and whether the first kind survives contact with approvals, controls, and the question of who is accountable when it is wrong is not yet decided.

My call anyway: within eighteen months, the deployments that show up in the P&L will not be the ones with the highest usage. They will be the ones that deleted the most handoffs. Usage is going to look, in hindsight, like a milestone somebody mistook for a result.

Where does that break? If you have watched an AI rollout move a real business number without a single handoff disappearing, that is the counterexample, and I would like to hear it.

— Isaac

Reply

Avatar

or to participate

Keep Reading