The most useful sentence anyone has given me about AI this month was said in a governance meeting, by one of our delivery leads, and nobody in the room was trying to be profound.
"AI has just helped us move the bottleneck further along. All we've done is accelerate the three months into three hours at the front, but then the backside problem is still the same."
Three months into three hours is a genuine, verifiable result. It is also, on its own, worth nothing. That observation was the most valuable thing produced in our business that week — more valuable than the automation it was describing.
The market is still arguing about whether AI pilots work. That argument is a distraction, because in the mid-market the pilots increasingly do work — and that is precisely where the trouble starts. A successful automation does not remove a constraint. It relocates it, usually to a step that was never measured, never documented and never owned, and almost always to the step that is hardest to automate next.
Nobody sells that step. There is no licence for it, no SKU, and no demo. Which is exactly why your provider automated the other end.
Here is the shape it takes. If you run a service operation of any size, you will recognise it.
Somebody identifies a genuine risk — say, a security tool that needs checking every morning for red flags. Somebody builds a correct automation to raise a ticket for that check. The automation fires perfectly, every morning, into a team that has never been shown where in that tool the red flags actually appear. The tickets arrive. They cannot be actioned. They sit.
Then the second-order effect arrives, and it is the one nobody models. A queue that fills with work nobody can complete does not stay still. It gets managed rather than worked — moved, reassigned, closed — because that is the only rational response available to people who are measured on queue depth and given no instruction. The risk that prompted the automation is now less likely to be caught than it was before anyone automated anything. The dashboard, meanwhile, reports that the check runs daily.
Read that sequence again, because it is the entire failure mode in miniature. A real risk was identified. A correct automation was built. It ran flawlessly. And the net effect on the organisation was negative.
I have seen versions of this inside my own operation, which is why I can describe it so precisely — and it is why we changed how we work. We now treat an automation as unfinished until the person receiving its output can demonstrably act on it, and we are rewriting the downstream work instructions ahead of the automations that feed them rather than after. That work is slow and it is not finished. It is also the only part of this that actually creates value, which is why we would rather be measured on it than on how much we have shipped.
The reason those queues cannot be actioned is almost never a people problem. It is a substrate problem. The work instruction for the downstream step — where one exists at all — has very often been copied from somewhere else, never checked, never owned, and written as plain undescriptive text with no screenshots and no named author. That is the material both a human and an agent have to act on.
This is the thing the industry does not want to say plainly. We have all been selling the front end — summarise this, draft that, triage the other — because the front end demonstrates beautifully in a forty-minute meeting. The back end is unglamorous, unbillable and slow, and it is where the value actually gets realised or lost. Documentation is the least loved asset in this industry, mine included, and it is now the one everything else depends on.
I saw the same geometry on the client side twice in one week. At a large utility, the head of cloud walked me through a reporting process that runs data out of a service management platform into a spreadsheet and then into an AI assistant, by hand, every cycle. The AI step is instant. The join either side of it is a person. At an aged care provider, an admissions team spends roughly three hours in conversation with a family and two hours writing it up. Automate the write-up and the three-hour conversation becomes the constraint — and no model is going to shorten a conversation with a family.
I am not arguing that none of this works. Our monthly Microsoft 365 environment health check used to cost a full engineering day per client, every month. It now takes about ten minutes. That saving is real, it is measured, and it has held.
What matters is why it held when the ticket automation in the example above does not. The output had an owner who already knew what to do with it. It fed a report that already existed. It sat in a process that was already documented. The automation slotted into a chain that was complete before we touched it.
That is the whole difference. Not model quality. Not prompt craft. Whether the next link in the chain was owned before you automated the previous one. Every genuine win I can point to has that property, and every disappointment I can point to is missing it.
All of this applies to my own business, which is why I am willing to write it down. We shipped an AI triage tool recently. The delivery lead who built it refused to call it a win until he had gone back to the people using it and asked whether it was helping. He was right to refuse. That instinct is the thing I would most like to buy more of, and nobody sells it.
Our industry measures the wrong thing. It shipped, it did not break, the licence is live — so we call it value. It is not value. Measure shipping and you will keep automating the easy front end, keep handing clients a longer queue at the back, and keep calling that transformation.
So do not ask your provider what they automated. Ask what happened to the step straight after it. If they cannot answer with a number, about a step they do not bill you for, they have sold you a demo with a monthly fee attached.
Ask them where the bottleneck went. If nobody knows, it went somewhere.