A few months ago I asked a category manager a simple question. Her AI tool had recommended consolidating a category onto two suppliers instead of five. The output looked confident. The slide was clean. There was a savings number attached to it.

I asked how she knew the recommendation was actually right. Not plausible. Not well formatted. Right.

She paused for a while. Then she told me nobody had gone back to check. The recommendation had been actioned three months earlier. Nobody had compared what the tool predicted to what actually happened. Nothing in the process asked anyone to. The tool moved on to its next recommendation. So did she.

That gap, between a decision being made and anyone finding out whether it was the right one, is what this last piece in the series is about. It's the one I found hardest to write. Not because it's complicated. Because it isn't. It's just something nobody's doing.

The verification step that quietly disappeared

Every procurement decision used to carry a checkpoint. A category manager recommended a supplier switch. A manager asked why. Someone with experience pressure tested the logic before it went anywhere. It was slow. It was human. It was imperfect. But it was real.

AI didn't remove that filter on purpose. It removed it by being fast enough and confident enough that the filter felt unnecessary. A recommendation that arrives instantly, formatted cleanly, with a number attached, doesn't invite the same scrutiny a colleague's hunch would. It looks like it's already been checked. It hasn't. It's been calculated.

Calculated and checked are not the same thing. Procurement has spent the better part of two years quietly treating them as synonyms.

The three places verification used to happen

Before the decision. Someone used to sanity check the inputs. Is this the right supplier list? Is this cost baseline current? Does this category actually behave the way the model assumes? That step still technically exists in most organisations. It's just been quietly demoted from required to optional, if time allows. Time rarely allows.

At the decision. A second person used to review the recommendation before it became an action. Procurement has relied on this maker checker control for decades. AI recommendations increasingly skip it. Not through a deliberate governance decision. Through accumulated convenience. The tool is trusted because it's usually right, and usually right quietly becomes assumed right.

After the decision. This is the one that's disappeared most completely. It's also what this article is really about. Nobody goes back. Nobody compares the predicted savings to the realised savings, or the predicted risk to what actually happened. The loop that should close, from recommendation to action to outcome to comparison to correction, simply doesn't.

WHAT HAPPENS AFTER THE RECOMMENDATION The Open Loop What most organisations actually run AI recommendation generated Action taken Outcome happens somewhere ✕ nobody looks Loop ends here The model never learns it was wrong. Confident errors compound quietly, recommendation after recommendation. The Closed Loop What verification actually requires AI recommendation generated Action taken, outcome logged Predicted vs. actual compared Gap owned by a named person The model, and the team, actually get better over time. A recommendation is a hypothesis. Only the closed loop tells you if it was true.

Why nobody's checking

This doesn't happen because people are careless. It happens for three reasons that are almost reasonable on their own. That's exactly why it's hard to catch.

Speed became the metric. Cycle time reduction is one of the easiest wins to show after an AI deployment, so it's the one everyone measures. Nobody tracks accuracy over the same period. Accuracy takes longer to observe. It doesn't fit neatly into a quarterly update. What gets measured gets managed. Speed gets measured. Correctness doesn't.

Trust arrived without a way to test it. The first few times a tool gets something right, confidence builds. That part is reasonable. What doesn't happen is a deliberate decision about how much trust has actually been earned, on what basis, and how it will keep being tested. Trust accumulates by default rather than by design. Defaults don't include an ongoing audit.

Nobody owns the outcome. The person who generated the recommendation isn't necessarily the person accountable for whether it worked. The category manager moves to the next sourcing event. The savings figure gets reported up as projected. By the time actuals would be available, attention has often moved on too. Verification needs someone still looking three, six, twelve months later. Almost nobody's job description says that.

A recommendation is a hypothesis wearing the outfit of a conclusion. The only thing that turns it into a conclusion is someone checking.

What checking actually looks like

This doesn't need a research function or a data science team. It needs verification treated as a designed step, not a hoped for byproduct of someone's diligence.

Log the prediction, not just the action. When an AI tool recommends a savings figure, a risk score, or a supplier ranking, write that number down at the moment it's made. Do it before the action. Do it before hindsight quietly adjusts anyone's memory of what was actually predicted.

Set a return date. Every significant AI influenced decision should carry a date to come back and compare prediction against reality. Not eventually. A real date, on a calendar, assigned to a real person.

Track the gap, not just the hits. Organisations love reporting the wins, the times an AI recommendation matched or beat expectations. Almost nobody builds a habit of tracking the gap between prediction and reality as its own metric. That gap, tracked over time, tells you far more about whether to trust the next recommendation than any accuracy claim in a vendor deck.

Rotate the reviewer. The person checking the outcome shouldn't always be the person who acted on the recommendation. Self verification has an obvious blind spot. Nobody enjoys finding out their own call was wrong, and that quiet reluctance shapes how thoroughly anyone checks their own work.

The specific danger of agentic systems

Everything above gets harder as procurement moves toward the agentic systems I wrote about in part two of this series. When AI only recommends, a human decision still sits between the model and the consequence. It's an imperfect checkpoint, but a real one. When AI acts directly, that checkpoint disappears. Verification has to be rebuilt somewhere else on purpose, because it won't appear on its own.

An agent that routes a purchase request, selects a supplier, and triggers a workflow on its own is making dozens or hundreds of small decisions a day that no human ever individually reviews. That's the entire point of deploying it. But nobody reviews each decision and nobody reviews any decisions are supposed to be different things. In practice, for a lot of organisations right now, they've quietly become the same thing.

Sampling matters here more than anywhere else. Not every agentic decision needs a human check. A statistically meaningful sample of them does. Reviewed regularly, by someone with the authority and the actual time to flag a pattern before it becomes a habit the agent keeps repeating at scale.

The uncomfortable question for leadership

I'd put this to any procurement leader currently scaling an AI supported process. If I asked you right now for the ten highest value recommendations your AI tool made this quarter, could you tell me what actually happened for each one? Not what was projected. What happened.

Most leaders I've asked can't answer that. Not because the outcomes were bad. Because nobody built the mechanism to know either way. That's a more serious gap than a bad outcome would be. A bad outcome at least generates a signal. Silence generates nothing. Nothing is exactly what most organisations are currently getting back from their AI investment, in terms of genuine, checked evidence.

This isn't an argument against AI in procurement. Every article in this series has rested on the opposite premise, that AI, done properly, changes procurement for the better. It's an argument that "done properly" has a piece almost nobody budgets for, staffs for, or puts on a governance agenda. The discipline of finding out, deliberately and on a schedule, whether the machine was right.

Closing the series

Four articles, one thread running underneath all of them. Data quality is the blind spot nobody funds. Agentic AI needs a fundamentally different kind of data than analytics AI ever required. Most procurement functions are still visiting their data rather than living in it. And even where the data and the maturity are both in decent shape, almost nobody is closing the loop to check whether the AI's confidence was actually earned.

None of these are model problems. All four are design and discipline problems. The unglamorous, unfunded, rarely celebrated work that decides whether AI in procurement becomes a genuine capability, or just an expensive way to be wrong faster.

Fix the data. Understand what kind your AI actually needs. Move from visiting the numbers to living in them. And then, the part this article was written to insist on, go back and check.

Fourth and final piece in the series on data, AI, and procurement. Part one looked at data quality as procurement AI's blind spot. Part two looked at why agentic AI needs fundamentally different data. Part three looked at the shift from analytics tourist to analytics resident. Thank you for reading the series through. The inbox is open if any of it landed.