Zendoric
← Back to the day · July 26, 2026

Agentic AI: 60% of spending goes to correcting its own answers, and 93% of companies are already over budget

🕒 Published on Zendoric: July 26, 2026 · 00:23

McKinsey reveals that 93% of companies are already over their agentic AI budget, even though the price of a token has collapsed. The cause is not the model: 60% of spending goes to the review and self-correction cycles agents run before delivering a useful answer.

By Zendoric · July 26, 2026.

The opening figure dismantles the most widespread intuition about the cost of AI. According to Stanford HAI's AI Index 2025, cited in a new McKinsey report, processing one million tokens (the smallest unit of text a language model bills for) with capability equivalent to GPT-3.5 cost 20 dollars in early 2024 and fell to 0.07 dollars before the year was out. The token, in theory, has become almost free.

And yet corporate spending on AI has tripled in the twelve months to the end of 2025, according to a Menlo Ventures study cited in the same report. In May 2026 McKinsey surveyed 75 qualified participants across five major sectors for its Enterprise AI FinOps Survey (FinOps being the discipline of managing and controlling cloud spending, now applied to AI), and the result is emphatic: 93% of organisations have already blown past their AI budget. One in five participants in McKinsey's forthcoming State of AI survey, with 1,719 responses collected between May and June 2026, admits to having actively restricted AI use for reasons of pure cost, not for lack of technical capability or internal resistance.

The report, published in McKinsey Quarterly in July 2026 by researchers at QuantumBlack (McKinsey's analytics and AI arm), names the exact culprit: 60% of agentic spending does not go on the first inference call, but on what McKinsey calls "response refinement": the iterative cycles in which an agent reviews, corrects and improves its own output before delivering it as a usable result. It is the part of the process for which most corporate budgeting models never created a line of their own.

McKinsey identifies two structural causes. First, the shift from subscription pricing to consumption pricing at the big model providers: when you pay per token generated, every longer answer is more expensive, and the system's own design rewards verbosity. Second, the habit of routing trivial tasks through frontier models, sized and priced for complex problems the task at hand did not require. Added to that is the scale effect: a pilot with a controlled token volume does not behave like an agent in production orchestrating several subagents, each refining its own output before passing it to the next. The cost curve stops being linear at the exact moment the company thinks it has finished experimenting and starts deploying for real.

David Tepper, CEO of Pay-i (cited in the report), sums up the shift in focus McKinsey proposes with a line that works as a headline in itself: tokens are not value, tokens are the bill. The question McKinsey puts to senior executives is not how much the token costs, but whether what the agent produces is worth more than the total cost of producing it, including human oversight and downstream error correction. An agent three times more expensive per task but requiring no human review can work out cheaper in real terms than a cheap one that demands constant correction. Neither the token price nor the model tier tells the whole story on its own.

Our reading is that this report documents, with figures, the first serious economic symptom of agentic AI, and it fits a pattern we have seen before in other infrastructure technologies: capacity falls in price faster than the governance discipline needed to exploit it well. It is not proof that agents do not work, nor that the promise of agentic AI is smoke; it is proof that companies are buying a new product with another product's accounting tools. The cloud went through the same thing a decade ago, when the elasticity of spending on AWS or Azure forced the invention of FinOps as a corporate function in its own right. Here it has to be reinvented for agents, with metrics such as cost per task correctly completed, not per token consumed.

There are predictable winners and losers in this phase. The winners are the finance and procurement teams that start demanding cost attribution at the agent level, not aggregate LLM spending, and the vendors of cost-governance software that, like Pay-i, sell precisely that instrumentation. The losers, for now, are the organisations that treated agentic AI as just another software licence: sign the contract, switch on the agent and wait for the savings. Related coverage published the same day by the same outlet points to a sibling problem: many companies are getting better business insights with AI, but not the cost or time savings they expected, which confirms that the problem is not one of capability but of end-to-end governance.

In the short term, this is going to slow the pace of deployment at a number of organisations, and it is honest to acknowledge it: the AI budget conversation has moved up from the IT department to the boardroom, and it is not going back down. But in the medium term, McKinsey itself makes it clear in its framework of real value variables (correctness of the output, human oversight burden, compute cost in reasoning): once fine-grained instrumentation and per-agent cost attribution exist, the abundance AI promises will run its course, only with grown-up accounting. The token price will keep falling; what has been missing until now is the discipline to properly measure what that token, multiplied by a thousand self-correction cycles, really costs to produce.

🔗 Related on Zendoric

Sources & references