Agentic AI Costs: How to Measure and Control Them


The pattern repeats from one company to the next. A team puts an agent into production, the results hold up, and the consumption only surfaces on the following statement: tens of thousands of euros across six weeks. No incident, no technical failure. That is precisely what makes the problem difficult to address.
It is not an isolated case: in a September 2026 Futurum Research survey of enterprise technology buyers, 46.9% of enterprises reported running over budget on AI, against 5.6% that came in under plan. This article sets out how to measure agentic AI costs, and why the blocker is almost never technical.
It comes out of the latest episode of the AI Partners podcast, on our YouTube channel, where Thomas Spitz, CEO of AI Partners, hosted Théo Martin, Ecosystem Lead at Dust. Three questions were on the table: cost, governing a fleet of agents, and provider lock-in. Half the hour went on the first.
Because you are buying consumption, not access. A licence costs the same whether someone opens the tool once a month or lives in it, while an AI agent pursuing a goal bills for every step it takes to get there.
All the arithmetic that made software budgeting predictable goes with that model. A thousand people at 30 euros a month is a figure you set a year ahead and it holds. Now the amount depends on who adopts, how deeply, on which tasks, and how many attempts an agent runs through before it lands. A loop that retries twenty times costs twenty times, and nothing in the system objects.
On top of that sits an effect few companies planned for. The share of organisations running several agents together on one process doubled from 9% to 18% in a single quarter, per KPMG research published in June 2026. Each additional agent in the chain multiplies the calls behind one task. So the bill accelerates at precisely the point where adoption starts delivering.
The invoice arrives on time, it is accurate, and it is unusable. Providers bill by account and by user, never by team or by use case, so a finance director ends up with a correct number and no way to act on it.
The absence of a baseline runs deeper than most leaders assume. Still in the Futurum survey, 10% of enterprises have no formal AI budget to measure against and a further 5.6% do not know where their spending stands. One in six, then, cannot even say whether it is over.
Which produces a sequence seen across the market: the blindness comes before the overrun. You cannot correct a number you cannot decompose, and you cannot ask whether an agent met its objective when you were never able to see what that objective was.
By team. Per-person tracking is too granular to support a conversation about return, and a single company-wide line is too coarse to support a decision.
The model already exists inside the business, and few organisations think to reuse it: payroll. An HR function does not arbitrate every salary. It sets the frame, distributes the budget, and pushes the trade-offs down to the managers who know what their team is trying to achieve. A department head accountable for human resources is accountable for agentic ones too.
What changes in practice is who has the conversation. Instead of unpicking thousands of individual usage patterns, the finance director turns to the commercial director and asks them to account for their consumption. That request only means something if the measurement exists at that level. Without it, there is a total and nobody to discuss it with.
The market is already moving this way, sometimes without having decided to. Among over-budget enterprises, 23.1% shifted spending to a non-IT business-unit budget. The moment marketing or support pays part of the bill, they become the buyer, and they are who you will be negotiating with at renewal.
Then come the thresholds, and that is where most of the work sits. On a recent engagement we built them as a matrix of business function against AI maturity, because neither axis works alone. Setting a per-function threshold forces you to name the use cases first: you interview the teams, examine comparable practice in the market, and give more headroom to the three or four functions that genuinely need it.
The distribution tends to surprise leadership teams. The ten heaviest users absorbed close to 10% of the budget. Below them, a large soft middle sat well under its allowance, and roughly a fifth of licensed users barely consumed anything. An average flattens that reality, which makes it a poor basis for a forecast.
One precaution before you build yours: the heaviest users are often advanced users whose activity sits outside their actual remit. They are experimenting. Projecting their profile across the company overstates the need.
Take a 100,000 euro invoice. Knowing that 80,000 of it sits on four or five shared agents tied to approved use cases is a workable position: each one can be examined and decided on. Knowing it is scattered across dozens of individually built agents is not. Each would have to be understood separately, and no organisation has that capacity.
That is the whole difference between a shared agent and a personal prompt history. The first attaches to an accounting line and a named use case, the only form in which a cost becomes arguable. The second attaches to nothing.
Concentration also makes optimisation cheap, and that is where it pays for itself. Plenty of companies deployed their most capable model everywhere, for want of knowing what actually mattered. Once usage is visible, it emerges that most of it never needed that. Moving those cases to a lighter model is a single setting change when the use case lives on a shared agent, against fifty separate operations when it lives on fifty personal ones.
One limit, for completeness: an agent built to be general-purpose, used across very different tasks, stays hard to assess. High-value and low-value work run through the same object. Shared agents make the bill readable, they do not make it obvious.
They also surface what governance misses. 30% of executives report that poorly governed AI usage has led to cost overruns, in a July 2026 WitnessAI survey of 300 business leaders. The clearest case remains an employee who built a personal agent to run a private investment project on the company's budget. Nobody approved it. Nobody could see it. It surfaced the day the company started tracking spend by team.
Three levels, in order, because each is unreadable without the one before it.
A cap that halts work without warning is not a control, it is an outage. And distribution should be uneven by design: a function like customer service, delegating repetitive high-volume work, has an appetite that a department still exploring does not.
In practice most companies skip this step and fund the overrun. Among those affected, 47.6% sought supplemental funding and 43.3% absorbed it into the next planning cycle. Absorbing it quietly lifts next year's baseline, and the exception becomes the norm without anyone deciding it.
The starting point is unambiguous. In the WitnessAI study, only 9% of respondents said more than three-quarters of their AI initiatives delivered a measurable financial return. The report's own explanation is the useful part: the impact scatters across budgets, productivity reports and business units instead of landing in a single metric.
So measure where the objectives already exist, which is the team, with the instrument you already apply to payroll. Give a team an agentic wage bill alongside its human one and hold the sum against the targets it was already accountable for. You avoid the trap that catches most AI business cases, where the gain is measured on a task that was cheap to begin with.
This reading also leaves room for second-order effects. A sales team that beats its targets does not imply fewer salespeople: the rational decision is often to hire more and exploit the advantage.
And it protects against an error nobody talks about, because it never appears on an invoice. Anchoring the ceiling to last year's subscription price and demanding proof of return before every increase has the appearance of prudence. What it actually does is cap adoption at whatever the most cautious function will tolerate. If a competitor's commercial team becomes twice as effective, the budget you protected is smaller than the return you never collected. Both risks are real, but only one of them shows up in the accounts.
With attribution, not optimisation. Make spend visible at team level, attach it to named shared agents, then set the allowances, in that order. Optimisation techniques are abundant and well documented; if they go unused, it is because nobody can tell which workload to apply them to.
That is what we frame during the audit and roadmap, alongside the access scope, action log and escalation thresholds set out in our guide to agentic AI governance from pilot to production. A cost ceiling is a guardrail like any other, and a programme with that one alone is not governed, just metered. Because an allowance nobody knows how to use is spend without return, upskilling belongs inside the control system rather than beside it.
Thomas Spitz, CEO of AI Partners, hosts Théo Martin, Ecosystem Lead at Dust, on controlling agentic AI costs, governing a fleet of agents, and staying flexible when model providers change their pricing.
Should the AI budget sit with IT or with the business units?
With the business units, with IT or finance providing the measurement layer. A central team cannot assess the return on a use case it does not run, and a business unit cannot manage a budget it cannot see.
Does switching to cheaper models actually reduce the bill?
Yes, and it is usually the fastest lever, but only where the platform makes a model change a setting rather than a project. The effort to capture the saving depends entirely on how concentrated your agents are.
Are agentic AI costs ever predictable?
Within a range, not to the euro. Model a best, expected and worst case with retries and context growth included, then plan against the worst case.
Where do you start with no visibility at all?
By splitting the invoice by team. Until spend is attributed, no optimisation is manageable, because nobody can tell which workload to apply it to.
Want to know what your agents actually cost, and which team is paying for what? Book a 30-minute call and we'll review your current attribution, with no preparation needed on your side.