Why almost nobody can forecast their AI bill
Only 11% of companies can forecast AI costs within 10%. Why the number keeps moving, and what the teams that hit their budget do differently.
This spring, Mavvrik and Benchmarkit asked 396 companies how close their AI costs landed to their forecast.
11% got within 10%.
A year earlier it was 15%. Companies have more experience with AI now, more dashboards, and more people watching the bill, and they're getting worse at predicting it.
The misses aren't small. In the same survey, 62% said an unexpected AI cost changed a business decision in the past year. 40% had to take it to the board. A quarter delayed or cancelled an AI project outright.
Pigment's latest CFO survey found the same thing from the finance side. Of 2,000 finance leaders, 83% said consumption-based AI costs had pushed spending past what they expected. 13% stayed inside their AI budget.
Your AI bill works like a contractor billing by the hour, and nobody tells you the hours up front.
Software used to be a line item you set once a year. Seats times price. Most AI budgets are still written that way, and that's why they break.
"We'll just add a buffer"
This is the usual first fix:
We know the number will be off, so we'll forecast what we expect and add 20% on top.
That works if the misses are consistent. They aren't. In the Mavvrik survey, 89% of companies missed by more than 10%. In Pigment's, the typical overrun was 11 to 25%, more than a third were over by more than 25%, and 17% were over by more than half.
A buffer covers a predictable error. This error moves with the work. To budget for it, you need to know where it comes from.
Why the number won't sit still
There are 4 reasons, and they stack:
- The same task costs a different amount each time
- The cheaper model isn't always cheaper
- Model fees are the smaller part of the bill
- AI does work on its own, around the clock
The same task costs a different amount each time
A seat costs the same every month. A task doesn't.
Researchers from Stanford, Carnegie Mellon, UC Berkeley, and Microsoft Research ran 8 frontier reasoning models through 12 kinds of tasks and measured what each run actually cost. Running the same query on the same model again produced up to 9.7x more thinking tokens.
Same question, same model, same price per token. Close to a 10x spread in what you pay.
Multiply that across every person and every agent in a company and the monthly total gets very hard to call in advance. What you can do is watch it as it happens and know which team or tool is moving it.
The cheaper model isn't always cheaper
The obvious way to cut cost is to move work to a model with a lower price per token. The same study tested whether that holds.
Across 336 pairs of models, the one with the lower listed price cost more in total 106 times. That's 32%. Some reversals ran as high as 28x.
Their sharpest example: Gemini 3 Flash was listed 80% cheaper than GPT-5.4, and cost 38% more across the study's tasks. A cheaper model that has to think longer, retry, or take more steps to finish can burn through the savings and then some.
Price alone should not be used to infer which model is actually cheaper.
Lingjiao Chen, lead author
Price per token is a bad unit for a budget. Cost per finished task is the one that matters, and you only get it by measuring.
Model fees are the smaller part of the bill
Most AI forecasts start from token pricing. The FinOps Foundation estimates that token and API costs are only 10 to 25 percent of total AI spend once you count the infrastructure, software, and people around them.
Much of that is bought piece by piece. Vector databases, orchestration platforms, observability tools, and prompt management are often bought by different teams, billed separately from the model spend they support, and renewed on different cycles. On top of that, there are the AI tools people expense on a corporate card.
None of it shows up in the line that says "AI". The forecast was wrong before the first token was spent.
AI does work on its own, around the clock
A seat sits idle when its owner goes home. An agent doesn't.
Agents in CI, scheduled jobs, and coding tools left running decide how many steps to take, call tools, loop, retry when something fails, and run overnight. A small change to a prompt or a model upgrade can multiply what one run costs, and you find out on the invoice.
Mavvrik found that 15% of companies running agents can't attribute agent costs at any level. And 39% said their AI coding tools cost more than the license or usage they planned for.
What the teams that hit their number do differently
Nobody in these surveys forecasts AI perfectly. The ones that land close have stopped trying to predict one number in January and started controlling the number as the month goes.
They put a hard limit on every person and team
A budget in a spreadsheet tells you about an overrun after it happens. A limit set in the tool stops it. KPMG's latest AI Pulse found only 36% of organizations have direct token or usage controls in place.
The teams that do give each team a budget, give each person a monthly limit inside it, and keep a small pool for the people who need more. That turns an open-ended cost into a ceiling you chose.
This is what Tokenize budgets do: one budget per person that covers Claude, ChatGPT, and Cursor, applied as real spend limits in each tool.
They match the model to the work
The study's lesson is to measure before you switch. Teams that do this well keep the top model for work that needs it and start routine work on something cheaper, then check cost per task to see if the choice held up.
With model control, you pick which models each team can use and which one new Claude Code and Codex sessions start on.
They watch spend while the month is still going
The forecasts that hold are the ones updated every day. These teams see where spend is heading before month end and get an alert when one API key or one agent starts burning faster than usual, so a runaway job costs them a day and not a quarter.
Tokenize shows spend for every API key and agent, with a month-end forecast and alerts when usage jumps or a budget is running low.
They give it an owner
The FinOps Foundation found that organizations with clear ownership of AI costs were 3.7 times more likely to show AI's value to their CFO. Ownership is usually split. Engineering picks the tools, finance pays the bills, and nobody owns the total.
Someone has to own the whole number: every tool, every key, every renewal. Then it's possible to put what you spent next to what you got.
The numbers in one place
| Finding | Number | Source |
|---|---|---|
| Companies that forecast AI costs within 10% | 11% (15% in 2025) | Mavvrik and Benchmarkit, 396 companies |
| Unexpected AI cost changed a business decision | 62% | Mavvrik and Benchmarkit |
| Consumption-based AI costs went past expectations | 83% | Pigment, 2,000 finance leaders |
| Lower-priced model cost more in total | 32% of model pairs | Chen et al., 8 models, 12 tasks |
| Cost spread on repeated runs of the same query | up to 9.7x thinking tokens | Chen et al. |
| Token and API share of total AI spend | 10 to 25% | FinOps Foundation |
| Organizations with direct token or usage controls | 36% | KPMG AI Pulse, via PointFive |
Sources: Mavvrik and Benchmarkit, 2026 State of AI Cost Governance Report, surveyed April to May 2026, released July 29, 2026. Pigment, Q3 CFO Index, surveyed August 19 to September 21, 2026. Lingjiao Chen et al., The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More, arXiv 2603.23971, revised May 2026. FinOps Foundation, AI Spend Budget Busters, October 2026. PointFive, Why Companies Overspend on AI, citing KPMG AI Pulse Q2 2026.
Questions to ask before next quarter's budget
If all of your AI is a handful of seats on flat plans, a spreadsheet is fine and you can stop here. The more of these that apply, the less a spreadsheet will hold.
How much of your AI spend is usage-based?
Seats are easy to forecast. API keys, usage-based enterprise plans, and coding tools that charge for overage aren't. If usage-based spend is a growing share of the bill, your forecast error grows with it.
Do you know your cost per task, or only your cost per token?
If you're choosing models by listed price, a third of those choices may be costing you more. Pick a few common workflows and measure what one finished run costs on each model you allow.
Can anyone stop a team from going over, or only report it after?
A dashboard that shows last month's overrun is a history lesson. If no limit sits inside the tools themselves, the budget is a hope.
Do you know what your agents spent last night?
If agents run in CI or on a schedule, they spend while nobody is watching. You should be able to see spend by key, and hear about a spike the same day.
Does one person own the total?
Add up model fees, AI tools on corporate cards, supporting infrastructure, and coding tool overage. If nobody can give you that number today, start there.
Budgeting for work
AI costs won't become as steady as seats. The work changes, the models change, and agents keep running. The companies that hit their number accept that and change how they manage it: limits in every tool, costs measured per task, and a forecast that updates every day.
Once the cost is under control, the conversation with finance changes too. Instead of explaining last month's surprise, you can show what the money produced and argue for more of it.
If you'd like to see what that looks like for your teams, book a demo and we'll walk through it with your own tools and spend.