Somewhere in your enterprise right now, an AI agent is quietly working. It's retrieving data, building context, reasoning through a problem and, when it reaches a dead end, trying another approach. Each of these invisible steps consumes tokens, and each of these tokens has a cost. Nobody can predict exactly what the final cost will be.
New AI spend management and the value imperative
For decades, enterprise technology spending was predictable: software by the seat, infrastructure by the server and cloud services by the instance. Finance could forecast annual costs with confidence, while procurement negotiated contracts based on users, licenses, or infrastructure capacity.
Agentic AI changes that model entirely.
Unlike traditional software, AI agents pursue objectives rather than simply responding to prompts, collecting information, evaluating options and executing multi-step workflows with varying autonomy. Their behavior is dynamic, and so is what they consume.
This introduces an entirely new economic model for the enterprise. Instead of purchasing software licenses, organizations are increasingly purchasing cognitive capacity and paying for it by the token.
Yet the model invoice tells only part of the economic story. Independent analyses of production AI deployments suggest that more than 70% of total AI costs sit outside the model bill itself, in orchestration, retrieval, retries, observability and the surrounding application architecture. As organizations deploy increasingly sophisticated agentic systems, governing AI spend means managing not only token consumption, but also the entire operating model that enables AI to deliver business value.
When lower AI token costs drive higher spend
Independent analyses of enterprise application programming interface (API) usage show the blended price of frontier models falling roughly 67% year over year, yet estimates of average enterprise AI budgets show them rising from about $1.2 million to roughly $7 million over a similar period. In short, cheaper tokens are not producing cheaper bills. Economists call this pattern Jevons Paradox—the observation that when the cost of using a resource falls, consumption of that resource often rises faster than the savings, so total spending increases rather than decreases. Cheaper tokens do not shrink the AI bill, but change what becomes economically viable to attempt.
This feels like the next evolution of FinOps: cloud computing turned infrastructure into a consumption-based service and created a new discipline for governing that spend. Agentic AI is doing the same, shifting the question from how much AI costs to how organizations estimate, manage and govern an increasingly invisible form of consumption.
This is the AI token economy, or what we at CGI more broadly refer to as agentic AI cost management: the discipline of estimating, managing and controlling variable AI consumption before it controls you.
The organizations that succeed won't necessarily be those that spend the least on AI. They'll be the ones that understand where tokens create value, where they create waste and how to govern consumption as AI becomes embedded across every part of the business.
Why agentic AI costs are difficult to forecast
Agentic AI isn't only an IT cost. AI consumption is spreading across every business function and increasingly outside centralized control, which is exactly why agentic AI costs are so difficult to forecast.
One of the biggest misconceptions we encounter is that AI costs will become easier to predict as the technology matures. The opposite may be true.
While providers continue to reduce the price per token, the volume of tokens consumed by increasingly sophisticated AI systems continues to grow. Advanced models rely on larger context windows, more complex reasoning and autonomous decision-making.
Finance, procurement and IT leaders are experiencing this challenge in different ways: approving forecasts they can't pin down, negotiating usage-based contracts and choosing among models with no single right answer for every task.
The conversation therefore needs to shift. Instead of asking, “How much will AI cost?” organizations should be asking, “How do we govern AI consumption so that every token contributes measurable business value?”
A framework for governing the AI token economy
From our experience working with clients, organizations don't need perfect predictions. They need a practical framework for making better decisions. At CGI, we think about this through three connected disciplines: estimate, manage and control.
- Estimate means accepting that AI consumption is variable. Rather than relying on a single forecast, organizations should consider the use of scenario-based planning, which reflects different workloads, user behaviors and levels of adoption. In practice, that means building P50 and P90 consumption ranges for each workload archetype, including simple chat, retrieval-augmented generation and multi-agent orchestration, rather than relying on a single point estimate. We also recommend revisiting those ranges quarterly as usage patterns shift.
- Manage focuses on reducing unnecessary consumption without limiting innovation. It includes selecting the right model for the task, optimizing prompts, reusing cached responses where appropriate and designing workflows that minimize waste. Two of the most underused levers are prompt caching, which can cut the cost of repeated context by 50 to 90% depending on the provider, and batch or asynchronous processing, typically discounted around 50% for workloads that do not require an immediate response. Additionally, routing work across a tier of models by task complexity, rather than defaulting every request to the most capable model, is often worth more than any single vendor discount.
- Control is about governance. Organizations need visibility into where tokens are being consumed, which business units are generating costs and which AI use cases are delivering measurable value. Practically, this means per-team and per-agent budget ceilings, real-time anomaly detection for token spend and chargeback or showback so that the business unit generating the cost also sees it.
Implementing this requires an operating-model change touching architecture, finance, procurement, governance and culture all at once. Organizations that build this discipline now will be best positioned as adoption accelerates.
This is no longer a niche concern. In early 2026, the Linux Foundation announced plans for a Tokenomics Foundation to build open standards for measuring and governing AI token cost and efficiency, echoing what FinOps did for cloud spend a decade earlier.
The challenge now isn't simply adopting AI. It's learning how to govern the invisible economy that powers it.
Before your next AI budget conversation, it is worth asking three questions.
- Do you know which three business units or use cases generate the majority of your token spend?
- Do any of your production agents run without a budget ceiling or a named owner?
- And could your finance team forecast next quarter's AI spend within 20%, using actual usage data rather than vendor list prices?
If the honest answer to any of these is no, the token economy is already managing you, rather than the other way around.
CGI works with executives across industries to establish the methodologies and guardrails needed to manage AI costs effectively.
Sources: This blog draws on the following report and articles.
- NavyaAI Research, "AI Cost Report 2026: Token Prices & Rising AI Bills," last modified May 26, 2026, https://www.navyaai.com/reports/ai-cost-report-token-prices-vs-ai-bill.
- Jake Angelo, "Uber Burned Through Its Entire 2026 AI Budget in Four Months. Now Its COO Is Questioning Whether It's Worth It," Fortune, last modified May 26, 2026, https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/.
- Jake Angelo, "Microsoft Reports Are Exposing AI's Real Cost Problem: Using the tech is more expensive than paying human employees,” Fortune, last modified May 22, 2026, https://fortune.com/2026/05/22/microsoft-ai-cost-problem-tokens-agents/.

