Many organisations are moving quickly to adopt AI, and there is a lesson from cloud that should not be missed: consumption models need governance early, not after usage has scaled. That matters for both IT and finance. AI spend is already showing up across models, copilots, agents, developer tools, and the cloud services around them.
Without the right operating model, it can quickly become another growing cost line with limited visibility and unclear accountability. Many of the decisions shaping those costs are also being made early through architecture, platform, data, and deployment choices. Across AI services more broadly, usage is often metered in tokens rather than users or servers. In practice, tokenomics is about understanding how tokens are used, how they are measured, and what that consumption costs at scale. But the token bill is only the most visible part of the economics.
The clearer unit of cost is the AI transaction. That transaction is not just model inference. It also includes retrieval, orchestration, tool use, integration, observability, and the platform and security services needed to run AI safely in production. Tokens may be the easiest part to meter, but they are only one part of the true cost profile. The real commercial question is therefore not just how many tokens are being used, but what it costs to complete a meaningful AI transaction and whether that transaction creates enough value.
The old software controls are not enough. AI consumption needs the same discipline organisations now apply to cloud, but with a clearer link between operating model, cost controls, and business value.
Why AI consumption economics matters now
For CFOs, AI consumption economics is more than simply a technical consideration. It is a control issue. Cloud already showed what happens when consumption runs ahead of governance. AI can move even faster because the gap between pilot and production is often blurred and demand can rise quickly once teams start using it in anger. Across AI platforms, cost is shaped by architectural decisions made well before finance sees the consumption bill. Model selection, context strategy, RAG design, routing, deployment type, orchestration, integration, observability, and agent execution design can all determine the eventual run cost profile. The bill is not just a reflection of usage; it reflects a series of design choices. That is why finance needs early visibility. The real question is not how much AI costs in the abstract, but which architectural decisions are driving the full transaction cost, who is making them, and whether the return stacks up.
What makes this more urgent is how quickly AI costs can build. In plenty of organisations, the early spend barely registers, so it does not get much attention. Then usage becomes real, more teams start leaning on it, and suddenly it is a meaningful monthly cost line. That tends to happen fastest in developer and productivity use cases, where demand can build quickly and not always in a predictable way. Finance leaders must be mindful that AI spend can move from manageable to material more quickly than many expect.
Agentic AI makes this even more important. A simple prompt-and-response interaction has a relatively straightforward cost profile. An agentic workflow can be very different. It may break a task into multiple steps, retrieve information, call tools, check outputs, retry when something fails, and keep context across a longer interaction. Agent costs can therefore be highly variable. One transaction may complete cleanly, while another may involve retries, recursive planning, additional tool calls, or validation loops that make it materially more expensive. Each of those steps can add consumption.
For production forecasting, averages can be misleading. Leaders need to understand the normal cost of a transaction, but also the higher-cost cases that occur often enough to matter. A practical way to frame this is to look at the typical transaction cost alongside the cost of the more demanding transactions, rather than relying only on average consumption. That gives finance and IT a more realistic view of how agentic AI will behave once it is embedded in core business processes.
Why AI is reshaping the finance conversation
AI combines strategic investment decisions with a live consumption model. Decisions made early on data, architecture, security, and operating model have a direct bearing on how costs develop over time and where they land in the P&L. That is why both IT and finance need to be involved earlier. This is not just about signing off spend, but making sure the organisation has an operating model that can support AI at scale without losing sight of cost, accountability, or value.
What matters is that finance understands enough about how AI deployment works before the organisation commits too far, too fast, and that IT can explain the commercial implications clearly enough for informed decisions to be made. The organisation must understand what it is taking on once adoption moves beyond experimentation. That means being clear on the deployment model, how consumption is likely to build, how costs will flow through the P&L, and where accountability will sit over time. Without that joined-up view, it becomes much easier to overcommit early and much harder to course-correct later.
What IT and finance need to align on
IT must explain AI deployment choices in a way finance can work with, and finance must understand the commercial implications well enough to help shape the right guardrails. The first issue is not which model a team prefers, but the economic profile the use case actually needs. That means being clear on whether the workload justifies premium model spend, what likely demand looks like before anything scales, and how architectural decisions such as model selection, context strategy, RAG design, routing, deployment type, and agent execution design will shape cost over time. It also means deciding whether the application needs one model for every task, or whether model routing can send simpler, lower-risk work to lower-cost models and escalate only when confidence, complexity, sensitivity, or business impact requires it. If that is not articulated clearly early on, organisations can end up with avoidable blockers later when spend, accountability, and value all come under pressure.
AI deployment proposals need a credible business case before funding is approved. That should include a clear consumption envelope: the expected cost per transaction, anticipated volumes, likely peak or worst-case cost scenarios, and the controls that will prevent abnormal or runaway consumption. It should also set out the likely consumption pattern, the broad cost range, where accountability will sit, and how spending will be tracked as usage grows.
More importantly, it should show how the operating model will work in practice: who owns the deployment choices, who owns the run cost, how the P&L impact will be understood, and how finance and IT will work together to keep the controls proportionate as adoption grows. The role of finance is to help put the right guardrails around the approach, while IT guides how the deployment should be shaped and scaled.
Seven factors that drive AI transaction costs
-
Model choice and routing: different models carry different costs, so not every developer or agent use case should default to the most advanced option. Architectures can route simpler, lower-risk tasks to lower-cost models and escalate only where confidence, complexity, sensitivity, or business impact justifies it.
-
Prompt and context size: longer prompts, larger instructions, and excessive conversation history all increase token consumption, so context should be purposeful rather than simply expansive.
-
Retrieval and context quality: poor retrieval design can increase the cost of search, re-ranking, and orchestration, then inflate the model context with unnecessary content. Good retrieval design improves both answer quality and cost control by making sure the model receives the information it actually needs.
-
Response length: longer outputs drive more token usage, which can increase costs quickly at scale.
-
Agentic workflow design: multi-step reasoning, retrieval, tool calls, retries, recursive planning, and validation loops can make one transaction materially more expensive than another, so forecasting should look beyond simple averages.
-
Transaction architecture: the full cost of an AI interaction is shaped by architectural choices including model selection, context strategy, RAG design, routing, deployment type, agent execution design, integration, observability, security, and the platform services needed to support it.
-
Workload shape: interactive requests, agent flows, batch processing, and sustained high-volume traffic all create different consumption patterns and should be managed differently.
This is where the comparison with cloud consumption still helps. Finance teams learned to manage cloud through tagging, chargeback, reserved capacity, budget alerts, and better design choices. AI needs the same discipline, just applied to a different unit of consumption. A badly written prompt is not that different from an oversized virtual machine. However, reducing prompt or model costs should not be pursued in isolation. A cheaper AI transaction that produces lower-quality answers, increases retries, or drives avoidable human escalation can ultimately cost the organisation more. The objective is therefore not the lowest-cost transaction, but the optimal balance of cost, quality, latency, risk, and business value across the end-to-end process.
A practical control model for AI unit economics
A practical AI unit economics model starts with visibility, but it should not stop at reporting. Finance needs to see consumption by service, workload, and business dimension, not just another cloud line item that keeps growing. That means understanding how tokenisation works in the context of the services being used, tracking input tokens such as prompts, instructions, and context, and tracking output tokens such as generated responses or agent actions. From there, organisations can calculate token and platform cost, but they should treat those as supporting signals rather than the primary metric. The more useful measure is cost per successful business outcome or completed transaction, with token and platform consumption traceable underneath it. For example, cost per service request resolved successfully rather than tokens, credits, or model calls consumed.
That distinction matters. The control model should capture the total run cost around the AI service, but it should also connect that cost to the outcome the service is meant to deliver. That includes inference, orchestration, data retrieval, tool use, observability, security controls, integration, and the wider cloud services used to support production workloads. Once that visibility exists, AI usage should be tied back to departments, products, teams, or use cases so the cost sits where the demand comes from.
For agentic workloads, FinOps controls also need to become part of the execution architecture, not just retrospective reporting. That means setting practical runtime controls such as maximum model calls, tool calls, retries, agent steps, token or reasoning budgets, time-outs, and potentially a maximum transaction cost threshold. These controls help prevent abnormal consumption before it becomes a finance problem, while still allowing teams to design workflows that produce good-quality outcomes.
For agentic workloads, the control model should also show the spread of transaction costs, not just the average. In practice, that means understanding the typical cost of a transaction and the cost of the higher-consumption cases that occur often enough to affect budgets. This gives leaders a better basis for production forecasting because it captures the variability created by retries, tool calls, longer workflows, and validation loops.
Model routing should also be treated as a practical cost-control lever. The aim is not to use the cheapest model in every case, but to match the model to the task and escalate only when the level of complexity, confidence, risk, or value justifies the additional cost. That gives IT a way to control run cost through design, while giving finance a clearer explanation of why different transactions have different cost profiles.
Then the operating model becomes critical. IT and finance need to be clear on who owns which decisions, how costs will be allocated, where thresholds sit, and how optimisation will be handled as usage grows. That includes better prompt design, tighter model selection, model routing, more use of batch, more use of prompt caching, runtime limits for agentic workloads, and reserved or provisioned capacity where demand is predictable. This puts the right guardrails around AI so scale does not come at the expense of control.
Microsoft examples of cost control in practice
Microsoft provides several capabilities that can support this approach across Azure and Microsoft 365. Azure Cost Management can track the Azure services behind Microsoft Foundry workloads, which means finance can apply budgets, alerts, and chargeback instead of relying on hindsight. Cost management and application observability serve different purposes. Azure Cost Management shows where cloud spend sits across services, resources, and workloads. Application-level observability explains why a particular agent or business transaction incurred that spend, including the model calls, retrieval steps, tool use, retries, latency, and outcomes involved. The value comes from connecting these perspectives. Cloud cost data shows where money is being spent, while application traces explain what happened inside the transaction and whether the spend contributed to a successful business outcome.
This is where AI FinOps becomes particularly useful, linking platform consumption to business value through the full transaction lifecycle. Within Microsoft Foundry, teams can choose between standard deployments, batch, and provisioned throughput based on the shape of the workload rather than defaulting to the most flexible option. They can also design applications to route different classes of task to different models, so premium models are reserved for the work that genuinely needs them.
Prompt caching can reduce the cost of repeated long prompts. Batch can bring down the cost of high-volume workloads that do not need an instant response. Provisioned throughput can create more stability where demand is predictable, with Azure Reservations helping on sustained production workloads. At the same time, Microsoft 365 Copilot pay-as-you-go and Copilot Chat agents create another metered cost line that finance needs to watch just as closely. If controls are split across Azure and Microsoft 365, accountability can end up split as well.
What leaders should do next
What IT and finance leaders should do next is build enough shared understanding to stay in control as AI adoption grows. That means five things:
-
Define AI unit economics: expected cost per successful business outcome or completed transaction, with token usage, credits, actions, retrieval, model execution, tool calls, platform services, and infrastructure costs traceable underneath.
-
Explain the architectural decisions: how model selection, model routing, context strategy, RAG design, deployment type, agent execution design, retrieval, integration, and observability shape cost and value.
-
Clarify accountability: who owns the AI services, platforms, business demand, and P&L impact so spend does not disappear between budgets.
-
Set the operating model: where decisions sit, where approvals are needed, and how cost controls will work as usage grows.
-
Review spend against outcomes: focus on business value, not usage metrics alone, so AI can scale without losing grip on the economics.
The next phase of AI adoption will be shaped by organisations that put the right operating model, cost controls, and commercial discipline around consumption early enough. Tokenomics still helps explain an important part of the underlying mechanics, but on its own it is becoming too narrow for executive decision-making. The more useful shift is towards AI unit economics: understanding the cost per successful business outcome, with tokens, model calls, retrieval, tools, platform services, and infrastructure traceable underneath. Microsoft has put enough tooling in place across Microsoft Foundry, Azure Cost Management, application observability, and Microsoft 365 Copilot pay-as-you-go to support that model, but the underlying point is broader than any one platform. With the right approach, IT and finance can stay in control of AI adoption instead of reacting to the cost only after it shows up.