Insights

Before you cap AI spend, attribute it

Most organizations see AI cost late, and agents make it harder to predict. Runtime controls are the right destination, and every one of them depends on a step that comes first: knowing whose spend it is.

Paul Turner

VP, Market Strategy, Tray.ai

Every AI governance conversation eventually reaches cost, and when it does, the first instinct is a ceiling. Set a limit, stop the overrun, move on. It’s the right instinct, aimed one step too early.

A ceiling has to belong to someone: a team, an app, a person who answers for it. Most organizations can’t yet say who that someone is for most of their AI spend, and until they can, a limit has nothing to attach to.

So this is about the order of operations. The controls worth building are well understood. The step that makes them possible usually gets skipped.

Most organizations see AI cost late

Gartner published research in September 2026 arguing that cost governance has to become part of AI governance, and the survey behind it isn’t reassuring.1 A small group of organizations only learn what their AI costs when the bill arrives. Most of the rest know their costs in some form but have nothing in place to prevent an overrun. The share that can stop one before it happens is small.

The same research found that the business is rarely the party accountable for the cost and value of AI deployments. Think about what that means. IT holds the bill for spend it didn’t decide on, in apps and agents it didn’t build. When the invoice is higher than expected, the conversation that follows has no natural target, so it lands on whoever pays.

Why agents make it worse

Traditional software had a cost profile you could forecast: licences, seats, a hosting bill that moved slowly. AI spend behaves differently, and agents push it further. Gartner lists the reasons, and each one is worth spelling out because each one breaks a different forecasting habit.1

  • Token pricing. You pay for what goes in and what comes out, per request, so cost tracks behaviour.
  • Autonomous workflows. An agent decides how many steps a task takes. Two runs of the same task can cost very different amounts.
  • Retry loops. An agent that fails and tries again keeps paying for each attempt. An unwatched loop can run all night.
  • Expanding context windows. Each step in a long task can carry everything before it, so later steps cost more than earlier ones.
  • Model routing. The model that handles a request may change, and with it the price, without anyone choosing that on purpose.
  • Tool overhead. Every call an agent makes to a tool adds its own tokens on top of the work itself.

The research also describes an organization that spent its whole annual AI budget in a small part of the year. All it takes is a lot of small, reasonable decisions made by software, and a meter left unread until the end of the month.

A dashboard after the fact is a lagging indicator

The common first response is a dashboard: pull the invoices, break them down by provider and model, review them monthly. That’s better than nothing, and Gartner describes it accurately: a report after the fact is a lagging indicator.1 By the time the chart shows a spike, the money is spent.

Worse, a provider invoice broken down by model tells you which model was expensive. What you can act on is which app, which team or which decision drove the cost, and the invoice is silent on that.

The fix, in Gartner’s framing, belongs in the runtime path: controls that act while the work is running, enforced at the gateway, the agent runtime, the tool layer or the platform where the AI app is built and run.

The destination: controls at runtime

Gartner’s minimum policy for AI cost describes four controls, and they are a sensible picture of where a mature program ends up.1

  1. 1

    Tag every run

    Record which agent or app used which tokens, for whom, and where, as the work happens.

  2. 2

    Budgets by user and team

    Give each user, project or group a budget, with a warning before the ceiling and a stop at it.

  3. 3

    Route by budget

    Send non-critical work to cheaper models so the expensive ones are kept for the work that needs them.

  4. 4

    Limit a single run

    Make sure no one runaway agent run can consume a whole budget on its own.

Summarized from Gartner’s minimum viable AI FinOps policy. The order is the argument of this post: the first step is what the other three are built on.

Gartner adds a caveat that’s easy to miss: if the cost data isn’t close to real time, a ceiling can’t be enforced. A budget checked against yesterday’s numbers is a report with a threshold on it.

Attribution comes first

Look at those four controls again and ask what each one needs before it can work.

A budget by team needs every unit of spend assigned to a team. A budget by user needs it assigned to a user. Routing non-critical work to cheaper models needs someone to have decided which work is non-critical, and that decision belongs to whoever owns the app doing the work. A limit on a single run needs the run to be identifiable as a run, belonging to something, so the limit can be set at the right level for that thing.

Every control on the list assumes the first step is already done. That’s the core of the argument: you can’t set a budget for spend you can’t assign to an app, a team and an owner. A ceiling on unattributed spend is a ceiling on the whole organization, and the only lever it gives you is to stop everything at once. That’s a lever you’ll never pull, so in practice the limit never gets set.

Attribution also changes the conversation before any control exists. When spend arrives already assigned to a named app with a named owner, the monthly review turns into a short list of questions for specific people. Is this app worth what it costs? Did that spike come from a change you made? Should this workload run on a cheaper model? Those are questions an owner can answer in a day. An unattributed invoice can’t be answered by anyone.

There is a second-order benefit too. The research found the business is rarely accountable for AI cost and value. Attribution is how that changes. Once spend is assigned to an owner in a business team, the value side of the conversation has somewhere to go as well: the person who answers for the cost is the person who can show what the app returns.

Proportionality

Gartner’s other point on cost is proportionality: give teams delegated authority to spend inside approved budgets, and add oversight only as spend crosses thresholds, the way a procurement policy works.1 Small experiments should be cheap to start and cheap to stop. Large, growing workloads should get more attention as they grow.

That model also depends on attribution. Delegated authority has to be delegated to someone, and a threshold has to be measured against something. Without an owner and an app on each line of spend, proportionality collapses into one of two bad options: review everything, which slows the experiments that should be cheap, or review nothing, which is how a year’s budget runs out early.

So the order is: attribute first, then delegate, then set thresholds, then enforce. Skipping to the last step is what most programs try, and it’s why so many of them end up with a policy that names a ceiling and a dashboard that shows when it was crossed.

Where Tray Helix fits

Tray Helix is the governed runtime for apps built in Claude Code, Codex and Cursor. Every app deployed through it has a named owner and a registry entry, and its AI and compute spend is visible per app, so spend arrives already assigned to an app, a team and the person who answers for it. That’s the attribution step this post argues for. Billing is per workspace plus usage, so how you’re charged stays simple while the view of where the spend goes stays specific.

Footnotes

  1. Gartner, “Cost Governance Is a Non-Negotiable Addition to AI Governance,” G00861207, Kornutick et al., 24 September 2026, drawing on the 2026 Gartner AI Governance Survey. GARTNER is a registered trademark and service mark of Gartner, Inc. and/or its affiliates in the U.S. and internationally and is used herein with permission. All rights reserved. Gartner does not endorse any vendor, product or service depicted in its research publications and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner research publications consist of the opinions of Gartner’s research organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this research, including any warranties of merchantability or fitness for a particular purpose. Back Back Back Back Back

  • governance
  • ai cost
  • finops
  • agents

Helix is the governed runtime for AI-built apps

Deploy what your teams build, put SSO in front of it, connect it with managed credentials, and give every app a named owner.

Want to talk to someone first? Contact us