Token Bleed: How to Control Runaway AI Token Costs with a 5-Layer Control Stack

“Agentic tasks are uniquely expensive, consuming 1000x more tokens than code reasoning and code chat.”
Stanford Digital Economy Lab, 2026

That number captures a problem enterprises are only beginning to see clearly. AI models may be getting cheaper per token, but the systems being built on top of them are consuming vastly more tokens. Gartner expects AI inference costs per agentic workflow to increase more than fivefold through 2028, even as model economics improve.

The problem is that AI spending can grow continuously, at machine speed, with little connection between what is being consumed and the business value being created. We call this token bleed: uncontrolled AI spending caused by employees treating token consumption as productivity, stolen API credentials, agents trapped in execution loops, and workloads repeatedly sending far more context than a task requires.

Agentic AI makes token bleed an operational risk rather than just another cloud bill. A chatbot may make one model call for a request. An agent can reason, invoke tools, inspect results and call the model repeatedly before completing the same task. Stanford found that even repeated runs of the same agentic task can vary by as much as 30× in token consumption, with higher token use not necessarily producing better results.

This is where AI cost governance becomes necessary. It is the combination of technical and financial controls that makes AI spending attributable, bounded and measurable against business outcomes. A practical control model has five layers:

  1. Identity: scoped workload identities and protected API keys
  2. Budgets: limits that stop requests before runaway spend occurs
  3. Efficiency: model routing, prompt caching and context management
  4. Agent guardrails: step, token and repetition limits that stop runaway execution
  5. Value measurement: cost per successful outcome instead of raw token consumption

Key Takeaways

  • Cheaper tokens don't mean smaller bills. Gartner predicts AI inference costs per agentic workflow will rise more than fivefold through 2028.
  • Agents are expensive by design. A Stanford study found agentic coding tasks use about 1,000 times more tokens than code chat, and identical runs can differ by up to 30 times.
  • Alerts are not controls. An alert tells you after the money is spent. A pre-execution budget check stops the next call.
  • Measure outcomes, not tokens. Token leaderboards reward waste, as one large tech company found out in 2026.

What Is Token Bleed?

Token bleed is AI spending that grows faster than the value it produces, because nobody sees or limits token use in real time. It shows up as a monthly bill that surprises finance, security and engineering at the same time.

A token is the unit that AI models use to read and write text. Providers charge per token, and usually charge more for output tokens than for input tokens. Three things make token spend hard to control:

  • It doesn't feel like money. A developer sees "tokens," not dollars, so there is little friction before each request.
  • It runs at machine speed. An agent can make thousands of calls in an hour with no person in the loop.
  • An API key works like a credit card with no limit. Anyone who holds the key can spend.

The security community already names this risk. The OWASP Top 10 for LLM Applications lists LLM10:2025 Unbounded Consumption. It covers uncontrolled inference that leads to "denial of service (DoS), economic losses, model theft, and service degradation." Its financial form is called denial of wallet: attackers, or runaway workloads, exploit pay-per-use pricing until the bill becomes unsustainable.


Why Do AI Bills Rise When Token Prices Fall?

AI bills rise because workloads consume tokens faster than prices fall. Agentic AI makes many model calls per task, and each call resends a growing context. The drop in unit price is small compared with the rise in volume.

Gartner calls this the inference paradox. In August 2026, Gartner predicted that AI inference costs per agentic workflow will increase more than fivefold through 2028. It also found that routing a task to agentic reasoning models raises inference cost by at least five times compared with a basic chatbot interaction.

"Product leaders cannot rely on more efficient token economics to rationalize AI costs."

The pattern also shows up in software development. In June 2026, Gartner predicted that AI coding costs will be higher than the average developer's salary by 2028, as token consumption grows and pricing moves from seats to consumption.

"Token discipline will not emerge through developer choice alone, as developers tend to optimize for speed and convenience over cost efficiency."


The Four Sources of Token Bleed

Token bleed has four common sources: incentives that reward consumption, stolen credentials, agent loops and context bloat. Each one needs a different control.

SourceWhat happensReal examplePrimary control
TokenmaxxingEmployees inflate token use to look productiveA Meta employee's internal "Claudeonomics" leaderboard tracked more than 60 trillion tokens in 30 days before it was taken down in April 2026Value metrics (Layer 5)
LLMjackingAttackers steal cloud or API credentials and run models on your accountSysdig reported attacks costing over $100,000 per day with a frontier model in 2024Scoped keys and anomaly shutdown (Layers 1 and 2)
Agent loopsAgents repeat tool calls or hand tasks back and forth without endA widely cited agent-failure case describes a four-agent LangChain workflow that remained in a feedback loop for 11 days and accumulated about $47,000 in API charges.Loop guardrails (Layer 4)
Context bloatEvery call resends long histories, documents and tool outputStanford researchers found input tokens, not output, drive agent costCaching and context limits (Layer 3)

Post-mortem: when a token leaderboard backfired

Tokenmaxxing is the purest example of Goodhart's Law: when a measure becomes a target, it stops being a good measure.

According to reporting by The Pragmatic Engineer, internal leaderboards and spending targets at several large tech companies encouraged engineers to use more tokens, not better ones. One engineer described asking AI "to prototype features I have no intention of working on."

The same reporting describes a better pattern. Shopify renamed its leaderboard a "usage dashboard" and added circuit breakers to catch runaway agents. It changed the metric and added a hard stop, which is exactly what Layers 4 and 5 below describe.


Worked Example: What an AI Agent Really Costs per Task

An AI agent can cost more than 80 times as much as a chatbot for the same customer request, because it resends its growing context on every turn. Caching and a turn limit can cut that cost by about 85%.

The model below is illustrative. The prices are assumptions, so replace them with your provider's current rates.

Assumptions:

  • A support workload of 50,000 tickets per month
  • Input tokens at $3 per million, output tokens at $15 per million
  • Cached input tokens billed at 10% of the input price (cache-write premiums not included)
  • Chatbot: one call per ticket, 2,500 input tokens and 500 output tokens
  • Agent: starts with 4,000 tokens of context and adds 1,500 tokens per turn. It writes 500 output tokens per turn and runs 20 turns per ticket

Because the agent resends the whole conversation on every turn, its input grows with each step:

Total input tokens = (turns × starting context) + (added tokens per turn × turns × (turns − 1) ÷ 2)

For 20 turns: (20 × 4,000) + (1,500 × 20 × 19 ÷ 2) = 365,000 input tokens per ticket

ScenarioInput tokens per ticketCost per ticketCost per month
Chatbot2,500$0.015$750
Agent, 20 turns, no caching365,000$1.245$62,250
Agent, 20 turns, with prompt caching365,000 (332,500 cached)$0.347$17,363
Agent, capped at 12 turns, with prompt caching147,000 (126,500 cached)$0.189$9,473

Three lessons come out of the table:

  • The agent costs 83 times as much as the chatbot before any controls are applied.
  • Caching cuts the agent's cost by 72%, because most of each call is repeated context.
  • A turn cap combined with caching cuts it by 85%. First, confirm that 12 turns still solve the task.

The 5-Layer Token Control Stack

The 5-Layer Token Control Stack stops token bleed by placing controls at each point where spend can escape:

  • Identity limits who can spend.
  • Budgets stop spend before it happens.
  • Efficiency lowers the cost of each task.
  • Agent guardrails stop loops.
  • Value metrics keep incentives honest.
LayerControlStops
1. Identity and keysOne scoped key per workload, stored in a secrets manager and rotatedLLMjacking and shared-key sprawl
2. Budgets and rate limitsPre-execution budget checks, quotas and automatic shutdown on anomaliesSpikes from any source
3. EfficiencyModel routing, prompt caching and context limitsContext bloat and overpowered models
4. Agent guardrailsStep limits, token limits and repetition detection per taskAgent loops
5. Value metrics and governanceCost per successful outcome, clear ownership, weekly reviewsTokenmaxxing and silent drift

Reference architecture: route every call through an AI gateway

The simplest way to apply Layers 1 to 3 is to send every model call through one gateway, instead of letting each application call providers directly.

 Apps, agents, IDE tools
          │  (workload identity, not shared keys)
          ▼
 ┌───────────────────────── AI Gateway ─────────────────────────┐
 │ 1. Authenticate workload   → reject unknown or revoked keys  │
 │ 2. Check budget + rate     → block BEFORE the call if over   │
 │ 3. Route model             → small model first, escalate     │
 │ 4. Check prompt cache      → reuse repeated context          │
 │ 5. Log tokens + cost       → per user, team, workload, task  │
 └───────────────────────────────────────────────────────────────┘
          │                                   │
          ▼                                   ▼
 Model providers                    Cost + anomaly dashboard
                                    (finance, security, engineering)

Layer 1: Scope identity and protect keys

Treat every AI API key like a production credential. LLMjacking starts with a leaked key, and a shared key makes it impossible to tell who spent what.

  • Issue one key per workload, never per team or company.
  • Store keys in a secrets manager, never in code or notebooks.
  • Run secret scanning on repositories, and rotate keys on a schedule.
  • Set quotas at the cloud provider level too, as a second line of defense.
  • Scope the credentials used by agent tools, such as MCP servers, to the minimum permissions. This limits both the security impact and the spend if an agent misbehaves.

Layer 2: Enforce budgets before execution

A budget alert tells you about spend after it has happened. A budget check blocks the next request. The $47,000 loop above had visibility but no enforcement, so it ran until someone read the bill.

  • Set budgets per workload, per user and per task.
  • Check the budget before each call, in the gateway.
  • Shut down automatically when spend exceeds a set multiple of the normal hourly rate.
  • Aim to stop a spike in minutes, not at the next billing review.

Layer 3: Make each task cheaper

Efficiency controls lower the cost of work you want to keep:

  • Model routing sends simple requests to smaller, cheaper models and escalates only hard ones. In LMSYS's RouteLLM research, routing cut costs by over 85% on the MT Bench benchmark while keeping 95% of GPT-4's performance.
  • Prompt caching reuses repeated context, such as system prompts, documents and conversation history, at a discounted rate.
  • Context limits summarize or trim old turns instead of resending everything.

Gartner recommends the same approach: implement "inference-tiering, routing and orchestration" to match complex tasks with cost-efficient models.

Layer 4: Put guardrails on every agent

Agents need hard limits, because their cost varies widely even on the same task. This small guard stops a loop on step count, token budget or repeated actions:

import hashlib

class AgentBudgetExceeded(Exception):
    pass

class AgentGuard:
    def __init__(self, max_steps=12, max_tokens=150_000, max_repeats=3):
        self.max_steps = max_steps
        self.max_tokens = max_tokens
        self.max_repeats = max_repeats
        self.steps = 0
        self.tokens = 0
        self.seen = {}

    def check_before_call(self, estimated_tokens: int, action: str) -> None:
        """Call BEFORE every model or tool call. Raises to stop the agent."""
        self.steps += 1
        if self.steps > self.max_steps:
            raise AgentBudgetExceeded(f"step limit {self.max_steps} reached")
        if self.tokens + estimated_tokens > self.max_tokens:
            raise AgentBudgetExceeded(f"token budget {self.max_tokens} reached")

        key = hashlib.sha256(action.encode()).hexdigest()
        self.seen[key] = self.seen.get(key, 0) + 1
        if self.seen[key] > self.max_repeats:
            raise AgentBudgetExceeded("same action repeated: possible loop")

    def record_usage(self, actual_tokens: int) -> None:
        self.tokens += actual_tokens

When the guard stops an agent, it should do three things:

  • save the task state
  • return a clear message to the user
  • log the event for review

A stopped agent is a cheap failure. An unstopped one is an expensive one.

Layer 5: Measure value and assign ownership

Layer 5 keeps the other four honest. It replaces "tokens used" with "cost per successful outcome," and it gives each type of risk a named owner. The next two sections describe both.


Which Metrics Show Whether AI Spend Is Working?

The right metric is cost per successful outcome, not tokens consumed. Tokens measure effort. Outcomes measure value.

MetricWhat it tells youWarning sign
Cost per successful taskThe true unit cost of AI workRising while success rate stays flat
Tokens per task (median and 95th percentile)How variable the workload isA 95th percentile far above the median, which suggests loops or bloat
Cache hit rateHow much repeated context is reusedLow on workloads with long, stable prompts
Share of requests on smaller modelsWhether routing worksEverything goes to the most expensive model
Agent stops by guardrailHow often loops are caughtZero, which may mean no guardrails are active
Time to stop a spend spikeSpeed of your response to an anomalyMeasured in days, not minutes

Report these weekly to engineering, finance and security, per workload and per team. A quarterly review of the bill is too slow for spend that runs at machine speed.


Who Owns Token Risk?

Token risk needs shared ownership. Finance owns the budget, security owns the credentials and engineering owns efficiency. Without a named owner for each, the risk falls between teams.

ResponsibilityFinanceSecurityEngineering
Set budgets per team and workloadOwnerConsultedConsulted
Protect and rotate API keysInformedOwnerResponsible
Detect and stop anomaliesInformedOwnerResponsible
Routing, caching and context designInformedConsultedOwner
Define value metricsOwnerInformedResponsible
Weekly cost reviewOwnerConsultedResponsible

The FinOps discipline is moving quickly to fill this gap. In the FinOps Foundation's State of FinOps 2026 survey of 1,192 respondents, 58% named AI cost management as the skill their teams need most. In 2025, 63% of respondents were managing AI spend. By 2026, nearly all of them were.


The Token Governance Maturity Model

Most enterprises move through four stages of token governance. The goal is not the lowest spend. It is spend that is visible, enforced and tied to value.

StageWhat it looks likeMain risk
1. UnmeteredTeams experiment freely, with shared keys and no trackingSurprise bills and stolen keys
2. CappedBlanket spending limits after the first surprise billLimits block valuable work along with waste
3. MeasuredUsage is tracked per team, and training covers how to use AIPeople learn to use more tokens, not fewer
4. Value-governedThe 5-layer stack is in place, metrics track outcomes, and ownership is sharedKeeping controls current as models and prices change

Most companies stall at stage 3. They train people to use AI, but not to use it efficiently. Token literacy fills that gap. It means knowing when a smaller model is enough, how caching works, and when an agent is the wrong tool.


What Comes Next for AI Cost Governance

Tokens are becoming a managed unit of spend, like cloud compute before them. In June 2026, the Linux Foundation announced its intent to create the Tokenomics Foundation in partnership with the FinOps Foundation. It formally launched in August 2026 to develop vendor-neutral frameworks, specifications and practices for measuring AI cost, value and ROI.

"Tokens have become the new unit of technology spend." Jim Zemlin, CEO, Linux Foundation

Problem: AI spend grows at machine speed, through incentives, stolen keys, agent loops and bloated context, while controls lag behind.

Solution: A 5-layer control stack that scopes identity, enforces budgets before each call, lowers the cost per task, stops agent loops and measures value per dollar.

Vision: Token budgets become as standard as cloud budgets. Every agent ships with limits, every workload reports its cost per outcome, and shared standards let finance compare AI spend across providers the way it compares cloud spend today.

If your AI spend is growing faster than your ability to explain it, talk to Tarento's Generative and Agentic AI team.


Frequently Asked Questions

What is token bleed in AI? Token bleed is AI spending that grows faster than the value it delivers, because no one sees or limits token use in real time. It usually comes from agent loops, stolen API keys, oversized prompts or incentives that reward heavy usage instead of results.

Why do AI bills go up when token prices go down? Workloads grow faster than prices fall. AI agents make many calls per task and resend their growing context each time. Gartner predicts inference costs per agentic workflow will rise more than fivefold through 2028, even as unit prices keep dropping.

What is tokenmaxxing? Tokenmaxxing is when employees deliberately use more AI tokens to look productive, often because of leaderboards or usage targets. It inflates cost without improving output. The fix is to measure cost per successful outcome and remove rankings based on raw token consumption.

What is LLMjacking, and how do you prevent it? LLMjacking is when attackers steal cloud or API credentials and run expensive AI models on your account. Prevent it with one scoped key per workload, a secrets manager, secret scanning, regular key rotation, and automatic shutdown when spending jumps well above normal levels.

How do you stop an AI agent from running in an infinite loop? Give every agent hard limits: a maximum number of steps, a token budget per task and a rule that stops repeated identical actions. Check these limits before each model call, not after, so the agent stops before the next request is billed.

What is the difference between a budget alert and budget enforcement? A budget alert tells you that spending has already crossed a line. Budget enforcement blocks the next request before it runs. Alerts help with reporting, but only enforcement stops runaway spending in time, which matters when agents make thousands of calls per hour.

What is the best metric for AI cost management? Cost per successful outcome is the most useful metric. It shows what each resolved ticket, merged pull request or finished report really costs. Track it alongside cache hit rate, model routing share and how quickly your team can stop a spending spike.

< previous
Tarento: A Premier Delivery Partner for Infor ERPs
Next >
LakeBridge vs SnowConvert AI vs DataVolve: Data Migration Tool Coverage Compared (2026)
Next >
logo
Thor Bot Avatar