Goldman: Agent Token Use to Multiply 24x by 2030 as Enterprises Hit the First AI Cost Ceiling
Goldman Sachs has put a number on the agent token economy: consumption will multiply 24 times by 2030, with the bullish scenario reaching 120 quadrillion tokens per month. The forecast arrives at the same moment that early enterprise deployments are producing the first concrete evidence of what agentic AI actually costs in production — and the math is worse than most finance teams expected.
The Mechanics of Agent Token Demand
The key distinction from chatbot usage is loop depth. A chatbot answers once. An agent plans, calls tools, retrieves data, evaluates results, decides next steps, corrects errors, and repeats. A single user request in an agentic workflow can consume 10x to 50x the tokens of an equivalent chatbot interaction, and in complex multi-agent pipelines, the multiplier can run higher.
Goldman’s 24x projection is built on two compounding trends: AI agent adoption expanding from early-adopter enterprise deployments to broad workflow integration across industries, and per-task compute increasing as agents are given longer context, more tool calls, and more verification steps. The offsetting force is inference cost, which Goldman forecasts will fall 60-70% per year — consistent with the trajectory of the past three years, where frontier API prices have dropped roughly 97% since 2023.
The net result is a token volume explosion that does not necessarily translate to equivalent spending growth, provided inference costs continue falling on schedule.
Where the Cost Ceiling Is Appearing Now
The projection is forward-looking, but the cost reckoning is already live in 2026. Uber’s COO said publicly that the company is scrutinizing AI token spend more carefully after costs ran ahead of productivity gains. Microsoft has been revoking Claude Code access from its developer workforce, with full cutover to Copilot CLI planned by June 30 — framed as product consolidation but timed to fiscal year-end in a way that points to cost reduction as a motivation.
The most extreme case is anecdotal but illustrative: a company reportedly burned through $500M in Claude credits in a single month after failing to configure per-user spending limits. The number is large enough to be an outlier, but the failure mode — no usage governance on agentic deployments — is structural, not exceptional.
CNBC reporting indicates roughly 95% of enterprise AI inference still runs on frontier-tier models even for simple, low-value tasks. The obvious correction — routing cheap tasks to cheaper models — delivers 20-30% cost reduction in deployments like Augment Code’s Prism system, which released data in April 2026 showing per-turn model routing at frontier quality.
The Trade-Off
Goldman’s framing is blunt: the agent economy is a fight between productivity and token waste. If agents are completing tasks that previously took hours of human time, the token cost is irrelevant at the margin. If agents are consuming tokens on failed attempts, redundant verification loops, and poorly-scoped tasks that produce outputs nobody uses, the economics invert.
The enterprises currently rethinking their AI spend appear to be confronting a version of the second scenario. Agent deployments rolled out without granular cost attribution, task-level ROI measurement, or tiered model routing are burning through budgets at rates that are hard to justify against productivity data — like the NBER survey showing 90% of firms reporting zero measurable productivity impact over three years of AI use.
What 120 Quadrillion Tokens Looks Like
For context: a trillion tokens is roughly 750 billion words of text, or about 1.5 million novels. 120 quadrillion tokens monthly — Goldman’s bullish 2030 case — is 120 million trillion tokens. At current frontier pricing of around $5-15/M input tokens, that would imply spending in the hundreds of trillions per year, which is obviously impossible. The scenario only works if inference costs fall by several more orders of magnitude, which Goldman’s -65%/year forecast supports but does not guarantee.
The practical planning implication for enterprise AI buyers is straightforward: the cost-per-token curve is your runway. Agents that are not productive at $5/M input will be productive at $0.05/M. The question is whether the governance and tooling infrastructure is in place to capture that value as costs fall, or whether the budget pressure arrives first.