GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Goldman: Agent Token Use to Multiply 24x by 2030 as Enterprises Hit the First AI Cost Ceiling

Goldman Sachs has put a number on the agent token economy: consumption will multiply 24 times by 2030, with the bullish scenario reaching 120 quadrillion tokens per month. The forecast arrives at the same moment that early enterprise deployments are producing the first concrete evidence of what agentic AI actually costs in production — and the math is worse than most finance teams expected.

The Mechanics of Agent Token Demand

The key distinction from chatbot usage is loop depth. A chatbot answers once. An agent plans, calls tools, retrieves data, evaluates results, decides next steps, corrects errors, and repeats. A single user request in an agentic workflow can consume 10x to 50x the tokens of an equivalent chatbot interaction, and in complex multi-agent pipelines, the multiplier can run higher.

Goldman’s 24x projection is built on two compounding trends: AI agent adoption expanding from early-adopter enterprise deployments to broad workflow integration across industries, and per-task compute increasing as agents are given longer context, more tool calls, and more verification steps. The offsetting force is inference cost, which Goldman forecasts will fall 60-70% per year — consistent with the trajectory of the past three years, where frontier API prices have dropped roughly 97% since 2023.

The net result is a token volume explosion that does not necessarily translate to equivalent spending growth, provided inference costs continue falling on schedule.

Where the Cost Ceiling Is Appearing Now

The projection is forward-looking, but the cost reckoning is already live in 2026. Uber’s COO said publicly that the company is scrutinizing AI token spend more carefully after costs ran ahead of productivity gains. Microsoft has been revoking Claude Code access from its developer workforce, with full cutover to Copilot CLI planned by June 30 — framed as product consolidation but timed to fiscal year-end in a way that points to cost reduction as a motivation.

The most extreme case is anecdotal but illustrative: a company reportedly burned through $500M in Claude credits in a single month after failing to configure per-user spending limits. The number is large enough to be an outlier, but the failure mode — no usage governance on agentic deployments — is structural, not exceptional.

CNBC reporting indicates roughly 95% of enterprise AI inference still runs on frontier-tier models even for simple, low-value tasks. The obvious correction — routing cheap tasks to cheaper models — delivers 20-30% cost reduction in deployments like Augment Code’s Prism system, which released data in April 2026 showing per-turn model routing at frontier quality.

The Trade-Off

Goldman’s framing is blunt: the agent economy is a fight between productivity and token waste. If agents are completing tasks that previously took hours of human time, the token cost is irrelevant at the margin. If agents are consuming tokens on failed attempts, redundant verification loops, and poorly-scoped tasks that produce outputs nobody uses, the economics invert.

The enterprises currently rethinking their AI spend appear to be confronting a version of the second scenario. Agent deployments rolled out without granular cost attribution, task-level ROI measurement, or tiered model routing are burning through budgets at rates that are hard to justify against productivity data — like the NBER survey showing 90% of firms reporting zero measurable productivity impact over three years of AI use.

What 120 Quadrillion Tokens Looks Like

For context: a trillion tokens is roughly 750 billion words of text, or about 1.5 million novels. 120 quadrillion tokens monthly — Goldman’s bullish 2030 case — is 120 million trillion tokens. At current frontier pricing of around $5-15/M input tokens, that would imply spending in the hundreds of trillions per year, which is obviously impossible. The scenario only works if inference costs fall by several more orders of magnitude, which Goldman’s -65%/year forecast supports but does not guarantee.

The practical planning implication for enterprise AI buyers is straightforward: the cost-per-token curve is your runway. Agents that are not productive at $5/M input will be productive at $0.05/M. The question is whether the governance and tooling infrastructure is in place to capture that value as costs fall, or whether the budget pressure arrives first.