GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

OpenAI Enterprise: 64% of Output Tokens Now Run Agent Loops — Chat Has Become the Minority Workload

OpenAI’s Enterprise Signals report, updated August 12, contains a number that reframes how frontier AI compute is actually being used: as of June 2026, agentic AI — defined as Codex tokens — accounts for 64% of combined Codex and ChatGPT output tokens across enterprise customers.

Chat is now the minority workload.

What This Measures

OpenAI draws a deliberate distinction in its framing. Both ChatGPT and Codex use the same underlying frontier models. The difference is execution mode: ChatGPT tokens go toward a human reading a reply; Codex tokens go toward an agent completing a multi-step task — writing code, running tests, managing files, handling errors, iterating.

The 64% figure captures that shift in aggregate. In dollar terms, most enterprise frontier compute is now funding loops, not responses.

Why Token Composition Matters

Agent loops consume tokens differently from chat. A chat session averages a few hundred tokens per turn. A Codex task — debugging a production regression, migrating a codebase section, building a feature end-to-end — can run tens of thousands of tokens across dozens of tool calls. The token-per-useful-output ratio is orders of magnitude higher.

That means the 64% share of output tokens almost certainly understates the share of useful enterprise work done by agents. When 64% of tokens are agentic and each agentic token does more work per unit than a chat token, the effective proportion of business output attributable to agent execution is higher still.

The Frontier Model Economics Shift

For model providers, this changes the revenue structure. Enterprise accounts that started as seat-based ChatGPT contracts now consume the bulk of their compute through Codex, which is metered differently, runs at higher token volumes, and has different latency tolerance. A customer whose finance team uses ChatGPT for summaries and whose engineering team runs Claude Code and Codex simultaneously is two different revenue lines with very different growth curves.

At 64%, agentic is already large enough to set pricing strategy. The next question is what the number looks like at 70%, 80%. OpenAI notes enterprise customers can request customised benchmarks against the dataset, suggesting the aggregate figure masks meaningful variation by industry vertical. Financial services and software companies are almost certainly above 64%; sectors earlier in AI adoption are below it.

The Floor

June 2026 is the reference point. The trend from every observable metric — Codex user counts, Claude Code growth, GitHub Copilot token volume, AWS AgentCore adoption — points the same direction. The 64% figure is a floor, not a target.

The composition of enterprise AI token use has flipped. What started as an experiment with chat interfaces has reorganised around autonomous task execution. Frontier model providers that didn’t build for agent workloads have been repricing and rebundling since Q1 2026. The June data confirms that repricing was warranted.