GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

OpenAI Codex and ChatGPT Work Hit 8M Users — Usage Limits Reset for the Fifth Time in a Week

OpenAI reset usage limits on GPT-5.6 Sol for the fifth time since public launch on July 10, as Codex and ChatGPT Work together crossed 8 million active users. Five resets in five days is the clearest public signal that demand is outrunning provisioned inference capacity.

CEO Sam Altman acknowledged the pressure directly on X: “5.6 Sol growth is insane. The inference team has done heroic work to continue to scale, but it is possible there are some hiccups soon.” OpenAI rarely flags capacity risk publicly before problems occur.

What a Reset Means

Each usage reset removes the five-hour per-session cap on Sol access, temporarily giving all Plus and Pro subscribers unrestricted agentic sessions. OpenAI is using resets as a pressure valve — a workaround for the gap between committed capacity and actual demand.

The demand profile is non-linear. Codex and ChatGPT Work are used primarily for multi-step coding tasks, which consume far more tokens per session than conversational chat. At 750 tokens per second on Cerebras-routed infrastructure — up from 70-100 TPS for GPT-5.5 xHigh — sessions complete faster, but per-session compute cost remains significant. Users can run more ambitious tasks per session, which drives up total cluster load even as individual token efficiency improves.

Altman’s Efficiency Numbers

In a CNBC interview on July 14, Altman disclosed that Sol is 54% more token-efficient than GPT-5.5 xHigh on agentic coding tasks. The 750 TPS Cerebras figure means a coding session that took 40 minutes at GPT-5.5 xHigh now completes in under 10 minutes, while consuming roughly half the tokens.

For enterprise customers running autonomous coding pipelines, the efficiency gain is material. For OpenAI’s infrastructure team, the math is harder: more efficient tokens per task means each user can complete more ambitious tasks in the same session window, which may be driving higher total utilisation per cohort even as per-task token consumption falls.

ChatGPT Work as Enterprise Layer

The 8 million figure combines Codex — the autonomous background coding agent — with ChatGPT Work, which launched simultaneously as a managed-environment version of the desktop app. Work adds workspace isolation, IT admin controls, and enterprise audit logging. The merged product positions OpenAI directly against Cursor and Claude Code for enterprise coding budgets.

The original ChatGPT Work launch article covered the product announcement. The usage numbers published this week describe the business: 8 million active users in five days puts Codex adoption faster than ChatGPT’s own early growth curve by most historical comparisons.

Infrastructure Headroom

Altman said OpenAI will “move mountains” to scale capacity but gave no specific hardware timelines. The recurring resets suggest managed demand queuing rather than hard capacity walls — limits reset when the system determines short-term throughput has cleared a backlog.

OpenAI’s current footprint includes 220,000 GPUs at SpaceX Colossus 1, with Colossus 2 expansion reported but not yet operational. At 8 million active users running sustained agentic sessions, the gap between demand and supply is the kind of problem the company would prefer to have — and also the one that most urgently needs a resolution.