GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
← Back to feed

OpenAI Launches ChatGPT Images 2.0: 99% Text Accuracy, Reasoning-Enabled Generation, 4K Output

OpenAI has launched ChatGPT Images 2.0, powered by a new model (gpt-image-2), available now across ChatGPT, Codex for Mac, and the developer API. The rollout extends to all users without requiring plan changes; advanced outputs tied to thinking and reasoning models are gated to Plus, Pro, Business, and Enterprise tiers.

What Changed

The headline improvement is text rendering. Prior GPT Image models could produce visually compelling results while producing garbled or illegible text — a limitation that excluded them from professional design, publishing, and marketing workflows. GPT-Image-2 posts 99% accuracy on standard typography benchmarks, and handles dense small-font layouts, multilingual scripts (Japanese, Korean, Chinese, Hindi, Bengali), iconography, and UI mockups. OpenAI is explicitly positioning the model for magazine covers, storyboards, product catalogues, and software interface prototyping.

Resolution caps at 4096×4096 pixels at full quality. API developers can output up to 2K. Generation speed is roughly twice that of GPT-Image-1.5.

Aspect ratio support now runs from 3:1 panoramas down to 1:3 vertical — adding banners, mobile assets, and slide formats that were poorly served by the previous fixed-ratio options.

Reasoning Integration

The more consequential addition is thinking-model coupling. When a user selects a thinking or pro model in ChatGPT, Images 2.0 can pull live web context, cross-check its own outputs, and generate up to eight distinct but thematically consistent images from a single prompt. A content team can, in one request, get a set of social assets at different dimensions with matching characters and visual language.

This is the first time OpenAI has wired its reasoning stack into image generation. The practical effect is a pipeline that can ground images in current information — a product launch image that incorporates today’s branding guidelines fetched from a URL, for example — rather than relying solely on the model’s static knowledge.

Key Numbers

  • Text rendering accuracy: ~99% on typography benchmarks
  • Maximum output: 4096×4096 (full quality), 2K in API
  • Generation speed: ~2x faster than GPT-Image-1.5
  • Aspect ratios: 3:1 to 1:3
  • Max outputs per thinking prompt: 8 coherent images
  • Knowledge cutoff: December 2025
  • API model ID: gpt-image-2

What It Competes With

The practical benchmark is Nano Banana Pro (Google’s Gemini 3 Pro Image), which held the Arena image generation top spot through early April after the leaked tape models appeared on LM Arena. GPT-Image-2 is the model that was tested under codenames maskingtape, gaffertape, and packingtape before being pulled from the leaderboard in April. Microsoft’s MAI-Image-2-Efficient, launched April 14 at $19.50/M output tokens, competes primarily on throughput and cost for volume tasks.

OpenAI’s differentiation is the reasoning integration and text fidelity. For professional users who need legible type and current-context grounding, the value proposition is distinct from the aesthetic-first tools.

Acknowledged limitations: complex physical-world reasoning (origami, spatial puzzles), highly repetitive textures, and detailed diagrams still require review.