GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

xAI Imagine Image 2.0 Enters Arena at #2 in Both Image Categories — 24 Points Behind GPT-Image-2 in Editing

xAI released Imagine Image 2.0 on August 8 as the new Quality Mode on grok.com/imagine and across its iOS and Android apps. Arena benchmarks published the same day place the model second globally in both image arenas, a substantial jump from the previous variant.

Arena Numbers

CategoryEloGlobal RankLeader (GPT-Image-2)
Image Edit Arena1,439#21,463
Text-to-Image Arena1,320#21,380

The previous xAI image model (grok-imagine-image-quality) ranked approximately 14th in Text-to-Image with around 1,228 Elo and 1,390 in Image Edit. The 2.0 jump is the largest single-generation Arena improvement from any image model this year.

The rankings were recorded using the “low” speed variant of Image 2.0. The model appears on Arena under the SpaceXAI label following xAI’s dissolution into SpaceX in July.

Other models that rank further down: Reve 2.1, Meta Muse-Image, Alibaba Qwen-Image-3.0-Pro, Google Gemini, ByteDance SeedDream.

What Changed

Image 2.0 is built for controlled, dense compositions rather than one-off generation. The capability additions:

  • Region editing: Magic wand and segmentation tools target a specific area while preserving the rest of the image unchanged.
  • Transparent export: Background removal to a clean alpha channel, without needing an external tool.
  • Multi-reference composition: Up to five reference images combined in a single generation.
  • Smart Resize: Frame extension with ratios from 1:2 through 2:1, filling new space without distorting the original content.
  • Typography control: Layout planning for dense text compositions, small text kept legible across the frame.

API access is not yet available. The model is Grok app and web only for now. xAI said API access is “coming soon.”

Gap to GPT-Image-2

The 24-point editing gap is within realistic striking range. The 60-point Text-to-Image gap is larger and puts Image 2.0 in a different tier for pure generation quality despite the identical rank position.

GPT-Image-2 has led both Arena image categories since its April 2026 launch. No other model has consistently challenged it at the top of both tables simultaneously. Image 2.0 is the first model from any lab to hold second place across both categories in the same evaluation cycle.

The image editing market is where the real commercial application sits: product photography, marketing asset production, and layered design workflows that require precision edits rather than full regeneration. On that dimension, the 24-point gap suggests Image 2.0 is a credible alternative for editing use cases even before API access arrives.