Luma uni-1.1-max Enters Arena Image Leaderboards at #1 Human Preference Elo — $0.10 Per Image
Luma Labs’ uni-1.1-max entered Arena’s Text-to-Image and Image Edit leaderboards on May 5, 2026, landing at #1 on Human Preference Elo across overall, style and editing, and reference-based generation categories. The Arena Leaderboard Changelog confirms both uni-1.1 and uni-1.1-max were added simultaneously.
The model positions Luma as a top-3 lab on Arena’s image benchmarks — the first time the company has placed at that tier across both generation and editing tasks.
Architecture
Uni-1.1 takes a different approach from most image generation models. Reasoning and image generation run inside the same model architecture rather than in separate systems stitched together at inference time. Luma calls it a unified understanding and generation model: the same weights that process reasoning and visual understanding also drive generation.
The practical result, per Luma’s benchmarks: tighter adherence to multi-constraint prompts, cleaner reference image grounding, and editing that responds to intent rather than to prompt syntax.
On RISEBench — a benchmark evaluating Reasoning-Informed Visual Editing across Temporal, Causal, Spatial, and Logical categories — uni-1.1 achieves state-of-the-art results, leading on both overall reasoning and spatial logic. The RISEBench result is notable because spatial reasoning in image editing (accurate shadows, correct perspective, physically plausible object layouts) is a consistent failure mode for generation models that don’t integrate visual understanding tightly with the generation process.
Capabilities
Both uni-1.1 and uni-1.1-max support the same feature set through a unified API endpoint:
- Text rendering — readable text on signs, labels, and surfaces
- Spatial reasoning — accurate shadows, perspective, and object physics
- Reference-guided generation — up to 9 reference images for text-to-image, 8 for editing
- Multi-panel output — storyboards and sequential frames with consistent style
- Web search grounding — searches for visual references before generating when enabled
- Cultural styles — manga, ukiyo-e, film noir, and other visual traditions
Image editing uses a source parameter — the model modifies the source image based on the prompt while preserving unmentioned parts. Style transfer, background replacement, and object swaps are the primary use cases.
Pricing
| Model | Per image (text-to-image, 2K) |
|---|---|
| uni-1.1 | $0.0404 |
| uni-1.1-max | $0.1000 |
Luma claims the models run at less than half the price and latency of comparable models. Generation time is approximately 31 seconds per image at the uni-1.1 tier. Provisioned throughput is available for production workloads requiring guaranteed capacity.
Context
Uni-1 launched in March 2026. The 1.1 release is the first update, adding improvements to output quality without changing the API surface — the same prompts, parameters, and aspect ratio options work identically on both model versions. Developers switch by changing only the model field.
Arena added both uni-1.1 variants on May 5, the same day gpt-image-2 (medium) received an updated score reflecting performance across Arena’s full user base at scale — Arena’s standard methodology for score stabilisation as sample sizes grow.
The image generation arena now has three distinct competitive tiers: proprietary frontier models (OpenAI, Google), specialised generation labs (Luma, Stability), and open-weight providers. Luma’s #1 Human Preference Elo position makes it the top-ranked non-frontier-lab entrant on the leaderboard.