GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Google Launches Gemini Omni at I/O 2026: World Model Targets Any Input, Any Output

Google used its I/O 2026 keynote to announce Gemini Omni, a new model family positioned as the lab’s answer to multimodal generative AI at platform scale. DeepMind CEO Demis Hassabis described it as a model that can “create anything from any input” — combining Gemini’s reasoning and world understanding with Google’s existing media generation stack.

The first model in the family, Gemini Omni Flash, is rolling out globally through the Gemini app, Google Flow, and YouTube Shorts. API access for developers and enterprise customers is coming soon.

What Omni Actually Is

Gemini Omni is not a simple video-generation bolt-on. According to Google, it merges Gemini’s core intelligence with Veo (video), Nano Banana (image), and Genie (world simulation) into a unified model capable of generating media grounded in real-world knowledge.

The architecture is designed around physical world simulation — understanding what happens next based on a user’s actions. That lineage traces back to DeepMind’s years of world-model research, and the Omni framing is explicit: this is meant to simulate environments, not just render pixels.

Today’s launch focuses on video output. Inputs can be any combination of text, images, video, and audio. Google demonstrated conversational video editing — users can describe changes in natural language and have Omni apply them. Image and audio output modalities will follow in subsequent rollouts.

Why the Nano Banana Reference Matters

Google described Gemini Omni Flash as “the Nano Banana of video” — meaning it targets Nano Banana’s position in image generation: a fast, capable, widely-deployed model that can run at scale without the cost of the frontier tier. The implication is that Omni Flash is optimized for everyday video creation tasks that don’t require full Veo quality.

This mirrors how Google split the Gemini family into Flash and Pro tiers. Omni Flash is the workhorse. A higher-tier Omni model is implied but unannounced.

Platform Deployment

Omni Flash availability at launch:

  • Gemini app — accessible to Google AI subscribers
  • Google Flow — Google’s AI video production tool
  • YouTube Shorts — integration for short-form content creation
  • Developer API — arriving shortly after the keynote

The integration with YouTube Shorts is the highest-volume deployment. Shorts serves over 70 billion daily views. Putting a video generation model directly in the creation flow gives Google a dataset and usage moat that no other lab can replicate at that scale.

Context: 3.2 Quadrillion Tokens

In his I/O opening, Sundar Pichai disclosed that Google now processes 3.2 quadrillion AI tokens per month, up from 480 trillion at I/O 2025 — a 6.7x increase in a year. That scale sits underneath every Gemini product including Omni, and it sets the benchmark for what deploying a multimodal world model at infrastructure level actually looks like.

What It Competes With

Omni Flash enters a market where OpenAI has retired Sora (unprofitable at $1M/day) and replaced it with a video generation API layer. Runway and Kling dominate independent video generation. Google’s angle is different: Omni Flash is not positioned as a standalone video product but as infrastructure embedded in apps that already have billions of users.

The world model framing also differentiates it from every current competitor. Runway generates video. Omni, in Google’s framing, understands video as a simulation of physical reality — a distinction that will matter more as robotic and autonomous system applications come online.