Google Launches Gemini Omni at I/O 2026: World Model Targets Any Input, Any Output
Google used its I/O 2026 keynote to announce Gemini Omni, a new model family positioned as the lab’s answer to multimodal generative AI at platform scale. DeepMind CEO Demis Hassabis described it as a model that can “create anything from any input” — combining Gemini’s reasoning and world understanding with Google’s existing media generation stack.
The first model in the family, Gemini Omni Flash, is rolling out globally through the Gemini app, Google Flow, and YouTube Shorts. API access for developers and enterprise customers is coming soon.
What Omni Actually Is
Gemini Omni is not a simple video-generation bolt-on. According to Google, it merges Gemini’s core intelligence with Veo (video), Nano Banana (image), and Genie (world simulation) into a unified model capable of generating media grounded in real-world knowledge.
The architecture is designed around physical world simulation — understanding what happens next based on a user’s actions. That lineage traces back to DeepMind’s years of world-model research, and the Omni framing is explicit: this is meant to simulate environments, not just render pixels.
Today’s launch focuses on video output. Inputs can be any combination of text, images, video, and audio. Google demonstrated conversational video editing — users can describe changes in natural language and have Omni apply them. Image and audio output modalities will follow in subsequent rollouts.
Why the Nano Banana Reference Matters
Google described Gemini Omni Flash as “the Nano Banana of video” — meaning it targets Nano Banana’s position in image generation: a fast, capable, widely-deployed model that can run at scale without the cost of the frontier tier. The implication is that Omni Flash is optimized for everyday video creation tasks that don’t require full Veo quality.
This mirrors how Google split the Gemini family into Flash and Pro tiers. Omni Flash is the workhorse. A higher-tier Omni model is implied but unannounced.
Platform Deployment
Omni Flash availability at launch:
- Gemini app — accessible to Google AI subscribers
- Google Flow — Google’s AI video production tool
- YouTube Shorts — integration for short-form content creation
- Developer API — arriving shortly after the keynote
The integration with YouTube Shorts is the highest-volume deployment. Shorts serves over 70 billion daily views. Putting a video generation model directly in the creation flow gives Google a dataset and usage moat that no other lab can replicate at that scale.
Context: 3.2 Quadrillion Tokens
In his I/O opening, Sundar Pichai disclosed that Google now processes 3.2 quadrillion AI tokens per month, up from 480 trillion at I/O 2025 — a 6.7x increase in a year. That scale sits underneath every Gemini product including Omni, and it sets the benchmark for what deploying a multimodal world model at infrastructure level actually looks like.
What It Competes With
Omni Flash enters a market where OpenAI has retired Sora (unprofitable at $1M/day) and replaced it with a video generation API layer. Runway and Kling dominate independent video generation. Google’s angle is different: Omni Flash is not positioned as a standalone video product but as infrastructure embedded in apps that already have billions of users.
The world model framing also differentiates it from every current competitor. Runway generates video. Omni, in Google’s framing, understands video as a simulation of physical reality — a distinction that will matter more as robotic and autonomous system applications come online.