@GoogleDeepMind: We’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video…
Summary
Google DeepMind announces Gemini Omni, a new model that combines Gemini's intelligence with generative media systems to create video from any input, marking a significant step in multimodal AI.
View Cached Full Text
Cached at: 05/19/26, 06:50 PM
We’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video.
It combines Gemini’s intelligence with our generative media systems - representing a leap forward in world understanding, multimodality, and editing
Omni brings together an improved understanding of physics with Gemini’s knowledge of history, biology, and culture, bridging the gap from photorealism to meaningful storytelling.
Actions have consequences, environments respond to events, and narratives evolve logically.
Define a character once - then place them in any scene, and they’ll stay consistent across locations, actions and lighting.
Apply styles, motion, or effects by using input references, or just describe it with natural language.
You can even reimagine the action in a video you took by asking Gemini Omni.
Transform your world instantly - change the environment, add new objects, or create something completely unexpected.
You can try Gemini Omni Flash - the first model in the Omni family - in the @GeminiApp, @FlowbyGoogle and @YouTube Shorts.
In the coming weeks, we’ll also be rolling it out via APIs. #GoogleIO
Similar Articles
Introducing Gemini Omni: Create Anything from Anything
Google introduces Gemini Omni, a new multimodal AI model capable of processing and generating content across text, images, audio, and video from any input type.
Gemini Omni
Gemini Omni is a new AI model from Google DeepMind that combines reasoning with creative capabilities, enabling multimodal understanding, video editing, and content generation, with built-in safety measures and digital watermarking.
Google’s Gemini Omni turns images, audio, and text into video — and that’s just the start
Google announces Gemini Omni, a family of multimodal models that can generate video from images, audio, and text, reasoning across inputs to produce consistent, high-quality outputs. The first model, Gemini Omni Flash, rolls out at Google I/O to the Gemini app, YouTube Shorts, and Flow.
Gemini Omni 1.1 Flash lets you build with more control
Google DeepMind has released Gemini Omni 1.1 Flash, an updated AI model for generative video with enhanced creative controls, scene extension, and faster prototyping, now available via APIs in Google AI Studio.
@GoogleDeepMind: Build your next story with Gemini Omni.
Google DeepMind announces Gemini Omni, a new AI model for building stories.