Skip to content
Google DeepMind@GoogleDeepMind · May 19, 2026

We’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with…

5 tweets1 min read8.4Koriginal

Summary

Google announces Gemini Omni, a new AI model that combines advanced intelligence with generative media capabilities to create videos from text descriptions. The model maintains character consistency across scenes, understands physics and logic, and allows editing through natural language or style references.

Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.

  1. #1

    We’re dropping Gemini Omni: our first step towards a model that can create anything from anything - starting with video. It combines Gemini’s intelligence with our generative media systems - representing a leap forward in world understanding, multimodality, and editing 🧵

  2. #2

    Omni brings together an improved understanding of physics with Gemini's knowledge of history, biology, and culture, bridging the gap from photorealism to meaningful storytelling. Actions have consequences, environments respond to events, and narratives evolve logically.

  3. #3

    🔘 Define a character once - then place them in any scene, and they’ll stay consistent across locations, actions and lighting. 🔘 Apply styles, motion, or effects by using input references, or just describe it with natural language.

  4. #4

    You can even reimagine the action in a video you took by asking Gemini Omni. Transform your world instantly - change the environment, add new objects, or create something completely unexpected.

  5. #5

    You can try Gemini Omni Flash - the first model in the Omni family - in the @GeminiApp, @FlowbyGoogle and @YouTube Shorts. In the coming weeks, we'll also be rolling it out via APIs. #GoogleIO