Skip to content
Google@Google · Sep 1, 2026

We’re introducing a new capability to our latest Gemini models: agentic video understanding. This allows developers to…

4 tweets1 min read5.8Koriginal

Summary

Google is launching agentic video understanding for Gemini models, enabling dynamic processing of long-form video content that reduces token usage by 88% while improving accuracy by 7%. The feature is available now via the Gemini API and coming soon to the Gemini app and YouTube.

Summarized by ThreadOut AI from the full thread. May miss nuance — read the thread below.

  1. #1

    We’re introducing a new capability to our latest Gemini models: agentic video understanding. This allows developers to process long-form video content with more accuracy, while using up to 88% less tokens. See how it works 🧵

  2. #2

    Today, most AI models use “static” processing to analyze videos, looking at just one frame-per-second by default. Agentic video understanding allows Gemini to dynamically process and reason across the video file — scanning and inspecting visual frames, audio, and transcripts while using native tools to adjust its processing speed — making it easier to find what you need, faster.

  3. #3

    With agentic video understanding, developers can dynamically search, scan, and inspect long-form video with up to: 📉 88% fewer tokens 📉 66% lower cost ✅ 7% better accuracy

  4. #4

    Agentic video understanding is available now across our latest models via the Gemini API in @GoogleAIStudio and the Gemini Enterprise Agent Platform. Coming soon to the @GeminiApp and to @YouTube's “Ask YouTube” feature on the video watch page. Learn more ↓ goo.gle/4x5Knd1