Product Updates Google DeepMind Blog

Introducing agentic video understanding with Gemini

Geminivideo understandingagentic AIGoogle DeepMind

Video analysis traditionally uses static processing, where the model ingests video at a fixed frames-per-second rate (default 1 FPS). This forces developers to choose between high token costs or techniques that drop critical details, especially for long-form content like 10-minute how-to guides, 90-minute lectures, and multi-hour recordings.

Agentic video understanding instead pairs Gemini's core reasoning with native video tools, enabling the model to take an active, goal-directed role. It decides what to watch, at what speed, and through which modality (frames, audio, or transcript), invoking an internal tool to load only the relevant segments. This unlocks capabilities like sub-second moment retrieval, more accurate anomaly detection, and precise counting.

The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Across standard benchmarks, Gemini models with agentic video understanding cut token consumption by up to 88%, reduce analysis costs by up to 66%, and improve accuracy by up to 7%. Gemini 3.7 Flash with agentic understanding offers the best overall quality and the best quality-to-cost combination, placing it at the accuracy-to-cost Pareto frontier among tested video understanding models.

Read original →

← Back to home