Google Launches Agentic Video Understanding for Gemini Models with 88% Lower Token Usage

2 Sources

Share

Google introduced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, enabling dynamic video processing that cuts token usage by up to 88% and boosts accuracy by 7%. The feature moves beyond static frame-by-frame processing to intelligently analyze videos across frames, audio, and transcripts.

Google has introduced agentic video understanding across its latest Gemini models, marking a shift from traditional static processing to intelligent, dynamic video analysis.

1

2

The feature is now available for Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite, delivering substantial improvements in both efficiency and accuracy for video analysis tasks.

Source: Google

Source: Google

Dramatic Reduction in Token Usage and Costs

The agentic video understanding capability reduces token usage by up to 88% while cutting analysis costs by up to 66%.

2

These efficiency gains represent a breakthrough for developers and enterprises processing video content at scale. Google reports that the feature also improves accuracy by up to 7% across standard video analysis benchmarks, making it both more cost-effective video analysis and more reliable than previous approaches.

How Agentic Video Understanding Works

Unlike static processing, which splits videos into individual frames at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding allows Gemini models to decide what to watch and at what speed.

1

The AI can dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts.

2

This intelligent approach pairs the model's core reasoning with native video tools, similar to how agentic vision combines code execution with image understanding capabilities.

New Capabilities Unlock Advanced Use Cases

The dynamic video processing approach enables Gemini models to pinpoint split-second changes and answer complex questions across multi-hour videos.

1

Specific capabilities include sub-second moment retrieval, more accurate anomaly detection, and precise counting and tracking of physical movements and objects. Users can now inspect videos for visual artifacts with greater precision, opening possibilities for quality control, content moderation, and detailed video analytics.

Benefits for Long-Form Video Analysis

The efficiency gains prove especially pronounced for long-form video analysis, from 10-minute how-to guides to 90-minute lectures and multi-hour recordings.

2

Previously, static processing forced developers to choose between high token costs or techniques that dropped critical details. The agentic approach eliminates this trade-off by intelligently selecting which elements to analyze and when.

Availability and Future Rollout Plans

Agentic video understanding is currently available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform for both video uploads and YouTube videos.

2

Google announced that the feature will roll out to the Gemini app soon, expanding access beyond API users.

1

The company also plans to integrate agentic video understanding into YouTube's "Ask YouTube" feature, potentially making it easier for creators to analyze their videos with AI-driven video analysis tools. This integration could transform how content creators understand viewer engagement and optimize their productions across the platform.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved