2 Sources
[1]
Gemini now analyzes your videos more accurately, and at a lower cost
It can also track physical movement and count distinct objects in videos. Google has been steadily improving Gemini to make it more useful. We've already spotted the company working on new features for Gemini on mobile, while also releasing new Gemini models such as Gemini 3.7 Flash. However, Google is now giving Gemini the ability to analyze videos in a more cost-effective and efficient manner. Google announced today that it's bringing agentic video understanding to the latest Gemini models: Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. This new capability allows users to upload videos to Gemini and have the AI analyze them. It also reduces token usage by up to 88% and offers up to 7% better accuracy. Users have been able to upload videos to Gemini for a while now, but the AI could only perform what Google calls "static" processing on them. This meant it would split the video into individual frames, resulting in slower performance and higher costs. With agentic video understanding, Gemini can now decide what to watch and at what speed. It can also choose between frames, audio, and the transcript to analyze videos more efficiently. This means Gemini can pinpoint split-second changes, answer complex questions across multi-hour videos, inspect videos for visual artifacts, and even count and track physical movements and objects. Right now, agentic video understanding is only available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. However, the company also said the feature will roll out to the Gemini app soon. Google will also soon start using agentic video understanding to power YouTube's "Ask YouTube" feature, which could make it easier for creators to analyze their videos with Gemini.
[2]
Introducing agentic video understanding with Gemini
Summaries were generated by Google AI. Generative AI is experimental. Today, we're launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to agentic vision, which combines code execution with Gemini models' native image understanding, agentic video understanding uses Gemini's native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more. The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Benchmarks Unlike current 'static' processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding pairs the model's core reasoning with native video tools to dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts. Across standard video analysis benchmarks, Gemini models with agentic video understanding reduce analysis costs by up to 66% and token consumption by up to 88%, while improving accuracy by up to 7%. These efficiency gains are especially pronounced on long-form video (from 10-minute how-to guides to 90-minute lectures and multi-hour recordings), where static processing forces developers to choose between high token costs or techniques that drop critical details.
Share
Copy Link
Google introduced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, enabling dynamic video processing that cuts token usage by up to 88% and boosts accuracy by 7%. The feature moves beyond static frame-by-frame processing to intelligently analyze videos across frames, audio, and transcripts.
Google has introduced agentic video understanding across its latest Gemini models, marking a shift from traditional static processing to intelligent, dynamic video analysis.
1
2
The feature is now available for Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite, delivering substantial improvements in both efficiency and accuracy for video analysis tasks.
Source: Google
The agentic video understanding capability reduces token usage by up to 88% while cutting analysis costs by up to 66%.
2
These efficiency gains represent a breakthrough for developers and enterprises processing video content at scale. Google reports that the feature also improves accuracy by up to 7% across standard video analysis benchmarks, making it both more cost-effective video analysis and more reliable than previous approaches.Unlike static processing, which splits videos into individual frames at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding allows Gemini models to decide what to watch and at what speed.
1
The AI can dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts.2
This intelligent approach pairs the model's core reasoning with native video tools, similar to how agentic vision combines code execution with image understanding capabilities.The dynamic video processing approach enables Gemini models to pinpoint split-second changes and answer complex questions across multi-hour videos.
1
Specific capabilities include sub-second moment retrieval, more accurate anomaly detection, and precise counting and tracking of physical movements and objects. Users can now inspect videos for visual artifacts with greater precision, opening possibilities for quality control, content moderation, and detailed video analytics.Related Stories
The efficiency gains prove especially pronounced for long-form video analysis, from 10-minute how-to guides to 90-minute lectures and multi-hour recordings.
2
Previously, static processing forced developers to choose between high token costs or techniques that dropped critical details. The agentic approach eliminates this trade-off by intelligently selecting which elements to analyze and when.Agentic video understanding is currently available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform for both video uploads and YouTube videos.
2
Google announced that the feature will roll out to the Gemini app soon, expanding access beyond API users.1
The company also plans to integrate agentic video understanding into YouTube's "Ask YouTube" feature, potentially making it easier for creators to analyze their videos with AI-driven video analysis tools. This integration could transform how content creators understand viewer engagement and optimize their productions across the platform.Summarized by
Navi
[1]
27 Jan 2026•Technology

19 Jun 2025•Technology

19 May 2026•Technology

1
Technology

2
Policy and Regulation

3
Health