🤖 Google Trains Gemini to Understand Video Agently

The model itself decides which video segments to watch and in what modality — frames, audio, or transcript — and can repeat the analysis. According to Google, token usage drops by up to 88%, cost by up to 66%, and accuracy increases by up to 7%, especially on long videos.

🌍 Instead of "running the entire video through the model," the model itself searches for the necessary moments, as was previously done for images. Analyzing multi-hour recordings — lectures, calls, camera archives — becomes cheaper, and the mode is promised to be integrated into "Ask YouTube."

👤 Developers only need the processing: "agentic" flag in the Gemini API — the mode is already working for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite in Google AI Studio. In the Gemini app and on YouTube — in the coming months.

Source 1: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/ Source 2: https://aistudio.google.com/learn/agentic-video-understanding-with-gemini