---
title: "Gemini's new agent-based video analysis cuts token use up to 88%"
url: https://www.parallelquant.com/posts/gemini-s-new-agent-based-video-analysis-cuts-token-use-up-to-88-b27512
source_name: "The Decoder"
source_url: https://the-decoder.com/google-geminis-new-agent-based-video-analysis-cuts-token-usage-by-up-to-88-percent/
published: 2026-09-02T08:21:38.000Z
topics: ["llms", "products"]
publisher: "Parallel Quant"
---

# Gemini's new agent-based video analysis cuts token use up to 88%

*2026-09-02 · Source: [The Decoder](https://the-decoder.com/google-geminis-new-agent-based-video-analysis-cuts-token-usage-by-up-to-88-percent/)*

Google is rolling out agent-based video analysis to Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Instead of scanning video frame-by-frame at a fixed rate, the model decides which segments to examine and at what resolution, which Google says cuts token usage by up to 88% while improving accuracy on multi-hour footage.

**Why it matters:** Video has been the most token-expensive input for multimodal models, so letting the model control its own sampling density is a meaningful efficiency trick that could make long-form video analysis, like surveillance or lecture archives, economically viable at scale. It also reflects a broader industry shift toward giving models agentic control over their own inference process rather than just their outputs.

**Topics:** llms, products

---
Read the original: https://the-decoder.com/google-geminis-new-agent-based-video-analysis-cuts-token-usage-by-up-to-88-percent/
Canonical: https://www.parallelquant.com/posts/gemini-s-new-agent-based-video-analysis-cuts-token-use-up-to-88-b27512
Published by Parallel Quant — https://www.parallelquant.com
