@googledevs: https://x.com/googledevs/status/2100299269261635948
Summary
Google's agentic video understanding capability is being utilized by three companies to reduce costs and enhance accuracy in video processing tasks using Gemini Flash models.
View Cached Full Text
Cached at: 09/16/26, 08:10 PM
How 3 Companies Are Building with Agentic Video Understanding
Standard video processing with AI often makes processing long footage cumbersome and misses crucial, split-second details. Earlier this month, we launched agentic video understanding to solve this problem. However, to truly put this capability to the test, we gave @mosaic_so, @PonderStudioAI, and @tryrevyl early access to our newest Gemini Flash models to see if it indeed led to fewer costs and increased accuracy.
Here’s what they built:
@Mosaic_so w/ Production Video Editing
For raw, multi-hour production video editing, @mosaic_so built an adaptive routing layer into their video agent @motion_so. Instead of flooding the prompt with every frame, their agent uses agentic video understanding to build an index, set targeted review windows, and pull only high-density frames where edit constraints require verification. This agentic approach cut median token usage by 97% and nearly doubled its ability to handle complex edits compared to static processing.
@PonderStudioAI w/ Selecting B-Roll
Instead of a complex perception loop, @PonderStudioAI uses a single agentic call to evaluate cinematic motion and find the best stable camera shots from raw B-roll. By using agentic video understanding, their work achieves a near-perfect 0.967 F1 score, while cutting token costs by ~72%.
@TryRevyl w/ Mobile Cloud Dev
Static screenshot testing (typically done at 1FPS) often misses dropped frames and UI stutters, but @tryrevyl solves this by wiring agentic video understanding directly into their assertion loop. Gemini helps their agent scan animations more closely when it matters, catching quick bugs and improving accuracy by 65%.
0:27
Get started
Agentic video understanding is available today for video uploads and @YouTube videos. It’s supported by Gemini 3.8 Flash, 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite via the Gemini API in @GoogleAIStudio and the Gemini Enterprise Agent Platform.
Read the developer guide.
Similar Articles
Introducing agentic video understanding with Gemini
Google DeepMind introduces agentic video understanding for Gemini models, reducing token consumption by up to 88% and improving accuracy in video analysis.
@Saboo_Shubham_: Agentic video understanding with Gemini 3.7 Flash is HUGE. Just reduced cost, less tokens, and higher accuracy. No bigg…
The tweet highlights that agentic video understanding with Gemini 3.7 Flash significantly reduces costs, lowers token usage, and improves accuracy.
Google's Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut (13 minute read)
Google releases Gemini 3.7 Flash, a new AI model focused on coding and agentic workflows, with a 50% introductory price cut on API tokens until end of 2026. The model improves planning, instruction following, and error recovery to reduce retries and manual oversight.
@googledevs: Building creative tools or media editing software? Gemini Omni 1.1 Flash is now available through the Gemini API in @Go…
Gemini Omni 1.1 Flash is now available through the Gemini API in Google AI Studio, providing production-ready control for generative video with features like scene extension, 4K upscaling, and faster prototyping.
With Gemini 3.5 Flash, Google bets its next AI wave on agents, not chatbots
Google launched Gemini 3.5 Flash, a new AI model optimized for coding and autonomous agents, shifting focus from chatbots to agentic AI. It outperforms previous models and powers new products like Antigravity 2.0 and Gemini Spark.