Tag
This paper introduces Grounded Entity Biographies (GEB), a framework for augmenting long-video memory to track entities across events, demonstrating improved performance on question-answering benchmarks.
The paper proposes VideoLoop, a multimodal agent with looped working memory to address semantic thrashing in long-form video understanding, demonstrating performance gains on benchmarks like VideoMME.
TRACE introduces a condition-aware benchmark and evaluation framework for streaming video understanding, explicitly accounting for temporal validity, execution conditions, and multidimensional performance metrics beyond accuracy.
This paper introduces PlaylistEval, an agentic framework for benchmarking video-language judges on day-scale videos, revealing that frontier models achieve only 75.4% accuracy and highlighting the need for multi-modal approaches.
VideoMM introduces an adaptive macro-micro inference framework that reduces visual token overhead in video MLLMs, achieving a 6.13× speedup and 7.4% accuracy gain for efficient long-form video understanding.
The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.
The paper introduces BVB, a benchmark for agentic video understanding via programmatic reconstruction in Blender, evaluating models on perceptual similarity and spatiotemporal fact retention.
GPT-6 Astra has successfully replicated the landing gear retraction and extension mechanism of the Cessna 337 Skymaster from a YouTube video segment, overcoming limitations of previous models.
VLX-VR is an agentic-aware video reasoning model that uses a Think–Memory–Observation loop and reinforcement learning to adaptively gather evidence, achieving state-of-the-art performance on the MINERVA benchmark.
Google's update to Gemini 3.7 Flash introduces video token reduction by up to 88%, lowering costs significantly, though accuracy improvements are questionable.
This paper presents a controlled study on visual-token allocation for long-video multimodal language models, finding that frame selection significantly drives accuracy, while spatial compression is nearly free when savings are reinvested into more frames, highlighting the need for a unified comparison harness.
Gemini 3.7 Flash showcases a new agentic video understanding feature that accurately counts rapid movements like claps by dynamically adjusting processing speed.
The tweet highlights that agentic video understanding with Gemini 3.7 Flash significantly reduces costs, lowers token usage, and improves accuracy.
Google DeepMind introduces agentic video understanding for Gemini models, reducing token consumption by up to 88% and improving accuracy in video analysis.
Isaac 0.5 is an open-source AI model that teaches robots to learn movements by watching massive amounts of videos, significantly reducing training costs, improving efficiency by 210 times, and using MoE architecture.
Video-IFBench introduces a comprehensive benchmark for evaluating how well multimodal large language models follow diverse instructions in video understanding tasks, covering various constraints and task types.
Tencent introduces WeMM-Embedding, a series of universal multimodal embedding models in 9B, 4B, and 2B sizes, built on Qwen3.5, supporting text, images, videos, and visual documents for embedding generation.
MoTE introduces task-specific expert routing to replace dense decoder feed-forward networks in multi-task video understanding, improving accuracy and efficiency with interpretable, sparse computation.
This paper presents MoE-ViE, a mixture-of-experts vision encoder that efficiently scales for image and video understanding, outperforming larger dense models with lower inference latency.
StreamOPD enhances streaming video understanding through a post-training recipe using on-policy distillation and spatio-temporal cue gating, achieving significant benchmark improvements without requiring inference-time memory.