video-understanding

Tag

Cards List
#video-understanding

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Hugging Face Daily Papers ↗ · yesterday Cached

This paper introduces Grounded Entity Biographies (GEB), a framework for augmenting long-video memory to track entities across events, demonstrating improved performance on question-answering benchmarks.

0 favorites 0 likes
#video-understanding

VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents

Hugging Face Daily Papers ↗ · yesterday Cached

The paper proposes VideoLoop, a multimodal agent with looped working memory to address semantic thrashing in long-form video understanding, demonstrating performance gains on benchmarks like VideoMME.

0 favorites 0 likes
#video-understanding

TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

arXiv cs.CL ↗ · 2d ago Cached

TRACE introduces a condition-aware benchmark and evaluation framework for streaming video understanding, explicitly accounting for temporal validity, execution conditions, and multidimensional performance metrics beyond accuracy.

0 favorites 0 likes
#video-understanding

PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?

Hugging Face Daily Papers ↗ · 2d ago Cached

This paper introduces PlaylistEval, an agentic framework for benchmarking video-language judges on day-scale videos, revealing that frontier models achieve only 75.4% accuracy and highlighting the need for multi-modal approaches.

0 favorites 0 likes
#video-understanding

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

arXiv cs.AI ↗ · 2026-09-16 Cached

VideoMM introduces an adaptive macro-micro inference framework that reduces visual token overhead in video MLLMs, achieving a 6.13× speedup and 7.4% accuracy gain for efficient long-form video understanding.

0 favorites 0 likes
#video-understanding

Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning

arXiv cs.LG ↗ · 2026-09-16 Cached

The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.

0 favorites 0 likes
#video-understanding

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

The paper introduces BVB, a benchmark for agentic video understanding via programmatic reconstruction in Blender, evaluating models on perceptual similarity and spatiotemporal fact retention.

0 favorites 0 likes
#video-understanding

@FinanceYF5: GPT-6 Astra has accomplished a mechanical structure that other models have always failed to accurately replicate: The l…

X AI KOLs Following ↗ · 2026-09-11

GPT-6 Astra has successfully replicated the landing gear retraction and extension mechanism of the Cessna 337 Skymaster from a YouTube video segment, overcoming limitations of previous models.

0 favorites 0 likes
#video-understanding

VLX-VR: An Agentic-Aware Video Reasoning Model

arXiv cs.CL ↗ · 2026-09-10 Cached

VLX-VR is an agentic-aware video reasoning model that uses a Think–Memory–Observation loop and reinforcement learning to adaptively gather evidence, achieving state-of-the-art performance on the MINERVA benchmark.

0 favorites 0 likes
#video-understanding

Gemini's 88% video token cut landed on 3.7 Flash, not 3.8

Reddit r/ArtificialInteligence ↗ · 2026-09-03

Google's update to Gemini 3.7 Flash introduces video token reduction by up to 88%, lowering costs significantly, though accuracy improvements are questionable.

0 favorites 0 likes
#video-understanding

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

This paper presents a controlled study on visual-token allocation for long-video multimodal language models, finding that frame selection significantly drives accuracy, while spatial compression is nearly free when savings are reinvested into more frames, highlighting the need for a unified comparison harness.

0 favorites 0 likes
#video-understanding

@googleaidevs: Can Gemini count the number of claps? Accurately counting rapid movements is a notoriously tricky task for AI. Because …

X AI KOLs Timeline ↗ · 2026-09-01 Cached

Gemini 3.7 Flash showcases a new agentic video understanding feature that accurately counts rapid movements like claps by dynamically adjusting processing speed.

0 favorites 0 likes
#video-understanding

@Saboo_Shubham_: Agentic video understanding with Gemini 3.7 Flash is HUGE. Just reduced cost, less tokens, and higher accuracy. No bigg…

X AI KOLs Following ↗ · 2026-09-01 Cached

The tweet highlights that agentic video understanding with Gemini 3.7 Flash significantly reduces costs, lowers token usage, and improves accuracy.

0 favorites 0 likes
#video-understanding

Introducing agentic video understanding with Gemini

Google DeepMind Blog ↗ · 2026-09-01 Cached

Google DeepMind introduces agentic video understanding for Gemini models, reducing token consumption by up to 88% and improving accuracy in video analysis.

0 favorites 0 likes
#video-understanding

This is amazing – it solves the problem of teaching robots human movements being too costly and slow. Previously, teaching robotic arms to grasp objects required thousands of hours of manual teleoperation, with high costs. Isaac allows robots to learn skills by watching massive amounts of ordinary online videos.

X AI KOLs Timeline ↗ · 2026-08-27 Cached

Isaac 0.5 is an open-source AI model that teaches robots to learn movements by watching massive amounts of videos, significantly reducing training costs, improving efficiency by 210 times, and using MoE architecture.

0 favorites 0 likes
#video-understanding

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Hugging Face Daily Papers ↗ · 2026-08-26 Cached

Video-IFBench introduces a comprehensive benchmark for evaluating how well multimodal large language models follow diverse instructions in video understanding tasks, covering various constraints and task types.

0 favorites 0 likes
#video-understanding

tencent/WeMM-Embedding 9B/4B/2B

Reddit r/LocalLLaMA ↗ · 2026-08-25

Tencent introduces WeMM-Embedding, a series of universal multimodal embedding models in 9B, 4B, and 2B sizes, built on Qwen3.5, supporting text, images, videos, and visual documents for embedding generation.

0 favorites 0 likes
#video-understanding

MoTE: Mixture of Task Experts for Multi-Task Video Understanding

Hugging Face Daily Papers ↗ · 2026-08-25 Cached

MoTE introduces task-specific expert routing to replace dense decoder feed-forward networks in multi-task video understanding, improving accuracy and efficiency with interpretable, sparse computation.

0 favorites 0 likes
#video-understanding

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

Hugging Face Daily Papers ↗ · 2026-08-18 Cached

This paper presents MoE-ViE, a mixture-of-experts vision encoder that efficiently scales for image and video understanding, outperforming larger dense models with lower inference latency.

0 favorites 0 likes
#video-understanding

StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding

Hugging Face Daily Papers ↗ · 2026-08-17 Cached

StreamOPD enhances streaming video understanding through a post-training recipe using on-policy distillation and spatio-temporal cue gating, achieving significant benchmark improvements without requiring inference-time memory.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback