video-understanding

Tag

Cards List
#video-understanding

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

Hugging Face Daily Papers · 2026-08-06 Cached

StreamArena is a benchmark for hour-scale interactive streaming video understanding, paired with a two-tier architecture called StreamMind that outperforms existing streaming baselines across real-time perception, historical recall, proactive interaction, and tool use.

0 favorites 0 likes
#video-understanding

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Hugging Face Daily Papers · 2026-08-06 Cached

GST-Bench is a new VQA benchmark for evaluating global spatial awareness in video understanding, testing whether VLMs can build coherent global scene representations from long-horizon egocentric video. Evaluation of 22 state-of-the-art VLMs shows a large gap versus humans, with the best model scoring 42.68 vs 79.08.

0 favorites 0 likes
#video-understanding

Frame selection is the whole game: notes on making LLMs watch video

Hacker News Top · 2026-08-03 Cached

Engineering notes on optimizing frame selection for feeding video to LLMs, covering scene detection, deduplication strategies, and token budget management.

0 favorites 0 likes
#video-understanding

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

arXiv cs.AI · 2026-08-03 Cached

ViSAGE is a multimodal agentic memory framework for long-form video understanding that builds self-correcting, entity-centric memories via cross-modal binding, bidirectional memory refinement, and multi-agent cross-verification, achieving 5.9% higher accuracy than baselines.

0 favorites 0 likes
#video-understanding

@tom_doerr: VideoAgent is an all-in-one open-source framework for comprehensive video intelligence, combining understanding, editin…

X AI KOLs Timeline · 2026-08-02 Cached

VideoAgent is an all-in-one open-source framework for comprehensive video intelligence, combining understanding, editing, and creative generation through a unified agentic workflow.

0 favorites 0 likes
#video-understanding

@HuggingModels: Meet Mage-VL, a game changer for multimodal AI! It handles images, text, and even video understanding in one model. Per…

X AI KOLs Timeline · 2026-07-31 Cached

Mage-VL is introduced as a multimodal AI model that handles images, text, and video understanding in a single model, enabling richer interactive applications.

0 favorites 0 likes
#video-understanding

RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning

arXiv cs.CL · 2026-07-31 Cached

This paper introduces Reflective Retrieval Memory (RRM), a memory framework that distills procedural retrieval experience from historical task trajectories to improve evidence retrieval for long-horizon multimodal reasoning. RRM matches or exceeds prior state-of-the-art on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long benchmarks.

0 favorites 0 likes
#video-understanding

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

arXiv cs.AI · 2026-07-31 Cached

Introduces Pegasus, a low-resource framework that translates human demonstration videos into robot-executable data using graph-based task representation, hierarchical affordance latent space, and closed-loop physics verification, aiming to turn hardware data collection into scalable knowledge transfer.

0 favorites 0 likes
#video-understanding

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Google DeepMind Blog · 2026-07-30 Cached

Google DeepMind launches Gemini Robotics ER 2, an embodied reasoning model that enables robots to understand video, orchestrate tasks, and collaborate with multiple robots, now publicly available via the Gemini API.

0 favorites 0 likes
#video-understanding

microsoft/Mage-VL · Hugging Face - An Efficient Codec-Native Streaming Multimodal Foundation Model

Reddit r/LocalLLaMA · 2026-07-28 Cached

Microsoft introduces Mage-VL, a codec-native streaming multimodal foundation model for image and video understanding that achieves up to 3.5x inference speedup by using a sparsity pattern inspired by video codecs, cutting visual tokens by over 75%.

0 favorites 0 likes
#video-understanding

@HaiyuWu1: Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to a…

X AI KOLs Following · 2026-07-24 Cached

Induction Labs introduces imagination models, a new foundation model architecture that learns from internet-scale video. Their first model, Photon-1, learns to use a computer by watching 18 years of screen recordings without action labels, achieving better results at 30× lower pretraining cost than Gemini 3.1 Flash.

0 favorites 0 likes
#video-understanding

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

Hugging Face Daily Papers · 2026-07-19 Cached

TimeLens2 introduces a generalist video temporal grounding method using multimodal LLMs, treating temporal evidence as an interval set and achieving state-of-the-art performance across multiple benchmarks.

0 favorites 0 likes
#video-understanding

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Hugging Face Daily Papers · 2026-07-17 Cached

Audio-Visual Flamingo (AV-Flamingo) is a fully open audio-visual large language model designed for understanding and reasoning over long and complex videos, outperforming similarly sized open models and competitive with larger models.

0 favorites 0 likes
#video-understanding

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Papers with Code Trending · 2026-07-16 Cached

VideoChat3 is a fully open, efficient, and generalist video-centric multimodal large language model that introduces Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for streaming video perception, along with scalable video data synthesis pipelines, achieving superior performance with only 4B parameters.

0 favorites 0 likes
#video-understanding

@heyshrutimishra: BREAKING: Chinese cybersecurity company just found a way around one of AI's most expensive problems. Meta built an AI m…

X AI KOLs Timeline · 2026-07-14 Cached

360 AI Research presents MoSA, a method that learns object recognition from 10,000 hours of unlabeled video, generating 21 million+ self-labels without human intervention, solving AI's expensive labeling bottleneck.

0 favorites 0 likes
#video-understanding

OpenMOSS-Team/MOSS-VL-Realtime

Hugging Face Models Trending · 2026-07-14 Cached

MOSS-VL-Realtime is a realtime streaming vision-language model that processes continuous video frames, supports interruptible interaction, proactive silence, and dynamic correction, with timestamp-aware encoding and a 256K context window.

0 favorites 0 likes
#video-understanding

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory

Hugging Face Daily Papers · 2026-07-06 Cached

Light-Omni is a multimodal agent framework for efficient video understanding that uses dual contextual states (global state and parametric latent state) to avoid iterative reasoning, achieving faster and more accurate processing with significant speedup and memory savings.

0 favorites 0 likes
#video-understanding

Best way for agents to "watch" a video?

Reddit r/AI_Agents · 2026-07-03

The article discusses optimal methods for AI agents to process and understand video content, exploring various techniques for video analysis.

0 favorites 0 likes
#video-understanding

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

Hugging Face Daily Papers · 2026-07-03 Cached

This paper introduces PadCaptioner, a 3B parameter model for omni-modal dense video captioning that uses parallelized autoregressive decoding to achieve high efficiency and quality, outperforming 7B counterparts. A latent planning mechanism enables lossless parallel generation by exploiting weak local dependencies among events.

0 favorites 0 likes
#video-understanding

Video-Oasis: Rethinking Evaluation of Video Understanding

Hugging Face Daily Papers · 2026-07-02 Cached

Video-Oasis reveals that 55% of existing video benchmarks can be solved without visual input, exposing significant capability gaps in current video understanding models. State-of-the-art models perform only marginally above random guessing on the remaining video-native challenges.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback