video-understanding

Tag

Cards List
#video-understanding

M^3Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks

Hugging Face Daily Papers ↗ · 2026-06-03 Cached

M^3Eval is a comprehensive evaluation framework and benchmark for probing memory capabilities in multi-modal models, grounded in cognitive psychology. Experiments reveal consistent weaknesses in memory maintenance, interference patterns, and spatial-temporal grounding.

0 favorites 0 likes
#video-understanding

Benchmarking Visual State Tracking in Multimodal Video Understanding

Hugging Face Daily Papers ↗ · 2026-06-02 Cached

Introduces VSTAT, a benchmark for evaluating visual state tracking in multimodal large language models (MLLMs) using 834 clips and 1,500 questions. Current MLLMs perform poorly compared to humans, failing at visual perception rather than reasoning.

0 favorites 0 likes
#video-understanding

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

Hugging Face Daily Papers ↗ · 2026-06-01 Cached

X-Stream introduces the first benchmark for multi-stream video understanding, evaluating MLLMs as multiplexers across multiple concurrent streams. The study reveals that current MLLMs achieve only about 50% accuracy, exposing significant limitations in handling multiple streams.

0 favorites 0 likes
#video-understanding

Linear Scaling Video VLMs for Long Video Understanding

Hugging Face Daily Papers ↗ · 2026-05-29 Cached

StateKV is an inference-time method that enables linear-time video prefill for long-video vision-language models by carrying cross-frame context in a fixed-capacity recurrent state, maintaining accuracy close to full self-attention without fine-tuning.

0 favorites 0 likes
#video-understanding

EarlyTom: Early Token Compression Completes Fast Video Understanding

Hugging Face Daily Papers ↗ · 2026-05-28 Cached

EarlyTom is a training-free framework that compresses visual tokens early in the vision encoder to reduce time-to-first-token and computational costs while maintaining accuracy, achieving up to 2.65x TTFT reduction.

0 favorites 0 likes
#video-understanding

Kwai-Keye/Keye-VL-2.0-30B-A3B

Hugging Face Models Trending ↗ · 2026-05-25 Cached

Kwai-Keye releases Keye-VL-2.0-30B-A3B, a 30B-class vision-language model with advanced video understanding, sparse attention, and agent capabilities, achieving top benchmarks.

0 favorites 0 likes
#video-understanding

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Hugging Face Daily Papers ↗ · 2026-05-25 Cached

LLaVA-OneVision-2 introduces codec-stream tokenization and windowed attention for efficient video understanding, achieving state-of-the-art performance across multiple multimodal benchmarks including video, spatial, and tracking tasks.

0 favorites 0 likes
#video-understanding

MetaphorVU: Towards Metaphorical Video Understanding

Hugging Face Daily Papers ↗ · 2026-05-25 Cached

This paper introduces MetaphorVU-Bench, the first systematic benchmark for metaphorical video understanding, and proposes MetaphorBoost, an inference-time enhancement framework that improves cross-domain mapping in multimodal large language models.

0 favorites 0 likes
#video-understanding

Show HN: Lance – image/video generation and understanding in one model

Hacker News Top ↗ · 2026-05-20 Cached

ByteDance releases Lance, a 3B parameter unified multimodal model supporting image and video generation, understanding, and editing, trained from scratch with a multi-task recipe.

0 favorites 0 likes
#video-understanding

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

Hugging Face Daily Papers ↗ · 2026-05-20 Cached

Introduces Flat-Pack Bench, a benchmark for evaluating fine-grained spatio-temporal reasoning in large vision-language models using furniture assembly tasks. Experiments show current LVLMs struggle with tracking and spatial interactions.

0 favorites 0 likes
#video-understanding

@HappyyPablo: open sourcing Marlin-2B a tiny VLM to extract structured information from videos Marlin is finetuned for two questions …

X AI KOLs Timeline ↗ · 2026-05-19 Cached

Open-sourcing Marlin-2B, a tiny VLM for extracting structured information from videos, fine-tuned to answer 'what is happening and when'. Best open model in its weight class, competitive with Gemini-2.5-flash.

1 favorites 1 likes
#video-understanding

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

Hugging Face Daily Papers ↗ · 2026-05-19 Cached

ParaVT introduces the first multi-agent end-to-end RL framework for parallel video tool calling, addressing the Tool Prior Paradox with PARA-GRPO, and fully open-sources the paper, code, weights, and data.

0 favorites 0 likes
#video-understanding

@elonmusk: Grok groks videos

X AI KOLs Following ↗ · 2026-05-18 Cached

Grok now supports full video analysis, including summarization, translation, scene explanation, and context extraction, becoming natively multimodal with strong vision capabilities.

0 favorites 0 likes
#video-understanding

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

Hugging Face Daily Papers ↗ · 2026-05-18 Cached

SWIM is a novel training strategy that aligns vision and language representations for fine-grained object understanding using only textual prompts, leveraging mask supervision during training to improve cross-modal attention. It introduces the NL-Refer dataset and achieves superior performance over visual-prompt-based methods.

0 favorites 0 likes
#video-understanding

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Hugging Face Daily Papers ↗ · 2026-05-18 Cached

OmniPro is the first benchmark for evaluating proactive streaming video understanding in omni-modal large language models, featuring 2,700 samples covering diverse tasks and dual-mode evaluation protocols.

0 favorites 0 likes
#video-understanding

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs

Hugging Face Daily Papers ↗ · 2026-05-17 Cached

LiteFrame proposes a lightweight video encoder with Compressed Token Distillation training that reduces latency and enables processing 8x more frames for long-form video understanding in Video LLMs, improving accuracy while reducing compute.

0 favorites 0 likes
#video-understanding

bytedance-research/Lance

Hugging Face Models Trending ↗ · 2026-05-15 Cached

ByteDance Research introduces Lance, a 3B-parameter unified multimodal model trained from scratch on 128 A100 GPUs, capable of image and video understanding, generation, and editing within a single framework.

0 favorites 0 likes
#video-understanding

VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation

Hugging Face Daily Papers ↗ · 2026-05-15 Cached

VideoSeeker introduces a paradigm for instance-level video understanding that integrates agentic reasoning with visual prompts, achieving superior performance through automated data synthesis and reinforcement learning, outperforming GPT-4o and Gemini-2.5-Pro.

0 favorites 0 likes
#video-understanding

@VincentLogic: NVIDIA really went all out this time, directly releasing an open-source video understanding monster Nemotron 3 Nano Omni that processes video at an insane speed: 1 hour to handle 10 hours of video content, 10 times faster than playback speed. The core relies on 3D convolution technology, no longer scanning frame by frame, but instead…

X AI KOLs Timeline ↗ · 2026-05-14

NVIDIA has open-sourced the video understanding model Nemotron 3 Nano Omni, which uses 3D convolution technology and processes video 10 times faster than playback speed. It excels at audio-video analysis, surveillance retrieval, and asset tagging, but is not suitable for code or text inference tasks.

0 favorites 0 likes
#video-understanding

ViMU: Benchmarking Video Metaphorical Understanding

Hugging Face Daily Papers ↗ · 2026-05-14 Cached

ViMU is the first benchmark designed to evaluate video understanding models' ability to interpret metaphorical, ironic, and social meanings beyond literal visual comprehension, using hint-free open-ended and multiple-choice questions.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback