@ma_nanye: VSTAT highlights the substantial perceptual gap between humans and MLLMs, but it goes far beyond that. Its diverse task…

X AI KOLs Following Papers

Summary

VSTAT is a new benchmark for visual state tracking in videos that reveals perceptual gaps between humans and multimodal LLMs.

VSTAT highlights the substantial perceptual gap between humans and MLLMs, but it goes far beyond that. Its diverse tasks are designed not merely to assess simple pixel-space tracking, but to evaluate how well models capture and understand evolving world states in the latent space of videos. Text is only one way to probe this capability, and we are excited to see future evaluations explore new modalities such as pixels, actions, and beyond! Working on this benchmark has been a lot of fun along the way—huge shout-out to my amazing collaborators!
Original Article
View Cached Full Text

Cached at: 06/03/26, 11:56 PM

VSTAT highlights the substantial perceptual gap between humans and MLLMs, but it goes far beyond that. Its diverse tasks are designed not merely to assess simple pixel-space tracking, but to evaluate how well models capture and understand evolving world states in the latent space of videos. Text is only one way to probe this capability, and we are excited to see future evaluations explore new modalities such as pixels, actions, and beyond!

Working on this benchmark has been a lot of fun along the way—huge shout-out to my amazing collaborators!

Sihyun Yu (@sihyun_yu): Can MLLMs actually track what’s happening in a video? Introducing VSTAT 🎯, our new benchmark for visual state tracking.

The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don’t.

https://t.co/ZKqIDH5PcN

🧵 [1/11]

Similar Articles

Benchmarking Visual State Tracking in Multimodal Video Understanding

Hugging Face Daily Papers

Introduces VSTAT, a benchmark for evaluating visual state tracking in multimodal large language models (MLLMs) using 834 clips and 1,500 questions. Current MLLMs perform poorly compared to humans, failing at visual perception rather than reasoning.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Hugging Face Daily Papers

GST-Bench is a new VQA benchmark for evaluating global spatial awareness in video understanding, testing whether VLMs can build coherent global scene representations from long-horizon egocentric video. Evaluation of 22 state-of-the-art VLMs shows a large gap versus humans, with the best model scoring 42.68 vs 79.08.