visual-state-tracking

Tag

Cards List
#visual-state-tracking

@ma_nanye: VSTAT highlights the substantial perceptual gap between humans and MLLMs, but it goes far beyond that. Its diverse task…

X AI KOLs Following · 2026-06-03 Cached

VSTAT is a new benchmark for visual state tracking in videos that reveals perceptual gaps between humans and multimodal LLMs.

0 favorites 0 likes
#visual-state-tracking

Benchmarking Visual State Tracking in Multimodal Video Understanding

Hugging Face Daily Papers · 2026-06-02 Cached

Introduces VSTAT, a benchmark for evaluating visual state tracking in multimodal large language models (MLLMs) using 834 clips and 1,500 questions. Current MLLMs perform poorly compared to humans, failing at visual perception rather than reasoning.

0 favorites 0 likes
← Back to home

Submit Feedback