active-vision

Tag

Cards List
#active-vision

An Exam for Active Observers

arXiv cs.CL · yesterday Cached

This paper introduces ActiveVision, a benchmark to evaluate active observation in multimodal large language models. Frontier models like GPT-5.5 and Claude Fable 5 perform poorly, solving only 10.6% and 3.5% of tasks respectively, compared to human 96.1%, highlighting a lack of iterative visual perception.

0 favorites 0 likes
← Back to home

Submit Feedback