active-observation

Tag

Cards List
#active-observation

An Exam for Active Observers

arXiv cs.CL · 2d ago Cached

This paper introduces ActiveVision, a benchmark to evaluate active observation in multimodal large language models. Frontier models like GPT-5.5 and Claude Fable 5 perform poorly, solving only 10.6% and 3.5% of tasks respectively, compared to human 96.1%, highlighting a lack of iterative visual perception.

0 favorites 0 likes
← Back to home

Submit Feedback