visual-states

Tag

Cards List
#visual-states

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Hugging Face Daily Papers · 2026-07-29 Cached

This paper introduces See2Think, a unified evaluation framework to test whether multimodal models genuinely rely on intermediate visual states during reasoning. It finds that visual reasoning is highly model- and environment-dependent, with faithful rendering being the main bottleneck.

0 favorites 0 likes
← Back to home

Submit Feedback