Tag
This paper proposes a modality transfer task for Large Multimodal Models (LMMs) in GIS workflows and finds that current OpenAI LMMs struggle to transfer spatial information between image and text modalities, highlighting a critical bottleneck for autonomous GIS agents.
The paper introduces SeePhys Pro, a benchmark to diagnose modality transfer issues in multimodal RL for physics reasoning, revealing that models struggle with representation-invariant reasoning and often rely on residual textual cues rather than visual evidence.