Tag
This paper presents OpenVisTool, an open framework for synthesizing instructive visual tool-use trajectories, along with a dataset (OpenVisTool-42K) and benchmark. It shows that fine-tuning on causally grounded supervision improves visual tool-use performance across multiple model backbones.
This paper audits visual tool-use in multimodal LLMs via causal interventions, revealing that returned observations often lack causal effect despite aggregate accuracy gains. It identifies failure modes like 'Calling Without Looking' and 'Looking Without Planning', introducing the concept of the 'illusion of visual tool-use'.