Tag
This paper audits visual tool-use in multimodal LLMs via causal interventions, revealing that returned observations often lack causal effect despite aggregate accuracy gains. It identifies failure modes like 'Calling Without Looking' and 'Looking Without Planning', introducing the concept of the 'illusion of visual tool-use'.
This paper introduces a causal audit method to evaluate whether latent communication channels between LLM agents actually transmit task-relevant information, decomposing performance effects into message presence, content, and agent-specific contributions. Applied to Qwen3 models, it shows that aggregate accuracy alone cannot identify the causal role of latent messages.