intervention-framework

Tag

Cards List
#intervention-framework

When Vision Speaks for Sound

Hugging Face Daily Papers · 2026-05-13 Cached

This paper identifies that video-capable multimodal LLMs often appear to understand audio but actually rely on visual cues, a failure mode termed the audio-visual Clever Hans effect. It introduces Thud, an intervention-driven probing framework to diagnose this issue, and proposes an alignment recipe that improves audio-visual consistency by 28 percentage points.

0 favorites 0 likes
← Back to home

Submit Feedback