Tag
This paper studies how vision-language models shift their reliance between image and text sources when one modality is degraded, revealing task-dependent behavior and introducing a new conflict benchmark for evaluation.