Tag
This paper identifies the 'question-first paradox' in vision-language models, where placing the question before the image hinders answer accuracy despite improving visual attention. It proposes a training-free technique called 'question echoing'—restating the question before and after the image—which improves performance across multiple benchmarks without architecture changes.