Tag
The paper examines how vision-language models handle conflicts between in-context and parametric knowledge, revealing asymmetric biases where models favor in-context information for text entities but parametric information for image entities due to slower visual processing.