Tag
Introduces ENTRAP-VL, a taxonomically structured dataset of 1,500 items to probe contextual entrainment in vision-language models, examining how textual and visual context independently influence model outputs.
This paper extends contextual entrainment from token-level to sentence-level, showing that even counterfactual sentences in prompts increase their probability during inference. The effect decreases with model size and is driven by 2-4% of attention heads, which can be ablated without performance loss.