Tag
This paper studies how irrelevant text context biases predictions in multimodal large language models, showing that context-induced decision margins follow an affine transformation of context-free margins, offering insights into model sensitivity.
An open evaluation setup with 55 LLMs blind-grading each other reveals statistically significant same-family rating bias across 8 model families, with Mistral penalizing its own models most severely. The study highlights issues with aggregate leaderboards and proposes improvements like within-response mixed-effects models.