Tag
The paper introduces Tri-PvP, a benchmark that exposes visual bias and asymmetric evidence-form preferences in omni-modal large language models, revealing deep-seated modality biases that are linearly decodable from early layers and resistant to surface mitigation.
Introduces C3PO, a benchmark of 3,404 samples for evaluating cross-modal composition and counterfactual reasoning in multimodal LLMs. It finds modality dominance causes most failures, with even the best model (Gemini-3.1-Pro) far below human accuracy.