Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts
Summary
The paper introduces Tri-PvP, a benchmark that exposes visual bias and asymmetric evidence-form preferences in omni-modal large language models, revealing deep-seated modality biases that are linearly decodable from early layers and resistant to surface mitigation.
View Cached Full Text
Cached at: 09/23/26, 03:33 PM
Paper page - Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts
Source: https://huggingface.co/papers/2609.06011
Abstract
Tri-PvP benchmark reveals visual bias and asymmetric evidence-form preferences in omni-modal language models, with early-layer decodable modality effects resistant to surface mitigation.
Omni-modal large language models(OLLMs) jointly process vision, audio, and text, yet their modality bias undercross-modal conflictremains underexplored. Existing benchmarks conflate two distinct forms of evidence within a single modality:perceptual signals(e.g., a photograph or recording of a dog) andpropositional signals(e.g., the declarative claim “this is a dog”), such that any measured modality bias is inherently confounded withevidence-form bias, precluding clean attribution to either source. To address this, we introduceTri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, where vision and audio each take perceptual or propositional form. Evaluating five OLLMs, we find robustvisual biasacross most models and evidence-type conditions. Crucially, we reveal a systematic asymmetry inevidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio. Further analyses vialayer-wise linear probingandcontrastive decodingreveal that modality bias is already linearly decodable from early representation layers and can only be partially mitigated, calling for mitigation strategies beyond surface-level interventions.
View arXiv pageView PDFGitHub2Add to collection
Get this paper in your agent:
hf papers read 2609\.06011
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.06011 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.06011 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.06011 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models
This paper investigates modality preference in omni-modal large language models (OLLMs), revealing a paradigm shift from text-dominance to visual preference. The authors introduce a conflict-based benchmark and layer-wise probing to diagnose cross-modal hallucinations using internal model signals.
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Introduces OPD-V, a visual on-policy self-distillation paradigm for multimodal large language models that leverages positive and negative teachers to exploit modality balance as privileged information, improving reasoning performance across benchmarks while reducing training cost.
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
This paper investigates why multimodal LLMs fail on vision-centric tasks when visual evidence conflicts with language priors, introducing the WhatIfVis benchmark and showing that models often encode visual evidence but cannot reliably control their reliance on it.
Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts
This paper examines how the order of conflicting evidence in multimodal large language models affects judgments, revealing cross-modal evidence noncommutativity where placing perceptual evidence later increases model reliance on it.
Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
The paper introduces multimodal contextual sycophancy in large language models, where external text overrides visual evidence, and proposes a diagnostic method using System-2 Visual Arbitration to improve performance.