VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
Summary
VIBE is a framework that evaluates generative bias in Large Audio-Language Models using open-ended tasks with human-recorded speech, revealing systematic biases triggered by gender and accent cues.
View Cached Full Text
Cached at: 07/09/26, 07:52 AM
Paper page - VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
Source: https://huggingface.co/papers/2604.17248
Abstract
Large Audio-Language Models exhibit systematic generative biases in realistic scenarios when evaluated through open-ended tasks using human-recorded speech, with bias magnitude varying significantly by task and triggered by gender and accent cues.
Large Audio-Language Models(LALMs) are increasingly integrated into daily applications, yet theirgenerative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluatesgenerative biasthroughopen-ended taskssuch as personalized recommendations, usinghuman-recorded speech. Unlike MCQs, our method allowsstereotypical associationsto manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 12 state-of-the-art LALMs reveals systematic biases in realistic scenarios. Both gender and accent cues trigger statistically significantdistributional shifts, and bias magnitude is strongly task-dependent.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2604\.17248
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2604.17248 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2604.17248 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2604.17248 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
VibeVoice Technical Report
VibeVoice is a new model from Microsoft that synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer. It achieves superior fidelity and compression, supporting up to 90 minutes of audio with multiple speakers.
ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood
ChildVox presents a comprehensive benchmark for analyzing children's acoustic communication across developmental stages, integrating over 20 sub-tasks from 17 child-centered audio and speech datasets.
VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
Introduces VIABench, a comprehensive video benchmark for evaluating multimodal large language models in real-world visual assistance for blind and visually impaired individuals, covering 761 videos and 14,526 annotations across three tasks.
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
Researchers introduce Voice of India, a 536-hour closed benchmark of unscripted telephonic conversations across 15 Indian languages and 139 regional clusters, exposing geographic and demographic ASR performance disparities.
Introducing Real World VoiceEQ: Measuring the human quality of voice AI
Real World VoiceEQ is a new benchmark for evaluating the human quality of voice AI, based on over a million human ratings, assessing models across speech recognition, synthesis, and understanding in real-world conditions.