Tag
This paper introduces an instruction-free alignment-only method for building large audio-language models by freezing the LLM and audio encoder, training only a lightweight projector on self-generated data, achieving competitive performance with less data than traditional multi-stage pipelines.
A first neuron-level interpretability study of how large audio-language models encode multilingual emotion, introducing Consistency-Regularized Fusion to identify Multilingual Emotion Neurons across 12 languages and showing cross-lingual transfer benefits.
Introduces counterfactual audits to test whether audio language model judges actually use paralinguistic evidence when evaluating speech-to-speech responses, finding that contrastive success often overstates native reliability and similar accuracies can hide different failure modes across Gemini, GPT, and open models.
A paper introducing ILL, an inaudible low-frequency red-teaming method to expose safety vulnerabilities in audio-language models, and DRG, a defense that detects distribution shifts and requests a second recording to recover accuracy.
Introduces ESCUCHA, the first Spanish speech understanding benchmark for evaluating large audio language models across heterogeneous acoustic conditions and reasoning abilities, comprising 1,000 curated questions from diverse real-world sources.
Introduces a reinforcement learning with verifiable rewards recipe for data-efficient adaptation of audio-language models to code-switched ASR, achieving significant gains across 10 language pairs with minimal data.
VIBE is a framework that evaluates generative bias in Large Audio-Language Models using open-ended tasks with human-recorded speech, revealing systematic biases triggered by gender and accent cues.
This paper introduces Afrispeech Semantics, a benchmark for evaluating audio language models on semantic reasoning tasks including entailment, consistency, plausibility, accent drift, and accent restraint across diverse domains and accents.
KoALa-Bench introduces a Korean-focused benchmark suite for evaluating large audio language models on six tasks, including novel measures of speech faithfulness and Korea-specific cultural content.