@Chenyang_Lyu: Excited to publicly release LongSpeech, which will be presented at #ICASSP2026 ! Most Audio LLMs are at short audio but…
Summary
Researchers release LongSpeech, a 100k-segment dataset of ~10-min clips to benchmark long-form audio understanding across 8 tasks, to be presented at ICASSP 2026.
Similar Articles
FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following
This paper describes FBK's submission to the IWSLT 2026 Instruction Following shared task, developing SpeechLLMs for short-form and long-form speech instruction following, exploring segmentation methods and achieving robust long-form performance with fixed 30-second segmentation.
Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
Swanbench-Speech is a comprehensive benchmark for evaluating long-form speech generation across diverse scenarios, using multi-dimensional metrics covering acoustics, semantics, and expressiveness, revealing limitations of current models.
NAVER LABS Europe Submission to the Instruction-following 2026 Short Track
This paper describes NAVER LABS Europe's submission to the IWSLT 2026 instruction-following short track, improving upon their previous winning system by using a new speech projector (SpeechMapper) trained solely on ASR data and augmenting training with a synthetic SQA dataset (fakACL). The resulting system ties for first place in the constrained track while using a weaker LLM backbone.
SpeechDx: A Multi-Task Benchmark for Clinical Speech AI
SpeechDx is a large-scale benchmark for clinical speech AI spanning 12 datasets and 27 tasks across diverse health conditions, structured by stages of speech production. It evaluates 12 state-of-the-art audio encoders and shows that current models do not generalize reliably across the clinical speech landscape.
Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge
This paper introduces techniques for the second MLC-SLM Challenge, including random leading-silence cropping and synthetic data generation to enhance multilingual conversational speech tasks, achieving improved accuracy and reduced error rates.