@Chenyang_Lyu: Excited to publicly release LongSpeech, which will be presented at #ICASSP2026 ! Most Audio LLMs are at short audio but…

X AI KOLs Following Papers

Summary

Researchers release LongSpeech, a 100k-segment dataset of ~10-min clips to benchmark long-form audio understanding across 8 tasks, to be presented at ICASSP 2026.

Excited to publicly release LongSpeech, which will be presented at #ICASSP2026 ! Most Audio LLMs are at short audio but struggle with long-form recordings. Our new dataset features 100,000+ segments (~10 mins each) to benchmark long-speech understanding across 8 tasks,
Original Article

Similar Articles

FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following

arXiv cs.CL

This paper describes FBK's submission to the IWSLT 2026 Instruction Following shared task, developing SpeechLLMs for short-form and long-form speech instruction following, exploring segmentation methods and achieving robust long-form performance with fixed 30-second segmentation.

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track

arXiv cs.CL

This paper describes NAVER LABS Europe's submission to the IWSLT 2026 instruction-following short track, improving upon their previous winning system by using a new speech projector (SpeechMapper) trained solely on ASR data and augmenting training with a synthetic SQA dataset (fakACL). The resulting system ties for first place in the constrained track while using a weaker LLM backbone.

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

arXiv cs.AI

SpeechDx is a large-scale benchmark for clinical speech AI spanning 12 datasets and 27 tasks across diverse health conditions, structured by stages of speech production. It evaluates 12 state-of-the-art audio encoders and shows that current models do not generalize reliably across the clinical speech landscape.