Tag
Microsoft has released VibeVoice-ASR-Streaming, a unified streaming ASR model that transcribes who said what with support for customized hotwords and 10 languages.
This paper introduces HEAR, a benchmark for evaluating speaker-attributed reasoning in speech language models, and presents A2R, a 30B model optimized with counterfactual data to improve performance on multi-speaker tasks.
This paper analyzes multimodal systems for the CHiME-9 MCoRec cocktail-party scenario, comparing design strategies such as audio-visual target speech separation, improved recognition, and LLM-based conversational grouping, finding that speech overlap alone does not explain performance differences.