speaker-attribution

Tag

Cards List
#speaker-attribution

Microsoft VibeVoice-ASR-Streaming Released

Reddit r/LocalLLaMA · 6h ago Cached

Microsoft has released VibeVoice-ASR-Streaming, a unified streaming ASR model that transcribes who said what with support for customized hotwords and 10 languages.

0 favorites 0 likes
#speaker-attribution

HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding

arXiv cs.CL · 2d ago Cached

This paper introduces HEAR, a benchmark for evaluating speaker-attributed reasoning in speech language models, and presents A2R, a 30B model optimized with counterfactual data to improve performance on multi-speaker tasks.

0 favorites 0 likes
#speaker-attribution

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

arXiv cs.CL · 2026-08-11 Cached

This paper analyzes multimodal systems for the CHiME-9 MCoRec cocktail-party scenario, comparing design strategies such as audio-visual target speech separation, improved recognition, and LLM-based conversational grouping, finding that speech overlap alone does not explain performance differences.

0 favorites 0 likes
← Back to home

Submit Feedback