multi-speaker

Tag

Cards List
#multi-speaker

**Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**

Hugging Face Blog · 9h ago Cached

NVIDIA releases Nemotron 3 Diarization, an open-weight 100M-parameter model that achieves state-of-the-art speaker diarization with a 14.72% error rate, supporting real-time and offline processing for up to eight speakers.

0 favorites 0 likes
#multi-speaker

HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding

arXiv cs.CL · 2026-09-01 Cached

This paper introduces HEAR, a benchmark for evaluating speaker-attributed reasoning in speech language models, and presents A2R, a 30B model optimized with counterfactual data to improve performance on multi-speaker tasks.

0 favorites 0 likes
#multi-speaker

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Hugging Face Daily Papers · 2026-08-03 Cached

This paper introduces SwanTale, a unified multi-speaker speech and audio generation model supporting both zero-shot and instruct tasks, along with SwanData-Caption for data annotation and SwanVAE for high-quality multi-audio-modality generation.

0 favorites 0 likes
#multi-speaker

@MosiAI_Official: MOSS-Transcribe-Diarize-0.9B is now open source on @huggingface. Built with an end-to-end audio-to-structured-transcrip…

X AI KOLs Following · 2026-07-09 Cached

MOSS-Transcribe-Diarize-0.9B is an open-source end-to-end audio understanding model for long-form multi-speaker transcription, diarization, and timestamp generation, released by Mosi AI under Apache 2.0.

0 favorites 0 likes
#multi-speaker

Fish Audio S2 Technical Report

Papers with Code Trending · 2026-03-09 Cached

Fish Audio S2 is an open-source text-to-speech system featuring multi-speaker capabilities, multi-turn generation, and instruction-following control, backed by a production-ready inference engine with low latency.

0 favorites 0 likes
#multi-speaker

VibeVoice Technical Report

Papers with Code Trending · 2025-08-26 Cached

VibeVoice is a new model from Microsoft that synthesizes long-form multi-speaker speech using next-token diffusion and a highly efficient continuous speech tokenizer. It achieves superior fidelity and compression, supporting up to 90 minutes of audio with multiple speakers.

0 favorites 0 likes
← Back to home

Submit Feedback