speaker-diarization

Tag

Cards List
#speaker-diarization

@rohanpaul_ai: Love this, another huge release from Meta. Lunched Muse Voice Transcribe for real-time voice dictation, with the lowest…

X AI KOLs Timeline · 2026-09-01 Cached

Meta has released Muse Voice Transcribe, a real-time speech-to-text model with a 3.1% word error rate and adaptive streaming capabilities, making it suitable for voice agents.

0 favorites 0 likes
#speaker-diarization

Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge

arXiv cs.CL · 2026-08-17 Cached

This paper introduces techniques for the second MLC-SLM Challenge, including random leading-silence cropping and synthetic data generation to enhance multilingual conversational speech tasks, achieving improved accuracy and reduced error rates.

0 favorites 0 likes
#speaker-diarization

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

arXiv cs.CL · 2026-07-28 Cached

Indic DiarBench is a multilingual joint diarization and ASR benchmark covering all 22 scheduled languages of India with 108 hours of human-corrected multi-speaker audio, capturing conversational nuances like code-mixing and speaker overlap.

0 favorites 0 likes
#speaker-diarization

@oliviscusAI: microsoft just released a tool that transcribes a full hour of audio at once, tracking who spoke and when. it's called …

X AI KOLs Timeline · 2026-07-20 Cached

Microsoft released vibevoice, a 7B model that transcribes up to an hour of audio in one shot with built-in speaker diarization and timestamps, supporting 50+ languages and running locally without API costs.

0 favorites 0 likes
#speaker-diarization

@YichiZ03: https://x.com/YichiZ03/status/2078588932191895976

X AI KOLs Timeline · 2026-07-18 Cached

MOSS-TD, a speaker-aware ASR system, is optimized within the SGLang-Omni serving stack, enabling 38-minute multi-speaker audio to be transcribed in about 49 seconds on a single H100, with concurrent processing of 16 meetings.

0 favorites 0 likes
#speaker-diarization

Quantifying the Sources of Instability in LLM-Based Stance Analysis of Public Discourse

arXiv cs.CL · 2026-07-14 Cached

This paper proposes a diagnostic framework to separate preprocessing pipeline instability from measurement method instability in LLM-based stance analysis of public discourse, finding that cross-method disagreement is larger and more systematic than pipeline effects, and that aggregate metrics can mask these instabilities.

0 favorites 0 likes
#speaker-diarization

Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech

arXiv cs.CL · 2026-07-10 Cached

This paper presents a system for the MLC-SLM 2026 Challenge that combines speaker diarization with fine-tuned Qwen-ASR using supervised full fine-tuning, LoRA on synthetic speech, and GRPO reinforcement learning to achieve a 17.97 tcpMER on the final evaluation set.

0 favorites 0 likes
#speaker-diarization

@KKaWSB: Open-source meeting note-taking tool in the privacy-first track, stars reached 12.9k+: meetily — 100% local AI meeting assistant, real-time transcription + speaker diarization + auto-summary, never touches the cloud. A single-file desktop app built with Rust + Tauri, MIT license, supports …

X AI KOLs Timeline · 2026-07-05 Cached

Meetily is a privacy-first open-source meeting note-taking tool that runs 100% locally with real-time transcription, speaker diarization, and auto-summary. Built with Rust + Tauri under the MIT license, it is ideal for industries with strict privacy requirements such as law and healthcare.

0 favorites 0 likes
#speaker-diarization

@FeitengLi: Next week, after adding speaker labeling and speech generation, it won't be this cheap early bird price anymore.

X AI KOLs Timeline · 2026-07-03 Cached

EdgeSpeak officially launched, a local-first, privacy-preserving accurate transcription tool, supporting semantic segmentation and timestamps, compatible with OpenAI Audio API, etc. It will later add speaker labeling and speech generation features.

0 favorites 0 likes
#speaker-diarization

@uniswap12: Microsoft open-sourced a voice AI that can transcribe 60 minutes of long audio in one go, handling 4 people speaking simultaneously. VibeVoice, open-sourced by Microsoft, 24.8k stars, I only found out about it today. For converting recordings to text, I've been using Whisper, but it often times out on long meeting recordings and struggles with multi-speaker recognition...

X AI KOLs Timeline · 2026-06-04 Cached

Microsoft open-sourced the VibeVoice speech AI framework, which supports one-shot transcription of 60-minute long audio, multi-speaker diarization and timestamp labeling, and also provides multi-role TTS synthesis capabilities. It is based on Qwen2.5 and comes with a 0.5B lightweight real-time version. It has received 24.8k stars on GitHub.

0 favorites 0 likes
#speaker-diarization

MUSCAT: MUltilingual, SCientific ConversATion Benchmark

arXiv cs.CL · 2026-04-20 Cached

MUSCAT is a new multilingual, scientific conversation benchmark dataset for evaluating ASR systems on challenging multilingual scenarios including code-switching, domain-specific vocabulary, and mixed language input. The dataset consists of bilingual discussions on scientific papers between speakers using different languages, with results showing current state-of-the-art systems struggle with these multilingual challenges.

0 favorites 0 likes
← Back to home

Submit Feedback