Tag
Meta has released Muse Voice Transcribe, a real-time speech-to-text model with a 3.1% word error rate and adaptive streaming capabilities, making it suitable for voice agents.
This paper introduces techniques for the second MLC-SLM Challenge, including random leading-silence cropping and synthetic data generation to enhance multilingual conversational speech tasks, achieving improved accuracy and reduced error rates.
Indic DiarBench is a multilingual joint diarization and ASR benchmark covering all 22 scheduled languages of India with 108 hours of human-corrected multi-speaker audio, capturing conversational nuances like code-mixing and speaker overlap.
Microsoft released vibevoice, a 7B model that transcribes up to an hour of audio in one shot with built-in speaker diarization and timestamps, supporting 50+ languages and running locally without API costs.
MOSS-TD, a speaker-aware ASR system, is optimized within the SGLang-Omni serving stack, enabling 38-minute multi-speaker audio to be transcribed in about 49 seconds on a single H100, with concurrent processing of 16 meetings.
This paper proposes a diagnostic framework to separate preprocessing pipeline instability from measurement method instability in LLM-based stance analysis of public discourse, finding that cross-method disagreement is larger and more systematic than pipeline effects, and that aggregate metrics can mask these instabilities.
This paper presents a system for the MLC-SLM 2026 Challenge that combines speaker diarization with fine-tuned Qwen-ASR using supervised full fine-tuning, LoRA on synthetic speech, and GRPO reinforcement learning to achieve a 17.97 tcpMER on the final evaluation set.
Meetily is a privacy-first open-source meeting note-taking tool that runs 100% locally with real-time transcription, speaker diarization, and auto-summary. Built with Rust + Tauri under the MIT license, it is ideal for industries with strict privacy requirements such as law and healthcare.
EdgeSpeak officially launched, a local-first, privacy-preserving accurate transcription tool, supporting semantic segmentation and timestamps, compatible with OpenAI Audio API, etc. It will later add speaker labeling and speech generation features.
Microsoft open-sourced the VibeVoice speech AI framework, which supports one-shot transcription of 60-minute long audio, multi-speaker diarization and timestamp labeling, and also provides multi-role TTS synthesis capabilities. It is based on Qwen2.5 and comes with a 0.5B lightweight real-time version. It has received 24.8k stars on GitHub.
MUSCAT is a new multilingual, scientific conversation benchmark dataset for evaluating ASR systems on challenging multilingual scenarios including code-switching, domain-specific vocabulary, and mixed language input. The dataset consists of bilingual discussions on scientific papers between speakers using different languages, with results showing current state-of-the-art systems struggle with these multilingual challenges.