Universal-3.5 Pro
Summary
Universal-3.5 Pro improves native code switching, diarization, and adds more languages, enhancing speech recognition capabilities.
Similar Articles
Supertone/supertonic-3
Supertonic 3 is a lightweight, open-weight text-to-speech model designed for fast on-device inference, expanding support to 31 languages with improved stability and expression tags.
STT That Can Challenge Dragon Professional on Windows
A new speech-to-text tool claims to rival Dragon Professional on Windows, offering a competitive alternative for voice recognition.
Gemini 3.5 Live Translate
Gemini 3.5 Live Translate is a new audio model for real-time speech-to-speech translation.
nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face
NVIDIA released Nemotron-Labs-Audex-30B-A3B, a unified audio-text LLM built on a 30B MoE backbone with 3B activated parameters, offering strong performance on audio understanding, speech recognition/translation, and generation while preserving text reasoning and alignment capabilities.
@multimodalart: UniSE: Unified Speech Enhancement high quality open source model for making an audio crisp & isolating speakers in mult…
UniSE is a unified, prompt-free, autoregressive speech enhancement model based on a decoder-only language model, supporting multiple tasks like speech restoration, target speaker extraction, and speech separation in a single model.