@HuggingApps: NVIDIA Nemotron just dropped an audio-native model that hears the world, not just words transcription, translation, sou…
Summary
NVIDIA released Nemotron, an audio-native model capable of transcription, translation, sound recognition, audio Q&A, TTS, and full speech-to-speech, with open weights in 2B and 30B sizes.
View Cached Full Text
Cached at: 07/21/26, 02:45 PM
NVIDIA Nemotron just dropped an audio-native model that hears the world, not just words 🔥
transcription, translation, sound recognition & audio Q&A, TTS and full speech-to-speech, hears, thinks, and talks back natively
open weights, 2B and 30B
▶️ https://t.co/jUv9HVSfr3 https://t.co/uyfA9D4wIR
Nemotron-Labs-Audex - a Hugging Face Space by nvidia
Source: https://huggingface.co/spaces/nvidia/Nemotron-Labs-Audex Fetching metadata from the HF Docker repository...
Similar Articles
@kwindla: https://x.com/kwindla/status/2062544580105359686
NVIDIA released Nemotron 3.5 ASR, an open-source multilingual speech-to-text model with the lowest latency tested, available in multilingual and English-only variants, ideal for voice agents and self-hosted deployments.
@DataChaz: @NVIDIA just quietly dropped an incredibly impressive speech recognition model that completely changes the math for loc…
NVIDIA quietly released Nemotron-3.5-ASR, a lightweight 0.6B parameter open-source speech recognition model designed for real-time streaming with support for 40+ languages, low latency, and cache-aware architecture.
nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face
NVIDIA released Nemotron-Labs-Audex-30B-A3B, a unified audio-text LLM built on a 30B MoE backbone with 3B activated parameters, offering strong performance on audio understanding, speech recognition/translation, and generation while preserving text reasoning and alignment capabilities.
NVIDIA Launches Nemotron 3 Nano Omni Model, Unifying Vision, Audio and Language for up to 9x More Efficient AI Agents
NVIDIA announces Nemotron 3 Nano Omni, an open multimodal model that unifies vision, audio, and language processing to enable faster and more efficient AI agents, achieving up to 9x higher throughput compared to other open omni models.
nvidia/nemotron-3.5-asr-streaming-0.6b
NVIDIA releases Nemotron 3.5 ASR, a 600M parameter multilingual streaming speech recognition model supporting 40 language-locales with a Cache-Aware FastConformer-RNNT architecture for low-latency transcription. The model supports configurable chunk sizes and is ready for commercial use under the OpenMDW-1.1 license.