@HuggingApps: NVIDIA Nemotron just dropped an audio-native model that hears the world, not just words transcription, translation, sou…

X AI KOLs Following Models

Summary

NVIDIA released Nemotron, an audio-native model capable of transcription, translation, sound recognition, audio Q&A, TTS, and full speech-to-speech, with open weights in 2B and 30B sizes.

NVIDIA Nemotron just dropped an audio-native model that hears the world, not just words 🔥 transcription, translation, sound recognition & audio Q&A, TTS and full speech-to-speech, hears, thinks, and talks back natively open weights, 2B and 30B ▶️ https://t.co/jUv9HVSfr3 https://t.co/uyfA9D4wIR
Original Article
View Cached Full Text

Cached at: 07/21/26, 02:45 PM

NVIDIA Nemotron just dropped an audio-native model that hears the world, not just words 🔥

transcription, translation, sound recognition & audio Q&A, TTS and full speech-to-speech, hears, thinks, and talks back natively

open weights, 2B and 30B

▶️ https://t.co/jUv9HVSfr3 https://t.co/uyfA9D4wIR


Nemotron-Labs-Audex - a Hugging Face Space by nvidia

Source: https://huggingface.co/spaces/nvidia/Nemotron-Labs-Audex Fetching metadata from the HF Docker repository...

Similar Articles

@kwindla: https://x.com/kwindla/status/2062544580105359686

X AI KOLs Timeline

NVIDIA released Nemotron 3.5 ASR, an open-source multilingual speech-to-text model with the lowest latency tested, available in multilingual and English-only variants, ideal for voice agents and self-hosted deployments.

nvidia/Nemotron-Labs-Audex-30B-A3B · Hugging Face

Reddit r/LocalLLaMA

NVIDIA released Nemotron-Labs-Audex-30B-A3B, a unified audio-text LLM built on a 30B MoE backbone with 3B activated parameters, offering strong performance on audio understanding, speech recognition/translation, and generation while preserving text reasoning and alignment capabilities.

nvidia/nemotron-3.5-asr-streaming-0.6b

Hugging Face Models Trending

NVIDIA releases Nemotron 3.5 ASR, a 600M parameter multilingual streaming speech recognition model supporting 40 language-locales with a Cache-Aware FastConformer-RNNT architecture for low-latency transcription. The model supports configurable chunk sizes and is ready for commercial use under the OpenMDW-1.1 license.