Tag
NVIDIA highlights how its Nemotron open models let teams build specialized, trustworthy AI tailored to their business data and workflows.
NVIDIA shares five lessons from over 5,000 Kagglers who fine-tuned reasoning models using LoRA adapters and synthetic chain-of-thought data in the Nemotron Model Reasoning Challenge, focusing on verifiable data, token budget, and infrastructure.
This paper presents an end-to-end adaptation of NVIDIA's Nemotron retrieval stack for Modern Greek, including a new benchmark HERA and models fine-tuned for retrieval, reranking, and grounded generation across specialist domains.
NVIDIA's Digital Marketing team built an AI-powered localization platform using NVIDIA Nemotron Speech, achieving ~70% reduction in translation turnaround time and ~25% cost savings, processing over 11 million words across 20,000 files.
This paper presents an engineering study adapting NVIDIA Nemotron 3.5 ASR Streaming 0.6B to Kikuyu, Dholuo, and Kalenjin, achieving 42.97% and 33.98% WER on internal sets for Kikuyu and Dholuo, respectively, through data-centric techniques including corpus auditing, normalization, and streaming evaluation.
User shares benchmark results running the 550B Nemotron Ultra model across two machines using RPC, achieving impressive throughput on older AMD MI50 and Nvidia P40 GPUs.
NVIDIA released Nemotron, an audio-native model capable of transcription, translation, sound recognition, audio Q&A, TTS, and full speech-to-speech, with open weights in 2B and 30B sizes.
NVIDIA releases Nemotron 3 Embed, a collection of open embedding models that top the RTEB leaderboard, featuring an 8B flagship model and efficient 1B variants for production-scale retrieval.
A detailed guide on running the quantized NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B model on two RTX 3090s using vLLM with full 262K context, achieving high inference speeds without CPU offloading.
LangChain shows a 7-minute tutorial by Partner Engineer Srimanth Tangedipalli on running Deep Agents Code inside a governed NemoClaw OpenShell Sandbox with NVIDIA Nemotron 3 Ultra via Baseten.
Successfully ran the 75B Nemotron Puzzle model locally on a 64GB M2 Max Mac, demonstrating large model inference on consumer hardware.
NVIDIA introduces TwoTower, a method that decouples context representation and denoising in diffusion language models, achieving 2.42x throughput while retaining 98.7% of autoregressive quality on a 30B MoE backbone.
Fuji Kanaeda announces departure from Nvidia after a year, highlighting contributions to synthetic data generation (NeMo Data Designer) and Nemotron LLM builds, praising the team's work on open-source AI.
NVIDIA discusses the importance of open and synthetic data for building robust AI agents, highlighting their Nemotron open datasets for training, reasoning, and tool-use.
NVIDIA Nemotron 3 Ultra achieves benchmark-leading performance with LangChain Deep Agents harness, offering higher accuracy at lower cost than closed models without retraining.
NVIDIA AI releases a 75B MoE model (9.3B active) compressed from Nemotron-3-Super-120B using the Iterative Puzzle framework, with 1M token context support.
NVIDIA releases Nemotron-Labs-3-Puzzle-75B-A9B, a compressed hybrid MoE LLM derived from Nemotron-3-Super, achieving approximately 2× higher server throughput and improved concurrency while maintaining strong accuracy across reasoning, coding, and long-context benchmarks.
NVIDIA released Nemotron-Labs-Audex-30B-A3B, a unified audio-text LLM built on a 30B MoE backbone with 3B activated parameters, offering strong performance on audio understanding, speech recognition/translation, and generation while preserving text reasoning and alignment capabilities.
NVIDIA highlights how open frontier models and AI infrastructure are driving AI research, as reflected in accepted papers at ICML 2026, with contributions spanning robotics, life sciences, and synthetic data.
OpenMed privacy-filter v2 using nemotron and MLX 8-bit achieves 755 tokens per second on a Mac, redacting 1,152 PII identifiers across 22 categories from a 13,000-token clinical file without data leaving the machine.