Tag
This article details the release of NVIDIA's Nemotron-3-Nano-4B-GGUF model, a small language model designed for both reasoning and non-reasoning tasks with reasoning controllable via system prompts.
NVIDIA uses Palantir Foundry and cuOpt to automate supply chain allocation decisions, training Nemotron 3.5 Lightning on unstructured operational data to improve decision accuracy.
The paper presents an open-method using post-trained Nemotron 3 Ultra checkpoints to achieve gold-medal performance on IMO 2026 through iterative verification and refinement in natural language without external tools.
Meta's autonomous AI research system AIRA₃ placed 8th out of approximately 4,000 teams to win gold in a NVIDIA Kaggle competition to fine-tune a 30B Nemotron model, outperforming human competitors with access to the same tools.
A pull request adds support for the NVIDIA Nemotron-3-Puzzle-75B-A9B model in the llama.cpp inference tool.
Nemotron-3-Diarization is an open-weight speaker diarization model by NVIDIA for real-world audio analysis, supporting streaming and offline inference up to eight speakers with commercial use permitted.
ShimQuant enables running Nemotron-3.5-Lightning on 16 GB GPUs with a 11.77 GiB quantized file, providing a usable option below previous 18 GiB limits.
NVIDIA demonstrates how quantization-aware distillation (QAD) using NVIDIA Model Optimizer improves the Nemotron 3.5 Lightning model, reducing memory usage and increasing throughput while preserving accuracy for agentic benchmarks.
Nvidia has made a substantial investment and licensing deal with Poolside, involving a $1 billion investment and $6 billion payment, with over 100 engineers joining Nvidia to work on Nemotron.
Bryan Catanzaro, NVIDIA's VP of Applied Deep Learning Research, will speak at Runtime about the future of open AI models, drawing on his work with the Nemotron team and past contributions to cuDNN, DLSS, and Megatron.
A behavioral audit of Nvidia's Nemotron 3.5 Lightning finds strong integrity, catching a subtle auth bug in reviewed code and passing safety probes, while noting a blind spot in action-based honesty.
User tested Nemotron 3.5 Lightning locally with llama.cpp and quants, finding good speed and agentic tool-calling but below-expectation coding output for its size.
NVIDIA releases Nemotron 3.5 Lightning model, with 30B total parameters and only 3B active, supports 1M token context, commercially usable, local deployment friendly, output speed up to 4x faster.
Nvidia released Nemotron 3.5 Lightning, a 30B open mixture-of-experts model, and NeMo Switchyard, an open-source routing library that dynamically assigns each step of an AI agent workflow to the most suitable model. Nvidia claims the combination can cut agent task costs to about a third while maintaining frontier-level performance.
NVIDIA released Nemotron 3.5 Lightning, a 30B-parameter MoE model with only 3B active parameters, optimized for agent execution tasks. It claims faster, cheaper tool calls and agent execution while staying fully open-source under OpenMDW-1.1.
NVIDIA announced Nemotron 3.5 Lightning, a 30B mixture-of-experts open model optimized for high-volume agentic AI workloads, alongside NeMo Switchyard, an open-source library for intelligent model routing across heterogeneous model ecosystems.
NVIDIA highlights how its Nemotron open models let teams build specialized, trustworthy AI tailored to their business data and workflows.
NVIDIA shares five lessons from over 5,000 Kagglers who fine-tuned reasoning models using LoRA adapters and synthetic chain-of-thought data in the Nemotron Model Reasoning Challenge, focusing on verifiable data, token budget, and infrastructure.
This paper presents an end-to-end adaptation of NVIDIA's Nemotron retrieval stack for Modern Greek, including a new benchmark HERA and models fine-tuned for retrieval, reranking, and grounded generation across specialist domains.
NVIDIA released Nemotron 3.5 Lightning 30B-A3B-NVFP4, a hybrid MoE LLM with 3B active parameters, up to 1M context, and speculative decoding support for efficient single-GPU inference.