@FinanceYF5: NVIDIA releases Nemotron 3.5 Lightning with 30B total parameters, only 3B active, supports 1M token context, commercially usable. Local deployment friendly: BF16 and smaller NVFP4 versions, can run on a single H100 or DGX Spark...
Summary
NVIDIA releases Nemotron 3.5 Lightning model, with 30B total parameters and only 3B active, supports 1M token context, commercially usable, local deployment friendly, output speed up to 4x faster.
View Cached Full Text
Cached at: 08/12/26, 10:24 AM
NVIDIA releases Nemotron 3.5 Lightning
- 30B total parameters, only 3B active, supports 1M token context, commercially usable.
- Local deployment friendly: BF16 and a smaller NVFP4 version, runs on a single H100 or DGX Spark; RTX 5090 is also on the supported list.
- Official claims up to 4x higher output speed, 86% accuracy on PinchBench, and 30% faster than Qwen3.6 35B at the same accuracy. https://t.co/KbbCAWfuF5
Similar Articles
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
NVIDIA released Nemotron 3.5 Lightning 30B-A3B-NVFP4, a hybrid MoE LLM with 3B active parameters, up to 1M context, and speculative decoding support for efficient single-GPU inference.
@heyshrutimishra: NVIDIA just dropped Nemotron 3.5 Lightning 30 billion parameters. Only 3 billion active. Built for the execution layer …
NVIDIA released Nemotron 3.5 Lightning, a 30B-parameter MoE model with only 3B active parameters, optimized for agent execution tasks. It claims faster, cheaper tool calls and agent execution while staying fully open-source under OpenMDW-1.1.
@ctnzr: We've gone even farther: Nemotron 3 Super is 120B and pretrained on 25T tokens in NVFP4. Nemotron 3 Ultra is ~500B and …
NVIDIA announces Nemotron 3 Super (120B) and Nemotron 3 Ultra (~500B) models, pretrained on 25T tokens using NVFP4 precision, emphasizing accelerated computing and efficiency improvements.
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
NVIDIA releases Nemotron-3-Ultra, a 550B-parameter open-weight model with a hybrid architecture combining Mamba-2, MoE, and attention, supporting up to 1M token context and configurable reasoning mode.
@FinanceYF5: Source:
NVIDIA announces the upcoming release of Nemotron 3 Ultra this week.