Nemotron-3-Super-120B-A12B (hybrid Mamba+MoE) holds perfect needle retrieval to 504K tokens on 4×3090
Summary
NVIDIA's Nemotron-3-Super-120B-A12B, a hybrid Mamba and mixture-of-experts model, achieves perfect needle-in-haystack retrieval at 504K tokens using only four RTX 3090 GPUs.
Similar Articles
NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B on 2x3090s
A detailed guide on running the quantized NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B model on two RTX 3090s using vLLM with full 262K context, achieving high inference speeds without CPU offloading.
@ctnzr: We've gone even farther: Nemotron 3 Super is 120B and pretrained on 25T tokens in NVFP4. Nemotron 3 Ultra is ~500B and …
NVIDIA announces Nemotron 3 Super (120B) and Nemotron 3 Ultra (~500B) models, pretrained on 25T tokens using NVFP4 precision, emphasizing accelerated computing and efficiency improvements.
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
NVIDIA releases Nemotron-3-Ultra, a 550B-parameter open-weight model with a hybrid architecture combining Mamba-2, MoE, and attention, supporting up to 1M token context and configurable reasoning mode.
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
NVIDIA released Nemotron 3.5 Lightning 30B-A3B-NVFP4, a hybrid MoE LLM with 3B active parameters, up to 1M context, and speculative decoding support for efficient single-GPU inference.
NVIDIA Puzzle-75B-A9B NVFP4 at 132 t/s on 3×3090 — Why is this size category a desert otherwise?
NVIDIA's Puzzle-75B-A9B model achieves 132 tokens per second using NVFP4 quantization on three RTX 3090 GPUs, raising discussion about the lack of competition in this model size category.