Nemotron-3-Super-120B-A12B (hybrid Mamba+MoE) holds perfect needle retrieval to 504K tokens on 4×3090

Reddit r/LocalLLaMA Models

Summary

NVIDIA's Nemotron-3-Super-120B-A12B, a hybrid Mamba and mixture-of-experts model, achieves perfect needle-in-haystack retrieval at 504K tokens using only four RTX 3090 GPUs.

No content available
Original Article

Similar Articles

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B on 2x3090s

Reddit r/LocalLLaMA

A detailed guide on running the quantized NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B model on two RTX 3090s using vLLM with full 262K context, achieving high inference speeds without CPU offloading.

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4

Hugging Face Models Trending

NVIDIA releases Nemotron-3-Ultra, a 550B-parameter open-weight model with a hybrid architecture combining Mamba-2, MoE, and attention, supporting up to 1M token context and configurable reasoning mode.

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Hugging Face Models Trending

NVIDIA released Nemotron 3.5 Lightning 30B-A3B-NVFP4, a hybrid MoE LLM with 3B active parameters, up to 1M context, and speculative decoding support for efficient single-GPU inference.