I got Nemotron Puzzle 75B running smoothly on a 64GB M2 Max
Summary
Successfully ran the 75B Nemotron Puzzle model locally on a 64GB M2 Max Mac, demonstrating large model inference on consumer hardware.
Similar Articles
@NVIDIAAI: You're welcome
NVIDIA AI releases a 75B MoE model (9.3B active) compressed from Nemotron-3-Super-120B using the Iterative Puzzle framework, with 1M token context support.
Running local models on an M4 with 24GB memory
A guide on running local AI models like Qwen 3.5-9B on an M4 MacBook with 24GB RAM using tools like LM Studio, Ollama, and pi, including specific configuration tips for optimal performance.
2x 512gb ram M3 Ultra mac studios
A user shares their $25k hardware setup of two 512GB RAM M3 Ultra Mac Studios for running large language models locally, having tested DeepSeek V3 Q8 and GLM 5.1 Q4 via the exo distributed inference backend, while awaiting Kimi 2.6 MLX optimization.
nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4
NVIDIA releases Nemotron-Labs-3-Puzzle-75B-A9B, a compressed version of Nemotron-3-Super with improved inference efficiency and strong benchmark performance.
NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B on 2x3090s
A detailed guide on running the quantized NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B model on two RTX 3090s using vLLM with full 262K context, achieving high inference speeds without CPU offloading.