Nemotron-3-Super-120B-A12B(混合Mamba+MoE)在4×3090上实现504K token的完美针检索

Reddit r/LocalLLaMA 模型

摘要

英伟达的Nemotron-3-Super-120B-A12B,一种混合Mamba和混合专家模型,仅使用四块RTX 3090 GPU就实现了在504K token下的完美大海捞针检索。

暂无内容
查看原文

相似文章

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4

Hugging Face Models Trending

NVIDIA 发布 Nemotron-3-Ultra,一个拥有 5500 亿参数的开源权重模型,采用结合 Mamba-2、MoE 和注意力的混合架构,支持高达 100 万 token 的上下文长度和可配置的推理模式。

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

Hugging Face Models Trending

NVIDIA released Nemotron 3.5 Lightning 30B-A3B-NVFP4, a hybrid MoE LLM with 3B active parameters, up to 1M context, and speculative decoding support for efficient single-GPU inference.