@natolambert: Thinky with a ~1T param, 41B active, apache-2 model Benchmarks are a clear step up from Nemotron Ultra (55B active), ne…
Summary
Thinky is a ~1T parameter mixture-of-experts model with 41B active parameters, released under Apache-2 license. It achieves new best results among American models, benchmark improvements over Nemotron Ultra, with omni-modal input.
View Cached Full Text
Cached at: 07/16/26, 10:19 PM
Thinky with a ~1T param, 41B active, apache-2 model Benchmarks are a clear step up from Nemotron Ultra (55B active), new best American model, and omni input. A bit behind GLM 5.2 on agentic benchies, and Kimi K 2.6 on multi modal Super exciting! Thank you @johnschulman2 & team https://t.co/NtkO7vmfSy
Similar Articles
Nemotron 3 Ultra. 550 billion parameters, 55B active. 1 million context
NVIDIA releases Nemotron 3 Ultra, a massive 550 billion parameter mixture-of-experts model with 55B active parameters and a 1 million token context window.
VibeThinker: 3B param model that beats Opus 4.5 on reasoning with novel SFT+GRPO
This technical report introduces VibeThinker-3B, a 3B parameter dense model that achieves frontier-level reasoning performance on benchmarks like AIME26 and LiveCodeBench, matching or exceeding much larger models such as DeepSeek V3.2 and GLM-5 through a combination of curriculum-based SFT, multi-domain RL, and offline self-distillation.
@heyshrutimishra: NVIDIA just dropped Nemotron 3.5 Lightning 30 billion parameters. Only 3 billion active. Built for the execution layer …
NVIDIA released Nemotron 3.5 Lightning, a 30B-parameter MoE model with only 3B active parameters, optimized for agent execution tasks. It claims faster, cheaper tool calls and agent execution while staying fully open-source under OpenMDW-1.1.
New open model from Tencent Hy: Hy3 (295B total 21B active - apache 2.0)
Tencent releases Hy3, a 295B-parameter Mixture-of-Experts model with 21B active parameters and Apache 2.0 license, achieving strong benchmark performance comparable to larger flagship models.
VibeThinker-3B: what is this witchcraft? Killing it at MathQA like it has ~30B parameters
VibeThinker-3B is a small 3B parameter model that achieves performance comparable to ~30B parameter models on the MathQA benchmark, demonstrating significant efficiency.