@eliebakouch: Kimi K3 (2.8T total parameters) is an open weight model competing with fable and gpt 5.6 sol while being much cheaper, …
Summary
Kimi K3 is an open-weight 2.8 trillion parameter LLM with innovations like linear attention, latent MoE, and new activation functions, offering competitive performance at lower cost.
View Cached Full Text
Cached at: 07/17/26, 12:24 AM
Kimi K3 (2.8T total parameters) is an open weight model competing with fable and gpt 5.6 sol while being much cheaper, this is just insane
i don’t think we realize how impressive it is to scale to this parameter count and ship a banger model. they scaled more than 2x from their previous model while keeping research bets such as linear attention (KDA), quantile load balancing, attention residual..
the arch uses (stable?) latent moe from nvidia with 16 experts out of 896 (same expert sparsity as DSv4), per head muon (like in glm 5), new activation function called “Sigmoid Tanh Unit (SiTU)”, QB load balancing by the goat @Jianlin_S
waiting to see the tech report (they mentioned they will release one ) more test time compute curve and test it my own research tasks, but overall this looks really good and this is only the beginning
the gap between K2.5 and K2.7 was impressive, it will likely be the same for K3 and K3.2!!
Kimi.ai (@Kimi_Moonshot): Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional
Similar Articles
unsloth/Kimi-K3
Kimi K3 is an open-weight 2.8T-parameter native multimodal agentic model with a novel architecture (KDA, AttnRes), 1M token context, and open-source MoE framework.
@svpino: I've been testing Kimi K3, and holy smokes, this is the best open-weight model the world has seen. It's a 2.8T-paramete…
Santiago Valdarrama reports that Kimi K3 is the best open-weight model he has tested, a 2.8T-parameter vision model with tool calling, reasoning, and a 1M context window.
Kimi K3 Architecture Overview and Notes
Sebastian Raschka provides an architectural overview of the open-weight Kimi K3 model, highlighting its scaling from 48B to 2.8T parameters, new LatentMoE and attention residual components, removal of RoPE in favor of NoPE, and native multimodal support. The model emphasizes inference efficiency and matches frontier performance.
How Kimi K3 Engineered Its Way to the Frontier [R]
Kimi K3 by Moonshot is an open-weight model ranking fourth among 580 models, featuring innovations like Kimi Delta Attention to reduce KV cache memory, Quantile Balancing for expert load balancing, and AgentENV for efficient RL training sandboxing.
On Kimi K3: Its Capabilities And Related Discontents (70 minute read)
Kimi K3 is a 2.8T parameter open model from Moonshot AI, showing strong benchmark performance but likely over-optimized and lagging behind top closed models by months. It is distilled from Claude and its release may precede an IPO.