Releasing the model weights and technical report of Kimi K3 (2 minute read)

TLDR AI Models

Summary

Kimi Moonshot released Kimi K3, a 2.8-trillion-parameter multimodal model with a 1M context window and architectural innovations like Kimi Delta Attention and Attention Residuals, claiming significant efficiency gains and outperforming Claude Opus 4.8 and GPT-5.5 on internal benchmarks.

Moonshot has released the model weights for Kimi K3, along with a technical report. Kimi K3 is a 2.8T Mixture-of-Experts model with native visual understanding. It has a 1-million-token context window and a new model architecture that gives it 2.5x the intelligence per unit of compute. Alongside Kimi K3, Moonshot is opening up more of the stack behind it β€” high-performance attention kernels, a MoE communication library, and infrastructure for running agent environments at scale.
Original Article
View Cached Full Text

Cached at: 07/28/26, 06:22 PM

# Thread by @Kimi_Moonshot on Thread Reader App Source: [https://threadreaderapp.com/thread/2081760186235289764.html](https://threadreaderapp.com/thread/2081760186235289764.html) ## More from @Kimi\_Moonshot [![Kimi.ai Profile picture](https://pbs.twimg.com/profile_images/1910294000927645696/QseOV0uF_bigger.png)](https://threadreaderapp.com/user/Kimi_Moonshot) Jul 16 Introducing Kimi K3: Open Frontier Intelligence πŸ”Ή 2\.8 Trillion Parameters, 1 Million Context, Native Multimodal πŸ”Ή Kimi Delta Attention enables up to 6\.3x faster decoding in million\-token contexts πŸ”Ή Attention Residuals deliver ~25% higher training efficiency at <2% additional cost πŸ”Ή Built for long\-horizon agentic coding and self\-evolving workflows Kimi K3 is now live on on[Kimi\.com](http://kimi.com/), Kimi Work, Kimi Code, and the Kimi API\. Open Weights by July 27, 2026\. πŸ”— API:[platform\.kimi\.ai](http://platform.kimi.ai/) πŸ”— Tech blog:[kimi\.com/blog/kimi\-k3](http://kimi.com/blog/kimi-k3)[![Image](https://threadreaderapp.com/images/1px.png)](https://pbs.twimg.com/media/HNXu0kobMAAPljb.png) [![Image](https://threadreaderapp.com/images/1px.png)](https://pbs.twimg.com/media/HNXu2GWaYAAwH4w.png) K3 is built on Kimi Delta Attention \(KDA\) and Attention Residuals \(AttnRes\), two architectural updates designed to improve how information flows across sequence length and model depth\. We have also scaled up Mixture of Experts \(MoE\) sparsity, effectively activating 16 out of 896 experts when paired with a Stable LatentMoE framework\. Together with refined training and data recipes, these structural changes yield an approximate 2\.5Γ— improvement in overall scaling efficiency compared to K2, allowing the model to convert compute into intelligence more effectively\.[![Image](https://threadreaderapp.com/images/1px.png)](https://pbs.twimg.com/media/HNXu7q4akAABQAM.jpg) Internal knowledge work bench Beyond public benchmarks, Kimi K3 Max also shows consistent gains on our internal benchmarks, which are built from recurring patterns and challenges in real\-world user\-agent workflows\. It scores 75\.5 on Online Exp Bench, 73\.5 on DECK\-Bench, and 62\.6 on Finance\-Bench, outperforming Claude Opus 4\.8 \(max\) and GPT\-5\.5 \(xhigh\) across all three\. These results reflect broad improvements in Kimi K3's agentic knowledge work capabilities, enabling more capable and reliable performance in real\-world use cases\.[![Image](https://threadreaderapp.com/images/1px.png)](https://pbs.twimg.com/media/HNXvBPrbgAABDsU.jpg) Read 6 tweets [![Kimi.ai Profile picture](https://pbs.twimg.com/profile_images/1910294000927645696/QseOV0uF_bigger.png)](https://threadreaderapp.com/user/Kimi_Moonshot) May 14 Meet Kimi Web Bridge \- Kimi's browser extension\. Agent can now interact with websites like a human: search, scroll, click, type and complete tasks\. Supports Kimi Code CLI, Claude Code, Cursor, Codex, Hermes, and more\. Available now on and the Chrome Web Store\.[kimi\.com/features/webbr…](http://kimi.com/features/webbridge) ![Video Poster](https://pbs.twimg.com/amplify_video_thumb/2054909059720228865/img/CyAnLs_pDh7XpVH8.jpg) Search across multiple platforms at scale and auto\-fill results directly into your spreadsheet\. ![Video Poster](https://pbs.twimg.com/amplify_video_thumb/2054897363790217216/img/nbniEtocC-P--Ux9.jpg) With K2\.6's multimodal capability, your agent will open a website, navigate through it, and replicate it\. ![Video Poster](https://pbs.twimg.com/amplify_video_thumb/2054898469773692928/img/5y2IzclItQN1wex5.jpg) Read 6 tweets [![Kimi.ai Profile picture](https://pbs.twimg.com/profile_images/1910294000927645696/QseOV0uF_bigger.png)](https://threadreaderapp.com/user/Kimi_Moonshot) Apr 20 Meet Kimi K2\.6 agent \- Video hero section, WebGL shaders, real backends\. From one prompt\. πŸ”Ή Video hero sections \- cinematic aesthetic, auto\-composited πŸ”Ή WebGL shader animations \- native GLSL / WGSL, liquid metal, caustics, raymarching πŸ”Ή Motion design \- GSAP \+ Framer Motion πŸ”Ή Backend database: Kimi wires up auth \+ database \+ backend in one pass\. πŸ”Ή Website stack \- React 19 \+ TypeScript \+ Vite \+ Tailwind \+ shadcn/ui πŸ”Ή 3D w/ physically\-based lighting \- Three\.js \+ React Three Fiber ![Video Poster](https://pbs.twimg.com/amplify_video_thumb/2046263820600082432/img/xglCOQdVjPwjJgua.jpg) Video hero sections, built right in\. K2\.6 agent calls video generation APIs to create real cinematic footage for your hero, not stock placeholders\. Composited into the page, synced to scroll, with shader overlays\. ![Video Poster](https://pbs.twimg.com/amplify_video_thumb/2046264273031315456/img/IGRuKMtRfUcI94Ha.jpg) Speaks fluent WebGL shader\. Writes GLSL / WGSL directly \- fragment shaders, vertex shaders, noise, SDF, raymarching\. Prompt: "a liquid\-metal hero with soft caustics\." ![Video Poster](https://pbs.twimg.com/amplify_video_thumb/2046264325741105152/img/gvgQ8lpOt4aR7RHd.jpg) Read 7 tweets [![Kimi.ai Profile picture](https://pbs.twimg.com/profile_images/1910294000927645696/QseOV0uF_bigger.png)](https://threadreaderapp.com/user/Kimi_Moonshot) Mar 16 Introducing π‘¨π’•π’•π’†π’π’•π’Šπ’π’ π‘Ήπ’†π’”π’Šπ’…π’–π’‚π’π’”: Rethinking depth\-wise aggregation\. Residual connections have long relied on fixed, uniform accumulation\. Inspired by the duality of time and depth, we introduce Attention Residuals, replacing standard depth\-wise recurrence with learned, input\-dependent attention over preceding layers\. πŸ”Ή Enables networks to selectively retrieve past representations, naturally mitigating dilution and hidden\-state growth\. πŸ”Ή Introduces Block AttnRes, partitioning layers into compressed blocks to make cross\-layer attention practical at scale\. πŸ”Ή Serves as an efficient drop\-in replacement, demonstrating a 1\.25x compute advantage with negligible \(<2%\) inference latency overhead\. πŸ”Ή Validated on the Kimi Linear architecture \(48B total, 3B activated parameters\), delivering consistent downstream performance gains\. πŸ”—Full report: [github\.com/MoonshotAI/Att…](https://github.com/MoonshotAI/Attention-Residuals/blob/master/Attention_Residuals.pdf)[![Image](https://threadreaderapp.com/images/1px.png)](https://pbs.twimg.com/media/HDgCpkHb0AA0a7_.jpg) Scaling law experiments reveal a consistent 1\.25Γ— compute advantage across varying model sizes\.[![Image](https://threadreaderapp.com/images/1px.png)](https://pbs.twimg.com/media/HDgCt49a8AA14V_.png) Analysis of training dynamics demonstrates how AttnRes naturally mitigates hidden\-state magnitude growth and yields a more uniform gradient distribution across depth\.[![Image](https://threadreaderapp.com/images/1px.png)](https://pbs.twimg.com/media/HDgCxm0bUAA7KrD.jpg) Read 4 tweets [![Kimi.ai Profile picture](https://pbs.twimg.com/profile_images/1910294000927645696/QseOV0uF_bigger.png)](https://threadreaderapp.com/user/Kimi_Moonshot) Nov 28, 2025 Meet Kimi Agentic Slides\! Now with Nano Banana Pro 🍌 🎁 Thanksgiving Gift: 48H FREE & UNLIMITED ACCESS πŸ”Έ Agentic search \(Kimi K2\) πŸ”Έ Files β†’ Slides \(PDFs, images, docs\+\) πŸ”Έ Fully editable \+ PPTX export πŸ”Έ Designer\-level visuals \(infographics, illustrations\) Try now:[kimi\.com/slides](https://www.kimi.com/slides) Here's a quick guide\. πŸ‘‡ Research paper \-\> Presentation Ready Deck[![Image](https://threadreaderapp.com/images/1px.png)](https://pbs.twimg.com/media/G61dkYva0AARFcB.jpg) Read 5 tweets

Similar Articles

Kimi-K3 Technical Report [pdf]

Hacker News Top

MoonshotAI releases Kimi-K3, a 2.8T-parameter open-weight multimodal agentic model with a 1M-token context window, built on new Kimi Delta Attention and Attention Residuals architecture, achieving significant scaling improvements.

@levie: The k3 weights have arrived

X AI KOLs Timeline

Kimi.ai released the model weights and technical report for Kimi K3, a 2.8T parameter MoE model with native visual understanding and a 1M-token context window, claiming 2.5x intelligence per unit of compute.