@ProfTomYeh: Kimi 3 seminar recording is uploaded http://byhand.ai/v/kimi3 ~ Prof. Tom Yeh
Summary
Prof. Tom Yeh uploads a seminar recording on Kimi 3, featuring special guest Nathan Lambert, with discussions on RLHF, model architecture, and frontier AI topics.
View Cached Full Text
Cached at: 08/22/26, 07:22 AM
Kimi 3 seminar recording is uploaded 👉 https://t.co/hxfYtPg0X7
~ Prof. Tom Yeh https://t.co/6CbTa3Wf5l
Kimi 3 ~ New Seminar Recording
Source: https://www.byhand.ai/p/library-videos-seminars-2026-kimi-3-seminar-recording
Library›Seminar Series 2026
- Manifold-Constrained Hyper Connections (mHC) from DeepSeek (Jan 9, 2026)
- How Small Models Learn Tool Use (Jan 11, 2026)
- Generative AI (Jan 11, 2026)
- Attention (Jan 15, 2026)
- Google Ironwood TPU: From Bits to HBM (Jan 19, 2026)
- 9 AI Eval Formulas You Must Know (Jan 24, 2026)
- Meta Superintelligence Labs vs Facebook AI Research (Jan 30, 2026)
- Transformer: Six Levels of Understanding (Feb 13, 2026)
- OpenClaw Seminar (Feb 23, 2026)
- PPO → DPO → GRPO → Rubrics (Mar 2, 2026)
- Gemma 4 (May 7, 2026)
- Qwen 3.6 (May 21, 2026)
- Kimi 3 (Aug 6, 2026)
Thank you to everyone who joined the Frontier AI Seminar on Kimi 3.
It has been a while since I gave the last seminar on Qwen. Happy to get together with many of you again to study frontier model math and architecture by hand in Excel!
We are blessed withNathan Lambert, the author ofInterconnects, some of the most authoritative writing on open models, as a special guest. Nathan also wrote theRLHF bookwhich has just beenpublished by Manning.
Here are some of my favorite moments:
- Nathan shared his behind-the-scenes story about his trip to Moonshot AI.
- Nathan’s famous dog Phoebe had a surprise appearance on camera.
- I got to finally talk about Linear Transformer line of frontier model advances.
- I got to mention LeBron James in my explanation of the attention mechanism.
Thanks for your feedback!
Before the summer, I did plan to talk about Linear Transformer line and wanted to use NVIDIA Nemotron as the backdrop. However, when Kimi 3 was released, I found an even stronger storyline, one where I can track the evolution much further from Linear Attention, DeltaNet, to the latest Kimi Delta Attention (KDA).
- Interview and Discussion - Nathan Lambert - Moonshot AI Tour - Open vs Closed Models - Audience Q&A
- Kimi 3 by Hand - Attention - Quadratic Attention - Inference and KV Cache - Local Attention - Linear Attention - DeltaNet - Kimi Delta Attention (KDA) - Convolution and Gating - Chunkwise Parallel Computation - Seq2Seq LSTM
Limited-time free: the full recording and the Excel workbook are open to everyone for now.
Discussion about this post
Ready for more?
Similar Articles
Kimi K3 Architecture Overview and Notes
Sebastian Raschka provides an architectural overview of the open-weight Kimi K3 model, highlighting its scaling from 48B to 2.8T parameters, new LatentMoE and attention residual components, removal of RoPE in favor of NoPE, and native multimodal support. The model emphasizes inference efficiency and matches frontier performance.
Kimi-K3 is published on HuggingFace
Moonshot AI has released Kimi-K3, a 2.8T-parameter mixture-of-experts model with 1M token context window, available on HuggingFace under a permissive license with commercial limitations.
On Kimi K3: Its Capabilities And Related Discontents (70 minute read)
Kimi K3 is a 2.8T parameter open model from Moonshot AI, showing strong benchmark performance but likely over-optimized and lagging behind top closed models by months. It is distilled from Claude and its release may precede an IPO.
How Kimi K3 Engineered Its Way to the Frontier [R]
Kimi K3 by Moonshot is an open-weight model ranking fourth among 580 models, featuring innovations like Kimi Delta Attention to reduce KV cache memory, Quantile Balancing for expert load balancing, and AgentENV for efficient RL training sandboxing.
Kimi K3 on HF Viewer!
Kimi K3, a new AI model, is now available on the Hugging Face Viewer for easy access and exploration.