@ProfTomYeh: Kimi 3 seminar recording is uploaded http://byhand.ai/v/kimi3 ~ Prof. Tom Yeh

X AI KOLs Timeline Events

Summary

Prof. Tom Yeh uploads a seminar recording on Kimi 3, featuring special guest Nathan Lambert, with discussions on RLHF, model architecture, and frontier AI topics.

Kimi 3 seminar recording is uploaded 👉 https://t.co/hxfYtPg0X7 ~ Prof. Tom Yeh https://t.co/6CbTa3Wf5l
Original Article
View Cached Full Text

Cached at: 08/22/26, 07:22 AM

Kimi 3 seminar recording is uploaded 👉 https://t.co/hxfYtPg0X7

~ Prof. Tom Yeh https://t.co/6CbTa3Wf5l


Kimi 3 ~ New Seminar Recording

Source: https://www.byhand.ai/p/library-videos-seminars-2026-kimi-3-seminar-recording

LibrarySeminar Series 2026

  1. Manifold-Constrained Hyper Connections (mHC) from DeepSeek (Jan 9, 2026)
  2. How Small Models Learn Tool Use (Jan 11, 2026)
  3. Generative AI (Jan 11, 2026)
  4. Attention (Jan 15, 2026)
  5. Google Ironwood TPU: From Bits to HBM (Jan 19, 2026)
  6. 9 AI Eval Formulas You Must Know (Jan 24, 2026)
  7. Meta Superintelligence Labs vs Facebook AI Research (Jan 30, 2026)
  8. Transformer: Six Levels of Understanding (Feb 13, 2026)
  9. OpenClaw Seminar (Feb 23, 2026)
  10. PPO → DPO → GRPO → Rubrics (Mar 2, 2026)
  11. Gemma 4 (May 7, 2026)
  12. Qwen 3.6 (May 21, 2026)
  13. Kimi 3 (Aug 6, 2026)

Thank you to everyone who joined the Frontier AI Seminar on Kimi 3.

It has been a while since I gave the last seminar on Qwen. Happy to get together with many of you again to study frontier model math and architecture by hand in Excel!

We are blessed withNathan Lambert, the author ofInterconnects, some of the most authoritative writing on open models, as a special guest. Nathan also wrote theRLHF bookwhich has just beenpublished by Manning.

Here are some of my favorite moments:

  • Nathan shared his behind-the-scenes story about his trip to Moonshot AI.
  • Nathan’s famous dog Phoebe had a surprise appearance on camera.
  • I got to finally talk about Linear Transformer line of frontier model advances.
  • I got to mention LeBron James in my explanation of the attention mechanism.

Thanks for your feedback!

Before the summer, I did plan to talk about Linear Transformer line and wanted to use NVIDIA Nemotron as the backdrop. However, when Kimi 3 was released, I found an even stronger storyline, one where I can track the evolution much further from Linear Attention, DeltaNet, to the latest Kimi Delta Attention (KDA).

  1. Interview and Discussion - Nathan Lambert - Moonshot AI Tour - Open vs Closed Models - Audience Q&A
  2. Kimi 3 by Hand - Attention - Quadratic Attention - Inference and KV Cache - Local Attention - Linear Attention - DeltaNet - Kimi Delta Attention (KDA) - Convolution and Gating - Chunkwise Parallel Computation - Seq2Seq LSTM

Limited-time free: the full recording and the Excel workbook are open to everyone for now.

⬇ Excel Workbook

Discussion about this post

Ready for more?

Similar Articles

Kimi K3 Architecture Overview and Notes

Hacker News Top

Sebastian Raschka provides an architectural overview of the open-weight Kimi K3 model, highlighting its scaling from 48B to 2.8T parameters, new LatentMoE and attention residual components, removal of RoPE in favor of NoPE, and native multimodal support. The model emphasizes inference efficiency and matches frontier performance.

Kimi-K3 is published on HuggingFace

Reddit r/artificial

Moonshot AI has released Kimi-K3, a 2.8T-parameter mixture-of-experts model with 1M token context window, available on HuggingFace under a permissive license with commercial limitations.

How Kimi K3 Engineered Its Way to the Frontier [R]

Reddit r/MachineLearning

Kimi K3 by Moonshot is an open-weight model ranking fourth among 580 models, featuring innovations like Kimi Delta Attention to reduce KV cache memory, Quantile Balancing for expert load balancing, and AgentENV for efficient RL training sandboxing.

Kimi K3 on HF Viewer!

Reddit r/LocalLLaMA

Kimi K3, a new AI model, is now available on the Hugging Face Viewer for easy access and exploration.