@berryxia: Moonshot AI founder Yang Zhilin recently released a 40-minute video. Born in 1992, valedictorian of Tsinghua CS undergrad, PhD from CMU, co-author of Transformer-XL and XLNet, former researcher at Google Brain and Meta, he calmly deconstructs Kimi K2 in front of the camera...
Summary
Moonshot AI founder Yang Zhilin released a 40-minute video detailing the training process of the Kimi K2 model, which cost only $4.6 million. In an 8-model real-time programming competition, Kimi K2 took first place, defeating GPT-5.5 and others, demonstrating how a small team can overturn the traditional compute-stacking paradigm through architecture optimization.
Similar Articles
@FinanceYF5: One of the key technologies for Kimi to beat US models may trace back to founder Yang Zhilin’s doctoral thesis written ten years ago. He is 34, with a bachelor's from Tsinghua and a PhD from CMU, and during his PhD, he worked at Meta AI and Google Brain. That XLNet paper, cited over 10,000 times, later evolved into Kimi K2's trillion-parameter Mo…
The article points out that one key technology for Kimi defeating US models may originate from founder Yang Zhilin’s doctoral thesis ten years ago, mentioning the connection between XLNet and Kimi K2's trillion-parameter MoE architecture.
@gnotuy: We open sourced Kimi K2.6. The next frontier in test-time compute isn't bigger models. It's better organizations of int…
Moonshot AI has open sourced Kimi K2.6 and argues that the next frontier in test-time compute is better organization of intelligence rather than simply building bigger models.
@zhang_benita: https://x.com/zhang_benita/status/2078716535548600458
This article features an interview with Yang Zhilin, founder of Moonshot AI, discussing the challenges and vision of building foundation models and the AI assistant Kimi, reflecting on the past year of development.
Kimi K3, and what we can still learn from the pelican benchmark
Chinese AI lab Moonshot AI announced Kimi K3, a 2.8 trillion parameter open-weights model, claiming it is the first open 3T-class model and beating several leading models on benchmarks. The article also discusses the model's pricing and a fun pelican SVG benchmark test.
@starmexxx: a moonshot engineer leaked the benchmark anthropic, openai and xai all buried the same week: kimi k3 beat opus 5, gpt-5…
A leaked benchmark suggests Kimi K3 outperforms major AI models like Opus 5 and GPT-5.6 at a fraction of the cost, leading companies to remove comparison charts from their sites.