@FinanceYF5: One of the key technologies for Kimi to beat US models may trace back to founder Yang Zhilin’s doctoral thesis written ten years ago. He is 34, with a bachelor's from Tsinghua and a PhD from CMU, and during his PhD, he worked at Meta AI and Google Brain. That XLNet paper, cited over 10,000 times, later evolved into Kimi K2's trillion-parameter Mo…
Summary
The article points out that one key technology for Kimi defeating US models may originate from founder Yang Zhilin’s doctoral thesis ten years ago, mentioning the connection between XLNet and Kimi K2's trillion-parameter MoE architecture.
View Cached Full Text
Cached at: 07/21/26, 06:48 PM
One of the key technologies behind Kimi’s victory over American models may trace back to a doctoral thesis written by its founder, Yang Zhilin, a decade ago.
At 34, he holds a bachelor’s degree from Tsinghua University and a PhD from CMU, with stints at Meta AI and Google Brain during his doctoral studies. XLNet, which has been cited over 10,000 times, later evolved into the trillion-parameter MoE architecture of Kimi K2. https://t.co/QraSavtKQt
Similar Articles
@berryxia: Moonshot AI founder Yang Zhilin recently released a 40-minute video. Born in 1992, valedictorian of Tsinghua CS undergrad, PhD from CMU, co-author of Transformer-XL and XLNet, former researcher at Google Brain and Meta, he calmly deconstructs Kimi K2 in front of the camera...
Moonshot AI founder Yang Zhilin released a 40-minute video detailing the training process of the Kimi K2 model, which cost only $4.6 million. In an 8-model real-time programming competition, Kimi K2 took first place, defeating GPT-5.5 and others, demonstrating how a small team can overturn the traditional compute-stacking paradigm through architecture optimization.
@YRSM_Simon: This is big news! Kimi 2.6 is a generative-level model. In this age of overflowing LLM capabilities, speed will become the deciding factor in competition. Is the chip sector about to see another 'sector rotation'? 😅
Cerebras is now running Kimi K2.6, a trillion-parameter model, in enterprise trials at ~1,000 tokens/s, the fastest frontier model performance ever measured by Artificial Analysis.
Implementing Kimi K3 from scratch in PyTorch [P]
This paper details the architecture of the Kimi K3 model (featuring 2.8 trillion parameters and a 1 million token context) released by Moonshot AI, and demonstrates how to implement its seven core innovations from scratch using PyTorch, including KDA recurrent attention, gated MLA, and stable latent MoE.
@interjc: The market still needs a disruptive force; whether you use Kimi or not, the major companies' reset cycles are increasing.
Kimi releases the K3 model, featuring 2.8 trillion parameters, a million-token context window, and native multimodal capabilities. It leverages Kimi Delta Attention and Attention Residuals to enhance inference speed and training efficiency.
Kimi K2.7 Code: 1T MoE, $0.95/M tokens, MIT license, beats Opus 4.8 on MCP tool-calling
Moonshot AI 发布了专注于编程的开放式权重模型 Kimi K2.7 Code,拥有1万亿参数和384个专家,性能在MCP工具调用上超越Opus 4.8,成本仅为十分之一。