@FinanceYF5: One of the key technologies for Kimi to beat US models may trace back to founder Yang Zhilin’s doctoral thesis written ten years ago. He is 34, with a bachelor's from Tsinghua and a PhD from CMU, and during his PhD, he worked at Meta AI and Google Brain. That XLNet paper, cited over 10,000 times, later evolved into Kimi K2's trillion-parameter Mo…

X AI KOLs Following News

Summary

The article points out that one key technology for Kimi defeating US models may originate from founder Yang Zhilin’s doctoral thesis ten years ago, mentioning the connection between XLNet and Kimi K2's trillion-parameter MoE architecture.

One of the key technologies for Kimi to beat US models may trace back to founder Yang Zhilin’s doctoral thesis written ten years ago. He is 34, with a bachelor's from Tsinghua and a PhD from CMU, and during his PhD, he worked at Meta AI and Google Brain. That XLNet paper, cited over 10,000 times, later evolved into Kimi K2's trillion-parameter MoE architecture. https://t.co/QraSavtKQt
Original Article
View Cached Full Text

Cached at: 07/21/26, 06:48 PM

One of the key technologies behind Kimi’s victory over American models may trace back to a doctoral thesis written by its founder, Yang Zhilin, a decade ago.

At 34, he holds a bachelor’s degree from Tsinghua University and a PhD from CMU, with stints at Meta AI and Google Brain during his doctoral studies. XLNet, which has been cited over 10,000 times, later evolved into the trillion-parameter MoE architecture of Kimi K2. https://t.co/QraSavtKQt

Similar Articles

@berryxia: Moonshot AI founder Yang Zhilin recently released a 40-minute video. Born in 1992, valedictorian of Tsinghua CS undergrad, PhD from CMU, co-author of Transformer-XL and XLNet, former researcher at Google Brain and Meta, he calmly deconstructs Kimi K2 in front of the camera...

X AI KOLs Timeline

Moonshot AI founder Yang Zhilin released a 40-minute video detailing the training process of the Kimi K2 model, which cost only $4.6 million. In an 8-model real-time programming competition, Kimi K2 took first place, defeating GPT-5.5 and others, demonstrating how a small team can overturn the traditional compute-stacking paradigm through architecture optimization.