@jxmnop: ok sorry everyone apparently they did distill lol. but only a tiny bit
Summary
Jack Morris corrects his earlier claim about an open-weight model being trained without distillation from OpenAI or Anthropic, acknowledging that it actually did use a small amount of distillation.
View Cached Full Text
Cached at: 07/16/26, 12:00 AM
ok sorry everyone apparently they did distill lol. but only a tiny bit
Jack Morris (@jxmnop): people are underestimating what a big deal this is
this is the ONLY open-weight model that’s trained without distilling from OpenAI or Anthropic
• Kimi distills • GLM distills • Qwen distills • Nemotron distills (Kimi & DeepSeek, which counts)
basically a fully different
Similar Articles
@Miles_Brundage: I am not sure I have seen a good analysis of how much distillation reduces this gap - people have very different views …
Miles Brundage comments on the lack of quantitative analysis on how distillation affects the capability gap between open-weight and proprietary AI models, referencing a claim by Epoch AI that open-weight models lag by four months.
Distilling The Moat (6 minute read)
The article argues that AI companies' competitive moat, built on expensive model training, is easily undermined by distillation—replicating models through repeated API queries—as demonstrated by industry practices like xAI training Grok on OpenAI models and Anthropic accusing Chinese labs of mining Claude.
@zhaisf: These were some magical results from distillation by @geoffreyhinton that really shocked me when I first saw them, and …
The article discusses surprising robustness of model distillation with respect to training distribution, even with little overlap with target distribution, and its implications for on/off-policy distillation.
@ben_burtenshaw: before model distillation was an attack vector. it was. pretty handy way of improving model performance on a task you c…
Ben Burtenshaw announces a live stream on July 7th covering knowledge distillation in post-training, showing how to implement it using small models to approach large model performance.
@SergioPaniego: before model distillation was an attack vector. it was. pretty handy way of improving model performance on a task you c…
The article provides a brief history of model distillation in AI and announces an upcoming live stream class on distilling open models using TRL (Transformer Reinforcement Learning).