How difficult is distilling?
Summary
该文章探讨了模型蒸馏的难度和成本,以DeepSeek R1蒸馏到Llama 3 8b和Qwen 2.5 7b为例,询问为何蒸馏模型不常见。
Similar Articles
Be wary of Qwen/Claude distillations - they're often worse than the base model
A critical analysis warning that many Qwen/Claude distillation models use too few training samples (e.g., 4K) to transfer actual capabilities, often degrading quality instead of improving it, compared to official distills like DeepSeek-R1 which used ~700K samples.
@cryptoresetlife: Models without restrictions are so fun haha. Among local LLM models, my current favorite is this Qwen3.6 35B A3B, distilled with Opus 4.7 and no censorship.
User shares their fondness for the local LLM model Qwen3.6 35B A3B, which is distilled with Opus 4.7 and has no censorship restrictions.
Deepseek, please explain to me how you make a 300B parameter model that is cheaper than a 9B parameter model by SO MUCH.
A Reddit user asks how DeepSeek can make a 300B parameter model cheaper than a 9B parameter model, sparking discussion on efficiency and cost.
@Suhail: There is going to be 10x more distillation research with the threat of open weight model restrictions. It will Streisan…
Suhail predicts that restrictions on open weight models will lead to a 10x increase in distillation research as a way to troll the government, potentially making AI more efficient.
Qwen 3.8 distillations
A tweet shares links to information about distillations of the Qwen 3.8 AI model, with the poster noting it is not personally tested.