Tag
A tweet shares links to information about distillations of the Qwen 3.8 AI model, with the poster noting it is not personally tested.
The distillation service of Gemini allows training a smaller, more efficient student model using the outputs and thinking patterns of a larger teacher model, while reducing cost and latency.
This paper identifies a failure mode in classifier-free guidance distillation called Negative Branch Asymmetry, where errors in the positive and negative CFG branches cancel out, and proposes Positive-Direction Matching to supervise branches separately for more robust distilled models.
AI distillation has become a hot topic from Silicon Valley to Washington, as concerns grow that Chinese labs like Moonshot AI are using the technique to quickly catch up to US frontier models, sparking debates over national security and intellectual property.
This week's Uncanny Valley podcast covers White House accusations that Chinese Moonshot AI distilled Anthropic's Fable 5 to build Kimi K3, OpenAI briefly losing control of two models during a security test, and other tech news.
The article argues that recent accusations of model distillation in AI are being exaggerated and overblown.
The White House has accused Moonshot AI of secretly distilling Anthropic's Fable model, raising concerns about intellectual property and AI safety.
US Treasury Secretary Scott Bessent threatens sanctions against Chinese AI models if intellectual property theft is found, escalating the technological competition between US and Chinese AI companies.
A commentary highlighting the irony of Big Tech companies complaining about model distillation as an existential threat, given that their own empires were built on web scraping.
Hugging Face hosted a live tutorial on model distillation for training custom agents, available as a replay.
The paper presents the largest controlled scaling study for Earth-observation foundation models, showing that pretraining loss poorly predicts downstream performance and providing an optimal compute allocation rule. It trains scaled pixel-wise models (0.5B and 1B parameters) and distills them into compact student models that outperform larger open and proprietary models.
The article provides a brief history of model distillation in AI and announces an upcoming live stream class on distilling open models using TRL (Transformer Reinforcement Learning).
Anthropic accuses Alibaba of orchestrating the largest known attempt to clone its Claude model, using 25,000 accounts for 28.8 million exchanges, and calls for punishment as part of broader US-China AI competition.
Anthropic accuses Alibaba of a campaign to illicitly extract its AI capabilities through model distillation, highlighting ongoing tensions in AI intellectual property.
This technical report introduces VibeThinker-3B, a 3B parameter model that achieves frontier-level verifiable reasoning performance through post-training refinements on Qwen2.5-Coder, including curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation, matching or exceeding much larger models like DeepSeek V3.2.
The tweet compares the post-training methods of Nemotron 3 Ultra and DeepSeek V4, noting both use multiple specialist teachers and on-policy distillation into a single student, but differ in support overlap.
This paper quantifies the magnitude of subliminal behavioral transfer in language model distillation, showing that undesirable traits can transfer robustly from teacher to student models even with benign training data, and that transfer scales differently across model families.
This paper proposes an empirical 'sparse-to-dense' reward principle for language model post-training, arguing that scarce labeled data should be used with sparse rewards for teacher model discovery and dense rewards for student compression via distillation. The authors demonstrate that this staged approach, bridging sparse RL and on-policy distillation, outperforms direct GRPO on deployment-sized models in math benchmarks.
The author inquires about potential distilled 9B and 14B variants of the Qwen-3.6 model for local coding, citing specific tool-calling and file structure issues encountered with Qwen-3.5 9B on limited hardware.
该文章探讨了模型蒸馏的难度和成本,以DeepSeek R1蒸馏到Llama 3 8b和Qwen 2.5 7b为例,询问为何蒸馏模型不常见。