Tag
Wired reports on the deep divide in Silicon Valley over Chinese open-weight AI models, with some startups opposing a ban while AI giants like Anthropic push for regulation due to concerns about IP theft and lack of safety guardrails.
Argues that accusations of knowledge distillation by China from US AI models are legally baseless and amount to propaganda to protect overvalued US AI companies.
An article simplifying the concept of LLM distillation for a political audience, explaining how smaller models learn from larger ones.
Suhail predicts that restrictions on open weight models will lead to a 10x increase in distillation research as a way to troll the government, potentially making AI more efficient.
Experts are skeptical that Chinese company Moonshot achieved Kimi K3's advanced capabilities by copying Anthropic's Fable model via distillation, citing insufficient time and technical limitations.
A claim suggests that a distilled AI model can outperform its larger original model, which is counterintuitive.
VCSD removes the need for external teachers, privileged answers, or visual evidence in on-policy self-distillation by using content-erased control images to produce contrastive signals, consistently outperforming existing methods on vision-language benchmarks.
The White House is divided on how to respond to China's rapid AI progress, with debates over stricter controls versus workable restrictions, particularly regarding distillation attacks on US models like Anthropic's Claude Fable 5.
The U.S. Treasury threatens sanctions against Chinese AI company Moonshot for allegedly distilling Anthropic's Fable model, escalating tensions over AI intellectual property theft and export control violations.
Session-Adaptive Orthogonal Distillation (SAOD) is a technique that compresses a 744-billion parameter model (1.5TB) to under 100GB, greatly reducing storage and inference costs.
Commentary criticizing Anthropic and OpenAI for claiming their closed models were hacked/distilled by Chinese model Kimi3, questioning the inconsistency with earlier security claims.
The former White House Science Advisor claimed that Kimi K3 was distilled from Anthropic's Fable model.
Moonshot AI allegedly distilled Anthropic's Fable model for its K3 using sophisticated methods to evade detection, raising concerns about large-scale industrial espionage in AI development.
MUX proposes a method for lossless continuous reasoning by distilling discrete reasoning steps into multiplexed latent tokens that encode a superposition of subwords, achieving higher bandwidth and enabling parallel exploration in language model reasoning tasks.
The article argues that AI companies' competitive moat, built on expensive model training, is easily undermined by distillation—replicating models through repeated API queries—as demonstrated by industry practices like xAI training Grok on OpenAI models and Anthropic accusing Chinese labs of mining Claude.
A quote from @enoreyes of @factoryAI suggests that distillation is unstoppable and that model labs may already recognize the shape of intelligence.
An analysis of output similarity suggests that Chinese AI frontier models like Kimi K3, DeepSeek V4, and GLM 5.2 show stylistic differences indicating meaningful independent development, challenging claims that their progress relies heavily on distillation from US models.
A paper proposing Trace-Based On-Policy Distillation (TOPD), a teacher-supervised framework for transferring reasoning abilities to masked diffusion language models without reward estimation, achieving comparable accuracy to RL-trained counterparts with significant compute speedup.
Microsoft open-sourced Resource2Skill, which automatically distills executable agent skills from human resources such as tutorial videos, articles, and code, enabling real-world applications like web pages, PPTs, and Excel.
Kimi K3 is a 2.8T parameter open model from Moonshot AI, showing strong benchmark performance but likely over-optimized and lagging behind top closed models by months. It is distilled from Claude and its release may precede an IPO.