Tag
This article argues the AI race narrative has shifted: Western labs now quietly adopt open Chinese advances like DeepSeek's KV cache compression (437x smaller than V1), which has driven inference and cached-token pricing at Anthropic and OpenAI sharply down.
OpenAI disclosed a coordinated adversarial model-distillation campaign that sought to extract protected reasoning from its models, attributing a core cluster of the activity to individuals associated with Moonshot AI (developer of Kimi).
The paper introduces Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a multi-step flow-matching or diffusion teacher into an efficient one-step student while suppressing generation of forgotten training data, requiring only the teacher and forget-set samples without access to retained data or auxiliary classifiers.
Jevstiller is a tool that distills a larger AI model into a local model with a configurable disagreement bound, enabling faster inference while ensuring agreement with the original model on a specified percentage of requests.
Nvidia CEO Jensen Huang commented on AI model distillation, calling it "competition" and asserting that testing and distilling competitors' models is fair play — a notable stance amid industry debate over model distillation practices.
The article questions the threat posed by Chinese AI given their reliance on model distillation and inferior resources compared to US companies like OpenAI and Anthropic, suggesting they are not an immediate concern.
OpenAI reveals that 80–90% of their research focuses on future models like GPT-7 and GPT-8, which are then distilled into smaller, cheaper models for widespread use.
The article discusses concerns about US AI labs maintaining their lead amid competitors' mass model distillation and use of American training data suppliers.
A distilled student model of Qwen-Image-2.1, trained by Viggle using Distribution Matching Distillation, enabling text-to-image generation and image editing in 6 steps instead of 40, resulting in about 5× faster performance with competitive quality.
A 2B vision-language model distilled from a tool-using agent achieves first-place performance on long egocentric video question answering by pruning its multilingual embedding table to meet parameter limits, reaching 89% accuracy of the larger pipeline with only 1.1% of parameters.
US government agencies have accused six Chinese AI companies of stealing capabilities from leading US AI models through distillation attacks, recommending mitigations that could impact AI users.
This paper presents Wayu-Paxa-TTS-Edge, an 82M-parameter Thai TTS model trained on synthetic speech from a voice-cloning teacher, achieving high accuracy and prosody for on-device use without reference audio.
Research indicates that distilling AI models into architectures similar to the teacher model results in better performance compared to distilling into very different architectures.
Elon Musk admitted under oath that xAI was distilling OpenAI's models, then expressed outrage when OpenAI cut off access to the Cursor tool.
Elon Musk testified that xAI used OpenAI's models to train Grok through model distillation, highlighting the ethical and legal controversies surrounding such practices in the AI industry.
The paper introduces Matryoshka Language Model Suites, which nests multiple model sizes (500M, 1.5B, 3B) into a single architecture for joint training, allowing smaller models to benefit from distillation with the largest model.
Quantization-Aware Healing is a method that recovers compressed 4-bit language models by distilling directly from the original uncompressed model, offering faster and more stable performance than Quantization-Aware Training.
The author expresses concern about the ease of accessing uncensored AI models like Qwen 3.8 27B and questions the effectiveness of safety regulations in the face of open source advancements.
This paper presents SimpleOPD, a method for on-policy distillation from long-context reasoning teachers to short-context students, overcoming tokenizer mismatch and training instability to enhance mathematical reasoning capabilities.
A tweet shares links to information about distillations of the Qwen 3.8 AI model, with the poster noting it is not personally tested.