model-distillation

Tag

Cards List
#model-distillation

The AI Race Just Got Awkward

Hacker News Top ↗ · 3d ago Cached

This article argues the AI race narrative has shifted: Western labs now quietly adopt open Chinese advances like DeepSeek's KV cache compression (437x smaller than V1), which has driven inference and cached-token pricing at Anthropic and OpenAI sharply down.

0 favorites 0 likes
#model-distillation

Disrupting a coordinated model-distillation campaign

OpenAI Blog ↗ · 3d ago Cached

OpenAI disclosed a coordinated adversarial model-distillation campaign that sought to extract protected reasoning from its models, attributing a core cluster of the activity to individuals associated with Moonshot AI (developer of Kimi).

0 favorites 0 likes
#model-distillation

Data Unlearning via Inverse Distillation

arXiv cs.LG ↗ · 3d ago Cached

The paper introduces Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a multi-step flow-matching or diffusion teacher into an efficient one-step student while suppressing generation of forgotten training data, requiring only the teacher and forget-set samples without access to retained data or auxiliary classifiers.

0 favorites 0 likes
#model-distillation

Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound

Hacker News Top ↗ · 4d ago Cached

Jevstiller is a tool that distills a larger AI model into a local model with a configurable disagreement bound, enabling faster inference while ensuring agreement with the original model on a specified percentage of requests.

0 favorites 0 likes
#model-distillation

@wallstengine: Nvidia CEO Jensen Huang on AI model distillation: “That’s called competition. You’re allowed to test somebody else’s pr…

X AI KOLs Following ↗ · 5d ago Cached

Nvidia CEO Jensen Huang commented on AI model distillation, calling it "competition" and asserting that testing and distilling competitors' models is fair play — a notable stance amid industry debate over model distillation practices.

0 favorites 0 likes
#model-distillation

Are the Chinese really that big of a threat, if their main source of innovation is distilling models?

Reddit r/artificial ↗ · 6d ago

The article questions the threat posed by Chinese AI given their reliance on model distillation and inferior resources compared to US companies like OpenAI and Anthropic, suggesting they are not an immediate concern.

0 favorites 0 likes
#model-distillation

OpenAI says 80–90% of research is aimed at GPT-7, GPT-8 and beyond, then distilled into cheap small models

Reddit r/singularity ↗ · 6d ago

OpenAI reveals that 80–90% of their research focuses on future models like GPT-7 and GPT-8, which are then distilled into smaller, cheaper models for widespread use.

0 favorites 0 likes
#model-distillation

With the new Opus/Fable models out, how will US labs keep their lead if competitors can train on them?

Reddit r/ArtificialInteligence ↗ · 2026-09-23

The article discusses concerns about US AI labs maintaining their lead amid competitors' mass model distillation and use of American training data suppliers.

0 favorites 0 likes
#model-distillation

Viggle/Qwen-Image-2.1-viggle-turbo

Hugging Face Models Trending ↗ · 2026-09-22 Cached

A distilled student model of Qwen-Image-2.1, trained by Viggle using Distribution Matching Distillation, enabling text-to-image generation and image editing in 6 steps instead of 40, resulting in about 5× faster performance with competitive quality.

0 favorites 0 likes
#model-distillation

Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model

Hugging Face Daily Papers ↗ · 2026-09-10 Cached

A 2B vision-language model distilled from a tool-using agent achieves first-place performance on long egocentric video question answering by pruning its multilingual embedding table to meet parameter limits, reaching 89% accuracy of the larger pipeline with only 1.1% of parameters.

0 favorites 0 likes
#model-distillation

Six Chinese AI firms accused of aggressively copying US frontier models

Ars Technica ↗ · 2026-09-09 Cached

US government agencies have accused six Chinese AI companies of stealing capabilities from leading US AI models through distillation attacks, recommending mitigations that could impact AI users.

0 favorites 0 likes
#model-distillation

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

This paper presents Wayu-Paxa-TTS-Edge, an 82M-parameter Thai TTS model trained on synthetic speech from a voice-cloning teacher, achieving high accuracy and prosody for on-device use without reference audio.

0 favorites 0 likes
#model-distillation

@seclink: We have found that distilling the model to a target model with a highly similar architecture yields performance far sup…

X AI KOLs Following ↗ · 2026-09-01 Cached

Research indicates that distilling AI models into architectures similar to the teacher model results in better performance compared to distilling into very different architectures.

0 favorites 0 likes
#model-distillation

@VraserX: Elon Musk admitting under oath that xAI was distilling OpenAI’s models, then acting outraged when OpenAI cuts off a Mus…

X AI KOLs Following ↗ · 2026-08-29 Cached

Elon Musk admitted under oath that xAI was distilling OpenAI's models, then expressed outrage when OpenAI cut off access to the Cursor tool.

0 favorites 0 likes
#model-distillation

@VraserX: https://x.com/VraserX/status/2093563301330346314

X AI KOLs Timeline ↗ · 2026-08-29 Cached

Elon Musk testified that xAI used OpenAI's models to train Grok through model distillation, highlighting the ethical and legal controversies surrounding such practices in the AI industry.

0 favorites 0 likes
#model-distillation

@willdepue: great paper, don’t buy the width mixing (i think fixed width probably better) and feel like they should be able to get …

X AI KOLs Following ↗ · 2026-08-21 Cached

The paper introduces Matryoshka Language Model Suites, which nests multiple model sizes (500M, 1.5B, 3B) into a single architecture for joint training, allowing smaller models to benefit from distillation with the largest model.

0 favorites 0 likes
#model-distillation

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

Hugging Face Daily Papers ↗ · 2026-08-21 Cached

Quantization-Aware Healing is a method that recovers compressed 4-bit language models by distilling directly from the original uncompressed model, offering faster and more stable performance than Quantization-Aware Training.

0 favorites 0 likes
#model-distillation

@AlexFinn: This is scary. I downloaded an uncensored version of Qwen 3.8 27B onto my Mac It literally does anything you want. Firs…

X AI KOLs Timeline ↗ · 2026-08-20 Cached

The author expresses concern about the ease of accessing uncensored AI models like Qwen 3.8 27B and questions the effectiveness of safety regulations in the face of open source advancements.

0 favorites 0 likes
#model-distillation

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

arXiv cs.CL ↗ · 2026-08-17 Cached

This paper presents SimpleOPD, a method for on-policy distillation from long-context reasoning teachers to short-context students, overcoming tokenizer mismatch and training instability to enhance mathematical reasoning capabilities.

0 favorites 0 likes
#model-distillation

Qwen 3.8 distillations

Reddit r/LocalLLaMA ↗ · 2026-08-16

A tweet shares links to information about distillations of the Qwen 3.8 AI model, with the poster noting it is not personally tested.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback