distillation

Tag

Cards List
#distillation

MiMo-V2.6 distilled themselves into Qwen 9B!

Reddit r/LocalLLaMA · yesterday

MiMo-V2.6 has been distilled into Qwen 9B, creating a more efficient version of the Qwen model released on Hugging Face.

0 favorites 0 likes
#distillation

AI doesnt have to fail for the bubble to burst

Reddit r/ArtificialInteligence · yesterday

The article argues that the AI bubble could burst without AI failure due to competition from cheaper models and cost-reduction techniques like distillation, which may erode profits from large industry spending.

0 favorites 0 likes
#distillation

GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation

arXiv cs.AI · 2d ago Cached

The paper introduces GUARD, a method for natural forgetting in large reasoning models that uses guided answer-reasoning distillation to suppress unsafe or private content in chain-of-thought traces while preserving reasoning utility.

0 favorites 0 likes
#distillation

ACLArena: Agent Continue Learning in Multi-stage Post-training

Hugging Face Daily Papers · 2d ago Cached

The paper presents ACLArena, a framework for evaluating Agent Continual Learning in multi-stage post-training, analyzing forgetting and generalization mechanisms, and proposing an improved ACL recipe using offline replay and LoRA experts.

0 favorites 0 likes
#distillation

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

Hugging Face Daily Papers · 2d ago Cached

The paper introduces THAW-VLA, a method that distills world-model representations into Vision-Language-Action models for robotics, enhancing robustness and performance on simulation and real hardware without additional inference overhead.

0 favorites 0 likes
#distillation

Anthropic just named seven Chinese AI labs for stealing Claude's reasoning. I read the full report and the Qwen part doesn't hold up the way the headline does.

Reddit r/LocalLLaMA · 4d ago

The article critiques Anthropic's report accusing seven Chinese AI labs of illicitly distilling Claude's reasoning, highlighting weak evidence for Alibaba's Qwen and suggesting political framing in the accusations.

0 favorites 0 likes
#distillation

@EmperoAI: Qwen3.8-35B-A3B-Distill - our first MoE. Qwen3.8 reasoning distilled into Qwen3.6-35B-A3B: 35B total, ~3B active. ARC-C…

X AI KOLs Following · 6d ago Cached

Empero AI releases their first MoE model, Qwen3.8-35B-A3B-Distill, which distills reasoning from Qwen3.8 into a smaller, efficient architecture, showing improved benchmarks on ARC-Challenge.

0 favorites 0 likes
#distillation

Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning

arXiv cs.LG · 2026-09-16 Cached

The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.

0 favorites 0 likes
#distillation

Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation

arXiv cs.CL · 2026-09-16 Cached

This paper introduces MIFS, a pipeline for synthesizing RL-ready multimodal data to enhance instruction following in MLLMs, achieving an 8.13% average improvement on benchmarks and faster training convergence.

0 favorites 0 likes
#distillation

Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control

Hugging Face Daily Papers · 2026-09-15 Cached

The paper introduces Zing-0.5, a 5B parameter autoregressive world model for generating playable worlds with real-time user interaction through combined keyboard and text controls. It achieves high performance in navigation tasks and demonstrates low-cost real-time inference at 24 FPS.

0 favorites 0 likes
#distillation

@Flopsie4: Hopefully in future work it's also explore how these byte models respond to quantization. If they respond the same or e…

X AI KOLs Following · 2026-09-14 Cached

A discussion on future work exploring how byte-level models respond to quantization for potential improvements in local AI, based on a Meta paper showing byte models outperforming token models as compute scales.

0 favorites 0 likes
#distillation

@MindsAI_Jack: Step in the right direction. As far as I know, pioneered at Google originally with ByT5 that was more resilient against…

X AI KOLs Following · 2026-09-14 Cached

A research paper from Meta demonstrates that byte-level language models, initially inferior to token-based models, can outperform them as computational resources increase, shown through distilled 1B models trained on up to 1 trillion bytes.

0 favorites 0 likes
#distillation

@galoisextn: Holy shit there’s only bots out here spreading slop, not one comment about the graphic

X AI KOLs Following · 2026-09-14 Cached

A paper from Meta shows that byte-level models start behind token models but surpass them with increasing compute, demonstrated with distilled 1B models trained on up to 1 trillion bytes.

0 favorites 0 likes
#distillation

UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed while keeping the accuracy of xhigh

Reddit r/LocalLLaMA · 2026-09-14

UkisAI has post-trained the Qwen 3.8 27B model to reduce unnecessary thinking tokens by 58% and achieve 1.95x speed up with less than 1% accuracy loss, open-sourcing the model and offering a free research API.

0 favorites 0 likes
#distillation

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

arXiv cs.CL · 2026-09-14 Cached

This paper presents a large-scale study comparing distilled byte and token models, finding that byte models achieve higher performance ceilings with more compute and are more data-efficient than token models.

0 favorites 0 likes
#distillation

KuaiRP Series Role-playing Models Technical Report

arXiv cs.AI · 2026-09-12 Cached

This paper introduces the KuaiRP series of role-playing models, detailing a multi-stage training pipeline that addresses the trade-off between deep domain knowledge injection and general capability preservation, achieving state-of-the-art performance with low deployment costs.

0 favorites 0 likes
#distillation

Qwen, Kimi and DeepSeek ran industrial Claude distillation: Alibaba 151M exchanges, Moonshot 23M, DeepSeek 12M in 14 days. Kimi and DeepSeek also secretly served Opus to their own users and harvested the chain-of-thought

Reddit r/singularity · 2026-09-10 Cached

Anthropic revealed that Chinese AI companies Alibaba, Moonshot AI, and DeepSeek executed extensive distillation attacks on its Claude model to extract chain-of-thought reasoning for training their own models.

0 favorites 0 likes
#distillation

Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout

Hugging Face Daily Papers · 2026-09-08 Cached

Mask Forcing mitigates mode collapse in autoregressive video diffusion distillation by injecting masked cleaner signals during self-rollout, improving visual quality without extra training data.

0 favorites 0 likes
#distillation

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Hugging Face Daily Papers · 2026-09-08

AuK is an open-source foundational model that unifies speech generation and editing through natural-language instructions, achieving leading performance with efficient inference via distillation.

0 favorites 0 likes
#distillation

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Hugging Face Daily Papers · 2026-09-08 Cached

The paper introduces On-Policy Reverse Distillation (OPRD), a method that enables stronger AI models to exceed weaker supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, achieving higher performance with fewer updates in distillation scenarios.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback