knowledge-distillation

Tag

Cards List
#knowledge-distillation

Learning from Teacher Continuations at Student States

arXiv cs.CL ↗ · 16h ago Cached

The paper introduces OLIVE, an online knowledge distillation method where the student generates prefixes and the teacher continues them, addressing covariate shift and fragmented supervision in existing distillation approaches. OLIVE outperforms on-policy distillation and offline SFT on reasoning and agentic tasks under comparable training budgets.

0 favorites 0 likes
#knowledge-distillation

Selective Amortization of Full-Budget Counterfactual Reasoning for Visual Token Communication

arXiv cs.AI ↗ · 2d ago Cached

ACV-Gate is an adaptive candidate evaluation framework that enhances reconstruction quality in generative image communication by selectively applying counterfactual reasoning to informative visual tokens, thereby reducing computation costs.

0 favorites 0 likes
#knowledge-distillation

Persistent Negatives for Adversarial Black-Box On-Policy Distillation

arXiv cs.CL ↗ · 2d ago Cached

This paper introduces persistent-negative adversarial distillation to address the moving-target problem in black-box on-policy distillation, improving discriminator stability and performance across benchmarks.

0 favorites 0 likes
#knowledge-distillation

Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment

arXiv cs.CL ↗ · 2d ago Cached

This paper analyzes how prompt template selection during knowledge distillation affects the safety alignment of student large language models, finding that chat templates lead to greater degradation compared to non-chat templates across multiple models and benchmarks.

0 favorites 0 likes
#knowledge-distillation

SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages

arXiv cs.CL ↗ · 2d ago Cached

SEA-CLIP-Tiny is a compact multilingual text-vision embedding model for Southeast Asian languages with fewer than 50M parameters, achieving strong retrieval performance across seven languages. The model weights, datasets, and code are publicly released.

0 favorites 0 likes
#knowledge-distillation

Rank-Reliable Teacher-Guided Fitness Approximation for Expensive Evolutionary Optimization: A TinyML Architecture Search Study

arXiv cs.AI ↗ · 2d ago Cached

The paper introduces TGL-NSGA-II, a teacher-guided fitness approximation framework for expensive evolutionary optimization in constrained TinyML neural architecture search, achieving improved efficiency and reliability over standard methods.

0 favorites 0 likes
#knowledge-distillation

PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

Hugging Face Daily Papers ↗ · 2d ago Cached

PMOPD 提出一种几何感知的多教师在线策略蒸馏方法,通过构建任务子空间记忆并对梯度和优化器更新做投影来缓解能力跷跷板效应,并加入冲突探测与任务循环策略,在 Code/Reason/Math 任务上于 Qwen2.5-7B 和 Llama-3.1-8B 上分别平均提升 2.54 和 2.09 分。

0 favorites 0 likes
#knowledge-distillation

Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders

Hugging Face Daily Papers ↗ · 2d ago Cached

This paper investigates on-policy distillation in large language models using sparse crosscoders, revealing that it reweights existing features rather than creating new ones, with SFT warm-up playing a role in pre-reweighting features.

0 favorites 0 likes
#knowledge-distillation

I trained a 500M VLM that answers typed questions about an image (choice / score / yes-no) with calibrated probabilities. ~400 ms on an M1 Pro, no text generation [P]

Reddit r/MachineLearning ↗ · 3d ago Cached

A small vision-language model that answers typed questions about images with calibrated probabilities, optimized for fast inference on consumer hardware.

0 favorites 0 likes
#knowledge-distillation

Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

Hugging Face Daily Papers ↗ · 3d ago Cached

This paper investigates whether KL divergence is necessary for on-policy distillation of large language models, showing that preserving update direction suffices and introducing Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to enhance multi-teacher learning.

0 favorites 0 likes
#knowledge-distillation

Scaling Properties of Same-Family On-Policy Distillation

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper studies the scaling properties of on-policy distillation (OPD) across weak-to-strong, same-base, and strong-to-weak teacher-student LLM setups, finding a universal 'useful-transfer' regime where held-out accuracy rises linearly with KL divergence, and fitting power laws showing smaller teachers can transfer better at matched gold scores.

0 favorites 0 likes
#knowledge-distillation

Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD

arXiv cs.LG ↗ · 5d ago Cached

This paper introduces Selective Supervision for Direct-OPD (S2D-OPD), a method that improves knowledge distillation by masking low-divergence states, enhancing accuracy on math reasoning benchmarks without extra computation.

0 favorites 0 likes
#knowledge-distillation

LastOPD: Taming Collapse in Latent On-Policy Distillation

arXiv cs.LG ↗ · 5d ago Cached

The paper proposes LastOPD, a method to prevent collapse in latent on-policy distillation by applying latent signals only at the last layer during a short crossfade period, leading to improved performance on benchmarks like MATH-500.

0 favorites 0 likes
#knowledge-distillation

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

Hugging Face Daily Papers ↗ · 5d ago Cached

MOPD-Router proposes a token-level routing framework for multi-teacher on-policy distillation, enabling the use of unlabeled data and cross-domain supervision without domain labels. Experiments demonstrate significant performance improvements over standard methods in both unlabeled and labeled scenarios.

0 favorites 0 likes
#knowledge-distillation

On Mitigation of Subliminal Learning in Large Language Models

arXiv cs.CL ↗ · 2026-09-22 Cached

This paper examines subliminal learning in large language models and introduces liminal training as a method to reduce unintended trait acquisition during fine-tuning while preserving task performance.

0 favorites 0 likes
#knowledge-distillation

On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation

arXiv cs.LG ↗ · 2026-09-21 Cached

This paper introduces contrastive self-distillation to improve reasoning in AI models by separating correctness from behavioral shifts, demonstrating enhanced performance and stability.

0 favorites 0 likes
#knowledge-distillation

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

arXiv cs.AI ↗ · 2026-09-21 Cached

The paper introduces Calibrated On-Policy Distillation (Cal-OPD), a method that estimates the teacher's self-deviation to calibrate teacher-student discrepancies, improving on-policy distillation for mathematical reasoning tasks.

0 favorites 0 likes
#knowledge-distillation

Layer-wise Curriculum Learning for Efficient LLM Compression

arXiv cs.LG ↗ · 2026-09-18 Cached

The paper introduces a layer-wise curriculum learning method for efficient LLM compression, achieving state-of-the-art performance with significant reductions in GPU memory usage and training time.

0 favorites 0 likes
#knowledge-distillation

ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

arXiv cs.AI ↗ · 2026-09-16 Cached

ReDraft is a reference-driven revision method for continual post-training of large vision-language models that balances learning new tasks and preserving old ones, achieving higher accuracy and less forgetting than standard approaches like SFT.

0 favorites 0 likes
#knowledge-distillation

OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation

arXiv cs.LG ↗ · 2026-09-16 Cached

OPD-Aha introduces a novel on-policy distillation method that leverages visual preference to correct hallucinations in multimodal reasoning, achieving consistent improvements on various benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback