Tag
The paper introduces OLIVE, an online knowledge distillation method where the student generates prefixes and the teacher continues them, addressing covariate shift and fragmented supervision in existing distillation approaches. OLIVE outperforms on-policy distillation and offline SFT on reasoning and agentic tasks under comparable training budgets.
ACV-Gate is an adaptive candidate evaluation framework that enhances reconstruction quality in generative image communication by selectively applying counterfactual reasoning to informative visual tokens, thereby reducing computation costs.
This paper introduces persistent-negative adversarial distillation to address the moving-target problem in black-box on-policy distillation, improving discriminator stability and performance across benchmarks.
This paper analyzes how prompt template selection during knowledge distillation affects the safety alignment of student large language models, finding that chat templates lead to greater degradation compared to non-chat templates across multiple models and benchmarks.
SEA-CLIP-Tiny is a compact multilingual text-vision embedding model for Southeast Asian languages with fewer than 50M parameters, achieving strong retrieval performance across seven languages. The model weights, datasets, and code are publicly released.
The paper introduces TGL-NSGA-II, a teacher-guided fitness approximation framework for expensive evolutionary optimization in constrained TinyML neural architecture search, achieving improved efficiency and reliability over standard methods.
PMOPD 提出一种几何感知的多教师在线策略蒸馏方法,通过构建任务子空间记忆并对梯度和优化器更新做投影来缓解能力跷跷板效应,并加入冲突探测与任务循环策略,在 Code/Reason/Math 任务上于 Qwen2.5-7B 和 Llama-3.1-8B 上分别平均提升 2.54 和 2.09 分。
This paper investigates on-policy distillation in large language models using sparse crosscoders, revealing that it reweights existing features rather than creating new ones, with SFT warm-up playing a role in pre-reweighting features.
A small vision-language model that answers typed questions about images with calibrated probabilities, optimized for fast inference on consumer hardware.
This paper investigates whether KL divergence is necessary for on-policy distillation of large language models, showing that preserving update direction suffices and introducing Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to enhance multi-teacher learning.
This paper studies the scaling properties of on-policy distillation (OPD) across weak-to-strong, same-base, and strong-to-weak teacher-student LLM setups, finding a universal 'useful-transfer' regime where held-out accuracy rises linearly with KL divergence, and fitting power laws showing smaller teachers can transfer better at matched gold scores.
This paper introduces Selective Supervision for Direct-OPD (S2D-OPD), a method that improves knowledge distillation by masking low-divergence states, enhancing accuracy on math reasoning benchmarks without extra computation.
The paper proposes LastOPD, a method to prevent collapse in latent on-policy distillation by applying latent signals only at the last layer during a short crossfade period, leading to improved performance on benchmarks like MATH-500.
MOPD-Router proposes a token-level routing framework for multi-teacher on-policy distillation, enabling the use of unlabeled data and cross-domain supervision without domain labels. Experiments demonstrate significant performance improvements over standard methods in both unlabeled and labeled scenarios.
This paper examines subliminal learning in large language models and introduces liminal training as a method to reduce unintended trait acquisition during fine-tuning while preserving task performance.
This paper introduces contrastive self-distillation to improve reasoning in AI models by separating correctness from behavioral shifts, demonstrating enhanced performance and stability.
The paper introduces Calibrated On-Policy Distillation (Cal-OPD), a method that estimates the teacher's self-deviation to calibrate teacher-student discrepancies, improving on-policy distillation for mathematical reasoning tasks.
The paper introduces a layer-wise curriculum learning method for efficient LLM compression, achieving state-of-the-art performance with significant reductions in GPU memory usage and training time.
ReDraft is a reference-driven revision method for continual post-training of large vision-language models that balances learning new tasks and preserving old ones, achieving higher accuracy and less forgetting than standard approaches like SFT.
OPD-Aha introduces a novel on-policy distillation method that leverages visual preference to correct hallucinations in multimodal reasoning, achieving consistent improvements on various benchmarks.