knowledge-distillation

Tag

Cards List
#knowledge-distillation

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion

arXiv cs.LG ↗ · 2026-08-28 Cached

The paper identifies strong-modality collapse in multimodal learning where fusion degrades the dominant modality's performance, and proposes Inverted Asymmetric Fusion (IAF) to preserve it, improving over unimodal baselines.

0 favorites 0 likes
#knowledge-distillation

GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval

arXiv cs.LG ↗ · 2026-08-27 Cached

GreenLeaf Law Embed Tiny is a compact 0.6B parameter embedding model for legal domain retrieval, achieving competitive performance on benchmarks like MLEB and MTEB(Law, v1) with efficient inference for resource-constrained environments.

0 favorites 0 likes
#knowledge-distillation

Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition

arXiv cs.LG ↗ · 2026-08-27 Cached

This paper introduces Dynamic Influence Weighting (DIW), a knowledge distillation method that improves single-IMU activity recognition by dynamically weighting teacher targets from multiple IMUs during training, achieving significant performance gains.

0 favorites 0 likes
#knowledge-distillation

A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

arXiv cs.AI ↗ · 2026-08-25 Cached

This paper presents a reproducible, license-aware knowledge distillation method for creating efficient safety classifiers for large language models that can run on CPU hardware, achieving performance comparable to larger models while reducing false alarms.

0 favorites 0 likes
#knowledge-distillation

D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

Hugging Face Daily Papers ↗ · 2026-08-25 Cached

D³-MOPD is a zero-overhead scheduler for multi-teacher distillation that dynamically adjusts domain sampling ratios based on per-domain reverse-KL trajectories, improving convergence efficiency and closing most of the student-to-teacher performance gap.

0 favorites 0 likes
#knowledge-distillation

Bern2Edge: A Neurosymbolic Compiler for Edge Deployment via Bernstein Polynomial Networks

arXiv cs.LG ↗ · 2026-08-24 Cached

Bern2Edge is a neurosymbolic compiler that converts pre-trained neural networks into Bernstein polynomial representations for efficient and interpretable edge deployment on FPGAs, achieving significant reductions in latency and resource usage while maintaining accuracy.

0 favorites 0 likes
#knowledge-distillation

Sparse Token Routing in Efficient Transformers

arXiv cs.CL ↗ · 2026-08-24 Cached

This paper evaluates adaptive computation in transformers using SEWN, a two-stream model with a learned gate for token routing, demonstrating that SEWN-sparse achieves significant throughput improvements over BERT-base and DistilBERT with modest accuracy trade-offs while providing interpretable token-importance signals.

0 favorites 0 likes
#knowledge-distillation

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

arXiv cs.AI ↗ · 2026-08-21 Cached

The paper proposes Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD) to address the mismatch between teacher-derived rewards and genuine reasoning progress, improving reasoning performances in language model training.

0 favorites 0 likes
#knowledge-distillation

Accelerating Visual On-Policy Distillation with Batched Speculative Jacobi Rollouts

arXiv cs.LG ↗ · 2026-08-20 Cached

This paper introduces HB-SJD, a batched speculative Jacobi decoding method for visual on-policy distillation that accelerates rollout generation by processing multiple tokens in parallel, reducing training time while preserving generation quality.

0 favorites 0 likes
#knowledge-distillation

@aikangarooking: Still making videos! The step-by-step distillation tutorial is here~ Hand-holding you through distilling a book (highly recommended to save). My open-source CangjieSkill has also hit 8.3K stars! It's just that I've been busy lately, and the distillation tutorial everyone has been asking for is finally ready. With this method, you can distill any book, making the knowledge and methodology work for you. Not just books, ...

X AI KOLs Timeline ↗ · 2026-08-19

Introduces a comprehensive step-by-step tutorial on distilling knowledge, teaching you how to extract knowledge from books and other materials, and mentions that the open-source CangjieSkill project has received 8.3K stars.

0 favorites 0 likes
#knowledge-distillation

Tail-Aware Top-$k$ On-Policy Distillation

arXiv cs.LG ↗ · 2026-08-18 Cached

This paper proposes Tail-Aware Top-k On-Policy Distillation (TA-OPD) to address the loss of tail probability in on-policy distillation for language models, improving downstream accuracy on benchmarks.

0 favorites 0 likes
#knowledge-distillation

Rethinking Reverse KL as Adaptive Entropy Distillation

arXiv cs.LG ↗ · 2026-08-18 Cached

This paper proposes Adaptive Entropy Distillation (AED), a method that dynamically calibrates token-level imitation strength in knowledge distillation using teacher entropy, achieving superior performance on instruction-following and mathematical reasoning benchmarks.

0 favorites 0 likes
#knowledge-distillation

STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs

arXiv cs.CL ↗ · 2026-08-11 Cached

The paper introduces STEMMA, a multi-agent framework that adversarially probes self-identity consistency in LLMs, motivated by concerns that knowledge distillation may transfer behavioral traits like identity representation from teacher to student models.

0 favorites 0 likes
#knowledge-distillation

Making Knowledge Distillation Cheap Enough to Run at Scale

Hugging Face Blog ↗ · 2026-08-10 Cached

Multiverse Computing announces a paper on making LLM knowledge distillation cheaper via offline top-K logits and a fused chunked KL loss, cutting VRAM usage for distillation at scale.

0 favorites 0 likes
#knowledge-distillation

FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding

arXiv cs.CL ↗ · 2026-08-10 Cached

FutureBridge introduces a token reranker for collaborative decoding that ranks LLM-SLM candidates based on how well the SLM can continue reasoning from them, improving the Qwen3-1.7B SLM's math accuracy by 35.1% over greedy decoding.

0 favorites 0 likes
#knowledge-distillation

Simple-OPD: Demystifying Warm-up for On-policy Distillation

arXiv cs.CL ↗ · 2026-08-10 Cached

This paper investigates the warm-up stage for on-policy distillation (OPD), showing that teacher-compatible chain-of-thought supervision and LoRA-based training with near-saturation duration improve OPD effectiveness. It introduces Simple-OPD, a plug-and-play initialization method that boosts OPD performance across diverse settings.

0 favorites 0 likes
#knowledge-distillation

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

arXiv cs.LG ↗ · 2026-08-06 Cached

Introduces SPOT, a method for on-policy distillation that uses sparse probing and outcome calibration to improve reasoning performance in smaller student models while balancing solution quality and coverage.

0 favorites 0 likes
#knowledge-distillation

Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation

arXiv cs.AI ↗ · 2026-08-06 Cached

This paper studies self-distillation with privileged information (PI) as a lone post-training objective for LLMs, reproducing reported gains on easy tasks but showing it fails on difficult reasoning tasks: per-token loss drops while validation accuracy stagnates or degrades. The authors trace the failure to PI bias, which pulls teacher targets toward a reference trajectory and trains students to be flatter and less decisive without improving reasoning.

0 favorites 0 likes
#knowledge-distillation

Stuck on "A": Diagnosing and Repairing Interface Injury in Attention-to-KDA Linearization of a 0.6B Language Model

arXiv cs.CL ↗ · 2026-08-05 Cached

This paper investigates the failure modes of converting Qwen3-0.6B-Base attention layers to KDA linear attention on a single GPU, identifying an 'interface injury' where the model predicts option labels rather than content, and proposes a format-targeted KL stage to repair it.

0 favorites 0 likes
#knowledge-distillation

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance

arXiv cs.CL ↗ · 2026-08-04 Cached

This paper introduces RSTG, a method that selectively applies on-policy distillation to recover learning signals from zero-variance GRPO groups during LLM post-training, achieving substantial gains on math and code reasoning benchmarks.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback