Tag
The paper identifies strong-modality collapse in multimodal learning where fusion degrades the dominant modality's performance, and proposes Inverted Asymmetric Fusion (IAF) to preserve it, improving over unimodal baselines.
GreenLeaf Law Embed Tiny is a compact 0.6B parameter embedding model for legal domain retrieval, achieving competitive performance on benchmarks like MLEB and MTEB(Law, v1) with efficient inference for resource-constrained environments.
This paper introduces Dynamic Influence Weighting (DIW), a knowledge distillation method that improves single-IMU activity recognition by dynamically weighting teacher targets from multiple IMUs during training, achieving significant performance gains.
This paper presents a reproducible, license-aware knowledge distillation method for creating efficient safety classifiers for large language models that can run on CPU hardware, achieving performance comparable to larger models while reducing false alarms.
D³-MOPD is a zero-overhead scheduler for multi-teacher distillation that dynamically adjusts domain sampling ratios based on per-domain reverse-KL trajectories, improving convergence efficiency and closing most of the student-to-teacher performance gap.
Bern2Edge is a neurosymbolic compiler that converts pre-trained neural networks into Bernstein polynomial representations for efficient and interpretable edge deployment on FPGAs, achieving significant reductions in latency and resource usage while maintaining accuracy.
This paper evaluates adaptive computation in transformers using SEWN, a two-stream model with a learned gate for token routing, demonstrating that SEWN-sparse achieves significant throughput improvements over BERT-base and DistilBERT with modest accuracy trade-offs while providing interpretable token-importance signals.
The paper proposes Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD) to address the mismatch between teacher-derived rewards and genuine reasoning progress, improving reasoning performances in language model training.
This paper introduces HB-SJD, a batched speculative Jacobi decoding method for visual on-policy distillation that accelerates rollout generation by processing multiple tokens in parallel, reducing training time while preserving generation quality.
Introduces a comprehensive step-by-step tutorial on distilling knowledge, teaching you how to extract knowledge from books and other materials, and mentions that the open-source CangjieSkill project has received 8.3K stars.
This paper proposes Tail-Aware Top-k On-Policy Distillation (TA-OPD) to address the loss of tail probability in on-policy distillation for language models, improving downstream accuracy on benchmarks.
This paper proposes Adaptive Entropy Distillation (AED), a method that dynamically calibrates token-level imitation strength in knowledge distillation using teacher entropy, achieving superior performance on instruction-following and mathematical reasoning benchmarks.
The paper introduces STEMMA, a multi-agent framework that adversarially probes self-identity consistency in LLMs, motivated by concerns that knowledge distillation may transfer behavioral traits like identity representation from teacher to student models.
Multiverse Computing announces a paper on making LLM knowledge distillation cheaper via offline top-K logits and a fused chunked KL loss, cutting VRAM usage for distillation at scale.
FutureBridge introduces a token reranker for collaborative decoding that ranks LLM-SLM candidates based on how well the SLM can continue reasoning from them, improving the Qwen3-1.7B SLM's math accuracy by 35.1% over greedy decoding.
This paper investigates the warm-up stage for on-policy distillation (OPD), showing that teacher-compatible chain-of-thought supervision and LoRA-based training with near-saturation duration improve OPD effectiveness. It introduces Simple-OPD, a plug-and-play initialization method that boosts OPD performance across diverse settings.
Introduces SPOT, a method for on-policy distillation that uses sparse probing and outcome calibration to improve reasoning performance in smaller student models while balancing solution quality and coverage.
This paper studies self-distillation with privileged information (PI) as a lone post-training objective for LLMs, reproducing reported gains on easy tasks but showing it fails on difficult reasoning tasks: per-token loss drops while validation accuracy stagnates or degrades. The authors trace the failure to PI bias, which pulls teacher targets toward a reference trajectory and trains students to be flatter and less decisive without improving reasoning.
This paper investigates the failure modes of converting Qwen3-0.6B-Base attention layers to KDA linear attention on a single GPU, identifying an 'interface injury' where the model predicts option labels rather than content, and proposes a format-targeted KL stage to repair it.
This paper introduces RSTG, a method that selectively applies on-policy distillation to recover learning signals from zero-variance GRPO groups during LLM post-training, achieving substantial gains on math and code reasoning benchmarks.