Tag
The paper proposes FLoKD, an adaptive knowledge-distillation framework for federated LoRA fine-tuning of LLMs over wireless networks, reducing communication overhead by 50-65% while maintaining competitive performance.
This paper proposes adaptive reciprocal knowledge distillation (AR-KD), a novel method that improves knowledge transfer from teacher to student models by simplifying the teacher's output distribution through relational alignment, achieving up to 7.13% accuracy gain on CIFAR-100 and ImageNet-1k datasets.
This paper presents a three-stage training pipeline for developing compact dense retrievers, introducing PolDense for Polish and EuroDense for European languages, which achieve strong performance with significantly reduced parameters compared to larger models.
The paper examines correctness-gated multi-teacher distillation, finding decision shifts, lost label functionality, and an inconclusive grounding audit, with no incremental benefit over hard filtering.
This paper proposes a knowledge distillation framework to compress vision-language models for on-device fire detection, showing that compact models can retain most of their teacher's capability while being deployable on embedded hardware.
RAPID introduces a reliability-aware pair importance distillation method to improve knowledge distillation efficiency and performance in text classification tasks.
The article introduces FANS, a hypernetwork-based framework for heterogeneous federated learning that learns a shared architecture space and uses parallel training with self-distillation to optimize model selection across diverse devices.
CALM proposes a method for decentralized federated learning that uses class-wise agreement and label-gated disagreement modulation to handle non-IID data, improving teacher weighting and distillation trust without additional communication overhead.
ZipDepth is a lightweight zero-shot monocular depth estimation model that achieves the best accuracy-efficiency trade-off, running in real time on any device from mobile phones to server GPUs, and has been accepted at ECCV 2026.
This paper provides convergence theory for knowledge distillation in asynchronous peer-to-peer gossip learning networks, demonstrating that it contracts function disagreement and analyzing theoretical convergence rates.
This paper introduces a reinforcement learning-based distillation framework for training compact instruction-following rerankers, using off-policy GRPO for teacher enhancement and on-policy distillation for student learning, demonstrating superior performance under distribution shift.
The paper investigates on-policy distillation of large language models, demonstrating that a single training query can achieve substantial state coverage and alignment, suggesting the method is algorithm-starved rather than data-starved.
The paper proposes CRAD, a class-wise reliability-aware distillation method for decentralized federated learning to handle heterogeneous architectures and non-IID data, achieving improved accuracy on image classification benchmarks.
The paper analyzes the gap between teacher mimicry and true task performance in knowledge distillation under teacher misspecification using order-parameter methods, showing that mimicry metrics can be invariant while true errors increase with mismatch.
This paper proposes a sample-wise adaptive temperature scaling method for Transformed Teacher Matching in knowledge distillation, improving performance on image classification benchmarks by locally minimizing KL divergence between teacher and student distributions.
This paper introduces Switch Distillation, a novel mid-training objective that selectively applies knowledge distillation based on teacher confidence to improve reasoning and preserve factual recall in smaller language models.
The paper introduces Influence-Directed Adaptive On-Policy Distillation (IDA-OPD) to solve the diversity bottleneck in sampled-token on-policy distillation, enhancing diversity transfer in reasoning models without full-vocabulary teacher data.
The paper identifies strong-modality collapse in multimodal learning where fusion degrades the dominant modality's performance, and proposes Inverted Asymmetric Fusion (IAF) to preserve it, improving over unimodal baselines.
GreenLeaf Law Embed Tiny is a compact 0.6B parameter embedding model for legal domain retrieval, achieving competitive performance on benchmarks like MLEB and MTEB(Law, v1) with efficient inference for resource-constrained environments.
This paper introduces Dynamic Influence Weighting (DIW), a knowledge distillation method that improves single-IMU activity recognition by dynamically weighting teacher targets from multiple IMUs during training, achieving significant performance gains.