knowledge-distillation

Tag

Cards List
#knowledge-distillation

FLoKD: Adaptive Knowledge Distillation for Federated Low-Rank LLM over Wireless Networks

arXiv cs.AI ↗ · 2026-09-15 Cached

The paper proposes FLoKD, an adaptive knowledge-distillation framework for federated LoRA fine-tuning of LLMs over wireless networks, reducing communication overhead by 50-65% while maintaining competitive performance.

0 favorites 0 likes
#knowledge-distillation

Discovering and Preserving Category Correlation Knowledge via Adaptive Reciprocal Knowledge Distillation

arXiv cs.LG ↗ · 2026-09-15 Cached

This paper proposes adaptive reciprocal knowledge distillation (AR-KD), a novel method that improves knowledge transfer from teacher to student models by simplifying the teacher's output distribution through relational alignment, achieving up to 7.13% accuracy gain on CIFAR-100 and ImageNet-1k datasets.

0 favorites 0 likes
#knowledge-distillation

Parameter-Efficient Retrievers for Polish and European Languages

arXiv cs.CL ↗ · 2026-09-14 Cached

This paper presents a three-stage training pipeline for developing compact dense retrievers, introducing PolDense for Polish and EuroDense for European languages, which achieve strong performance with significantly reduced parameters compared to larger models.

0 favorites 0 likes
#knowledge-distillation

Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation

arXiv cs.AI ↗ · 2026-09-11 Cached

The paper examines correctness-gated multi-teacher distillation, finding decision shifts, lost label functionality, and an inconclusive grounding audit, with no incremental benefit over hard filtering.

0 favorites 0 likes
#knowledge-distillation

Distilling Vision-Language Models for On-Device Fire Understanding

arXiv cs.AI ↗ · 2026-09-10 Cached

This paper proposes a knowledge distillation framework to compress vision-language models for on-device fire detection, showing that compact models can retain most of their teacher's capability while being deployable on embedded hardware.

0 favorites 0 likes
#knowledge-distillation

RAPID: Reliability-Aware Pair Importance Distillation

arXiv cs.AI ↗ · 2026-09-10 Cached

RAPID introduces a reliability-aware pair importance distillation method to improve knowledge distillation efficiency and performance in text classification tasks.

0 favorites 0 likes
#knowledge-distillation

FANS: Federated Adaptive Network Search Learning for Heterogeneous Devices

arXiv cs.LG ↗ · 2026-09-10 Cached

The article introduces FANS, a hypernetwork-based framework for heterogeneous federated learning that learns a shared architecture space and uses parallel training with self-distillation to optimize model selection across diverse devices.

0 favorites 0 likes
#knowledge-distillation

CALM: Class-wise Agreement and Label-gated Disagreement Modulation for Decentralized Federated Learning

arXiv cs.LG ↗ · 2026-09-10 Cached

CALM proposes a method for decentralized federated learning that uses class-wise agreement and label-gated disagreement modulation to handle non-IID data, improving teacher weighting and distillation trust without additional communication overhead.

0 favorites 0 likes
#knowledge-distillation

@rsasaki0109: ZipDepth [ECCV 2026] Official implementation of "ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on …

X AI KOLs Timeline ↗ · 2026-09-06 Cached

ZipDepth is a lightweight zero-shot monocular depth estimation model that achieves the best accuracy-efficiency trade-off, running in real time on any device from mobile phones to server GPUs, and has been accepted at ECCV 2026.

0 favorites 0 likes
#knowledge-distillation

Convergence Theory of Knowledge Distillation in Asynchronous P2P Gossip Learning Network

arXiv cs.LG ↗ · 2026-09-03 Cached

This paper provides convergence theory for knowledge distillation in asynchronous peer-to-peer gossip learning networks, demonstrating that it contracts function disagreement and analyzing theoretical convergence rates.

0 favorites 0 likes
#knowledge-distillation

On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers

arXiv cs.LG ↗ · 2026-09-03 Cached

This paper introduces a reinforcement learning-based distillation framework for training compact instruction-following rerankers, using off-policy GRPO for teacher enhancement and on-policy distillation for student learning, demonstrating superior performance under distribution shift.

0 favorites 0 likes
#knowledge-distillation

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

The paper investigates on-policy distillation of large language models, demonstrating that a single training query can achieve substantial state coverage and alignment, suggesting the method is algorithm-starved rather than data-starved.

0 favorites 0 likes
#knowledge-distillation

CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning

arXiv cs.LG ↗ · 2026-09-02 Cached

The paper proposes CRAD, a class-wise reliability-aware distillation method for decentralized federated learning to handle heterogeneous architectures and non-IID data, achieving improved accuracy on image classification benchmarks.

0 favorites 0 likes
#knowledge-distillation

Knowledge Distillation under Teacher Misspecification: An Order-Parameter Analysis of the Gap between Teacher Mimicry and Task Performance

arXiv cs.LG ↗ · 2026-09-01 Cached

The paper analyzes the gap between teacher mimicry and true task performance in knowledge distillation under teacher misspecification using order-parameter methods, showing that mimicry metrics can be invariant while true errors increase with mismatch.

0 favorites 0 likes
#knowledge-distillation

Temperature-Adaptive Transformed Teacher Matching

arXiv cs.LG ↗ · 2026-09-01 Cached

This paper proposes a sample-wise adaptive temperature scaling method for Transformed Teacher Matching in knowledge distillation, improving performance on image classification benchmarks by locally minimizing KL divergence between teacher and student distributions.

0 favorites 0 likes
#knowledge-distillation

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Hugging Face Daily Papers ↗ · 2026-09-01 Cached

This paper introduces Switch Distillation, a novel mid-training objective that selectively applies knowledge distillation based on teacher confidence to improve reasoning and preserve factual recall in smaller language models.

0 favorites 0 likes
#knowledge-distillation

Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation

Hugging Face Daily Papers ↗ · 2026-08-30 Cached

The paper introduces Influence-Directed Adaptive On-Policy Distillation (IDA-OPD) to solve the diversity bottleneck in sampled-token on-policy distillation, enhancing diversity transfer in reasoning models without full-vocabulary teacher data.

0 favorites 0 likes
#knowledge-distillation

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion

arXiv cs.LG ↗ · 2026-08-28 Cached

The paper identifies strong-modality collapse in multimodal learning where fusion degrades the dominant modality's performance, and proposes Inverted Asymmetric Fusion (IAF) to preserve it, improving over unimodal baselines.

0 favorites 0 likes
#knowledge-distillation

GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval

arXiv cs.LG ↗ · 2026-08-27 Cached

GreenLeaf Law Embed Tiny is a compact 0.6B parameter embedding model for legal domain retrieval, achieving competitive performance on benchmarks like MLEB and MTEB(Law, v1) with efficient inference for resource-constrained environments.

0 favorites 0 likes
#knowledge-distillation

Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition

arXiv cs.LG ↗ · 2026-08-27 Cached

This paper introduces Dynamic Influence Weighting (DIW), a knowledge distillation method that improves single-IMU activity recognition by dynamically weighting teacher targets from multiple IMUs during training, achieving significant performance gains.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback