adaptive-training

Tag

Cards List
#adaptive-training

D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation

Hugging Face Daily Papers ↗ · 2026-08-25 Cached

D³-MOPD is a zero-overhead scheduler for multi-teacher distillation that dynamically adjusts domain sampling ratios based on per-domain reverse-KL trajectories, improving convergence efficiency and closing most of the student-to-teacher performance gap.

0 favorites 0 likes
#adaptive-training

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction

arXiv cs.CL ↗ · 2026-08-04 Cached

This paper introduces AdaMTP, an adaptive training paradigm for multi-token prediction that dynamically aligns prediction horizons with sequence predictability using entropy-based segmentation, consistently outperforming standard MTP on math, code, and general benchmarks across three LLM backbones.

0 favorites 0 likes
#adaptive-training

LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents

arXiv cs.LG ↗ · 2026-06-18 Cached

LLMZero uses LLM agents to search over training trajectories via tree search, discovering adaptive multi-parameter transitions for RL post-training that outperform fixed schedules and grid search across diverse tasks.

0 favorites 0 likes
#adaptive-training

Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning

Hugging Face Daily Papers ↗ · 2026-05-12 Cached

Adaptive Teacher Exposure for Self-Distillation (ATESD) improves LLM reasoning by dynamically adjusting how much of the reference reasoning the teacher shows the student during training, using a learnable policy controller and a discounted learning-progress reward. Experiments on math benchmarks show consistent improvements over existing self-distillation and RL baselines.

0 favorites 0 likes
← Back to home

Submit Feedback