online-distillation

Tag

Cards List
#online-distillation

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

arXiv cs.AI · 2026-07-29 Cached

This paper introduces RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection into multimodal large language models. It improves injected knowledge accuracy while limiting drift in non-updated behaviors.

0 favorites 0 likes
#online-distillation

@Xudong07452910: A classic challenge in RL training of LLM agents: after a long task fails, where should the model start learning? The final reward can usually only tell the agent 'success' or 'failure', but it's hard to pinpoint which intermediate judgments are worth keeping and which actions led the entire trajectory astray. This paper proposes SEED, using 'self-evolving online distillation...'

X AI KOLs Timeline · 2026-07-20 Cached

This paper proposes SEED, a method that internalizes post-hoc skills from trajectories into model parameters through self-evolving online distillation, solving the reward sparsity problem in long-horizon RL training, achieving significant improvements on benchmarks such as ALFWorld.

0 favorites 0 likes
#online-distillation

@qingke_ai: https://x.com/qingke_ai/status/2073975637904380059

X AI KOLs Timeline · 2026-07-06 Cached

This paper investigates the position bias phenomenon in online distillation, finding that early tokens provide more useful supervision signals, and proposes the importance-weighted IW-OPD method to improve OPD training.

0 favorites 0 likes
#online-distillation

DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

Hugging Face Daily Papers · 2026-05-14 Cached

DiffusionOPD proposes a multi-task training paradigm for diffusion models that uses online policy distillation to efficiently combine task-specific teachers into a unified student, achieving state-of-the-art results on all evaluated benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback