Tag
This paper introduces RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection into multimodal large language models. It improves injected knowledge accuracy while limiting drift in non-updated behaviors.
This paper proposes SEED, a method that internalizes post-hoc skills from trajectories into model parameters through self-evolving online distillation, solving the reward sparsity problem in long-horizon RL training, achieving significant improvements on benchmarks such as ALFWorld.
This paper investigates the position bias phenomenon in online distillation, finding that early tokens provide more useful supervision signals, and proposes the importance-weighted IW-OPD method to improve OPD training.
DiffusionOPD proposes a multi-task training paradigm for diffusion models that uses online policy distillation to efficiently combine task-specific teachers into a unified student, achieving state-of-the-art results on all evaluated benchmarks.