on-policy-training

Tag

Cards List
#on-policy-training

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

arXiv cs.AI · 2026-08-21 Cached

The paper proposes Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD) to address the mismatch between teacher-derived rewards and genuine reasoning progress, improving reasoning performances in language model training.

0 favorites 0 likes
#on-policy-training

Rethinking Reverse KL as Adaptive Entropy Distillation

arXiv cs.LG · 2026-08-18 Cached

This paper proposes Adaptive Entropy Distillation (AED), a method that dynamically calibrates token-level imitation strength in knowledge distillation using teacher entropy, achieving superior performance on instruction-following and mathematical reasoning benchmarks.

0 favorites 0 likes
#on-policy-training

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

Hugging Face Daily Papers · 2026-08-04 Cached

ReflectRL is a framework that learns from 'golden negative trajectories' (failed reasoning attempts by expert models) by reflecting on them, then transfers this reflective reasoning back to direct reasoning, improving LLM performance across benchmarks.

0 favorites 0 likes
#on-policy-training

Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

Hugging Face Daily Papers · 2026-05-11 Cached

This paper introduces On-Policy Data Evolution (ODE) and a visual-native agent harness to improve multimodal deep search agents. By enabling reusable visual evidence and closed-loop data generation, ODE significantly boosts the performance of Qwen3-VL agents across multiple benchmarks, surpassing Gemini 2.5 Pro.

0 favorites 0 likes
← Back to home

Submit Feedback