on-policy-learning

Tag

Cards List
#on-policy-learning

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

arXiv cs.AI · 2026-07-03 Cached

Introduces Φ-Nav, a unified on-policy framework that uses hindsight reasoning to synthetically generate path-level instructions from exploratory trajectories, bridging the semantic supervision gap in Vision-Language Navigation and achieving competitive results on R2R-CE and RxR-CE benchmarks with fewer expert demonstrations.

0 favorites 0 likes
#on-policy-learning

OPRD: On-Policy Representation Distillation

Hugging Face Daily Papers · 2026-06-04

OPRD proposes a new knowledge distillation method that aligns student and teacher hidden states across layers during on-policy rollouts, eliminating sampling variance from token-space KL estimation. Empirically, OPRD outperforms output-space baselines on math reasoning benchmarks (AIME 2024/2025, AIMO) while being 1.44x faster and using 54% less memory.

0 favorites 0 likes
#on-policy-learning

Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning

arXiv cs.CL · 2026-05-12 Cached

This paper introduces On-Policy Harness Self-Distillation (OPHSD), a method that internalizes the capabilities of inference-time reasoning harnesses into the base model through self-distillation. The approach improves standalone performance on complex reasoning tasks, allowing the model to retain reasoning scaffolds without permanent external dependencies.

0 favorites 0 likes
#on-policy-learning

Rubric-based On-policy Distillation

Hugging Face Daily Papers · 2026-05-08 Cached

This paper introduces ROPD, a rubric-based on-policy distillation framework that achieves superior sample efficiency compared to traditional logit-based methods. It enables model alignment in black-box scenarios by using structured semantic rubrics instead of teacher logits.

0 favorites 0 likes
#on-policy-learning

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models

Hugging Face Daily Papers · 2026-05-06 Cached

This paper introduces D-OPSD, a novel training paradigm for step-distilled diffusion models that enables on-policy self-distillation during supervised fine-tuning. It allows models to learn new concepts or styles without compromising their efficient few-step inference capabilities.

0 favorites 0 likes
#on-policy-learning

How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data

Hugging Face Daily Papers · 2026-03-23 Cached

This paper introduces TESSY, a teacher-student cooperative framework for fine-tuning reasoning models that generates on-policy SFT data by decoupling generation into capability tokens (from teacher) and style tokens (from student), addressing catastrophic forgetting issues when using off-policy teacher data.

0 favorites 0 likes
← Back to home

Submit Feedback