token-level-supervision

Tag

Cards List
#token-level-supervision

DOPD: Dual On-policy Distillation

Hugging Face Daily Papers · 2026-06-29 Cached

DOPD proposes a dual on-policy distillation paradigm that dynamically routes token-level supervision between privileged teacher and student policies based on advantage gaps and probabilities, addressing privilege illusion and improving capability transfer in LLMs and VLMs.

0 favorites 0 likes
#token-level-supervision

@VukRosic99: When a small model learns from a big one, half the lesson is wasted The setup: a small "student" model writes an answer…

X AI KOLs Timeline · 2026-06-28 Cached

The paper identifies position bias in on-policy distillation for language models, where later tokens in student-generated answers receive degraded supervision. The proposed Importance-Weighted On-Policy Distillation (IW-OPD) weights corrections based on accumulated drift, improving learning speed and final performance.

0 favorites 0 likes
#token-level-supervision

Trust Region On-Policy Distillation

Hugging Face Daily Papers · 2026-05-31 Cached

The paper proposes Trust Region On-Policy Distillation (TrOPD) to stabilize on-policy distillation of large language models by using trust regions, outlier estimation, and off-policy guidance, outperforming existing methods on reasoning and code generation benchmarks.

0 favorites 0 likes
#token-level-supervision

When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning

arXiv cs.LG · 2026-05-22 Cached

This paper identifies that teacher token reliability in reasoning distillation is trajectory-structured and proposes Position-Weighted On-Policy Self-Distillation (PW-OPSD), which applies increasing position weights to improve performance without additional teacher computation.

0 favorites 0 likes
← Back to home

Submit Feedback