subspace-locking

Tag

Cards List
#subspace-locking

On the Geometry of On-Policy Distillation

Hugging Face Daily Papers · 2026-06-05 Cached

This paper characterizes the unique parameter space dynamics of on-policy distillation (OPD) for large language models, showing that it exhibits relaxed off-principal updates and subspace locking, distinguishing it from supervised fine-tuning and reinforcement learning with verifiable rewards.

0 favorites 0 likes
← Back to home

Submit Feedback