Tag
This thread introduces β-OPSD, a post-training abstraction that frames RL as a compiler for supervision, using policy optimization to derive targets and distillation for training.