beta-opsd

Tag

Cards List
#beta-opsd

@furongh: RL vs. distillation may be a false dichotomy. A post-training abstraction: RL as a compiler for supervision. Use policy…

X AI KOLs Timeline · 6d ago Cached

This thread introduces β-OPSD, a post-training abstraction that frames RL as a compiler for supervision, using policy optimization to derive targets and distillation for training.

0 favorites 0 likes
← Back to home

Submit Feedback