Tag
This tweet discusses OPD, OPSD, RLSD, and OPD^2 distillation techniques for language models, emphasizing the use of log probability differences to isolate post-training improvements.