Tag
The author discusses a paper that demystifies the warm-up process for OPD (likely on-policy distillation), explaining how warm-up enables well-defined student-sampled sequences and educational token-level dense rewards from the teacher.
Introduces SyRuP, a decoding-time framework that trains a cross-attention reward head to produce token-level adherence scores for system prompts, improving LLM following of complex prompts without model tuning.