token-level-reward

Tag

Cards List
#token-level-reward

@wu_taiqiang: How to maximize OPD performance? One important thing is warm-up. Then the student-sampled sequence is well defined in t…

X AI KOLs Following · 2026-08-13 Cached

The author discusses a paper that demystifies the warm-up process for OPD (likely on-policy distillation), explaining how warm-up enables well-defined student-sampled sequences and educational token-level dense rewards from the teacher.

0 favorites 0 likes
#token-level-reward

SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding

arXiv cs.CL · 2026-07-28 Cached

Introduces SyRuP, a decoding-time framework that trains a cross-attention reward head to produce token-level adherence scores for system prompts, improving LLM following of complex prompts without model tuning.

0 favorites 0 likes
← Back to home

Submit Feedback