contrastive-supervision

Tag

Cards List
#contrastive-supervision

@sheriyuo: Qwen Tongyi Lab proposes RLCSD, a simple but important critique of on-policy self-distillation. Their key observation i…

X AI KOLs Timeline · 2026-06-11 Cached

Qwen Tongyi Lab proposes RLCSD to address the style drift problem in on-policy self-distillation, where the learning signal focuses on style tokens rather than task-critical reasoning tokens. Their method uses contrastive supervision to focus on task-relevant tokens, achieving consistent improvements over prior methods on reasoning benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback