Tag
A research paper introducing Canon, a label-free self-distillation method that uses consensus among sampled solutions to provide dense token-level supervision for training large language models on reasoning tasks, improving pass@1 by up to 12 points and outperforming label-free reinforcement learning at a fraction of the compute.
The paper identifies 'temporal credit dilution' in learned dynamics models where global readouts focus on spurious correlates rather than brief physical events. It proposes CREST, a training-free method that re-anchors pooled representations using event core estimates, improving out-of-distribution robustness.
Proposes Cross-Model Entropy (CME) as a label-free reward signal for reinforcement learning post-training of large language models, enabling open-ended instruction following without ground-truth verifiers or human preference labels.