Tag
The paper introduces On-Policy Reverse Distillation (OPRD), a method that enables stronger AI models to exceed weaker supervisors by amplifying verifier-supported policy gradients along the teacher's shift direction, achieving higher performance with fewer updates in distillation scenarios.
CritICL is an inference-time framework that enhances large language model reasoning by leveraging structured failure patterns from smaller models as critique-based guidance, outperforming standard in-context learning with reduced generation and token costs.
Trust functions enable near-lossless weak-to-strong generalization by identifying reliable weak labels for training, achieving performance comparable to ground-truth supervision across multiple domains.