Tag
The paper introduces OPDVR, a method that combines on-policy distillation with verifiable rewards using a ReLU-gated reformulation to enhance reasoning in large language models without extra hyperparameters, outperforming standard approaches on six benchmarks.