verifiable-reward

Tag

Cards List
#verifiable-reward

On-policy Distillation with Verifiable Reward

Hugging Face Daily Papers · 2026-08-25 Cached

The paper introduces OPDVR, a method that combines on-policy distillation with verifiable rewards using a ReLU-gated reformulation to enhance reasoning in large language models without extra hyperparameters, outperforming standard approaches on six benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback