process-reward-models

Tag

Cards List
#process-reward-models

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

arXiv cs.LG · 2026-08-11 Cached

This paper formulates PRM stress testing as a quality-diversity search problem using MAP-Elites, characterizing what archive coverage can and cannot certify. Experiments on Qwen2.5-Math-PRM-7B reveal aggregation-dependent vulnerabilities, and a LoRA repair protocol reduces exploit rates.

0 favorites 0 likes
#process-reward-models

From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning

arXiv cs.CL · 2026-06-08 Cached

This paper introduces Prefix Utility Model (PUM), which evaluates LLM reasoning prefixes based on their utility (improvement in solve rate) rather than local correctness. PUM shows strong performance in mathematical reasoning tasks across selection, search, and reinforcement learning.

0 favorites 0 likes
#process-reward-models

Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport

arXiv cs.LG · 2026-05-11 Cached

This paper introduces Distributional Process Reward Models, using conditional optimal transport to calibrate PRMs for more accurate success probability estimates in inference-time scaling. It demonstrates improved calibration and downstream performance on mathematical reasoning benchmarks like MATH-500 and AIME.

0 favorites 0 likes
#process-reward-models

Unsupervised Process Reward Models

Hugging Face Daily Papers · 2026-05-11 Cached

This paper proposes unsupervised Process Reward Models (uPRM) that eliminate the need for human annotations by using LLM next-token probabilities to identify erroneous reasoning steps, achieving up to 15% accuracy improvements over LLM-as-a-Judge and performing comparably to supervised PRMs as verifiers and reward signals.

0 favorites 0 likes
← Back to home

Submit Feedback