Tag
RLCD is explained as a schema-conditioned Plackett–Luce objective that advances reward modeling from scalar rewards to pairwise preferences to multiway calibrated decisions, simplifying the understanding of Jev.
RewardVerse presents a rubric-based framework for video reward modeling to mitigate scalar drift and provide stable evaluation criteria, enhancing reinforcement learning in video generation.
This paper introduces Gradient-Aligned Reward (GAR), a method to enhance reinforcement learning for LLM reasoning by aligning gradients with expert solutions, showing improvements on math and general benchmarks.
WorldReward introduces a vision-language reward model for camera-conditioned world models that unifies action-consistency and visual-quality evaluation through chunk decomposition and preference aggregation, outperforming existing methods like GPT-5.5 on benchmarks.
FIRM-Video introduces a checklist-driven verification framework for constructing reliable reward models in text-to-video tasks, using temporal visual evidence to improve evaluation and alignment.
VA-Judger is the first reward model for joint video-audio generation that uses human preference feedback to evaluate holistic quality, including a dataset and benchmark, and demonstrates significant improvements in human preference rates when applied to models like LTX-2.
This paper proposes a proxemics-based reward formulation for deep reinforcement learning social navigation, modeling human personal space as Gaussian-mixture fields to improve social compliance while maintaining navigation efficiency.
This paper identifies procedural fairness failures in RLHF caused by averaging heterogeneous preferences, where majority groups dominate reward learning and minority preferences are under-represented. It proposes Preference-Aware RLHF (PA-RLHF), which improves alignment accuracy and reduces the fairness gap in controlled experiments.
This paper introduces LEMUR, a framework that combines multi-objective reinforcement learning with preference-based learning from multiple human feedback to learn Pareto-optimal policies without predefined reward functions.
This paper presents a system for CLEF 2026 CheckThat! Task 2 that uses LLM-based trace ranking and grouped reward modeling for verifying numerical claims in English and Arabic, comparing fine-tuned verifiers with lightweight reward models.
This survey provides a unified view of progress reward modeling for robotic learning, organizing the field into three steps: interface, methods, and data/benchmarks.
This paper introduces a regularization technique for Direct Alignment Algorithms (DAAs) that maintains normalized response probabilities, mitigating over-optimization and likelihood displacement. The method improves generation quality and benchmark performance, achieving over 20% relative increase on AlpacaEval2 and 9% gains on general benchmarks for Llama-3.1-8B-Instruct.
This paper identifies and formalizes rater state bias in RLHF preference data, where annotator emotional state can confound preference labels. It proposes an audit framework with falsifiable predictions to detect such biases.
This paper introduces DROPJ, a human-centred method for safely training and deploying agent policies by learning a world model from real-world trajectories, then eliciting human preferences with justifications to train a reward model for model predictive control. Experiments show that using human-generated simulated trajectories and justifications improves safety and reduces computational cost.
This paper introduces Progressive Code-Switching (PCS), a reinforcement learning approach with curriculum learning that gradually increases code-switching in LLMs to efficiently transfer multilingual reasoning capabilities.
Harvey partnered with Applied Compute to train a legal agent, optimizing the agent stack and post-training the GLM-5.1 model using reward signals from their Legal Agent Benchmark.
This paper proposes PAFO, a Pareto fairness optimization framework to mitigate personalized reward bias in reward models for LLMs, improving accuracy for minority user groups without harming majority groups.
This paper introduces Multimodal Multi-Dimensional Scalarization Process Reward Modeling (MMS-PRM), which enforces the worst dimension's robustness in multimodal reasoning to prevent failures like visual hallucinations from being masked by strong text logic.
Eval-Skill is an exploration-guided method that synthesizes reusable evaluation skills for reward modeling, achieving significant gains on RewardBench 2 over existing backbones.
Z-Reward is a teacher-student framework that decouples complex reasoning from efficient reward deployment for text-to-image training. It achieves 89.6% human preference accuracy with a 27B teacher and 88.6% with a 9B student, outperforming prior methods.