reward-modeling

Tag

Cards List
#reward-modeling

What Is RLCD? The Secret Behind Jev

Hacker News Top ↗ · 3d ago Cached

RLCD is explained as a schema-conditioned Plackett–Luce objective that advances reward modeling from scalar rewards to pairwise preferences to multiway calibrated decisions, simplifying the understanding of Jev.

0 favorites 0 likes
#reward-modeling

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

Hugging Face Daily Papers ↗ · 2026-09-19 Cached

RewardVerse presents a rubric-based framework for video reward modeling to mitigate scalar drift and provide stable evaluation criteria, enhancing reinforcement learning in video generation.

0 favorites 0 likes
#reward-modeling

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

arXiv cs.LG ↗ · 2026-09-04 Cached

This paper introduces Gradient-Aligned Reward (GAR), a method to enhance reinforcement learning for LLM reasoning by aligning gradients with expert solutions, showing improvements on math and general benchmarks.

0 favorites 0 likes
#reward-modeling

WorldReward: Reward Modeling for Camera-Conditioned World Models

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

WorldReward introduces a vision-language reward model for camera-conditioned world models that unifies action-consistency and visual-quality evaluation through chunk decomposition and preference aggregation, outperforming existing methods like GPT-5.5 on benchmarks.

0 favorites 0 likes
#reward-modeling

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

Hugging Face Daily Papers ↗ · 2026-08-22 Cached

FIRM-Video introduces a checklist-driven verification framework for constructing reliable reward models in text-to-video tasks, using temporal visual evidence to improve evaluation and alignment.

0 favorites 0 likes
#reward-modeling

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

Hugging Face Daily Papers ↗ · 2026-08-19 Cached

VA-Judger is the first reward model for joint video-audio generation that uses human preference feedback to evaluate holistic quality, including a dataset and benchmark, and demonstrates significant improvements in human preference rates when applied to models like LTX-2.

0 favorites 0 likes
#reward-modeling

Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper proposes a proxemics-based reward formulation for deep reinforcement learning social navigation, modeling human personal space as Gaussian-mixture fields to improve social compliance while maintaining navigation efficiency.

0 favorites 0 likes
#reward-modeling

Procedural Fairness Failures in RLHF from Preference Averaging

arXiv cs.LG ↗ · 2026-08-12 Cached

This paper identifies procedural fairness failures in RLHF caused by averaging heterogeneous preferences, where majority groups dominate reward learning and minority preferences are under-represented. It proposes Preference-Aware RLHF (PA-RLHF), which improves alignment accuracy and reduces the fairness gap in controlled experiments.

0 favorites 0 likes
#reward-modeling

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

arXiv cs.AI ↗ · 2026-08-03 Cached

This paper introduces LEMUR, a framework that combines multi-objective reinforcement learning with preference-based learning from multiple human feedback to learn Pareto-optimal policies without predefined reward functions.

0 favorites 0 likes
#reward-modeling

DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification

arXiv cs.CL ↗ · 2026-07-29 Cached

This paper presents a system for CLEF 2026 CheckThat! Task 2 that uses LLM-based trace ranking and grouped reward modeling for verifying numerical claims in English and Arabic, comparing fine-tuned verifiers with lightweight reward models.

0 favorites 0 likes
#reward-modeling

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

arXiv cs.CL ↗ · 2026-07-27 Cached

This survey provides a unified view of progress reward modeling for robotic learning, organizing the field into three steps: interface, methods, and data/benchmarks.

0 favorites 0 likes
#reward-modeling

Normalized Rewards for Preference Optimization

arXiv cs.LG ↗ · 2026-07-21 Cached

This paper introduces a regularization technique for Direct Alignment Algorithms (DAAs) that maintains normalized response probabilities, mitigating over-optimization and likelihood displacement. The method improves generation quality and benchmark performance, achieving over 20% relative increase on AlpacaEval2 and 9% gains on general benchmarks for Llama-3.1-8B-Instruct.

0 favorites 0 likes
#reward-modeling

Rater State Bias in RLHF Preference Data: An Audit Framework

arXiv cs.AI ↗ · 2026-07-21 Cached

This paper identifies and formalizes rater state bias in RLHF preference data, where annotator emotional state can confound preference labels. It proposes an audit framework with falsifiable predictions to detect such biases.

0 favorites 0 likes
#reward-modeling

Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models

arXiv cs.AI ↗ · 2026-07-16 Cached

This paper introduces DROPJ, a human-centred method for safely training and deploying agent policies by learning a world model from real-world trajectories, then eliciting human preferences with justifications to train a reward model for model predictive control. Experiments show that using human-generated simulated trajectories and justifications improves safety and reduces computational cost.

0 favorites 0 likes
#reward-modeling

Efficient Multilingual Reasoning Transfer via Progressive Code-Switching

arXiv cs.CL ↗ · 2026-07-02 Cached

This paper introduces Progressive Code-Switching (PCS), a reinforcement learning approach with curriculum learning that gradually increases code-switching in LLMs to efficiently transfer multilingual reasoning capabilities.

0 favorites 0 likes
#reward-modeling

@gabepereyra: Harvey partnered with @appliedcompute to train a legal agent. We optimized each part of the agent stack, including the …

X AI KOLs Following ↗ · 2026-06-22 Cached

Harvey partnered with Applied Compute to train a legal agent, optimizing the agent stack and post-training the GLM-5.1 model using reward signals from their Legal Agent Benchmark.

0 favorites 0 likes
#reward-modeling

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling

arXiv cs.AI ↗ · 2026-06-09 Cached

This paper proposes PAFO, a Pareto fairness optimization framework to mitigate personalized reward bias in reward models for LLMs, improving accuracy for minority user groups without harming majority groups.

0 favorites 0 likes
#reward-modeling

Improving Multimodal Reasoning via Worst Dimension Optimization

arXiv cs.AI ↗ · 2026-06-09 Cached

This paper introduces Multimodal Multi-Dimensional Scalarization Process Reward Modeling (MMS-PRM), which enforces the worst dimension's robustness in multimodal reasoning to prevent failures like visual hallucinations from being masked by strong text logic.

0 favorites 0 likes
#reward-modeling

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling

arXiv cs.CL ↗ · 2026-06-08 Cached

Eval-Skill is an exploration-guided method that synthesizes reusable evaluation skills for reward modeling, achieving significant gains on RewardBench 2 over existing backbones.

0 favorites 0 likes
#reward-modeling

Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

Hugging Face Daily Papers ↗ · 2026-06-08 Cached

Z-Reward is a teacher-student framework that decouples complex reasoning from efficient reward deployment for text-to-image training. It achieves 89.6% human preference accuracy with a 27B teacher and 88.6% with a 9B student, outperforming prior methods.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback