reward-model

Tag

Cards List
#reward-model

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

Hugging Face Daily Papers · 2026-08-10 Cached

RynnValue introduces an open-source robotic value foundation model that uses temporal distance as a scalable supervision target for reward learning, surpassing preference-supervised state-of-the-art on RBM-EVAL-OOD and improving real-world policy success rates.

0 favorites 0 likes
#reward-model

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

Hugging Face Daily Papers · 2026-07-31 Cached

This paper introduces a Multi-dimensional Evaluation-Verification Reward (EVR) for reinforcement learning fine-tuning of multi-reference image editing models, improving visual consistency and harmony.

0 favorites 0 likes
#reward-model

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

arXiv cs.AI · 2026-07-28 Cached

Introduces VlogReward, a reward model for evaluating vlog editing plans across six dimensions, along with a large-scale dataset and benchmark, achieving state-of-the-art results against GPT-5 and Gemini-3-Pro.

0 favorites 0 likes
#reward-model

Codifying the Judge: Scalable Evaluation via Program Distillation

arXiv cs.AI · 2026-07-28 Cached

This paper introduces PAJAMA, a system that distills LLM-as-a-judge decision logic into programmatic judges, eliminating per-sample API costs while matching the performance of a 13B LLM judge, and providing transparency and efficiency improvements for scalable evaluation.

0 favorites 0 likes
#reward-model

How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF

Hugging Face Daily Papers · 2026-07-22 Cached

This paper presents a systems study comparing C++ and PyTorch inference runtimes for reward model scoring in RLHF pipelines, finding that ONNXRuntime provides speedups on CPU while torch.compile leads on GPU, with batching strategy mattering more than language or runtime.

0 favorites 0 likes
#reward-model

@sheriyuo: TRACE assigns dense credit at tool boundaries by asking a frozen reference model whether each new observation raises th…

X AI KOLs Timeline · 2026-07-16 Cached

TRACE is a dense credit assignment method for multi-turn agentic reinforcement learning that uses a frozen reference model to compute per-action rewards from log-probability changes at tool boundaries, eliminating the need for a critic or process reward model. It significantly improves long-horizon tool-use performance on benchmarks like BrowseComp-Plus.

0 favorites 0 likes
#reward-model

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

Hugging Face Daily Papers · 2026-07-13 Cached

This paper introduces SpectraReward, a training-free reward function that leverages pretrained multimodal large language models (MLLMs) as zero-shot reward models for reinforcement learning in text-to-image generation, demonstrating consistent improvements over prior methods.

0 favorites 0 likes
#reward-model

PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration

arXiv cs.LG · 2026-06-29 Cached

Introduces PEBS, a per-rater empirical-Bayes shrinkage estimator for calibrating reward models in RLHF, reducing within-user RMSE by over 8.5% on PRISM and over 9.6% on PluriHarms.

0 favorites 0 likes
#reward-model

TuneJury: An Open Metric for Improving Music Generation Preference Alignment

Hugging Face Daily Papers · 2026-06-15 Cached

TuneJury is an open-source pairwise reward model for text-to-music generation that provides calibrated preference scoring and generalizes across multiple downstream applications.

0 favorites 0 likes
#reward-model

@neural_avb: Locally generating GRPO-like rollouts with my SLM, and using this tiny RM as the rubric. Next I'll be RL training on fr…

X AI KOLs Timeline · 2026-06-11 Cached

Neural_avb releases a lightweight Answer-eq Reward Model for RL training on QA tasks, claiming 80% agreement with external judge LM and faster than F1/ROUGE/BertScore.

0 favorites 0 likes
#reward-model

From Long News to Accurate Forecast: Importance-Aware Fusion and PRM-Guided Reflection for Time Series Forecasting

arXiv cs.AI · 2026-06-03 Cached

This paper introduces a framework for time series forecasting that uses importance-aware news compression and process reward model-guided retrieval to incorporate long news articles within fixed context limits, improving prediction accuracy across finance, energy, traffic, and Bitcoin benchmarks.

0 favorites 0 likes
#reward-model

Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics

arXiv cs.CL · 2026-06-03 Cached

This paper introduces a method to predict best-of-N inference scaling gains for language models using cheap statistics from a single labeled validation-set sampling pass. A compact predictor with three core features achieves Spearman ρ=0.90 with actual gains, enabling screening of configurations before expensive reward-model scoring.

0 favorites 0 likes
#reward-model

Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs

arXiv cs.AI · 2026-06-02 Cached

Introduces Latent Reward Steering (Lrs), an adaptive inference-time framework that uses sparse autoencoder latent states and a learned reward model to implicitly promote cognitive behaviors like verification and backtracking in reasoning LLMs, improving performance across multiple models and benchmarks.

0 favorites 0 likes
#reward-model

Configurable Reward Model for Balanced Safety Alignment

arXiv cs.CL · 2026-06-01 Cached

This paper introduces the Configurable Safety Reward Model (CSRM), a reward model that can be configured to accommodate heterogeneous and evolving safety requirements for LLM alignment. CSRM achieves state-of-the-art results on configurable safety benchmarks and improves the helpfulness-safety tradeoff.

0 favorites 0 likes
#reward-model

KARMA: Karma-Aligned Reward Model Adaptation

arXiv cs.CL · 2026-05-27 Cached

Introduces KARMA, a framework that trains a reward model on Reddit conversations to improve LLMs' context-sensitive conversational behavior via reinforcement learning, finding that the best reward model for predicting karma does not yield the best downstream alignment.

0 favorites 0 likes
#reward-model

Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation

arXiv cs.AI · 2026-05-27 Cached

This paper identifies weighting noise in LLM judges for multi-stakeholder tasks and proposes DecompR, a method that decouples utility estimation from aggregation using counterfactually calibrated weights.

0 favorites 0 likes
#reward-model

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

arXiv cs.CL · 2026-05-27 Cached

This paper introduces CroCo, a method for cross-lingual contrastive preference tuning on self-generated responses, showing that a reward model trained on English preferences can effectively rank responses in other languages, improving model performance across 14 languages without language-specific annotations.

0 favorites 0 likes
#reward-model

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

Hugging Face Daily Papers · 2026-05-26 Cached

This paper introduces alignment tampering, a vulnerability in RLHF where language models can manipulate preference datasets to amplify misaligned biases, demonstrating experimentally across biases like sexism, brand promotion, and goal-seeking, and showing that existing mitigation techniques are insufficient.

0 favorites 0 likes
#reward-model

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment

Hugging Face Daily Papers · 2026-05-20 Cached

AutoRubric-T2I automatically generates and selects explicit rubrics to guide Vision-Language Model judges for text-to-image generation, achieving high-quality reward signals with minimal human annotation and improving generation quality in downstream tasks.

0 favorites 0 likes
#reward-model

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Hugging Face Daily Papers · 2026-04-17 Cached

VEFX-Bench introduces a large-scale human-annotated video editing dataset (5,049 examples) with multi-dimensional quality labels and a specialized reward model for standardized evaluation of video editing systems. The paper addresses the lack of comprehensive benchmarks in AI-assisted video creation by providing VEFX-Dataset, VEFX-Reward, and a 300-video-prompt benchmark that reveals gaps in current editing models.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback