reward-overoptimization

Tag

Cards List
#reward-overoptimization

NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

arXiv cs.LG · 2026-06-29 Cached

This paper proposes NormGuard, a norm-budget regularizer that suppresses reward-irrelevant velocity norm inflation during RL post-training of flow-matching generative models, improving perceptual image quality without sacrificing reward.

0 favorites 0 likes
#reward-overoptimization

EvalStop: Using World Feedback to Detect and Correct Reward Overoptimization in Multi-Tenant RLHF Platforms

arXiv cs.LG · 2026-06-04 Cached

EvalStop is a scheduling primitive for multi-tenant RLHF platforms that detects and corrects reward overoptimization by monitoring downstream evaluation scores and terminating jobs on consecutive declines, achieving 98% precision and 99% recall while improving job completion time by 9% and cutting wasted compute by 22%.

0 favorites 0 likes
← Back to home

Submit Feedback