Tag
This paper proposes NormGuard, a norm-budget regularizer that suppresses reward-irrelevant velocity norm inflation during RL post-training of flow-matching generative models, improving perceptual image quality without sacrificing reward.
EvalStop is a scheduling primitive for multi-tenant RLHF platforms that detects and corrects reward overoptimization by monitoring downstream evaluation scores and terminating jobs on consecutive declines, achieving 98% precision and 99% recall while improving job completion time by 9% and cutting wasted compute by 22%.