langevin-correction

Tag

Cards List
#langevin-correction

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper introduces LC-GRPO, a flow-based GRPO framework with Langevin correction that bridges the train-inference gap by aligning stochastic training rollouts with deterministic ODE sampling, improving reward optimization on models like SD3.5, FLUX.1-Dev, and HunyuanVideo.

0 favorites 0 likes
← Back to home

Submit Feedback