Tag
This paper proposes CSDG, a method for offline reinforcement learning that expresses Bellman backups as in-sample targets plus a convex-hull-neighborhood local correction, controlling OOD action estimation errors and improving value stability.
The paper presents World Value Model (WVM), a generalist robotic value model that combines world models with value estimation to accurately assess task progression and improve robotic policy learning from mixed-quality data, achieving state-of-the-art results on standard benchmarks and a new suboptimal data benchmark.
This paper introduces POISE, a method for stable policy optimization in large reasoning models by estimating baselines using the model's own internal states, reducing computational overhead compared to PPO and GRPO.