ode-solver

Tag

Cards List
#ode-solver

GenPO++: Generative Policy Optimization with Jacobian-free Likelihood Ratios

arXiv cs.LG ↗ · 2026-06-08 Cached

GenPO++ proposes a reversible generative policy optimization framework that uses history states as auxiliary memory in a high-order reversible ODE solver, enabling exact inversion and Jacobian-free likelihood-ratio computation for flow-based policies in reinforcement learning. It achieves competitive performance on large-scale control, fine-tuning, and real-world robotic tasks while improving stability and efficiency.

0 favorites 0 likes
← Back to home

Submit Feedback