Tag
This paper proposes a Hessian-free hypergradient-based bilevel reinforcement learning algorithm that achieves state-of-the-art sample complexity and removes the PL condition assumption in convergence analysis.