Tag
This paper introduces adaptive probabilistic shielding for safe reinforcement learning, where the shield is computed from an online learned MDP model, adapting as the model becomes more accurate.
This paper introduces a constant-aware comparison protocol for average-reward reinforcement learning regret bounds, deriving an explicit finite lower certificate for communicating MDPs and improving published coefficients.
This paper provides the first finite-time convergence guarantees for the Natural Policy Gradient algorithm in finite-horizon Markov Decision Processes, proving sublinear and linear convergence rates under different step size regimes.