mdps

Tag

Cards List
#mdps

Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning

arXiv cs.LG ↗ · 2026-08-21 Cached

This paper introduces adaptive probabilistic shielding for safe reinforcement learning, where the shield is computed from an online learned MDP model, adapting as the model becomes more accurate.

0 favorites 0 likes
#mdps

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

arXiv cs.LG ↗ · 2026-08-11 Cached

This paper introduces a constant-aware comparison protocol for average-reward reinforcement learning regret bounds, deriving an explicit finite lower certificate for communicating MDPs and improving published coefficients.

0 favorites 0 likes
#mdps

Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

arXiv cs.LG ↗ · 2026-07-28 Cached

This paper provides the first finite-time convergence guarantees for the Natural Policy Gradient algorithm in finite-horizon Markov Decision Processes, proving sublinear and linear convergence rates under different step size regimes.

0 favorites 0 likes
← Back to home

Submit Feedback