Tag
The paper introduces a vector Bellman theory for robust average-reward Markov decision processes, enabling optimal performance under transition uncertainty with finite model solvability and an approximately shifted Halpern planning algorithm.
This paper introduces a constant-aware comparison protocol for average-reward reinforcement learning regret bounds, deriving an explicit finite lower certificate for communicating MDPs and improving published coefficients.
This paper studies the sample complexity of robust average-reward Markov decision processes, deriving minimax-optimal learning rates via plug-in reductions under total-variation uncertainty sets.