Tag
The paper shows that the Bernstein CVaR-UCBVI algorithm achieves a near-minimax leading-order regret bound for CVaR reinforcement learning without continuity assumptions on return laws.