Tag
The paper shows that the Bernstein CVaR-UCBVI algorithm achieves a near-minimax leading-order regret bound for CVaR reinforcement learning without continuity assumptions on return laws.
This paper studies the joint effect of memory width and batch depth in stochastic Lipschitz bandits, characterizing the minimax pseudo-regret tradeoff up to logarithmic factors and showing that state width and update depth are not interchangeable.