标签
本文表明,Bernstein CVaR-UCBVI算法在CVaR强化学习中实现了无回报律连续性假设的近极小极大领先阶遗憾界限。
This paper studies the joint effect of memory width and batch depth in stochastic Lipschitz bandits, characterizing the minimax pseudo-regret tradeoff up to logarithmic factors and showing that state width and update depth are not interchangeable.