bandit-problem

Tag

Cards List
#bandit-problem

Continuity-Free Near-Minimax Leading-Order Regret for CVaR-UCBVI

arXiv cs.LG · 2026-09-01 Cached

The paper shows that the Bernstein CVaR-UCBVI algorithm achieves a near-minimax leading-order regret bound for CVaR reinforcement learning without continuity assumptions on return laws.

0 favorites 0 likes
← Back to home

Submit Feedback