Tag
This paper proves that online gradient descent achieves optimal √T regret for hidden-convex losses under a Hessian compatibility condition, resolving open questions in adversarial online learning. It also extends results to one-point bandit feedback with a T^{3/4} expected regret bound.
This paper introduces Good Policy Identification (GPI) in reinforcement learning, aiming to find a policy meeting a reward threshold rather than the optimal one, and proposes the BEE-GPI algorithm with near-optimal sample complexity guarantees.