bandit-feedback

Tag

Cards List
#bandit-feedback

Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback

arXiv cs.LG · 2026-05-27 Cached

This paper proves that online gradient descent achieves optimal √T regret for hidden-convex losses under a Hessian compatibility condition, resolving open questions in adversarial online learning. It also extends results to one-point bandit feedback with a T^{3/4} expected regret bound.

0 favorites 0 likes
#bandit-feedback

Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback

arXiv cs.LG · 2026-05-25 Cached

This paper introduces Good Policy Identification (GPI) in reinforcement learning, aiming to find a policy meeting a reward threshold rather than the optimal one, and proposes the BEE-GPI algorithm with near-optimal sample complexity guarantees.

0 favorites 0 likes
← Back to home

Submit Feedback