good-policy-identification

Tag

Cards List
#good-policy-identification

Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback

arXiv cs.LG · 2026-05-25 Cached

This paper introduces Good Policy Identification (GPI) in reinforcement learning, aiming to find a policy meeting a reward threshold rather than the optimal one, and proposes the BEE-GPI algorithm with near-optimal sample complexity guarantees.

0 favorites 0 likes
← Back to home

Submit Feedback