Tag
This paper studies KL-regularized contextual bandits and shows that greedy sampling can achieve logarithmic regret without explicit eluder-dimension dependence for both reward and preference feedback.