Tag
This paper studies a human-AI service system with an automated chatbot and human agents, proposing a UCB-DPP policy that learns unknown parameters and achieves regret Õ(K√T) while stabilizing queues.
OpenAI presents a novel exploration strategy for deep reinforcement learning using ensembles of Q-functions with upper-confidence bounds (UCB), demonstrating significant performance improvements on the Atari benchmark.