Tag
This paper studies a human-AI service system with an automated chatbot and human agents, proposing a UCB-DPP policy that learns unknown parameters and achieves regret Õ(K√T) while stabilizing queues.