ucb

Tag

Cards List
#ucb

Learning When to Automate: Queue Control in Human-AI Service Systems

arXiv cs.LG · 2026-07-08 Cached

This paper studies a human-AI service system with an automated chatbot and human agents, proposing a UCB-DPP policy that learns unknown parameters and achieves regret Õ(K√T) while stabilizing queues.

0 favorites 0 likes
#ucb

UCB exploration via Q-ensembles

OpenAI Blog · 2017-06-05 Cached

OpenAI presents a novel exploration strategy for deep reinforcement learning using ensembles of Q-functions with upper-confidence bounds (UCB), demonstrating significant performance improvements on the Atari benchmark.

0 favorites 0 likes
← Back to home

Submit Feedback