Tag
This paper studies multi-agent multi-armed bandits with heavy-tailed rewards under three information-asymmetry regimes, proposing robust decentralized algorithms with regret guarantees nearly matching centralized rates, and validating them on Pareto-distributed reward environments.
This paper introduces an online variant of assistance games and provides the first provably efficient learning algorithms for both the human and assistant agents, achieving near-optimal regret bounds.