标签
This paper studies decentralized multi-player Q-learning in episodic Markov decision processes under three forms of information asymmetry, proposing algorithms that achieve regret bounds matching the single-agent Q-learning rate up to logarithmic factors.
本文介绍了 MEMOA,这是一种针对大规模在线智能体的去中心化策略。该策略通过平均场纳什均衡实现最优性,在超越贪婪基线的同时,比中心化方法具有更好的扩展性。