Tag
This paper studies decentralized multi-player Q-learning in episodic Markov decision processes under three forms of information asymmetry, proposing algorithms that achieve regret bounds matching the single-agent Q-learning rate up to logarithmic factors.
This paper studies multi-agent multi-armed bandits with heavy-tailed rewards under three information-asymmetry regimes, proposing robust decentralized algorithms with regret guarantees nearly matching centralized rates, and validating them on Pareto-distributed reward environments.
This paper introduces InfoDelphi, a framework that uses information asymmetry (partitioning evidence into shared public and disjoint private subsets) to improve multi-agent LLM deliberation and forecasting. On the PolyGym benchmark, it outperforms single-agent and multi-agent baselines by 12-18% in Brier score and 4-8 percentage points in accuracy, demonstrating that diverse evidence is key to effective multi-agent reasoning.
The paper introduces GPTNT, a benchmark built on Keep Talking and Nobody Explodes that requires two multimodal agents to collaborate in real-time under time pressure and information asymmetry, revealing critical weaknesses in state-of-the-art systems.
This paper formalizes communication policy for LLM agents and proposes Communication Policy Evolution (CPE), a self-evolution framework that refines communication policies through rollout and prompt-level evolving, achieving best task success across multiple settings.
The post mentions that AI companies like Moonshot AI and DeepSeek are hiring agent harness talent, and promotes the job aggregation platform jobleap.cn, believing that mastering the latest job opportunities is key.