Tag
This paper studies decentralized multi-player Q-learning in episodic Markov decision processes under three forms of information asymmetry, proposing algorithms that achieve regret bounds matching the single-agent Q-learning rate up to logarithmic factors.
This paper proposes VGG-MADiffRL, a value-gradient-guided multi-agent diffusion reinforcement learning algorithm, and MDCA, a hierarchical control architecture, for cooperative target tracking in multi-AUV ad-hoc networks under constrained acoustic communication and dynamic underwater disturbances.
Proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework for heterogeneous USV cooperative pursuit in constrained port waterways. The MASAC instantiation achieves a 75% capture rate and shows promising zero-shot transfer to a real map scenario.
This paper introduces PLATO, a pointer-network-based actor with a graph neural network critic for multi-agent reinforcement learning that handles both agent and task openness without retraining, evaluated in a wildfire suppression domain.
This paper presents a deep Q-network-based multi-agent reinforcement learning framework for decentralized conflict resolution among heterogeneous small UAVs and eVTOL aircraft operating under degraded surveillance conditions, evaluating policies across 90 combinations of traffic density and separation thresholds.
This paper proposes EffRank/n and D_act as low-overhead diagnostics to measure effects of reward attribution in cooperative multi-agent RL, and tests on SMACv2, finding that observation explains geometry while reward attribution mainly affects behavior.
This paper compares contextual combinatorial bandits and policy gradient algorithms for decentralized smart charging of large EV fleets, using a realistic simulation with dynamic pricing and renewable energy data.
HyPOLE introduces a framework for multi-agent reinforcement learning under partial observability that uses hyperproperty-guided learning via HyperLTL temporal logic, integrated with centralized training for decentralized execution, and demonstrates improvements over baselines on SMAC, MessySMAC, and WildFire benchmarks.
Introduces R2D-RL, a reinforcement learning environment that connects the RoboCup 2D Soccer Simulation server to Python-based MARL workflows via shared-memory communication, supporting full-field and scenario-based training with configurable opponents and reward shaping.
TRIDENT is a novel multi-agent reinforcement learning framework that breaks the coupling between hybrid discrete-continuous actions, hard safety constraints, and physics-governed dynamics, achieving provably safe coordination with a convergence guarantee to a constrained Nash equilibrium and significant reductions in training-time violations.
A method for contract-based compositional shielding that ensures global safety in multi-agent reinforcement learning without centralized runtime control, using local LTL obligations and a multi-armed bandit to optimize team reward.
This paper introduces a framework for two-sided matching with temporally extended feedback, formulating it as a partially observable Markov game with costly screening, noisy observations, and evolving latent profiles. The authors present Learn2Match, a multi-agent reinforcement learning benchmark, and show that independent PPO outperforms bandit baselines in social welfare but incurs higher information-friction loss.
This paper presents a distributed approach for constrained multi-agent reinforcement learning that uses state-augmented policy learning and neighbor-to-neighbor consensus over dual variables to satisfy global resource constraints while scaling linearly with the number of agents. Experiments on smart grid demand response demonstrate that consensus coordination is essential for feasibility, scaling to thousands of agents unlike centralized training approaches.
This paper introduces Differentiable Belief-based Opponent Shaping (D-BOS), a first-order method that treats observer beliefs as the shaped state and differentiates through belief update dynamics, allowing optimal strategies to emerge naturally from the environment's reward structure in hidden-role multi-agent settings.
This paper proposes a multi-agent reinforcement learning framework that co-trains an autonomous vehicle and pedestrians with personality-driven jaywalking behavior, achieving a 30% reduction in collisions compared to single-agent approaches and demonstrating more realistic interaction scenarios.
This paper introduces SLIM, a minimal architecture that decouples communication from policy representation in multi-agent reinforcement learning, achieving state-of-the-art performance under bandwidth constraints with minimal degradation.
This paper presents empirical evidence that quantum entanglement provides a measurable advantage in multi-agent reinforcement learning, using the CHSH game and cooperative navigation tasks to demonstrate performance improvements over classical baselines.
The paper introduces Diamond Attention, a method for multi-agent reinforcement learning that uses structured randomness to break symmetry and enable role differentiation among homogeneous agents, achieving perfect coordination in symmetric tasks like the XOR game.