Tag
This paper examines why evolutionarily stable cooperative outcomes in multi-agent systems may not be reachable through decentralized reinforcement learning algorithms, demonstrating distinct properties between stability and learning accessibility.
The paper proposes FRAC-MARL, a decentralized actor-critic multi-agent reinforcement learning method that achieves full Byzantine resilience by leveraging redundancy in communication, ensuring convergence to optimal parameters even under adversarial attacks.
DualSQL proposes a multi-agent reinforcement learning framework for Text-to-SQL, using a single shared model to jointly optimize schema linking and SQL generation, achieving state-of-the-art accuracy with smaller model sizes.
This paper proposes an LLM-enhanced multi-agent reinforcement learning framework to simultaneously optimize electric vehicle charging scheduling, station profitability, and grid stability, using LLMs for feature selection and adaptive weighting, outperforming state-of-the-art methods with reduced training time.
JaxAHT is an open-source JAX-based library that accelerates and standardizes Ad Hoc Teamwork research, providing a unified framework for teammate generation, training, and evaluation with significant performance improvements and a suite of evaluation teammates.
A hierarchical multi-agent reinforcement learning framework combining graph attention and dynamic role assignment improves tactical coordination and win rates in air combat.
This paper explores the impact of network topology and opponent information on the emergence of cooperation in multi-agent reinforcement learning systems, specifically in the Iterated Prisoner's Dilemma, finding that graph structure and information availability significantly influence cooperative strategies.
This paper proposes SIGMA, a hierarchical collaboration framework for cooperative multi-agent reinforcement learning that learns robust representations under noisy observations by exploiting cooperation structures through density-based grouping and aggregation methods.
This paper studies decentralized multi-player Q-learning in episodic Markov decision processes under three forms of information asymmetry, proposing algorithms that achieve regret bounds matching the single-agent Q-learning rate up to logarithmic factors.
This paper proposes VGG-MADiffRL, a value-gradient-guided multi-agent diffusion reinforcement learning algorithm, and MDCA, a hierarchical control architecture, for cooperative target tracking in multi-AUV ad-hoc networks under constrained acoustic communication and dynamic underwater disturbances.
Proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework for heterogeneous USV cooperative pursuit in constrained port waterways. The MASAC instantiation achieves a 75% capture rate and shows promising zero-shot transfer to a real map scenario.
This paper introduces PLATO, a pointer-network-based actor with a graph neural network critic for multi-agent reinforcement learning that handles both agent and task openness without retraining, evaluated in a wildfire suppression domain.
This paper presents a deep Q-network-based multi-agent reinforcement learning framework for decentralized conflict resolution among heterogeneous small UAVs and eVTOL aircraft operating under degraded surveillance conditions, evaluating policies across 90 combinations of traffic density and separation thresholds.
This paper proposes EffRank/n and D_act as low-overhead diagnostics to measure effects of reward attribution in cooperative multi-agent RL, and tests on SMACv2, finding that observation explains geometry while reward attribution mainly affects behavior.
This paper compares contextual combinatorial bandits and policy gradient algorithms for decentralized smart charging of large EV fleets, using a realistic simulation with dynamic pricing and renewable energy data.
HyPOLE introduces a framework for multi-agent reinforcement learning under partial observability that uses hyperproperty-guided learning via HyperLTL temporal logic, integrated with centralized training for decentralized execution, and demonstrates improvements over baselines on SMAC, MessySMAC, and WildFire benchmarks.
Introduces R2D-RL, a reinforcement learning environment that connects the RoboCup 2D Soccer Simulation server to Python-based MARL workflows via shared-memory communication, supporting full-field and scenario-based training with configurable opponents and reward shaping.
TRIDENT is a novel multi-agent reinforcement learning framework that breaks the coupling between hybrid discrete-continuous actions, hard safety constraints, and physics-governed dynamics, achieving provably safe coordination with a convergence guarantee to a constrained Nash equilibrium and significant reductions in training-time violations.
A method for contract-based compositional shielding that ensures global safety in multi-agent reinforcement learning without centralized runtime control, using local LTL obligations and a multi-armed bandit to optimize team reward.
This paper introduces a framework for two-sided matching with temporally extended feedback, formulating it as a partially observable Markov game with costly screening, noisy observations, and evolving latent profiles. The authors present Learn2Match, a multi-agent reinforcement learning benchmark, and show that independent PPO outperforms bandit baselines in social welfare but incurs higher information-friction loss.