Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints
Summary
This paper introduces SLIM, a minimal architecture that decouples communication from policy representation in multi-agent reinforcement learning, achieving state-of-the-art performance under bandwidth constraints with minimal degradation.
View Cached Full Text
Cached at: 05/26/26, 02:44 PM
Paper page - Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints
Source: https://huggingface.co/papers/2605.21085
Abstract
Researchers propose a novel communication architecture for multi-agent reinforcement learning that decouples policy representation from communication pathways, enabling better performance under bandwidth constraints.
Communication enables coordination in multi-agent reinforcement learning (MARL), but many real-world applications, e.g., search-and-rescue with drone swarms, operate under severebandwidth constraints. Manycommunication architecturesstill expose a coupled bottleneck in which a sharedlatent representationis used for bothpolicy executionandinter-agent communication. Consequently, reducing message size directly limits the policy’s latent space, often leading to significant performance degradation. We address this with two contributions. First, we introduce β, a normalised per-agent bandwidth budget that unifiessparsity,rounds, andmessage dimensioninto a single comparable constraint. Second, we provideSLIM, a minimal architecture that decouples the communication pathway from the policy’slatent representation, allowing us to isolate the effect of bandwidth from the effect of policy capacity while benefiting from in-step communication. We evaluate our method on severalpartially-observable MARLbenchmarks, where communication is essential. Our approach achieves state-of-the-art performance and exhibits scalability and robustness under limited communication, with only marginal degradation as bandwidth is reduced.
View arXiv pageView PDFGitHub1Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.21085 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.21085 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.21085 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Fully Byzantine-Resilient Multi-Agent Reinforcement Learning
The paper proposes FRAC-MARL, a decentralized actor-critic multi-agent reinforcement learning method that achieves full Byzantine resilience by leveraging redundancy in communication, ensuring convergence to optimal parameters even under adversarial attacks.
When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs
This paper studies when end-to-end reinforcement learning training improves multi-agent LLM workflows, comparing shared-policy and isolated-policy training across different workflows, tasks, and model scales, revealing conditional tradeoffs.
Scalable Constrained Multi-Agent Reinforcement Learning via State Augmentation and Consensus for Separable Dynamics
This paper presents a distributed approach for constrained multi-agent reinforcement learning that uses state-augmented policy learning and neighbor-to-neighbor consensus over dual variables to satisfy global resource constraints while scaling linearly with the number of agents. Experiments on smart grid demand response demonstrate that consensus coordination is essential for feasibility, scaling to thousands of agents unlike centralized training approaches.
SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing
SLIM-RL introduces a risk-budgeted random-masking reinforcement learning method for diffusion LLMs that avoids trajectory slicing, achieving state-of-the-art results on math and code benchmarks with significantly fewer training samples.
Safety of Latent Communication in Multi-Agent Systems
This paper shows that training lightweight latent communication links between LLM agents in multi-agent systems can increase harmful compliance, even when agents remain safety-aligned, and demonstrates an RL-based attack raising harmful compliance from 27.9 to 76.9 across topologies and benchmarks, while also proposing a way to repair compromised links without updating the agents.