Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints

Hugging Face Daily Papers Papers

Summary

This paper introduces SLIM, a minimal architecture that decouples communication from policy representation in multi-agent reinforcement learning, achieving state-of-the-art performance under bandwidth constraints with minimal degradation.

Communication enables coordination in multi-agent reinforcement learning (MARL), but many real-world applications, e.g., search-and-rescue with drone swarms, operate under severe bandwidth constraints. Many communication architectures still expose a coupled bottleneck in which a shared latent representation is used for both policy execution and inter-agent communication. Consequently, reducing message size directly limits the policy's latent space, often leading to significant performance degradation. We address this with two contributions. First, we introduce β, a normalised per-agent bandwidth budget that unifies sparsity, rounds, and message dimension into a single comparable constraint. Second, we provide SLIM, a minimal architecture that decouples the communication pathway from the policy's latent representation, allowing us to isolate the effect of bandwidth from the effect of policy capacity while benefiting from in-step communication. We evaluate our method on several partially-observable MARL benchmarks, where communication is essential. Our approach achieves state-of-the-art performance and exhibits scalability and robustness under limited communication, with only marginal degradation as bandwidth is reduced.
Original Article
View Cached Full Text

Cached at: 05/26/26, 02:44 PM

Paper page - Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints

Source: https://huggingface.co/papers/2605.21085

Abstract

Researchers propose a novel communication architecture for multi-agent reinforcement learning that decouples policy representation from communication pathways, enabling better performance under bandwidth constraints.

Communication enables coordination in multi-agent reinforcement learning (MARL), but many real-world applications, e.g., search-and-rescue with drone swarms, operate under severebandwidth constraints. Manycommunication architecturesstill expose a coupled bottleneck in which a sharedlatent representationis used for bothpolicy executionandinter-agent communication. Consequently, reducing message size directly limits the policy’s latent space, often leading to significant performance degradation. We address this with two contributions. First, we introduce β, a normalised per-agent bandwidth budget that unifiessparsity,rounds, andmessage dimensioninto a single comparable constraint. Second, we provideSLIM, a minimal architecture that decouples the communication pathway from the policy’slatent representation, allowing us to isolate the effect of bandwidth from the effect of policy capacity while benefiting from in-step communication. We evaluate our method on severalpartially-observable MARLbenchmarks, where communication is essential. Our approach achieves state-of-the-art performance and exhibits scalability and robustness under limited communication, with only marginal degradation as bandwidth is reduced.

View arXiv pageView PDFGitHub1Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2605.21085 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2605.21085 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.21085 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Fully Byzantine-Resilient Multi-Agent Reinforcement Learning

arXiv cs.LG

The paper proposes FRAC-MARL, a decentralized actor-critic multi-agent reinforcement learning method that achieves full Byzantine resilience by leveraging redundancy in communication, ensuring convergence to optimal parameters even under adversarial attacks.

Scalable Constrained Multi-Agent Reinforcement Learning via State Augmentation and Consensus for Separable Dynamics

arXiv cs.LG

This paper presents a distributed approach for constrained multi-agent reinforcement learning that uses state-augmented policy learning and neighbor-to-neighbor consensus over dual variables to satisfy global resource constraints while scaling linearly with the number of agents. Experiments on smart grid demand response demonstrate that consensus coordination is essential for feasibility, scaling to thousands of agents unlike centralized training approaches.

Safety of Latent Communication in Multi-Agent Systems

Hugging Face Daily Papers

This paper shows that training lightweight latent communication links between LLM agents in multi-agent systems can increase harmful compliance, even when agents remain safety-aligned, and demonstrates an RL-based attack raising harmful compliance from 27.9 to 76.9 across topologies and benchmarks, while also proposing a way to repair compromised links without updating the agents.