@XuXander24218: StreamMA: Making Multi-Agent Systems Faster and More Accurate! Hey everyone! Our team just released StreamMA. It is a n…
Summary
StreamMA is a multi-agent reasoning system that streams intermediate results step-by-step to improve latency and accuracy, achieving up to 26.9× speedup and +7.3% performance improvement on benchmarks.
View Cached Full Text
Cached at: 06/05/26, 01:17 PM
StreamMA: Making Multi-Agent Systems Faster and More Accurate!
Hey everyone! Our team just released StreamMA. It is a new way to make multi-agent systems both faster and more accurate. The core idea is simple but powerful: Instead of waiting for one agent to finish its full response, we now send information step by step. Each reasoning step goes to the next agent right away.
Why does this work so well? The upstream agent sends one step, and the downstream agents start working immediately. This makes everything much faster. Surprisingly, the reasoning quality also gets better. Complex reasoning steps are not all the same quality. Early steps are usually reliable and clear. Later steps are more likely to have errors. StreamMA lets downstream agents start with the good early steps. By the time any mistakes arrive, the other agents have already built their own strong reasoning path. This naturally reduces the impact of errors.
Key Contributions: ① Stream Protocol: We change communication from full responses to single steps. This improves both speed and performance. ② Three closed-form theorems: We show when Stream works best, the theoretical speedup limit, and the cost ratio. ③ Step-level scaling law: With the same number of agents, adding more steps improves both accuracy and speed. It works well together with traditional agent scaling.
Strong Experimental Results: • Performance: Average +7.3 percentage points better on 8 math, science, and code benchmarks (Opus 4.6) • Speedup: With 64 agents and 64 steps, we get 26.9× faster inference (theoretical max is 32.3×, we reached 83% of it) (HMMT26 GPT5.4)
Want to learn more? Paper: https://huggingface.co/papers/2606.05158… Project Page: https://zhenyangcs.github.io/StreamMA-website… GitHub: https://github.com/EnVision-Research/StreamMA…
If you work on multi-agent systems, distributed reasoning, or long-chain inference, we would love to hear your thoughts! This streaming idea brings new possibilities. #StreamMA #MultiAgent #AI #Reasoning #LLM
Paper page - Streaming Communication in Multi-Agent Reasoning
Source: https://huggingface.co/papers/2606.05158
Abstract
StreamMA enables efficient multi-agent reasoning by streaming intermediate results and leveraging reliable early steps to improve both latency and effectiveness across various reasoning tasks.
Multi-agent reasoning systemsadopt a “generate-then-transfer” paradigm that forcesend-to-end latencyto scale linearly with pipeline depth. We introduce StreamMA, a multi-agent reasoning system that streams each reasoning step to downstream agents as soon as it is generated,pipeliningadjacent agents and thus reducing latency. Surprisingly, thispipeliningalso improves effectiveness: because multi-step reasoning quality is non-uniform and early steps are more reliable than later ones, working with these reliable early steps instead of the fullchainprevents error-prone late steps from misleading downstream agents. We formalize both advantages with the first closed-form joint analysis of stream, serial, andsingle protocols, deriving theeffectiveness ordering,speedup upper bound, andcost ratio. Across eightreasoning benchmarksspanning mathematics, science, and code, two frontierLLMs(Claude Opus 4.6 and GPT-5.4), and three topologies (Chain,Tree,Graph), StreamMA outperforms both baselines (avg. +7.3 pp, max +22.4 pp on HMMT 2026; Claude Opus 4.6-high). Beyond these contributions, we discover a “step-level scaling law”: increasing per-agent steps consistently improves both effectiveness and efficiency, a new scaling dimension orthogonal to and composable withagent-count scaling.
View arXiv pageView PDFProject pageGitHub25Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.05158 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.05158 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.05158 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Streaming Communication in Multi-Agent Reasoning
StreamMA introduces a streaming communication paradigm for multi-agent reasoning that pipelines intermediate results to reduce latency and improve effectiveness by leveraging more reliable early steps, outperforming baselines across benchmarks and revealing a step-level scaling law.
StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance
StreamMemBench is a new streaming benchmark that tests how well personal-agent memory systems use observed evidence and user feedback for future-oriented assistance. Experiments show current systems often fail to turn stored information into reliable follow-up behavior.
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
AgentStream introduces a unified framework to evaluate self-evolving LLM agents under streaming task scenarios, showing that self-evolution reliability varies across scenarios and is gated by model capability.
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
X-Stream introduces the first benchmark for multi-stream video understanding, evaluating MLLMs as multiplexers across multiple concurrent streams. The study reveals that current MLLMs achieve only about 50% accuracy, exposing significant limitations in handling multiple streams.
Recursive Multi-Agent Systems
This paper introduces RecursiveMAS, a framework that extends recursive scaling principles to multi-agent systems for improved collaborative reasoning efficiency and accuracy. It demonstrates significant speedups and token reduction across various benchmarks compared to standard baselines.