基于小型语言模型网络的强化学习通信
摘要
TalkMesh 是一个采用强化学习的小型语言模型智能体去中心化网络,能够学习何时及进行何种通信,在 GSM8K 和 MATH-500 等基准测试中显著优于自洽性方法,显著提升了准确性。
arXiv:2609.30578v1 Announce Type: new
Abstract: Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a problem can pass the key step to the others. We present TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate. Each agent samples a proposal and scores it with a trained confidence head. The most confident agent broadcasts a hint; agents below a confidence threshold revise, keeping each revision that outscores its proposal. Gossip consensus approximates the vote weighted by confidence without a coordinator. A talk policy, trained with group relative policy optimization on the change in correctness after revision, writes hints and revisions. With three agents, which together generate at most six outputs, the mesh reaches the accuracy of majority voting over 32 samples with each of three models. Trained with at most 8 agents and evaluated with 32, it raises accuracy from 0.568 under self-consistency to 0.705 (Qwen3.5-0.8B, GSM8K) and from 0.492 to 0.722 (SmolLM3-3B, MATH-500). When 4 of 8 agents collude on a wrong answer with fabricated confidence and poisoned hints, majority vote accuracy falls to 0.000 (Qwen3.5-0.8B, GSM8K). A defended mesh, whose agents rescore solutions with their own confidence heads, retains 0.507. Across reasoning, embodied coordination, and traffic signal control, messages improve a decision when the acting agent cannot observe the information it requires and another agent can send it.
查看缓存全文
缓存时间: 2026/09/29 09:40
# 1 Introduction
Source: [https://arxiv.org/html/2609.30578](https://arxiv.org/html/2609.30578)
Reinforcement Learning of Communication in a Mesh of Small Language Models
Mehmet Kerem Turkcan Department of Civil Engineering, Columbia University
mkt2126@columbia\.edu
Abstract
Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model’s most frequent answer\. Communication can add what sampling cannot: an agent that solves a problem can pass the key step to the others\. We present TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate\. Each agent samples a proposal and scores it with a trained confidence head\. The most confident agent broadcasts a hint; agents below a confidence threshold revise, keeping each revision that outscores its proposal\. Gossip consensus approximates the vote weighted by confidence without a coordinator\. A talk policy, trained with group relative policy optimization on the change in correctness after revision, writes hints and revisions\. With three agents, which together generate at most six outputs, the mesh reaches the accuracy of majority voting over 32 samples with each of three models\. Trained with at most 8 agents and evaluated with 32, it raises accuracy from 0\.568 under self\-consistency to 0\.705 \(Qwen3\.5\-0\.8B, GSM8K\) and from 0\.492 to 0\.722 \(SmolLM3\-3B, MATH\-500\)\. When 4 of 8 agents collude on a wrong answer with fabricated confidence and poisoned hints, majority vote accuracy falls to 0\.000 \(Qwen3\.5\-0\.8B, GSM8K\)\. A defended mesh, whose agents rescore solutions with their own confidence heads, retains 0\.507\. Across reasoning, embodied coordination, and traffic signal control, messages improve a decision when the acting agent cannot observe the information it requires and another agent can send it\.
Figure 1:Overview of TalkMesh\. \(a\) Concept: independent proposals, one hint from the most confident agent, revisions gated by the confidence head, and gossip consensus\. \(b\) Implementation: the frozen base model with two adapters, the five inference stages, and the training loop of the talk policy\.Self\-consistency\([Wang et al\., 2022](https://arxiv.org/html/2609.30578#bib.bib29)\)scales the compute of a language model at test time: it samples independent chains of thought and returns the majority answer\. Accuracy increases with the number of samples and saturates as the vote converges to the model’s most frequent answer\. Reading a peer’s reasoning lets an agent correct errors and recover omitted steps, so interaction adds information\. In multiagent debate and discussion\([Du et al\., 2023](https://arxiv.org/html/2609.30578#bib.bib6);[Liang et al\., 2024](https://arxiv.org/html/2609.30578#bib.bib16)\), a fixed set of agents reads one another’s solutions in rounds designed by hand\. Agents that converge on the same answers correlate their errors, which removes the independence the majority vote relies on\. Inference with interacting agents therefore requires communication that adds information and preserves independence\.
We present TalkMesh, a decentralized mesh of small language model agents on a sparse communication graph, whose agents learn when and what to communicate \(Figure[1](https://arxiv.org/html/2609.30578#S1.F1)\)\. Each agent first samples a proposalyiy\_\{i\}from the frozen base modelMM, as in self\-consistency\. We attach two LoRA adapters toMM: the confidence headVVscores every proposal, and the talk policyTTwrites hints and revisions\. The agent with the highest score, the speaker, writes a hint\. Every listener, an agent scoring below a threshold, reads the hint and rewrites its proposal as a revisionyi′y\_\{i\}^\{\\prime\}\. All other agents keep their proposals, which remain independent\. The acceptance test replaces a proposal with its revision only ifVVscores the revision higher\. Gossip consensus, in which neighbours repeatedly average answer tallies, approximates the vote weighted by confidence scores \(the weighted vote\) without a coordinator\. We trainTTinside the protocol with group relative policy optimization \(GRPO\) on a flip reward,\+1\+1for a revision that corrects a wrong proposal and−1\-1for the reverse; a hint receives the mean flip reward of its listeners\. A curriculum over network size raises the rollout sizeKK\(agents per training rollout\) through 3, 5, and 8\.
Figure[2](https://arxiv.org/html/2609.30578#S4.F2)reports accuracy against network sizeNN\. With three agents, the trained mesh reaches the accuracy of self\-consistency with 32 samples for all three models\. AtN=32N=32\(200 test problems per model\), the trained mesh raises accuracy from 0\.492 under self\-consistency to 0\.722 for SmolLM3\-3B on MATH\-500 and from 0\.568 to 0\.705 for Qwen3\.5\-0\.8B on GSM8K; without messages, the weighted vote \(Rerank\) reaches 0\.590 and 0\.655 \(Table[2](https://arxiv.org/html/2609.30578#S5.T2)\)\. For these two models and Qwen3\-1\.7B on MATH\-500, the trained mesh exceeds self\-consistency at everyNNfrom 3 to 32 after training with at mostK=8K=8agents\. When 4 ofN=8N\{=\}8agents are compromised, submitting a shared wrong answer with claimed confidence 0\.99 and a poisoned hint as speaker, majority vote accuracy on 150 test problems falls to 0\.000 for all three models\. The defended mesh rescores every solution with the confidence head and applies the acceptance test to every revision\. It retains 0\.507 with Qwen3\.5\-0\.8B on GSM8K \(0\.667 without attack\) and exceeds the undefended mesh, which omits both checks and falls to at most 0\.007, for every model \(p<10−3p<10^\{\-3\}, McNemar’s exact test; Section[5\.3](https://arxiv.org/html/2609.30578#S5.SS3)\)\. On the embodied coordination benchmark SwarmBench, REINFORCE training raises the Flocking score of all three models over twenty unseen seeds, and a Qwen3\.5\-0\.8B policy distilled from a centralized solver matches the solver on Synchronization \(Section[5\.4](https://arxiv.org/html/2609.30578#S5.SS4)\)\. We introduce a traffic signal control benchmark of sixteen intersections under New York City timing constraints and train language model meshes by behaviour cloning on traces of a communicating controller whose incident warnings reach intersections that cannot observe the incident\. Before any language model evaluation, we fix a severe subset of ten of the 70 scenarios, on which messages reduce the communicating controller’s mean person delay by 48\.2 s\. On this subset, delivering messages reduces the person delay of the Qwen3\.5\-0\.8B, 2B, and 4B meshes by23\.6±6\.723\.6\\pm 6\.7s \(standard error over 60 paired episodes\)\. Over all 70 scenarios, the reduction for the 4B mesh is11\.4±6\.911\.4\\pm 6\.9s \(Section[5\.5](https://arxiv.org/html/2609.30578#S5.SS5)\)\.
The hint and the incident warning illustrate the observability criterion: communication improves a decision when the required information is unobservable to the acting agent and observable to another agent that can send it\. Withholding messages lowers performance in reasoning and traffic, where the acting agent cannot observe the required information; in two SwarmBench settings where it can, scores with and without messages agree within one standard error \(Table[2](https://arxiv.org/html/2609.30578#S5.T2)\)\.
## 2 Related work
Self\-consistency\([Wang et al\., 2022](https://arxiv.org/html/2609.30578#bib.bib29)\)established majority voting over sampled chains of thought as the baseline for inference scaling\.[Cobbe et al\. \(2021\)](https://arxiv.org/html/2609.30578#bib.bib5)showed that reranking by a trained verifier outperforms finetuning alone, and[Li et al\. \(2024\)](https://arxiv.org/html/2609.30578#bib.bib14)showed that voting accuracy increases with the number of agents across models and tasks\. Without communication, the mesh reduces to the weighted vote \(Rerank\), which combines voting with verifier reranking\. Multiagent debate\([Du et al\., 2023](https://arxiv.org/html/2609.30578#bib.bib6);[Liang et al\., 2024](https://arxiv.org/html/2609.30578#bib.bib16)\), ReConcile\([Chen et al\., 2024](https://arxiv.org/html/2609.30578#bib.bib4)\), and Exchange\-of\-Thought\([Yin et al\., 2023](https://arxiv.org/html/2609.30578#bib.bib33)\)use handcrafted discussion rounds and an untrained communication policy among a fixed number of agents\. Decentralized agent networks\([Yang et al\., 2025](https://arxiv.org/html/2609.30578#bib.bib32);[Li et al\., 2025](https://arxiv.org/html/2609.30578#bib.bib15);[Ruan et al\., 2025](https://arxiv.org/html/2609.30578#bib.bib23);[Zhang et al\., 2026b](https://arxiv.org/html/2609.30578#bib.bib35);[Tian et al\., 2026](https://arxiv.org/html/2609.30578#bib.bib28)\)and consensus schemes robust to Byzantine agents\([Jo and Park, 2025](https://arxiv.org/html/2609.30578#bib.bib12);[Lee et al\., 2026](https://arxiv.org/html/2609.30578#bib.bib13)\)use frozen models\. Our training is closest to reinforcement learning of agent collaboration\([Liu et al\., 2026](https://arxiv.org/html/2609.30578#bib.bib18);[Sun et al\., 2024](https://arxiv.org/html/2609.30578#bib.bib27)\), learned multiround message protocols\([Zhang et al\., 2026a](https://arxiv.org/html/2609.30578#bib.bib34);[Fan et al\., 2026](https://arxiv.org/html/2609.30578#bib.bib7)\), communication guided by language models for cooperative control\([Bae et al\., 2026](https://arxiv.org/html/2609.30578#bib.bib2)\), and emergent communication in multiagent reinforcement learning\([Foerster et al\., 2016](https://arxiv.org/html/2609.30578#bib.bib9);[Sukhbaatar et al\., 2016](https://arxiv.org/html/2609.30578#bib.bib26);[Lowe et al\., 2017](https://arxiv.org/html/2609.30578#bib.bib20)\)\. In TalkMesh, \(i\) messages are natural language, \(ii\) the confidence head selects speakers and listeners, \(iii\) each hint receives the mean flip reward of its listeners, and \(iv\) we trainTTinside the inference protocol\.
Concurrent work with frontier models shows that agents sharing verified progress through a common workspace outperform the best independent agent\([Park et al\., 2026](https://arxiv.org/html/2609.30578#bib.bib22)\)\. Teams with collaboration strategies learned by reflective prompt search exceed a perfect router over their members’ answers in average accuracy on mathematics and physics benchmarks\([Pappu et al\., 2026](https://arxiv.org/html/2609.30578#bib.bib21)\)\. Both find that communication helps most when progress is verifiable; our acceptance test verifies each revision with the confidence head\. TalkMesh trains the talk policy with GRPO on at most 8 agents, reaches consensus by gossip without a judge model or common workspace, and exceeds self\-consistency at everyNNup to 32 with five open models of 0\.8B to 4B parameters \(Sections[5\.1](https://arxiv.org/html/2609.30578#S5.SS1)and[5\.2](https://arxiv.org/html/2609.30578#S5.SS2)\)\.
## 3 Method
Each agent in TalkMesh runs a frozen base modelMMwith two LoRA adapters: \(i\) a confidence headVV, which scores a solution, and \(ii\) a talk policyTT, which writes hints and revisions \(Figure[1](https://arxiv.org/html/2609.30578#S1.F1)\)\. Agents with the same base model share these weights; in a heterogeneous mesh \(Section[5\.3](https://arxiv.org/html/2609.30578#S5.SS3)\), each agent scores and revises with the adapters of its own model\. BecauseMMis never updated, training leaves the accuracy of a single proposal unchanged\.
Inference protocol\.Given a problemxxandNNagents on a connected communication graph, the mesh executes five stages\. \(i\)*Propose*: each agent samples a proposalyi∼M\(x\)y\_\{i\}\\sim M\(x\)with both adapters disabled and the sampler of self\-consistency, so the extracted answers are i\.i\.d\. Answers come from theANSWER:tag \(fallback: a boxed answer\) and match numerically; a proposal without an answer abstains\. \(ii\)*Score*: the confidence head assignsci=V\(x,yi\)=σ\(zYes−zNo\)c\_\{i\}=V\(x,y\_\{i\}\)=\\sigma\(z\_\{\\text\{Yes\}\}\-z\_\{\\text\{No\}\}\), the sigmoid of the logit gap between*Yes*and*No*after a fixed verification prompt\. \(iii\)*Speak*: the speaker is the agent with the highest confidence score,s=argmaxicis=\\arg\\max\_\{i\}c\_\{i\}\. Agents findsswithout a coordinator by gossiping the maximumcic\_\{i\}within a number of rounds equal to the graph diameter\. The speaker generates a hintm=T\(x,ys\)m=T\(x,y\_\{s\}\)of at most 112 tokens; the speaker prompt requests the key idea, the critical step, and the final answer\. \(iv\)*Gated revision*: every listener \(agenti≠si\\neq swithcic\_\{i\}below the thresholdt=0\.5t=0\.5or without an answer\) generates a revisionyi′=T\(x,yi,m\)y\_\{i\}^\{\\prime\}=T\(x,y\_\{i\},m\)\. The acceptance test keeps the revision only ifV\(x,yi′\)\>V\(x,yi\)V\(x,y\_\{i\}^\{\\prime\}\)\>V\(x,y\_\{i\}\)\. The other agents keep their proposals, which remain independent\. \(v\)*Consensus*: gossip consensus computes the mesh answer as the weighted votea^=argmaxa∑ici\[ai=a\]\\hat\{a\}=\\arg\\max\_\{a\}\\sum\_\{i\}c\_\{i\}\\,\\mathbf\{1\}\\\!\\left\[a\_\{i\}=a\\right\], with answersaia\_\{i\}and scorescic\_\{i\}after stage \(iv\)\.
Gossip consensus\.Each agentiiholds a sparse tally vectorτi\\tau\_\{i\}over answer strings, initialized as\{ai↦ci\}\\\{a\_\{i\}\\mapsto c\_\{i\}\\\}, and in each gossip roundrraverages it with the tallies of its neighbours:
τi\(r\+1\)=Wiiτi\(r\)\+∑j∈𝒩\(i\)Wijτj\(r\),Wij=11\+max\(di,dj\),Wii=1−∑j∈𝒩\(i\)Wij,\\begin\{gathered\}\\tau\_\{i\}^\{\(r\+1\)\}=W\_\{ii\}\\,\\tau\_\{i\}^\{\(r\)\}\+\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}W\_\{ij\}\\,\\tau\_\{j\}^\{\(r\)\},\\\\\[2\.0pt\] W\_\{ij\}=\\frac\{1\}\{1\+\\max\(d\_\{i\},d\_\{j\}\)\},\\qquad W\_\{ii\}=1\-\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}W\_\{ij\},\\end\{gathered\}\(1\)where agentiihas neighbours𝒩\(i\)\\mathcal\{N\}\(i\)and degreedid\_\{i\}\. These Metropolis weights form a symmetric, doubly stochastic mixing matrix, so on any connected graph everyτi\(r\)\\tau\_\{i\}^\{\(r\)\}converges to the network average\([Xiao et al\., 2005](https://arxiv.org/html/2609.30578#bib.bib31);[Boyd et al\., 2006](https://arxiv.org/html/2609.30578#bib.bib3)\)\. The answer of each agent,argmaxaτi\(r\)\(a\)\\arg\\max\_\{a\}\\tau\_\{i\}^\{\(r\)\}\(a\), therefore converges to the centralized weighted vote; Section[5\.1](https://arxiv.org/html/2609.30578#S5.SS1)measures their agreement\. OverRRgossip rounds, agentiisendsO\(diR\)O\(d\_\{i\}R\)tallies, independent ofNN; each tally holds one entry per distinct answer\. All runs use connected Watts\-Strogatz graphs with rewiring probability0\.10\.1, mean degree44\(22atN=3N\{=\}3\), andR=25R=25\.
Training the confidence head\.For each model and dataset, we trainVVas an outcome verifier\([Cobbe et al\., 2021](https://arxiv.org/html/2609.30578#bib.bib5)\)on eight solutions sampled fromMMper training problem, labelled by the ground truth answer\. Batches balance correct and incorrect solutions, and we keep the checkpoint with the highest validation accuracy\.
Training the talk policy\.We trainTT, one adapter for speaker and listener initialized to reproduceMM, with group relative policy optimization\([Shao et al\., 2024](https://arxiv.org/html/2609.30578#bib.bib25), GRPO;\)on rollouts of the inference protocol\. A rollout runs stages \(i\) to \(iii\) and selects the listeners\. For each problem, we sample a group ofG=4G\{=\}4hintsm1,…,mGm\_\{1\},\\dots,m\_\{G\}from the speaker context\(x,ys\)\(x,y\_\{s\}\), and every listener writes one revision under each hint\. The revisionyg,i′y\_\{g,i\}^\{\\prime\}of listeneriiunder hintggreceives the flip reward
rg,i=\[yg,i′correct\]−\[yicorrect\]∈\{−1,0,\+1\},r\_\{g,i\}=\\mathbf\{1\}\\\!\\left\[y\_\{g,i\}^\{\\prime\}\\text\{ correct\}\\right\]\-\\mathbf\{1\}\\\!\\left\[y\_\{i\}\\text\{ correct\}\\right\]\\in\\\{\-1,0,\+1\\\},\(2\)and hintggreceives the speaker rewardRg=1\|L\|∑i∈Lrg,iR\_\{g\}=\\frac\{1\}\{\|L\|\}\\sum\_\{i\\in L\}r\_\{g,i\}, the mean flip reward over listenersLL\. Both rewards score revisions before the acceptance test\. We obtain advantages by standardizing \(i\) speaker rewards over theGGhints of one problem and \(ii\) flip rewards over theGGrevisions of one listener\. Problems without listeners and groups with equal rewards contribute no gradient\. We penalize the KL divergence ofTTfromMMwith thek3k\_\{3\}estimator \(weight0\.10\.1\)\([Schulman, 2020](https://arxiv.org/html/2609.30578#bib.bib24)\)\. Only training uses ground truth answers, as in centralized training with decentralized execution\([Lowe et al\., 2017](https://arxiv.org/html/2609.30578#bib.bib20)\)\. A curriculum over network size raises the rollout sizeKKthrough 3, 5, and 8 for 20 steps each, with learning rate2×10−52\\times 10^\{\-5\}\. We keep the checkpoint with the highest mesh accuracy on validation problems, evaluated every 10 steps\.
## 4 Experimental setup
Figure 2:Accuracy against network sizeNNfor Qwen3\.5\-0\.8B on GSM8K and Qwen3\-1\.7B and SmolLM3\-3B on MATH\-500 \(pooled evaluation on 200 test problems, Section[4](https://arxiv.org/html/2609.30578#S4)\)\. Bands show one standard error\. All conditions except SC use the centralized weighted vote\. Labels give the margin of Trained talk over SC atN=32N\{=\}32, rounded to two decimals\.We use Qwen3\.5\-0\.8B, Qwen3\-1\.7B, and SmolLM3\-3B with thinking disabled and rank 16 LoRA adapters on all linear layers\([Hu et al\., 2021](https://arxiv.org/html/2609.30578#bib.bib11)\)\. We evaluate the 0\.8B model on GSM8K\([Cobbe et al\., 2021](https://arxiv.org/html/2609.30578#bib.bib5)\)and the larger models on the 326 numeric MATH\-500 problems\([Hendrycks et al\., 2021](https://arxiv.org/html/2609.30578#bib.bib10);[Lightman et al\., 2023](https://arxiv.org/html/2609.30578#bib.bib17)\)\. We use Qwen3\.5\-2B and Qwen3\.5\-4B for scaling within one model family \(Section[5\.2](https://arxiv.org/html/2609.30578#S5.SS2)\); Section[5\.3](https://arxiv.org/html/2609.30578#S5.SS3)states its own test sets\. Training uses the GSM8K training split and, for MATH\-500, 4,958 numeric MATH training problems; training, validation, and test sets are disjoint\.
We compare four conditions on identical proposals: \(i\)*SC*, self\-consistency by majority vote\([Wang et al\., 2022](https://arxiv.org/html/2609.30578#bib.bib29)\); \(ii\)*Rerank*, the weighted vote without communication; \(iii\)*Frozen talk*and \(iv\)*Trained talk*, the protocol withMMorTT, respectively, writing hints and revisions\. Reasoning agents sample at temperature 0\.7 \(nucleus 0\.8, or 0\.9 for SmolLM3\-3B; top\-kk20\); proposals and revisions use at most 320 \(GSM8K\) or 448 \(MATH\-500\) new tokens\.
Pooled evaluation uses 200 test problems per model \(of 326 on MATH\-500\) and two pools ofP=32P\{=\}32proposals per problem, scored once byVV\. EachN∈\{3,5,8,16,32\}N\\in\\\{3,5,8,16,32\\\}usesmin\(4,⌊P/N⌋\)\\min\(4,\\lfloor P/N\\rfloor\)random subsets per pool, each rerunning stages \(iii\) and \(iv\)\. We average accuracy over subsets, then pools; standard errors combine both\. Mesh accuracies use the centralized weighted vote\. Gossip consensus \(R=25R\{=\}25rounds, one graph perNN\) is compared with this vote and with SC \(McNemar’s exact test\)\.
## 5 Results
### 5\.1 Collective accuracy as a function of network size
We test whether Trained talk keeps its margin over SC asNNgrows and how much of it communication causes\. Figure[2](https://arxiv.org/html/2609.30578#S4.F2)reports all four conditions\. AtN=32N\{=\}32, Trained talk exceeds SC by\+0\.137\+0\.137\(0\.8B, GSM8K:0\.7050\.705against0\.5680\.568\),\+0\.085\+0\.085\(1\.7B, MATH\-500:0\.4630\.463against0\.3780\.378\), and\+0\.230\+0\.230\(3B, MATH\-500:0\.7220\.722against0\.4920\.492\)\. With three agents, Trained talk reaches SC with 32 samples for every model \(0\.5760\.576against0\.5680\.568,0\.4240\.424against0\.3780\.378,0\.6030\.603against0\.4920\.492\)\. Rerank reaches0\.6550\.655,0\.4170\.417, and0\.5900\.590atN=32N\{=\}32, and Frozen talk0\.6800\.680,0\.4500\.450, and0\.6600\.660\. Trained talk exceeds both at everyNN, so trainingTTadds accuracy beyond the weighted vote and beyond hints and revisions written byMM\. Without the curriculum, a talk policy trained at fixed rollout sizeK=5K\{=\}5\(not plotted\) exceeds SC only forN≤8N\\leq 8\. Gossip consensus agrees with the centralized weighted vote on 99\.4% of problems averaged over models,NN, and both pools \(at least 98% at everyNN\)\. Its margin over SC is significant at everyNNfor all three models in both pools \(p<0\.01p<0\.01, McNemar’s exact test\)\.
### 5\.2 Scaling within one model family
Table 1:GSM8K accuracy of Qwen3\.5\-0\.8B, Qwen3\.5\-2B, and Qwen3\.5\-4B \( pooled evaluation on 200 test problems, two pools of 32 proposals; standard errors at most 0\.020\)\. Trained talk accuracy increases with model size at everyNN\(bold: highest accuracy at eachNN\)\. Under self\-consistency, the 2B model trails the 0\.8B model atN≥5N\\geq 5\(italics: largest gap\)\.The models of Section[5\.1](https://arxiv.org/html/2609.30578#S5.SS1)differ in size and family\. We therefore repeat confidence head training, talk policy training, and pooled evaluation on GSM8K for Qwen3\.5\-2B and Qwen3\.5\-4B\. Table[1](https://arxiv.org/html/2609.30578#S5.T1)reports accuracy of SC and Trained talk\. Under SC, the 2B model trails the 0\.8B model at everyN≥5N\\geq 5\. Trained talk reaches0\.7050\.705,0\.8550\.855, and0\.8650\.865for 0\.8B, 2B, and 4B atN=32N\{=\}32and increases with model size at everyNN\. For the 2B model atN=32N\{=\}32, Trained talk exceeds SC by0\.3620\.362\(from unrounded means over both pools\), the largest margin over SC in this paper\.
### 5\.3 Compromised agents and heterogeneous meshes
Figure 3:Traffic signal control\. \(a\) The4×44\\times 4grid with a blocked segment, an incident warning relayed between neighbouring intersections, and upstream metering\. \(b\) Person delay of the Qwen3\.5 meshes, all 70 scenarios, messages delivered\. \(c\) The same meshes on the severe subset, messages withheld and delivered\.In a coordinated attack,k=4k\{=\}4ofN=8N\{=\}8agents submit a shared wrong answer with claimed confidence0\.990\.99and, as speaker, send a hint asserting that answer\. The defended mesh applies two checks: \(i\) agents rescore every solution with the shared confidence head, so all agents agree on the speaker and vote weights; \(ii\) revisions must pass the acceptance test\. The undefended mesh omits both checks\. With Trained talk on 150 test problems per model \(0\.8B on GSM8K; 1\.7B and 3B on MATH\-500\), majority vote over all eight ballots falls to0\.0000\.000accuracy for every model and the undefended mesh to at most0\.0070\.007\. The defended mesh retains 0\.507, 0\.340, and 0\.473 for 0\.8B, 1\.7B, and 3B \(0\.667, 0\.513, 0\.693 without attack\) and exceeds the undefended mesh for every model \(p<10−3p<10^\{\-3\}, McNemar’s exact test\)\. For 0\.8B and 3B it exceeds SC over the four honest agents \(0\.473, 0\.467\)\.
Heterogeneous meshes on GSM8K \(N=8N\{=\}8\) contain\(4,2,2\)\(4,2,2\),\(4,0,4\)\(4,0,4\), or\(0,4,4\)\(0,4,4\)agents of 0\.8B, 2B, and 4B; each agent uses the confidence head and talk policy of its own model, without retraining\. Over two runs per composition \(same 200 test problems, independent proposals\), the mesh averages 0\.808 accuracy against 0\.717 for SC; its lowest run \(0\.775\) exceeds the highest SC run \(0\.740\)\. Table[2](https://arxiv.org/html/2609.30578#S5.T2)gives the range of mean accuracy per composition\.
### 5\.4 Embodied coordination
On SwarmBench\([Ruan et al\., 2025](https://arxiv.org/html/2609.30578#bib.bib23)\), agents with5×55\{\\times\}5views on an8×88\{\\times\}8or10×1010\{\\times\}10grid broadcast to visible agents\. Flocking alone has dense reward, so REINFORCE trains moves and messages on the change in EMD score \(reduction in earth mover’s distance to the target shape\) plus cohesion\. With four agents, the EMD score over twenty unseen seeds rises from0\.050\.05to0\.600\.60\(Qwen3\.5\-0\.8B\),1\.651\.65to2\.302\.30\(Qwen3\-1\.7B\), and0\.800\.80to0\.950\.95\(SmolLM3\-3B\)\. The four other tasks count completions per episode and use distillation from a verifiable expert, a centralized solver: behaviour cloning on its traces, then GRPO on its scores of sampled moves, selecting checkpoints on separate seeds\. On eighty unseen seeds, distilled Qwen3\.5\-0\.8B matches the solver on Synchronization \(18/1818/18\) and raises Foraging, Pursuit, and Transport from zero to 55%, 48%, and 26% of the solver’s3\.83\.8,1\.91\.9, and5\.25\.2\.
### 5\.5 Traffic signal control under incidents and timing constraints
Table 2:Effect of withholding messages\. Hidden regime: the acting agent cannot observe what its decision requires; observed regime: it can\.Δ\\Delta: change from the control \(accuracy; completions or captures per episode; person delay in s, negative is better\)\. Ranges: compositions \(two runs of 200 problems\) or meshes \(20 paired episodes\)\. In parentheses: the largest reduction on one seed\.We simulate in SUMO\([Lopez et al\., 2018](https://arxiv.org/html/2609.30578#bib.bib19)\)sixteen signalized intersections on a4×44\\times 4grid of 200 m blocks, modeled on the grid scenario of the RESCO benchmark\([Ault and Sharon, 2021](https://arxiv.org/html/2609.30578#bib.bib1)\)\. Each intersection observes only its own camera\. Timing constraints modeled on New York City signals fix a 90 s cycle of two phases and a 12 s minimum green, above the 7 s minimum walk interval of the MUTCD\([Federal Highway Administration, 2009](https://arxiv.org/html/2609.30578#bib.bib8)\); the split \(the division of green time between the phases\) changes by at most±5\\pm 5s per cycle\. The objective is person delay: mean delay per person \(s\), counting 1\.2 persons per car and one per pedestrian\. In each of 40 incident seeds \(one simulated hour\), a random midblock avenue segment loses capacity for fifteen minutes\. The downstream intersection observes the incident: the feeder from the blocked segment \(an approach delivering vehicles into an intersection\) stops delivering\.
Engineered controllers\.We compare \(i\) fixed time control, an even split that never adapts, \(ii\) a local adaptive controller that moves its split toward Webster’s proportional split\([Webster, 1958](https://arxiv.org/html/2609.30578#bib.bib30)\), computed from the queues and waiting pedestrians it observes, and \(iii\) a communicating controller that adds messages to \(ii\)\. Each cycle, every intersection sends its neighbours its inbound queues, throughputs, and warnings\. A starvation detector warns when a feeder with a running average of at least five vehicles per cycle delivers at most one while its queue is empty despite available green\. Each intersection forwards a warning it receives to its neighbours; intersections on the flagged corridor apply metering, reducing their green toward it\. Over these 40 seeds, mean person delay is116\.9116\.9s \(communicating\),118\.1118\.1s \(local adaptive\), and157\.0157\.0s \(fixed time\)\. Against the local adaptive controller, the communicating controller has lower median \(75\.775\.7against80\.180\.1s\) and maximum \(302\.9302\.9against342\.8342\.8s\) person delay and reduces it by up to 126 s on a single seed \(Table[2](https://arxiv.org/html/2609.30578#S5.T2)\)\. The mean over the ten worst seeds \(CVaR25%\\mathrm\{CVaR\}\_\{25\\%\}\) is224\.9224\.9,238\.4238\.4, and309\.4309\.4s in the same order\. A confirmatory sweep of 40 new seeds reproduces every ordering, including the lower maximum delay of the communicating controller \(141 s below local adaptive\)\. An oracle controller metering from the true incident location and window has higher mean person delay \(124\.5124\.5s\) than the communicating controller\.
Language model meshes\.We train Qwen3\.5 meshes \(0\.8B, 2B, 4B\) by behaviour cloning on traces of the communicating controller\. Each prompt contains the camera summary, feeder deliveries against running averages, and received messages\. Each response is a split change and a message of one line naming the nearest blocked corridor while the intersection holds a warning, or stating that none is blocked\. We evaluate each mesh, controlling all sixteen intersections in closed loop, on 70 scenarios with three concurrent incidents of 25 minutes under two decoding strategies \(greedy decoding and sampling at temperature 0\.2\)\. Every episode runs twice, with messages delivered and withheld\. The severe subset is ten of these scenarios, fixed before any language model evaluation, on which the communicating controller reduces mean person delay by48\.248\.2s against the local adaptive controller \(Table[2](https://arxiv.org/html/2609.30578#S5.T2)\); over all 70 scenarios, the reduction is2\.12\.1s\. Training traces contain three concurrent incidents, and the training loss upweights responses with warnings\.
Table 3:Person delay \(s\) of the language model meshes\. All scenarios: 140 paired episodes per model, messages delivered\. Severe subset: 20 paired episodes per model\.Figure[3](https://arxiv.org/html/2609.30578#S5.F3)and Table[3](https://arxiv.org/html/2609.30578#S5.T3)report the results\. With messages delivered, mean person delay over all 70 scenarios is284\.3284\.3,282\.5282\.5, and274\.6274\.6s for the 0\.8B, 2B, and 4B meshes\. On the severe subset, delivering messages reduces the person delay of every mesh; the pooled reduction over 60 paired episodes is23\.6±6\.723\.6\\pm 6\.7s \(mean and standard error\)\. Over all 70 scenarios, messages reduce the person delay of the 4B mesh by11\.4±6\.911\.4\\pm 6\.9s\. For the 0\.8B and 2B meshes, the change matches the2\.12\.1s reduction of the communicating controller within one standard error\. Scored against the true incidents, warnings of the three meshes have mean precision 0\.76 \(share of warnings naming a blocked corridor during an incident or the six cycles after\) and recall 0\.98 \(share of incident cycles with a warning on a blocked corridor within two cycles\)\. Corridor corroboration, unused in Table[3](https://arxiv.org/html/2609.30578#S5.T3), delivers a warning only after two distinct senders flag one corridor within two cycles\. Under corroboration, one compromised intersection that keeps its control actions but warns of its own unblocked corridor every cycle leaves every mesh’s trajectory unchanged \(severe subset, greedy decoding\)\.
### 5\.6 An observability criterion for communication
Controls remove messages without retraining, in three hidden and two observed settings \(Table[2](https://arxiv.org/html/2609.30578#S5.T2)\)\. In reasoning, listeners see the speaker’s solution only in the hint: Rerank \(no hint, no revision\) lowers accuracy atN=32N\{=\}32by 0\.050, 0\.046, and 0\.132 \(0\.8B on GSM8K; 1\.7B and 3B on MATH\-500\)\. In heterogeneous meshes, hints cross model sizes; their control, SC, also omits confidence weighting\. Only the downstream intersection observes a traffic incident; upstream intersections act on it\. On the severe subset, messages reduce person delay for every language model mesh\. On Synchronization, Foraging, and Pursuit, each agent’s5×55\{\\times\}5view shows what its decision requires; Synchronization stays at its ceiling without messages, and Foraging and Pursuit scores with and without messages agree within one standard error\. Directive messages that assign moves to agents \(Pursuit,12×1212\{\\times\}12grid\) yield 1\.12 captures per episode, against 1\.16 without\. All outcomes match the criterion\.
## 6 Limitations and future work
All evaluated tasks have automatically scored outcomes\. Training requires supervision: \(i\) ground truth answers in reasoning, \(ii\) task rewards and a centralized solver on SwarmBench, and \(iii\) traces of the communicating controller in traffic\. The method therefore needs labelled problems near the target distribution\. Traffic results come from SUMO simulation under New York City timing constraints\. The scaling study covers one family, Qwen3\.5 from 0\.8B to 4B, our largest model\. Future work will extend the acceptance test and gossip consensus to \(i\) agents with separate owners and incentives and \(ii\) networks with partitions, churn, and asynchrony\.
#### Reproducibility\.
We release all environments, training scripts, evaluation protocols, seeds, result files, and the exact command for every reported number\.
## 7 Conclusion
With three agents, TalkMesh reaches the accuracy of SC with 32 samples; atN=32N\{=\}32, it raises accuracy from0\.4920\.492to0\.7220\.722\(SmolLM3\-3B, MATH\-500\) and from0\.5680\.568to0\.7050\.705\(Qwen3\.5\-0\.8B, GSM8K\), and communication and training each add accuracy beyond the weighted vote\. With four of eight agents compromised, majority vote accuracy falls to0\.0000\.000, while the defended mesh retains0\.5070\.507\(Qwen3\.5\-0\.8B, GSM8K\) and significantly exceeds the undefended mesh for every model\. Over 40 incident seeds, the communicating traffic controller has the lowest mean person delay \(116\.9116\.9s\), and on the severe subset messages reduce the person delay of every Qwen3\.5 mesh \(23\.623\.6s pooled over 60 paired episodes\)\. Across all settings, messages improve a decision when another agent holds the information it requires\.
## References
- J\. Ault and G\. SharonReinforcement Learning Benchmarks for Traffic Signal Control\.InNeurIPS Datasets and Benchmarks,Cited by:[§5\.5](https://arxiv.org/html/2609.30578#S5.SS5.p1.1)\.
- Baeet al\.\(2026\)S\. Bae, Y\. Park, S\. Lee, and S\. HanLLM\-Guided Communication for Cooperative Multi\-Agent Reinforcement Learning\.Note:arXiv:2605\.18077External Links:2605\.18077Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Boydet al\.\(2006\)S\. Boyd, A\. Ghosh, B\. Prabhakar, and D\. ShahRandomized gossip algorithms\.IEEE Transactions on Information Theory52\(6\),pp\. 2508–2530\.External Links:[Document](https://dx.doi.org/10.1109/tit.2006.874516)Cited by:[§3](https://arxiv.org/html/2609.30578#S3.p3.2)\.
- Chenet al\.\(2024\)J\. Chen, S\. Saha, and M\. BansalReConcile: Round\-Table Conference Improves Reasoning via Consensus among Diverse LLMs\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7066–7085\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.381)Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Cobbeet al\.\(2021\)K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. SchulmanTraining Verifiers to Solve Math Word Problems\.Note:arXiv:2110\.14168External Links:2110\.14168Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1),[§3](https://arxiv.org/html/2609.30578#S3.p4.1),[§4](https://arxiv.org/html/2609.30578#S4.p1.1)\.
- Duet al\.\(2023\)Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. MordatchImproving Factuality and Reasoning in Language Models through Multiagent Debate\.Note:arXiv:2305\.14325External Links:2305\.14325Cited by:[§1](https://arxiv.org/html/2609.30578#S1.p1.1),[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Fanet al\.\(2026\)W\. Fan, T\. Tognoli, H\. P\. Zou, C\. Miao, Y\. Wang, and X\. ZhangTodyComm: Task\-Oriented Dynamic Communication for Multi\-Round LLM\-based Multi\-Agent System\.Note:arXiv:2602\.03688External Links:2602\.03688Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Federal Highway Administration \(2009\)Federal Highway AdministrationManual on Uniform Traffic Control Devices for Streets and Highways, 2009 Edition\.Washington, DC\.Cited by:[§5\.5](https://arxiv.org/html/2609.30578#S5.SS5.p1.1)\.
- Foersteret al\.\(2016\)J\. N\. Foerster, Y\. M\. Assael, N\. de Freitas, and S\. WhitesonLearning to Communicate with Deep Multi\-Agent Reinforcement Learning\.Note:arXiv:1605\.06676External Links:1605\.06676Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring Mathematical Problem Solving With the MATH Dataset\.Note:arXiv:2103\.03874External Links:2103\.03874Cited by:[§4](https://arxiv.org/html/2609.30578#S4.p1.1)\.
- Huet al\.\(2021\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: Low\-Rank Adaptation of Large Language Models\.Note:arXiv:2106\.09685External Links:2106\.09685Cited by:[§4](https://arxiv.org/html/2609.30578#S4.p1.1)\.
- Jo and Park \(2025\)Y\. Jo and C\. ParkByzantine\-Robust Decentralized Coordination of LLM Agents\.Note:arXiv:2507\.14928External Links:2507\.14928Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Leeet al\.\(2026\)H\. Lee, V\. Yun, D\. Panagou, and S\. P\. KarimireddyRobust Multi\-Agent LLMs under Byzantine Faults\.Note:arXiv:2605\.09076External Links:2605\.09076Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Liet al\.\(2024\)J\. Li, Q\. Zhang, Y\. Yu, Q\. Fu, and D\. YeMore Agents Is All You Need\.Trans\. Mach\. Learn\. Res\.\.External Links:2402\.05120Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Liet al\.\(2025\)R\. Li, H\. Liu, L\. Zhao, Z\. Li, J\. Li, J\. Jiang, L\. Xu, C\. Zhao, M\. Fan, and C\. LiangSwarmSys: Decentralized Swarm\-Inspired Agents for Scalable and Adaptive Reasoning\.Note:arXiv:2510\.10047External Links:2510\.10047Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Lianget al\.\(2024\)T\. Liang, Z\. He, W\. Jiao, X\. Wang, Y\. Wang, R\. Wang, Y\. Yang, S\. Shi, and Z\. TuEncouraging Divergent Thinking in Large Language Models through Multi\-Agent Debate\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 17889–17904\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.992)Cited by:[§1](https://arxiv.org/html/2609.30578#S1.p1.1),[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Lightmanet al\.\(2023\)H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. CobbeLet’s Verify Step by Step\.Note:arXiv:2305\.20050External Links:2305\.20050Cited by:[§4](https://arxiv.org/html/2609.30578#S4.p1.1)\.
- Liuet al\.\(2026\)S\. Liu, Z\. Liang, X\. Lyu, and C\. AmatoLLM Collaboration with Multi\-Agent Reinforcement Learning\.Proceedings of the AAAI Conference on Artificial Intelligence40\(38\),pp\. 32150–32158\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i38.40487)Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Lopezet al\.\(2018\)P\. A\. Lopez, M\. Behrisch, L\. Bieker\-Walz, J\. Erdmann, Y\. Flötteröd, R\. Hilbrich, L\. Lücken, J\. Rummel, P\. Wagner, and E\. WießnerMicroscopic Traffic Simulation using SUMO\.In2018 21st International Conference on Intelligent Transportation Systems \(ITSC\),pp\. 2575–2582\.External Links:[Document](https://dx.doi.org/10.1109/itsc.2018.8569938)Cited by:[§5\.5](https://arxiv.org/html/2609.30578#S5.SS5.p1.1)\.
- Loweet al\.\(2017\)R\. Lowe, Y\. Wu, A\. Tamar, J\. Harb, P\. Abbeel, and I\. MordatchMulti\-Agent Actor\-Critic for Mixed Cooperative\-Competitive Environments\.InNeural Information Processing Systems,External Links:1706\.02275Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1),[§3](https://arxiv.org/html/2609.30578#S3.p5.2)\.
- Pappuet al\.\(2026\)A\. Pappu, M\. Suzgun, Y\. Kwon, F\. Bianchi, B\. El, M\. J\. Kochenderfer, H\. Cao, and J\. ZouSelf\-Organizing Agent Teams Learn to Reason Together\.Note:arXiv:2609\.22682External Links:2609\.22682Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p2.1)\.
- Parket al\.\(2026\)J\. Park, V\. Kontonis, S\. Garg, A\. Krishnamurthy, and D\. PapailiopoulosScaling Discovery through Test\-Time Communication\.Note:arXiv:2609\.21032External Links:2609\.21032Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p2.1)\.
- Ruanet al\.\(2025\)K\. Ruan, M\. Huang, J\. Wen, and H\. SunBenchmarking LLMs’ Swarm intelligence\.Note:arXiv:2505\.04364External Links:2505\.04364Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1),[§5\.4](https://arxiv.org/html/2609.30578#S5.SS4.p1.1)\.
- Schulman \(2020\)J\. SchulmanApproximating KL Divergence\.Note:[http://joschu\.net/blog/kl\-approx\.html](http://joschu.net/blog/kl-approx.html)Cited by:[§3](https://arxiv.org/html/2609.30578#S3.p5.2)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. GuoDeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models\.Note:arXiv:2402\.03300External Links:2402\.03300Cited by:[§3](https://arxiv.org/html/2609.30578#S3.p5.1)\.
- Sukhbaataret al\.\(2016\)S\. Sukhbaatar, A\. Szlam, and R\. FergusLearning Multiagent Communication with Backpropagation\.InNeural Information Processing Systems,External Links:1605\.07736Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Sunet al\.\(2024\)C\. Sun, S\. Huang, and D\. PompiliLLM\-based Multi\-Agent Reinforcement Learning: Current and Future Directions\.Note:arXiv:2405\.11106External Links:2405\.11106Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Tianet al\.\(2026\)C\. Tian, Y\. Yao, and J\. CuiQueenBee Planner: Skill\-Evolving Communication Topologies for Token\-Efficient LLM Multi\-Agent Systems\.Note:arXiv:2606\.27492External Links:2606\.27492Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Wanget al\.\(2022\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-Consistency Improves Chain of Thought Reasoning in Language Models\.Note:arXiv:2203\.11171External Links:2203\.11171Cited by:[§1](https://arxiv.org/html/2609.30578#S1.p1.1),[§2](https://arxiv.org/html/2609.30578#S2.p1.1),[§4](https://arxiv.org/html/2609.30578#S4.p2.1)\.
- Webster \(1958\)F\. V\. WebsterTraffic signal settings\.Road research technical paper,H\.M\.S\.O\.,London\.Cited by:[§5\.5](https://arxiv.org/html/2609.30578#S5.SS5.p2.1)\.
- Xiaoet al\.\(2005\)L\. Xiao, S\. Boyd, and S\. LallA scheme for robust distributed sensor fusion based on average consensus\.InIPSN 2005\. Fourth International Symposium on Information Processing in Sensor Networks, 2005,pp\. 63–70\.External Links:[Document](https://dx.doi.org/10.1109/ipsn.2005.1440896)Cited by:[§3](https://arxiv.org/html/2609.30578#S3.p3.2)\.
- Yanget al\.\(2025\)Y\. Yang, H\. Chai, S\. Shao, Y\. Song, S\. Qi, R\. Rui, and W\. ZhangAgentNet: Decentralized Evolutionary Coordination for LLM\-based Multi\-Agent Systems\.Note:arXiv:2504\.00587External Links:2504\.00587Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Yinet al\.\(2023\)Z\. Yin, Q\. Sun, C\. Chang, Q\. Guo, J\. Dai, X\. Huang, and X\. QiuExchange\-of\-Thought: Enhancing Large Language Model Capabilities through Cross\-Model Communication\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 15135–15153\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.936)Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Zhanget al\.\(2026a\)X\. Zhang, J\. Yu, and Z\. ZhongLearning Efficient Communication Protocols for Multi\-Agent Reinforcement Learning\.IEEE Transactions on Machine Learning in Communications and Networking4,pp\. 1335–1352\.External Links:[Document](https://dx.doi.org/10.1109/tmlcn.2026.3719222)Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.
- Zhanget al\.\(2026b\)Y\. Zhang, F\. Liu, Y\. Shan, X\. Huang, X\. Yang, Y\. Zhu, X\. Cheng, C\. Liu, K\. Zeng, T\. J\. Zhang, and W\. JiangSILO\-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi\-Agent LLM Systems\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 29379–29398\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.1354)Cited by:[§2](https://arxiv.org/html/2609.30578#S2.p1.1)\.相似文章
SLMs 作为多智能体路由器:一种渐进式 SFT 与强化学习方法
本文提出通过渐进式监督微调和强化学习训练小语言模型作为多智能体路由器,相比仅基于意图路由的LLM基线方法,实现了更好的检索相关性和更低的延迟。
SMAC-Talk:面向大语言模型的星际争霸多智能体挑战自然语言扩展
SMAC-Talk 是一个新的基准测试,在星际争霸多智能体挑战的基础上进行扩展,旨在评估基于 LLM 的智能体在具有自然语言通信的协作多智能体环境中的表现。该基准包含带有欺骗性通信者的场景,并使用 Qwen3.5 系列模型对智能体进行基准测试,以研究推理能力、记忆机制和模型规模对协调效果的影响。
面向小规模语言模型智能体的稳健强化学习
本文系统研究了面向小规模语言模型智能体(70-500M参数)的强化学习不稳定性问题,识别出三种失效模式,并提出了稳健技术,包括合并并重新初始化适配器方法及安全机制;该方法实现了稳定收敛并提升了胜率。
学习交流
OpenAI研究人员演示了协作型AI代理可以通过在简单世界中进行强化学习,发展出自己的有根据的和组合型语言。这些代理通过获得需要协调的目标奖励来学习交流,创建共享的符号语言以协调行为。
NeuroMAS:将多智能体系统视为具有联合强化学习的神经网络
NeuroMAS将多智能体语言系统视为可训练的类神经网络架构,以LLM代理作为节点,利用强化学习来学习通信和专业化。实验表明,其性能得到提升,并且从较小的系统逐步扩展比从头训练大型系统效果更好。