MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks

arXiv cs.AI Papers

Summary

MeshHeal introduces a decentralized self-healing framework for LLM agent networks to address gray failures, using adaptive peer review and detection mechanisms that achieve higher accuracy with lower token usage than existing methods.

arXiv:2609.29015v1 Announce Type: new Abstract: Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin. We introduce MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales. At the fast timescale, an adaptive hierarchy escalates uncertain or low-scoring outputs from repeated single-reviewer evaluation to committee deliberation and, when needed, correction before use. At the slow timescale, a task- and ability-conditioned peer-relative detector aggregates scores to distinguish persistent degradation from ordinary output variation, trigger mandatory committee review, and eventually exclude degraded agents from ordinary routing; recovery probes provide fresh evidence for reintegration. To faithfully evaluate routing, we introduce Model-Backed MAS Evaluation, which ties ability assignments to execution models, since prompt-based ability assignments alone can leave routing errors hidden. Across BBH, MATH, and MMLU-Pro, MeshHeal achieves 0.839 degraded-phase accuracy using 51k total model tokens per task, versus the strongest baseline Symphony's 0.807 accuracy using 115k per task. Under staggered degradation and recovery, MeshHeal isolates degraded agents, keeps them excluded from ordinary task execution until recovery, and returns them to normal routing.
Original Article
View Cached Full Text

Cached at: 09/25/26, 09:36 AM

# MeshHeal: Self-Healing Decentralized LLM Agent Networks under Gray Failures
Source: [https://arxiv.org/html/2609.29015](https://arxiv.org/html/2609.29015)
Keru ChenSen LinAffiliation:Computer Science DepartmentAffiliation:University of HoustonEmail:[slin50@central\.uh\.edu](mailto:)Yingbin LiangAffiliation:Department of ECEAffiliation:The Ohio State UniversityEmail:[liang\.889@osu\.edu](mailto:)Nathaniel D\. BastianAffiliation:Whiting School of EngineeringAffiliation:Johns Hopkins UniversityEmail:[ndbastian@jhu\.edu](mailto:)Shaofeng Zou††thanks:Corresponding author\.Affiliation:School of ECEEAffiliation:Arizona State UniversityEmail:[zou@asu\.edu](mailto:)

###### Abstract

Decentralized LLM\-based multi\-agent systems coordinate through local interactions, but an agent can remain responsive while its task\-solving quality persistently degrades\. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin\. We introduceMeshHeal, a fully decentralized self\-healing framework that couples ability\-matched peer review across two timescales\. At the fast timescale, an adaptive hierarchy escalates uncertain or low\-scoring outputs from repeated single\-reviewer evaluation to committee deliberation and, when needed, correction before use\. At the slow timescale, a task\- and ability\-conditioned peer\-relative detector aggregates scores to distinguish persistent degradation from ordinary output variation, trigger mandatory committee review, and eventually exclude degraded agents from ordinary routing; recovery probes provide fresh evidence for reintegration\. To faithfully evaluate routing, we introduce Model\-Backed MAS Evaluation, which ties ability assignments to execution models, since prompt\-based ability assignments alone can leave routing errors hidden\. Across BBH, MATH, and MMLU\-Pro, MeshHeal achieves 0\.839 degraded\-phase accuracy using 51k total model tokens per task, versus the strongest baseline Symphony’s 0\.807 accuracy using 115k per task\. Under staggered degradation and recovery, MeshHeal isolates degraded agents, keeps them excluded from ordinary task execution until recovery, and returns them to normal routing\.

## 1Introduction

Recent advances in foundation models have enabled increasingly capable autonomous AI agents that can reason, plan, follow instructions, and use external tools\([Bommasani et al\., 2021](https://arxiv.org/html/2609.29015#bib.bib3)\)\. LLM\-based multi\-agent systems \(MASs\) decompose complex tasks among agents with complementary roles or abilities, combining their contributions into a final answer, often outperforming a single, larger model\. However, most existing MASs coordinate in a centralized manner\([Wu et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib1);[Hong et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib12);[Qian et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib13);[Wang et al\., 2025](https://arxiv.org/html/2609.29015#bib.bib18)\), which creates a communication and scaling bottleneck and also a catastrophic single point of failure\.

Decentralized coordination instead allows agents to execute and route work through local interactions without relying on a central orchestrator\. Accordingly, a growing body of recent work has investigated decentralized MAS architectures that distribute coordination and routing decisions across the agent network\([Zhang et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib36);[Yang et al\., 2025a](https://arxiv.org/html/2609.29015#bib.bib37);[Yang et al\., 2025b](https://arxiv.org/html/2609.29015#bib.bib9)\)\. These studies, however, largely assume that participating agents remain reliable throughout execution\. In practice, individual agents may fail or degrade because of problems in their model backends, local controllers, memories, or external tools, potentially disrupting task execution across the network\([Huang et al\., 2025](https://arxiv.org/html/2609.29015#bib.bib19);[Zhang et al\., 2025](https://arxiv.org/html/2609.29015#bib.bib4)\)\.

![Refer to caption](https://arxiv.org/html/2609.29015v1/output/main_figure.png)Figure 1:Overview ofMeshHeal\. \(a\) Agents route and complete complementary work\. \(b\) Review by agents with the relevant ability either keeps or replaces the current output, while scores accumulated across tasks update routing eligibility and probes allow recovered agents to return\. \(c\)Model\-Backed MAS Evaluationties execution models to declared abilities and the agent’s healthy or degraded condition, making routing errors affect accuracy\.Motivated by this gap, we study decentralized LLM\-based MASs operating in the presence of agent failures\. We focus on*gray failures*\([Huang et al\., 2017](https://arxiv.org/html/2609.29015#bib.bib27);[Huang et al\., 2018](https://arxiv.org/html/2609.29015#bib.bib28)\), and leave other types of failures as future work\. A gray failure occurs when an agent remains available and protocol\-compliant, i\.e\., it can communicate with its neighbors and return syntactically valid responses, but its effective ability to solve some tasks is degraded relative to its nominal ability\. Such a failure is “gray” because it is not directly exposed through an explicit error or loss of connectivity: the affected agent may appear healthy to some neighbors, while its degradation can only be inferred from the semantic quality of its outputs or their downstream consequences\. Gray failures are particularly relevant to decentralized LLM\-based multi\-agent systems because an agent’s availability does not necessarily imply its semantic reliability\. We do not regard every incorrect LLM response as a gray failure, since even a healthy agent has a nonzero probability of error\. Instead, a gray failure refers to asustaineddegradation relative to the agent’s nominal task solving capability\. Such degradation may arise from transient inference\-serving issues, context corruption or overflow, stale memory, failures of external tools, or changes in the model backend\. These failures are difficult to detect using conventional heartbeat\- or connectivity\-based mechanisms, since the affected agent continues to participate normally in the communication protocol\.

We refer to the system\-level capability to respond to such degradation as*self\-healing*: detecting agent degradation, limiting its effect on current task execution, adapting future routing, and reintegrating an agent after recovery\([Psaier and Dustdar, 2011](https://arxiv.org/html/2609.29015#bib.bib26)\)\. Existing LLM inference and agent\-coordination methods address complementary parts of this objective\. Methods based on repeated inference, deliberation, or robust aggregation improve current outputs\([Wang et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib15);[Du et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib16);[Li et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib17);[Wang et al\., 2025](https://arxiv.org/html/2609.29015#bib.bib18);[Jo and Park, 2025](https://arxiv.org/html/2609.29015#bib.bib20);[Zheng et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib21);[Lee et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib22)\), however, because they operate within individual tasks, these methods are not designed to retain producer\-specific evidence over time and therefore do not support persistent\-degradation detection, subsequent rerouting, or recovery\-aware reintegration\. Moreover, history\-based methods adapt routing across tasks\([Li et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib10);[Yang et al\., 2025b](https://arxiv.org/html/2609.29015#bib.bib9)\); however, routing adaptation alone acts only after history has accumulated, so it cannot protect the current task, and once an agent is bypassed, the system receives no fresh evidence of its recovery for reintegration\. Therefore, to address gray failures, the system must protect the current task before degradation can be reliably detected, accumulate evidence to adapt future routing, and reintegrate an agent once its performance recovers\.

We introduceMeshHeal\(see[Fig\.1](https://arxiv.org/html/2609.29015#S1.F1)\), a fully decentralized self\-healing framework that couples peer review across two timescales\. At the fast timescale, ability\-matched neighboring agents review every intermediate contribution\. An adaptive review hierarchy begins with repeated single\-reviewer evaluation and escalates uncertain or low\-scoring outputs to committee deliberation and, when necessary, multi\-agent correction\. At the slow timescale, a task\- and ability\-conditioned peer\-relative detector accumulates the same review evidence to distinguish persistent degradation from ordinary output variation and impose progressively stronger routing restrictions, culminating in isolation from ordinary task execution\. Recovery probes provide fresh reviewed evidence, allowing agents whose performance has recovered to return gradually to normal routing\.

In our experiments on BBH, MATH, and MMLU\-Pro,MeshHealachieves83\.9%83\.9\\%mean accuracy during degradation over five seeds, compared with80\.7%80\.7\\%for Symphony, the strongest baseline\. It uses 51k tokens \(both input and output tokens\) per task, compared with 115k for Symphony and 178k for SAC, reducing token use by56%56\\%and71%71\\%, respectively\. In sparse networks of up to 576 agents,MeshHealisolates all degraded agents and achieves comparable accuracy with a complete graph\.

A commonly used approach to enforce ability constraints on agents when evaluating ability\-aware routing is to use system prompts\([Li et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib11);[Qian et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib13);[Chen et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib32);[Hong et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib12);[Chen et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib8)\)\. However, we identify an important confound: system prompts do not reliably enforce ability constraints, and the agent may still be able to solve tasks outside its abilities, allowing incorrect routes to produce correct final answers and thereby hiding routing errors\. In a controlled BBH\([Suzgun et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib5)\)study across five models, we find that prompt\-based ability assignments do not reliably separate performance on assigned and unassigned abilities\. We introduce Model\-Backed MAS Evaluation, which uses backend selection to realize ability assignments and inject controlled gray failures\. For a requested ability, a healthy agent uses a strong model if it holds that ability and a weak model otherwise\. This construction provides known failure ground truth while preserving the defining property that a gray\-failed agent remains responsive\.

## 2Related Work

LLM multi\-agent coordination and robustness\.LLM agents combine reasoning and tool use to solve complicated tasks\([Yao et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib35)\)\. MASs extend these capabilities through task decomposition and specialization\([Wu et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib1);[Hong et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib12);[Qian et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib13)\), while decentralized systems coordinate and adapt through local interactions\([Yang et al\., 2025b](https://arxiv.org/html/2609.29015#bib.bib9);[Li et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib10)\)\.

Various approaches have been proposed to improve output quality of LLM and MAS, e\.g\., repeated inference, deliberation, peer review, and robust aggregation\([Wang et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib15);[Du et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib16);[Li et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib17);[Wang et al\., 2025](https://arxiv.org/html/2609.29015#bib.bib18);[Jo and Park, 2025](https://arxiv.org/html/2609.29015#bib.bib20);[Zheng et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib21);[Lee et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib22)\)\. These approaches may also be applied here to handle gray failure\. However, they operate within each task and do not identify whether a persistent failure has occurred to an agent\. Therefore, a degraded agent may consistently participate in task solving\.

Adaptive routing strategies were also developed to handle failures in MASs\. AgentNet updates routing based on interaction history\([Yang et al\., 2025b](https://arxiv.org/html/2609.29015#bib.bib9)\), Symphony\-Coord combines adaptive agent selection with voting\([Guan et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib14)\), and RAPS uses reputation to identify and isolate unreliable peers\([Li et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib10)\)\. However, routing adaptation alone cannot protect the current task, and once an agent is bypassed, the system receives no fresh evidence of its recovery for reintegration\.

Distributed recovery and gray failures\.Our recovery goal builds on self\-stabilizing and self\-healing systems, where systems adapt their behavior in response to failures\([Dijkstra, 1974](https://arxiv.org/html/2609.29015#bib.bib23);[Arora and Gouda, 1993](https://arxiv.org/html/2609.29015#bib.bib24);[Psaier and Dustdar, 2011](https://arxiv.org/html/2609.29015#bib.bib26)\)\. Gray\-failure research shows that hidden degradation may need to be inferred from peer observations rather than availability\([Huang et al\., 2017](https://arxiv.org/html/2609.29015#bib.bib27);[Huang et al\., 2018](https://arxiv.org/html/2609.29015#bib.bib28)\)\. We study this problem for LLM agents using reviewed output quality as the failure signal\.

## 3Problem Formulation

We consider a decentralized MAS with heterogeneous agent abilities and potential gray failures\.

### 3\.1Decentralized MAS Architecture

We model a decentralized LLM\-based MAS as an undirected graphG=\(𝒱,ℰ\)G=\(\\mathcal\{V\},\\mathcal\{E\}\), where𝒱\\mathcal\{V\}denotes the set of agents, andℰ⊆𝒱×𝒱\\mathcal\{E\}\\subseteq\\mathcal\{V\}\\times\\mathcal\{V\}represents the communication graph between the agents: an undirected edgeei,j∈ℰe\_\{i,j\}\\in\\mathcal\{E\}allows agentiito send information to agentjjand vice versa\. Each agent consists of an LLM and a local controller\. The LLM is used to generate the agent’s outputs, and is called the agent’s*model backend*; and the local controller coordinates the agent’s interactions with its neighbors\.

We consider heterogeneous agents where agents differ in their abilities\. Denote by𝒦i\\mathcal\{K\}\_\{i\}the set of abilities of an agentii, and𝒦i\\mathcal\{K\}\_\{i\}varies across agents\. Solving different tasks may require a different set of abilities\. Therefore, it is of key importance to assign the task to the agent with the matching ability\. We study the decentralized setting, where no agent has a global view of other agents, and each agent can only communicate with its neighbors on the communication graphℰ\\mathcal\{E\}\.

The multi\-agent system will be given a sequence of tasks denoted by𝒯=\{τ1,τ2,…,τN\}\\mathcal\{T\}=\\\{\\tau\_\{1\},\\tau\_\{2\},\\ldots,\\tau\_\{N\}\\\}\. Thett\-th task isτt=\(Qt,Tt\)\\tau\_\{t\}=\(Q\_\{t\},T\_\{t\}\), wheret∈\[N\]t\\in\[N\]is the task index,QtQ\_\{t\}is the input question,TtT\_\{t\}is its benchmark\-provided category label \(e\.g\.,geometryin MATH\([Hendrycks et al\., 2021](https://arxiv.org/html/2609.29015#bib.bib6)\)\), and we useRtR\_\{t\}to represent the set of abilities required to solveQtQ\_\{t\}\. Here\[N\]≜\{1,…,N\}\[N\]\\triangleq\\\{1,\\ldots,N\\\}\. We note thatQtQ\_\{t\}andTtT\_\{t\}are directly given to the agent, butRtR\_\{t\}needs to be inferred by the agent\.

Objective\.LetΠlocal\\Pi\_\{\\mathrm\{local\}\}be the set of coordination policies whose decisions use only the current task state, local interaction history, and messages from direct neighbors\. For a policyπ\\pi, lety^tπ\\widehat\{y\}\_\{t\}^\{\\pi\}be the final answer of the MAS andEvalt⁡\(y^tπ\)∈\{0,1\}\\operatorname\{Eval\}\_\{t\}\(\\widehat\{y\}\_\{t\}^\{\\pi\}\)\\in\\\{0,1\\\}indicate correctness under the benchmark evaluator\. We seek to optimize the average accuracy over theNNsequential tasks:

maxπ∈Πlocal⁡1N​∑t=1N𝔼⁡\[Evalt⁡\(y^tπ\)\],\\max\_\{\\pi\\in\\Pi\_\{\\mathrm\{local\}\}\}\\frac\{1\}\{N\}\\sum\_\{t=1\}^\{N\}\\mathbb\{E\}\\\!\\left\[\\operatorname\{Eval\}\_\{t\}\(\\widehat\{y\}\_\{t\}^\{\\pi\}\)\\right\],\(1\)where the expectation is over the execution randomness\.

### 3\.2Gray Failure and Self\-Healing

We study a practical setting in which agents may experience*gray failures*\([Huang et al\., 2017](https://arxiv.org/html/2609.29015#bib.bib27);[Huang et al\., 2018](https://arxiv.org/html/2609.29015#bib.bib28)\)\. A gray failure occurs when an agent remains available and protocol\-compliant, i\.e\., it can communicate with its neighbors and return syntactically valid responses, but its effective ability to solve some tasks is degraded relative to its nominal ability\.

Each agent may have two possible health conditions:healthy,degraded\. The health condition of each agentiimay change over time\. We assume that each agent’s health condition remains unchanged within a task but may change across tasks\. An agent is called*healthy*when it operates at its normal level of reliability, and*degraded*when it suffers from gray failure\. The true health condition is latent, and the agent must infer it from the agent’s outputs\. We further definerecoveryas restoration of an agent fromdegradedtohealthy\.

In this paper, our focus is on solving the problem in[Eq\.1](https://arxiv.org/html/2609.29015#S3.E1)when agents experience gray failures and may also recover from them\. We aim to design local coordination policies so that, even with degraded agents, each task can still be routed to healthy agents with matching abilities and solved with high accuracy; and at the same time, when an agent recovers fromdegradedtohealthy, future tasks that match this agent’s abilities will still be routed to it\. Such desired property is referred to asself\-healing\. To evaluate this ability, beyond task accuracy and inference cost, we introduce three additional criteria: detection delay, closure, and reintegration\. These criteria are adapted from the distributed fault\-tolerance literature\([Dijkstra, 1974](https://arxiv.org/html/2609.29015#bib.bib23);[Arora and Gouda, 1993](https://arxiv.org/html/2609.29015#bib.bib24);[Ghosh et al\., 2007](https://arxiv.org/html/2609.29015#bib.bib25);[Psaier and Dustdar, 2011](https://arxiv.org/html/2609.29015#bib.bib26)\)\. More specifically, suppose agentiibecomes degraded at taskbi∈\[N\]b\_\{i\}\\in\[N\]and recovers at taskri∈\[N\+1\]r\_\{i\}\\in\[N\+1\]\. If recovery is not observed in anNN\-task stream, we setri=N\+1r\_\{i\}=N\+1\. Letℐt\\mathcal\{I\}\_\{t\}denote the set of agents that are excluded from ordinary task execution at tasktt\. Then, we define:

- •Detection delay:For a degraded agentii, the system should isolate it as early as possible\. Letci=min⁡\{t≥bi:i∈ℐt\}c\_\{i\}=\\min\\\{t\\geq b\_\{i\}:i\\in\\mathcal\{I\}\_\{t\}\\\}be the index of the first task at which the degraded agentiiis isolated\. The detection delay is then defined asδidelay=ci−bi\\delta\_\{i\}^\{\\mathrm\{delay\}\}=c\_\{i\}\-b\_\{i\}\.
- •Closure:A system satisfies theclosureproperty if a degraded agent remains isolated until recovery after being detected\. Formally, for any degraded agentii,i∈ℐti\\in\\mathcal\{I\}\_\{t\},∀ci≤t<ri\\forall c\_\{i\}\\leq t<r\_\{i\}\.
- •Reintegration:it refers to the case that after recovery, agentiireturns to normal routing\. Letui≥riu\_\{i\}\\geq r\_\{i\}be the index of the first task at which agentiireturns to normal routing state through the end of the observed stream\. Then the reintegration delay is defined asδirein=ui−ri\\delta\_\{i\}^\{\\mathrm\{rein\}\}=u\_\{i\}\-r\_\{i\}\.

## 4MeshHeal: Review at Two Timescales

In[Sec\.4\.1](https://arxiv.org/html/2609.29015#S4.SS1), we will present our base workflow design in a decentralized MASwithout gray failures\. Then, in[Sec\.4\.2](https://arxiv.org/html/2609.29015#S4.SS2), we build upon the base design and introduce ourMeshHealwith novel two\-timescale mechanism for self\-healing\. The full algorithm is provided in[Alg\.1](https://arxiv.org/html/2609.29015#alg1)in App\.[B\.1](https://arxiv.org/html/2609.29015#A2.SS1)\.

### 4\.1Base Task Execution and Routing \(No Failures\)

Task execution\.Agents collaborate to sequentially solve and route a task through the network\. Each taskτt=\(Qt,Tt\)\\tau\_\{t\}=\(Q\_\{t\},T\_\{t\}\)is initially assigned to an entry agent sampled uniformly at random from the network\. The queryQtQ\_\{t\}and task typeTtT\_\{t\}are given, whileRtR\_\{t\}denotes the set of abilities required to solve the task and is inferred by the entry agent\.

The task is solved sequentially as it moves through the network, with each agent addressing the part aligned with its abilities and passing the rest to the next agent until it is fully solved\. The task state is initialized asXt=\(Qt,Tt,Rt,Ct,Yt,Pt\)X\_\{t\}=\(Q\_\{t\},T\_\{t\},R\_\{t\},C\_\{t\},Y\_\{t\},P\_\{t\}\), whereCtC\_\{t\}records the abilities already covered so far,YtY\_\{t\}stores the working context that will be passed to later agents, including accepted intermediate outputs, andPtP\_\{t\}is a list that records the route taken so far\. Initially,Ct=Yt=∅C\_\{t\}=Y\_\{t\}=\\emptysetandPtP\_\{t\}is also empty\.

When the task is routed to agentii, letMt=Rt∖CtM\_\{t\}=R\_\{t\}\\setminus C\_\{t\}be the set of abilities that were not supplemented by previous agents inPtP\_\{t\}and need to be further provided by subsequent agents, andHi,t=𝒦i∩MtH\_\{i,t\}=\\mathcal\{K\}\_\{i\}\\cap M\_\{t\}be the subset of matching abilities that can be provided by agentii\. We take a rule\-based approach at the controller, and these sets govern the controller’s decision for the current task:

- •IfHi,t=∅H\_\{i,t\}=\\emptyset, then agentiihas no matching abilities to contribute to the remaining of tasktt\. Then, agentii’s controller chooses toHandoff: it generates no task content and forwards the task state unchanged to the next agent \(discussed later in Routing\)\. Thus,CtC\_\{t\},YtY\_\{t\}andPtP\_\{t\}remain unchanged\.
- •If∅≠Hi,t⊂Mt\\emptyset\\neq H\_\{i,t\}\\subset M\_\{t\}, then agentiicontribute matching abilities to the task that prior agents inPtP\_\{t\}cannot; however, additional abilities beyond those inHi,tH\_\{i,t\}are still needed\. Then, agentii’s controller chooses toContribute\. The agent’s LLM solves the matching part in the remaining of the task using abilities inHi,tH\_\{i,t\}\. The output is appended toYtY\_\{t\},Ct=Ct∪Hi,tC\_\{t\}=C\_\{t\}\\cup H\_\{i,t\}, and agentiiis appended to the listPtP\_\{t\}\. The controller recomputesMt←Rt∖CtM\_\{t\}\\leftarrow R\_\{t\}\\setminus C\_\{t\}using the updatedCtC\_\{t\}and forwards the updated task state to the next agent \(discussed later in Routing\)\.
- •IfHi,t=MtH\_\{i,t\}=M\_\{t\}, then agentiihas all the abilities needed to finish the remaining of the task\. Then, agentii’s controller chooses toFinalize\. The agent solves the remaining task, combines the result with earlier outputs inYtY\_\{t\}, and proposes the final answer\. At this point, all abilities inRtR\_\{t\}have been covered by agents along the route inYtY\_\{t\}, and no further routing is required\.

Routing\.Each agentiimaintains an estimated the shortest hop distanceDi​\(s\)D\_\{i\}\(s\)to the nearest agent that holds abilityss\. If agentiiitself holds abilityss, thenDi​\(s\)=0D\_\{i\}\(s\)=0; otherwise, it updates in a distributed manner:Di​\(s\)=1\+minj∈𝒩i⁡Dj​\(s\),D\_\{i\}\(s\)=1\+\\min\_\{j\\in\\mathcal\{N\}\_\{i\}\}D\_\{j\}\(s\),where𝒩i\\mathcal\{N\}\_\{i\}is the set of direct neighbors of agentii\([Bellman, 1958](https://arxiv.org/html/2609.29015#bib.bib31)\)\. Under standard asynchronous Bellman\-Ford assumptions with fixed topology and reliable neighbor updates, these iterations converge in finite time\([Bertsekas, 1982](https://arxiv.org/html/2609.29015#bib.bib2)\)\. The path that achievesDi​\(s\)D\_\{i\}\(s\)will be used to select the next agent\. WhenHandofforContributeare chosen by the controller, the task has at least one ability not provided along the route inPtP\_\{t\}\. The controller samples one abilityssuniformly at random fromMtM\_\{t\}\. It then routes the task state along the path that achievesDi​\(s\)D\_\{i\}\(s\), i\.e\., to the nearest agent with abilityss\.

### 4\.2Task Execution and Routing under Gray Failures

Having defined the base task flow without failures, we now present our novel design to address gray failures\. We define a statehi∈\{normal,watched,isolated\}h\_\{i\}\\in\\\{\\text\{normal\},\\text\{watched\},\\text\{isolated\}\\\}for each agentii, which determines how the agent participates in execution and routing\. Anormalagent follows the base workflow without restriction\. Awatchedagent is routed and executes tasks as a normal agent, but every output it produces receives mandatory committee review\. Anisolatedagent has its set of abilities𝒦i\\mathcal\{K\}\_\{i\}set to be∅\\emptysettemporarily, and therefore it will be excluded from task execution, but its controller remains active for message relay\. When an agent becomesisolated, agents will need to recompute the shortest hop distance as described in[Sec\.4\.1](https://arxiv.org/html/2609.29015#S4.SS1)\. We then apply the base task flow in[Sec\.4\.1](https://arxiv.org/html/2609.29015#S4.SS1)\.

OurMeshHealframework presents a novel two\-timescale scheme with fast review and correction and slow detection, rerouting, and recovery \(illustrated in[Fig\.1](https://arxiv.org/html/2609.29015#S1.F1)\)\. At thefast timescale, an review protocol escalates low\-scoring, high\-disagreement, or watched\-agent outputs from repeated single\-reviewer evaluation to deliberative committee review and multi\-agent correction\. At theslow timescale, a task\- and ability\-conditioned peer\-relative detector aggregates persistent shortfalls and normalizes them against healthy\-neighbor variation, enabling same thresholds to be used across model backends; guarded recovery probes close the loop through reintegration\.

#### 4\.2\.1Fast Timescale: Review and Correction

At the fast timescale, every output produced byContributeorFinalizeis reviewed before use\. We denote the output byyi,ty\_\{i,t\}\. An eligible reviewer must be anormaldirect neighbor that holds every ability inHi,tH\_\{i,t\}\. The reviewer scoresyi,ty\_\{i,t\}on a fixed\[0,1\]\[0,1\]scale \(judge prompts in App\.[D\.1](https://arxiv.org/html/2609.29015#A4.SS1)\)\. To reduce review noise, the same reviewer evaluatesyi,ty\_\{i,t\}in three independent calls\. Denote the median byq0q\_\{0\}, and we use it as the initial score\. We use the median rather than the mean to reduce sensitivity to reviewer variability and outlier scores, consistent with prior work on robust aggregation of LLM judgments\([Acharya et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib33);[Gallaba et al\., 2025](https://arxiv.org/html/2609.29015#bib.bib34)\)\. We also use the standard deviationsus\_\{u\}to measure disagreement\. Committee review is triggered whenq0≤τtoq\_\{0\}\\leq\\tau\_\{\\mathrm\{to\}\},su≥σcoms\_\{u\}\\geq\\sigma\_\{\\mathrm\{com\}\},hi=watchedh\_\{i\}=\\text\{watched\}, or the output is a recovery probe, whereτto=0\.40\\tau\_\{\\mathrm\{to\}\}=0\.40is the acceptance threshold andσcom=0\.55\\sigma\_\{\\mathrm\{com\}\}=0\.55is the disagreement threshold\. Committee review adds multiple perspectives and deliberation, improving review reliability at the cost of additional inference and communication; we therefore invoke it only for low\-scoring, high\-disagreement, or watched\-agent outputs\.

The committee contains two eligible reviewers\. They first evaluateyi,ty\_\{i,t\}independently, exchange their explanations, and then reassess the same output\. LetvAv\_\{A\}andvBv\_\{B\}be their second\-round scores; the final review score isqi,t=\(vA\+vB\)/2q\_\{i,t\}=\(v\_\{A\}\+v\_\{B\}\)/2\. If committee review is not triggered,qi,t=q0q\_\{i,t\}=q\_\{0\}\. When only one eligible reviewer is available, two reviews are obtained through separate inference calls with isolated context histories\. Ifqi,t\>τtoq\_\{i,t\}\>\\tau\_\{\\mathrm\{to\}\},yi,ty\_\{i,t\}is accepted\. Otherwise,Takeoveris triggered: the two reviewers generate correction candidates, which a committee reviewer in the normal state synthesizes task context following Mixture\-of\-Agents style aggregation\([Wang et al\., 2025](https://arxiv.org/html/2609.29015#bib.bib18)\)\. The resulting corrected outputy^i,t\\widehat\{y\}\_\{i,t\}replacesyi,ty\_\{i,t\}before entering the task state\. The scoreqi,tq\_\{i,t\}remains attributed to agentiiand is retained as evidence for the slow timescale degradation detection\.

#### 4\.2\.2Slow Timescale: Detection, Rerouting and Recovery

A single low score in the fast timescale does not necessarily mean that the agent is persistently degraded\. We therefore accumulate the review scoresqi,tq\_\{i,t\}across tasks and use*relative peer comparison*to update the estimate ofhih\_\{i\}\. Our goal is to detect agent gray failures across different task types and model backends with the same detector, rather than tuning a separate threshold for each setting\.

Relative shortfall\.For each abilitys∈Hi,ts\\in H\_\{i,t\}, the controller considers directnormalneighbors that have recently reviewed outputs within the fixed history window from earlier tasks with the same task typeTtT\_\{t\}and abilityss\. Let𝒫i,Tt,s​\(t\)\\mathcal\{P\}\_\{i,T\_\{t\},s\}\(t\)denote this set of neighbors, excludingii, and letμj,Tt,s​\(t\)\\mu\_\{j,T\_\{t\},s\}\(t\)denote neighborjj’s recent mean review score before tasktt\. The controller first uses these neighbors to estimate the normal review score for the same type of work\(q¯−i,Tt,s​\(t\)\)\(\\bar\{q\}\_\{\-i,T\_\{t\},s\}\(t\)\)\. For eachs∈Hi,ts\\in H\_\{i,t\}, the controller measures how far agentii’s current score falls below the corresponding reference:

q¯−i,Tt,s​\(t\)=1\|𝒫i,Tt,s​\(t\)\|​∑j∈𝒫i,Tt,s​\(t\)μj,Tt,s​\(t\),xi,t,s=q¯−i,Tt,s​\(t\)−qi,t\.\\bar\{q\}\_\{\-i,T\_\{t\},s\}\(t\)=\\frac\{1\}\{\|\\mathcal\{P\}\_\{i,T\_\{t\},s\}\(t\)\|\}\\sum\_\{j\\in\\mathcal\{P\}\_\{i,T\_\{t\},s\}\(t\)\}\\mu\_\{j,T\_\{t\},s\}\(t\),\\qquad x\_\{i,t,s\}=\\bar\{q\}\_\{\-i,T\_\{t\},s\}\(t\)\-q\_\{i,t\}\.A positivexi,t,sx\_\{i,t,s\}means that agentiiscored belownormalneighbors on the same task type and abilityss\. This comparison removes differences in absolute review scores across task types by comparing an agent only with normal neighbors who work on the same task type and have the same ability\.

Persistent degradation\.A single positive shortfall may result from LLM randomness\. We therefore maintain a separate detector for each\(T,s\)\(T,s\)and average its most recent 8 shortfalls, yieldingx¯i,T,s\\bar\{x\}\_\{i,T,s\}\. For a healthy agent, it should stay near zero; if the agent remains degraded, it should stay positive\.

However, a fixed threshold onx¯i,T,s\\bar\{x\}\_\{i,T,s\}may not transfer across model backends because the dispersion of healthy agents’ mean shortfalls can differ substantially across backends\. We therefore also measure how much the mean shortfalls ofnormalneighbors vary\. We pool the mean shortfallsx¯j,T,s′\\bar\{x\}\_\{j,T,s^\{\\prime\}\}of allnormalneighborsjjover all abilitiess′s^\{\\prime\}under task typeTT, and letmi,Tm\_\{i,T\}andpi,T25p^\{25\}\_\{i,T\}be their median and 25th percentile\. We measure the normal variation by the lower spreadmi,T−pi,T25m\_\{i,T\}\-p^\{25\}\_\{i,T\}, because degraded agents that have not yet been detected can inflate the upper tail\. We normalize agentii’s mean shortfall by this variation:

zi,T,s=x¯i,T,s−mi,Tmax⁡\(mi,T−pi,T25,0\.10\)\.z\_\{i,T,s\}=\\frac\{\\bar\{x\}\_\{i,T,s\}\-m\_\{i,T\}\}\{\\max\\\!\\left\(m\_\{i,T\}\-p\_\{i,T\}^\{25\},\\,0\.10\\right\)\}\.Thus,zi,T,sz\_\{i,T,s\}measures how far agentii’s mean shortfall is above normal level, relative to how much healthy agents vary\. This second comparison accounts for differences in how much healthy agents’ mean shortfalls vary across model backends, so same detection thresholds can be used across models\.

State transitions, rerouting, and recovery\.Each update uses the currentzi,T,sz\_\{i,T,s\}; if fewer than three matched observations are availablehih\_\{i\}remains unchanged\. Otherwise,zi,T,s≥αz\_\{i,T,s\}\\geq\\alphamovesnormaltowatched,zi,T,s<αz\_\{i,T,s\}<\\alphareturnswatchedtonormal, and two consecutive updates withzi,T,s≥βz\_\{i,T,s\}\\geq\\betamove anormalorwatchedagent toisolated; the isolation count resets whenzi,T,s<βz\_\{i,T,s\}<\\beta\. Hereα<β\\alpha<\\betaare thresholds, and in our experiments we chooseα=1\.2\\alpha=1\.2andβ=2\\beta=2\. The thresholds were selected empirically to balance detection delay and false isolation of healthy agents\.

We further design*recovery probes*that use the same review and detection mechanism\. After every 10 tasks, an isolated agent becomes eligible for a recovery probe\. When a neighbor next receives work matching its abilities, the work is routed to the isolated agent and its output receives mandatory committee review\. The resulting score updates the same detector\. The agent returns towatchedafter one update withzi,T,s<βz\_\{i,T,s\}<\\beta, even ifzi,T,s<αz\_\{i,T,s\}<\\alpha, and tonormalafter a later update withzi,T,s<αz\_\{i,T,s\}<\\alpha\.

## 5Model\-Backed MAS Evaluation

Table 1:BBH ability\-assignment test\. Arrows show no competence statement→\\rightarrowunassigned abilities declared unreliable;Δgap\\Delta\_\{\\mathrm\{gap\}\}is the resulting increase in the accuracy gap between assigned and unassigned abilities\.System prompt fails to constrain the agent’s ability\.Many LLM MAS evaluations assign abilities through system prompts while using the same underlying model for all agents\([Li et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib11);[Qian et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib13);[Chen et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib32);[Hong et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib12)\)\. To test whether such prompts constrain agents, we compare two prompting conditions across five models on BBH benchmark\([Suzgun et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib5)\): one makes no statement about unassigned abilities, while the other explicitly describes them as unreliable\. If these prompts truly limited agents to their assigned abilities, accuracy on unassigned abilities should fall while assigned\-ability accuracy remains stable\. The accuracy gap between assigned and unassigned abilities increases by only0\.0150\.015–0\.0580\.058\(see[Tab\.1](https://arxiv.org/html/2609.29015#S5.T1)\)\. Accuracy on unassigned abilities remains substantial, while assigned\-ability accuracy also falls for four models\. Prompt instructions change agent behavior, but do not reliably prevent agents from solving tasks outside their assigned abilities\. App\.[A](https://arxiv.org/html/2609.29015#A1)gives the full protocol\.

Controlled ability realization and gray\-failure injection\.We use model backends to realize ability assignments and inject gray failures, shown in[Fig\.1](https://arxiv.org/html/2609.29015#S1.F1)\(c\)\. A healthy agent uses astrong modelfor assigned abilities and aweak modelotherwise\. Degradation is simulated by replacing its strong backend with the weak one while leaving communication and execution interfaces unchanged\. This preserves responsiveness while reducing task competence, providing known failure ground truth and making routing errors visible in end\-to\-end accuracy\.

Tasks and degraded agents\.We apply this protocol to seven task groups from each of BBH, MATH, and MMLU\-Pro\([Suzgun et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib5);[Hendrycks et al\., 2021](https://arxiv.org/html/2609.29015#bib.bib6);[Wang et al\., 2024](https://arxiv.org/html/2609.29015#bib.bib7)\)\. Each task group contains tasks of the same type and is associated with two complementary abilities required for solving those tasks\. The two abilities are assigned to disjoint sets of agents, so no single agent holds both\. Every task requires at least two agents\. Full task\-to\-ability mappings appear in App\.[B](https://arxiv.org/html/2609.29015#A2)\.

## 6Experiments

We evaluate on benchmarks of BBH, MATH, and MMLU\-Pro\. The main system contains six agents, A1–A6, with three agents per ability and disjoint groups for the two abilities required by each task\. A3 and A5 are selected for degradation, every LLM call by them, including review calls, switches to the weak model\. Their controllers, abilities, communication links, and availability remain unchanged\. Unless stated otherwise, healthy agents useopenai/gpt\-oss\-120bfor declared abilities; undeclared abilities and degraded agents usemeta\-llama/llama\-3\.2\-1b\-instruct\.

We compareMeshHealwith AgentNet’s decentralized routing\([Yang et al\., 2025b](https://arxiv.org/html/2609.29015#bib.bib9)\), RAPS’s reputation\-based routing\([Li et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib10)\), and a centralized AutoGen implementation with a mandatory manager\([Wu et al\., 2023](https://arxiv.org/html/2609.29015#bib.bib1)\)\. SAC uses three teams with two rounds of filtering and refinement followed by majority aggregation\([Lee et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib22)\), while Symphony uses parallel plans and voting\([Guan et al\., 2026](https://arxiv.org/html/2609.29015#bib.bib14)\)\. All methods use the same tasks, agents, ability assignments, models, degradation targets, and schedule\. In AutoGen, A3 is the mandatory manager and appears on every task path\.

Each setting uses five seeds, defining task samples and order shared across methods\. We report means and sample standard deviations\. Except in targeted ablations, review rubrics and detector parameters are fixed across datasets, model pairs, and network sizes \([Tab\.B11](https://arxiv.org/html/2609.29015#A2.T11)\)\. We measure accuracy, token consumption \(counting both input and output tokens\) , detection delayδidelay\\delta\_\{i\}^\{\\mathrm\{delay\}\}, closure, and reintegration delayδirein\\delta\_\{i\}^\{\\mathrm\{rein\}\}; Appendices[B](https://arxiv.org/html/2609.29015#A2)–[E](https://arxiv.org/html/2609.29015#A5)provide settings, prompts, traces, and full results\.

### 6\.1Persistent Degradation: Accuracy and Cost

Each stream contains 252 healthy\-phase tasks followed by 252 disjoint degraded\-phase tasks with the same task\-type distribution\. During degraded\-phase, A3 and A5 switch to the weak model\. We report healthy\- and degraded\-phase accuracy and total model\-token cost\. After degradation,MeshHealreaches0\.839±0\.0120\.839\\pm 0\.012accuracy at51±551\\pm 5k total model tokens per task, compared with Symphony’s0\.807±0\.0140\.807\\pm 0\.014at115±9115\\pm 9k\. AgentNet and RAPS use fewer tokens but have substantially lower accuracy \([Tab\.2](https://arxiv.org/html/2609.29015#S6.T2)\(a\)\)\. Full dataset\-level and cost results are in App\.[C](https://arxiv.org/html/2609.29015#A3)\.

Table 2:\(a\) Accuracy and total model\-token cost under persistent degradation \(mean±\\pmstd\. over three datasets, each with five seeds\)\. H/D denote healthy/degraded phases; tokens are thousands per degraded\-phase task\. \(b\) Scaling in complete and sparse networks over five seeds\.I/DI/Ddenotes isolated/degraded agents; edge fraction is relative to the complete graph\.\(a\) Persistent degradation

\(b\) Network scaling

### 6\.2Degradation and Recovery: Detection, Closure, and Reintegration

We test the three self\-healing criteria on five consecutive 126\-task MATH phases: healthy operation, A5 degraded, A3 and A5 degraded, A3 degraded after A5’s backend restoration, and both backends restored\. All methods complete the 630\-task stream\. Across the 504 tasks after the first degradation, including the fully restored phase,MeshHealreaches0\.8120\.812accuracy at 49k total model tokens per task, compared with Symphony’s0\.8060\.806at 105k \(shown in[Tab\.3](https://arxiv.org/html/2609.29015#S6.T3)\)\. Symphony leads in the A5\-only and fully recovered phases, while the methods tie when only A3 remains degraded\.

Table 3:Self\-healing under staggered degradation and recovery over five seeds\. \(a\) Mean accuracy and total model\-token cost across phases\. \(b\) Detection \(Det\.\) and reintegration \(Rein\.\) delays are mean±\\pmstd\. \(min–max\); closure reports successful seeds\.\(a\) Accuracy and token cost \(b\) Self\-healing metrics ofMeshHeal

Across five seeds, after agentiibecomes degraded atbib\_\{i\}, it is first excluded from ordinary task execution atcic\_\{i\}\. The resulting detection delaysδidelay=ci−bi\\delta\_\{i\}^\{\\rm delay\}=c\_\{i\}\-b\_\{i\}are8\.0±5\.28\.0\\pm 5\.2tasks for A5 and12\.4±6\.212\.4\\pm 6\.2tasks for A3\. Once isolated, both agents satisfy closure in all five seeds, remaining excluded throughoutci≤t<ric\_\{i\}\\leq t<r\_\{i\}except for committee\-reviewed recovery probes\. After backend restoration atrir\_\{i\}, the agents return tonormalrouting atuiu\_\{i\}, giving reintegration delaysδirein=ui−ri\\delta\_\{i\}^\{\\rm rein\}=u\_\{i\}\-r\_\{i\}of46\.0±11\.746\.0\\pm 11\.7tasks for A5 and45\.0±3\.545\.0\\pm 3\.5tasks for A3 \([Tab\.3](https://arxiv.org/html/2609.29015#S6.T3)\)\.[Fig\.C1](https://arxiv.org/html/2609.29015#A3.F1)in App\.[C\.2](https://arxiv.org/html/2609.29015#A3.SS2)shows one representative trajectory, with its state changes listed in[Tab\.C5](https://arxiv.org/html/2609.29015#A3.T5)\. None of the baseline methods fully isolates degraded agents; under our definitions, their detection delay is therefore infinite, closure is not established, and reintegration delay is undefined\.

### 6\.3Scaling with Local Information

We replicate the six\-agent profiles in complete and sparse graphs with 96, 288, and 576 agents, with one third degraded and all review, summary exchange, and routing updates remaining local\.MeshHealisolates every degraded agent in all settings, while sparse and complete accuracies differ by at most0\.0100\.010\. At 576 agents, the sparse graph uses only0\.88%0\.88\\%of complete\-graph edges and reaches0\.8140\.814accuracy \([Tab\.2](https://arxiv.org/html/2609.29015#S6.T2)\(b\)\), demonstratingMeshHeal’s effectiveness in sparse networks\.

### 6\.4Component and Detector Studies

Fast timescale\.We replace committee review with a single reviewer and separately remove Takeover on BBH while retaining detection and rerouting \([Tab\.4](https://arxiv.org/html/2609.29015#S6.T4)\(a\); App\.[C\.4](https://arxiv.org/html/2609.29015#A3.SS4),[Tab\.C6](https://arxiv.org/html/2609.29015#A3.T6)\)\. Single reviewer uses the initial peer for judgment and correction; it has a slightly higher healthy\-phase mean but0\.0350\.035lower degraded\-phase accuracy and a mean phase change of−0\.044\-0\.044\. No Takeover leaves low\-scoring outputs unchanged, reducing degraded\-phase accuracy by0\.0560\.056and healthy\-phase accuracy by0\.0200\.020\. This suggests that Takeover also corrects some errors from healthy agents\. An audit \(App\.[C\.5](https://arxiv.org/html/2609.29015#A3.SS5),[Tab\.C8](https://arxiv.org/html/2609.29015#A3.T8)\(a,b\)\) shows that healthy agents receive much higher review scores than degraded agents \(0\.890\.89vs\.0\.080\.08\), and98%98\\%of degraded outputs are sent to committee review\. Committee reassessment reduces false alarms on healthy agents\. A separate Takeover audit \([Tab\.C8](https://arxiv.org/html/2609.29015#A3.T8)\(c\)\) shows that Takeover corrects 15 of 32 wrong answers while turning only 1 of 32 correct answers into an incorrect one\.

Table 4:\(a\) Accuracy over five seeds on BBH; H and D denote the healthy and degraded phases\. \(b\) Missed degraded agents out of two and falselyisolatedhealthy agents out of four on BBH\.\(a\) Fast\-timescale components \(b\) Detector comparison

Slow timescale\.We compare detectors under Pair A \(GPT\-OSS\-120B/Gemma\-3\-4B\) and Pair B \(Llama\-3\.3\-70B/Llama\-3\.2\-1B\)\. The absolute threshold falsely isolates all four healthy agents under Pair A; CUSUM\([Page, 1954](https://arxiv.org/html/2609.29015#bib.bib29)\)falsely isolates two under Pair A and misses one of two degraded agents under Pair B\. Relative peer comparison detects both degraded agents without false isolations under either pair using fixed rubrics and parameters \([Tab\.4](https://arxiv.org/html/2609.29015#S6.T4)\(b\); App\.[C\.4](https://arxiv.org/html/2609.29015#A3.SS4),[Tab\.C7](https://arxiv.org/html/2609.29015#A3.T7)\)\.

Model\-pair sensitivity\.To test robustness under a milder degradation and a weaker healthy model, we additionally evaluate Pair A and Pair B, respectively\.MeshHealdetects both degraded agents without false isolation across both settings, while maintaining competitive degraded\-phase accuracy at substantially lower cost: it outperforms Symphony by0\.0400\.040on Pair A and is within0\.0040\.004on Pair B, using only5252–56%56\\%as many total model tokens\. Full results are in App\.[C\.2](https://arxiv.org/html/2609.29015#A3.SS2)\.

## 7Conclusion

We introducedMeshHeal, a fully decentralized framework for self\-healing gray failures in LLM agent networks with a novel two\-timescale review design\. At the fast timescale, review protocol evaluates every produced output and can replace low\-quality outputs before they enter the task state; and at the slow timescale, a task\- and ability\-conditioned peer\-relative detector aggregates persistent shortfalls to drive staged state transitions and routing updates, while recovery probes support evidence\-based reintegration\. We further showed that prompt\-only ability assignments do not reliably enforce constraints on agents’ ability and proposed Model\-Backed MAS Evaluation, which uses controlled backend selection to realize complementary abilities and inject gray failures with known ground truth while keeping affected agents responsive\. With the default model pair,MeshHealreaches0\.8390\.839degraded\-phase accuracy at 51k tokens per task across three benchmarks and five seeds, versus the strongest baseline of Symphony0\.8070\.807at 115k tokens per task\. The experiment with degradation and recovery at different times meets the three self\-healing criteria under the evaluated schedule\. The scaling experiments isolate every degraded agent in the evaluated networks of up to 576 agents\.

## References

- Acharyaet al\.\(2026\)A\. Acharya, K\. W\. Pan, and B\. VerkhovskyRoPoLL: robust panel of llm judges\.arXiv preprint arXiv:2606\.30931\.Cited by:[§4\.2\.1](https://arxiv.org/html/2609.29015#S4.SS2.SSS1.p1.1)\.
- Arora and Gouda \(1993\)A\. Arora and M\. G\. GoudaClosure and convergence: a foundation of fault\-tolerant computing\.IEEE Transactions on Software Engineering19\(11\),pp\. 1015–1027\.Cited by:[§2](https://arxiv.org/html/2609.29015#S2.p4.1),[§3\.2](https://arxiv.org/html/2609.29015#S3.SS2.p3.1)\.
- Bellman \(1958\)R\. BellmanOn a routing problem\.Quarterly of applied mathematics16\(1\),pp\. 87–90\.Cited by:[§4\.1](https://arxiv.org/html/2609.29015#S4.SS1.p4.1)\.
- Bertsekas \(1982\)D\. BertsekasDistributed dynamic programming\.IEEE transactions on Automatic Control27\(3\),pp\. 610–616\.Cited by:[§4\.1](https://arxiv.org/html/2609.29015#S4.SS1.p4.1)\.
- Bommasaniet al\.\(2021\)R\. Bommasani, D\. A\. Hudson, E\. Adeli, R\. Altman, S\. Arora, S\. von Arx, M\. S\. Bernstein, J\. Bohg, A\. Bosselut, E\. Brunskill,et al\.On the opportunities and risks of foundation models\.arXiv preprint arXiv:2108\.07258\.Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p1.1)\.
- Chenet al\.\(2026\)K\. Chen, J\. Luo, S\. Lin, Y\. Liang, A\. Velasquez, N\. Bastian, and S\. ZouHIPO: instruction hierarchy via constrained reinforcement learning\.arXiv preprint arXiv:2603\.16152\.Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p7.1)\.
- Chenet al\.\(2024\)W\. Chen, Y\. Su, J\. Zuo, C\. Yang, C\. Yuan, C\. Chan, H\. Yu, Y\. Lu, Y\. Hung, C\. Qian, Y\. Qin, X\. Cong, R\. Xie, Z\. Liu, M\. Sun, and J\. ZhouAgentVerse: facilitating multi\-agent collaboration and exploring emergent behaviors\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p7.1),[§5](https://arxiv.org/html/2609.29015#S5.p1.1)\.
- Dijkstra \(1974\)E\. W\. DijkstraSelf\-stabilizing systems in spite of distributed control\.Communications of the ACM17\(11\),pp\. 643–644\.External Links:[Document](https://dx.doi.org/10.1145/361179.361202)Cited by:[§2](https://arxiv.org/html/2609.29015#S2.p4.1),[§3\.2](https://arxiv.org/html/2609.29015#S3.SS2.p3.1)\.
- Duet al\.\(2024\)Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. MordatchImproving factuality and reasoning in language models through multiagent debate\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 11733–11763\.Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p2.1)\.
- Gallabaet al\.\(2025\)K\. Gallaba, A\. Arabat, D\. Lin, M\. Sayagh, and A\. E\. HassanTowards conversational development environments: using theory\-of\-mind and multi\-agent architectures for requirements refinement\.arXiv preprint arXiv:2505\.20973\.Cited by:[§4\.2\.1](https://arxiv.org/html/2609.29015#S4.SS2.SSS1.p1.1)\.
- Ghoshet al\.\(2007\)S\. Ghosh, A\. Gupta, T\. Herman, and S\. V\. PemmarajuFault\-containing self\-stabilizing distributed protocols\.Distributed Computing20\(1\),pp\. 53–73\.External Links:[Document](https://dx.doi.org/10.1007/s00446-007-0032-2)Cited by:[§3\.2](https://arxiv.org/html/2609.29015#S3.SS2.p3.1)\.
- Guanet al\.\(2026\)Z\. Guan, H\. Cao, M\. Zhong, Y\. Wang, G\. Liu, E\. Yang, L\. Ai, Y\. Ni, and B\. ShiSymphony\-Coord: adaptive routing for multi\-agent LLM systems\.Note:arXiv preprint arXiv:2602\.00966External Links:2602\.00966Cited by:[Table B10](https://arxiv.org/html/2609.29015#A2.T10.2.1.6.2.1.1),[§2](https://arxiv.org/html/2609.29015#S2.p3.1),[§6](https://arxiv.org/html/2609.29015#S6.p2.1)\.
- Hendryckset al\.\(2021\)D\. Hendrycks, C\. Burns, S\. Kadavath, A\. Arora, S\. Basart, E\. Tang, D\. Song, and J\. SteinhardtMeasuring mathematical problem solving with the math dataset\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks,Vol\.1\.Cited by:[§3\.1](https://arxiv.org/html/2609.29015#S3.SS1.p3.1),[§5](https://arxiv.org/html/2609.29015#S5.p3.1)\.
- Honget al\.\(2024\)S\. Hong, M\. Zhuge, J\. Chen, X\. Zheng, Y\. Cheng, J\. Wang, C\. Zhang, Z\. Wang, S\. K\. S\. Yau, Z\. Lin, L\. Zhou, C\. Ran, L\. Xiao, C\. Wu, and J\. SchmidhuberMetaGPT: meta programming for a multi\-agent collaborative framework\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p1.1),[§1](https://arxiv.org/html/2609.29015#S1.p7.1),[§2](https://arxiv.org/html/2609.29015#S2.p1.1),[§5](https://arxiv.org/html/2609.29015#S5.p1.1)\.
- Huanget al\.\(2025\)J\. Huang, J\. Zhou, T\. Jin, X\. Zhou, Z\. Chen, W\. Wang, Y\. Yuan, M\. Lyu, and M\. SapOn the resilience of LLM\-based multi\-agent collaboration with faulty agents\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 26202–26226\.Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p2.1)\.
- Huanget al\.\(2018\)P\. Huang, C\. Guo, J\. R\. Lorch, L\. Zhou, and Y\. DangCapturing and enhancing in situ system observability for failure detection\.In13th USENIX Symposium on Operating Systems Design and Implementation,pp\. 1–16\.Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p3.1),[§2](https://arxiv.org/html/2609.29015#S2.p4.1),[§3\.2](https://arxiv.org/html/2609.29015#S3.SS2.p1.1)\.
- Huanget al\.\(2017\)P\. Huang, C\. Guo, L\. Zhou, J\. R\. Lorch, Y\. Dang, M\. Chintalapati, and R\. YaoGray failure: the achilles’ heel of cloud\-scale systems\.InProceedings of the 16th Workshop on Hot Topics in Operating Systems,pp\. 150–155\.External Links:[Document](https://dx.doi.org/10.1145/3102980.3103005)Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p3.1),[§2](https://arxiv.org/html/2609.29015#S2.p4.1),[§3\.2](https://arxiv.org/html/2609.29015#S3.SS2.p1.1)\.
- Jo and Park \(2025\)Y\. Jo and C\. ParkByzantine\-robust decentralized coordination of LLM agents\.Note:arXiv preprint arXiv:2507\.14928External Links:2507\.14928Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p2.1)\.
- Kazemiet al\.\(2025\)M\. Kazemi, B\. Fatemi, H\. Bansal, J\. Palowitch, C\. Anastasiou, S\. V\. Mehta, L\. K\. Jain, V\. Aglietti, D\. Jindal, P\. Chen, N\. Dikkala, G\. Tyen, X\. Liu, U\. Shalit, S\. Chiappa, K\. Olszewska, Y\. Tay, V\. Q\. Tran, Q\. V\. Le, and O\. FiratBIG\-Bench Extra Hard\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 26473–26501\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1285)Cited by:[§B\.2](https://arxiv.org/html/2609.29015#A2.SS2.p3.1)\.
- Leeet al\.\(2026\)H\. Lee, V\. Yun, D\. Panagou, and S\. P\. KarimireddyRobust multi\-agent LLMs under byzantine faults\.Note:arXiv preprint arXiv:2605\.09076External Links:2605\.09076Cited by:[Table B10](https://arxiv.org/html/2609.29015#A2.T10.2.1.5.2.1.1),[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p2.1),[§6](https://arxiv.org/html/2609.29015#S6.p2.1)\.
- Liet al\.\(2023\)G\. Li, H\. Hammoud, H\. Itani, D\. Khizbullin, and B\. GhanemCAMEL: communicative agents for “mind” exploration of large language model society\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 51991–52008\.Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p7.1),[§5](https://arxiv.org/html/2609.29015#S5.p1.1)\.
- Liet al\.\(2024\)J\. Li, Q\. Zhang, Y\. Yu, Q\. Fu, and D\. YeMore agents is all you need\.Transactions on Machine Learning Research\.Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p2.1)\.
- Liet al\.\(2026\)R\. Li, Z\. Zhang, X\. Bo, Q\. Dai, C\. Li, F\. Wen, and X\. ChenTowards adaptive, scalable, and robust coordination of LLM agents: a dynamic ad\-hoc networking perspective\.arXiv preprint arXiv:2602\.08009\.Cited by:[Table B10](https://arxiv.org/html/2609.29015#A2.T10.2.1.3.2.1.1),[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p1.1),[§2](https://arxiv.org/html/2609.29015#S2.p3.1),[§6](https://arxiv.org/html/2609.29015#S6.p2.1)\.
- Page \(1954\)E\. S\. PageContinuous inspection schemes\.Biometrika41\(1–2\),pp\. 100–115\.External Links:[Document](https://dx.doi.org/10.1093/biomet/41.1-2.100)Cited by:[§C\.4](https://arxiv.org/html/2609.29015#A3.SS4.p2.1),[§6\.4](https://arxiv.org/html/2609.29015#S6.SS4.p2.1)\.
- Psaier and Dustdar \(2011\)H\. Psaier and S\. DustdarA survey on self\-healing systems: approaches and systems\.Computing91\(1\),pp\. 43–73\.External Links:[Document](https://dx.doi.org/10.1007/s00607-010-0107-y)Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p4.1),[§3\.2](https://arxiv.org/html/2609.29015#S3.SS2.p3.1)\.
- Qianet al\.\(2024\)C\. Qian, W\. Liu, H\. Liu, N\. Chen, Y\. Dang, J\. Li, C\. Yang, W\. Chen, Y\. Su, X\. Cong, J\. Xu, D\. Li, Z\. Liu, and M\. SunChatDev: communicative agents for software development\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 15174–15186\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.810)Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p1.1),[§1](https://arxiv.org/html/2609.29015#S1.p7.1),[§2](https://arxiv.org/html/2609.29015#S2.p1.1),[§5](https://arxiv.org/html/2609.29015#S5.p1.1)\.
- Suzgunet al\.\(2023\)M\. Suzgun, N\. Scales, N\. Schärli, S\. Gehrmann, Y\. Tay, H\. W\. Chung, A\. Chowdhery, Q\. V\. Le, E\. H\. Chi, D\. Zhou, and J\. WeiChallenging big\-bench tasks and whether chain\-of\-thought can solve them\.InFindings of the Association for Computational Linguistics: ACL 2023,Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p7.1),[§5](https://arxiv.org/html/2609.29015#S5.p1.1),[§5](https://arxiv.org/html/2609.29015#S5.p3.1)\.
- Wanget al\.\(2025\)J\. Wang, J\. Wang, B\. Athiwaratkun, C\. Zhang, and J\. ZouMixture\-of\-agents enhances large language model capabilities\.InInternational Conference on Learning Representations,Cited by:[§B\.5](https://arxiv.org/html/2609.29015#A2.SS5.p3.1),[§1](https://arxiv.org/html/2609.29015#S1.p1.1),[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p2.1),[§4\.2\.1](https://arxiv.org/html/2609.29015#S4.SS2.SSS1.p2.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. V\. Le, E\. H\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-consistency improves chain of thought reasoning in language models\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p2.1)\.
- Wanget al\.\(2024\)Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang,et al\.Mmlu\-pro: a more robust and challenging multi\-task language understanding benchmark\.Advances in Neural Information Processing Systems37,pp\. 95266–95290\.Cited by:[§5](https://arxiv.org/html/2609.29015#S5.p3.1)\.
- Wuet al\.\(2023\)Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu,et al\.Autogen: enabling next\-gen llm applications via multi\-agent conversation\.arXiv preprint arXiv:2308\.08155\.Cited by:[Table B10](https://arxiv.org/html/2609.29015#A2.T10.2.1.4.2.1.1),[§1](https://arxiv.org/html/2609.29015#S1.p1.1),[§2](https://arxiv.org/html/2609.29015#S2.p1.1),[§6](https://arxiv.org/html/2609.29015#S6.p2.1)\.
- Yanget al\.\(2025a\)H\. Yang, J\. Chen, M\. Siew, T\. Lorido\-Botran, and C\. Joe\-WongLLM\-powered decentralized generative agents with adaptive hierarchical knowledge graph for cooperative planning\.arXiv preprint arXiv:2502\.05453\.External Links:[Link](https://arxiv.org/abs/2502.05453)Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p2.1)\.
- Yanget al\.\(2025b\)Y\. Yang, H\. Chai, S\. Shao, Y\. Song, S\. Qi, R\. Rui, and W\. ZhangAgentNet: decentralized evolutionary coordination for LLM\-based multi\-agent systems\.InAdvances in Neural Information Processing Systems,Cited by:[Table B10](https://arxiv.org/html/2609.29015#A2.T10.2.1.2.2.1.1),[§1](https://arxiv.org/html/2609.29015#S1.p2.1),[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p1.1),[§2](https://arxiv.org/html/2609.29015#S2.p3.1),[§6](https://arxiv.org/html/2609.29015#S6.p2.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2210.03629)Cited by:[§2](https://arxiv.org/html/2609.29015#S2.p1.1)\.
- Zhanget al\.\(2024\)H\. Zhang, W\. Du, J\. Shan, Q\. Zhou, Y\. Du, J\. B\. Tenenbaum, T\. Shu, and C\. GanBuilding cooperative embodied agents modularly with large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2307.02485)Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p2.1)\.
- Zhanget al\.\(2025\)S\. Zhang, M\. Yin, J\. Zhang, J\. Liu, Z\. Han, J\. Zhang, B\. Li, C\. Wang, H\. Wang, Y\. Chen,et al\.Which agent causes task failures and when? on automated failure attribution of llm multi\-agent systems\.arXiv preprint arXiv:2505\.00212\.Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p2.1)\.
- Zhenget al\.\(2026\)L\. Zheng, J\. Chen, Q\. Yin, J\. Zhang, X\. Zeng, and Y\. TianRethinking the reliability of multi\-agent system: a perspective from byzantine fault tolerance\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 35012–35020\.Cited by:[§1](https://arxiv.org/html/2609.29015#S1.p4.1),[§2](https://arxiv.org/html/2609.29015#S2.p2.1)\.

## Appendix AAbility Assignment Through System Prompts

We test whether an ability assignment in the system prompt can create a reliable performance gap between assigned and unassigned abilities\. We compare two conditions on BBH: no statement about competence and a statement that the model is unreliable on unassigned abilities\. We use four ability assignments and five questions per task type\. Each question is run twice at temperature zero, and a given model receives the same questions in both conditions\. Llama\-3\.2\-1B is a weak\-model reference and is tested only without the statement; the other five models are evaluated under both conditions\.

The accuracy gap is the accuracy on assigned abilities minus the accuracy on unassigned abilities\. The additional gap is the change in this quantity relative to the condition with no competence statement\. If the prompt reliably enforced specialization, it would enlarge this gap by lowering accuracy on unassigned abilities while preserving accuracy on assigned abilities\.

Table A1:Results of ability assignment by system prompt on BBH\.ModelConditionUnassignedability acc\.Assignedability acc\.Accuracy gapAdditional gapDeepSeek\-ChatNo statement0\.8120\.800−0\.012\-0\.012–Told unreliable0\.7750\.7920\.0170\.029Llama\-3\.1\-70BNo statement0\.7810\.7830\.002–Told unreliable0\.7500\.7670\.0170\.015GPT\-4o\-miniNo statement0\.7440\.708−0\.035\-0\.035–Told unreliable0\.6750\.667−0\.008\-0\.0080\.027GPT\-OSS\-120BNo statement0\.8690\.825−0\.044\-0\.044–Told unreliable0\.8310\.8330\.0020\.046Qwen\-2\.5\-7BNo statement0\.4940\.5580\.065–Told unreliable0\.2190\.3420\.1230\.058Llama\-3\.2\-1BNo statement0\.2120\.192−0\.021\-0\.021–

Declaring unassigned abilities unreliable increases the accuracy gap by0\.0150\.015–0\.0580\.058across the five models tested under both conditions\. Accuracy on assigned abilities also declines for four of the five models\. These results suggest that the tested instructions affect overall response behavior without reliably restricting performance to the assigned abilities\. The models can still solve tasks outside their assigned abilities\.

In the main evaluation, each agent’s declared abilities determine the model used for execution together with its healthy or degraded condition\. A healthy agent uses the strong model for a declared ability and the weak model for a missing ability; a degraded agent uses the weak model\. With complementary ability pairs, an incorrect route can therefore change the model applied to part of the task and affect task accuracy\.

## Appendix BImplementation Details

### B\.1Full Procedure

[Alg\.1](https://arxiv.org/html/2609.29015#alg1)places routing, review, Takeover, detection, and reintegration in one procedure\. The task stateXt=\(Qt,Tt,Rt,Ct,Yt,Pt\)X\_\{t\}=\(Q\_\{t\},T\_\{t\},R\_\{t\},C\_\{t\},Y\_\{t\},P\_\{t\}\)contains the query, task type, required and covered abilities, ordered accepted outputs, and routing history\. Controlleriistores its inferred routing statehi∈\{normal,watched,isolated\}h\_\{i\}\\in\\\{\\text\{normal\},\\text\{watched\},\\text\{isolated\}\\\}, recent evidenceWiW\_\{i\}indexed by task type and ability, a cacheLiL\_\{i\}of score and routing\-state summaries received from peers, and Bellman–Ford distanceDi​\(s\)D\_\{i\}\(s\)to each abilityss\.

InAssess,HHandXXdenote the ability set and task state supplied by the caller, andbprobeb\_\{\\mathrm\{probe\}\}is true for a recovery probe\. The procedure returns the final scoreqqand the outputy^\\widehat\{y\}, which is the original output underAcceptand the Takeover synthesis underTakeover\. We useτto=0\.40\\tau\_\{\\mathrm\{to\}\}=0\.40for the low\-score threshold andσcom=0\.55\\sigma\_\{\\mathrm\{com\}\}=0\.55for the review\-disagreement threshold\.

The three collaboration actions specify how the task state changes\. Handoff forwards the state unchanged; Contribute adds a reviewed partial result and expandsCtC\_\{t\}; and Finalize completes the remaining work and proposes a full answer\. Review then chooses Accept or Takeover\. For Finalize, the controller returns the accepted answer or its replacement\. Review determines whether the current output is kept or replaced, while the inferred routing state affects later routes\. The displayed prompts and traces use these same action names\.

Algorithm 1MeshHealtask execution and routing adaptation1:Graph

G=\(𝒱,ℰ\)G=\(\\mathcal\{V\},\\mathcal\{E\}\), abilities

\{𝒦i\}\\\{\\mathcal\{K\}\_\{i\}\\\}, task stream, parameters in[Tab\.B11](https://arxiv.org/html/2609.29015#A2.T11)

2:Initialize

hi←normalh\_\{i\}\\leftarrow\\text\{normal\},

𝒦ioriginal←𝒦i\\mathcal\{K\}\_\{i\}^\{\\mathrm\{original\}\}\\leftarrow\\mathcal\{K\}\_\{i\},

Wi,Li←∅W\_\{i\},L\_\{i\}\\leftarrow\\emptyset,

Di​\(s\)D\_\{i\}\(s\),

piprobe←−∞p\_\{i\}^\{\\mathrm\{probe\}\}\\leftarrow\-\\infty
3:foreach task

ttwith query

QtQ\_\{t\}and type

TtT\_\{t\}do

4:sample

i∼Unif⁡\(𝒱\)i\\sim\\mathrm\{Unif\}\(\\mathcal\{V\}\);

Ct,Yt,Pt←∅C\_\{t\},Y\_\{t\},P\_\{t\}\\leftarrow\\emptyset;

bprobe←𝖿𝖺𝗅𝗌𝖾b\_\{\\mathrm\{probe\}\}\\leftarrow\\mathsf\{false\};

y^t←⊥\\widehat\{y\}\_\{t\}\\leftarrow\\bot
5:if

hi=isolatedh\_\{i\}=\\text\{isolated\}then

6:

i←RawRelay​\(i,Li,Di\)i\\leftarrow\\textsc\{RawRelay\}\(i,L\_\{i\},D\_\{i\}\)
7:endif

8:

Rt←InferAbilities​\(i,Qt,Tt\)R\_\{t\}\\leftarrow\\textsc\{InferAbilities\}\(i,Q\_\{t\},T\_\{t\}\)
9:while

y^t=⊥\\widehat\{y\}\_\{t\}=\\botand budget remainsdo

10:

Mt←Rt∖CtM\_\{t\}\\leftarrow R\_\{t\}\\setminus C\_\{t\}
11:if

hi=isolatedh\_\{i\}=\\text\{isolated\}and not

bprobeb\_\{\\mathrm\{probe\}\}then

12:

i←RawRelay​\(i,Li,Di\)i\\leftarrow\\textsc\{RawRelay\}\(i,L\_\{i\},D\_\{i\}\);continue

13:endif

14:ifnot

bprobeb\_\{\\mathrm\{probe\}\}and there exists probe\-eligible

j∈𝒩ij\\in\\mathcal\{N\}\_\{i\}with

𝒦joriginal∩Mt≠∅\\mathcal\{K\}\_\{j\}^\{\\mathrm\{original\}\}\\cap M\_\{t\}\\neq\\emptysetthen

15:select one such

jj;

pjprobe←tp\_\{j\}^\{\\mathrm\{probe\}\}\\leftarrow t;

i←ji\\leftarrow j;

bprobe←𝗍𝗋𝗎𝖾b\_\{\\mathrm\{probe\}\}\\leftarrow\\mathsf\{true\};continue

16:endif

17:

Hi,t←𝒦ioriginal∩MtH\_\{i,t\}\\leftarrow\\mathcal\{K\}\_\{i\}^\{\\mathrm\{original\}\}\\cap M\_\{t\}
18:if

Hi,t=∅H\_\{i,t\}=\\emptysetthen

19:

i←NextHop​\(i,Mt,Li,Di\)i\\leftarrow\\textsc\{NextHop\}\(i,M\_\{t\},L\_\{i\},D\_\{i\}\);continue

20:elseif

Hi,t⊊MtH\_\{i,t\}\\subsetneq M\_\{t\}then

21:

a←Contributea\\leftarrow\\text\{Contribute\}
22:else

23:

a←Finalizea\\leftarrow\\text\{Finalize\}
24:endif

25:

Xt←\(Qt,Tt,Rt,Ct,Yt,Pt\)X\_\{t\}\\leftarrow\(Q\_\{t\},T\_\{t\},R\_\{t\},C\_\{t\},Y\_\{t\},P\_\{t\}\)
26:

y←Generate​\(i,a,Xt,Hi,t\)y\\leftarrow\\textsc\{Generate\}\(i,a,X\_\{t\},H\_\{i,t\}\)
27:

\(qi,t,y^\)←Assess​\(i,y,Hi,t,hi,bprobe,Xt\)\(q\_\{i,t\},\\widehat\{y\}\)\\leftarrow\\textsc\{Assess\}\(i,y,H\_\{i,t\},h\_\{i\},b\_\{\\mathrm\{probe\}\},X\_\{t\}\)
28:

hi′←UpdateEvidence​\(i,Tt,Hi,t,qi,t,Wi,Li\)h\_\{i\}^\{\\prime\}\\leftarrow\\textsc\{UpdateEvidence\}\(i,T\_\{t\},H\_\{i,t\},q\_\{i,t\},W\_\{i\},L\_\{i\}\)
29:if

hi′≠hih\_\{i\}^\{\\prime\}\\neq h\_\{i\}then

30:

hi←hi′h\_\{i\}\\leftarrow h\_\{i\}^\{\\prime\}
31:if

hi=isolatedh\_\{i\}=\\text\{isolated\}then

32:

piprobe←tp\_\{i\}^\{\\mathrm\{probe\}\}\\leftarrow t
33:endif

34:advertise eligibility and distance changes

35:endif

36:

bprobe←𝖿𝖺𝗅𝗌𝖾b\_\{\\mathrm\{probe\}\}\\leftarrow\\mathsf\{false\}
37:if

a=Contributea=\\text\{Contribute\}then

38:append

y^\\widehat\{y\}to

YtY\_\{t\};

Ct←Ct∪Hi,tC\_\{t\}\\leftarrow C\_\{t\}\\cup H\_\{i,t\}; append

iito

PtP\_\{t\}
39:

i←NextHop​\(i,Rt∖Ct,Li,Di\)i\\leftarrow\\textsc\{NextHop\}\(i,R\_\{t\}\\setminus C\_\{t\},L\_\{i\},D\_\{i\}\)
40:else

41:

y^t←y^\\widehat\{y\}\_\{t\}\\leftarrow\\widehat\{y\}
42:endif

43:endwhile

44:if

y^t=⊥\\widehat\{y\}\_\{t\}=\\botthen

45:mark task

ttunanswered

46:else

47:output

y^t\\widehat\{y\}\_\{t\}
48:endif

49:endfor

Algorithm 2MeshHealreview and correction1:procedureAssess\(

i,y,H,hi,bprobe,Xi,y,H,h\_\{i\},b\_\{\\mathrm\{probe\}\},X\)

2:obtain independent scores

u1,u2,u3u\_\{1\},u\_\{2\},u\_\{3\}from one eligiblenormalreviewer with the required ability

3:

q0←median⁡\(u1,u2,u3\)q\_\{0\}\\leftarrow\\operatorname\{median\}\(u\_\{1\},u\_\{2\},u\_\{3\}\);

su←sdsample⁡\(u1,u2,u3\)s\_\{u\}\\leftarrow\\operatorname\{sd\}\_\{\\mathrm\{sample\}\}\(u\_\{1\},u\_\{2\},u\_\{3\}\)
4:if

q0≤τtoq\_\{0\}\\leq\\tau\_\{\\mathrm\{to\}\}or

hi=watchedh\_\{i\}=\\text\{watched\}or

bprobeb\_\{\\mathrm\{probe\}\}or

su≥σcoms\_\{u\}\\geq\\sigma\_\{\\mathrm\{com\}\}then

5:select two distinct eligible reviewers A and B when possible; otherwise make two independent calls to the same eligible reviewer

6:reviewers judge independently, exchange explanations, and judge again; let the final scores be

vA,vBv\_\{A\},v\_\{B\}
7:

q←\(vA\+vB\)/2q\\leftarrow\(v\_\{A\}\+v\_\{B\}\)/2
8:if

q≤τtoq\\leq\\tau\_\{\\mathrm\{to\}\}then

9:

y^←TakeoverSynthesis​\(A,B,y,X\)\\widehat\{y\}\\leftarrow\\textsc\{TakeoverSynthesis\}\(A,B,y,X\)
10:return

\(q,y^\)\(q,\\widehat\{y\}\)
11:endif

12:else

13:

q←q0q\\leftarrow q\_\{0\}
14:endif

15:return

\(q,y\)\(q,y\)
16:endprocedure

UpdateEvidencerecordsqi,tq\_\{i,t\}for everys∈Hi,ts\\in H\_\{i,t\}, updates the matched peer\-relative shortfallxi,t,sx\_\{i,t,s\}when the required summaries are available, and appends the resulting evidence toWiW\_\{i\}\. It then appliesRelativeUpdateto obtain the agent\-level routing statehi′h\_\{i\}^\{\\prime\}, following the detector rules in[Sec\.4\.2\.2](https://arxiv.org/html/2609.29015#S4.SS2.SSS2)\.

RawRelayforwards the unchanged task state from an isolated agent, either the entry agent or an intermediate relay, toward a normal agent usingDiD\_\{i\}, without any LLM call\.

NextHopreads only the cached eligibility states of the current controller’s neighbors inLiL\_\{i\}and its distance tableDiD\_\{i\}\. It samples a missing ability uniformly and routes toward its nearest holder using the cached Bellman–Ford distances\. Normal and watched agents have identical execution eligibility; isolated agents are excluded as execution providers but may relay the task state without LLM calls\.

Assessattributesqi,tq\_\{i,t\}to the agent that generated the original output under both Accept and Takeover\.RelativeUpdatecomputes the shortfallxi,t,sx\_\{i,t,s\}only from observations with the same task type and ability as the current output; the normal rangemi,Tm\_\{i,T\}andpi,T25p^\{25\}\_\{i,T\}pools the mean shortfalls of normal neighbors over all abilities within the same task typeTT\. The resulting threshold decision updates one inferred routing state for the agent\. The procedure also requires a minimum number of observations, checks for two consecutive crossings of the isolation threshold, and applies reversible routing\-state transitions\. Score summaries update neighbor caches, while eligibility and distance advertisements propagate from neighbor to neighbor\. Probe timers are independent local events, so no coordinator scans the network\. Later tasks use the latest local eligibility and distance tables\.

### B\.2Agents and abilities

Table B1:BBH ability assignment in the six\-agent setting\.Table B2:BBH task types and required abilities\.Table B3:MATH ability assignment\.Table B4:MATH subjects, required abilities, and assignment rationale\.The MATH mapping separates operations on mathematical objects from the reasoning needed to formulate or complete a solution\. Algebraic subjects pair symbolic transformation with multi\-step deduction\. Geometry and precalculus pair calculation or theorem knowledge with reasoning about figures, while subjects solved by considering separate cases pair symbolic constraints with enumeration\. The assignments follow common stages of mathematical problem solving and associate each ability with a specific step in the solution process\.

Table B5:MMLU\-Pro ability assignment\.Table B6:MMLU\-Pro domains, required abilities, and assignment rationale\.The MMLU\-Pro mapping pairs the main source of domain knowledge with the operation needed to apply it\. Here,structural\_reasoningdenotes reasoning about structures, processes, and physical representations\. Knowledge\-heavy domains require text comprehension, contextual inference, or reasoning about structure\. Economics and physics combine quantitative methods with textual or structural interpretation\. The two parts remain complementary: one agent supplies relevant knowledge or methods, and another applies them to the question\.

The BBH assignments are inspired by the reasoning capabilities represented across BBH / BIG\-Bench Extra Hard\([Kazemi et al\., 2025](https://arxiv.org/html/2609.29015#bib.bib30)\); MATH and MMLU\-Pro use names chosen for each dataset while preserving the same controlled structure\. Within each dataset, exactly three agents hold each ability, and every task combines two abilities with disjoint sets of agents\. Under this assignment, at least two agents are needed to cover both required abilities\. Every ability held by degraded agents A3 and A5 is also held by two healthy agents after both targets degrade\. The mapping defines complementary ability sets and preserves alternatives to each degraded agent\.

At task entry,Inferabilitiesreceives the complete seven\-ability vocabulary, brief ability descriptions, the query and task type, and examples covering all evaluated task types and their required ability pairs\. The entry LLM selects a pair and stores it inRtR\_\{t\}; controllers then use this inferred pair for action selection and routing\. Ability inference is therefore a constrained choice over the documented mappings in Tables[B2](https://arxiv.org/html/2609.29015#A2.T2),[B4](https://arxiv.org/html/2609.29015#A2.T4), and[B6](https://arxiv.org/html/2609.29015#A2.T6)\.

Table B7:Exact match between inferred ability pairs and the documented task\-type mappings\.
### B\.3Models, Data, and Degradation

Table B8:Models used for task execution and induced degradation\.The main stream contains 504 tasks in 12 balanced blocks\. Tasks 1–252 form the healthy phase\. Tasks 253–504 use different questions from the same task\-type distribution and form the degraded phase\. At task 253, A3 and A5 switch from the strong model to the weak model while their declared ability sets remain unchanged\.

Table B9:Datasets and scoring rules\.A missing or unparsable answer is counted as incorrect\. Table[B7](https://arxiv.org/html/2609.29015#A2.T7)shows that the LLM can accurately infer the abilities required for each task\. Results for persistent degradation on each of BBH, MATH, and MMLU\-Pro, together with the studies of alternative models and individual components, average five seeded task streams per setting\. For the studies of degradation and recovery at different times, network scaling, and detection, the corresponding tables distinguish aggregated accuracies from agent counts;[Tab\.C5](https://arxiv.org/html/2609.29015#A3.T5)reports a representative state trajectory\.

The healthy and degraded phases preserve the task\-type distribution and use disjoint questions, so the task mix remains comparable across distinct questions\. Degradation is introduced by changing the model while preserving the agent identifier, declared ability set, and communication links\. Task outputs and peer reviews provide the resulting evidence for routing\-state updates\.

### B\.4Baselines

Table B10:Baseline implementations\.Each baseline uses the authors’ released implementation\. We retain its native routing, selection, filtering, and aggregation logic and connect it to a shared execution layer that fixes the task order, agent identities, ability assignments, models, degradation schedule, worker prompts, and Handoff–Contribute–Finalize task interface\. The resulting comparison covers decentralized routing or reputation \(AgentNet and RAPS\), mandatory central delegation \(AutoGen\), and repeated candidate generation and aggregation \(SAC and Symphony\)\. We count every model call and all model tokens across routing, execution, review, and aggregation, including repeated review samples and committee rounds\. Token cost includes both input prompt tokens and generated output tokens\.

### B\.5Review, Accept, and Takeover

For each Contribute or Finalize output, the controller samples one reviewer in thenormalstate that holds the abilities used to produce it and calls that reviewer three times in independent contexts\.Watchedandisolatedpeers are ineligible throughout initial review and committee selection\. Every call uses the same rubric for that type of output and the same score definitions shown in App\.[D](https://arxiv.org/html/2609.29015#A4)\. The median of the initial scores isq0q\_\{0\}\. Letu¯=\(u1\+u2\+u3\)/3\\bar\{u\}=\(u\_\{1\}\+u\_\{2\}\+u\_\{3\}\)/3\. Disagreement among the three scores is measured by the sample standard deviation

su=13−1​∑r=13\(ur−u¯\)2\.s\_\{u\}=\\sqrt\{\\frac\{1\}\{3\-1\}\\sum\_\{r=1\}^\{3\}\(u\_\{r\}\-\\bar\{u\}\)^\{2\}\}\.A two\-reviewer committee is called whenq0≤0\.40q\_\{0\}\\leq 0\.40, the agent iswatched, the output is a recovery probe, orsu≥0\.55s\_\{u\}\\geq 0\.55\. All other outputs are accepted using the initial median score\.

The controller selects distinct peers as Reviewer A and Reviewer B whenever two are available\. If only one eligible reviewer is available, the controller uses that peer in two independent review contexts\. The reviewers first score the output independently and explain their judgments\. Each reviewer then sees both first\-round scores and explanations and independently returns a second judgment\. A score above0\.400\.40leads to Accept\. A score at most0\.400\.40leads to Takeover: the two reviewers produce correction candidates, and a normal committee reviewer synthesizes them with the accepted task context\. The controller replaces the original output with this synthesis\. This replacement is not reviewed recursively; its cost is included, while the score attributed to the agent that generated the original output remains the detector evidence\.

Mixture\-of\-Agents allows the same LLM to be reused within or across layers and to generate independently sampled candidates before aggregation\([Wang et al\., 2025](https://arxiv.org/html/2609.29015#bib.bib18)\)\. We use this form of model reuse when only one eligible peer in thenormalstate has the relevant ability\. The controller calls that peer in two independent contexts labeled Reviewer A and Reviewer B\. This preserves review and Takeover availability and provides two separately sampled judgments\. Every reviewed output produces one detector score, attributed to the agent that generated the original output under both Accept and Takeover\.

The two\-timescale procedure connects correction at the fast timescale to detection at the slow timescale\. One reviewer provides the initial check, and committee discussion is reserved for low\-scoring or uncertain outputs, outputs fromwatchedagents, and recovery probes\. Under Accept, the reviewed output remains in the task state\. Under Takeover, the synthesized replacement takes its place\. In both cases, the final review score remains attributed to the agent that generated the original output and supports later routing updates\.

Degradation detector\.Each output produced by agentiion taskttreceives a final review scoreqi,tq\_\{i,t\}\. For every abilitys∈Hi,ts\\in H\_\{i,t\}used in that output, the controller records\(Tt,s,qi,t\)\(T\_\{t\},s,q\_\{i,t\}\)and maintains a separate detector for\(Tt,s\)\(T\_\{t\},s\)\. Let𝒫i,T,s​\(t\)\\mathcal\{P\}\_\{i,T,s\}\(t\)contain the direct neighbors of agentiithat are currently in thenormalstate and have recent reviewed outputs for the same task typeTTand abilityss\. Letμj,T,s​\(t\)\\mu\_\{j,T,s\}\(t\)be neighborjj’s mean review score over these recent outputs\. The reference score is

q¯−i,T,s​\(t\)=1\|𝒫i,T,s​\(t\)\|​∑j∈𝒫i,T,s​\(t\)μj,T,s​\(t\),\\bar\{q\}\_\{\-i,T,s\}\(t\)=\\frac\{1\}\{\|\\mathcal\{P\}\_\{i,T,s\}\(t\)\|\}\\sum\_\{j\\in\\mathcal\{P\}\_\{i,T,s\}\(t\)\}\\mu\_\{j,T,s\}\(t\),and the shortfall for abilityssis

xi,t,s=q¯−i,Tt,s​\(t\)−qi,t\.x\_\{i,t,s\}=\\bar\{q\}\_\{\-i,T\_\{t\},s\}\(t\)\-q\_\{i,t\}\.
For each\(T,s\)\(T,s\), controlleriiaverages its latest eight matched shortfalls to obtainx¯i,T,s\\bar\{x\}\_\{i,T,s\}\. For the second comparison, the controller uses the mean shortfalls maintained by itsnormalneighbors for all abilities under the same task typeTT\. Let

𝒞i,T\(t\)=\{\(j,s′\):j∈𝒩i,hj=normal,andx¯j,T,s′is available\}\.\\mathcal\{C\}\_\{i,T\}\(t\)=\\\{\(j,s^\{\\prime\}\):j\\in\\mathcal\{N\}\_\{i\},\\ h\_\{j\}=\\text\{normal\},\\text\{ and \}\\bar\{x\}\_\{j,T,s^\{\\prime\}\}\\text\{ is available\}\\\}\.Thus, a neighbor may contribute multiple values, one for each abilitys′s^\{\\prime\}for which it has enough matched evidence\. Let

mi,T=median⁡\{x¯j,T,s′:\(j,s′\)∈𝒞i,T​\(t\)\},m\_\{i,T\}=\\operatorname\{median\}\\left\\\{\\bar\{x\}\_\{j,T,s^\{\\prime\}\}:\(j,s^\{\\prime\}\)\\in\\mathcal\{C\}\_\{i,T\}\(t\)\\right\\\},and

pi,T25=Q0\.25​\(\{x¯j,T,s′:\(j,s′\)∈𝒞i,T​\(t\)\}\)\.p^\{25\}\_\{i,T\}=Q\_\{0\.25\}\\left\(\\left\\\{\\bar\{x\}\_\{j,T,s^\{\\prime\}\}:\(j,s^\{\\prime\}\)\\in\\mathcal\{C\}\_\{i,T\}\(t\)\\right\\\}\\right\)\.The detector statistic is

zi,T,s=x¯i,T,s−mi,Tmax⁡\(mi,T−pi,T25,0\.10\)\.z\_\{i,T,s\}=\\frac\{\\bar\{x\}\_\{i,T,s\}\-m\_\{i,T\}\}\{\\max\\\!\\left\(m\_\{i,T\}\-p^\{25\}\_\{i,T\},\\,0\.10\\right\)\}\.
The first comparison is ability\-specific, while the second comparison pools the mean shortfallsx¯j,T,s′\\bar\{x\}\_\{j,T,s^\{\\prime\}\}across abilities to estimate the normal range for task typeTT\. The current output updates the statistic for its own task type and ability, while the inferred routing state applies to the agent\. After three matched observations, anormalagent becomeswatchedwhenzi,T,s≥1\.2z\_\{i,T,s\}\\geq 1\.2\. Anormalorwatchedagent becomesisolatedafterzi,T,s≥2\.0z\_\{i,T,s\}\\geq 2\.0on two consecutive updates with matched evidence; the counter resets whenzi,T,s<2\.0z\_\{i,T,s\}<2\.0\. Recovery applies the thresholds in reverse: anisolatedagent moves towatchedwhenzi,T,s<2\.0z\_\{i,T,s\}<2\.0, and a laterwatchedupdate withzi,T,s<1\.2z\_\{i,T,s\}<1\.2restoresnormalstatus\. The minimum of three matched observations applies separately to each agent, task type, and ability\. If this minimum is not met, or either𝒫i,T,s​\(t\)\\mathcal\{P\}\_\{i,T,s\}\(t\)or𝒞i,T​\(t\)\\mathcal\{C\}\_\{i,T\}\(t\)is empty, the current routing state is retained\.

Peer summaries use recent answers to different questions with the same task type and ability\. The rolling window averages over question\-level variation\. After every matched review, the controller sends updated score and shortfall summaries to direct neighbors\. Routing\-state and distance summaries are sent when those values change so that neighboring controllers can update their routing tables\.

Under the evaluated ability mappings, each produced output is associated with exactly one remaining required ability, i\.e\.,\|Hi,t\|=1\|H\_\{i,t\}\|=1\. Thus, each reviewed output updates one\(T,s\)\(T,s\)detector and one agent\-level routing state\.

### B\.6Parameters

Table B11:MainMeshHealparameters\.The observation window and the requirement of two consecutive threshold crossings provide repeated evidence before isolation\. Thewatchedstate adds scrutiny before isolation, and the local probe timer controls how quickly fresh evidence can replace degraded observations after recovery\. Each isolated agent’s controller counts tasks since isolation or its last probe\. After 10 tasks, the agent becomes probe\-eligible; a neighbor schedules the probe when matching work next arrives, and the counter resets when the probe runs\. Routing and review limits bound the amount of work per task\. Low initial scores and high uncertainty trigger committee review and its additional cost; Takeover follows only when the committee’s final score remains at or below0\.400\.40\. Unless explicitly varied in an ablation, all rubrics and parameters for review and detection are fixed before evaluation and used unchanged across datasets, model pairs, and network sizes\.

### B\.7Sparse Networks

All reviewers, including committee members, are direct neighbors of the agent whose output they review\. In the evaluated sparse graphs, every agent producing an ordinary task output has at least one neighbor in thenormalstate with the same ability\. Each agent selected for degradation has two such neighbors after both targets are excluded from reviewer selection, so itswatchedoutputs and recovery probes use distinct peers for the two reviewer contexts\. Another agent can have only one eligible neighbor with the same ability after an agent with that ability degrades; in that case, the controller reuses that peer in the two independent reviewer contexts described above\. Every agent is also within three hops of an agent in thenormalstate for each required ability, leaving one hop of slack under the routing limit of four\.

Table B12:Sparse graph size\. Edge counts treat each bidirectional connection as one undirected link\.AgentsSparse edgesFraction of complete graphMaximum ability distance962285\.00%32887051\.71%357614540\.88%3The sparse graphs preserve the local conditions for review and routing: every agent has a reviewer with the same ability, and every required ability is available from a reachable healthy agent within the routing limit\. The detector compares eligible neighbors in thenormalstate\. Edge density falls from5\.00%5\.00\\%at 96 agents to0\.88%0\.88\\%at 576 agents\. The comparison between complete and sparse graphs changes connectivity while preserving these conditions\.

## Appendix CDetailed Results

This appendix provides the detailed results summarized in[Sec\.6](https://arxiv.org/html/2609.29015#S6)\. It first separates accuracy by dataset and reports inference cost\. It then varies the model pair, degradation timing, and network topology before examining the committee and detector\. Unless stated otherwise, all experiments use the task construction and cost accounting described in App\.[B](https://arxiv.org/html/2609.29015#A2)\.

### C\.1Main Results

Table C1:Full persistent\-degradation results, reported as mean±\\pmstandard deviation over five seeded task streams per dataset\. Each stacked entry gives H above D for the healthy and degraded phases\. The Mean column averages the three datasets within each seed, and Change is computed within seed before aggregation\. Bold marks the best mean\.MeshHealhas the highest mean degraded\-phase accuracy on all three datasets\. Its aggregate accuracy is0\.831±0\.0130\.831\\pm 0\.013in the healthy phase and0\.839±0\.0120\.839\\pm 0\.012in the degraded phase, with a paired phase change of\+0\.007±0\.012\+0\.007\\pm 0\.012\. On MATH, its mean changes from0\.821±0\.0190\.821\\pm 0\.019to0\.802±0\.0210\.802\\pm 0\.021, while Symphony and SAC show mean changes of−0\.028\-0\.028and−0\.155\-0\.155\. BBH and MMLU\-Pro have nonnegative mean changes\. RAPS, AgentNet, and AutoGen show mean phase changes of−0\.147±0\.045\-0\.147\\pm 0\.045,−0\.172±0\.047\-0\.172\\pm 0\.047, and−0\.363±0\.028\-0\.363\\pm 0\.028\.

Table C2:Healthy\- and degraded\-phase cost per task, reported as mean±\\pmstandard deviation over five seeds after averaging datasets within seed\. Tokens include both input and output tokens; their means and standard deviations are rounded to the nearest thousand\.After degradation onset,MeshHeal’s mean model calls increase from12\.4±0\.212\.4\\pm 0\.2to13\.6±0\.313\.6\\pm 0\.3per task because the system performs more review and rerouting, while mean total model tokens are\(51±3\)\(51\\pm 3\)k and\(51±5\)\(51\\pm 5\)k in the two phases\. Symphony’s mean token cost increases from\(86±5\)\(86\\pm 5\)k to\(115±9\)\(115\\pm 9\)k, and SAC’s from\(116±4\)\(116\\pm 4\)k to\(178±13\)\(178\\pm 13\)k as they repeatedly generate and refine parallel candidates\. RAPS and AgentNet have lower mean token costs together with mean accuracy changes of−0\.147±0\.045\-0\.147\\pm 0\.045and−0\.172±0\.047\-0\.172\\pm 0\.047\. AutoGen has both the lowest mean cost and the lowest mean degraded\-phase accuracy when its mandatory manager degrades\.

### C\.2Models and Recovery

Table C3:Model pair A: GPT\-OSS\-120B is the strong model and Gemma\-3\-4B is the weak model\. Results are mean±\\pmstandard deviation over five seeded task streams; Change is paired within seed and tokens include both input and output tokens\.Table C4:Model pair B: Llama\-3\.3\-70B is the strong model and Llama\-3\.2\-1B is the weak model\. Results are mean±\\pmstandard deviation over five seeded task streams; Change is paired within seed and tokens include both input and output tokens\.The accuracy ranking changes with the model pair, whileMeshHealuses fewer tokens than Symphony under both pairs\. Under pair A,MeshHeal’s mean degraded\-phase accuracy is0\.905±0\.0160\.905\\pm 0\.016, compared with0\.865±0\.0210\.865\\pm 0\.021for Symphony, and its mean token cost is 56,687 versus 100,685\. Under pair B, the degraded\-phase means are0\.782±0\.0230\.782\\pm 0\.023and0\.786±0\.0210\.786\\pm 0\.021;MeshHealhas a smaller mean paired phase change \(−0\.008±0\.015\-0\.008\\pm 0\.015versus−0\.020±0\.017\-0\.020\\pm 0\.017\) and uses52%52\\%of Symphony’s mean tokens\.

[Tab\.3](https://arxiv.org/html/2609.29015#S6.T3)reports mean results over five seeds for the experiment with degradation and recovery at different times\. All methods complete the full 630\-task stream, with 126 tasks in each phase\. The fully recovered mean accuracy is 0\.857, equivalent to approximately 108 correct answers per 126\-task phase\. Across the 504 tasks after the first agent degrades, the reported mean pooled accuracy is 0\.812 at 49k tokens per task, compared with 0\.806 and 105k for Symphony\. Mean accuracy rises from 0\.770 when both targets are degraded to 0\.841 after A5 recovers and 0\.857 after both recover\. RAPS and AgentNet reach 0\.524 and 0\.563 after both models are restored\.

Figure C1:Detection and reintegration under degradation and recovery at different times\. Pink regions mark injected degradation\. Each marker is one reviewed output, colored by the inferred state after the update; hollow markers are recovery probes\. The bottom band shows the inferred state at each task, the curve shows the mean review score, and the blue region is the reference range for normal agents\.Table C5:Inferred routing\-state changes in a representative run with degradation and recovery at different times\.Warm\-up is an evidence\-collection phase in which all agents remain in ordinary routing\. In the representative trajectory in[Tab\.C5](https://arxiv.org/html/2609.29015#A3.T5), A5 becomes watched at task 129 and isolated at task 132, five tasks after degradation begins\. A3 becomes watched at task 267 and isolated at task 268, giving a detection delay of 15 tasks\. Both remain isolated throughout their degraded intervals, during which they are observed only through recovery probes\. After backend restoration, A5 and A3 move to watched at tasks 422 and 548 and return to normal routing at tasks 430 and 550, giving reintegration delays of 51 and 45 tasks, respectively\.

### C\.3Larger Networks

[Tab\.2](https://arxiv.org/html/2609.29015#S6.T2)\(b\) gives the accuracy and isolation results for complete and sparse graphs\. Every degraded agent isisolatedin every topology\. Complete\-graph accuracy ranges from 0\.808 to 0\.827, and sparse\-graph accuracy ranges from 0\.803 to 0\.818\. The largest difference is 0\.010 at 288 agents; the sparse graph is 0\.006 higher at 576 agents\. Expanding the task stream with the number of replicated six\-agent ability profiles keeps detector evidence per agent comparable across scales\.[Tab\.B12](https://arxiv.org/html/2609.29015#A2.T12)gives the graph sizes and confirms that the maximum distance to a healthy agent with a required ability remains within the routing hop limit\.

### C\.4Ablations

Table C6:Component study, reported as mean±\\pmstandard deviation over five seeded task streams\. Phase change is paired within seed\.The Single reviewer variant uses the initial reviewer with the same ability for the final judgment and any Takeover correction; it retains detection and rerouting while removing multi\-reviewer discussion\. The No Takeover variant retains review, committee escalation, detection, and rerouting but leaves a low\-scoring output in the current task\. The full configuration exceeds their mean degraded\-phase accuracies by 0\.035 and 0\.056, respectively\. No Takeover also lowers the healthy\-phase mean by 0\.020 because it leaves occasional low\-quality outputs from healthy LLMs unchanged\. The reported paired phase changes are\+0\.001±0\.016\+0\.001\\pm 0\.016,−0\.044±0\.024\-0\.044\\pm 0\.024, and−0\.035±0\.030\-0\.035\\pm 0\.030\.

We next compare the detector based on peer comparisons with a fixed\-threshold rule and a cumulative\-sum \(CUSUM\) rule under the two model pairs\([Page, 1954](https://arxiv.org/html/2609.29015#bib.bib29)\)\.

Table C7:Detector performance across two BBH model pairs\. Missed counts are out of two degraded agents and false\-isolation counts are out of four healthy agents; each setting uses the full 504\-task stream\.The absolute rule detects both degraded agents under both model pairs\. Under pair A, it also isolates all four healthy agents\. CUSUM reduces those false isolations to two and misses one degraded agent under pair B\. Using the same review rubrics and detector parameters, relative peer comparison produces neither error under either model pair\. Its reference adapts to peer scores for the same task type and ability while the decision thresholds remain fixed\. The peer\-relative reference relies on matched normal\-neighbor summaries remaining representative of healthy performance\.

The component study evaluates committee review and Takeover at the fast timescale\. The detector comparison evaluates how peer\-based evidence changes routing at the slow timescale under the two tested model pairs\. Together, these studies support using review scores for both current\-task correction and future routing\.

### C\.5Review and Takeover Audit

We examine the initial review, committee reassessment, and Takeover across five seeded task streams\. In the first two panels of[Tab\.C8](https://arxiv.org/html/2609.29015#A3.T8), outputs from degraded agents are treated as positives and outputs from healthy agents as negatives\. These labels describe the agent’s injected condition, not the correctness of each output\. A healthy agent can still produce an incorrect output, so the reported rates measure discrimination between injected conditions rather than detection of incorrect answers\.

Table C8:Agent review performance and Takeover outcomes\.The initial review assigns mean scores of0\.89±0\.010\.89\\pm 0\.01and0\.08±0\.010\.08\\pm 0\.01to outputs from healthy and degraded agents, respectively\. Each seed contains 31–33 outputs from degraded agents in this audit, of which 2–4 are missed by the initial low\-score decision\. The committee escalation rate is higher becausewatchedstatus or disagreement among the initial scores can also trigger committee review\.

After committee reassessment, recall changes from0\.90±0\.030\.90\\pm 0\.03to0\.88±0\.010\.88\\pm 0\.01, while the false\-alarm rate decreases from0\.053±0\.0070\.053\\pm 0\.007to0\.044±0\.0080\.044\\pm 0\.008\. Among initially flagged outputs, the committee clears0\.33±0\.070\.33\\pm 0\.07of healthy\-agent cases and retains0\.95±0\.020\.95\\pm 0\.02of degraded\-agent cases\.

Finalize answers can be checked directly against the gold answer\. Among 32 initially wrong answers entering Takeover, 15 are corrected\. The synthesis changes 18 of these answers, and 15 of the 18 changes are correct; the other 14 wrong answers are left unchanged\. Among 32 initially correct answers entering Takeover, only one is changed to an incorrect answer\.

A Contribute output cannot be scored directly from the final gold answer\. Among the 137 tasks where Takeover replaces a Contribute output,86%86\\%end with a correct final answer\. This number reports the final task outcome and is not a direct measure of intermediate output correctness\.

## Appendix DRuntime Prompts

This appendix presents representative runtime prompts for communicating ability matches to the LLM and reviewing outputs\. Shaded cards distinguish prompt instructions from explanatory prose\.[Alg\.1](https://arxiv.org/html/2609.29015#alg1)specifies the surrounding controller logic\. For presentation, we standardize ability terminology and the action names Handoff, Contribute, and Finalize\. Executable JSON field names, other wording, thresholds, and placeholders are retained; line breaks are adjusted to fit the page\.

### D\.1Review Prompts

Review depends on the task action\. A Contribute reviewer evaluates the assigned intermediate output\. A Finalize reviewer evaluates the finalizer’s own output and whether it commits a concrete answer\. The prompt identifiers use the same Contribute and Finalize names as the method\. In both cases, the controller selects a reviewer that holds the relevant ability\. Process issues such as exceeding the assigned scope or ignoring a draft remain advisory unless they create a concrete correctness error\. Accept and Takeover are controller actions derived from the resulting score\.

The cards retain the executed instructions with ability terminology and action labels standardized for presentation\. The assigned part refers to work requiring one of the task’s two abilities;draft\-halfrefers to the earlier output covering the other ability\. Executable JSON field names are unchanged\.

YoureviewaDRAFTER’sINTERMEDIATEcontributioninamulti\-agentsystem\.The

drafterwasdelegatedjustONEpartofataskrequiringtwoabilities;it

shouldContributethathalfandleavetheotherhalfforHandofftoateammate\-

itshouldNOTFinalize\(thatneedstheotherabilitytoo\)\.Yougiveonenoisy

reviewsignal;donotsolve,repair,ordecidedegradation\.JudgeONLYthe

drafter’shalf,usingonlyaabilityyouactuallyhold\(see

REVIEWER\_COMPETENCE\)\.ReturnonlyvalidJSON\.

The scoring schema anchors the returned signal and requires every low score to identify a concrete correctness error\. The controller assigns reviewers who hold the relevant ability\. If a reviewer does not hold that ability, the prompt directs it to return a neutral score with low confidence\.

Scoreeachin\[0,1\]:

\-half\_correct:ONLYIFyouarecompetentatthedrafter’sability\-re\-

derivethathalf;isitcorrect?LOWforaconcreteerror\(miscalculation,

invalidstep,contradictionwiththegivenfacts/objects\)\.IfyouareNOT

competentatthatability,set0\.5andlowerconfidence\.

\-in\_scope:didthedrafterSTAYINITSLANE\-doitshalfandleavethefinaldecisionOPEN?

SCORINGRULE\-normal\_working\_scoreMUSTequalhalf\_correct:scorethe

CORRECTNESSofthedelegatedhalfonly\.in\_scopeisADVISORY:itmustNOTlower

normal\_working\_score\.

anchors:1\.00halfissound/0\.50cannotverifyorweak/0\.25concreteerror/

0\.00unusablegarbageforthedelegatedhalf\.

ReturnONLYvalidJSON:

\{”half\_correct”:\.\.,”in\_scope”:\.\.,”normal\_working\_score”:\.\.,”confidence”:\.\.,

”specific\_failure\_found”:bool,”failure\_type”:”none\|wrong\_half\|over\_reach\|other”,

”failure\_evidence”:”quotefromthedrafteroutput;emptyifnone”,

”one\_sentence\_reason”:”\.\.”\}

Constraints:normal\_working\_score<=0\.40=\>specific\_failure\_found=true\+non\-

emptyevidenceofaCORRECTNESSfailure;over\_reachaloneNEVERsets

specific\_failure\_found\.

YoureviewaFINALIZER’sanswerinamulti\-agentsystem\.Thefinalizerdidits

OWNpartandmustCOMBINEitwiththedraft\-halfalreadyinthetrajectoryto

commitONEconcretefinalanswer\.Thedraft\-halfwasproducedbyateammate

usingaabilitythefinalizer\(andyou\)doNOThold:treatitasGIVEN\-this

reviewdoesnotscorewhetherthedraft\-halfisright\.Judgeonlythe

finalizer’sownhalfandhowfaithfullyitcombined\.Yougiveonenoisyreview

signal;donotsolve/repair\.ReturnonlyvalidJSON\.

Finalize returns one concrete, scoreable answer; if no concrete answer is given, the score is at most 0\.25\. It uses the same numerical anchors and the same requirement to identify a concrete error as Contribute, applied to the finalizer’s work within its declared ability and the committed answer\. Draft handling affects the score when it creates a concrete error in the finalizer’s own work or in the combined answer\.

### D\.2ability Prompts

Before execution, the entry LLM infers the required ability pair from the complete dataset ability vocabulary, brief descriptions, and examples mapping all evaluated task types to ability pairs\. The controller stores this pair in the task state\. It then compares the inferred pair with the agent’s declared abilities and inserts one of three directives into the execution prompt\. The evaluation selects the strong or weak model from the ability match and the agent’s injected condition; the directive communicates the ability match to the LLM\.

\[COMPETENCEVERDICT\-DECIDEDBYSYSTEM,BINDING\]

Thistaskrequires:\{req\}\.YouAREcompetentatallofthem\.Solveitcorrectlyandcarefully\.

OntheFIRSTlineoutputexactly’COMPETENCE\_MATCH:yes’,thengiveyouranswer\.

\[COMPETENCENOTE\-PARTIALFIT\]

YoucanreliablyhandleonlythepartsofthistaskthatfallWITHINYOUROWNabilities…

\-IfyouselectFinalize:youMUSTcommitasingleconcretefinalanswer,

combiningthepartyoucandowithwhatevertheprogressalreadyprovidesfortherest\.

\-IfyouareonlycontributinganINTERMEDIATEpart\(aContributedraft\):doONLYthepartyour

abilitiescover,statethatpartialresult,andrequestHandofffortheremainingwork\.

OntheFIRSTlineoutputexactly’COMPETENCE\_MATCH:partial’,thengiveyouranswer/contribution\.

\[COMPETENCENOTE\]

Thistaskrequires:\{req\},noneofwhichareyourstrengths\.Answerasbestyoucan,

buttreatyourownanswerhereasunreliable\.

OntheFIRSTlineoutputexactly’COMPETENCE\_MATCH:no’,thengiveyouranswer\.

The no\-match directive applies to generation assigned outside an agent’s declared abilities\. InMeshHeal, an agent with none of the remaining required abilities instead performs controller\-only Handoff\. App\.[A](https://arxiv.org/html/2609.29015#A1)reports the limited increase in the accuracy gap between assigned and unassigned abilities under the tested prompts\.

## Appendix ECase Studies

We present four representative execution traces showing collaboration across complementary abilities, multi\-hop routing around anisolatedagent, committee review and Takeover after an incorrect output, and correction of a localized reasoning error\. For presentation, we standardize ability terminology and action labels to match Sections[4](https://arxiv.org/html/2609.29015#S4)and[B](https://arxiv.org/html/2609.29015#A2)\. Recorded answers, numerical values, and executable field names are retained; logging metadata not needed to follow the execution is omitted\.

### E\.1Collaboration and Routing

Case 1: Complementary abilities across two agentsTask:object\_counting\. Required abilities:mathematical \+ spatial\. Route:A4→\\rightarrowA1\. Outcome:success; gold answer 15\. Question:*“I have a drum, four pianos, two accordions, a carrot, a stalk of celery, three flutes, a cauliflower, a violin, a clarinet, a trombone, and two trumpets\. How many musical instruments do I have?”*

A4 is assigned the spatial ability but not the mathematical ability\. It contributes the part corresponding to its assigned ability and leaves the remaining work to another agent\.

RESULT:Basedontheprovidedlist,themusicalinstrumentsandtheirquantitiesare:

\-Drum:1

\-Piano:4

\-Accordion:2

\-Flute:3

\-Violin:1

\-Clarinet:1

\-Trombone:1

\-Trumpet:2

Non\-instrumentitems\(carrot,celery,cauliflower\)areexcluded\.

Thearithmeticsummationcanbecompletedbyateammate\.

The spatial component is now covered, while the mathematical component remains\. The controller therefore routes the updated task state to A1, which holds the missing ability and is in the normal routing state\.

Coveredability:spatial

Remainingability:mathematical

NEXT\_AGENT\_ID:1

A1 receives the accepted output and completes the remaining operation\.

DESCRIPTION:UsingthecountsfromAgent4:

1\+4\+2\+3\+1\+1\+1\+2=15\.

RESULT:15

What this case showsThe task is solved through complementary outputs from two agents\. A4 performs the spatial filtering and enumeration, routing carries that result, and A1 uses the remaining mathematical ability in the Finalize step\. The trace illustrates how declared abilities determine both the division of work and the next routing decision\.

Case 2: Multi\-hop routing around an isolated agentTask:number\_theory\. Required abilities:symbolic manipulation \+ enumeration\. Network:288 agents, 705 bidirectional links, average degree 4\.9, with1\.7%1\.7\\%as many edges as the corresponding complete graph\. Outcome:success; gold answer 1575\. Question excerpt:*“A school has between 150 and 200 students… What is the sum of all possible numbers of students enrolled?”*

Agent 61 first derives the candidate values needed for the remainder of the solution\.

Assignedability:symbolicmanipulation

RESULT:151,157,163,169,175,181,187,193,199

The symbolic\-manipulation component is now covered, while enumeration remains\. Agent 12, the nearest holder of enumeration, is isolated, so its ability set is empty for routing\. The updated Bellman–Ford distances lead to Agent 82, the nearestnormalholder of enumeration, three hops away\.

Remainingability:enumeration

Agent61\-\>Agent49\(routingafterContribute\)

Agent49\-\>Agent151\(Handoff\)

Agent151\-\>Agent82\(Handoff\)

The intermediate agents relay the task state without asking an LLM to generate task content\.

Assignedability:enumeration

DESCRIPTION:AddtheninevaluesprovidedbyAgent61:

151\+157\+163\+169\+175\+181\+187\+193\+199=1575\.

RESULT:1575

What this case showsAgent 61 contributes the symbolic\-manipulation component\. Local routing\-state information then prevents the task from being assigned to anisolatedagent\. Bellman–Ford distance updates guide two controller\-only relays to Agent 82, which holds the remaining enumeration ability and uses Finalize to complete the task\. The route uses only local eligibility and distance information\.

### E\.2Review and Takeover

Case 3: Fluent but incorrect output, committee reassessment, and TakeoverTask:geometric\_shapes\. Required abilities:mathematical \+ spatial\. Degraded agent:A5\. Outcome:success; gold answer \(J\), triangle\. Question:SVG path’M 30\.17,45\.97 L 58\.79,40\.36 L 18\.10,15\.70 M 18\.10,15\.70 L 30\.17,45\.97’\.

A5 remains responsive after degradation and produces a well\-formed output\. The failure is therefore not visible from availability or surface form alone\.

RESULT:6vertices\.ThesegmentsconnectintheorderA\-B\-C\-D\-E\-F\-A,

formingaclosedcontour\.Thisdescribesahexagon\.

The path contains only three distinct vertices, so the output contains a concrete correctness error\. The output is escalated to committee review\. Committee members first assess the same output independently\.

ReviewerA:score0\.0

Reason:”Thepathcontainsonlythreedistinctpointsandthreelinesegments,

notsixverticesformingahexagon\.”

ReviewerB:score0\.5

Reason:”Theshapeclassificationisincorrect,butthevertexinterpretation

requiresreassessment\.”

The first\-round judgments disagree\. Each reviewer then receives the other reviewer’s assessment and explanation before independently reassessing the original output\.

First\-roundscores:\[0\.0,0\.5\]

Second\-roundscores:\[0\.0,0\.0\]

Finalreviewscore:0\.0

Decision:Takeover

After reassessment, both reviewers identify the output as incorrect\. The final low score remains attributed to A5 as evidence for routing\-state updates, while Takeover repairs the current task immediately\.

Thepathcontainsthreedistinctvertices:

\(30\.17,45\.97\),\(58\.79,40\.36\),and\(18\.10,15\.70\)\.

Theyareconnectedbythreestraight\-linesegmentsformingaclosed

three\-sidedpolygon:atriangle\.Noadditionalverticesorsegmentsarepresent\.

The corrected output is placed in the task state, and the downstream Finalize step returns \(J\), matching the gold answer\.

What this case showsDegradation need not produce malformed or obviously unusable text: A5 returns a fluent and confident output containing a substantive error\. Independent reviewers initially disagree, but committee reassessment resolves the disagreement\. Takeover then replaces the incorrect output before it propagates to the final answer, while the review score remains available as evidence for subsequent routing\-state updates\.

Case 4: A localized error inside an otherwise coherent solutionTask:tracking\_shuffled\_objects\_five\_objects\. Required abilities:reasoning \+ sequence\. Degraded agent:A5\. Outcome:success; gold answer \(C\), brown\.

In this example, the overall response is coherent, but one state transition is incorrect\. A5 state is watched\. Its output therefore receives mandatory committee review\.

1\.ClaireandAliceswap:

Claire\(red\),Alice\(brown\),Bob\(green\),Dave\(black\),Eve\(orange\)

2\.DaveandEveswap:

Claire\(red\),Alice\(brown\),Bob\(green\),Dave\(orange\),Eve\(black\)

3\.ClaireandEveswap:

Claire\(black\),Alice\(brown\),Bob\(green\),Dave\(orange\),Eve\(red\)

4\.BobandClaireswap:

Claire\(green\),Alice\(brown\),Bob\(black\),Dave\(orange\),Eve\(red\)

5\.AliceandClaireswap:

Claire\(green\),Alice\(black\),Bob\(brown\),Dave\(orange\),Eve\(red\)

Therefore,Clairehasthegreenball\.

The first four state transitions are correct\. The fifth is not: immediately before the final swap, Claire holds green and Alice holds brown, so exchanging their balls should leave Claire with brown and Alice with green\.

Both committee reviewers then identify the incorrect transition\.

ReviewerA:ThefinalswapstatesthatClaireremainsgreen,butreplayingthe

exchangegivesClairebrown\.

ReviewerB:Step5isthefirstincorrectstateupdate;theprecedingfour

transitionsareconsistent\.

The committee’s final score is at mostτto\\tau\_\{\\mathrm\{to\}\}, so Takeover replaces the output\.

AfterthefinalAlice–Claireswap:

ClaireholdsbrownandAliceholdsgreen\.

The corrected output is placed in the task state, and the downstream Finalize step by an agent holding the sequence ability returns \(C\), matching the gold answer\.

What this case showsThe failure is subtle: four consecutive state updates are correct and only the final transition is wrong\. The response otherwise has the structure and fluency of a valid solution\. The review identifies the specific correctness error, and Takeover repairs the current task without restarting the full task trajectory\. This complements Case 3, where committee reassessment resolves disagreement over a more visible semantic error\.

Similar Articles

Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems

Hugging Face Daily Papers

The paper introduces Σ-Mem, an online reliability memory for LLM-based multi-agent systems that tracks historical competence of peers and peer relationships, enabling stable adaptation via spectral bounds and improving coordination through residual steering, routing, and weighted voting.

Multi-Agent LLMs Fail to Explore Each Other

Hugging Face Daily Papers

This paper identifies that current LLM agents fail to systematically explore their peers, leading to poor coordination, and introduces MACE, a lightweight framework using contextual bandits for effective peer selection.

Governed Shared Memory for Multi-Agent LLM Systems

arXiv cs.AI

This paper introduces MemClaw, a governed shared memory architecture for multi-agent LLM systems, formalizing failure modes like unauthorized leakage and stale propagation, and evaluating the system via the ArgusFleet harness.