Conflict-Resilient Multi-Agent Reasoning via Signed Graph Modeling

arXiv cs.AI Papers

Summary

This paper introduces SIGMA, a signed graph-informed multi-agent reasoning framework that explicitly models trust, conflict, and neutral relations among LLM agents to achieve conflict-resilient and globally consistent predictions, outperforming state-of-the-art baselines on six benchmarks.

arXiv:2605.19418v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) have demonstrated strong reasoning and decision-making capabilities that consistently surpass those of single LLM agents. However, their performance often suffers from naive aggregation mechanisms that assume uniformly cooperative interactions. Upon close inspection, we observe that existing graph-based MAS frameworks (1) propagate errors when conflicting signals arise without control, and (2) lack explicit modeling of conflicting inter-agent relations as well as structural awareness, failing to identify reliable interaction patterns. To bridge this gap, we introduce SIGMA, a novel SIgned Graph-informed Multi-Agent reasoning framework that explicitly captures trust, conflict, and neutral relations among agents via a signed relational graph. Specifically, given a query, SIGMA first selects a set of relevant and diverse agents, then constructs a structured signed interaction graph with confidence-weighted edges. Reasoning proceeds through conflict-aware signed message passing, which reinforces information from trustworthy agents while suppressing conflicting signals, and terminates with a structure- and conflict-aware weighted aggregation to yield globally consistent and conflict-resilient predictions. Extensive experiments on six benchmark datasets, across multiple LLM backbones and diverse multi-agent configurations, demonstrate that SIGMA consistently outperforms state-of-the-art baselines, achieving notable gains in both accuracy and conflict-resilient performance.
Original Article
View Cached Full Text

Cached at: 05/20/26, 08:29 AM

# Conflict-Resilient Multi-Agent Reasoning via Signed Graph Modeling
Source: [https://arxiv.org/html/2605.19418](https://arxiv.org/html/2605.19418)
Longgang He1∗Longzhu He2∗Daojing He1†Chaozhuo Li2† 1Harbin Institute of Technology \(Shenzhen\) 2Beijing University of Posts and Telecommunications hedaojing@hit\.edu\.cn lichaozhuo@bupt\.edu\.cn \*Equal contribution†\\daggerCorresponding authors

###### Abstract

LLM\-based multi\-agent systems \(MAS\) have demonstrated strong reasoning and decision\-making capabilities that consistently surpass those of single LLM agents\. However, their performance often suffers from naive aggregation mechanisms that assume uniformly cooperative interactions\. Upon close inspection, we observe that existing graph‑based MAS frameworks ① propagate errors when conflicting signals arise without control, and ② lack explicit modeling of conflicting inter‑agent relations as well as structural awareness, failing to identify reliable interaction patterns\. To bridge this gap, we introduceSIGMA, a novelSIgnedGraph\-informedMulti\-Agent reasoning frameworkthat explicitly capturestrust,conflict, andneutralrelations among agents via a signed relational graph\. Specifically, given a query,SIGMAfirst selects a set of relevant and diverse agents, then constructs a structured signed interaction graph with confidence‑weighted edges\. Reasoning proceeds through conflict‑aware signed message passing, which reinforces information from trustworthy agents while suppressing conflicting signals, and terminates with a structure\- and conflict\-aware weighted aggregation to yield globally consistent and conflict\-resilient predictions\. Extensive experiments on six benchmark datasets, across multiple LLM backbones and diverse multi‑agent configurations, demonstrate thatSIGMAconsistently outperforms state\-of\-the\-art baselines, achieving notable gains in both accuracy and conflict\-resilient performance\.

## 1Introduction

As Large Language Models \(LLMs\) continue to reshape the landscape of artificial intelligence,LLM\-driven agentshave demonstrated remarkable capabilities in reasoningFerraget al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib2)\); Wuet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib8)\); Weiet al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib9)\); Tranet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib10)\), planningBougie and Watanabe \([2025](https://arxiv.org/html/2605.19418#bib.bib11)\); Heet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib12)\); Belleet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib13)\), and decision makingFerraget al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib2)\); Dolant and Kumar \([2025](https://arxiv.org/html/2605.19418#bib.bib14)\); Liet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib15)\), exhibiting increasing autonomy and adaptability across a wide range of applications, including code generationWanget al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib16)\); Jianget al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib17)\), data analysisHonget al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib18)\), and embodied intelligenceWanget al\.\([2023a](https://arxiv.org/html/2605.19418#bib.bib19)\)\. Building upon the success of single\-agent systems, recent studies have shown that LLM\-based Multi\-Agent Systems \(MAS\) can further enhance performance by leveraging collaborative intelligenceDuet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib59)\); Heet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib20)\); Suet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib21)\)\. By orchestrating multiple agents with diverse expertiseHanet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib22)\); Zhanget al\.\([2025a](https://arxiv.org/html/2605.19418#bib.bib23)\), MAS enables more robust and scalable problem\-solving, shifting from isolated reasoning to collaborative intelligence\.

Existing MAS methods can be broadly categorized into three types based on the interaction topology among agents:Chain\-basedHonget al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib63)\); Holtet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib64)\); Qianet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib5)\),Tree\-basedWanget al\.\([2025a](https://arxiv.org/html/2605.19418#bib.bib45)\); Liet al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib46)\); Wuet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib65)\), andGraph\-basedZhugeet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib47)\); Yunet al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib49)\); Zhanget al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib50)\); Fenget al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib4)\)approaches, as illustrated in[Figure˜1](https://arxiv.org/html/2605.19418#S1.F1)\(Left\)\.Chain\-based MASorganize agents sequentially, forming a simple reasoning pipeline, whereasTree\-based MASemploy hierarchical coordination to aggregate outputs from multiple branches\.Graph\-based MASrepresent agents as nodes and interactions as edges, explicitly modeling inter\-agent dependencies and enabling flexible information flowTranet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib10)\); Guoet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib36)\)\. Among these three types,graph\-based MAShas attracted considerable attention due to its capability to capture complex interactions and support more adaptive multi\-agent reasoning\.

However, existing graph\-based multi\-agent system methods still exhibit notable limitations\. First, they cannot explicitly model conflicting relationships between agents\. Most methods rely solely on generic similarity or communication links to construct the graphZhugeet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib47)\); Yunet al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib49)\); Liet al\.\([2021](https://arxiv.org/html/2605.19418#bib.bib48)\), which leads to implicit conflict signals propagating uncontrollably across multiple reasoning steps, thereby amplifying errors\. Second, they lack logical consistency and structural awareness, making it difficult to identify and leverage the latent patterns in agent interactionsZhanget al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib50)\); Fenget al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib4)\), which in turn undermines system robustness and reliability in complex tasks\. Together, these limitations leave existing graph\-based MAS methods particularly vulnerable in scenarios involving noisy or adversarial agents\.

![Refer to caption](https://arxiv.org/html/2605.19418v1/x1.png)Figure 1:\(Left\) Prior MAS treat all agents as equally reliable, includingchain,tree, andgraph\-basedstructures\. \(Right\)SIGMAmodelstrust,conflict, andneutralrelations via signed graph modeling, enabling the system to identify which agents to trust or challenge for conflict\-resilient reasoning\.To address this, we propose aSIgnedGraph\-informedMulti\-Agent Reasoning frameworkfor LLM\-based multi\-agent systems, dubbedSIGMA\. As illustrated in[Figure˜1](https://arxiv.org/html/2605.19418#S1.F1)\(Right\), unlike prior graph\-based MAS approaches,SIGMAcaptures not only the existence but also the polarity of agent interactions via signed graph modelingZaslavsky \([1982](https://arxiv.org/html/2605.19418#bib.bib1)\); Liuet al\.\([2021](https://arxiv.org/html/2605.19418#bib.bib85)\); Zhanget al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib86)\)\. Edges encode not only the existence of interactions but also their polarityDerret al\.\([2018](https://arxiv.org/html/2605.19418#bib.bib84)\)\(e\.g\.,trustorconflict\), allowingSIGMAto characterize multifaceted inter\-agent relationships through three complementary types of relations: ①Trust, capturing supportive interactions; ②Conflict, modeling contradictory interactions; and ③Neutral, representing weak or uncertain interactions\. By explicitly modeling interaction polarity,SIGMAidentifies trustworthy agents and those to challenge, enabling robust and consistent multi\-agent reasoning\.

Although the approach is promising, introducing signed graph modeling is non\-trivial and involves three key modeling challenges\. ①More complex interactions\.Agent interactions are highly dynamic and multifaceted, generating supportive, conflicting, or neutral signals that propagate across multiple reasoning steps\. If not properly handled, these interactions can amplify errors or suppress valuable dissenting information\. ②More vulnerability to noise or conflicts\.The presence of low\-confidence or conflicting agent outputs increases the likelihood of misleading signals affecting the final consensus\. Without careful management, noise can propagate through the network, reducing overall reliability\. ③More challenging information aggregation\.Aggregating positive and negative signals across multi\-hop neighborhoods is challenging\. Improper handling may lead to information collapse, overlooked conflicts, or inconsistent global representations, ultimately compromising multi\-agent reasoning\.

To tackle these challenges,SIGMAfirst performs query\-guided agent selection, ensuring that the chosen agents are both semantically relevant and diverse\. It then constructs a signed heterogeneous relational graph, estimating pairwise agreement signals and annotating each relationship as trust, conflict, or neutral, with corresponding confidence weights\. Reasoning proceeds via a conflict\-aware signed message passing mechanism, which reinforces information from trusted agents while suppressing conflicting signals, thereby mitigating the influence of unreliable agents\. Finally, a structure\- and conflict\-aware weighted aggregation integrates all agent outputs to maximize agreement, minimize inconsistency, and produce a globally consistent, high\-quality prediction\. By explicitly modelingtrust,conflict, andneutral,SIGMAtransforms inter\-agent interactions from uniform cooperation into structured reasoning dynamics, enabling more reliable and robust multi\-agent LLM reasoning\.

Contributions\.We summarize our main contributions as follows: ① We highlight a key limitation in existing multi\-agent systems: assuming uniformly cooperative interactions without modeling trust or conflict, which reduces robustness under contradictory or noisy outputs\. ② We proposeSIGMA, asigned graph\-informed MAS frameworkthat modelstrust,conflict, andneutralrelations, enabling conflict\-resilient multi\-agent reasoning\. ③ Extensive experiments show thatSIGMAconsistently outperforms baselines, achieving robust predictions even with conflicting or low\-quality agents\.

Organization\.The rest of this paper is organized as follows\. In[Section˜2](https://arxiv.org/html/2605.19418#S2), we introduce the notations and preliminaries\. Section[3](https://arxiv.org/html/2605.19418#S3)details our proposed methodSIGMA\. Section[4](https://arxiv.org/html/2605.19418#S4)presents comprehensive experimental results\. In Section[5](https://arxiv.org/html/2605.19418#S5), we discuss related work\. Finally, Section[6](https://arxiv.org/html/2605.19418#S6)concludes the paper\.

## 2Preliminary

In this section, we formalizeLLM\-based multi\-agent systemsand extend the interaction modeling with asigned graph representationto explicitly model both trustworthy and conflicting interactions\.

LLM\-based Multi\-Agent System Formalization\.Consider a multi\-agent system of LLM\-based agentsFerraget al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib2)\); Wuet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib8)\); Duet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib59)\)modeled as a directed interaction graph𝒢=\(𝒱,ℰ\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\), where\|𝒱\|=N\|\\mathcal\{V\}\|=Ndenotes the number of agents andℰ⊆𝒱×𝒱\\mathcal\{E\}\\subseteq\\mathcal\{V\}\\times\\mathcal\{V\}represents communication links\. Let𝑨∈ℝN×N\\boldsymbol\{A\}\\in\\mathbb\{R\}^\{N\\times N\}denote the interaction matrix, where𝑨i​j\\boldsymbol\{A\}\_\{ij\}quantifies the strength of influence from agentvjv\_\{j\}to agentviv\_\{i\}\. Each agentvi∈𝒱v\_\{i\}\\in\\mathcal\{V\}is instantiated by a LLM with distinct roles, prompts, or reasoning strategiesHeet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib20)\); Zhugeet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib47)\)\. We abstract each agent as a computational unit capturing its reasoning and interaction processes:

vi=\(ℳi,ℛi,𝒫i\),v\_\{i\}=\(\\mathcal\{M\}\_\{i\},\\mathcal\{R\}\_\{i\},\\mathcal\{P\}\_\{i\}\),\(1\)whereℳi\\mathcal\{M\}\_\{i\}denotes the underlying LLM,ℛi\\mathcal\{R\}\_\{i\}specifies the role configuration, and𝒫i\\mathcal\{P\}\_\{i\}defines the prompting strategy\. Given a queryQQ, the system evolves overTTinteraction rounds\. At each iterationtt, each agent maintains a latent reasoning state𝒉i\(t\)∈ℝd\\boldsymbol\{h\}\_\{i\}^\{\(t\)\}\\in\\mathbb\{R\}^\{d\}, which captures its intermediate reasoning output\. These states are updated via interactions with neighbors\. The update rule is defined as:

𝒉i\(t\)=fi​\(Q,𝒉i\(t−1\),∑j∈𝒩​\(i\)𝑨i​j​𝒉j\(t−1\)\),\\boldsymbol\{h\}\_\{i\}^\{\(t\)\}=f\_\{i\}\\Bigl\(Q,\\boldsymbol\{h\}\_\{i\}^\{\(t\-1\)\},\\sum\\nolimits\_\{j\\in\\mathcal\{N\}\(i\)\}\\boldsymbol\{A\}\_\{ij\}\\boldsymbol\{h\}\_\{j\}^\{\(t\-1\)\}\\Bigr\),\(2\)wherefi​\(⋅\)f\_\{i\}\(\\cdot\)denotes the reasoning function parameterized by the LLM, and𝒩​\(i\)\\mathcal\{N\}\(i\)denotes the neighborhood of agentviv\_\{i\}\. Let𝑯\(t\)=\[𝒉1\(t\),…,𝒉N\(t\)\]⊤∈ℝN×d\\boldsymbol\{H\}^\{\(t\)\}=\[\\boldsymbol\{h\}\_\{1\}^\{\(t\)\},\\dots,\\boldsymbol\{h\}\_\{N\}^\{\(t\)\}\]^\{\\top\}\\in\\mathbb\{R\}^\{N\\times d\}denote the global matrix\. The system dynamics can be expressed in compact matrix form across all agents over interaction rounds as𝑯\(t\)=ℱ​\(Q,𝑯\(t−1\),𝑨\)\\boldsymbol\{H\}^\{\(t\)\}=\\mathcal\{F\}\\bigl\(Q,\\boldsymbol\{H\}^\{\(t\-1\)\},\\boldsymbol\{A\}\\bigr\), whereℱ​\(⋅\)\\mathcal\{F\}\(\\cdot\)aggregates agent\-wise updates in a unified iterative message\-passing process\. AfterTTrounds, a global aggregation operator𝒜​\(⋅\)\\mathcal\{A\}\(\\cdot\)produces the final prediction:y=𝒜​\(𝑯\(T\)\)y=\\mathcal\{A\}\(\\boldsymbol\{H\}^\{\(T\)\}\)\. Most existing methodsZhugeet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib47)\); Yunet al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib49)\); Zhanget al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib50)\)assume a non\-negative interaction matrix \(𝑨i​j≥0\\boldsymbol\{A\}\_\{ij\}\\geq 0\), implicitly assuming all interactions are cooperative This assumption neglects the presence of conflicting or unreliable signals among agents, which may lead to unbounded error accumulation, degraded reasoning robustness, and unstable consensus formation in complex reasoning scenarios\.

![Refer to caption](https://arxiv.org/html/2605.19418v1/x2.png)Figure 2:Visualization of Balanced Triad types\. First two are balanced, last two are imbalanced\.Signed Graph Representation\.Inspired by Balance TheoryHeider \([1946](https://arxiv.org/html/2605.19418#bib.bib71)\), signed graphs represent relationships by polarity and magnitude\. Specifically, each interaction is given by𝑨i​j=𝒔i​j⋅𝒘i​j\\boldsymbol\{A\}\_\{ij\}=\\boldsymbol\{s\}\_\{ij\}\\cdot\\boldsymbol\{w\}\_\{ij\}, where𝒔i​j∈\{−1,0,\+1\}\\boldsymbol\{s\}\_\{ij\}\\in\\\{\-1,0,\+1\\\}indicates polarity and𝒘i​j≥0\\boldsymbol\{w\}\_\{ij\}\\geq 0its magnitude, as shown in[Definition˜1](https://arxiv.org/html/2605.19418#Thmdefinition1)\.

###### Definition 1\(Balance TheoryHeider \([1946](https://arxiv.org/html/2605.19418#bib.bib71)\)\)\. A triad of nodesvi,vj,vkv\_\{i\},v\_\{j\},v\_\{k\}is balanced if the product of its edge signs∏\(p,q\)∈\{\(i,j\),\(j,k\),\(k,i\)\}𝐬p​q=\+1\\prod\_\{\(p,q\)\\in\\\{\(i,j\),\(j,k\),\(k,i\)\\\}\}\\boldsymbol\{s\}\_\{pq\}=\+1is positive; otherwise, it is imbalanced\. Balanced triads reflect consistent configurations\. For instance, ifviv\_\{i\}has a positive relationship withvjv\_\{j\}and a negative relationship withvkv\_\{k\}, the triad is imbalanced\. To restore balance,viv\_\{i\}may either form a positive relationship withvkv\_\{k\}or a negative relationship withvjv\_\{j\}\.[Figure˜2](https://arxiv.org/html/2605.19418#S2.F2)shows all four triad types\.

Accordingly, the signed matrix admits a decomposition𝑨=𝑨\+−𝑨−\\boldsymbol\{A\}=\\boldsymbol\{A\}^\{\+\}\-\\boldsymbol\{A\}^\{\-\}, where𝑨±=max⁡\{±𝑨,0\}\\boldsymbol\{A\}^\{\\pm\}=\\max\\\{\\pm\\boldsymbol\{A\},0\\\}, with the max operator applied element\-wise\. This separates positive and negative interactions and induces a partition of each agent’s neighborhood into𝒩i±=\{j:±𝑨i​j\>0\}\\mathcal\{N\}\_\{i\}^\{\\pm\}=\\\{\\,j:\\pm\\boldsymbol\{A\}\_\{ij\}\>0\\,\\\}\. Under this signed structure, propagation distinguishes cooperative and conflicting signalsLiet al\.\([2025a](https://arxiv.org/html/2605.19418#bib.bib89)\)\. A signed aggregation is:

𝒉i\(t\)=∑j∈𝒩i\+αi​j​𝒉j\(t−1\)−∑j∈𝒩i−βi​j​𝒉j\(t−1\)\.\\boldsymbol\{h\}\_\{i\}^\{\(t\)\}=\\sum\\nolimits\_\{j\\in\\mathcal\{N\}\_\{i\}^\{\+\}\}\\alpha\_\{ij\}\\boldsymbol\{h\}\_\{j\}^\{\(t\-1\)\}\-\\sum\\nolimits\_\{j\\in\\mathcal\{N\}\_\{i\}^\{\-\}\}\\beta\_\{ij\}\\boldsymbol\{h\}\_\{j\}^\{\(t\-1\)\}\.\(3\)whereαi​j\\alpha\_\{ij\}andβi​j\\beta\_\{ij\}are normalized attention weights induced byA\+A^\{\+\}and𝑨−\\boldsymbol\{A\}^\{\-\}, respectively\. To ensure stable propagation, we further normalize the signed interaction matrix as𝑨~i​j=𝑨i​j/\(∑k\|𝑨i​k\|\+ϵ\)\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}=\{\\boldsymbol\{A\}\_\{ij\}\}/\(\{\\sum\_\{k\}\|\\boldsymbol\{A\}\_\{ik\}\|\+\\epsilon\}\), which controls the scale of incoming interactions and improves numerical stability across layers\. This yields the following compact formulation for signed aggregation under normalized propagation:

𝑯\(t\)=𝑨~\+​ϕ​\(𝑯\(t−1\)\)−𝑨~−​ϕ​\(𝑯\(t−1\)\)\.\\boldsymbol\{H\}^\{\(t\)\}=\\tilde\{\\boldsymbol\{A\}\}^\{\+\}\\phi\(\\boldsymbol\{H\}^\{\(t\-1\)\}\)\-\\tilde\{\\boldsymbol\{A\}\}^\{\-\}\\phi\(\\boldsymbol\{H\}^\{\(t\-1\)\}\)\.\(4\)whereϕ​\(⋅\)\\phi\(\\cdot\)denotes a shared transformation over agent representations, implemented via LLM reasoning or prompt\-driven updates\. Although written in continuous form, it operates through textual interactions, where signed aggregation is realized via prompting\. Positive edges reinforce consistent signals, while negative edges suppress conflicting ones, a critical mechanism absent in existing LLM\-based multi\-agent frameworks and prone to propagating misleading signals\.

![Refer to caption](https://arxiv.org/html/2605.19418v1/x3.png)Figure 3:Overview ofSIGMA, through four stages enabling robust multi\-agent reasoning: \(I\) Query\-Guided Agent Selection, leveraging multi\-dimensional attributes; \(II\) Signed Relational Graph Construction, explicitly modeling heterogeneous inter\-agent relations; \(III\) Conflict\-Aware Signed Message Passing; and \(IV\) Signed Consensus Readout, whereSIGMAintegrates agent representations to yield a globally coherent consensus according to net supportive strength in the signed graph\.
## 3Methodology

This section presentsSIGMA, a signed graph\-informed multi\-agent reasoning framework, as illustrated in[Figure˜3](https://arxiv.org/html/2605.19418#S2.F3)\. Unlike conventional approaches that assume uniform cooperation,SIGMAexplicitly captures multifaceted and potentially conflicting inter\-agent interactions\. Given a query,SIGMAfirst selects a subset of relevant agents \(⊳\\triangleright[Section˜3\.1](https://arxiv.org/html/2605.19418#S3.SS1)\), then constructs a signed relational graph to encode trust and conflict relationships \(⊳\\triangleright[Section˜3\.2](https://arxiv.org/html/2605.19418#S3.SS2)\), and subsequently performs conflict\-aware signed message passing to iteratively refine agent representations and drive consensus \(⊳\\triangleright[Section˜3\.3](https://arxiv.org/html/2605.19418#S3.SS3)\) and finally generates a global prediction \(⊳\\triangleright[Section˜3\.4](https://arxiv.org/html/2605.19418#S3.SS4)\)\. By disentangling cooperative and adversarial signals within a unified propagation framework,SIGMAachieves more stable and robust reasoning under noisy or unreliable interactions\. The detailed algorithm is provided in Appendix[Algorithm˜1](https://arxiv.org/html/2605.19418#algorithm1)\.

### 3\.1Query\-Guided Agent Selection

Effective multi\-agent collaboration begins with selecting a high\-quality subset of agents tailored to the specific query\. Involving all available agents indiscriminately often leads to irrelevant participants, redundant reasoning that promotes groupthink, and low\-reliability outputs that propagate errors\. To address these issues,SIGMAintroduces aquery\-guided agent selectionmechanism that jointly considers three complementary dimensions, which have been shown to be critical for robust collective intelligence: ① semantic relevance, ② structural diversity, and ③ agent confidence\.

Specifically, semantic relevance ensures that the selected agents possess knowledge and capabilities aligned with the queryLewiset al\.\([2020](https://arxiv.org/html/2605.19418#bib.bib72)\); Karpukhinet al\.\([2020](https://arxiv.org/html/2605.19418#bib.bib73)\), thereby minimizing irrelevant computation and reducing noise from the outset; structural diversity promotes complementary perspectives among agents, which has been shown to significantly enhance robustness and mitigate correlated errors in multi\-agent systems and ensemble methodsNguyenet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib39)\); Wanget al\.\([2025a](https://arxiv.org/html/2605.19418#bib.bib45)\); Yunet al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib49)\); and agent confidence provides a reliability prior by favoring agents exhibiting higher self\-consistency or lower uncertainty in their initial reasoningWanget al\.\([2023b](https://arxiv.org/html/2605.19418#bib.bib62)\); Gal and Ghahramani \([2016](https://arxiv.org/html/2605.19418#bib.bib74)\), directly addressing the well\-known hallucination and inconsistency issues in LLMs\.

To balance these objectives efficiently without introducing excessive hyperparameters,SIGMAcombines diversity and confidence into a composite metric with equal weighting\. This 50/50 allocation is principled and well\-supported: in ensemble learning theory, diversity and individual reliability are two largely orthogonal factors \(e\.g\., addressing variance and bias, respectively\) that contribute comparably to the overall performance of collective systemsWoodet al\.\([2023](https://arxiv.org/html/2605.19418#bib.bib75)\); Webb and Zheng \([2004](https://arxiv.org/html/2605.19418#bib.bib76)\)\. When no strong prior knowledge about their relative importance is available, equal weighting has been widely shown to be a robust and effective strategy across both traditional ensembles and modern LLM\-based multi\-agent frameworksNguyenet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib39)\); Jinet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib77)\)\. The resulting unified scoring function, which jointly balances semantic relevance, structural diversity, and agent confidence to provide a principled and reliable measure of agent suitability, is defined as:

𝐬​\(vi\)=λ​𝗋𝖾𝗅​\(vi,Q\)⏟relevance\+1−λ2​\(𝔼j≠i​\[1−sim​\(𝒉i\(0\),𝒉j\(0\)\)\]⏟diversity\+𝖼𝗈𝗇𝖿​\(vi\)⏟confidence\),\\mathbf\{s\}\(v\_\{i\}\)=\\lambda\\,\\underbrace\{\\mathsf\{rel\}\(v\_\{i\},Q\)\}\_\{\\begin\{subarray\}\{c\}\\text\{relevance\}\\end\{subarray\}\}\+\\tfrac\{1\-\\lambda\}\{2\}\\,\\Big\(\\underbrace\{\\mathbb\{E\}\_\{j\\neq i\}\[1\-\\mathrm\{sim\}\(\\boldsymbol\{h\}\_\{i\}^\{\(0\)\},\\boldsymbol\{h\}\_\{j\}^\{\(0\)\}\)\]\}\_\{\\begin\{subarray\}\{c\}\\text\{diversity\}\\end\{subarray\}\}\+\\underbrace\{\\mathsf\{conf\}\(v\_\{i\}\)\}\_\{\\begin\{subarray\}\{c\}\\text\{confidence\}\\end\{subarray\}\}\\Big\),\(5\)where𝗋𝖾𝗅​\(vi,Q\)\\mathsf\{rel\}\(v\_\{i\},Q\)quantifies semantic alignment \(via embedding similarity or LLM\-as\-a\-judge scoring\); the diversity term𝖽𝗂𝗏​\(vi\)=𝔼j≠i​\[1−sim​\(𝒉i\(0\),𝒉j\(0\)\)\]\\mathsf\{div\}\(v\_\{i\}\)=\\mathbb\{E\}\_\{j\\neq i\}\[1\-\\mathrm\{sim\}\(\\boldsymbol\{h\}\_\{i\}^\{\(0\)\},\\boldsymbol\{h\}\_\{j\}^\{\(0\)\}\)\]penalizes redundancy in initial representations𝒉i\(0\)\\boldsymbol\{h\}\_\{i\}^\{\(0\)\}to encourage structurally complementary agents; and𝖼𝗈𝗇𝖿​\(vi\)\\mathsf\{conf\}\(v\_\{i\}\)measures reliability through self\-consistency checks, uncertainty estimation, or historical performance\. The top\-kkagents are then selected as𝒱s=TopKi∈𝒱⁡𝐬​\(vi\)\\mathcal\{V\}\_\{s\}=\\operatorname\{TopK\}\_\{i\\in\\mathcal\{V\}\}\\mathbf\{s\}\(v\_\{i\}\), with\|𝒱s\|=k\|\\mathcal\{V\}\_\{s\}\|=k, ensuring that subsequent reasoning stages operate on a diverse, relevant, and reliable subset of agents\. This unified scoring achieves a principled trade\-off among relevance, diversity, and confidence while maintaining simplicity\. By balancing these objectives,SIGMAconstructs a compact agent pool that is task\-relevant and robust, providing a strong foundation for subsequent signed graph construction and conflict\-aware reasoning\.

### 3\.2Signed Relational Graph Construction

Following query\-guided agent selection,SIGMAconstructs a signed adjacency matrix𝑨∈ℝk×k\\boldsymbol\{A\}\\in\\mathbb\{R\}^\{k\\times k\}that explicitly models heterogeneous inter\-agent interactions\. Unlike conventional graph\-based MAS frameworks that assume homogeneous or implicitly cooperative interactionsZhugeet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib47)\); Liet al\.\([2021](https://arxiv.org/html/2605.19418#bib.bib48)\); Yunet al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib49)\),SIGMAdistinguishes supportive, conflicting, and neutral relations, which is critical for robust reasoning\. The sign encodes interaction polarity and the magnitude reflects confidence strength, disentangling directional agreement from interaction intensity\. To achieve this,SIGMAdefines a unified pairwise evaluation function that simultaneously captures the polarity and strength of each inter\-agent relation in the initial reasoning space, providing a basis for robust reasoning and controlled signal propagation:

si​j=sign⁡\(feval​\(𝒉i\(0\),𝒉j\(0\)\)\),wi​j=\|feval​\(𝒉i\(0\),𝒉j\(0\)\)\|,𝑨i​j=si​j​wi​j\.s\_\{ij\}=\\operatorname\{sign\}\\\!\\bigl\(f\_\{\\text\{eval\}\}\(\\boldsymbol\{h\}\_\{i\}^\{\(0\)\},\\boldsymbol\{h\}\_\{j\}^\{\(0\)\}\)\\bigr\),\\quad w\_\{ij\}=\\bigl\|f\_\{\\text\{eval\}\}\(\\boldsymbol\{h\}\_\{i\}^\{\(0\)\},\\boldsymbol\{h\}\_\{j\}^\{\(0\)\}\)\\bigr\|,\\quad\\boldsymbol\{A\}\_\{ij\}=s\_\{ij\}\\,w\_\{ij\}\.\(6\)wherefeval​\(⋅,⋅\)f\_\{\\text\{eval\}\}\(\\cdot,\\cdot\)measures relational compatibility between agents in the initial reasoning space\. The variablesi​j∈\{−1,\+1\}s\_\{ij\}\\in\\\{\-1,\+1\\\}encodes interaction polarity \(trustvs\.conflict\), whilewi​j≥0w\_\{ij\}\\geq 0captures confidence strength\. This decomposition explicitly separates directional agreement from interaction intensity, enabling structured modeling of supportive and conflicting interactions\.

This formulation induces a natural partition of each agent’s neighborhood into positive and negative sets,𝒩i±=\{j:±Ai​j\>0\},\\mathcal\{N\}\_\{i\}^\{\\pm\}=\\\{\\,j:\\pm A\_\{ij\}\>0\\,\\\},corresponding to supportive and conflicting neighbors\. This partition naturally extends to multi\-hop neighborhoods, capturing higher\-order dependencies in agreement and conflict\. Positive edges reinforce consistent signals, whereas negative edges provide corrective evidence to mitigate error amplification\. To capture higher\-order dependencies, we define balanced and unbalanced multi\-hop neighborhoods that propagate agreement and disagreement recursively\. Specifically, letBi\(1\)=𝒩i\+B\_\{i\}^\{\(1\)\}=\\mathcal\{N\}\_\{i\}^\{\+\}andUi\(1\)=𝒩i−U\_\{i\}^\{\(1\)\}=\\mathcal\{N\}\_\{i\}^\{\-\}\. Forℓ≥1\\ell\\geq 1, we recursively define:

Bi\(ℓ\+1\)=⋃vk∈Bi\(ℓ\)𝒩k\+∪⋃vk∈Ui\(ℓ\)𝒩k−,Ui\(ℓ\+1\)=⋃vk∈Ui\(ℓ\)𝒩k\+∪⋃vk∈Bi\(ℓ\)𝒩k−\.B\_\{i\}^\{\(\\ell\+1\)\}=\\bigcup\\nolimits\_\{v\_\{k\}\\in B\_\{i\}^\{\(\\ell\)\}\}\\mathcal\{N\}\_\{k\}^\{\+\}\\;\\cup\\;\\bigcup\\nolimits\_\{v\_\{k\}\\in U\_\{i\}^\{\(\\ell\)\}\}\\mathcal\{N\}\_\{k\}^\{\-\},\\quad U\_\{i\}^\{\(\\ell\+1\)\}=\\bigcup\\nolimits\_\{v\_\{k\}\\in U\_\{i\}^\{\(\\ell\)\}\}\\mathcal\{N\}\_\{k\}^\{\+\}\\;\\cup\\;\\bigcup\\nolimits\_\{v\_\{k\}\\in B\_\{i\}^\{\(\\ell\)\}\}\\mathcal\{N\}\_\{k\}^\{\-\}\.\(7\)This recursion follows a polarity\-consistent propagation principle, akin to signed graph balance: agreement preserves polarity, while disagreement flips it\. Consequently,Bi\(ℓ\)B\_\{i\}^\{\(\\ell\)\}captures multi\-hop consistent reasoning chains, whereasUi\(ℓ\)U\_\{i\}^\{\(\\ell\)\}encodes higher\-order conflict structures, enablingSIGMAto jointly model local interactions and long\-range agreement and contradiction patterns\.

### 3\.3Conflict\-Aware Signed Message Passing

We distinguish interaction iterationst=1,…,Tt=1,\\dots,Tfrom message\-passing layersℓ=1,…,L\\ell=1,\\dots,L, where each iteration appliesLLlayers of signed propagation\. Building upon the signed relational graph and the recursively defined multi\-hop balanced \(Bi\(ℓ\)B\_\{i\}^\{\(\\ell\)\}\) and unbalanced \(Ui\(ℓ\)U\_\{i\}^\{\(\\ell\)\}\) neighborhoods from the previous stage,SIGMAperforms iterative refinement of agent representations through conflict\-aware signed message passing\. This stage is critical because symmetric aggregation of supportive and conflicting signals would collapse semantically distinct reasoning paths and undermine robustness\.

To explicitly disentangle agreement from disagreement, each agent maintains separate positive and negative representation vectors\. Formally, at iterationttand layerℓ\\ell, the hidden state of each agentviv\_\{i\}is represented as a concatenated pair𝒉i\(t,ℓ\)=\[𝒉ipos​\(t,ℓ\),𝒉ineg​\(t,ℓ\)\]\\boldsymbol\{h\}\_\{i\}^\{\(t,\\ell\)\}=\\bigl\[\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(t,\\ell\)\},\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(t,\\ell\)\}\\bigr\], where𝒉ipos​\(t,ℓ\)\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(t,\\ell\)\}and𝒉ineg​\(t,ℓ\)\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(t,\\ell\)\}capture the agent’s latent semantic state under supportive and conflicting contexts, respectively\. The update rules adopt a two\-part aggregation scheme:

𝒉ipos​\(t,ℓ\)\\displaystyle\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(t,\\ell\)\}=𝒞\(ℓ\)​\(𝒉ipos​\(t,ℓ−1\),𝒜​𝒢​𝒢\(ℓ\)​\(\{𝒉jpos​\(t,ℓ−1\):vj∈Bi\(ℓ\)\},\{𝒉jneg​\(t,ℓ−1\):vj∈Ui\(ℓ\)\}\)\),\\displaystyle=\\mathcal\{C\}^\{\(\\ell\)\}\\\!\\Bigl\(\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(t,\\ell\-1\)\},\\,\\mathcal\{AGG\}^\{\(\\ell\)\}\(\\\{\\boldsymbol\{h\}\_\{j\}^\{\\mathrm\{pos\}\(t,\\ell\-1\)\}:v\_\{j\}\\in B\_\{i\}^\{\(\\ell\)\}\\\},\\\{\\boldsymbol\{h\}\_\{j\}^\{\\mathrm\{neg\}\(t,\\ell\-1\)\}:v\_\{j\}\\in U\_\{i\}^\{\(\\ell\)\}\\\}\)\\Bigr\),\(8\)𝒉ineg​\(t,ℓ\)\\displaystyle\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(t,\\ell\)\}=𝒞\(ℓ\)​\(𝒉ineg​\(t,ℓ−1\),𝒜​𝒢​𝒢\(ℓ\)​\(\{𝒉jneg​\(t,ℓ−1\):vj∈Bi\(ℓ\)\},\{𝒉jpos​\(t,ℓ−1\):vj∈Ui\(ℓ\)\}\)\)\.\\displaystyle=\\mathcal\{C\}^\{\(\\ell\)\}\\\!\\Bigl\(\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(t,\\ell\-1\)\},\\,\\mathcal\{AGG\}^\{\(\\ell\)\}\(\\\{\\boldsymbol\{h\}\_\{j\}^\{\\mathrm\{neg\}\(t,\\ell\-1\)\}:v\_\{j\}\\in B\_\{i\}^\{\(\\ell\)\}\\\},\\\{\\boldsymbol\{h\}\_\{j\}^\{\\mathrm\{pos\}\(t,\\ell\-1\)\}:v\_\{j\}\\in U\_\{i\}^\{\(\\ell\)\}\\\}\)\\Bigr\)\.\(9\)where𝒞\(ℓ\)\\mathcal\{C\}^\{\(\\ell\)\}and𝒜​𝒢​𝒢\(ℓ\)\\mathcal\{AGG\}^\{\(\\ell\)\}denote aggregation operations \(e\.g\., averaging or implicit attention via LLM reasoning\), and the neighborhoodsBi\(ℓ\)B\_\{i\}^\{\(\\ell\)\}andUi\(ℓ\)U\_\{i\}^\{\(\\ell\)\}facilitate the propagation of both local pairwise interactions and higher\-order agreement and disagreement patterns\. Positive representations aggregate supportive signals from balanced neighborhoods while incorporating corrective, polarity\-flipped inputs from unbalanced neighborhoods; the negative representations perform the symmetric, polarity\-reversed operation\. This design leverages conflicting information to refine representations\. In practice, aggregation is implemented via text\-level interaction, where agents condition on others’ responses through structured prompts rather than explicit embedding exchange\.

AfterTTinteraction iterations, each withLLlayers of signed message passing, the final representation of each agent is obtained by fusing the positive and negative components as𝒉i\(T\)=𝒉ipos​\(T,L\)∥𝒉ineg​\(T,L\)\\boldsymbol\{h\}\_\{i\}^\{\(T\)\}=\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(T,L\)\}\\;\\\|\\;\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(T,L\)\}, where∥\\\|denotes the vector concatenate operation\. By explicitly separating and balancing supportive and conflicting signals at both local and multi\-hop scales, this conflict\-aware signed message passing produces robust, context\-aware, and semantically enriched agent representations that serve as high\-quality inputs for the subsequent signed consensus readout stage\.

### 3\.4Signed Consensus Readout

AfterTTiterations of conflict\-aware signed message passing,SIGMAperforms a signed consensus readout to produce a coherent prediction from𝒉i\(T\)\\boldsymbol\{h\}\_\{i\}^\{\(T\)\}\. Specifically, we compute a weighted signed aggregationy=∑i=1k𝒘i​𝒉i\(T\)y=\\sum\_\{i=1\}^\{k\}\\boldsymbol\{w\}\_\{i\}\\,\\boldsymbol\{h\}\_\{i\}^\{\(T\)\}, where the weight for each agent is given by𝒘i=∑j𝑨i​j/\(∑p\|∑q𝑨p​q\|\)\\boldsymbol\{w\}\_\{i\}=\{\\sum\_\{j\}\\boldsymbol\{A\}\_\{ij\}\}/\(\{\\sum\_\{p\}\\bigl\|\\sum\_\{q\}\\boldsymbol\{A\}\_\{pq\}\\bigr\|\}\)\. This structure\-aware weighting captures the net supportive strength of agentviv\_\{i\}within the signed graph, emphasizing agents consistently endorsed by others while down\-weighting those in conflicting neighborhoods\. In practice, this aggregation is implemented via text\-level interaction, where𝒘i\\boldsymbol\{w\}\_\{i\}modulates the influence of each agent’s response rather than directly aggregating continuous embeddings, and the final answer is generated from the aggregated outputs \(e\.g\., via LLM\-based generation or simple post\-processing\), yielding an accurate and conflict\-resilient consensus across the agent set\.

## 4Experiment

We conduct a series of experiments to thoroughly evaluate the effectiveness ofSIGMA, first describing the experimental setup \(⊳\\triangleright[Section˜4\.1](https://arxiv.org/html/2605.19418#S4.SS1)\) and then addressing the following five research questions:

- •RQ1:How doesSIGMAcompare with different single\- and multi\-agent baselines? \(⊳\\triangleright[Section˜4\.2](https://arxiv.org/html/2605.19418#S4.SS2)\)
- •RQ2:What is the contribution of each key component to the performance ofSIGMA? \(⊳\\triangleright[Section˜4\.3](https://arxiv.org/html/2605.19418#S4.SS3)\)
- •RQ3:How doesSIGMAperform under different conflicts, noise, and agent setups? \(⊳\\triangleright[Section˜4\.4](https://arxiv.org/html/2605.19418#S4.SS4)\)
- •RQ4:How sensitive isSIGMAto different hyperparameter settings and performance? \(⊳\\triangleright[Section˜4\.5](https://arxiv.org/html/2605.19418#S4.SS5)\)
- •RQ5:Can case studies clearly demonstrateSIGMA’s modeling of trust and conflict? \(⊳\\triangleright[Section˜4\.6](https://arxiv.org/html/2605.19418#S4.SS6)\)

### 4\.1Experiment Setup

Datasets and Metrics\.To evaluate the performance ofSIGMA, we conduct experiments on six benchmark datasets spanning two categories\. The detailed statistics and characteristics of these datasets are summarized in[Table˜3](https://arxiv.org/html/2605.19418#A4.T3)and described in Appendix[D\.1](https://arxiv.org/html/2605.19418#A4.SS1)\. Specifically, the multi\-domain category includes three general reasoning datasets:MMLUHendryckset al\.\([2021](https://arxiv.org/html/2605.19418#bib.bib78)\),MMLU\-ProWanget al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib79)\), andGPQAReinet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib80)\), covering knowledge across diverse subjects\. The domain\-specific category includes two mathematical reasoning datasets,GSM8KCobbeet al\.\([2021](https://arxiv.org/html/2605.19418#bib.bib81)\)andMultiArithRoy and Roth \([2015](https://arxiv.org/html/2605.19418#bib.bib82)\), along with one code reasoning dataset,HumanEvalChenet al\.\([2021](https://arxiv.org/html/2605.19418#bib.bib83)\)\.

Baselines\.We select a diverse set of baselines for evaluation\. For single\-agent approaches, we select CoTWeiet al\.\([2022](https://arxiv.org/html/2605.19418#bib.bib91)\), ComplexCoTFuet al\.\([2023](https://arxiv.org/html/2605.19418#bib.bib61)\), Self\-Consistency \(SC\)Wanget al\.\([2023b](https://arxiv.org/html/2605.19418#bib.bib62)\), and PHPZhenget al\.\([2023](https://arxiv.org/html/2605.19418#bib.bib92)\)\. For multi\-agent approaches, we select MoAWanget al\.\([2025a](https://arxiv.org/html/2605.19418#bib.bib45)\), Self\-MoALiet al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib46)\), Complete Graph, Random Graph, DyLANLiuet al\.\([2023](https://arxiv.org/html/2605.19418#bib.bib69)\), AutoGenWuet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib65)\), GPTSwarmZhugeet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib47)\), G\-DesignerZhanget al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib50)\), and GoAYunet al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib49)\)\. Detailed information is provided in Appendix[D\.2](https://arxiv.org/html/2605.19418#A4.SS2)\.

Implementation Details\.We access the LLMs via the OpenAI API, and mainly test ongpt\-5\.4\(GPT\-5\) andgpt\-5\.4\-mini\(GPT\-5 mini\) as the backbone models\. We use temperature0\.70\.7for stochastic decoding and0for deterministic baselines\. In query\-guided agent selection, the number of top\-kkagents is set individually for each dataset, withkkroughly ranging from33to 7\. By default,kkis set to44or55for different datasets, whileλ\\lambdais tuned from\{0\.3,0\.4,0\.5,0\.6,0\.7\}\\\{0\.3,0\.4,0\.5,0\.6,0\.7\\\}, usingλ=0\.5\\lambda=0\.5as the default\. The relational evaluatorfeval​\(⋅,⋅\)f\_\{\\rm eval\}\(\\cdot,\\cdot\)in[Equation˜6](https://arxiv.org/html/2605.19418#S3.E6)is implemented via cosine similarity on embeddings produced byall\-MiniLM\-L6\-v2\(𝒟=384\\mathcal\{D\}=384\)\. We performL∈\{1,2,3\}L\\in\\\{1,2,3\\\}layers of conflict\-aware signed message passing and useL=2L=2by default\. A lightweight consensus readout module aggregates the final agent representations into the global prediction\. We provide explicit agent profiling for multi\-agent methods with diverse role configurations, generated bygpt\-5\.4\. For all benchmarks, we useB′∈\{40,80\}B^\{\\prime\}\\in\\\{40,80\\\}queries for hyperparameter validation\. We additionally conduct comparative experiments usingDeepSeek\-V3\.2, reported in Appendix[E](https://arxiv.org/html/2605.19418#A5)\. Although formulated in a continuous space,SIGMAis implemented via text\-based LLM interactions, with signed weights applied during aggregation without explicit embedding\-level message passing\.

Table 1:Performance comparison with single\-agent and multi\-agent baselines across multiple benchmarks\. The best results are inbold, and the runner\-ups areunderlined\. “Mul\.”, “Rel\.”, and “Conf\.” indicate whether the method supports multi\-agent collaboration, models inter\-agent relations, and handles conflicting signals, respectively\. Hollow, half\-filled, and filled circles denote no, partial, and full support in these aspects\.†\\daggernotably indicates papers with over one hundred citations\.MethodMul\.Rel\.Conf\.MMLUMMLU\-ProGPQAGSM8KMultiArithHumanEvalAvg\.Vanilla∘\\circ ∘\\circ ∘\\circ 89\.5884\.1745\.9692\.8097\.4293\.7583\.95CoT\(⊳\\rhdNeurIPS’22†\)∘\\circ ∘\\circ ∘\\circ 89\.72↑0\.1487\.58↑3\.4147\.21↑1\.2594\.47↑1\.6797\.81↑0\.3995\.25↑1\.5085\.34ComplexCoT\(⊳\\rhdICLR’24†\) ∘\\circ ∘\\circ ∘\\circ 89\.81↑0\.2388\.76↑4\.5947\.48↑1\.5294\.74↑1\.9498\.23↑0\.8195\.10↑1\.3585\.69SC\(⊳\\rhdICLR’23†\)∘\\circ ∘\\circ ∘\\circ 89\.87↑0\.2991\.43↑7\.2647\.20↑1\.2494\.76↑1\.9698\.45↑1\.0395\.37↑1\.6286\.18PHP\(⊳\\rhdICML’24†\) ∙\\bullet ∘\\circ ∘\\circ 90\.22↑0\.6491\.32↑7\.1548\.32↑2\.3695\.87↑3\.0798\.50↑1\.0895\.52↑1\.7786\.62MoA\(⊳\\rhdICLR’25†\)∙\\bullet ∘\\circ ∘\\circ 89\.93↑0\.3588\.57↑4\.4052\.54↑6\.5894\.10↑1\.3098\.22↑0\.8095\.42↑1\.6786\.46Self\-MoA\(⊳\\rhdarXiv’25†\) ∙\\bullet ∘\\circ ∘\\circ 90\.45↑0\.8788\.69↑4\.5252\.73↑6\.7794\.34↑1\.5498\.42↑1\.0095\.48↑1\.7386\.69Complete Graph∙\\bullet ∙\\bullet ∘\\circ 90\.60↑1\.0294\.73↑10\.5650\.41↑4\.4594\.21↑1\.4198\.50↑1\.0895\.82↑2\.0787\.38Random Graph∙\\bullet ∙\\bullet ∘\\circ 89\.72↑0\.1493\.42↑9\.2549\.33↑3\.3794\.15↑1\.3598\.43↑1\.0195\.10↑1\.3586\.69DyLAN\(⊳\\rhdarXiv’23†\)∙\\bullet ∙\\bullet 89\.12↓0\.4691\.43↑7\.2649\.72↑3\.7694\.04↑1\.2497\.54↑0\.1096\.13↑2\.3886\.33AutoGen\(⊳\\rhdCOLM’24†\) ∙\\bullet ∙\\bullet 89\.31↓0\.2793\.29↑9\.1250\.45↑4\.4995\.13↑2\.3398\.12↑0\.7096\.01↑2\.2687\.05GPTSwarm\(⊳\\rhdICML’24†\)∙\\bullet ∙\\bullet 90\.87↑1\.2994\.87↑10\.7051\.01↑5\.0595\.31↑2\.5198\.51↑1\.0996\.45↑2\.7087\.84G\-Designer\(⊳\\rhdICML’25†\) ∙\\bullet ∙\\bullet 91\.44↑1\.8695\.31↑11\.1452\.72↑6\.7696\.10↑3\.3098\.72↑1\.3096\.88↑3\.1388\.53GoA\(⊳\\rhdICLR’26\)∙\\bullet ∙\\bullet 91\.73↑2\.1595\.10↑10\.9353\.16↑7\.2095\.25↑2\.4598\.23↑0\.8196\.32↑2\.5788\.30SIGMA∙\\bullet ∙\\bullet ∙\\bullet 92\.53↑2\.9595\.71↑11\.5454\.51↑8\.5596\.23↑3\.4398\.87↑1\.4597\.14↑3\.3989\.17

### 4\.2Main Results \(RQ1\)

In this section, we comprehensively evaluateSIGMAagainst a diverse set of single\-agent and multi\-agent baselines across six benchmark datasets, including three general reasoning datasets \(MMLU,MMLU\-Pro,GPQA\) and three domain\-specific datasets \(GSM8K,MultiArith, andHumanEval\), as reported in[Table˜1](https://arxiv.org/html/2605.19418#S4.T1)\. All experiments are conducted using multi\-agent configurations based ongpt\-5\.4\(GPT\-5\) andgpt\-5\.4\-mini\(GPT\-5 mini\) for a fair comparison\.

Overall,SIGMAconsistently achieves the best performance across all six datasets, attaining an average accuracy of89\.1789\.17%\. Multi\-agent methods generally outperform single\-agent baselines such as CoT, ComplexCoT, and Self\-Consistency, particularly on challenging benchmarks such asMMLU\-ProandGPQA, highlighting the benefits of collaborative reasoning\. However, existing approaches remain fundamentally limited by their reliance on implicitly cooperative interactions\. Although graph\-based methods model inter\-agent dependencies and capture structured interactions, they do not explicitly account for conflicting signals, which makes them vulnerable to error propagation under noisy or inconsistent agent outputs\. In contrast,SIGMAaddresses this limitation by explicitly modeling both supportive and conflicting interactions, enabling more reliable and robust multi\-agent reasoning\.

By contrast,SIGMAachieves state\-of\-the\-art performance across all benchmarks, attaining92\.5392\.53% onMMLU,95\.7195\.71% onMMLU\-Pro,54\.5154\.51% onGPQA,96\.2396\.23% onGSM8K,98\.8798\.87% onMultiArith, and97\.1497\.14% onHumanEval\. These consistent gains highlight the effectiveness of explicitly modeling both supportive and conflicting interactions\. The improvements stem from three key factors: ① signed relational modeling distinguishing trust and conflict, ② conflict\-aware message passing suppressing misleading signals while preserving disagreement, and ③ structure\-aware aggregation prioritizing consistent agents\. Together, these enable reliable multi\-agent reasoning under noisy conditions\.

![Refer to caption](https://arxiv.org/html/2605.19418v1/x4.png)
![Refer to caption](https://arxiv.org/html/2605.19418v1/x5.png)
![Refer to caption](https://arxiv.org/html/2605.19418v1/x6.png)
![Refer to caption](https://arxiv.org/html/2605.19418v1/x7.png)

Figure 4:Panels \(a\) and \(b\) present the ablation study, illustrating the contribution of each module inSIGMAon theMMLU\-ProandHumanEval\. Panels \(c\) and \(d\) show the robustness analysis ofSIGMA\.![Refer to caption](https://arxiv.org/html/2605.19418v1/x8.png)
![Refer to caption](https://arxiv.org/html/2605.19418v1/x9.png)
![Refer to caption](https://arxiv.org/html/2605.19418v1/x10.png)
![Refer to caption](https://arxiv.org/html/2605.19418v1/x11.png)

Figure 5:Hyperparameter sensitivity\. Panels \(a\) and \(b\) show the effect ofλ\\lambdaonMMLUandHumanEval, respectively\. Panels \(c\) and \(d\) show the effect ofkkonMMLUandHumanEvaldatasets, respectively\.
### 4\.3Ablation Study \(RQ2\)

To further evaluate the contribution of each component in theSIGMAframework, we conducted a single\-module ablation study using theDeepSeek\-V3\.2model on theMMLU\-ProandHumanEvaldatasets, in which key components were systematically ablated to assess their impact on performance\. Specifically, ①w/oQAS, which replacesquery\-guided agent selectionwith random selection to examine the effect of relevance, structural diversity, and reliability; ②w/oSRG, which replaces thesigned relational graphwith a non\-negative adjacency matrix, ignoring trust, conflict, and neutral polarities, to evaluate the role of heterogeneous interaction modeling; ③w/oCMP, which disablesconflict\-aware message passingand aggregates positive and negative signals uniformly to analyze error propagation mitigation; and ④w/oSCR, which removesstructure\-aware weightingin the final aggregation, using simple averaging to assess its contribution to globally consistent decision making\.

As shown in[Figure˜4](https://arxiv.org/html/2605.19418#S4.F4), \(a\) and \(b\) summarizes the ablation results on theMMLU\-ProandHumanEvaldatasets\. ForMMLU\-Prodataset, removingQASmoderately reduces performance, SRG slightly lowers accuracy,CMPcauses a pronounced drop, andSCRresults in a moderate decline, highlighting the contribution of each module to error mitigation and output aggregation\. ForHumanEvaldataset, performance is more sensitive: removingQASorSRGsubstantially degrades accuracy, disablingCMPleads to the largest drop, and removingSCRsignificantly reduces correctness, confirming that all modules are crucial and that their relative importance varies with task characteristics\.

### 4\.4Robustness Analysis \(RQ3\)

We evaluateSIGMA’s robustness by single\-type injection of the four malicious agents defined in Appendix[D\.3](https://arxiv.org/html/2605.19418#A4.SS3): Random\-Noise \(RandomNoiseAgent\), Low\-Quality Reasoning \(LowQualityAgent\), Conflict\-Inducing \(ConflictAgent\), and Blind\-Conformity \(CopycatAgent\)\. With 8 agents in total, the malicious ratio increases from0% to5050%\. On bothMMLUandHumanEval,SIGMAexhibits only minor performance degradation across all attack types\. OnMMLU, the maximum drop remains below4\.14\.1points, and under the most subtle Blind\-Conformity attack, accuracy still reaches89\.6589\.65%, slightly higher than88\.5588\.55% under Conflict\-Inducing, indicating stronger suppression of implicit herding behaviors\. OnHumanEval, even at5050% malicious ratio, the worst\-case Conflict\-Inducing attack still achieves91\.0091\.00%, while Blind\-Conformity reaches92\.5092\.50%, further demonstrating robustness in generative settings\. As the malicious ratio increases, negative edges toward malicious agents rise from<5%<5\\%to\>65%\>65\\%, and their weights approach zero or become negative, suppressing harmful outputs\. This aligns with Appendix[C](https://arxiv.org/html/2605.19418#A3), showing trust\-conflict modeling isolates untrustworthy agents\.

### 4\.5Hyperparameter Sensitivity \(RQ4\)

We analyze the sensitivity ofSIGMAto two key hyperparameters inquery\-guided agent selection: the weightλ\\lambdabalancing semantic relevance and diversity/confidence, and the number of top\-kkselected agentskk\. As shown in Figure[5](https://arxiv.org/html/2605.19418#S4.F5), \(a\) and \(b\) illustrate the effect of varyingλ\\lambdaon theMMLUandHumanEvaldatasets\. OnMMLU, accuracy peaks atλ=0\.6\\lambda=0\.6, indicating that a balanced emphasis on semantic relevance and diversity/confidence yields the best performance, while values that are too low or too high lead to a slight decrease in accuracy\. OnHumanEval, accuracy reaches its maximum atλ=0\.4\\lambda=0\.4, reflecting the higher importance of diversity and confidence in code generation tasks\. \(c\) and \(d\) show the effect of top\-kk\. Accuracy gradually increases askkgrows from33to55, and stabilizes afterk=4k=4, suggesting that a moderate number of selected agents is sufficient for robust performance across both datasets, further validating the effectiveness ofSIGMA\.

### 4\.6Case Study \(RQ5\)

![Refer to caption](https://arxiv.org/html/2605.19418v1/x12.png)Figure 6:Case study\.Figure[6](https://arxiv.org/html/2605.19418#S4.F6)provides qualitative insights into howSIGMAmodels trust and conflict among agents\. In theMMLU\-ProPsychology classification task, four agents correctly classify Pascale as a developmental psychologist, while a fifth gives an ambiguous output;SIGMAconstructs a signed graph with positive edges linking consistent agents and negative edges connecting the conflicting agent, amplifying supportive signals and suppressing misleading contributions\. In theMultiArithCarol carrot\-picking problem, five agents estimate the number of bad carrots; the signed graph emphasizes consistent outputs and attenuates minor conflicts, thus producing the correct consensus of77\. The heatmap shows net support weights and the network diagram shows agent connections to the aggregation node, demonstrating robust, interpretable reasoning\.

## 5Related Works

LLM\-based Multi\-Agent Reasoning\.Recent studies leverage multiple LLMs as agents to perform collaborative reasoning and decision\-making at test time\. Early approaches adopt weak interaction paradigms, such as LLM\-debateDuet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib59)\); Chanet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib60)\)and Self\-ConsistencyWanget al\.\([2023b](https://arxiv.org/html/2605.19418#bib.bib62)\), where agents generate diverse responses with limited coordination\. Subsequent frameworks introduce more structured communication, including chain\-based pipelines \(e\.g\., MetaGPTHonget al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib63)\)\), star\-based coordination \(e\.g\., AutoGenWuet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib65)\)and MiniGridZhouet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib67)\)\), and hierarchical or tree\-based organizationsIshibashi and Nishimura \([2024](https://arxiv.org/html/2605.19418#bib.bib68)\)\. More recent works further model agent interactions via graph structures, such as GPTSwarmZhugeet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib47)\), DyLANLiuet al\.\([2023](https://arxiv.org/html/2605.19418#bib.bib69)\), MacNetQianet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib70)\), G\-DesignerZhanget al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib50)\)and MasRouterYueet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib51)\), enabling flexible multi\-agent communication\. Despite these advances, most existing methods rely on predefined or input\-independent topologies and assume homogeneous cooperation, overlooking conflicting, unreliable, or adversarial reasoning among agents\. This highlights a fundamental limitation: current frameworks lack an inductive bias to model heterogeneous interactions, including both agreement and disagreement\.

Graph for Multi\-Agent Systems\.Graphs, as a fundamental data structure for representing relationships among multiple agents, have been widely adopted in the pre\-LLM era, particularly in multi\-agent reinforcement learning \(MARL\) for facilitating communication and coordinationWenet al\.\([2022](https://arxiv.org/html/2605.19418#bib.bib42)\); Vinyalset al\.\([2019](https://arxiv.org/html/2605.19418#bib.bib43)\)\. With the proliferation of LLM\-based agentsFerraget al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib2)\); Weiet al\.\([2026](https://arxiv.org/html/2605.19418#bib.bib9)\), researchers have similarly recognized that interactions among multiple LLMs can naturally be modeled from a graph\-based perspectiveZhugeet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib47)\); Zhanget al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib50)\); Yueet al\.\([2025](https://arxiv.org/html/2605.19418#bib.bib51)\)\. Early approaches are often implicit, such as ChatEvalChanet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib60)\), and AutoGenWuet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib65)\), while more recent frameworks, including DyLANLiuet al\.\([2023](https://arxiv.org/html/2605.19418#bib.bib69)\), GPTSwarmZhugeet al\.\([2024](https://arxiv.org/html/2605.19418#bib.bib47)\), and G\-DesignerZhanget al\.\([2025b](https://arxiv.org/html/2605.19418#bib.bib50)\), explicitly represent multiple agents as nodes in a graph to capture their interactions\. However, these methods generally rely on predefined or input\-independent topologies, failing to adaptively model heterogeneous interactions, including trust and conflict among agents\. Consequently, they lack the inductive bias necessary for conflict\-aware multi\-agent reasoning, motivating the introduction of signed graph modeling in our approach\. Further discussions on related works can be found in Appendix[F](https://arxiv.org/html/2605.19418#A6)\.

## 6Conclusion

In this paper, we analyze LLM\-based multi\-agent systems \(MAS\) and show that the lack of explicit modeling of diverse inter\-agent interactions, particularly the relations oftrust,conflict, andneutral, fundamentally limits system robustness, especially in the presence of conflicting or noisy agent outputs\. To address this gap, we proposeSIGMA, asigned graph\-informed multi\-agent reasoning frameworkthat models agent interactions through a structured signed relational graph and employs conflict\-aware message passing to ensure prediction consistency\. Extensive experiments demonstrate thatSIGMAconsistently outperforms state\-of\-the\-art single\- and multi\-agent baselines, achieving higher accuracy and robust, conflict\-resilient performance across diverse reasoning benchmarks\.

## References

- Minimization of functions having lipschitz continuous first partial derivatives\.Pacific Journal of Mathematics \(PJM\)16\(1\),pp\. 1–3\.Cited by:[§C\.2](https://arxiv.org/html/2605.19418#A3.SS2.p1.pic1.3.3.3.3.3.3.3.3.3.3.3.3.3.3.3.3.3.3)\.
- N\. Belle, D\. Barnes, A\. Amayuelas, I\. Bercovich, X\. E\. Wang, and W\. Wang \(2025\)Agents of change: self\-evolving llm agents for strategic planning\.arXiv preprint arXiv:2506\.04651\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- N\. Bougie and N\. Watanabe \(2025\)Citysim: modeling urban behaviors and city dynamics with large\-scale llm\-driven agent simulation\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track \(EMNLP\),pp\. 215–229\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- G\. Box \(1988\)Signal\-to\-noise ratios, performance criteria, and transformations\.Technometrics30\(1\),pp\. 1–17\.Cited by:[§C\.1](https://arxiv.org/html/2605.19418#A3.SS1.p1.pic1.4.4.4.4.4.4.4.4.4.4.4.4.4.4.4.4.4.4)\.
- C\. Chan, W\. Chen, Y\. Su, J\. Yu, W\. Xue, S\. Zhang, J\. Fu, and Z\. Liu \(2024\)ChatEval: towards better llm\-based evaluators through multi\-agent debate\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§5](https://arxiv.org/html/2605.19418#S5.p1.1),[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.
- M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. D\. O\. Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman,et al\.\(2021\)Evaluating large language models trained on code\.arXiv preprint arXiv:2107\.03374\.Cited by:[5th item](https://arxiv.org/html/2605.19418#A4.I1.i5.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p1.1)\.
- Y\. Choi, T\. Ko, J\. Choi, and C\. Kim \(2025\)Beyond binary: improving signed message passing in graph neural networks for multi\-class graphs\.IEEE Transactions on Pattern Analysis and Machine Intelligence \(TPAMI\)\.Cited by:[§F\.1](https://arxiv.org/html/2605.19418#A6.SS1.p1.1)\.
- K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano,et al\.\(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[4th item](https://arxiv.org/html/2605.19418#A4.I1.i4.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p1.1)\.
- T\. Derr, Y\. Ma, and J\. Tang \(2018\)Signed graph convolutional networks\.InIEEE International Conference on Data Mining \(ICDM\),pp\. 929–934\.Cited by:[§F\.1](https://arxiv.org/html/2605.19418#A6.SS1.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p4.1)\.
- A\. Dolant and P\. Kumar \(2025\)Agentic llm framework for adaptive decision discourse\.arXiv preprint arXiv:2502\.10978\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. Mordatch \(2024\)Improving factuality and reasoning in language models through multiagent debate\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1),[§2](https://arxiv.org/html/2605.19418#S2.p2.8),[§5](https://arxiv.org/html/2605.19418#S5.p1.1)\.
- T\. Feng, H\. Zhang, Z\. Lei, P\. Han, and J\. You \(2026\)GraphPlanner: graph memory\-augmented agentic routing for multi\-agent llms\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p2.1),[§1](https://arxiv.org/html/2605.19418#S1.p3.1)\.
- M\. A\. Ferrag, N\. Tihanyi, and M\. Debbah \(2025\)From llm reasoning to autonomous ai agents: a comprehensive review\.arXiv preprint arXiv:2504\.19678\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1),[§2](https://arxiv.org/html/2605.19418#S2.p2.8),[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.
- Y\. Fu, H\. Peng, A\. Sabharwal, P\. Clark, and T\. Khot \(2023\)Complexity\-based prompting for multi\-step reasoning\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[2nd item](https://arxiv.org/html/2605.19418#A4.I2.i2.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1)\.
- Y\. Gal and Z\. Ghahramani \(2016\)Dropout as a bayesian approximation: representing model uncertainty in deep learning\.InInternational Conference on Machine Learning \(ICML\),pp\. 1050–1059\.Cited by:[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p2.1)\.
- T\. Guo, X\. Chen, Y\. Wang, R\. Chang, S\. Pei, N\. V\. Chawla, O\. Wiest, and X\. Zhang \(2024\)Large language model based multi\-agents: a survey of progress and challenges\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence \(IJCAI\),pp\. 8048–8057\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p2.1)\.
- S\. Han, Q\. Zhang, W\. Jin, and Z\. Xu \(2024\)LLM multi\-agent systems: challenges and open problems\.arXiv preprint arXiv:2402\.03578\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- J\. He, S\. Chen, F\. Zhang, and Z\. Yang \(2024\)From words to actions: unveiling the theoretical underpinnings of llm\-driven autonomous systems\.InInternational Conference on Machine Learning \(ICML\),pp\. 17807–17841\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- J\. He, C\. Treude, and D\. Lo \(2025\)Llm\-based multi\-agent systems for software engineering: literature review, vision, and the road ahead\.ACM Transactions on Software Engineering and Methodology \(ACM TOSEM\)34\(5\),pp\. 1–30\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1),[§2](https://arxiv.org/html/2605.19418#S2.p2.8)\.
- F\. Heider \(1946\)Attitudes and cognitive organization\.The Journal of Psychology21\(1\),pp\. 107–112\.Cited by:[§2](https://arxiv.org/html/2605.19418#S2.p3.3),[Definition 1](https://arxiv.org/html/2605.19418#Thmdefinition1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[1st item](https://arxiv.org/html/2605.19418#A4.I1.i1.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p1.1)\.
- S\. Holt, M\. R\. Luyten, and M\. van der Schaar \(2024\)L2MAC: large language model automatic computer for extensive code generation\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p2.1)\.
- S\. Hong, Y\. Lin, B\. Liu, B\. Liu, B\. Wu, C\. Zhang, D\. Li, J\. Chen, J\. Zhang, J\. Wang,et al\.\(2025\)Data interpreter: an llm agent for data science\.InFindings of the Association for Computational Linguistics \(ACL Findings\),pp\. 19796–19821\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- S\. Hong, M\. Zhuge, J\. Chen, X\. Zheng, Y\. Cheng, J\. Wang, C\. Zhang, Z\. Wang, S\. K\. S\. Yau, Z\. Lin,et al\.\(2024\)MetaGPT: meta programming for a multi\-agent collaborative framework\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p2.1),[§5](https://arxiv.org/html/2605.19418#S5.p1.1)\.
- Y\. Ishibashi and Y\. Nishimura \(2024\)Self\-organized agents: a llm multi\-agent framework toward ultra large\-scale code generation and optimization\.arXiv preprint arXiv:2404\.02183\.Cited by:[§F\.2](https://arxiv.org/html/2605.19418#A6.SS2.p1.1),[§5](https://arxiv.org/html/2605.19418#S5.p1.1)\.
- J\. Jiang, F\. Wang, J\. Shen, S\. Kim, and S\. Kim \(2026\)A survey on large language models for code generation\.ACM Transactions on Software Engineering and Methodology \(ACM TOSEM\)35\(2\),pp\. 1–72\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- W\. Jin, H\. Du, B\. Zhao, X\. Tian, B\. Shi, and G\. Yang \(2025\)A comprehensive survey on multi\-agent cooperative decision\-making: scenarios, approaches, challenges and perspectives\.arXiv preprint arXiv:2503\.13415\.Cited by:[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p3.8)\.
- V\. Karpukhin, B\. Oguz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. Yih \(2020\)Dense passage retrieval for open\-domain question answering\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 6769–6781\.Cited by:[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p2.1)\.
- P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel,et al\.\(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.InProceedings of the 34th International Conference on Neural Information Processing Systems \(NeurIPS\),pp\. 9459–9474\.Cited by:[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p2.1)\.
- M\. Li, S\. Zhao, Q\. Wang, K\. Wang, Y\. Zhou, S\. Srivastava, C\. Gokmen, T\. Lee, L\. E\. Li, R\. Zhang,et al\.\(2024\)Embodied agent interface: benchmarking llms for embodied decision making\.InProceedings of the 38th International Conference on Neural Information Processing Systems \(NeurIPS\),pp\. 100428–100534\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- R\. Li, L\. Xu, S\. Liu, J\. Ji, L\. Li, Q\. Lin, and L\. Ma \(2025a\)Structure balance and gradient matching\-based signed graph condensation\.InProceedings of the AAAI Conference on Artificial Intelligence \(AAAI\),Vol\.39,pp\. 12121–12129\.Cited by:[§F\.1](https://arxiv.org/html/2605.19418#A6.SS1.p1.1),[§2](https://arxiv.org/html/2605.19418#S2.p5.3)\.
- W\. Li, Y\. Lin, M\. Xia, and C\. Jin \(2025b\)Rethinking mixture\-of\-agents: is mixing different large language models beneficial?\.arXiv preprint arXiv:2502\.00674\.Cited by:[6th item](https://arxiv.org/html/2605.19418#A4.I2.i6.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p2.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1)\.
- Y\. Li, S\. Ren, P\. Wu, S\. Chen, C\. Feng, and W\. Zhang \(2021\)Learning distilled collaboration graph for multi\-agent perception\.InProceedings of the 35th International Conference on Neural Information Processing Systems \(NeurIPS\),pp\. 29541–29552\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p3.1),[§3\.2](https://arxiv.org/html/2605.19418#S3.SS2.p1.1)\.
- W\. Lin and B\. Li \(2022\)Status\-aware signed heterogeneous network embedding with graph neural networks\.IEEE Transactions on Neural Networks and Learning Systems \(TNNLS\)35\(4\),pp\. 4580–4592\.Cited by:[§F\.1](https://arxiv.org/html/2605.19418#A6.SS1.p1.1)\.
- H\. Liu, Z\. Zhang, P\. Cui, Y\. Zhang, Q\. Cui, J\. Liu, and W\. Zhu \(2021\)Signed graph neural network with latent groups\.InProceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining \(SIGKDD\),pp\. 1066–1075\.Cited by:[§F\.1](https://arxiv.org/html/2605.19418#A6.SS1.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p4.1)\.
- Z\. Liu, Y\. Zhang, P\. Li, Y\. Liu, and D\. Yang \(2023\)Dynamic llm\-agent network: an llm\-agent collaboration framework with agent team optimization\.arXiv preprint arXiv:2310\.02170\.Cited by:[8th item](https://arxiv.org/html/2605.19418#A4.I2.i8.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1),[§5](https://arxiv.org/html/2605.19418#S5.p1.1),[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.
- D\. Nguyen, H\. Le, K\. Do, S\. Gupta, S\. Venkatesh, and T\. Tran \(2024\)Diversifying training pool predictability for zero\-shot coordination: a theory of mind approach\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence \(IJCAI\),pp\. 166–174\.Cited by:[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p2.1),[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p3.8)\.
- C\. Qian, W\. Liu, H\. Liu, N\. Chen, Y\. Dang, J\. Li, C\. Yang, W\. Chen, Y\. Su, X\. Cong,et al\.\(2024\)Chatdev: communicative agents for software development\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(volume 1: Long papers\) \(ACL\),pp\. 15174–15186\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p2.1)\.
- C\. Qian, Z\. Xie, Y\. Wang, W\. Liu, K\. Zhu, H\. Xia, Y\. Dang, Z\. Du, W\. Chen, C\. Yang,et al\.\(2025\)Scaling large language model\-based multi\-agent collaboration\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§5](https://arxiv.org/html/2605.19418#S5.p1.1)\.
- D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman \(2024\)Gpqa: a graduate\-level google\-proof q&a benchmark\.InFirst Conference on Language Modeling \(COLM\),Cited by:[3rd item](https://arxiv.org/html/2605.19418#A4.I1.i3.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p1.1)\.
- S\. Roy and D\. Roth \(2015\)Solving general arithmetic word problems\.InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 1743–1752\.Cited by:[4th item](https://arxiv.org/html/2605.19418#A4.I1.i4.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p1.1)\.
- H\. Su, R\. Chen, S\. Tang, Z\. Yin, X\. Zheng, J\. Li, B\. Qi, Q\. Wu, H\. Li, W\. Ouyang,et al\.\(2025\)Many heads are better than one: improved scientific idea generation by a llm\-based multi\-agent system\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\) \(ACL\),pp\. 28201–28240\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- R\. Sun, Y\. Wu, X\. Wang, C\. Chen, W\. Zhang, and X\. Lin \(2023\)Clique identification in signed graphs: a balance theory based model\.IEEE Transactions on Knowledge and Data Engineering \(TKDE\)35\(12\),pp\. 12513–12527\.Cited by:[§F\.1](https://arxiv.org/html/2605.19418#A6.SS1.p1.1)\.
- K\. Tran, D\. Dao, M\. Nguyen, Q\. Pham, B\. O’Sullivan, and H\. D\. Nguyen \(2025\)Multi\-agent collaboration mechanisms: a survey of llms\.arXiv preprint arXiv:2501\.06322\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p2.1)\.
- O\. Vinyals, I\. Babuschkin, W\. M\. Czarnecki, M\. Mathieu, A\. Dudzik, J\. Chung, D\. H\. Choi, R\. Powell, T\. Ewalds, P\. Georgiev,et al\.\(2019\)Grandmaster level in starcraft ii using multi\-agent reinforcement learning\.Nature575\(7782\),pp\. 350–354\.Cited by:[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.
- G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. Anandkumar \(2023a\)Voyager: an open\-ended embodied agent with large language models\.arXiv preprint arXiv:2305\.16291\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- J\. Wang, W\. Jue, B\. Athiwaratkun, C\. Zhang, and J\. Zou \(2025a\)Mixture\-of\-agents enhances large language model capabilities\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[5th item](https://arxiv.org/html/2605.19418#A4.I2.i5.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p2.1),[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p2.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1)\.
- X\. Wang, L\. Dong, S\. Rangasrinivasan, I\. Nwogu, S\. Setlur, and V\. Govindaraju \(2025b\)AutoMisty: a multi\-agent llm framework for automated code generation in the misty social robot\.In2025 IEEE/RSJ International Conference on Intelligent Robots and Systems \(IROS\),pp\. 9194–9201\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- X\. Wang, J\. Wei, D\. Schuurmans, Q\. V\. Le, E\. H\. Chi, S\. Narang, A\. Chowdhery, and D\. Zhou \(2023b\)Self\-consistency improves chain of thought reasoning in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[3rd item](https://arxiv.org/html/2605.19418#A4.I2.i3.p1.1),[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p2.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1),[§5](https://arxiv.org/html/2605.19418#S5.p1.1)\.
- Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang,et al\.\(2024\)MMLU\-pro: a more robust and challenging multi\-task language understanding benchmark\.InProceedings of the 38th International Conference on Neural Information Processing Systems \(NeurIPS\),pp\. 95266–95290\.Cited by:[2nd item](https://arxiv.org/html/2605.19418#A4.I1.i2.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p1.1)\.
- G\. I\. Webb and Z\. Zheng \(2004\)Multistrategy ensemble learning: reducing error by combining ensemble learning techniques\.IEEE Transactions on Knowledge and Data Engineering \(TKDE\)16\(8\),pp\. 980–991\.Cited by:[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p3.8)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. H\. Chi, Q\. V\. Le, and D\. Zhou \(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.InProceedings of the 36th International Conference on Neural Information Processing Systems \(NeurIPS\),pp\. 24824–24837\.Cited by:[1st item](https://arxiv.org/html/2605.19418#A4.I2.i1.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1)\.
- X\. Wei, Y\. Dong, X\. Wang, X\. Zhang, Z\. Zhao, D\. Shen, L\. Xia, and D\. Yin \(2026\)Beyond react: a planner\-centric framework for complex tool\-augmented llm reasoning\.InProceedings of the AAAI Conference on Artificial Intelligence \(AAAI\),Vol\.40,pp\. 33845–33853\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1),[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.
- M\. Wen, J\. G\. Kuba, R\. Lin, W\. Zhang, Y\. Wen, J\. Wang, and Y\. Yang \(2022\)Multi\-agent reinforcement learning is a sequence modeling problem\.InProceedings of the 36th International Conference on Neural Information Processing Systems \(NeurIPS\),pp\. 16509–16521\.Cited by:[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.
- D\. Wood, T\. Mu, A\. M\. Webb, H\. W\. Reeve, M\. Lujan, and G\. Brown \(2023\)A unified theory of diversity in ensemble learning\.Journal of Machine Learning Research \(JMLR\)24\(359\),pp\. 1–49\.Cited by:[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p3.8)\.
- J\. Wu, J\. Zhu, Y\. Liu, M\. Xu, and Y\. Jin \(2025\)Agentic reasoning: a streamlined framework for enhancing llm reasoning with agentic tools\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\) \(ACL\),pp\. 28489–28503\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1),[§2](https://arxiv.org/html/2605.19418#S2.p2.8)\.
- Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu,et al\.\(2024\)Autogen: enabling next\-gen llm applications via multi\-agent conversations\.InFirst Conference on Language Modeling \(COLM\),Cited by:[7th item](https://arxiv.org/html/2605.19418#A4.I2.i7.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p2.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1),[§5](https://arxiv.org/html/2605.19418#S5.p1.1),[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.
- Y\. Yue, G\. Zhang, B\. Liu, G\. Wan, K\. Wang, D\. Cheng, and Y\. Qi \(2025\)Masrouter: learning to route llms for multi\-agent systems\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\) \(ACL\),pp\. 15549–15572\.Cited by:[§F\.2](https://arxiv.org/html/2605.19418#A6.SS2.p1.1),[§5](https://arxiv.org/html/2605.19418#S5.p1.1),[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.
- S\. Yun, J\. Peng, P\. Li, W\. Fan, J\. Chen, J\. Zou, G\. Li, and T\. Chen \(2026\)Graph\-of\-agents: a graph\-based framework for multi\-agent llm collaboration\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[11st item](https://arxiv.org/html/2605.19418#A4.I2.i11.p1.1),[§F\.2](https://arxiv.org/html/2605.19418#A6.SS2.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p2.1),[§1](https://arxiv.org/html/2605.19418#S1.p3.1),[§2](https://arxiv.org/html/2605.19418#S2.p2.25),[§3\.1](https://arxiv.org/html/2605.19418#S3.SS1.p2.1),[§3\.2](https://arxiv.org/html/2605.19418#S3.SS2.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1)\.
- T\. Zaslavsky \(1982\)Signed graphs\.Discrete Applied Mathematics \(DAM\)4\(1\),pp\. 47–74\.Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p4.1)\.
- G\. Zhang, Y\. Yue, Z\. Li, S\. Yun, G\. Wan, K\. Wang, D\. Cheng, J\. X\. Yu, and T\. Chen \(2025a\)Cut the crap: an economical communication pipeline for llm\-based multi\-agent systems\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2605.19418#S1.p1.1)\.
- G\. Zhang, Y\. Yue, X\. Sun, G\. Wan, M\. Yu, J\. Fang, K\. Wang, T\. Chen, and D\. Cheng \(2025b\)G\-designer: architecting multi\-agent communication topologies via graph neural networks\.InInternational Conference on Machine Learning \(ICML\),pp\. 76678–76692\.Cited by:[10th item](https://arxiv.org/html/2605.19418#A4.I2.i10.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p2.1),[§1](https://arxiv.org/html/2605.19418#S1.p3.1),[§2](https://arxiv.org/html/2605.19418#S2.p2.25),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1),[§5](https://arxiv.org/html/2605.19418#S5.p1.1),[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.
- Z\. Zhang, L\. Li, S\. Wan, S\. Wang, Z\. Wang, Z\. Lu, D\. Hao, and W\. Li \(2024\)DropEdge not foolproof: effective augmentation method for signed graph neural networks\.InProceedings of the 38th International Conference on Neural Information Processing Systems \(NeurIPS\),pp\. 117041–117069\.Cited by:[§F\.1](https://arxiv.org/html/2605.19418#A6.SS1.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p4.1)\.
- C\. Zheng, Z\. Liu, E\. Xie, Z\. Li, and Y\. Li \(2023\)Progressive\-hint prompting improves reasoning in large language models\.arXiv preprint arXiv:2304\.09797\.Cited by:[4th item](https://arxiv.org/html/2605.19418#A4.I2.i4.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1)\.
- Z\. Zhou, B\. Hu, C\. Zhao, P\. Zhang, and B\. Liu \(2024\)Large language model as a policy teacher for training reinforcement learning agents\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence \(IJCAI\),pp\. 5671–5679\.Cited by:[§5](https://arxiv.org/html/2605.19418#S5.p1.1)\.
- M\. Zhuge, W\. Wang, L\. Kirsch, F\. Faccio, D\. Khizbullin, and J\. Schmidhuber \(2024\)Gptswarm: language agents as optimizable graphs\.InInternational Conference on Machine Learning \(ICML\),Cited by:[9th item](https://arxiv.org/html/2605.19418#A4.I2.i9.p1.1),[§F\.2](https://arxiv.org/html/2605.19418#A6.SS2.p1.1),[§1](https://arxiv.org/html/2605.19418#S1.p2.1),[§1](https://arxiv.org/html/2605.19418#S1.p3.1),[§2](https://arxiv.org/html/2605.19418#S2.p2.25),[§2](https://arxiv.org/html/2605.19418#S2.p2.8),[§3\.2](https://arxiv.org/html/2605.19418#S3.SS2.p1.1),[§4\.1](https://arxiv.org/html/2605.19418#S4.SS1.p2.1),[§5](https://arxiv.org/html/2605.19418#S5.p1.1),[§5](https://arxiv.org/html/2605.19418#S5.p2.1)\.

Conflict\-Resilient Multi\-Agent Reasoning via Signed Graph Modeling

Supplementary Materials

## Appendix Contents

## Appendix ANotations

We summarize the main notations used throughout this paper in Table[2](https://arxiv.org/html/2605.19418#A1.T2)\. These notations cover the representation of agents, their interactions, the signed relational graph, multi\-hop neighborhoods, agent embeddings, selection scores, and aggregation functions used inSIGMA, ensuring clarity and facilitating understanding of the methodological and experimental descriptions\.

Table 2:Notations used inSIGMA\.NotationDescription𝒢=\(𝒱,ℰ\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\)Multi\-agent interaction graph with agents𝒱\\mathcal\{V\}and edgesℰ\\mathcal\{E\}\.N=\|𝒱\|N=\|\\mathcal\{V\}\|Total number of agents\.vi∈𝒱v\_\{i\}\\in\\mathcal\{V\}Theii\-th agent\.𝑨∈ℝN×N\\boldsymbol\{A\}\\in\\mathbb\{R\}^\{N\\times N\}Interaction matrix;𝑨i​j\\boldsymbol\{A\}\_\{ij\}measures influence fromvjv\_\{j\}toviv\_\{i\}\.𝑨\+,𝑨−\\boldsymbol\{A\}^\{\+\},\\boldsymbol\{A\}^\{\-\}Positive \(trust\) and negative \(conflict\) adjacency matrices\.𝑨~\\tilde\{\\boldsymbol\{A\}\}Normalized signed interaction matrix\.𝒩​\(i\)\\mathcal\{N\}\(i\)Neighborhood of agentviv\_\{i\}\.𝒩i\+,𝒩i−\\mathcal\{N\}\_\{i\}^\{\+\},\\mathcal\{N\}\_\{i\}^\{\-\}Positive and negative neighbors ofviv\_\{i\}\.Bi\(ℓ\),Ui\(ℓ\)B\_\{i\}^\{\(\\ell\)\},U\_\{i\}^\{\(\\ell\)\}ℓ\\ell\-hop balanced and unbalanced neighborhoods ofviv\_\{i\}\.QQInput query\.TTNumber of interaction \(reasoning\) rounds\.ddDimension of agent representations\.𝒉i\(t\)∈ℝd\\boldsymbol\{h\}\_\{i\}^\{\(t\)\}\\in\\mathbb\{R\}^\{d\}Hidden state of agentviv\_\{i\}at iterationtt\.𝑯\(t\)∈ℝN×d\\boldsymbol\{H\}^\{\(t\)\}\\in\\mathbb\{R\}^\{N\\times d\}Global representation matrix at iterationtt\.𝒉ipos,𝒉ineg\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\},\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\}Positive and negative embeddings of agentviv\_\{i\}\.fi​\(⋅\)f\_\{i\}\(\\cdot\)LLM\-based reasoning function of agentviv\_\{i\}\.ℱ​\(⋅\)\\mathcal\{F\}\(\\cdot\)Global update function for all agents\.𝒜​\(⋅\)\\mathcal\{A\}\(\\cdot\)Global aggregation \(readout\) function\.𝗋𝖾𝗅​\(vi,Q\)\\mathsf\{rel\}\(v\_\{i\},Q\)Semantic relevance between agentviv\_\{i\}and queryQQ\.𝖼𝗈𝗇𝖿​\(vi\)\\mathsf\{conf\}\(v\_\{i\}\)Confidence \(reliability\) of agentviv\_\{i\}\.sim​\(⋅,⋅\)\\mathrm\{sim\}\(\\cdot,\\cdot\)Similarity function \(e\.g\., cosine similarity\)\.𝒔​\(vi\)\\boldsymbol\{s\}\(v\_\{i\}\)Unified agent selection score\.𝒱s\\mathcal\{V\}\_\{s\}Selected subset of agents\.kkNumber of selected agents \(\|𝒱s\|=k\|\\mathcal\{V\}\_\{s\}\|=k\)\.feval​\(⋅,⋅\)f\_\{\\text\{eval\}\}\(\\cdot,\\cdot\)Pairwise evaluation function for relational compatibility\.αi​j,βi​j\\alpha\_\{ij\},\\beta\_\{ij\}Normalized aggregation weights for positive and negative edges\.ϕ​\(⋅\)\\phi\(\\cdot\)Learnable transformation in message passing\.𝒜​𝒢​𝒢​\(⋅\)\\mathcal\{AGG\}\(\\cdot\)Aggregation function \(e\.g\., mean or attention\)\.𝒞​\(⋅\)\\mathcal\{C\}\(\\cdot\)Combination function for updating embeddings\.∥\\\|Concatenation or fusion function\.𝒘i\\boldsymbol\{w\}\_\{i\}Signed importance weight of agentviv\_\{i\}in readout\.yyFinal prediction \(global output\)\.λ\\lambdaTrade\-off hyperparameter in agent selection\.ϵ\\epsilonSmall constant for numerical stability\.

## Appendix BAlgorithm and Complexity Analysis

### B\.1Algorithm

We present the inference pipeline ofSIGMA\(Algorithm[1](https://arxiv.org/html/2605.19418#algorithm1)\), detailing how agents are selected, interactions are modeled via a signed relational graph, and conflict\-aware message passing is performed to derive the final prediction\. We then provide a comprehensive computational complexity analysis to characterize the efficiency and scalability ofSIGMAin practical multi\-agent reasoning scenarios\.

Input:Query

QQ, pool of agents

𝒱\\mathcal\{V\}, top\-

kkagents

kk, number of interaction iterations

TT, number of message\-passing layers

LL
Output:Final prediction

yy
1

Phase I: Query\-Guided Agent Selection

2for*each agentvi∈𝒱v\_\{i\}\\in\\mathcal\{V\}*do

3Compute semantic relevance:

𝗋𝖾𝗅​\(vi,Q\)\\mathsf\{rel\}\(v\_\{i\},Q\);

4Compute structural diversity:

𝖽𝗂𝗏​\(vi\)=𝔼j≠i​\[1−sim​\(𝒉i\(0\),𝒉j\(0\)\)\]\\mathsf\{div\}\(v\_\{i\}\)=\\mathbb\{E\}\_\{j\\neq i\}\[1\-\\mathrm\{sim\}\(\\boldsymbol\{h\}\_\{i\}^\{\(0\)\},\\boldsymbol\{h\}\_\{j\}^\{\(0\)\}\)\];

5Compute agent confidence:

𝖼𝗈𝗇𝖿​\(vi\)\\mathsf\{conf\}\(v\_\{i\}\);

6Compute composite score:

𝒔​\(vi\)=λ​𝗋𝖾𝗅​\(vi,Q\)\+1−λ2​\(𝖽𝗂𝗏​\(vi\)\+𝖼𝗈𝗇𝖿​\(vi\)\)\\boldsymbol\{s\}\(v\_\{i\}\)=\\lambda\\mathsf\{rel\}\(v\_\{i\},Q\)\+\\frac\{1\-\\lambda\}\{2\}\(\\mathsf\{div\}\(v\_\{i\}\)\+\\mathsf\{conf\}\(v\_\{i\}\)\);

7

8end for

9Select top\-

kkagents:

𝒱s←TopKvi∈𝒱⁡𝒔​\(vi\)\\mathcal\{V\}\_\{s\}\\leftarrow\\operatorname\{TopK\}\_\{v\_\{i\}\\in\\mathcal\{V\}\}\\boldsymbol\{s\}\(v\_\{i\}\);

10

Phase II: Signed Relational Graph Construction

11for*each pair\(vi,vj\)∈𝒱s\(v\_\{i\},v\_\{j\}\)\\in\\mathcal\{V\}\_\{s\}*do

12Compute pairwise evaluation:

fi​j=feval​\(𝒉i\(0\),𝒉j\(0\)\)f\_\{ij\}=f\_\{\\text\{eval\}\}\(\\boldsymbol\{h\}\_\{i\}^\{\(0\)\},\\boldsymbol\{h\}\_\{j\}^\{\(0\)\}\);

13Extract relation polarity:

si​j←sign⁡\(fi​j\)s\_\{ij\}\\leftarrow\\operatorname\{sign\}\(f\_\{ij\}\);

14Extract confidence magnitude:

wi​j←\|fi​j\|w\_\{ij\}\\leftarrow\|f\_\{ij\}\|;

15Assign signed adjacency:

Ai​j←si​j⋅wi​jA\_\{ij\}\\leftarrow s\_\{ij\}\\cdot w\_\{ij\};

16

17end for

18Partition neighborhoods:

𝒩i\+=\{j:𝑨i​j\>0\}\\mathcal\{N\}\_\{i\}^\{\+\}=\\\{j:\\boldsymbol\{A\}\_\{ij\}\>0\\\},

𝒩i−=\{j:𝑨i​j<0\}\\mathcal\{N\}\_\{i\}^\{\-\}=\\\{j:\\boldsymbol\{A\}\_\{ij\}<0\\\};

19for*ℓ=1\\ell=1toLL*do

20Compute balanced neighborhood

Bi\(ℓ\)B\_\{i\}^\{\(\\ell\)\}and unbalanced neighborhood

Ui\(ℓ\)U\_\{i\}^\{\(\\ell\)\}recursively;

21

22end for

23

Phase III: Conflict\-Aware Signed Message Passing

24Initialize agent states:\{𝒉i\(0\)\}\\\{\\boldsymbol\{h\}\_\{i\}^\{\(0\)\}\\\},\{𝒉ipos​\(0,0\)\}\\\{\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(0,0\)\}\\\},\{𝒉ineg​\(0,0\)\}\\\{\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(0,0\)\}\\\};

25

26for*t=1t=1toTT*do

27Initialize layer states:

𝒉ipos​\(t,0\)←𝒉ipos​\(t−1,L\)\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(t,0\)\}\\leftarrow\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(t\-1,L\)\},

𝒉ineg​\(t,0\)←𝒉ineg​\(t−1,L\)\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(t,0\)\}\\leftarrow\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(t\-1,L\)\};

28

29for*ℓ=1\\ell=1toLL*do

30for*each agentvi∈𝒱sv\_\{i\}\\in\\mathcal\{V\}\_\{s\}*do

31Update positive state:

𝒉ipos​\(t,ℓ\)←𝒞\(ℓ\)​\(𝒉ipos​\(t,ℓ−1\),𝒜​𝒢​𝒢\(ℓ\)​\(\{𝒉jpos​\(t,ℓ−1\):j∈Bi\(ℓ\)\},\{𝒉jneg​\(t,ℓ−1\):j∈Ui\(ℓ\)\}\)\)\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(t,\\ell\)\}\\leftarrow\\mathcal\{C\}^\{\(\\ell\)\}\\Bigl\(\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(t,\\ell\-1\)\},\\mathcal\{AGG\}^\{\(\\ell\)\}\\bigl\(\\\{\\boldsymbol\{h\}\_\{j\}^\{\\mathrm\{pos\}\(t,\\ell\-1\)\}:j\\in B\_\{i\}^\{\(\\ell\)\}\\\},\\\{\\boldsymbol\{h\}\_\{j\}^\{\\mathrm\{neg\}\(t,\\ell\-1\)\}:j\\in U\_\{i\}^\{\(\\ell\)\}\\\}\\bigr\)\\Bigr\);

32

33Update negative state:

𝒉ineg​\(t,ℓ\)←𝒞\(ℓ\)​\(𝒉ineg​\(t,ℓ−1\),𝒜​𝒢​𝒢\(ℓ\)​\(\{𝒉jneg​\(t,ℓ−1\):j∈Bi\(ℓ\)\},\{𝒉jpos​\(t,ℓ−1\):j∈Ui\(ℓ\)\}\)\)\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(t,\\ell\)\}\\leftarrow\\mathcal\{C\}^\{\(\\ell\)\}\\Bigl\(\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(t,\\ell\-1\)\},\\mathcal\{AGG\}^\{\(\\ell\)\}\\bigl\(\\\{\\boldsymbol\{h\}\_\{j\}^\{\\mathrm\{neg\}\(t,\\ell\-1\)\}:j\\in B\_\{i\}^\{\(\\ell\)\}\\\},\\\{\\boldsymbol\{h\}\_\{j\}^\{\\mathrm\{pos\}\(t,\\ell\-1\)\}:j\\in U\_\{i\}^\{\(\\ell\)\}\\\}\\bigr\)\\Bigr\);

34

35end for

36

37end for

38

39end for

40

41Fuse final representations:

𝒉i\(T\)←𝒉ipos​\(T,L\)∥𝒉ineg​\(T,L\)\\boldsymbol\{h\}\_\{i\}^\{\(T\)\}\\leftarrow\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{pos\}\(T,L\)\}\\;\\\|\\;\\boldsymbol\{h\}\_\{i\}^\{\\mathrm\{neg\}\(T,L\)\};

42

Phase IV: Signed Consensus Readout

43Compute agent weights:𝒘i←∑j𝑨i​j∑p\|∑q𝑨p​q\|\\boldsymbol\{w\}\_\{i\}\\leftarrow\\frac\{\\sum\_\{j\}\\boldsymbol\{A\}\_\{ij\}\}\{\\sum\_\{p\}\\left\|\\sum\_\{q\}\\boldsymbol\{A\}\_\{pq\}\\right\|\};

44

45Generate final prediction through signed consensus:

y←LLM​\-​Gen⁡\(\{\(𝒉i\(T\),𝒘i\)\}i=1k\)y\\leftarrow\\operatorname\{LLM\\text\{\-\}Gen\}\\bigl\(\\\{\(\\boldsymbol\{h\}\_\{i\}^\{\(T\)\},\\boldsymbol\{w\}\_\{i\}\)\\\}\_\{i=1\}^\{k\}\\bigr\);

46

47return

yy;

Algorithm 1Overall ofSIGMA

### B\.2Complexity Analysis

We analyze the computational complexity of each stage ofSIGMAin detail, carefully considering the main operations performed during agent selection, signed graph construction, conflict\-aware message passing, and the final signed consensus readout process:

- •Stage I: Agent Selection\.Computing the selection score in[Equation˜5](https://arxiv.org/html/2605.19418#S3.E5)involves: \(1\) relevance estimation, \(2\) pairwise similarity for diversity, and \(3\) confidence estimation\. The dominant cost arises from pairwise similarity computation, which takes𝒪​\(N2​d\)\\mathcal\{O\}\(N^\{2\}d\), whereNNis the number of candidate agents andddis the representation dimension\. Selecting top\-kkagents takes𝒪​\(N​log⁡N\)\\mathcal\{O\}\(N\\log N\)\.
- •Stage II: Signed Graph Construction\.Constructing the signed adjacency matrix requires evaluating all agent pairs in𝒱s\\mathcal\{V\}\_\{s\}, leading to a complexity of𝒪​\(k2​d\)\\mathcal\{O\}\(k^\{2\}d\)whenfevalf\_\{\\text\{eval\}\}is implemented via similarity functions\. Neighborhood partitioning takes𝒪​\(k2\)\\mathcal\{O\}\(k^\{2\}\)\.
- •Stage III: Signed Message Passing\.At each iteration, each agent aggregates information from its neighbors\. The per\-iteration cost is𝒪​\(k2​d\)\\mathcal\{O\}\(k^\{2\}d\)\. OverTTiterations, the total complexity is𝒪​\(T​k2​d\)\\mathcal\{O\}\(Tk^\{2\}d\)\.
- •Stage IV: Signed Consensus Readout\.Computing weights𝒘i\\boldsymbol\{w\}\_\{i\}requires summing over adjacency entries, taking𝒪​\(k2\)\\mathcal\{O\}\(k^\{2\}\)\. Final aggregation takes𝒪​\(k​d\)\\mathcal\{O\}\(kd\)\.

The overall computational complexity ofSIGMAis𝒪​\(N2​d\)\+𝒪​\(k2​d\)\+𝒪​\(T​k2​d\),\\mathcal\{O\}\(N^\{2\}d\)\+\\mathcal\{O\}\(k^\{2\}d\)\+\\mathcal\{O\}\(Tk^\{2\}d\),whereNNis the total number of agents,kkis the number of selected agents,TTis the number of message\-passing iterations, andddis the feature dimension\. Sincek≪Nk\\ll NandTTis small, the complexity is dominated by the agent selection stage, yielding an approximate cost of𝒪​\(N2​d\)\\mathcal\{O\}\(N^\{2\}d\)\. In practice, with smallkkandTT\(e\.g\.,k≤10k\\leq 10,T≤3T\\leq 3\),SIGMAremains efficient and scalable for multi\-agent reasoning tasks\.

## Appendix CTheoretical Analysis

In this part, we provide theoretical insights into the effectiveness of signed graph modeling in multi\-agent reasoning\. Specifically, we analyze: ① its ability to suppress error propagation, and ② the stability of signed message passing, thereby providing intuition for the observed robustness and conflict\-resilience\.

Remark on Theoretical Scope\.The following analysis is intended to provide qualitative insights into the behavior of signed aggregation rather than strict guarantees under all real\-world conditions\. In practice, LLM\-based agents may exhibit correlated and biased errors; however, the analysis highlights how explicit modeling of trust and conflict can mitigate error propagation under reasonable assumptions\. Consider a set ofkkagents with representations𝒉i=𝒉i∗\+ϵi\\boldsymbol\{h\}\_\{i\}=\\boldsymbol\{h\}\_\{i\}^\{\*\}\+\\boldsymbol\{\\epsilon\}\_\{i\}, where𝒉i∗\\boldsymbol\{h\}\_\{i\}^\{\*\}denotes the true signal andϵi\\boldsymbol\{\\epsilon\}\_\{i\}represents noise\. Let the signed adjacency matrix be𝑨=𝑨\+−𝑨−\\boldsymbol\{A\}=\\boldsymbol\{A\}^\{\+\}\-\\boldsymbol\{A\}^\{\-\}and define its normalized form as𝑨~i​j=𝑨i​j/\(∑k\|𝑨i​k\|\+ϵ\)\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}=\\boldsymbol\{A\}\_\{ij\}/\(\\sum\_\{k\}\|\\boldsymbol\{A\}\_\{ik\}\|\+\\epsilon\)\. The iterative propagation is then given by𝑯\(t\)=𝑨~​ϕ​\(𝑯\(t−1\)\)\\boldsymbol\{H\}^\{\(t\)\}=\\tilde\{\\boldsymbol\{A\}\}\\,\\phi\(\\boldsymbol\{H\}^\{\(t\-1\)\}\), whereϕ​\(⋅\)\\phi\(\\cdot\)is aLϕL\_\{\\phi\}\-Lipschitz continuous function\.

### C\.1Error Suppression Property

Lemma 1 \(Error Suppression via Signed Aggregation\)\.Under mild assumptions on signal alignment and noise correlation, signed aggregation improves the signal\-to\-noise ratio \(SNR\)\[[4](https://arxiv.org/html/2605.19418#bib.bib41)\]in expectation compared to unsigned aggregation\. Formally, for agentiiwith clean signal𝒔i\\boldsymbol\{s\}\_\{i\}and noiseϵi\\boldsymbol\{\\epsilon\}\_\{i\}, after one\-step aggregation:𝒉i\(t\)=∑j𝑨~i​j​\(𝒔j\+ϵj\),\\boldsymbol\{h\}\_\{i\}^\{\(t\)\}=\\sum\_\{j\}\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}\(\\boldsymbol\{s\}\_\{j\}\+\\boldsymbol\{\\epsilon\}\_\{j\}\),the resulting SNR satisfies:𝔼​\[SNRisigned\]≥𝔼​\[SNRiunsigned\],\\mathbb\{E\}\[\\mathrm\{SNR\}\_\{i\}^\{\\text\{signed\}\}\]\\geq\\mathbb\{E\}\[\\mathrm\{SNR\}\_\{i\}^\{\\text\{unsigned\}\}\],\(10\)provided that the polarity of edges is consistent with the underlying signal structure\.

###### Proof\.

We decompose each representation as𝒉i=𝒔i\+ϵi\\boldsymbol\{h\}\_\{i\}=\\boldsymbol\{s\}\_\{i\}\+\\boldsymbol\{\\epsilon\}\_\{i\}, where𝒔i\\boldsymbol\{s\}\_\{i\}denotes the underlying task\-relevant signal component, whileϵi\\boldsymbol\{\\epsilon\}\_\{i\}captures noise induced by imperfect reasoning, model uncertainty, and conflicting or unreliable information across agents\.

Noise Assumption\.We assume that𝔼​\[ϵi\]=0\\mathbb\{E\}\[\\boldsymbol\{\\epsilon\}\_\{i\}\]=0and allow weak cross\-agent correlation:

\|𝔼​\[ϵi⊤​ϵj\]\|≤ρ,i≠j,\|\\mathbb\{E\}\[\\boldsymbol\{\\epsilon\}\_\{i\}^\{\\top\}\\boldsymbol\{\\epsilon\}\_\{j\}\]\|\\leq\\rho,\\quad i\\neq j,\(11\)whereρ≥0\\rho\\geq 0is a small constant capturing limited dependence among agent noises\.

Signal Alignment Assumption\.For supportive edges\(i,j\)\(i,j\), we assume𝒔i⊤​𝒔j\>0\\boldsymbol\{s\}\_\{i\}^\{\\top\}\\boldsymbol\{s\}\_\{j\}\>0, while for conflicting edges,𝒔i⊤​𝒔j<0\\boldsymbol\{s\}\_\{i\}^\{\\top\}\\boldsymbol\{s\}\_\{j\}<0\. After aggregation, we obtain:

𝒉i\(t\)=∑j𝑨~i​j​𝒉j,ϵi\(t\)=∑j𝑨~i​j​ϵj\.\\boldsymbol\{h\}\_\{i\}^\{\(t\)\}=\\sum\_\{j\}\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}\\boldsymbol\{h\}\_\{j\},\\quad\\boldsymbol\{\\epsilon\}\_\{i\}^\{\(t\)\}=\\sum\_\{j\}\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}\\boldsymbol\{\\epsilon\}\_\{j\}\.\(12\)The resulting noise power can be bounded as:

𝔼​\[‖ϵi\(t\)‖22\]≤∑j𝑨~i​j2​σj2\+𝒪​\(ρ\),\\mathbb\{E\}\[\\\|\\boldsymbol\{\\epsilon\}\_\{i\}^\{\(t\)\}\\\|\_\{2\}^\{2\}\]\\leq\\sum\_\{j\}\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}^\{2\}\\sigma\_\{j\}^\{2\}\+\\mathcal\{O\}\(\\rho\),\(13\)where the second term accounts for residual effects due to weak noise correlations\. Meanwhile, the signal component evolves as:

𝒔i\(t\)=∑j𝑨~i​j​𝒔j,\\boldsymbol\{s\}\_\{i\}^\{\(t\)\}=\\sum\_\{j\}\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}\\boldsymbol\{s\}\_\{j\},\(14\)which benefits from polarity\-aware aggregation: supportive signals are reinforced, while conflicting signals are partially canceled\. Compared to unsigned aggregation, this mechanism suppresses the contribution of misaligned signals while preserving coherent ones\.

Under the alignment assumption, such polarity\-aware interactions increase the effective signal magnitude relative to noise, thereby improving the expected signal\-to\-noise ratio \(SNR\)\. ∎

### C\.2Stability of Signed Message Passing

Theorem 1 \(Stability\)\.Supposeϕ\\phiisLϕL\_\{\\phi\}\-Lipschitz continuous\[[1](https://arxiv.org/html/2605.19418#bib.bib44)\]and the normalized matrix satisfies∑j\|𝑨~i​j\|≤1\\sum\_\{j\}\|\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}\|\\leq 1, the signed message passing update is stable under the infinity norm:‖𝑯\(t\)‖∞≤Lϕ​‖𝑯\(t−1\)‖∞,\\\|\\boldsymbol\{H\}^\{\(t\)\}\\\|\_\{\\infty\}\\leq L\_\{\\phi\}\\\|\\boldsymbol\{H\}^\{\(t\-1\)\}\\\|\_\{\\infty\},\(15\)and the iteration is contractive whenLϕ<1L\_\{\\phi\}<1\.

###### Proof\.

By the Lipschitz continuity ofϕ\\phi, we have:

‖ϕ​\(𝑿\)‖∞≤Lϕ​‖𝑿‖∞\.\\\|\\phi\(\\boldsymbol\{X\}\)\\\|\_\{\\infty\}\\leq L\_\{\\phi\}\\\|\\boldsymbol\{X\}\\\|\_\{\\infty\}\.\(16\)Next, leveraging the row\-normalization property of𝑨~\\tilde\{\\boldsymbol\{A\}\}, we obtain:

‖𝑨~​𝑿‖∞=maxi⁡\|∑j𝑨~i​j​𝑿j\|≤maxi​∑j\|𝑨~i​j\|​‖𝑿j‖∞≤‖𝑿‖∞\.\\\|\\tilde\{\\boldsymbol\{A\}\}\\boldsymbol\{X\}\\\|\_\{\\infty\}=\\max\_\{i\}\\Big\|\\sum\_\{j\}\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}\\boldsymbol\{X\}\_\{j\}\\Big\|\\leq\\max\_\{i\}\\sum\_\{j\}\|\\tilde\{\\boldsymbol\{A\}\}\_\{ij\}\|\\\|\\boldsymbol\{X\}\_\{j\}\\\|\_\{\\infty\}\\leq\\\|\\boldsymbol\{X\}\\\|\_\{\\infty\}\.\(17\)Combining the above inequalities yields:

‖𝑯\(t\)‖∞=‖𝑨~​ϕ​\(𝑯\(t−1\)\)‖∞≤Lϕ​‖𝑯\(t−1\)‖∞\.\\\|\\boldsymbol\{H\}^\{\(t\)\}\\\|\_\{\\infty\}=\\\|\\tilde\{\\boldsymbol\{A\}\}\\phi\(\\boldsymbol\{H\}^\{\(t\-1\)\}\)\\\|\_\{\\infty\}\\leq L\_\{\\phi\}\\\|\\boldsymbol\{H\}^\{\(t\-1\)\}\\\|\_\{\\infty\}\.\(18\)This establishes the stability of the iterative update\. Furthermore, whenLϕ<1L\_\{\\phi\}<1, the mapping becomes contractive, implying that the sequence\{𝑯\(t\)\}\\\{\\boldsymbol\{H\}^\{\(t\)\}\\\}converges\. ∎

### C\.3Discussion

The above analysis provides theoretical insights into the design ofSIGMA\. Specifically:

- •Signed aggregation mitigates error propagation by reinforcing aligned signals while attenuating conflicting ones, leading to improved SNR under mild assumptions\.
- •The normalized signed propagation ensures stability of iterative updates, preventing uncontrolled amplification of representations\.

Importantly, these properties are directly enabled by the explicit modeling of interaction polarity inSIGMA, which distinguishes it from prior graph\-based MAS approaches that assume homogeneous positive interactions\. Together, they help explain the empirical robustness and conflict\-resilient performance ofSIGMA\.

## Appendix DDetails on Experimental Settings

Table 3:Descriptions and statistics of benchmark datasets used to evaluateSIGMA, including multi\-domain general reasoning datasets, domain\-specific mathematical reasoning datasets, and code generation datasets, along with their answer types, evaluation metrics, test set sizes, and licenses\.CategoryDatasetAnswer TypeMetricTestLicenseMulti\-domainMMLUMulti\-choiceAcc\.153MIT LicenseMulti\-domainMMLU\-ProMulti\-choiceAcc\.152MIT LicenseMulti\-domainGPQAMulti\-choiceAcc\.1,000MIT LicenseMathematical reasoningGSM8KNumberAcc\.1,319MIT LicenseMathematical reasoningMultiArithNumberAcc\.600UnspecifiedCode generationHumanEvalCodePass@1164MIT License

### D\.1Datasets

We evaluate theSIGMAframework on the following benchmark datasets, selected to cover both multi\-domain general reasoning and domain\-specific tasks\. Detailed statistics, including test set sizes, answer types, and evaluation metrics, are provided in[Table˜3](https://arxiv.org/html/2605.19418#A4.T3)to give a comprehensive overview of each dataset and facilitate reproducibility\. All datasets are publicly available\.MMLU,MMLU\-Pro, andGPQAcan be accessed via GitHub111[https://github\.com](https://github.com/)or Hugging Face222[https://huggingface\.co](https://huggingface.co/)Datasets, whileGSM8K,MultiArith, andHumanEvalare available through Hugging Face Datasets or their respective official repositories\.

- •MMLU\[[21](https://arxiv.org/html/2605.19418#bib.bib78)\]: A comprehensive multi\-domain benchmark with 57 diverse subjects \(which can be grouped into 4 broader categories\), designed to evaluate general reasoning and knowledge across humanities, STEM, and other domains\. In our experiments, we employed stratified sampling with 50 test samples per subject \(resulting in a subset of 153 questions\) to ensure balanced representation\.
- •MMLU\-Pro\[[50](https://arxiv.org/html/2605.19418#bib.bib79)\]: A professional\-level extension ofMMLU, comprising 14 categories focused on advanced reasoning tasks in professional and technical fields\. We sampled 150 test examples per category \(forming a representative subset\) to guarantee a balanced evaluation\.
- •GPQA\[[40](https://arxiv.org/html/2605.19418#bib.bib80)\]: A challenging benchmark targeting graduate\-level question answering, emphasizing deep understanding, logical reasoning, and domain\-specific knowledge\. It is particularly useful for evaluating multi\-hop reasoning capabilities of LLM\-based agents\.
- •GSM8K\[[8](https://arxiv.org/html/2605.19418#bib.bib81)\]andMultiArith\[[41](https://arxiv.org/html/2605.19418#bib.bib82)\]: Two widely used mathematical reasoning datasets designed to test step\-by\-step numerical problem solving\.GSM8Kcontains 1,319 examples, whileMultiArithhas 600 problems\. Both datasets are useful for evaluating symbolic reasoning and multi\-step calculation skills\.
- •HumanEval\[[6](https://arxiv.org/html/2605.19418#bib.bib83)\]: A code generation benchmark consisting of programming problems that assess functional correctness of generated code\. It contains 164 problems with diverse difficulty levels, testing the model’s ability to understand specifications and generate executable solutions\.

For completeness, we also report results on the fullMMLUtest set \(approximately 2,850 questions\) in the appendix, while the main results use the sampled subset for efficiency\.

### D\.2Baselines

We provide the baseline implementations with their respective licenses as follows\. For methods with publicly available code, we provide the GitHub links; for methods without released code, we provide the corresponding paper links\.

- •
- •
- •
- •
- •
- •
- •
- •
- •
- •
- •

For all baselines, we follow the recommended hyperparameters from the original papers or official implementations\. If unavailable or suboptimal, we apply careful tuning for best performance\. To ensure fairness, we standardize key settings \(e\.g\., hidden dimensions, layers\) to match our method, so that performance differences reflect model design rather than external factors\.

### D\.3Design of Malicious Agents

In this section, we introduce the design concepts for four types of problematic agents\. Each agent is crafted to simulate a distinct failure mode in multi\-agent math\-solving, allowing us to systematically evaluate system robustness under challenging scenarios:

- •Random\-Noise Agent \(RandomNoiseAgent\):Outputs completely random or unrelated text, simulating a fully malfunctioning or crashed agent that contributes zero useful information\.
- •Low\-Quality Reasoning Agent \(LowQualityAgent\):Operates with a system prompt that embeds cognitive biases, encouraging flawed strategies, premature conclusions, and missed details
- •Conflict\-Inducing Agent \(ConflictAgent\):Acts as a professional devil’s advocate, actively seeking flaws in the majority opinion and proposing alternative answers regardless of correctness\. It is particularly disruptive in multi\-turn settings where it observes other agents’ outputs\.
- •Blind\-Conformity Agent \(CopycatAgent\):Abandons independent reasoning and instead adopts answers proposed by peers, constructing post\-hoc justification that aligns with observed outputs rather than reflecting genuine understanding\.

The following subsections provide detailed system prompts and behavioral instructions for each of these agents, illustrated with coloredtcolorboxframes to highlight their distinct characteristics\.

RandomNoiseAgent Python Class[⬇](data:text/plain;base64,Y2xhc3MgUmFuZG9tTm9pc2VBZ2VudChBbmFseXplQWdlbnQpOgoKICAgIGRlZiBfX2luaXRfXyhzZWxmLCBkb21haW4sIGxsbV9uYW1lLCAqKmt3YXJncyk6CiAgICAgICAgIyBsbG1fbmFtZSBpcyBhIHBsYWNlaG9sZGVyOyB0aGlzIGFnZW50IGRvZXMgbm90IGludm9rZSB0aGUgTExNLgogICAgICAgIHN1cGVyKCkuX19pbml0X18oZG9tYWluLCBsbG1fbmFtZSwgKiprd2FyZ3MpCiAgICAgICAgc2VsZi5hZ2VudF90eXBlID0gImJhZCIKCiAgICBhc3luYyBkZWYgZ2VuZXJhdGVfcmVzcG9uc2Uoc2VsZiwgaW5wdXRfZGF0YSwgKiprd2FyZ3MpOgogICAgICAgIG5vaXNlX3RleHRzID0gWwogICAgICAgICAgICAiVGhlIGFuc3dlciBpcyA0MiBiZWNhdXNlIGl0J3MgdGhlIGFuc3dlciB0byBldmVyeXRoaW5nLiIsCiAgICAgICAgICAgICJJJ20gY29uZnVzZWQsIHBsZWFzZSBpZ25vcmUgbXkgcHJldmlvdXMgcmVhc29uaW5nLiIsCiAgICAgICAgICAgICJBY2NvcmRpbmcgdG8gbXkgY2FsY3VsYXRpb24sIHRoZSBwcm9iYWJpbGl0eSBpcyAwLjYxOC4iLAogICAgICAgICAgICAiSSB0aGluayB0aGUgcHJldmlvdXMgYWdlbnRzIGFyZSBjb3JyZWN0LiIsCiAgICAgICAgXQogICAgICAgIHJldHVybiByYW5kb20uY2hvaWNlKG5vaXNlX3RleHRzKQ==)classRandomNoiseAgent\(AnalyzeAgent\):def\_\_init\_\_\(self,domain,llm\_name,\*\*kwargs\):super\(\)\.\_\_init\_\_\(domain,llm\_name,\*\*kwargs\)self\.agent\_type="bad"asyncdefgenerate\_response\(self,input\_data,\*\*kwargs\):noise\_texts=\["Theansweris42becauseit’stheanswertoeverything\.","I’mconfused,pleaseignoremypreviousreasoning\.","Accordingtomycalculation,theprobabilityis0\.618\.","Ithinkthepreviousagentsarecorrect\.",\]returnrandom\.choice\(noise\_texts\)

System Prompt ofLowQualityAgentYou are a slightly careless and overconfident math problem solver\. Your task is to solve the given math problem step by step\.Key instructions:•You tend to skip detailed verification steps to save time\.•If a number looks "about right", you accept it without double\-checking\.•You often confuse similar mathematical operations \(e\.g\., multiplication vs addition\)\.•You rarely reconsider your initial approach, even when it leads to strange results\.•You don’t explicitly state unit conversions, and may forget them\.Your final answer should end with“The answer is \[result\]\.”

System Prompt ofConflictAgentYou are a professional devil’s advocate in a team of math solvers\. Your job is to critically examine other agents’ solutions and actively point out potential flaws or alternative interpretations\.Key instructions:•Always try to find a different answer or reasoning path from what others have proposed\.•If multiple agents agree on an answer, challenge it by exploring edge cases or questioning assumptions\.•Even if you secretly agree with the majority, you must present at least one credible counterargument\.•You value mathematical rigor over consensus\.•Phrase your disagreement constructively but firmly, e\.g\., "I respectfully disagree with the previous answer because…"Your final answer should end with“The answer is \[result\]\.”

System Prompt ofCopycatAgentYou are a highly agreeable collaborator in a math\-solving team\. You believe that the majority opinion is usually correct, and you prioritize team harmony over independent verification\.Key instructions:•First, check if any other agent has already proposed an answer\.•If there is a clear majority answer, adopt it directly and construct reasoning that supports that answer\.•If no consensus exists yet, provide your own reasoning, but state that you are open to changing your mind\.•Avoid contradicting teammates; find ways to align your reasoning with theirs\.•Use phrases like "I agree with the previous analysis…", "Building on what others have said…"Your final answer should end with“The answer is \[result\]\.”

## Appendix EAdditional Experiment Results

### E\.1Main Results

In this section, to further validate the effectiveness of our method, we present a carefully selected set of detailed experimental results forSIGMAacross four datasets: two general reasoning datasets and two domain\-specific datasets\. These experiments are conducted using multi\-agent setups based onDeepSeek\-V3\.2, allowing for comprehensive evaluation of our approach\.

Table 4:Performance comparison with single\-agent and multi\-agent baselines across multiple benchmarks\. The best results are inbold, and the runner\-ups areunderlined\. “Mul\.”, “Rel\.”, and “Conf\.” indicate whether the method supports multi\-agent collaboration, models inter\-agent relations, and handles conflicting signals, respectively\. Hollow, half\-filled, and filled circles denote no, partial, and full support in these aspects\.†\\daggernotably indicates papers with over one hundred citations\.MethodMul\.Rel\.Conf\.MMLUMMLU\-ProGSM8KHumanEvalAvg\.Vanilla∘\\circ ∘\\circ ∘\\circ 87\.4675\.7194\.4793\.1387\.69CoT\(⊳\\rhdNeurIPS’22†\)∘\\circ ∘\\circ ∘\\circ 88\.74↑1\.2885\.72↑10\.0194\.85↑0\.3893\.41↑0\.2890\.68ComplexCoT\(⊳\\rhdICLR’24†\)∘\\circ ∘\\circ ∘\\circ 89\.41↑1\.9586\.45↑10\.7495\.03↑0\.5693\.14↑0\.0291\.01MoA\(⊳\\rhdICLR’25†\)∙\\bullet ∘\\circ ∘\\circ 89\.77↑2\.3180\.76↑5\.0594\.77↑0\.3094\.03↑0\.9089\.83GPTSwarm\(⊳\\rhdICML’24†\)∙\\bullet ∙\\bullet 90\.03↑2\.5790\.12↑14\.4195\.24↑0\.7794\.16↑1\.0392\.39G\-Designer\(⊳\\rhdICML’25†\)∙\\bullet ∙\\bullet 90\.85↑3\.3990\.54↑14\.8396\.66↑2\.1994\.82↑1\.6993\.22GoA\(⊳\\rhdICLR’26\)∙\\bullet ∙\\bullet 91\.53↑4\.0790\.24↑14\.5395\.43↑0\.9694\.25↑1\.1292\.86SIGMA∙\\bullet ∙\\bullet ∙\\bullet 92\.16↑4\.7091\.43↑15\.7296\.81↑2\.3495\.22↑2\.0993\.91

Based on the results reported in Table[4](https://arxiv.org/html/2605.19418#A5.T4), several observations can be made\. First, compared to single\-agent baselines CoT and ComplexCoT, multi\-agent methods consistently achieve higher performance across all datasets, particularly onMMLUandMMLU\-Pro, demonstrating the benefits of collaborative reasoning in handling complex tasks\. Second, existing multi\-agent approaches such as MoA, GPTSwarm, G\-Designer and GoA differ in their support for multi\-agent interaction Mul\., relational modeling Rel\., and partial conflict handling Conf\., resulting in varying degrees of performance gains\. For instance, G\-Designer performs well onGSM8KandHumanEvalbut is still slightly behind GoA onMMLU\. Finally, our methodSIGMAachieves the best results across all four datasets, reaching 92\.16% and 91\.43% onMMLUandMMLU\-Pro, respectively, with an improvement of 0\.63 and 1\.90 points over GoA, and demonstrating stable gains onGSM8K96\.81% and HumanEval 95\.22%\. These results strongly validate that our explicit trust–conflict modeling effectively enhances the reliability and consistency of multi\-agent reasoning\. Moreover,SIGMAsimultaneously supports multi\-agent interaction, inter\-agent relational modeling and conflict\-aware signal processing, highlighting its comprehensive adaptability and robustness in complex reasoning scenarios\.

Experimental Analysis on GSM8K\.Taking the GSM8K dataset as an example, the experiments employed multi\-agent reasoning usingDeepSeek\-V3\.2\-based agents, with single\-agent methods serving as a baseline for comparison\. The dataset was divided into659659batches, with an average batch processing time of approximately73\.7973\.79seconds, and the overall average accuracy reached96\.81%96\.81\\%\. Outputs from each agent were integrated throughSIGMA’s Signed Consensus, where positive and negative edges correspond to trust and conflict relationships, respectively, effectively mitigating the influence of low\-confidence or conflicting information on the final answers\.

For each problem, the agents’ reasoning outputs and corresponding weights were recorded and used to generate the final consensus answer\. For instance, in theFollower growth task, the agents’ output weights were\[0\.35,−0\.20,0\.35,0\.10\]\[0\.35,\-0\.20,0\.35,0\.10\], and the weighted integration produced a total predicted follower count of180180, matching the ground truth; in theFrench fries problem, the final consensus output was4848, also consistent with the correct answer\. Overall,SIGMAefficiently combines reasoning from multiple agents, suppresses conflicting signals, and substantially improves accuracy on the GSM8K dataset, demonstrating robust and consistent multi\-agent collaborative reasoning performance\.

Experimental Analysis on HumanEval\.We evaluatedSIGMAon the HumanEval dataset usingDeepSeek\-V3\.2\-based agents for multi\-agent collaborative reasoning, with single\-agent execution serving as a baseline\. The dataset was divided into8080batches, with an average batch processing time of approximately65\.465\.4seconds, and the final average functional correctness accuracy reached95\.22%95\.22\\%\. Outputs from each agent were integrated usingSIGMA’s Signed Consensus, where positive and negative edges represent trust and conflict relationships, respectively, effectively mitigating the impact of low\-confidence or conflicting outputs on the final answers\.

For each problem, the reasoning outputs from all agents along with their corresponding weights were recorded and combined to generate the final consensus solution\. For example, in aFunction\-level task to determine whether a number is a simple power of another, the agents’ output weights were\[0\.45,−0\.15,0\.30,0\.10,−0\.10\]\[0\.45,\-0\.15,0\.30,0\.10,\-0\.10\], and after weighted integration, the consensus output correctly returnedTrueforis\_simple\_power\(8,2\)andFalseforis\_simple\_power\(3,2\), matching the expected results\. In this process, Signed Consensus assigns higher weights to more reliable agents while suppressing low\-confidence or conflicting outputs, effectively guiding the agents to reach the correct consensus\. Similarly, in another string manipulation task, the final consensus output passed all internal test cases, confirming correct functionality\. These results demonstrate thatSIGMAnot only integrates reasoning across multiple agents but also leverages the positive and negative edges in the signed graph to actively suppress erroneous signals, thereby improving the overall functional correctness of code generation\. Overall, our approach exhibits robust and consistent multi\-agent collaborative reasoning performance on the HumanEval dataset\.

### E\.2Hardware and Software Configurations\.

We conduct the experiments using the following hardware and software configurations via AutoDL333[https://autodl\.com](https://autodl.com/)with SSH access through PyCharm Professional:

- •Operating System:Ubuntu 20\.04 LTS
- •CPU:14\-core Intel\(R\) Xeon\(R\) Gold 6348 CPU @ 2\.60GHz with 120GB RAM
- •GPU:NVIDIA A800 80GB PCIe, CUDA 13\.0, Driver Version 580\.65\.06
- •Storage:System disk: 30GB; Data disk: 50GB SSD
- •Software:Python 3\.11, PyTorch 2\.11\.0, PyTorch Geometric 2\.7\.0, CUDA 13\.0
- •Access:Experiments run via SSH on AutoDL cloud instance using PyCharm Professional

## Appendix FAdditional Related Work

### F\.1Signed Graph Representation Learning\.

Signed graphs offer a principled framework to model interactions that can be either supportive or antagonistic, with edges encoding both polarity and magnitude\[[9](https://arxiv.org/html/2605.19418#bib.bib84),[35](https://arxiv.org/html/2605.19418#bib.bib85),[63](https://arxiv.org/html/2605.19418#bib.bib86)\]\. Rooted in social balance theory\[[34](https://arxiv.org/html/2605.19418#bib.bib87),[43](https://arxiv.org/html/2605.19418#bib.bib88)\], early work emphasized structural consistency, such as identifying balanced and unbalanced triads, while more recent advances in graph neural networks extend message passing to signed edges\[[31](https://arxiv.org/html/2605.19418#bib.bib89),[7](https://arxiv.org/html/2605.19418#bib.bib90)\], enabling learned representations that account for agreement and conflict\. These techniques have proven effective in social networks, recommendation systems, and opinion dynamics, capturing complex relational structures and mitigating propagation of misleading information\. Despite these successes, signed graph methods remain largely unexplored in LLM\-based multi\-agent systems \(MAS\) for reasoning, where agents may produce heterogeneous and sometimes conflicting outputs\. Leveraging signed graphs in MAS enables modeling of trust and conflict among agents, facilitating conflict\-aware aggregation, robust multi\-hop propagation, and improved consistency in multi\-agent reasoning under heterogeneous interactions\.

### F\.2Conflict\-Aware Multi\-Agent Learning

In LLM\-based multi\-agent systems \(MAS\), handling conflicting or unreliable agents has received growing attention\. Prior works investigate weighted aggregation\[[25](https://arxiv.org/html/2605.19418#bib.bib68),[58](https://arxiv.org/html/2605.19418#bib.bib51)\]and trust\-aware message passing\[[66](https://arxiv.org/html/2605.19418#bib.bib47),[59](https://arxiv.org/html/2605.19418#bib.bib49)\]to mitigate the impact of adversarial or low\-quality agents\. These approaches improve robustness in specific tasks, such as collaborative decision\-making or planning, by assigning higher influence to reliable agents while attenuating misleading signals from unreliable ones\. However, they are often limited to fixed topologies or task\-specific heuristics and rarely generalize to multi\-hop, reasoning\-intensive LLM\-based scenarios\. In contrast, our method leverages signed graph modeling to provide a principled and flexible framework that explicitly encodes both supportive and conflicting relationships, enabling conflict\-aware aggregation and robust multi\-agent reasoning under heterogeneous interactions\.

## Appendix GLimitation

Focus on reasoning benchmarks with clear answers\.SIGMAconsistently outperforms existing methods across six benchmarks spanning general reasoning, mathematical reasoning, and code generation, demonstrating the core value of explicit trust–conflict modeling in tasks that require precise and verifiable reasoning\. Currently, our evaluation focuses on these settings with well\-defined correct outputs, whileSIGMA’s performance on open\-ended tasks such as creative writing, multi\-turn dialogue, or strategic planning has not been specifically explored\. In such tasks, the notion of "correctness" is inherently more ambiguous, and disagreements among agents can sometimes be a source of creativity rather than noise to be eliminated\. Importantly, this scope does not diminishSIGMA’s advantages on mainstream reasoning tasks; rather, it opens up exciting opportunities for extending the signed graph approach to collaborative scenarios where open\-ended, constructive conflict plays a central role\.

## Appendix HBroader Impacts

SIGMAsignificantly enhances the robustness and reliability of multi\-agent reasoning by explicitly modeling trust, conflict, and neutral relations among agents through signed graphs\. This capability offers direct benefits for high\-stakes applications that demand precise reasoning, such as medical decision support, legal analysis, and scientific research, by effectively suppressing error propagation and dynamically isolating unreliable agents, thereby reducing the spread of misinformation in collaborative AI pipelines\. Moreover, the framework’s interpretable trust dynamics strengthen the transparency and auditability of automated decision\-making, providing meaningful support for the deployment of trustworthy AI\.

On the other hand, any technology aimed at improving group decision quality carries potential for misuse\. If improperly configured,SIGMA’s conflict\-suppression mechanism may inadvertently silence constructive dissent or minority viewpoints in open\-ended deliberation or opinion aggregation scenarios\. In addition, robust multi\-agent reasoning capabilities could be exploited to generate more persuasive disinformation\. It is worth emphasizing that the signed graph formalism underlyingSIGMAis inherently transparent: edge polarities and consensus weights can be inspected, adjusted, or overridden by human operators\. We encourage future work to accompany deployment with appropriate human oversight and usage guidelines, so as to maximize the benefits of collaborative AI while minimizing potential risks\.

Similar Articles

TMAS: Scaling Test-Time Compute via Multi-Agent Synergy

Hugging Face Daily Papers

TMAS introduces a multi-agent framework that enhances large language model reasoning by scaling test-time compute through structured collaboration and hierarchical memory systems. The approach uses specialized agents, cross-trajectory information flow, and hybrid reward reinforcement learning to improve iterative scaling and stability on challenging reasoning benchmarks.

Towards Security-Auditable LLM Agents: A Unified Graph Representation

arXiv cs.AI

This paper introduces Agent-BOM, a unified graph representation for security auditing in LLM-based agentic systems. It addresses the semantic gap in post-hoc auditing by modeling static capabilities and dynamic runtime states to detect complex attack chains like memory poisoning and tool misuse.