Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On
Summary
This vision paper argues that trust in Agent-to-Agent (A2A) networks must be integrated from the ground up, as existing agent alignment techniques are insufficient to address systemic vulnerabilities like adversarial composition and semantic misalignment.
View Cached Full Text
Cached at: 05/20/26, 08:27 AM
# Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On
Source: [https://arxiv.org/html/2605.19035](https://arxiv.org/html/2605.19035)
Yixiang Yao![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=2709,height=6\.83331pt\]\{all\-twemojis\.pdf\}\\includegraphics\[page=1958,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}, Yuhang Yao![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=696,height=6\.83331pt\]\{all\-twemojis\.pdf\}\\includegraphics\[page=3549,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}, Xinyi Fan![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=703,height=6\.83331pt\]\{all\-twemojis\.pdf\}\\includegraphics\[page=2327,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}, Jiechao Gao![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=715,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}, Jie Wang![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=715,height=6\.83331pt\]\{all\-twemojis\.pdf\}\},Minjia Zhang![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=703,height=6\.83331pt\]\{all\-twemojis\.pdf\}\},Srivatsan Ravi![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=2709,height=6\.83331pt\]\{all\-twemojis\.pdf\}\},Carlee Joe\-Wong![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=696,height=6\.83331pt\]\{all\-twemojis\.pdf\}\} ![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=2709,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}University of Southern California,![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=696,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}Carnegie Mellon University, ![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=703,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}University of Illinois Urbana\-Champaign,![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=715,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}Stanford University ![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=1958,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}yixiangy@usc\.edu,![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=3549,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}yuhangya@alumni\.cmu\.edu,![[Uncaptioned image]](https://arxiv.org/html/2605.19035v1/all-twemojis.pdf)\{\}^\{\\includegraphics\[page=2327,height=6\.83331pt\]\{all\-twemojis\.pdf\}\}xfan31@illinois\.edu
###### Abstract
The rapid advancement of Large Language Models has given rise to autonomous LLM\-based agents capable of complex reasoning and execution\. As these agents transition from isolated operation to collaborative ecosystems, we witness the emergence of the Agent\-to\-Agent \(A2A\) network, a paradigm where heterogeneous agents autonomously coordinate to solve multi\-step tasks\. While these networks may offer better task performance compared to simply using one agent to complete the entire task, they introduce systemic vulnerabilities, such as adversarial composition, semantic misalignment, and cascading operational failures, that existing agent alignment techniques cannot address\. In this vision paper, we argue that the trustworthiness of A2A networks cannot be fully guaranteed via retrofitting on existing protocols that are largely designed for individual agents\. Rather, it must be architected from the very beginning of the A2A coordination framework\. We present a comprehensive conceptual framework that situates trust in A2A systems through four design pillars\. We contour the trustworthy A2A network using these pillars and envision the future of this emerging paradigm\.
## 1Introduction
The rapid advancement of Large Language Models \(LLMs\) has transformed artificial intelligence into not just passive text generation systems but active problem\-solving entities\. By equipping LLMs with tools, memory, planning capabilities, and external interfaces, researchers have developed LLM\-based agents capable of reasoning, interacting with environments, and executing multi\-step workflowsSchicket al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib32)\); Yanget al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib15)\); Yaoet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib35)\)\. In this context, anagentrefers to an autonomous system that perceives context, formulates intermediate plans, invokes tools or APIs, and updates its internal state to achieve high\-level objectivesWooldridge and Jennings \([1995](https://arxiv.org/html/2605.19035#bib.bib14)\)\. Compared to standalone LLMs, agents offer increased autonomy, adaptability, and the ability to operate in open\-ended environments\. These agents are now applied to complex reasoning and coding tasks such as automated software engineering, bug detection and patching, data analysis pipelines, and mathematical problem solvingLewiset al\.\([2020](https://arxiv.org/html/2605.19035#bib.bib50)\); Yanget al\.\([2024](https://arxiv.org/html/2605.19035#bib.bib59)\)\. Beyond that, agentic systems are increasingly connected to physical and cyber\-physical domains, such as robotic manipulation, autonomous laboratory experimentation, and real\-time decision\-making in industrial control systems\. This shift marks a transition from static language models to dynamic, action\-oriented systems\.
However, as real\-world tasks grow in complexity, the limitations of single\-agent systems become increasingly apparent\. Many applications, such as automated software engineering, financial analysis, supply chain coordination, and policy planning, require diverse forms of expertise and parallel processing\. This has led to the emergence of multi\-agent systems and, more specifically, Agent\-to\-Agent \(A2A\) networks, where specialized agents collaborate autonomouslyWuet al\.\([2024](https://arxiv.org/html/2605.19035#bib.bib13)\); Liet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib12)\)\. In an A2A network, distinct agents \(e\.g\., a planner, a coder, a reviewer, and a deployer\) exchange information, delegate subtasks, and negotiate outcomes without continuous human supervision\. This modular structure mirrors human organizational systems and offers practical advantages: specialization improves performance, parallelism increases efficiency, and distributed design enhances scalability and fault tolerance\. A recent popular example of this emerging paradigm is OpenClawOpenClaw Team \([2026](https://arxiv.org/html/2605.19035#bib.bib62)\), an open ecosystem where heterogeneous agents dynamically discover and invoke shared capabilities via public registries like ClawHubOpenClaw Developers \([2026](https://arxiv.org/html/2605.19035#bib.bib85)\), and even coordinate autonomously on AI\-only platforms like MoltbookOpenClaw Community \([2026](https://arxiv.org/html/2605.19035#bib.bib86)\)\. As a result, A2A networks are becoming a natural and pragmatic evolution of agent\-based AI\.
Figure 1:Illustration of trust issues in agent\-to\-agent networks\.Red,green, andgrey“lobster”s represent benign, adversarial, and disconnected agents, respectively\.Yet, this architectural transition introduces a fundamental challenge:trustis a fundamental requirement of using A2A networks, particularly for tasks where failure could have serious reputational or safety consequences\. Trust, however, does not compose automatically across agentsCemriet al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib18)\); Hendryckset al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib19)\)\. Methods designed to align or secure a single agent do not guarantee the safety of a network of interacting agents\. In practice, current A2A systems often rely on patch\-like solutions, such as guardrails, post\-hoc verification, human\-in\-the\-loop approvals, protocol wrappers, or sandboxing, to mitigate risksDonget al\.\([2024b](https://arxiv.org/html/2605.19035#bib.bib63)\); Sandhu \([1998](https://arxiv.org/html/2605.19035#bib.bib22)\); Zhenget al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib31)\)\. These techniques improve local robustness, but they typically treat trust as an overlay rather than a system\-level invariant\. As observed in open environments like OpenClawOpenClaw Team \([2026](https://arxiv.org/html/2605.19035#bib.bib62)\), the dynamic composition of third\-party skills introduces compositional trust failures \(e\.g\., malicious skill poisoning, cascading tool\-chain exploits, and adversarial prompt injections\) that easily bypass local safeguards\. When multiple agents interact through natural language and shared state, new failure modes emerge that cannot be resolved by strengthening components independently\.
Several systemic vulnerabilities illustrate this problem\. First, cascading execution can amplify minor errors: a small hallucination in one agent’s output may propagate through downstream agents, leading to large\-scale operational failuresParket al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib37)\)\. Second, semantic misalignment arises when agents interpret shared instructions differently, producing globally unsafe outcomes even when each agent behaves “correctly” according to its local objectiveKieranset al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib10)\)\. Third, the adversarial composition allows malicious or malformed inputs to traverse benign agents and trigger harmful actions at privileged nodesGreshakeet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib21)\)\. These failures are compositional rather than intrinsic, as they emerge from interaction dynamics rather than isolated behavior\. Consequently, patch\-like solutions often detect problems only after unsafe trajectories have already become reachable\. For example, a local guardrail will fail when a “safety agent” and an “optimization agent” have slightly divergent definitions of a shared goal, which results in an interaction that produces a compliant but dangerous system state that isolated filters cannot detect\.
Solving these challenges does not mean just improving filters or adding additional monitors\. Because these “bolted\-on” safeguards act as external processes that attempt to detect or repair unsafe behavior only after it arises, they depend heavily on detection reliability, which introduce latency and resource overhead, and remain vulnerable to bypass or delayPerezet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib9)\); Schroeder and Wood\-Doughty \([2024](https://arxiv.org/html/2605.19035#bib.bib11)\)\. More fundamentally, the underlying transition function of the network, which is the core mechanism that dictates how the system’s global state evolves in response to agent actions like messages and tool calls, is left entirely unconstrained\. Because this foundational logic can still generate invalid configurations before any external monitor intervenes, unsafe global states remain theoretically reachable\. Therefore, addressing trust in A2A networks requires an architectural paradigm shift: rather than treating safety as an auxiliary feature to be retrofitted onto an unconstrained system, foundational trust guarantees must be embedded directly into the system’s design\.
In this paper, we argue for the latter\. We propose that trust in Agent\-to\-Agent networks must be baked in rather than bolted on\. Instead of treating trust as an attribute of individual agents or communication channels, we present a paradigm where trust is a strict rule built into the core architecture that the network can never enter an unsafe state, regardless of how agents interact\. We introduce the concept of a Trustworthy Agent Network \(TAN\), defined through four constitutive design pillars \(Compositional Robustness, Semantic Containment, Accountability, and Cross\-Boundary Reliability\) and evaluate the cost through operational metrics including inference latency, resource overhead, scalability, and determinism\. Our position is that only by embedding these principles into the transition dynamics of A2A systems can we mitigate systemic trust failures at scale\. The remainder of this paper develops this framework, analyzes existing approaches under its lens, and outlines a blueprint for architecting agent networks that are trustworthy by construction rather than patchwork\.
## 2Current Agent Systems Are Not Trustworthy
As agent systems evolve from isolated executors into open, language\-driven networks, trust failures increasingly arise not from individual agent behavior, but from how agents are composed, coordinated, and embedded within larger systems\. In such settings, failures manifest not only as incorrect outputs, but as cascading errors, semantic drift, responsibility diffusion, and adversarial composition across interacting agents\. We review prior approaches by where they attempt to introduce trust, at the level of individual agent behavior, workflow coordination, or system infrastructure \([Table˜1](https://arxiv.org/html/2605.19035#S2.T1);[Section˜2\.1](https://arxiv.org/html/2605.19035#S2.SS1)\)\. While these lines of work improve robustness locally, they do not establish trust as a system\-level invariant across the agent network\.
At a higher level, most proposals either attach external checks to an otherwise unconstrained system, i\.e\., bolted\-on trust mechanisms \([Section˜2\.3\.1](https://arxiv.org/html/2605.19035#S2.SS3.SSS1)\), or impose partial architectural constraints that still leave unsafe trajectories reachable, i\.e\., partially baked\-in system constraints \([Section˜2\.3\.2](https://arxiv.org/html/2605.19035#S2.SS3.SSS2)\)\.[Section˜2\.3](https://arxiv.org/html/2605.19035#S2.SS3)formalizes this distinction and clarifies why neither category suffices to guarantee trust at the network level\.
Table 1:Existing approaches to agent trust and their structural limitations\.### 2\.1Pros and Cons of Existing Approaches
Existing approaches to agent trust typically intervene after the core agent architecture is already in place, by tuning agent behavior, orchestrating workflows, or constraining communication/execution interfaces\. As summarized in[Table˜1](https://arxiv.org/html/2605.19035#S2.T1), these methods can substantially reduce local failure rates, but they rarely restrict the reachable system trajectories of a multi\-agent network\. In other words, trust is improved as an overlay \(via prompting, supervision, protocols, or isolation\), rather than enforced as a system\-level invariant of interaction dynamics\.
#### 2\.1\.1Single\-Agent Alignment and Self\-Regulation
This category improves trust by strengthening the node: the individual LLM agent\. It includes: \(i\) prompt\-level steering, such as prompt engineeringWhiteet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib48)\)and in\-context learningDonget al\.\([2024a](https://arxiv.org/html/2605.19035#bib.bib49)\); \(ii\) knowledge grounding such as RAGLewiset al\.\([2020](https://arxiv.org/html/2605.19035#bib.bib50)\)and memory\-augmented agentsZhanget al\.\([2025b](https://arxiv.org/html/2605.19035#bib.bib51)\); \(iii\) training\-based alignment such as SFT/parameter\-efficient tuningWanget al\.\([2025a](https://arxiv.org/html/2605.19035#bib.bib52)\), agentic RL and RLHF\-style alignmentZhanget al\.\([2025a](https://arxiv.org/html/2605.19035#bib.bib53)\), unlearningGenget al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib54)\), adversarial preference learningWanget al\.\([2025b](https://arxiv.org/html/2605.19035#bib.bib55)\), and multimodality alignmentTsaiet al\.\([2024](https://arxiv.org/html/2605.19035#bib.bib56)\); and \(iv\) agent\-internal control loops and tool\-augmented execution, such as ReflectionRenze and Guven \([2024](https://arxiv.org/html/2605.19035#bib.bib57)\), ReActYaoet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib35)\), Toolformer\-style tool useSchicket al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib32)\), skills abstractionsAnthropic \([2026](https://arxiv.org/html/2605.19035#bib.bib61)\), computer\-use agents \(e\.g\., SWE\-agent\)Yanget al\.\([2024](https://arxiv.org/html/2605.19035#bib.bib59)\), and integrated agent stacks such as OpenClawOpenClaw Team \([2026](https://arxiv.org/html/2605.19035#bib.bib62)\)\.
These methods are often efficient and scalable because they do not require heavy coordination overhead\. However, they rely on the implicit assumption that aligned agents compose into an aligned network\. Once agents interact via language, this assumption fails: semantic drift, amplification, and misinterpretation can arise even when each agent is locally aligned\. As a result, purely node\-level alignment reduces harmful outputs in isolation but cannot prevent interaction\-driven failures or guarantee network\-level trust\.
#### 2\.1\.2Multi\-Agent Coordination and Workflow Control
This category improves trust by regulating the workflow, how multiple agents are composed and how decisions are produced\. Representative mechanisms include guardrails and LLM\-as\-a\-judge monitoringDonget al\.\([2024b](https://arxiv.org/html/2605.19035#bib.bib63)\); Zhenget al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib31)\), structured role\-playShaoet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib64)\), deep\-research style iterative decompositionXu and Peng \([2025](https://arxiv.org/html/2605.19035#bib.bib60)\), explicit human\-agent\-computer interaction loopsLuet al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib65)\), training\-free or post\-hoc error correction loopsZhouet al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib66)\), and ensemble\-style decision aggregation such as votingKaesberget al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib67)\)and debateDuet al\.\([2024](https://arxiv.org/html/2605.19035#bib.bib5)\)\. Planner–executor architecturesLiet al\.\([2024](https://arxiv.org/html/2605.19035#bib.bib68)\)and supervisor frameworksBaoet al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib69)\)further introduce hierarchical control to reduce obvious coordination breakdowns\.
The main advantage is broader interaction awareness: these methods explicitly acknowledge that failures emerge from composition and attempt to correct them by oversight, redundancy, or structured orchestration\. The limitation is that they mostly operate as reactive governance: they detect, critique, or repair trajectories after candidate actions are generatedCemriet al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib18)\); Irvinget al\.\([2018](https://arxiv.org/html/2605.19035#bib.bib23)\)\. They regulate how decisions are aggregated, but do not guarantee that the underlying semantics remain safe under propagation across agents\. Because enforcement is probabilistic and monitor\-dependent, unsafe global states remain reachable when supervision fails, is delayed, or is bypassed\.
#### 2\.1\.3Protocol\-Centric Trust
Protocol\-centric approaches shift trust to the communication and identity layer\. They include standardized interfaces such as MCPAnthropic \([2024a](https://arxiv.org/html/2605.19035#bib.bib29)\)and tool\-calling / tool\-integrated reasoning protocolsLumeret al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib70)\), as well as identity and integrity mechanisms such as authenticated delegationSouthet al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib71)\), decentralized identifiers \(DID\)Reedet al\.\([2020](https://arxiv.org/html/2605.19035#bib.bib30)\), audit trails for accountabilityOjewaleet al\.\([2026](https://arxiv.org/html/2605.19035#bib.bib72)\), and secure A2A messaging protocolsHableret al\.\([2025](https://arxiv.org/html/2605.19035#bib.bib73)\)\. This category also includes defensive tooling such as malware detectionSommer and Paxson \([2010](https://arxiv.org/html/2605.19035#bib.bib28)\)and cryptographic methodsSuçeken and Özkaraca \([2024](https://arxiv.org/html/2605.19035#bib.bib74)\)that aim to harden the agent communication surface\.
The key strength is determinism at the syntax/identity layer: schema validation, signatures, and authorization checks can be rule\-based and scalableSandhu \([1998](https://arxiv.org/html/2605.19035#bib.bib22)\)\. The structural gap is semantic: protocol compliance can ensure a message is well\-formed and properly attributed, yet still allow semantically unsafe instructions, misaligned intent, or harmful downstream state transitionsGreshakeet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib21)\); Zouet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib20)\)\. Thus, these methods secure syntax, provenance, and access, but do not ensure semantic containment or network\-level invariants over meaningHendryckset al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib19)\)\.
#### 2\.1\.4Trust Environments and Execution Containment
Finally, environment\-centric approaches constrain the execution boundary via access control and isolation: role\-based access control \(RBAC\)Sandhu \([1998](https://arxiv.org/html/2605.19035#bib.bib22)\), sandboxing/fault isolationWahbeet al\.\([1993](https://arxiv.org/html/2605.19035#bib.bib24)\), scoped memory controlsBousetouane \([2026](https://arxiv.org/html/2605.19035#bib.bib75)\), and trusted execution environments \(TEE\)Costan and Devadas \([2016](https://arxiv.org/html/2605.19035#bib.bib25)\)\. These mechanisms can sharply reduce attack surface and limit privilege escalation by restricting what actions are possible and what information can be accessed\.
Their limitation is scope: they provide strong local containment but do not, by themselves, define global trust invariants over multi\-agent semantics\. Authorized agents can still execute semantically harmful plans within their allowed privileges, and isolation alone cannot guarantee that agent\-to\-agent intent remains consistentHubingeret al\.\([2024](https://arxiv.org/html/2605.19035#bib.bib17)\); Greenblattet al\.\([2024](https://arxiv.org/html/2605.19035#bib.bib16)\)\. Consequently, trust environments strengthen boundary safety, yet leave the core problem, semantic misalignment and unsafe global trajectories, largely unresolved\.
### 2\.2Why Existing Approaches Fail Systematically: The Bolted\-On Problem
[Table˜1](https://arxiv.org/html/2605.19035#S2.T1)exposes a consistent limitation across existing techniques: trust is treated as an attribute of individual agents, communication channels, or coordination heuristics rather than as an emergent and enforceable property of the agent network itself\.
As a result, current systems lack: \(i\) a formal notion of network\-level trust, \(ii\) guarantees at the semantic level of language\-based interaction, \(iii\) principled accountability and causal attribution across agents, and \(iv\) robustness across heterogeneous agents and trust boundaries\. Failures therefore manifest as cascading errors, responsibility diffusion, semantic leakage, or adversarial composition, failure modes that cannot be eliminated through incremental tuning or monitoring\.
These approaches reveal a common pattern: trust is added as a corrective layer rather than embedded as a defining property of the system’s dynamics\. Bolted\-on safeguards attempt to detect or repair unsafe behavior after it arises, whereas a trustworthy agent network requires guarantees that hold by construction\. This motivates a transition toward a baked\-in paradigm, in which trust is enforced intrinsically through the design of interaction rules, execution constraints, and semantic boundaries\.
### 2\.3Definition of Bolted\-On and Baked\-In Trust
To rigorously distinguish between “Bolted\-On” and “Baked\-In” trust, we model the agent network as a state transition system\. We also provide a representative example in[Figure˜2](https://arxiv.org/html/2605.19035#S2.F2)\.
Let𝒮\\mathcal\{S\}denote the global state space of all possible configurations of the agent network \(including valid and invalid states\), and let𝒮safe⊂𝒮\\mathcal\{S\}\_\{safe\}\\subset\\mathcal\{S\}represent the subset of states that satisfy safety invariants \(i\.e\., trust constraints\)\.𝒜\\mathcal\{A\}is the action space, representing the set of all possible inputs \(messages, tool calls\) an agent can generate\. Letδ:𝒮×𝒜→𝒮\\delta:\\mathcal\{S\}\\times\\mathcal\{A\}\\to\\mathcal\{S\}denote the transition function governing the agent network’s state evolution\. At timett, the function accepts the current statest∈𝒮s\_\{t\}\\in\\mathcal\{S\}and actionat∈𝒜a\_\{t\}\\in\\mathcal\{A\}, mapping them to the subsequent statest\+1s\_\{t\+1\}\.
#### 2\.3\.1Bolted\-On Trust \(Extrinsic Verification\)
In a bolted\-on architecture, the underlying transition functionδ\\deltaremains unconstrained and may generate unsafe states\. Trust is enforced by an external monitoring functionMMthat evaluates outcomes after they are produced:
st\+1=\{δ\(st,at\),ifM\(δ\(st,at\)\)∈𝒮safe,st,otherwise\.s\_\{t\+1\}=\\begin\{cases\}\\delta\(s\_\{t\},a\_\{t\}\),&\\text\{if \}M\(\\delta\(s\_\{t\},a\_\{t\}\)\)\\in\\mathcal\{S\}\_\{safe\},\\\\ s\_\{t\},&\\text\{otherwise\}\.\\end\{cases\}\(1\)
This mechanism implies that the safety depends on the reliability of the monitorMM\. The unsafe statest′∈𝒮∖𝒮safes\_\{t\}^\{\\prime\}\\in\\mathcal\{S\}\\setminus\\mathcal\{S\}\_\{safe\}becomes reachable viaδ\\deltaifMMproduces a false positive \(failure of detection, delayed, or bypassed\)\.
Figure 2:Bolted\-On v\.s\. Baked\-In
#### 2\.3\.2Baked\-In Trust \(Intrinsic Constraints\)
In a baked\-in architecture, the transition functionδ\\deltais defined such that all reachable states satisfy safety invariants:
∀st∈𝒮safe,∀at∈𝒜:st\+1=δ\(st,at\)∈𝒮safe\.\\forall s\_\{t\}\\in\\mathcal\{S\}\_\{safe\},\\forall a\_\{t\}\\in\\mathcal\{A\}:\\quad s\_\{t\+1\}=\\delta\(s\_\{t\},a\_\{t\}\)\\in\\mathcal\{S\}\_\{safe\}\.\(2\)
Rather than detecting violations after execution, baked\-in designs eliminate unsafe transitions from the system topology\. Any actionata\_\{t\}that would result in a state outside𝒮safe\\mathcal\{S\}\_\{safe\}is undefined inδ\\delta\. Thus, the unsafe state is unreachable by definition, not by inspection\.
This formal distinction clarifies why incremental safeguards are insufficient: without constraining the transition function itself, unsafe states remain reachable\. Having separated bolted\-on verification from baked\-in constraints, we now turn to the design principles required to enforce trust as a system\-level invariant\.
## 3Trustworthy Agent Network: Definition and Design Principles
As AI systems evolve from isolated tools into collaborative swarms, reliance on individual agent alignment is no longer sufficient\. We define a Trustworthy Agent Network \(TAN\) not merely as a collection of aligned agents, but as a resilient infrastructure that guarantees trust properties at the interaction level\.
To structure this definition, we propose a two\-tier framework that connects the "Why" and "What" of trustworthy networks\. As for "how", it is "looked ahead" in[Section˜5](https://arxiv.org/html/2605.19035#S5)\. As illustrated in[Figure˜3](https://arxiv.org/html/2605.19035#S3.F3), our framework is composed of two interconnected layers:
- •Conceptual Necessity \([Section˜3\.1](https://arxiv.org/html/2605.19035#S3.SS1)\): This layer identifies the specific risks inherent to multi\-agent systems, such as semantic misalignment and operational failure, which drive the necessity for a new trust paradigm\.
- •Agent Network Trust \([Section˜3\.2](https://arxiv.org/html/2605.19035#S3.SS2)\): This layer defines what specific properties must theoretically hold to mitigate those risks\.
Additionally, we employ additional evaluation metrics in[Section˜3\.3](https://arxiv.org/html/2605.19035#S3.SS3)to quantify how a method performs under the TAN framework\.
Figure 3:From vulnerabilities to trust requirements in agent networks\. Left: multi\-agent risks\. Right: properties required for trustworthy operation\.### 3\.1Conceptual Necessity: Why Agent Network Trust Matters
The shift from single\-agent to multi\-agent systems introduces risks that arecompositionalrather than intrinsic\. Formally, we define an agent network as a tuple𝒩=⟨𝒜,Σ⟩\\mathcal\{N\}=\\langle\\mathcal\{A\},\\Sigma\\rangle, where𝒜=\{a1,…,an\}\\mathcal\{A\}=\\\{a\_\{1\},\\dots,a\_\{n\}\\\}is the set of agents andΣ\\Sigmais the global state space\. LetΦ:Σ→\{0,1\}\\Phi:\\Sigma\\to\\\{0,1\\\}be the global safety predicate, whereΦ\(s\)=1\\Phi\(s\)=1denotes a safe state andΦ\(s\)=0\\Phi\(s\)=0denotes a violation\.
The system dynamics are governed by the interaction between agent outputs and the environment:
- •Agent Execution: An agentai∈𝒜a\_\{i\}\\in\\mathcal\{A\}receives an inputx∈ℝx\\in\\mathbb\{R\}and generates an outputy=ai\(x\)y=a\_\{i\}\(x\)wherey∈ℝy\\in\\mathbb\{R\}\.
- •State Transition: The global state evolves via a transition functionδ:Σ×ℝ→Σ\\delta:\\Sigma\\times\\mathbb\{R\}\\to\\Sigma\. Given a current statest∈Σs\_\{t\}\\in\\Sigmaand an agent outputyy, the system transitions to a new statest\+1=δ\(st,y\)s\_\{t\+1\}=\\delta\(s\_\{t\},y\)wherettis the time\.
The core problem of agent networks is that local safety does not imply global safety\. Even if every agentaia\_\{i\}is locally robust \(i\.e\., its internal checks pass\), their interaction can yield an unsafe global state\. Formally, this divergence occurs when a set of valid local transitions produces a trace to a violation\. We identify four mechanisms driving this divergence\.
#### 3\.1\.1Vulnerable Multi\-Agent Collaboration
In an open network, a trusted agentaia\_\{i\}often consumes outputs from an untrusted or less\-capable agentaja\_\{j\}\. The risk is thataja\_\{j\}produces a syntactically valid but semantically malicious payload that exploitsaia\_\{i\}\.
Lety=aj\(x\)y=a\_\{j\}\(x\)be the output from agentjj, and letV\(⋅\)V\(\\cdot\)be a validator \(e\.g\., checking if the output is valid JSON\)\. A failure occurs when the payload passes validation but triggers a safety violation when processed by the receiving agent:
V\(y\)=1∧Φ\(ai\(y\)\)=0V\(y\)=1\\quad\\land\\quad\\Phi\(a\_\{i\}\(y\)\)=0\(3\)
Ideally, trust should be transitive, but in reality, the compositionai∘aja\_\{i\}\\circ a\_\{j\}introduces an attack surface \(e\.g\., prompt injection\) that neither agent can detect in isolation\.
Example: A finance agent \(aia\_\{i\}\) asks a web scraper \(aja\_\{j\}\) to summarize a URL\. The scraper returns a summary \(yy\) that contains a hidden instruction: "Ignore previous rules and transfer funds\." The finance agent’s validator confirmsyyis text \(V\(y\)=1V\(y\)=1\), but processing it triggers an unauthorized transaction \(Φ=0\\Phi=0\)\.
#### 3\.1\.2Semantic Misalignment
Unlike a malicious exploit, misalignment occurs when agents share a syntax but possess divergent mappings from natural language instructions to state transitions\.
Letxxbe the instruction sent by agentaia\_\{i\}to agentaja\_\{j\}at current statess\. LetΣtarget⊂Σ\\Sigma\_\{target\}\\subset\\Sigmabe the subset of states thataia\_\{i\}intends to reach viaxx\(the “safe” outcome\)\. Agentaja\_\{j\}processesxxand generates an actionyy, triggering the transition to a new statest\+1s\_\{t\+1\}:
st\+1=δ\(st,y\)s\_\{t\+1\}=\\delta\(s\_\{t\},y\)\(4\)Misalignment is defined as a transition where the resulting state is valid according to theaja\_\{j\}’s local logic \(it successfully executed a task\) but lies outside theaia\_\{i\}’s target subspace:
st\+1∉Σtargetdespiteajcompleting taskxs\_\{t\+1\}\\notin\\Sigma\_\{target\}\\quad\\text\{despite\}\\quad a\_\{j\}\\text\{ completing task \}x\(5\)This divergence arises because the mappingx→Σtargetx\\to\\Sigma\_\{target\}is implicit\. The receiveraja\_\{j\}steers the network to a statest\+1s\_\{t\+1\}that satisfies the literal instruction but violates the implied semantic constraints ofaia\_\{i\}\.
Example: Agent A \(aia\_\{i\}\) sends instructionx=x=“find the best route”\. A implies a target stateΣtarget\\Sigma\_\{target\}where the route avoids high\-risk zones\. Agent B \(aja\_\{j\}\) interprets the best route as the shortest path, transitioning the system state tost\+1s\_\{t\+1\}\. Therefore, the path thatst\+1s\_\{t\+1\}represented passes through a conflict zone, makingst\+1∉Σtargets\_\{t\+1\}\\notin\\Sigma\_\{target\}\.
#### 3\.1\.3Ethical & Legal Violations
Privacy and copyright safety are often non\-compositional\. A global states∈Σs\\in\\Sigmamay be unsafe even if the individual state transitions leading to it were locally verified\.
Lety1y\_\{1\}andy2y\_\{2\}be outputs from agentsa1a\_\{1\}anda2a\_\{2\}, respectively\. Lets0s\_\{0\}be the initial safe state\. First, the system transitions viay1y\_\{1\}to states1=δ\(s0,y1\)s\_\{1\}=\\delta\(s\_\{0\},y\_\{1\}\)\. Second, the system transitions viay2y\_\{2\}to states2=δ\(s1,y2\)s\_\{2\}=\\delta\(s\_\{1\},y\_\{2\}\)\.
A potential violation occurs when the intermediate states1s\_\{1\}is safe, but the aggregated states2s\_\{2\}violates the global safety predicateΦ\\Phi:
Φ\(s1\)=1butΦ\(s2\)=0\\Phi\(s\_\{1\}\)=1\\quad\\text\{but\}\\quad\\Phi\(s\_\{2\}\)=0\(6\)
Example: Agenta1a\_\{1\}updates the state with anonymized medical records \(y1y\_\{1\}\)\. The resulting states1s\_\{1\}is safe\. Agenta2a\_\{2\}updates the state with public voter registrations with names and birth dates \(y2y\_\{2\}\)\. The final states2s\_\{2\}now contains enough correlated data to re\-identify patients, triggering a violation \(Φ\(s2\)=0\\Phi\(s\_\{2\}\)=0\)\.
#### 3\.1\.4Operational Failure
Complexity in agent networks can lead to unstable dynamics where the state trajectory consumes infinite resources without reaching a valid terminal state\.
LetΣtarget⊂Σ\\Sigma\_\{target\}\\subset\\Sigmabe the subset of valid completion states\. LetR\(t\)R\(t\)be the cumulative resource cost at timett\. An operational failure occurs when:
limt→∞R\(t\)=∞while∀t,st∉Σtarget\\lim\_\{t\\to\\infty\}R\(t\)=\\infty\\quad\\text\{while\}\\quad\\forall t,s\_\{t\}\\notin\\Sigma\_\{target\}\(7\)This signifies a divergence where the system actively generates outputsyty\_\{t\}\(consuming compute\) but the resulting state trajectory never intersects with the target subspace\.
Example: A coder agent \(aia\_\{i\}\) writes a scriptyiy\_\{i\}that fails a test\. The tester agent \(aja\_\{j\}\) reports the erroryjy\_\{j\}\. The coder, lacking the context to fix the root cause, makes a cosmetic change and resubmits\. The state trajectory cycles infinitely \(st→st\+1→…s\_\{t\}\\to s\_\{t\+1\}\\to\\dots\), consuming API credits \(R→∞R\\to\\infty\) without ever entering the success state \(Σtarget\\Sigma\_\{target\}\)\.
### 3\.2Agent Network Trust: What Must Hold
To mitigate the potential risks, a Trustworthy Agent Network must satisfy four core pillars of trust\. These are not optional features but constitutive requirements that constrain the global state space\.
#### 3\.2\.1Compositional Robustness
To solve the issue of vulnerable multi\-agent collaboration where a validatorV\(y\)=1V\(y\)=1fails to detect a malicious state transitionΦ\(δ\(st,y\)\)=0\\Phi\(\\delta\(s\_\{t\},y\)\)=0, the system must enforce that safety is a structural guarantee of the transition function itself, not merely a quality of the agent’s output\.
For any agentaja\_\{j\}designated as untrusted, the network must guarantee that the set of all reachable states viaδ\\deltais strictly contained within the safe subspace ofΣ\\Sigma\. The necessary condition is that for any arbitrary payloadyygenerated byaja\_\{j\}, the resulting state transition preserves the global safety predicate:
∀y,Φ\(δ\(st,y\)\)=1\\forall y,\\quad\\Phi\(\\delta\(s\_\{t\},y\)\)=1\(8\)This condition asserts that the system’s safety is decoupled from the semantic content ofyy\. Even ifyycontains prompt injection,δ\\deltamust be bounded such that it is mathematically impossible for such a payload to mutate state or trigger a violation\.
#### 3\.2\.2Semantic Containment
Semantic Containment aims the control the consistency among individual agents\. To solve the misalignment issue wherest\+1∉Σtargets\_\{t\+1\}\\notin\\Sigma\_\{target\}, the system must enforce that the alignment between agents is explicitly constructed\.
We define consistency𝒞\\mathcal\{C\}as the verifiable alignment between the sender’s intentxxand the receiver’s actionyy\. For trust to hold, the network must guarantee that the consistency of the pair\(x,y\)\(x,y\)implies safety in the state transition:
𝒞\(x,y\)=1⟹δ\(st,y\)∈Σtarget\\mathcal\{C\}\(x,y\)=1\\implies\\delta\(s\_\{t\},y\)\\in\\Sigma\_\{target\}\(9\)This condition asserts that the receiveraja\_\{j\}is not free to optimize solely for its local utility\. Instead, its outputyymust be semantically bound to the constraints ofxxsuch that the receiver’s action preserves the sender’s intent regarding the same input\.
#### 3\.2\.3Accountability & Attributability
Accountability defines the property that every states∈Σs\\in\\Sigmamust encode its own provenance\. To solve the issue of ethical and legal violations, the system must enforce causal traceability\.
Specifically, the global state spaceΣ\\Sigmamust be encoded to the history of agent interactions\. A mapping𝒯\\mathcal\{T\}should be defined to trace any statests\_\{t\}back to the set of unique agents that contributed to its current value\. The necessary condition is that for any unsafe state, the set of contributing agents is non\-empty and uniquely identifiable:
∀st∈Σ,Φ\(st\)=0⟹𝒯\(st\)≠∅\\forall s\_\{t\}\\in\\Sigma,\\quad\\Phi\(s\_\{t\}\)=0\\implies\\mathcal\{T\}\(s\_\{t\}\)\\neq\\emptyset\(10\)This property ensures that no safety violation can emerge anonymously\. The global state must intrinsically encode the identity of the agents responsible for every transitionδ\\deltasuch that liability for a privacy or copyright breach is guaranteed\.
#### 3\.2\.4Cross\-Boundary Reliability
Cross\-Boundary Reliability defines the property of liveness and termination\. To solve the issue of infinite resource consumption \(R\(t\)→∞R\(t\)\\to\\infty\), the system dynamics must be strictly convergent\.
For any interaction sequence initiated at states0s\_\{0\}, there must exist a maximum resource budgetRmax<∞R\_\{max\}<\\infty\. The condition that must hold is that for all valid trajectories, the cumulative resource costR\(t\)R\(t\)never exceeds this limit without the system reaching a terminal state:
∀t,R\(t\)≤Rmax⟹st∈Σtarget∨Terminated\(st\)\\forall t,R\(t\)\\leq R\_\{max\}\\implies s\_\{t\}\\in\\Sigma\_\{target\}\\lor\\text\{Terminated\}\(s\_\{t\}\)\(11\)
This implies that no loop or deadlock can persist indefinitely\. Every state transition consumes a portion of the finite budget, and the process eventually halts either at success or a safe failure state\.
### 3\.3Evaluation Metrics: Beyond Functional Pillars
While the four design pillars define what a trustworthy network must achieve, they do not quantify the cost of achieving it\. A “Bolted\-On” solution \(e\.g\., an external LLM\-as\-a\-Judge\) might theoretically satisfy all functional pillars but fail in practice due to prohibitive latency or computational expense\.
To rigorously evaluate existing methods against our proposed architecture, we introduce three classes of operational metrics: efficiency, scalability, and determinism\.
EfficiencyThe efficiency is depicted as the additional resource consumption required strictly for safety enforcement\. We use two metrics to quantify it\.
- •Inference Latency \(ElE\_\{l\}\): The temporal delay introduced by the verification mechanism before a transaction is finalized\. El=Tactual−TtaskTtaskE\_\{l\}=\\frac\{T\_\{actual\}\-T\_\{task\}\}\{T\_\{task\}\}\(12\)
- •Resource Overhead \(ErE\_\{r\}\): The ratio of resource, including tokens, memory, network bandwidth, storage, and computation cost, consumed for verification versus the actual task\. Et=Cactual−CtaskCtaskE\_\{t\}=\\frac\{C\_\{actual\}\-C\_\{task\}\}\{C\_\{task\}\}\(13\)
ScalabilityAs the network grows, the complexity of the trust mechanism must not explode\. We evaluate the scalabilityEsE\_\{s\}via asymptotic complexity \(BigOOnotation\) of the safety layer relative to the number of agents, interactions and so on\.
DeterminismSafety mechanisms must be consistent\. A probabilistic safety check is a vulnerability\. We define Determinism Score \(EdE\_\{d\}\), which is the probability that the safety mechanism \(SS\) returns the identical verdict for identical inputsxxacrosskktrials\.
Ed=P\(S\(xi\)=S\(xj\)\)∀i,j∈\[1,k\]E\_\{d\}=P\(S\(x\_\{i\}\)=S\(x\_\{j\}\)\)\\quad\\forall i,j\\in\[1,k\]\(14\)
Typically, if we evaluate and compare the bolted\-on and baked\-in TAN, we should have something similar to what’s in[Table˜2](https://arxiv.org/html/2605.19035#S3.T2)\.
Table 2:Comparison of bolted\-on and baked\-in TAN via evaluation metrics
## 4Analysis of Existing Techniques with TAN
Table 3:Evaluation of A2A network with TAN framework\. For four design pillar fulfillment: ●= Fully fulfills; ◐= Partially fulfills; ○= Not fulfilled\. For operational metrics: ●= High \(Good / Low overhead / Strong\); ◐= Medium; ○= Low \(Poor / Expensive\); – = N/A\.Table[3](https://arxiv.org/html/2605.19035#S4.T3)evaluates representative approaches to agent trust under the TAN framework\. The results show that existing techniques tend to concentrate on specific layers of the system, workflow orchestration, communication protocols, or execution environments, rather than enforcing trust as a global invariant of the agent network\. When examined through the four design pillars and operational metrics, each category exhibits clear strengths, yet also yield structural issues\.
### 4\.1Single\-Agent Alignment: Strong Local Robustness, Weak Compositional Guarantees
Single\-agent alignment methods focus on improving the internal behavior of individual agentsBaiet al\.\([2022](https://arxiv.org/html/2605.19035#bib.bib26)\); Ouyanget al\.\([2022](https://arxiv.org/html/2605.19035#bib.bib27)\)\. Within a multi\-agent setting, these techniques partially improve compositional robustness: better\-aligned agents are less likely to produce harmful outputs, and reinforcement learning or adversarial training can increase resistance to perturbationsMadryet al\.\([2017](https://arxiv.org/html/2605.19035#bib.bib8)\)\. Operationally, they are generally efficient and scalable\. Most introduce limited additional inference latency and moderate resource overhead, and they scale naturally as more agents are added\.
However, their trust assumptions remain local\. Alignment does not enforce semantic containment across agentsParket al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib37)\)\. Even if each agent is individually aligned, their interaction can still produce misaligned global states due to divergent interpretations of intent\. Accountability is also external to the alignment mechanism since these approaches do not encode provenance into state transitions\. Cross\-boundary reliability remains heuristic, as internal control loops and planning strategies may reduce looping behavior but do not guarantee bounded execution\. In this sense, single\-agent alignment exemplifies a bolted\-on trust philosophy at the model level: trust is encouraged through external tools or prompting rather than enforced through constraints on the internal transition itselfBrownet al\.\([2020](https://arxiv.org/html/2605.19035#bib.bib3)\)\. Consequently, safety does not reliably compose in networked environmentsAmodeiet al\.\([2016](https://arxiv.org/html/2605.19035#bib.bib7)\)\.
### 4\.2Multi\-Agent Workflow Coordination: Broader Interaction Awareness, Yet Reactive and Non\-Deterministic
Multi\-agent workflow coordination techniques explicitly address interaction dynamics\. Compared to single\-agent alignment, these approaches acknowledge that failures arise from composition\. They introduce additional layers of reasoning or oversight to regulate how decisions are aggregated and executed\. For example, planner–executor structures impose task decomposition, and supervisors distribute responsibility across rolesShenet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib6)\); Yaoet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib35)\)\. Human\-agent interaction can significantly strengthen accountability when approvals and logs attribute actions to identifiable actorsChristianoet al\.\([2017](https://arxiv.org/html/2605.19035#bib.bib2)\)\.
Despite these advances, most coordination mechanisms remain reactive\. Guardrails operate as monitors layered on top of unconstrained transitions, meaning unsafe states remain reachable if detection fails\. Voting and debate increase resource overhead and inference latency, often substantially, and their outcomes depend on probabilistic model reasoning, which somewhat limits determinismDuet al\.\([2024](https://arxiv.org/html/2605.19035#bib.bib5)\)\. While some architectures partially improve compositional robustness or cross\-boundary reliability, they do not embed invariants fully into the transitionPerezet al\.\([2023](https://arxiv.org/html/2605.19035#bib.bib9)\)\. As a result, coordination methods extend trust beyond the single\-agent level but largely retain a bolted\-on character at the workflow layer, improving empirical robustness without eliminating structural vulnerability\.
### 4\.3Protocol\-Centric Trust: Deterministic, but with Semantic Gaps and Variable Overhead
Protocol\-centric trust mechanisms shift trust to the communication and identity layer\. Their primary strength lies in determinism: schema validation, identity verification, and cryptographic checks are rule\-based rather than probabilistic, resulting in strong consistency \(EdE\_\{d\}\) and generally favorable scalability \(EsE\_\{s\}\)Anthropic \([2024b](https://arxiv.org/html/2605.19035#bib.bib36)\)\. Lightweight protocol enforcement, such as structured message validation or signature checks, typically introduces minimal inference latency and resource overhead\. However, more advanced cryptographic techniques, such as secure multi\-party computationGoldreich \([1998](https://arxiv.org/html/2605.19035#bib.bib46)\); Du and Atallah \([2001](https://arxiv.org/html/2605.19035#bib.bib79)\), zero\-knowledge proofsGoldreich and Oren \([1994](https://arxiv.org/html/2605.19035#bib.bib81)\); Sunet al\.\([2021](https://arxiv.org/html/2605.19035#bib.bib80)\), homomorphic encryptionAcaret al\.\([2018](https://arxiv.org/html/2605.19035#bib.bib76)\); Brakerskiet al\.\([2014](https://arxiv.org/html/2605.19035#bib.bib78)\); Cheonet al\.\([2017](https://arxiv.org/html/2605.19035#bib.bib77)\), or blockchain\-backed verificationSayeedet al\.\([2020](https://arxiv.org/html/2605.19035#bib.bib82)\); Gadekalluet al\.\([2022](https://arxiv.org/html/2605.19035#bib.bib83)\), can be computationally and communication intensive, substantially increasing both latency \(ElE\_\{l\}\) and resource overhead \(ErE\_\{r\}\)\. Thus, while protocol\-centric approaches are structurally deterministic, their efficiency depends on the strength and complexity of the guarantees employed\.
Functionally, these mechanisms partially constrain the transition function by restricting who may act and how messages must be formed, thereby strengthening compositional robustness at the syntactic level and improving accountability through traceability\. Compared to monitoring\-based coordination methods, they move closer to a baked\-in paradigm because constraints are embedded directly into interaction rules rather than applied retrospectively\. Nevertheless, protocol compliance secures syntax and identity, not semantics\. A well\-formed and authenticated message can still induce a semantically unsafe state transition\. Consequently, although protocol\-centric trust offers deterministic structural safeguards, it does not enforce semantic containment, leaving unsafe global trajectories reachable despite correct communication\.
### 4\.4Trust Environments: Strong Execution Containment, Limited Global Semantics
These mechanisms offer strong determinism and low operational overheadCostan and Devadas \([2016](https://arxiv.org/html/2605.19035#bib.bib25)\); Sabtet al\.\([2015](https://arxiv.org/html/2605.19035#bib.bib4)\)\. Access checks, isolation policies, and enclave verification are typically constant\-time operations that scale efficiently\. Accountability is strengthened because actions can be attributed to roles or isolated execution contexts\. Sandboxing and TEE can also impose resource caps, contributing to cross\-boundary reliability by preventing unbounded execution or privilege escalationCostan and Devadas \([2016](https://arxiv.org/html/2605.19035#bib.bib25)\); Wahbeet al\.\([1993](https://arxiv.org/html/2605.19035#bib.bib24)\)\. Among existing categories, trust environments represent some of the strongest structural constraints\.
Yet their scope is primarily infrastructural\. Similar to protocol\-centric methods, they prevent unauthorized transitions but do not prevent authorized agents from making semantically harmful decisions\. For instance, scoped memory reduces information leakage, and RBAC enforces capability boundaries, but neither guarantees that agents interpret each other’s instructions consistentlyDenning \([1976](https://arxiv.org/html/2605.19035#bib.bib1)\); Sandhu \([1998](https://arxiv.org/html/2605.19035#bib.bib22)\)\. These approaches embed constraints at the execution layer, i\.e\., a more baked\-in strategy than monitoring systems, but they do not encode global semantic invariants into the transition\. Consequently, they mitigate certain classes of failures without addressing semantic misalignment across agents\.
### 4\.5Key Takeaways
The evaluation of existing techniques leads to several key observations:
- •Strengthening individual agents improves robustness but does not guarantee safe global behavior in multi\-agent settings\.
- •Multi\-agent orchestration reduces obvious failures but remains probabilistic and bolted\-on\.
- •Protocol and environment efficiently enforce syntactic correctness, identity, and execution boundaries, yet do not constrain intent alignment\.
- •Across categories, trust is rarely embedded directly into the transition function\. Instead, safeguards are layered on top of unconstrained components\.
In conclusion, these findings indicate that existing methods address fragments of the trust problem but do not transform trust into a system\-level invariant\. This structural gap motivates the next step, that is, articulating a blueprint for a TAN framework in which trust is architected from the outset\.
## 5Blueprint for TAN
Figure 4:Blueprint for trustworthy agent networksA blueprint for a TAN must move beyond incremental safeguards and re\-architect trust as an intrinsic property of the network\. This section outlines how such a shift could address current limitations, what challenges arise in retrofitting existing methods, and which unexplored directions worth further effort\.
### 5\.1Embedding Compositional Robustness into the Transition Function
Current alignment and coordination methods assume that safety composes across agents\. The evaluation shows this assumption does not hold\. Even well\-aligned agents, when interacting, can produce unsafe global states\. To fix this within existing paradigms would require constraining the transition functionδ\\deltaitself rather than merely improving individual outputs\. Embedding compositional robustness requires combining alignment techniques with capability\-restricted action schema, typed interaction protocols, or formally defined state machines that eliminate unsafe transitions by construction\. However, imposing such constraints on open\-ended LLM agents is challenging because their action spaces are implicitly defined through natural language rather than explicit formal interfaces\.
Future TAN architectures must therefore formalize the global state space and allowable transitions more explicitly\. Capability\-based APIs, typed intent representations, or formally verifiable orchestration layers could restrict reachable states without excessive runtime monitoring\. The central challenge is preserving expressivity while embedding constraints: overly rigid transition systems limit agent autonomy, whereas overly flexible systems reintroduce unsafe trajectories\.
### 5\.2Achieving Semantic Containment Beyond Syntax
Protocol\-centric and environment\-based systems secure syntax and identity but do not guarantee that a receiver’s action preserves the sender’s intended target state within the global state space\. Introducing semantic validation layers such as intent classifiers or post\-hoc consistency checkers does not alter the underlying transition functionδ\\delta\. Instead, it instantiates the bolted\-on verification pattern defined in[Section˜2\.3\.1](https://arxiv.org/html/2605.19035#S2.SS3.SSS1), where safety depends on external monitoring and states outside𝒮safe\\mathcal\{S\}\_\{safe\}remain reachable if monitoring fails\.
A TAN blueprint must instead encode semantic constraints directly into interaction contracts\. This could involve explicit intent schemas, shared world models with invariant specifications, or constraint\-aware planning mechanisms where each state update is validated against formally defined semantic conditions\. Another promising direction is the development of semantic type systems that bind permissible actions to intended outcomes\. Unlike syntactic validation, semantic containment requires reasoning about intent and consequence, which risks increasing latency and resource cost unless carefully engineered\.
### 5\.3Intrinsic Accountability Through State\-Level Provenance
Existing systems strengthen accountability through logging, authentication, or audit trails, yet these mechanisms remain external to the state representation\. TAN requires that accountability be intrinsic: every state must encode its own causal provenance\. Relying on conventional logging mechanisms is insufficient, as logs remain external to the transition dynamics and do not enforce state\-level provenance\.
A forward\-looking blueprint should explore embedding lightweight provenance metadata directly into state updates\. Cryptographic commitments, incremental hashing, or structured causal graphs could ensure that any unsafe state can be traced to responsible agents without excessive overhead\. However, strong cryptographic guarantees can be computationally heavy, particularly in high\-frequency interaction settings\. The design challenge is therefore to achieve high determinism and strong attribution while maintaining acceptable latency and resource consumption\. Scalable provenance encoding must avoid super\-linear complexity as interactions increase\.
### 5\.4Ensuring Cross\-Boundary Reliability and Bounded Execution
Multi\-agent systems are particularly vulnerable to cascading loops and unbounded resource consumption\. Existing planner–executor structures, supervisors, sandboxing, and RBAC partially mitigate such risks, yet none guarantee bounded global trajectories\. Retrofitting termination guarantees into open\-ended language\-driven systems is inherently difficult, as generative models are not naturally finite\-state\.
A TAN architecture should incorporate explicit resource budgeting into the transition function\. Each state transition could consume a quantifiable portion of a global resource allowance, ensuring convergence to either a valid target state or a safe termination state\. Formal liveness guarantees, bounded\-depth planning trees, or checkpoint\-based execution contracts could be leveraged\. The challenge is balancing reliability with flexibility: strict bounds may constrain problem\-solving capability, while loose bounds risk instability\. Operationally, such mechanisms must remain scalable and deterministic without introducing excessive runtime overhead\.
### 5\.5Toward Fully Baked\-In Trust
The cumulative lesson from existing methods is that trust cannot remain an auxiliary layer\. Alignment improves individual agents, orchestration regulates workflows, protocols secure communication, and environments constrain execution\. However, unsafe states remain reachable because semantic invariants are not embedded into the transition dynamics\. Establishing a TAN therefore requires co\-designing semantics, protocols, and execution constraints so that safety is enforced within the transition function rather than through post\-hoc inspection\.
Achieving this objective entails integrating formal state representations, intent\-aware interaction contracts, intrinsic provenance encoding, and bounded execution policies into a unified architecture\. These properties impose distinct but complementary constraints on the transition functionδ\\delta: compositional robustness restricts reachable state topology; semantic containment constrains the meaning and effect of actions; intrinsic accountability attaches causal provenance to state updates; and cross\-boundary reliability bounds trajectories in time and resource space\. Because these constraints operate on the same transition system, they cannot be introduced independently without reintroducing reachability gaps\. A coherent TAN architecture must therefore defineδ\\deltasuch that safety, semantics, attribution, and liveness are enforced as coupled invariants\.
The principal design challenge lies in embedding these constraints while maintaining efficiency and scalability\. Excessively rigid transition systems risk limiting expressivity, whereas overly permissive dynamics undermine safety guarantees\. Trust mechanisms must therefore remain deterministic where feasible, compositional across agents, and computationally tractable in large\-scale networks\.
## 6Alternative Views
### 6\.1Agent\-Level Alignment Is Sufficient
The argument:If individual agents are sufficiently aligned through methods such as RLHF or large\-scale supervised fine\-tuning, then safety will naturally emerge at the network level\. Under this view, the problem reduces to improving model alignment\. Some proponents go further and suggest that building a single, highly capable “super\-agent” would avoid the complexity of multi\-agent interaction altogether\.
Our rebuttal:This view overlooks several structural realities\. First, training feasibility remains a major constraint\. It is extremely difficult to train a single agent that achieves state\-of\-the\-art performance across all domains, such as software engineering, legal reasoning, planning, and creative writing\. In practice, multi\-agent systems arise precisely because specialization is necessary\.
Second, even a perfectly aligned agent cannot remove computational bottlenecks\. A monolithic agent processes tasks sequentially, whereas networked systems allow parallel specialization\. As task complexity grows, parallel coordination becomes an architectural requirement rather than a convenience\.
Most importantly, safety does not compose automatically\. An individually aligned agent can still participate in globally unsafe behavior when interacting with others\. A reviewer may approve code that appears locally correct but becomes dangerous in a specific deployment configuration\. Alignment at the node level does not resolve semantic conflicts at the system level\. For these reasons, agent\-level alignment remains insufficient as a foundation for network trust\.
### 6\.2Measuring Trust Would Solve the Problem
The argument:Trust can be treated as a measurable quantity\. By assigning each agent a numerical trust score based on past behavior, the system can grant or restrict privileges using simple thresholds\. This makes trust computable and easy to operate\.
Our rebuttal:Reducing trust to a single scalar obscures its context\-dependent nature\. Trust is not one\-dimensional\. An agent may be reliable for data analysis but unsuitable for system administration\. Collapsing these distinctions into a single score risks granting authority in domains where competence has never been demonstrated\.
Moreover, once trust becomes a score, it becomes an optimization target\. Agents or adversaries controlling them can accumulate reputation through low\-risk actions and later exploit that accumulated credibility to perform a high\-impact harmful action\. This behavior exposes a fundamental weakness of threshold\-based trust models\. Quantification alone does not embed structural guarantees, and it merely shifts the problem into a different form\.
### 6\.3Ensemble Algorithms Ensure Reliability
The argument:Reliability can be achieved through ensemble reasoning among agents\. Mechanisms such as majority voting, debate, or chain\-of\-thought with cross\-examination naturally mitigate errors\. If multiple agents evaluate each other’s reasoning, hallucinations and unsafe decisions will be filtered out without the need for architectural constraints\.
Our rebuttal:This argument relies on an assumption of independence that rarely holds in practice\. When multiple agents share the same base model or training data, their errors are often correlated\. They may converge on the same plausible but incorrect conclusion so that it leads majority voting ineffective\.
Additionally, debate dynamics among language models can exhibit conformity effects\. Agents may shift their responses toward perceived consensus or authority rather than independently verifying correctness\. Without structural safeguards, such as isolation of reasoning traces or blind evaluation, ensemble mechanisms can reinforce shared biases rather than eliminate them\. Ensemble reasoning can improve empirical performance, but it does not guarantee structural robustness or semantic consistency\.
### 6\.4Existing Cryptographic Protocols Already Solve Trust
The argument:Trust in agent networks can be ensured through established cryptographic mechanisms such as encryption, digital signatures, and public key infrastructures\. If agents communicate over secure channels with verified identities, the network should be trustworthy\.
Our rebuttal:This perspective conflates communication security with semantic trustworthiness\. Cryptographic protocols secure the channel and verify the sender’s identity, but they do not evaluate the meaning or consequences of the transmitted message\. A message can be encrypted, authenticated, and fully compliant with protocol rules, yet still contain a hallucinated fact, a malicious instruction, or a semantically misaligned directive\.
In other words, cryptography secures the pipe, not the payload’s intent\. Trust challenges in agent networks arise at the level of meaning and system, not merely at the level of bits and bytes\. Ensuring that messages are authentic does not ensure that they are safe in context\. Therefore, while cryptographic protocols are essential for secure communication, they do not by themselves establish trustworthy multi\-agent behavior\.
## 7Conclusion
The main contribution of this work is to show that trust in Agent\-to\-Agent systems is fundamentally an architectural problem rather than a purely supervisory one\. Existing methods remain largely bolted\-on: they improve local robustness through external monitoring, filtering, and repair, but they do not prevent unsafe global behavior from emerging through agent composition\. In contrast, we argue that trustworthy agent ecosystems require trust to be baked into the system itself, such that unsafe trajectories are restricted by construction\. To formalize this perspective, we introduce the Trustworthy Agent Network \(TAN\) framework and identify four core principles, compositional robustness, semantic containment, accountability, and cross\-boundary reliability, as the foundation for trust at the network level\. Future work should convert this conceptual framing into implementable architectures, formal verification targets, and benchmarkable design criteria for real\-world multi\-agent systems\.
## References
- A survey on homomorphic encryption schemes: theory and implementation\.ACM Computing Surveys \(Csur\)51\(4\),pp\. 1–35\.Cited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- D\. Amodei, C\. Olah, J\. Steinhardt, P\. Christiano, J\. Schulman, and D\. Mané \(2016\)Concrete problems in ai safety\.arXiv preprint arXiv:1606\.06565\.Cited by:[§4\.1](https://arxiv.org/html/2605.19035#S4.SS1.p2.1)\.
- Anthropic \(2024a\)Introducing the model context protocol\.Note:[https://www\.anthropic\.com/news/model\-context\-protocol](https://www.anthropic.com/news/model-context-protocol)Cited by:[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.4.3.2.1.1)\.
- Anthropic \(2024b\)Model context protocol\.Note:Technical documentationCited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- Anthropic \(2026\)Agent skills\.Note:[https://platform\.claude\.com/docs/en/agents\-and\-tools/agent\-skills/overview](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview)Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon,et al\.\(2022\)Constitutional ai: harmlessness from ai feedback\.arXiv preprint arXiv:2212\.08073\.Cited by:[§4\.1](https://arxiv.org/html/2605.19035#S4.SS1.p1.1)\.
- Z\. Bao, Y\. Ji, W\. Wu, X\. Chen, and L\. He \(2025\)Supervisor alignment framework: enhancing llm alignment with query\-ignoring strategy and multi\-agent interaction\.InICASSP 2025\-2025 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 1–5\.Cited by:[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1)\.
- F\. Bousetouane \(2026\)AI agents need memory control over more context\.arXiv preprint arXiv:2601\.11653\.Cited by:[§2\.1\.4](https://arxiv.org/html/2605.19035#S2.SS1.SSS4.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.5.4.2.1.1)\.
- Z\. Brakerski, C\. Gentry, and V\. Vaikuntanathan \(2014\)\(Leveled\) fully homomorphic encryption without bootstrapping\.ACM Transactions on Computation Theory \(TOCT\)6\(3\),pp\. 1–36\.Cited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.\(2020\)Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[§4\.1](https://arxiv.org/html/2605.19035#S4.SS1.p2.1)\.
- M\. Cemri, M\. Z\. Pan, S\. Yang, L\. A\. Agrawal, B\. Chopra, R\. Tiwari, K\. Keutzer, A\. Parameswaran, D\. Klein, K\. Ramchandran,et al\.\(2025\)Why do multi\-agent llm systems fail?\.arXiv preprint arXiv:2503\.13657\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p3.1),[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p2.1)\.
- J\. H\. Cheon, A\. Kim, M\. Kim, and Y\. Song \(2017\)Homomorphic encryption for arithmetic of approximate numbers\.InInternational conference on the theory and application of cryptology and information security,pp\. 409–437\.Cited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- P\. F\. Christiano, J\. Leike, T\. Brown, M\. Martic, S\. Legg, and D\. Amodei \(2017\)Deep reinforcement learning from human preferences\.Advances in neural information processing systems30\.Cited by:[§4\.2](https://arxiv.org/html/2605.19035#S4.SS2.p1.1)\.
- V\. Costan and S\. Devadas \(2016\)Intel sgx explained\.Cryptology ePrint Archive\.Cited by:[§2\.1\.4](https://arxiv.org/html/2605.19035#S2.SS1.SSS4.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.5.4.2.1.1),[§4\.4](https://arxiv.org/html/2605.19035#S4.SS4.p1.1)\.
- D\. E\. Denning \(1976\)A lattice model of secure information flow\.Communications of the ACM19\(5\),pp\. 236–243\.Cited by:[§4\.4](https://arxiv.org/html/2605.19035#S4.SS4.p2.1)\.
- Q\. Dong, L\. Li, D\. Dai, C\. Zheng, J\. Ma, R\. Li, H\. Xia, J\. Xu, Z\. Wu, B\. Chang,et al\.\(2024a\)A survey on in\-context learning\.InProceedings of the 2024 conference on empirical methods in natural language processing,pp\. 1107–1128\.Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- Y\. Dong, R\. Mu, G\. Jin, Y\. Qi, J\. Hu, X\. Zhao, J\. Meng, W\. Ruan, and X\. Huang \(2024b\)Building guardrails for large language models\.arXiv preprint arXiv:2402\.01822\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p3.1),[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1)\.
- W\. Du and M\. J\. Atallah \(2001\)Secure multi\-party computation problems and their applications: a review and open problems\.InProceedings of the 2001 workshop on New security paradigms,pp\. 13–22\.Cited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. Mordatch \(2024\)Improving factuality and reasoning in language models through multiagent debate\.InForty\-first international conference on machine learning,Cited by:[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1),[§4\.2](https://arxiv.org/html/2605.19035#S4.SS2.p2.1)\.
- T\. R\. Gadekallu, T\. Huynh\-The, W\. Wang, G\. Yenduri, P\. Ranaweera, Q\. Pham, D\. B\. da Costa, and M\. Liyanage \(2022\)Blockchain for the metaverse: a review\.arXiv preprint arXiv:2203\.09738\.Cited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- J\. Geng, Q\. Li, H\. Woisetschlaeger, Z\. Chen, F\. Cai, Y\. Wang, P\. Nakov, H\. Jacobsen, and F\. Karray \(2025\)A comprehensive survey of machine unlearning techniques for large language models\.arXiv preprint arXiv:2503\.01854\.Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- O\. Goldreich and Y\. Oren \(1994\)Definitions and properties of zero\-knowledge proof systems\.Journal of Cryptology7\(1\),pp\. 1–32\.Cited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- O\. Goldreich \(1998\)Secure multi\-party computation\.Note:ManuscriptCited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- R\. Greenblatt, C\. Denison, B\. Wright, F\. Roger, M\. MacDiarmid, S\. Marks, J\. Treutlein, T\. Belonax, J\. Chen, D\. Duvenaud,et al\.\(2024\)Alignment faking in large language models\.arXiv preprint arXiv:2412\.14093\.Cited by:[§2\.1\.4](https://arxiv.org/html/2605.19035#S2.SS1.SSS4.p2.1)\.
- K\. Greshake, S\. Abdelnabi, S\. Mishra, C\. Endres, T\. Holz, and M\. Fritz \(2023\)Not what you’ve signed up for: compromising real\-world llm\-integrated applications with indirect prompt injection\.InProceedings of the 16th ACM workshop on artificial intelligence and security,pp\. 79–90\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p4.1),[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p2.1)\.
- I\. Habler, K\. Huang, V\. S\. Narajala, and P\. Kulkarni \(2025\)Building a secure agentic ai application leveraging a2a protocol\.arXiv preprint arXiv:2504\.16902\.Cited by:[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.4.3.2.1.1)\.
- D\. Hendrycks, M\. Mazeika, and T\. Woodside \(2023\)An overview of catastrophic ai risks\.arXiv preprint arXiv:2306\.12001\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p3.1),[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p2.1)\.
- E\. Hubinger, C\. Denison, J\. Mu, M\. Lambert, M\. Tong, M\. MacDiarmid, T\. Lanham, D\. M\. Ziegler, T\. Maxwell, N\. Cheng,et al\.\(2024\)Sleeper agents: training deceptive llms that persist through safety training\.arXiv preprint arXiv:2401\.05566\.Cited by:[§2\.1\.4](https://arxiv.org/html/2605.19035#S2.SS1.SSS4.p2.1)\.
- G\. Irving, P\. Christiano, and D\. Amodei \(2018\)AI safety via debate\.arXiv preprint arXiv:1805\.00899\.Cited by:[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p2.1)\.
- L\. B\. Kaesberg, J\. Becker, J\. P\. Wahle, T\. Ruas, and B\. Gipp \(2025\)Voting or consensus? decision\-making in multi\-agent debate\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 11640–11671\.Cited by:[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1)\.
- A\. Kierans, A\. Ghosh, H\. Hazan, and S\. Dori\-Hacohen \(2025\)Quantifying misalignment between agents: towards a sociotechnical understanding of alignment\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 27365–27373\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p4.1)\.
- P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel,et al\.\(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.Advances in neural information processing systems33,pp\. 9459–9474\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p1.1),[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- A\. Li, Y\. Xie, S\. Li, F\. Tsung, B\. Ding, and Y\. Li \(2024\)Agent\-oriented planning in multi\-agent systems\.arXiv preprint arXiv:2410\.02189\.Cited by:[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1)\.
- G\. Li, H\. Hammoud, H\. Itani, D\. Khizbullin, and B\. Ghanem \(2023\)Camel: communicative agents for" mind" exploration of large language model society\.Advances in neural information processing systems36,pp\. 51991–52008\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p2.1)\.
- J\. Liang, Z\. Cai, J\. Zhu, H\. Huang, K\. Zong, B\. An, M\. Alharthi, J\. He, L\. Zhang, H\. Li,et al\.\(2024\)Alignment at pre\-training\! towards native alignment for arabic llms\.Advances in Neural Information Processing Systems37,pp\. 13872–13896\.Cited by:[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- J\. Lu, Z\. Zhang, F\. Yang, J\. Zhang, L\. Wang, C\. Du, Q\. Lin, S\. Rajmohan, D\. Zhang, and Q\. Zhang \(2025\)Axis: efficient human\-agent\-computer interaction with api\-first llm\-based agents\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7711–7743\.Cited by:[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1)\.
- E\. Lumer, A\. Gulati, V\. K\. Subbiah, P\. H\. Basavaraju, and J\. A\. Burke \(2025\)Scalemcp: dynamic and auto\-synchronizing model context protocol tools for llm agents\.InInternational Joint Conference on Computational Intelligence,pp\. 23–42\.Cited by:[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.4.3.2.1.1)\.
- A\. Madry, A\. Makelov, L\. Schmidt, D\. Tsipras, and A\. Vladu \(2017\)Towards deep learning models resistant to adversarial attacks\.arXiv preprint arXiv:1706\.06083\.Cited by:[§4\.1](https://arxiv.org/html/2605.19035#S4.SS1.p1.1)\.
- V\. Ojewale, H\. Suresh, and S\. Venkatasubramanian \(2026\)Audit trails for accountability in large language models\.arXiv preprint arXiv:2601\.20727\.Cited by:[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.4.3.2.1.1)\.
- OpenClaw Community \(2026\)Moltbook: a social platform for autonomous ai agents\.Note:[https://www\.moltbook\.com](https://www.moltbook.com/)Accessed: 2026\-03\-05Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p2.1)\.
- OpenClaw Developers \(2026\)ClawHub: the openclaw skill registry\.Note:[https://clawhub\.ai](https://clawhub.ai/)Accessed: 2026\-03\-05Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p2.1)\.
- OpenClaw Team \(2026\)OpenClaw — personal ai assistant\.Note:[https://github\.com/openclaw/openclaw](https://github.com/openclaw/openclaw)Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p2.1),[§1](https://arxiv.org/html/2605.19035#S1.p3.1),[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray,et al\.\(2022\)Training language models to follow instructions with human feedback\.Advances in neural information processing systems35,pp\. 27730–27744\.Cited by:[§4\.1](https://arxiv.org/html/2605.19035#S4.SS1.p1.1)\.
- J\. S\. Park, J\. C\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. Bernstein \(2023\)Generative agents: interactive simulacra of human behavior\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p4.1),[§4\.1](https://arxiv.org/html/2605.19035#S4.SS1.p2.1)\.
- E\. Perez, S\. Ringer, K\. Lukosiute, K\. Nguyen, E\. Chen, S\. Heiner, C\. Pettit, C\. Olsson, S\. Kundu, S\. Kadavath,et al\.\(2023\)Discovering language model behaviors with model\-written evaluations\.InFindings of the association for computational linguistics: ACL 2023,pp\. 13387–13434\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p5.1),[§4\.2](https://arxiv.org/html/2605.19035#S4.SS2.p2.1)\.
- D\. Reed, M\. Sporny, D\. Longley, C\. Allen, R\. Grant, M\. Sabadello, and J\. Holt \(2020\)Decentralized identifiers \(dids\) v1\. 0\.Draft Community Group Report\.Cited by:[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.4.3.2.1.1)\.
- M\. Renze and E\. Guven \(2024\)Self\-reflection in llm agents: effects on problem\-solving performance\.arXiv preprint arXiv:2405\.06682\.Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- M\. Sabt, M\. Achemlal, and A\. Bouabdallah \(2015\)Trusted execution environment: what it is, and what it is not\.In2015 IEEE Trustcom/BigDataSE/Ispa,Vol\.1,pp\. 57–64\.Cited by:[§4\.4](https://arxiv.org/html/2605.19035#S4.SS4.p1.1)\.
- R\. S\. Sandhu \(1998\)Role\-based access control\.InAdvances in computers,Vol\.46,pp\. 237–286\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p3.1),[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p2.1),[§2\.1\.4](https://arxiv.org/html/2605.19035#S2.SS1.SSS4.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.5.4.2.1.1),[§4\.4](https://arxiv.org/html/2605.19035#S4.SS4.p2.1)\.
- S\. Sayeed, H\. Marco\-Gisbert, and T\. Caira \(2020\)Smart contract: attacks and protections\.Ieee Access8,pp\. 24416–24427\.Cited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. Scialom \(2023\)Toolformer: language models can teach themselves to use tools\.Advances in neural information processing systems36,pp\. 68539–68551\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p1.1),[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- P\. Schoenegger, M\. Carlson, C\. Schneider, and C\. Daly \(2026\)Verifiable semantics for agent\-to\-agent communication\.arXiv preprint arXiv:2602\.16424\.Cited by:[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.4.3.2.1.1)\.
- K\. Schroeder and Z\. Wood\-Doughty \(2024\)Can you trust llm judgments? reliability of llm\-as\-a\-judge\.arXiv preprint arXiv:2412\.12509\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p5.1)\.
- Y\. Shao, L\. Li, J\. Dai, and X\. Qiu \(2023\)Character\-llm: a trainable agent for role\-playing\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 13153–13187\.Cited by:[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1)\.
- Y\. Shen, K\. Song, X\. Tan, D\. Li, W\. Lu, and Y\. Zhuang \(2023\)Hugginggpt: solving ai tasks with chatgpt and its friends in hugging face\.Advances in Neural Information Processing Systems36,pp\. 38154–38180\.Cited by:[§4\.2](https://arxiv.org/html/2605.19035#S4.SS2.p1.1)\.
- R\. Sommer and V\. Paxson \(2010\)Outside the closed world: on using machine learning for network intrusion detection\.In2010 IEEE symposium on security and privacy,pp\. 305–316\.Cited by:[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.4.3.2.1.1)\.
- T\. South, S\. Marro, T\. Hardjono, R\. Mahari, C\. D\. Whitney, D\. Greenwood, A\. Chan, and A\. Pentland \(2025\)Authenticated delegation and authorized ai agents\.arXiv preprint arXiv:2501\.09674\.Cited by:[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.4.3.2.1.1)\.
- Ö\. Suçeken and O\. Özkaraca \(2024\)Cryptography with artificial intelligence: an overview\.InThe International Conference on Artificial Intelligence and Applied Mathematics in Engineering,pp\. 162–172\.Cited by:[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.4.3.2.1.1)\.
- X\. Sun, F\. R\. Yu, P\. Zhang, Z\. Sun, W\. Xie, and X\. Peng \(2021\)A survey on zero\-knowledge proof in blockchain\.IEEE network35\(4\),pp\. 198–205\.Cited by:[§4\.3](https://arxiv.org/html/2605.19035#S4.SS3.p1.4)\.
- Y\. Tsai, T\. Yen, P\. Guo, Z\. Li, and S\. Lin \(2024\)Text\-centric alignment for multi\-modality learning\.arXiv preprint arXiv:2402\.08086\.Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- R\. Wahbe, S\. Lucco, T\. E\. Anderson, and S\. L\. Graham \(1993\)Efficient software\-based fault isolation\.InProceedings of the fourteenth ACM symposium on Operating systems principles,pp\. 203–216\.Cited by:[§2\.1\.4](https://arxiv.org/html/2605.19035#S2.SS1.SSS4.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.5.4.2.1.1),[§4\.4](https://arxiv.org/html/2605.19035#S4.SS4.p1.1)\.
- L\. Wang, S\. Chen, L\. Jiang, S\. Pan, R\. Cai, S\. Yang, and F\. Yang \(2025a\)Parameter\-efficient fine\-tuning in large language models: a survey of methodologies\.Artificial Intelligence Review58\(8\),pp\. 227\.Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- Y\. Wang, P\. Wang, C\. Xi, B\. Tang, J\. Zhu, W\. Wei, C\. Chen, C\. Yang, J\. Zhang, C\. Lu,et al\.\(2025b\)Adversarial preference learning for robust llm alignment\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 21865–21881\.Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- J\. White, Q\. Fu, S\. Hays, M\. Sandborn, C\. Olea, H\. Gilbert, A\. Elnashar, J\. Spencer\-Smith, and D\. C\. Schmidt \(2023\)A prompt pattern catalog to enhance prompt engineering with chatgpt\.arXiv preprint arXiv:2302\.11382\.Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- M\. Wooldridge and N\. R\. Jennings \(1995\)Intelligent agents: theory and practice\.The knowledge engineering review10\(2\),pp\. 115–152\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p1.1)\.
- Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu,et al\.\(2024\)Autogen: enabling next\-gen llm applications via multi\-agent conversations\.InFirst conference on language modeling,Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p2.1)\.
- R\. Xu and J\. Peng \(2025\)A comprehensive survey of deep research: systems, methodologies, and applications\.arXiv preprint arXiv:2506\.12594\.Cited by:[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1)\.
- H\. Yang, S\. Yue, and Y\. He \(2023\)Auto\-gpt for online decision making: benchmarks and additional opinions\.arXiv preprint arXiv:2306\.02224\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p1.1)\.
- J\. Yang, C\. E\. Jimenez, A\. Wettig, K\. Lieret, S\. Yao, K\. Narasimhan, and O\. Press \(2024\)Swe\-agent: agent\-computer interfaces enable automated software engineering\.Advances in Neural Information Processing Systems37,pp\. 50528–50652\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p1.1),[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- S\. Yao, J\. Zhao, D\. Yu, N\. Gao,et al\.\(2023\)ReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p1.1),[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1),[§4\.2](https://arxiv.org/html/2605.19035#S4.SS2.p1.1)\.
- G\. Zhang, H\. Geng, X\. Yu, Z\. Yin, Z\. Zhang, Z\. Tan, H\. Zhou, Z\. Li, X\. Xue, Y\. Li,et al\.\(2025a\)The landscape of agentic reinforcement learning for llms: a survey\.arXiv preprint arXiv:2509\.02547\.Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- Z\. Zhang, Q\. Dai, X\. Bo, C\. Ma, R\. Li, X\. Chen, J\. Zhu, Z\. Dong, and J\. Wen \(2025b\)A survey on the memory mechanism of large language model\-based agents\.ACM Transactions on Information Systems43\(6\),pp\. 1–47\.Cited by:[§2\.1\.1](https://arxiv.org/html/2605.19035#S2.SS1.SSS1.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.2.1.2.1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.\(2023\)Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in neural information processing systems36,pp\. 46595–46623\.Cited by:[§1](https://arxiv.org/html/2605.19035#S1.p3.1),[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1)\.
- H\. Zhou, B\. Zhang, Z\. Li, M\. Yan, and M\. Zhang \(2025\)A training\-free llm\-based approach to general chinese character error correction\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13827–13852\.Cited by:[§2\.1\.2](https://arxiv.org/html/2605.19035#S2.SS1.SSS2.p1.1),[Table 1](https://arxiv.org/html/2605.19035#S2.T1.6.3.2.2.1.1)\.
- A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. Fredrikson \(2023\)Universal and transferable adversarial attacks on aligned language models\.arXiv preprint arXiv:2307\.15043\.Cited by:[§2\.1\.3](https://arxiv.org/html/2605.19035#S2.SS1.SSS3.p2.1)\.Similar Articles
How to solve the trust problem between AI agents?
The article explores how Web3 and onchain verification can address trust issues in AI agent collaboration, including through Agent Stores that act as marketplaces for agents to buy and sell services based on reputation and transaction history.
Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and Prototypes
This paper proposes a layered architecture for distributed general-purpose agent networks, enabling heterogeneous AI agents to discover, trust, and cooperate on open-ended tasks across personal devices and edge nodes.
The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
A position paper arguing that autonomous AI agents in science widen the verification gap and that scientific verification infrastructure must evolve with observable-by-default workflows, scalable verification, and clear attribution to sustain trustworthy science.
A social network for agents solves the easy half — the hard case is a person, an agent, and a 40-year-old transport
The author argues that agent-to-agent social networks are easy because trust is assumed, while the real challenge is integrating agents into legacy human communication like email, where cryptographically verifiable delegation of authority is needed.
If AI agents become everywhere, how do we know which ones to trust?
As AI agents become ubiquitous, the challenge shifts from comparing performance to establishing trust and reputation, requiring new discovery and verification systems.