Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

arXiv cs.AI Papers

Summary

This paper proposes isolation as a first-class principle for LLM-agent system safety, presenting a boundary-centric taxonomy to analyze failures and defenses. It systematically categorizes safety issues across five boundaries and outlines future research directions.

arXiv:2607.12406v1 Announce Type: new Abstract: The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequently, safety is no longer only about input--output content alignment. It also concerns system behavior and real-world execution outcomes. However, the current literature is fragmented across attack types, applications, and benchmarks. This makes it hard to explain why failures such as prompt injection, tool misuse, and memory poisoning often share the same structural cause, and how they spread through an agent workflow. In this survey, we treat isolation as a first-class principle for LLM-agent system safety. By isolation, we refer to the separation of user inputs, tool access, execution channels, inter-agent communication, and environment-originated context. We organize the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment. This view helps identify where the loss of isolation first occurs, how compromise propagates across boundaries, and which defenses are most relevant at each interface. We also summarize cross-boundary failure paths, discuss open challenges, and outline a research agenda for isolation-by-construction in future agent systems.
Original Article
View Cached Full Text

Cached at: 07/15/26, 04:20 AM

# Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions
Source: [https://arxiv.org/html/2607.12406](https://arxiv.org/html/2607.12406)
Huihao Jing![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x1.png),Wenbin Hu![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x2.png),Shaojin Chen![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x3.png),Haochen Shi![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x4.png),Sirui Zhang![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/assets/nyu.png)![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x5.png), Hanyu Yang![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/assets/swupl.png)![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x6.png),Changxuan Fan![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x7.png),Zhongwei Xie![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x8.png),Hongyu Luo![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x9.png), Wun Yu Chan![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x10.png),Wei Fan![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x11.png),Haoran Li![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x12.png),Yangqiu Song![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x13.png) ![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x14.png)HKUST,![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/assets/nyu.png)NYU,![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/assets/swupl.png)SWUPL,![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x15.png)MODEIO\.AI, hjingaa@connect\.ust\.hk

###### Abstract

The capability of LLM agents to function as the “brain” of a system fundamentally expands the scope of analysis beyond a standalone model\. Consequently, safety is no longer only about input–output content alignment\. It also concerns system behavior and real\-world execution outcomes\. However, the current literature is fragmented across attack types, applications, and benchmarks\. This makes it hard to explain why failures such as prompt injection, tool misuse, and memory poisoning often share the same structural cause, and how they spread through an agent workflow\. In this survey, we treat isolation as a first\-class principle for LLM\-agent system safety\. By isolation, we refer to the separation of user inputs, tool access, execution channels, inter\-agent communication, and environment\-originated context\. We organize the literature with a boundary\-centric taxonomy of five boundaries: user\-agent, agent\-tool, agent\-execution, agent\-agent, and system\-environment\. This view helps identify where the loss of isolation first occurs, how compromise propagates across boundaries, and which defenses are most relevant at each interface\. We also summarize cross\-boundary failure paths, discuss open challenges, and outline a research agenda for isolation\-by\-construction in future agent systems\.

Isolation as a First\-Class Principle for LLM\-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

Huihao Jing![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x16.png), Wenbin Hu![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x17.png), Shaojin Chen![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x18.png), Haochen Shi![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x19.png), Sirui Zhang![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/assets/nyu.png)![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x20.png),Hanyu Yang![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/assets/swupl.png)![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x21.png),Changxuan Fan![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x22.png),Zhongwei Xie![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x23.png),Hongyu Luo![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x24.png),Wun Yu Chan![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x25.png),Wei Fan![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x26.png),Haoran Li![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x27.png)††thanks:Corresponding author,Yangqiu Song![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x28.png)![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x29.png)HKUST,![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/assets/nyu.png)NYU,![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/assets/swupl.png)SWUPL,![[Uncaptioned image]](https://arxiv.org/html/2607.12406v1/x30.png)MODEIO\.AI,hjingaa@connect\.ust\.hk

## 1Introduction

![Refer to caption](https://arxiv.org/html/2607.12406v1/x31.png)Figure 1:A boundary\-centric taxonomy of LLM\-agent system safety organized around five isolation boundaries: user\-agent, agent\-tool, agent\-execution, agent\-agent, and system\-environment\.LLM agents are moving quickly from research prototypes to real systems\. Recent examples include CodexOpenAI \([2025](https://arxiv.org/html/2607.12406#bib.bib87)\), Claude CodeAnthropic \([2025](https://arxiv.org/html/2607.12406#bib.bib6)\), and OpenClawOpenClaw \([2026](https://arxiv.org/html/2607.12406#bib.bib88)\)\. Related work on managed agentsAnthropic \([2026](https://arxiv.org/html/2607.12406#bib.bib7)\)and agent harnessesMeng et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib84)\)shows the same shift\. These systems do more than generate text\. They read files, call tools, browse the web, edit state, coordinate with other agents, and take actions over long horizons\. As a result, the safety problem is no longer only about unsafe outputs\. It is also about how control, capability, memory, and context move through the system\.

This change has already shaped both industry practice and recent research\. A common direction is to decouple the brain from the hands\. Managed agentsAnthropic \([2026](https://arxiv.org/html/2607.12406#bib.bib7)\), AgentVisorYing et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib142)\), privilege\-separation proposalsJacob et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib52)\), CaMeL\-style design defensesDebenedetti et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib30)\), and ParallaxFokou \([2026](https://arxiv.org/html/2607.12406#bib.bib36)\)all push in this direction\. Different papers use different terms, such as virtualization, privilege separation, typed interfaces, and design\-level defense\. But the core idea is simple: agent safety improves when system boundaries are explicit and enforced structurally, rather than left to prompt instructions alone\.

Despite this progress, the literature remains fragmented\. Existing surveys often organize the space by attack type, application domain, or agent capabilityMeng et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib84)\); Li et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib68)\); Kim et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib63)\)\. That view is useful, but it leaves a gap\. It does not clearly explain why failures that look different on the surface often share the same cause: loss of isolation at a boundary\. It also makes cross\-boundary propagation harder to see\.

This survey addresses that gap\. We treat isolation as a first\-class principle for LLM\-agent safety and organize the literature around five boundaries: user\-agent, agent\-tool, agent\-execution, agent\-agent, and system\-environment\. Figure[1](https://arxiv.org/html/2607.12406#S1.F1)gives a compact view of this taxonomy\. At the center is the agent core\. Around it sit five interfaces where user inputs, tool access, execution channels, inter\-agent communication, and environment\-originated context may change status\. Theuser\-agent boundaryconcerns whether user content remains data or becomes control\. Theagent\-tool boundaryconcerns how external capabilities are accessed\. Theagent\-execution boundaryconcerns the transition from reasoning to action\. Theagent\-agent boundaryconcerns communication and coordination across multiple agents\. Thesystem\-environment boundaryconcerns how the agent system interacts with external content and state as a whole\. These boundaries are distinct, but closely related\. Together they provide a cleaner way to explain where compromise begins and how it later propagates\. This taxonomy is boundary\-centric, but not boundary\-exclusive\. Many papers naturally touch more than one boundary\. We therefore classify papers by their*primary safety boundary*, that is, the point where the loss of isolation first occurs\. This rule keeps the taxonomy simple and makes cross\-boundary analysis easier\.

Our contributions are summarized as follows:

- •We propose a boundary\-centric taxonomy for LLM\-agent safety based on five isolation boundaries, and map attacks/defenses to where isolation first breaks, showing how failures propagate across agent workflows\.
- •We unify prior work through an isolation\-centric view of failure propagation, and outline an isolation\-by\-construction agenda emphasizing trust separation, scoped capabilities, traceability, and recovery\.

Boundary\-Centric Taxonomy of LLM\-Agent SafetyUser\-AgentBoundary§[2](https://arxiv.org/html/2607.12406#S2)Direct promptinjection / jailbreaksAuthorityseparation failureMulti\-turnattacksMultimodal /long\-context attacksPersistentcompromiseDirect prompt injection remains the canonical failure mode, including instruction overridePerez and Ribeiro \([2022](https://arxiv.org/html/2607.12406#bib.bib93)\); Liu et al\. \([2023b](https://arxiv.org/html/2607.12406#bib.bib77)\), jailbreak suffixesZhu et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib164)\), and adaptive attacksMazeika et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib83)\)\.User content crosses privilege boundaries, behaving as policy or developer instructions rather than data\.Compromise can emerge gradually through dialogueRussinovich et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib102)\), contextual derailmentRen et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib98)\), and implicit steeringZhou et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib163)\)\.ImagesGong et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib38)\); Wu et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib132)\)and long contextsCheng et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib26)\)expand the attack surface beyond plain text\.Attacks can persist through memory injectionHe et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib44)\), in\-context corruptionWang et al\. \([2024e](https://arxiv.org/html/2607.12406#bib.bib126)\), and residual influenceXu et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib137)\)\.Agent\-ToolBoundary§[3](https://arxiv.org/html/2607.12406#S3)Tool outputs ascontrolTool misuse /selection failuresArgumentconstructionProtocol /metadata risksTrajectory\-levelorchestrationTool\-returned observations may be treated as trusted control rather than bounded evidence\. This observation\-control blur is central in indirect tool injectionYi et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib141)\)and tool feedback misuseFu et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib37)\)\.The system may select the wrong capabilityYe et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib140)\), route toward the wrong external serviceZhao et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib157)\), or trust a malicious tool advertisementZhang et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib151)\)\.Even when the correct tool is chosen, unsafe argument construction can convert local reasoning mistakes into external side effects\.MCP\-style ecosystems show that metadataJing et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib59)\), capability descriptions and protocol manipulationWang et al\. \([2026b](https://arxiv.org/html/2607.12406#bib.bib127)\), and benchmarked MCP attacksZhang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib148)\)are part of the interface attack surface\.Many failures only become visible at the trajectory level\. Unsafe workflows emerge in adaptive tool tracesWang et al\. \([2026a](https://arxiv.org/html/2607.12406#bib.bib117)\), trace\-level analysisChen et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib21)\), safer invocation studiesMou et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib85)\), and coding\-agent benchmarksSteinberg and Gal \([2026](https://arxiv.org/html/2607.12406#bib.bib108)\)\.Agent\-ExecutionBoundary§[4](https://arxiv.org/html/2607.12406#S4)Code / browser /GUI action risksUnsafe actionrealizationEmbodied agents /VLA systemsContainment /runtime mediationThis boundary covers the point where plans become real actions\. Examples include code execution attacksZhang et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib147)\), autonomous website hackingFang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib35)\), risky code generationGuo et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib41)\), and unsafe web interactionTur et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib115)\)\.The key shift is from unsafe text to unsafe side effects: the system may click, run, submit, or manipulate the wrong thing even when the language output looks plausible\.In embodied and VLA systems, perception errors, backdoors, and adversarial inputs can be turned into unsafe motion or manipulationZhang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib149)\); Robey et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib100)\), including adversarial VLA vulnerabilitiesWang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib121)\)\.The defense trend therefore emphasizes action containment through sandbox ecosystemsZhou et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib161)\), policy\-executable checksLu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib80)\), and realistic execution benchmarksCui et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib28)\)\.Agent\-AgentBoundary§[5](https://arxiv.org/html/2607.12406#S5)Prompt infection /communication attacksDebate /message propagationTopology /cascade / memoryDefense agents /attributionA compromised message can become a carrier of control and spread from one agent to another\. This is explicit in prompt infectionLee and Tiwari \([2024](https://arxiv.org/html/2607.12406#bib.bib65)\)and communication attacksHe et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib43)\)\.Discussion, critique, debateAmayuelas et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib4)\), knowledge floodingJu et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib60)\), and communication reuseHe et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib43)\)can amplify rather than reduce malicious influence\.Topology, routing, and shared memory determine whether failures stay local or become systemic\. This is visible in backdoor insertionWang et al\. \([2024c](https://arxiv.org/html/2607.12406#bib.bib124)\), recursive blockingZhou et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib162)\), topology\-aware analysisWang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib120)\), and memory weaponizationDas et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib29)\)\.Representative defenses rely on guard agentsXiang et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib133)\), verifiable safety rolesChen et al\. \([2025d](https://arxiv.org/html/2607.12406#bib.bib23)\), failure attributionZhang et al\. \([2025d](https://arxiv.org/html/2607.12406#bib.bib152)\), and hierarchical data managementMao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib82)\)\.System\-EnvBoundary§[6](https://arxiv.org/html/2607.12406#S6)Indirect promptAmbient injectionHostile web /interfacesRAG poisoning /retrieval corruptionEnv\-mediateddisclosureAuth / provenance /causal defenseExternal content meant to be observational becomes effective control\. This is the core of indirect prompt injectionAbdelnabi et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib1)\); Liu et al\. \([2024c](https://arxiv.org/html/2607.12406#bib.bib75)\),including retrieval\-surviving attacksChang et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib13)\); Zhu et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib165)\)\.Webpages, interfaces, and visual surfaces act as adversarial environments\. Includes WIPI\-style web threatsWu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib131)\), environmental privacy leakageLiao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib71)\), web\-agent injectionEvtimov et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib33)\); Wang et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib122)\), and visual injectionCao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib11)\)\.RAG systems are vulnerable to poisoned knowledge basesZou et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib166)\); Deng et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib31)\), retrieval corruptionXue et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib138)\); Zhang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib150)\), and compromised evidence selection turning context into faulty support\.The same interfaces introduce confidentiality risks, including data extractionQi et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib96)\); Jiang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib57)\), membership inferenceLi et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib70)\), and multimodal leakageAl\-Lawati and Wang \([2026](https://arxiv.org/html/2607.12406#bib.bib2)\)\.Defenses emphasize authenticationWang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib118)\), task\-alignment shieldingJia et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib54)\), temporal diagnosisZhang et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib153)\), causal attributionHe et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib45)\), and memory hygieneGowda \([2026](https://arxiv.org/html/2607.12406#bib.bib40)\)\.

Figure 2:A roadmap\-style view of our boundary\-centric taxonomy for LLM\-agent system safety\. Each boundary is organized into representative subtopics and example literature\.
## 2User\-Agent Boundary

### 2\.1Threat Model and Boundary Definition

The user\-agent boundary asks whether an agent can keep user content from becoming privileged control\. In a well\-isolated system, user input should remain a request, a query, or task data\. System and developer instructions should keep higher authority\. Failure begins when user\-controlled content starts to steer internal policy or behavior, as shown in early prompt\-injection and role\-confusion studiesPerez and Ribeiro \([2022](https://arxiv.org/html/2607.12406#bib.bib93)\); Liu et al\. \([2023b](https://arxiv.org/html/2607.12406#bib.bib77)\); Toyer et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib113)\); Pedro et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib90)\); Ren \([2024](https://arxiv.org/html/2607.12406#bib.bib99)\)\. This is usually the first exposed interface in an agent system\. Before an agent uses tools, acts, or coordinates with peers, it must decide what counts as instruction and what counts as data\. If that decision is unstable, later safeguards start from a weak position\.

### 2\.2Direct Prompt Injection and Automated Jailbreaks

Early work showed that natural\-language inputs can override hidden promptsPerez and Ribeiro \([2022](https://arxiv.org/html/2607.12406#bib.bib93)\); Liu et al\. \([2023b](https://arxiv.org/html/2607.12406#bib.bib77)\), leak system instructionsToyer et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib113)\), and confuse role hierarchyPedro et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib90)\)\. The core issue is simple: a low\-privilege source can behave like a high\-privilege one\. Recent work made this threat much more systematic\. Attackers now use adversarial suffixesZhu et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib164)\); Liu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib74)\), black\-box searchJia et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib55)\); Sitawarin et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib106)\), proxy\-guided optimizationJiang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib56)\), and automated red teamingPei et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib91)\)to find failures at scale\. Benchmarks such as HarmBenchMazeika et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib83)\), JailbreakBenchChao et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib15)\), and JailbreakEvalRan et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib97)\)also pushed the field toward broader and more reproducible evaluation\. Together, these papers show that the user\-agent boundary is vulnerable under repeated adaptive pressure\.

### 2\.3From Single\-Turn Attacks to Persistent Compromise

More recent papers show that this boundary is not challenged only by single\-turn prompts\. Multi\-turn jailbreaks exploit gradual steeringCheng et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib26)\); Russinovich et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib102)\), implicit cluesRen et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib98)\); Chang et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib14)\), and long\-context overloadUpadhayay et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib116)\); Zhou et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib163)\)\. Many systems look safer in one\-shot evaluation than they are in realistic sessions\. The attack surface also expands in multimodal settings\. User\-provided images can carry visual or typographic instructions that bypass text\-oriented safeguardsLiu et al\. \([2023a](https://arxiv.org/html/2607.12406#bib.bib76)\); Gong et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib38)\); Wu et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib132)\)\. This boundary is therefore not only about text\. It spans multiple input channels\. The strongest recent trend is persistence\. In\-context poisoningHe et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib44)\), dynamic soft promptingWang et al\. \([2024e](https://arxiv.org/html/2607.12406#bib.bib126)\), and memory injectionDong et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib32)\); Xu et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib137)\)show that user influence can remain after the original interaction ends\. Related work on poisoning alignment dataShao et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib103)\); Zhang et al\. \([2025e](https://arxiv.org/html/2607.12406#bib.bib154)\)and harmful fine\-tuningPathmanathan et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib89)\); Huang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib50)\)shows that the same weakness can also enter through post\-training pipelines\. The problem is no longer only prompt\-level override\. It becomes long\-term state corruption\.

### 2\.4Defenses, Evaluation, and Future Directions

Defenses at this boundary fall into three broad groups\. The first makes authority separation explicit through structured queriesPiet et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib94)\); Chen et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib20)\), signed promptsSuo \([2024](https://arxiv.org/html/2607.12406#bib.bib110)\), and DSL\-style interfacesSharma et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib104)\)\. The second hardens the model against jailbreaks through safety classifiersKim et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib62)\), semantic smoothingRobey et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib101)\), repetition\-based defensesJi et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib53)\); Zhang et al\. \([2024c](https://arxiv.org/html/2607.12406#bib.bib155)\), and inference\-time self\-protectionWang et al\. \([2025d](https://arxiv.org/html/2607.12406#bib.bib123),[2024d](https://arxiv.org/html/2607.12406#bib.bib125)\); Lin et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib72)\)\. The third focuses on repair after degradation, including unlearningLi et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib69)\); Zhao et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib156)\), editingLu et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib78)\); Zhao et al\. \([2024c](https://arxiv.org/html/2607.12406#bib.bib158)\), and refusal\-boundary controlXiong et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib135)\); Lu et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib79)\); Gou et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib39)\)\. Evaluation has also improved\. Recent work studies detection with perplexityAlon and Kamfonas \([2023](https://arxiv.org/html/2607.12406#bib.bib3)\), refusal\-loss signalsHu et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib48)\), and geometry\-aware methodsCandogan et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib10)\); Yung et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib144)\)\. Benchmarks such as SG\-BenchMou et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib86)\), Case\-BenchSun et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib109)\), and Sorry\-BenchXie et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib134)\)broaden the test space beyond a few popular jailbreak prompts\. Durability studies then ask a harder question: do safeguards still work under longer interaction and stronger adaptationQi et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib95)\); Tamirisa et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib111)\); Chen et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib19)\)?

Taken together, this literature suggests that user\-agent safety must move beyond short prompts and static refusal scores toward long\-session robustness, multimodal authority separation, and recovery from persistent compromise\.

## 3Agent\-Tool Boundary

### 3\.1Threat Model and Boundary Definition

The agent\-tool boundary concerns how an agent accesses external capabilities\. In a well\-isolated system, tools should extend what the agent can do without taking over how it decides\. Tool descriptions, tool outputs, and orchestration logic should remain constrained interfaces rather than hidden control channels\. Failure begins when external capabilities are exposed without clear separation between observation, instruction, and execution, as shown by early indirect injection and tool\-learning failuresYi et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib141)\); Ye et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib140)\); Fu et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib37)\)\.

This boundary matters because tools change the risk profile of an agent system\. A standalone model can generate unsafe text\. A tool\-using agent can search the web, call APIs, read files, write code, or trigger downstream actions\. Small control errors at this boundary can therefore turn into capability errors\. The key question is not only whether the model reasons correctly, but whether the interface keeps capability use scoped, typed, and resistant to manipulation\.

### 3\.2Tool Outputs, Tool Misuse, and Tool Selection Failures

The basic failure mode at this boundary is that tool\-returned content is treated as trusted instruction\. Early work on indirect prompt injection showed that hidden instructions in tool outputs can redirect downstream reasoning and actionYi et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib141)\)\. Later work made this more concrete\. ToolSword and Imprompter show that tool learning can fail at several stages, including interpretation, tool selection, argument generation, and feedback useYe et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib140)\); Fu et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib37)\)\. The result is not just poor task performance\. It can also produce confidentiality, integrity, and safety failures\. Recent papers push this line further by showing that manipulation can happen even before a tool is called correctly\. Attacks on third\-party APIsZhao et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib157)\), adversarial tool\-callingZhang et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib151)\), and prompt injection against tool selectionShi et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib105)\)show that the agent can be steered into using the wrong capability, or using the right capability in the wrong way\. Tool safety is therefore not one failure mode\. The system can fail by choosing the wrong tool, passing unsafe arguments, trusting malicious output, or composing plausible calls into an unsafe workflow\.

### 3\.3Protocol, Metadata, and MCP\-Style Risks

A major recent shift is that tool\-use risk is moving from content alone to protocol design\. In MCP\-style ecosystems, the model often sees tool descriptions, capability advertisements, and metadata before it uses the tool itself\. MCIP and MPMA show that this layer is already security\-criticalJing et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib59)\); Wang et al\. \([2026b](https://arxiv.org/html/2607.12406#bib.bib127)\)\. Metadata is not neutral from the model’s perspective\. It can shape preferences, change routing, and bias decisions before real tool output even appears\. This is why recent benchmark work on MCP and tool ecosystems matters\. MCP Security BenchZhang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib148)\)shows that attacks against model context protocol are not an edge case, but a natural extension of tool\-mediated prompt injection\. ToolSafeMou et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib85)\)makes the same point from the defense side: tool invocation needs stronger safety mediation, not just better general alignment\. The interface itself has become part of the attack surface\. As tool use becomes standardized through protocols and orchestrators, the boundary no longer sits only at the moment of API invocation\. It also sits at discovery, ranking, metadata interpretation, and workflow construction\. Protocol semantics and capability exposure thus become first\-class safety issues\.

### 3\.4Trajectory\-Level Risk and Safe Tool Orchestration

Recent work shows that tool attacks are often multi\-step\. A single call may look harmless, while the full trajectory becomes unsafe\. AdapTools and TraceSafe make this clear in multi\-step tool workflowsWang et al\. \([2026a](https://arxiv.org/html/2607.12406#bib.bib117)\); Chen et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib21)\)\. The security target is therefore not only one prompt or one API call, but the full trace of calls, arguments, observations, and intermediate state\. This is especially important for code interpreters and coding agents\. CIBER and MOSAIC\-Bench show that once tools support code execution or complex coding workflows, tool misuse can translate into direct downstream impact much more easilyBa et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib8)\); Steinberg and Gal \([2026](https://arxiv.org/html/2607.12406#bib.bib108)\)\. These systems are close to the execution boundary, but the first loss of isolation often still happens here, when a tool is selected, configured, or trusted too early\. Defenses follow the same shift from call\-level safety to interface\-level safety\. Some works harden the protocol layer through contextual integrityJing et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib59)\), metadata constraintsWang et al\. \([2026b](https://arxiv.org/html/2607.12406#bib.bib127)\), and benchmark\-based auditingZhang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib148)\)\. Others focus on safer invocation through least privilege, better routing, and orchestration\-aware evaluationChen and Cong \([2025](https://arxiv.org/html/2607.12406#bib.bib18)\); Mou et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib85)\)\. The broader lesson is simple: safer tool use depends less on the model inferring the right behavior from natural language, and more on interfaces that make capability scope, trust level, and action semantics explicit\. This also points to future work on trace\-level evaluation, privileged treatment of protocol metadata, and tighter links between tool control and execution control\.

## 4Agent\-Execution Boundary

### 4\.1Threat Model and Boundary Definition

The agent\-execution boundary concerns the point at which internal decisions become real actions\. In a well\-isolated system, planning and action should remain separate enough that unsafe decisions can still be checked, delayed, or blocked before they create side effects\. Failure begins when the system treats model output as ready\-to\-execute behavior without enough mediation\.

This boundary turns control errors into operational impact\. A model that produces unsafe text is one kind of problem\. An agent that clicks the wrong button, runs unsafe code, submits the wrong form, or issues unsafe physical commands is another, as recent cyber\-agent work makes clearZhang et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib147)\); Fang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib35),[a](https://arxiv.org/html/2607.12406#bib.bib34)\); Guo et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib41)\)\.

### 4\.2Code, Browser, and GUI Action Risks

Recent work on cyber and software agents makes this transition clear\. Agents can be redirected into malfunction amplificationZhang et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib147)\), risky code generationFang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib35)\), or direct offensive behavior such as website compromise and vulnerability exploitationFang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib34)\); Guo et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib41)\)\. The same pattern appears in browser and GUI agents\. ST\-WebAgentBenchLevy et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib66)\)and SafeArenaTur et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib115)\)show that web agents can produce unsafe outcomes through clicks, submissions, navigation, and action grounding\. Browser\-agent work also shows that refusal\-trained models remain vulnerable once they act through an interface rather than a chat windowKumar et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib64)\), while GUI\-agent studies show added risk from visual grounding, interface ambiguity, and hidden action consequencesChen et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib17)\)\. Execution safety therefore cannot be reduced to ordinary alignment\.

### 4\.3Embodied Agents and Vision\-Language\-Action Systems

The execution boundary becomes even sharper in embodied systems, where failures are often costlier and less reversible\. BADROBOTZhang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib149)\), robot jailbreakingRobey et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib100)\), and contextual backdoor attacksLiu et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib73)\); Jiao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib58)\)show that benign\-looking inputs can induce unsafe physical behavior or persistent actuation failures\. Benchmarks such as Embodied Agent InterfaceLi et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib67)\), VIVAHu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib49)\), HASARDTomilin et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib112)\), HEALChakraborty et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib12)\), and Embodied Red TeamingKarnik et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib61)\)make this risk more concrete\. Work on VLA vulnerabilities further shows that execution can be corrupted through perception, including visual perturbationsWu et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib130)\), adversarial patchesWang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib121)\), and action\-model weaknessesCheng et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib25)\); Zhou et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib160)\)\.

### 4\.4Containment, Runtime Mediation, and Future Directions

Because failures at this boundary create real side effects, defenses increasingly focus on containment rather than only prevention\. Recent work proposes constrained executionWang et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib119)\), zero\-trust architecturesLu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib80)\), active defenseZhou et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib161)\), and policy\-executable safeguardsMaiti \([2026](https://arxiv.org/html/2607.12406#bib.bib81)\)\. Evaluation reflects the same shift\. SafeArenaTur et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib115)\), ST\-WebAgentBenchLevy et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib66)\), and HWE\-BenchCui et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib28)\)ask whether the agent remains safe while interacting with real interfaces and workflows\. Overall, this boundary suggests that safe agent design must treat execution as a governed interface, not as the automatic continuation of model output\.

## 5Agent\-Agent Boundary

### 5\.1Threat Model and Boundary Definition

The agent\-agent boundary concerns what happens when multiple agents communicate, delegate, debate, and share intermediate state\. In a well\-isolated system, one agent’s message should remain a bounded contribution rather than an unverified control signal for others\. Failure begins when inter\-agent communication is treated as trusted reasoning by default\. Multi\-agent systems add channels that do not exist in single\-agent settings\. Messages can be forwarded, summarized, amplified, and stored in shared memory\. As a result, one local failure may become a system\-level failure, as shown by prompt\-infection, malicious\-agent, and propagation studiesLee and Tiwari \([2024](https://arxiv.org/html/2607.12406#bib.bib65)\); Wang et al\. \([2024c](https://arxiv.org/html/2607.12406#bib.bib124)\); He et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib43)\); Wang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib120)\)\.

### 5\.2Prompt Infection and Communication Attacks

The most direct failure mode at this boundary is communication\-borne prompt injection\.*Prompt Infection*showed that one compromised agent can pass malicious instructions to others, turning ordinary coordination into an attack channelLee and Tiwari \([2024](https://arxiv.org/html/2607.12406#bib.bib65)\)\. A message is not just information\. It can also be a carrier of control\. Recent work shows that this problem becomes stronger in realistic collaborative settings\. Debate\-based attacksAmayuelas et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib4)\), manipulated\-knowledge floodingJu et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib60)\), and explicit communication attacksHe et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib43)\)all show that harmful content can spread through discussion, critique, and consensus formation rather than through a single direct override\. CORBAZhou et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib162)\)pushes this further by studying recursive blocking behavior, where one malicious pattern can propagate and suppress useful coordination across the network\. The progression is clear: first one compromised message, then repeated propagation, then communication\-level cascade\.

### 5\.3Malicious Agents, Topology, and Cascade Failures

Another major line of work asks why some multi\-agent systems fail locally while others fail systemically\. One answer is topology\. Network structure, routing rules, and memory sharing determine how fast compromise travels and how hard it is to contain\. Work on malicious\-agent resiliencetse Huang et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib114)\), NetSafeYu et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib143)\), G\-SafeguardWang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib120)\), and communication\-aware multi\-agent risksHammond et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib42)\)supports this view\. Backdoor and memory attacks make this issue more concrete\. BadAgent and*Watch Out for Your Agents\!*show that malicious behavior can be implanted into agent workflows and activated later through interaction patternsWang et al\. \([2024c](https://arxiv.org/html/2607.12406#bib.bib124)\); Yang et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib139)\)\. Trojan Hippo then shows that shared memory is itself a high\-risk surface, because poisoned memory can be weaponized for later exfiltration or controlDas et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib29)\)\. The threat moves from bad messages, to bad agents, to bad shared state\. Once memory and topology are involved, rollback becomes much harder than in a single\-agent system\.

### 5\.4Defense Agents, Attribution, and Memory Partitioning

Recent defenses increasingly treat coordination as a security problem rather than only an efficiency problem\. GuardAgentXiang et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib133)\), AutoDefenseZeng et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib145)\), ShieldAgentChen et al\. \([2025d](https://arxiv.org/html/2607.12406#bib.bib23)\), and related multi\-agent defense pipelinesHossain et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib47)\)introduce dedicated safety roles that inspect, verify, or challenge other agents before outputs are accepted\. The main idea is simple: not all agents should have equal authority, and safety should be assigned explicit responsibility inside the workflow\. Another important trend is attribution\. Recent work asks not just whether a multi\-agent system failed, but which agent, which message, and which step caused the decisive failureZhang et al\. \([2025d](https://arxiv.org/html/2607.12406#bib.bib152)\)\. This matters because containment is difficult without diagnosis\. Structural defenses then push one step further\. AgentSafeMao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib82)\)uses hierarchical data management to reduce unsafe cross\-agent sharing, while G\-SafeguardWang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib120)\)monitors the interaction graph for anomalous propagation patterns\. This line of work suggests that memory partitioning, privilege separation, and topology\-aware monitoring may be more durable than generic prompt filtering alone\.

Overall, this boundary shows that collaboration is not automatically a safeguard\. Without isolation, it can become a mechanism for amplification\.

## 6System\-Environment Boundary

### 6\.1Threat Model and Boundary Definition

The system\-environment boundary concerns how an agent reads and reacts to the outside world\. In a well\-isolated system, webpages, retrieved passages, documents, emails, interface elements, and memory artifacts should remain observations rather than hidden commands\. Failure begins when environment\-originated content is absorbed into the agent context and then treated as if it had authority\.

The environment is often broad and weakly controlled\. Once outside content can steer reasoning, retrieval, or action, the outside world becomes part of the control loop, as shown by indirect injection and retrieval\-side studiesAbdelnabi et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib1)\); Zverev et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib167)\); Chang et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib13)\); Zhu et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib165)\)\.

### 6\.2Indirect Prompt Injection from Ambient Content

The starting point of this literature is indirect prompt injection\. Early work showed that malicious instructions can be hidden in webpages, emails, or documents and later executed by an LLM\-integrated system when the model reads themAbdelnabi et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib1)\)\. Recent work shows that this problem is broader and more persistent than first expected\. Automatic attacksLiu et al\. \([2024c](https://arxiv.org/html/2607.12406#bib.bib75)\), multimodal injectionsBagdasaryan et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib9)\); Wu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib131)\), and environment\-facing attacks on web agentsLiao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib71)\); Xu et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib136)\)all show that the environment can steer the agent without going through the nominal user channel\.*Overcoming the Retrieval Barrier*Chang et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib13)\)and*Your Agent is More Brittle Than You Think*Zhu et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib165)\)further show that indirect injection can survive realistic retrieval pipelines and remain effective even when the injected content appears peripheral\. The main challenge is therefore not only detecting bad strings, but preserving source separation after retrieval and context assembly\.

### 6\.3Web\-Agent Security and Hostile Interface Environments

This boundary becomes even sharper in web and computer\-use agents\. Here the environment is not just read\. It is also clicked, navigated, copied, and executed against\. INJECAGENTZhan et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib146)\), WASPEvtimov et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib33)\), and WebInjectWang et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib122)\)show that browser agents are highly vulnerable when malicious prompts are embedded in realistic websites and task flows\. Work on WIPIWu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib131)\), environmental injection for privacy leakageLiao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib71)\), and broader analyses of web\-agent fragilityChiang et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib27)\)reaches the same conclusion from slightly different angles: web environments are active adversarial surfaces, not passive information sources\. VPI\-BenchCao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib11)\)extends this picture to visual prompt injection, showing that computer\-use agents can also be misled through interface appearance rather than text alone\.

### 6\.4RAG Poisoning, Retrieval Corruption, and Disclosure Risks

RAG systems expose another major form of system\-environment failure\. Here the issue is often not explicit instruction, but corrupted evidence\. PoisonedRAGZou et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib166)\), PandoraDeng et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib31)\), BADRAGXue et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib138)\), and human\-imperceptible retrieval poisoningZhang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib150)\)show that attackers can manipulate knowledge bases or retrieved passages so that the agent reasons from compromised support\. Recent work also shows that retrieval systems create disclosure risks in addition to integrity risks\. Data extraction attacksPeng et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib92)\), backdoored retrieval databasesQi et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib96)\), agent\-based exfiltration attacksJiang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib57)\), membership inference studiesLi et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib70)\); Anderson et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib5)\), and multimodal leakage evaluationsAl\-Lawati and Wang \([2026](https://arxiv.org/html/2607.12406#bib.bib2)\)show that the environment can both push corrupted knowledge in and pull private knowledge out\. AgentPoison and MEMSAD further show that once retrieval outputs or memory stores are corrupted, the effect may continue across later interactions unless the system actively diagnoses and repairs the damageChen et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib24)\); Gowda \([2026](https://arxiv.org/html/2607.12406#bib.bib40)\)\.

### 6\.5Authentication, Provenance, and Causal Defenses

Recent defenses increasingly try to restore explicit trust structure at this boundary\. FATHWang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib118)\), SpotlightingHines et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib46)\), instruction\-detection methodsWen et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib129)\); Chen et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib22)\), and Task ShieldJia et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib54)\)all aim to mark, isolate, or filter environment\-originated instructions before they silently become control input\. More recent work moves toward stronger causal and provenance\-aware defenses\. AgentSentryZhang et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib153)\)uses temporal causal diagnostics, AttriGuardHe et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib45)\)uses causal attribution, and TrustRAGZhou et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib159)\)emphasizes trustworthy evidence selection and robustness\-aware retrieval\. The broader lesson is that system\-environment safety will likely depend less on stronger generic refusal and more on better source authentication, provenance tracking, memory hygiene, and context attribution\.

In short, this boundary suggests that safe agent systems need not only aligned models, but also environment interfaces that preserve isolation by construction\.

## 7Cross\-Boundary Challenges

### 7\.1How Failures Propagate Across Boundaries

Serious failures often cross boundaries\. A common pattern is sequential escalation: user input may first override control at the user\-agent boundary, then steer tool use, and finally trigger unsafe execution\. Another pattern starts from the environment\. A malicious webpageAbdelnabi et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib1)\), retrieved passage, or poisoned memory itemChen et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib24)\); Das et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib29)\)may enter through the system\-environment boundary and later propagate into tool calls, agent communicationLee and Tiwari \([2024](https://arxiv.org/html/2607.12406#bib.bib65)\); He et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib43)\), or action traces\. In multi\-agent systems, the spread can continue because compromised results may be forwarded or stored for reuseHe et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib43)\); Wang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib120)\); Mao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib82)\)\. The main unit of analysis is therefore often the full control path, not a single prompt, tool call, or action\. This is also why local robustness at one interface does not guarantee system\-level safety\.

### 7\.2Isolation\-by\-Construction

Safer agent systems need interfaces that preserve separation by design\. Inputs from users, tools, peer agents, and external content should remain distinguishable\. Capability access should be scoped\. Propagation paths should be observable through trace\-level monitoringMou et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib85)\), attributionChen et al\. \([2025d](https://arxiv.org/html/2607.12406#bib.bib23)\); He et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib45)\), and policy checksZhang et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib153)\)\. Recovery should also be a core requirement once compromise reaches memory or shared stateMao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib82)\); Gowda \([2026](https://arxiv.org/html/2607.12406#bib.bib40)\)\. In practice, this means that systems should make trust, authority, and capability boundaries explicit throughout the workflow\. In short, isolation\-by\-construction means building interfaces where the boundary remains visible to the system itself\.

### 7\.3Open Challenges

Open problems remain in evaluation, compositional defense, persistent state, and formalization\. Many benchmarks still test single boundaries, while real failures are cross\-boundaryTur et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib115)\); Evtimov et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib33)\); Cui et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib28)\)\. Defenses may fail after information is transformed downstream\. MemoryGowda \([2026](https://arxiv.org/html/2607.12406#bib.bib40)\), retrieval cachesChen et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib24)\), and shared agent stateDas et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib29)\)further obscure and persist compromises\. The field still lacks stable abstractions for authority, trust, and privilege in full workflowsSiu et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib107)\); Li et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib68)\); Kim et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib63)\)\.

## 8Conclusion

In this survey, we argue that LLM\-agent safety is best understood through boundaries\. Across user input, tool use, execution, inter\-agent communication, and environment interaction, failures often share a common pattern: loss of isolation at interfaces\. This helps explain where compromise begins and how it propagates\. Overall, safer agent systems require not only better alignment, but also clearer trust boundaries, scoped capabilities, runtime mediation, and stronger control of memory and context\.

## Limitations

This survey has several limitations\. First, the field is moving quickly, so any taxonomy can become incomplete as new systems and attacks appear\. Second, our taxonomy is boundary\-centric by design\. It is useful for isolation failures, but it is not the only valid way to organize the literature\. Third, many papers are naturally cross\-boundary\. We classify them by the boundary where the decisive loss of isolation first occurs, but this still requires judgment\. Finally, the literature is uneven across boundaries, so some parts of the survey are denser than others\.

## Ethical Considerations

All authors of this paper affirm their adherence to the ACM Code of Ethics and the ACL Code of Conduct\. This survey is intended to support safer research and system design for LLM agents\.

At the same time, organizing the attack and defense landscape of agent systems has a dual\-use aspect\. We therefore focus on high\-level mechanisms and design implications rather than operational attack instructions\.

## References

- Abdelnabi et al\. \(2023\)Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz\. 2023\.[Not What You’ve Signed Up For: Compromising real\-world llm\-integrated applications with indirect prompt injection\.](https://doi.org/10.1145/3605764.3623985)In*AISec@CCS*, pages 79–90\.
- Al\-Lawati and Wang \(2026\)Ali Al\-Lawati and Suhang Wang\. 2026\.[Do multimodal rag systems leak data? a comprehensive evaluation of membership inference and image caption retrieval attacks](https://arxiv.org/abs/2601.17644)\.*Preprint*, arXiv:2601\.17644\.
- Alon and Kamfonas \(2023\)Gabriel Alon and Michael Kamfonas\. 2023\.[Detecting language model attacks with perplexity\.](https://doi.org/10.48550/arXiv.2308.14132)*CoRR*\.
- Amayuelas et al\. \(2024\)Alfonso Amayuelas, Xianjun Yang, Antonis Antoniades, Wenyue Hua, Liangming Pan, and William Yang Wang\. 2024\.[MultiAgent Collaboration Attack: Investigating adversarial attacks in large language model collaborations via debate\.](https://doi.org/10.18653/v1/2024.findings-emnlp.407)In*EMNLP \(Findings\)*, pages 6929–6948\.
- Anderson et al\. \(2025\)Maya Anderson, Guy Amit, and Abigail Goldsteen\. 2025\.[Is my data in your retrieval database? membership inference attacks against retrieval augmented generation\.](https://doi.org/10.5220/0013108300003899)In*ICISSP \(2\)*, pages 474–485\.
- Anthropic \(2025\)Anthropic\. 2025\.Claude code: Best practices for agentic coding\.[https://www\.anthropic\.com/engineering/claude\-code\-best\-practices](https://www.anthropic.com/engineering/claude-code-best-practices)\.
- Anthropic \(2026\)Anthropic\. 2026\.Scaling managed agents: Decoupling the brain from the hands\.[https://www\.anthropic\.com/engineering/managed\-agents](https://www.anthropic.com/engineering/managed-agents)\.
- Ba et al\. \(2026\)Lei Ba, Qinbin Li, and Songze Li\. 2026\.[CIBER: A comprehensive benchmark for security evaluation of code interpreter agents\.](https://doi.org/10.48550/arXiv.2602.19547)*CoRR*\.
- Bagdasaryan et al\. \(2023\)Eugene Bagdasaryan, Tsung\-Yin Hsieh, Ben Nassi, and Vitaly Shmatikov\. 2023\.[\(ab\)using images and sounds for indirect instruction injection in multi\-modal llms\.](https://doi.org/10.48550/arXiv.2307.10490)*CoRR*\.
- Candogan et al\. \(2025\)Leyla Naz Candogan, Yongtao Wu, Elías Abad\-Rocamora, Grigorios Chrysos, and Volkan Cevher\. 2025\.[Single\-pass detection of jailbreaking input in large language models\.](https://openreview.net/forum?id=42v6I5Ut9a)*Trans\. Mach\. Learn\. Res\.*
- Cao et al\. \(2025\)Tri Cao, Bennett Lim, Yue Liu, Yuan Sui, Yuexin Li, Shumin Deng, Lin Lu, Nay Oo, Shuicheng Yan, and Bryan Hooi\. 2025\.[VPI\-Bench: Visual prompt injection attacks for computer\-use agents\.](https://doi.org/10.48550/arXiv.2506.02456)*CoRR*\.
- Chakraborty et al\. \(2025\)Trishna Chakraborty, Udita Ghosh, Xiaopan Zhang, Fahim Faisal Niloy, Yue Dong, Jiachen Li, Amit Roy\-Chowdhury, and Chengyu Song\. 2025\.[HEAL: An empirical study on hallucinations in embodied agents driven by large language models\.](https://aclanthology.org/2025.findings-emnlp.1158/)In*EMNLP \(Findings\)*, pages 21226–21243\.
- Chang et al\. \(2026\)Hongyan Chang, Ergute Bao, Xinjian Luo, and Ting Yu\. 2026\.[Overcoming the Retrieval Barrier: Indirect prompt injection in the wild for llm systems\.](https://doi.org/10.48550/arXiv.2601.07072)*CoRR*\.
- Chang et al\. \(2024\)Zhiyuan Chang, Mingyang Li, Yi Liu, Junjie Wang, Qing Wang, and Yang Liu\. 2024\.[Play Guessing Game with LLM: Indirect jailbreak attack with implicit clues\.](https://doi.org/10.18653/v1/2024.findings-acl.304)In*ACL \(Findings\)*, pages 5135–5147\.
- Chao et al\. \(2024\)Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J\. Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong\. 2024\.[JailbreakBench: An open robustness benchmark for jailbreaking large language models\.](http://papers.nips.cc/paper_files/paper/2024/hash/63092d79154adebd7305dfd498cbff70-Abstract-Datasets_and_Benchmarks_Track.html)In*NeurIPS*\.
- Chao et al\. \(2025\)Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J\. Pappas, and Eric Wong\. 2025\.[Jailbreaking black box large language models in twenty queries\.](https://doi.org/10.1109/SaTML64287.2025.00010)In*SaTML*, pages 23–42\.
- Chen et al\. \(2025a\)Chaoran Chen, Zhiping Zhang, Bingcan Guo, Shang Ma, Ibrahim Khalilov, Simret Araya Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, and Toby Jia\-Jun Li\. 2025a\.The Obvious Invisible Threat: Llm\-powered GUI agents’ vulnerability to fine\-print injections\.*CoRR*, abs/2504\.11281\.
- Chen and Cong \(2025\)Jizhou Chen and Samuel Lee Cong\. 2025\.[AgentGuard: Repurposing agentic orchestrator for safety evaluation of tool orchestration\.](https://doi.org/10.48550/arXiv.2502.09809)*CoRR*\.
- Chen et al\. \(2024a\)Kexin Chen, Yi Liu, Dongxia Wang, Jiaying Chen, and Wenhai Wang\. 2024a\.[Characterizing and evaluating the reliability of llms against jailbreak attacks\.](https://doi.org/10.48550/arXiv.2408.09326)*CoRR*\.
- Chen et al\. \(2025b\)Sizhe Chen, Julien Piet, Chawin Sitawarin, and David A\. Wagner\. 2025b\.[StruQ: Defending against prompt injection with structured queries\.](https://www.usenix.org/conference/usenixsecurity25/presentation/chen-sizhe)In*USENIX Security Symposium*, pages 2383–2400\.
- Chen et al\. \(2026\)Yen\-Shan Chen, Sian\-Yao Huang, Cheng\-Lin Yang, and Yun\-Nung Chen\. 2026\.[TraceSafe: A systematic assessment of llm guardrails on multi\-step tool\-calling trajectories\.](https://doi.org/10.48550/arXiv.2604.07223)*CoRR*\.
- Chen et al\. \(2025c\)Yulin Chen, Haoran Li, Yuan Sui, Yufei He, Yue Liu, Yangqiu Song, and Bryan Hooi\. 2025c\.[Can indirect prompt injection attacks be detected and removed?](https://aclanthology.org/2025.acl-long.890/)In*ACL \(1\)*, pages 18189–18206\.
- Chen et al\. \(2025d\)Zhaorun Chen, Mintong Kang, and Bo Li\. 2025d\.[ShieldAgent: Shielding agents via verifiable safety policy reasoning\.](https://proceedings.mlr.press/v267/chen25ae.html)In*ICML*\.
- Chen et al\. \(2024b\)Zhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song, and Bo Li\. 2024b\.[AgentPoison: Red\-teaming llm agents via poisoning memory or knowledge bases\.](http://papers.nips.cc/paper_files/paper/2024/hash/eb113910e9c3f6242541c1652e30dfd6-Abstract-Conference.html)In*NeurIPS*\.
- Cheng et al\. \(2024a\)Hao Cheng, Erjia Xiao, Chengyuan Yu, Zhao Yao, Jiahang Cao, Qiang Zhang, Jiaxu Wang, Mengshu Sun, Kaidi Xu, Jindong Gu, and Renjing Xu\. 2024a\.[Manipulation Facing Threats: Evaluating physical vulnerabilities in end\-to\-end vision language action models\.](https://doi.org/10.48550/arXiv.2409.13174)*CoRR*\.
- Cheng et al\. \(2024b\)Yixin Cheng, Markos Georgopoulos, Volkan Cevher, and Grigorios G\. Chrysos\. 2024b\.[Leveraging the context through multi\-round interactions for jailbreaking attacks\.](https://doi.org/10.48550/arXiv.2402.09177)*CoRR*\.
- Chiang et al\. \(2025\)Jeffrey Yang Fan Chiang, Seungjae Lee, Jia\-Bin Huang, Furong Huang, and Yizheng Chen\. 2025\.[Why are web ai agents more vulnerable than standalone llms? a security analysis\.](https://doi.org/10.48550/arXiv.2502.20383)*CoRR*\.
- Cui et al\. \(2026\)Fan Cui, Hongyuan Hou, Zizhang Luo, Chenyun Yin, and Yun Liang\. 2026\.[HWE\-Bench: Benchmarking llm agents on real\-world hardware bug repair tasks](https://arxiv.org/abs/2604.14709)\.*Preprint*, arXiv:2604\.14709\.
- Das et al\. \(2026\)Debeshee Das, Julien Piet, Darya Kaviani, Luca Beurer\-Kellner, Florian Tramèr, and David Wagner\. 2026\.[Trojan hippo: Weaponizing agent memory for data exfiltration](https://arxiv.org/abs/2605.01970)\.*Preprint*, arXiv:2605\.01970\.
- Debenedetti et al\. \(2025\)Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, and Florian Tramèr\. 2025\.[Defeating prompt injections by design\.](https://doi.org/10.48550/arXiv.2503.18813)*CoRR*\.
- Deng et al\. \(2024\)Gelei Deng, Yi Liu, Kailong Wang, Yuekang Li, Tianwei Zhang, and Yang Liu\. 2024\.[Pandora: Jailbreak gpts by retrieval augmented generation poisoning\.](https://doi.org/10.48550/arXiv.2402.08416)*CoRR*\.
- Dong et al\. \(2025\)Shen Dong, Shaochen Xu, Pengfei He, Yige Li, Jiliang Tang, Tianming Liu, Hui Liu, and Zhen Xiang\. 2025\.[A practical memory injection attack against llm agents\.](https://doi.org/10.48550/arXiv.2503.03704)*CoRR*\.
- Evtimov et al\. \(2025\)Ivan Evtimov, Arman Zharmagambetov, Aaron Grattafiori, Chuan Guo, and Kamalika Chaudhuri\. 2025\.[WASP: Benchmarking web agent security against prompt injection attacks\.](https://doi.org/10.48550/arXiv.2504.18575)*CoRR*\.
- Fang et al\. \(2024a\)Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang\. 2024a\.[Llm agents can autonomously exploit one\-day vulnerabilities\.](https://doi.org/10.48550/arXiv.2404.08144)*CoRR*\.
- Fang et al\. \(2024b\)Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang\. 2024b\.[Llm agents can autonomously hack websites\.](https://doi.org/10.48550/arXiv.2402.06664)*CoRR*\.
- Fokou \(2026\)Joel Fokou\. 2026\.[Parallax: Why ai agents that think must never act](https://arxiv.org/abs/2604.12986)\.*Preprint*, arXiv:2604\.12986\.
- Fu et al\. \(2024\)Xiaohan Fu, Shuheng Li, Zihan Wang, Yihao Liu, Rajesh K\. Gupta, Taylor Berg\-Kirkpatrick, and Earlence Fernandes\. 2024\.[Imprompter: Tricking llm agents into improper tool use\.](https://doi.org/10.48550/arXiv.2410.14923)*CoRR*\.
- Gong et al\. \(2025\)Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang\. 2025\.[FigStep: Jailbreaking large vision\-language models via typographic visual prompts\.](https://doi.org/10.1609/aaai.v39i22.34568)In*AAAI*, pages 23951–23959\.
- Gou et al\. \(2024\)Yunhao Gou, Kai Chen, Zhili Liu, Lanqing Hong, Hang Xu, Zhenguo Li, Dit\-Yan Yeung, James T\. Kwok, and Yu Zhang\. 2024\.[Eyes Closed, Safety on: Protecting multimodal llms via image\-to\-text transformation\.](https://doi.org/10.1007/978-3-031-72643-9_23)In*ECCV \(17\)*, pages 388–404\.
- Gowda \(2026\)Ishrith Gowda\. 2026\.[MEMSAD: Gradient\-coupled anomaly detection for memory poisoning in retrieval\-augmented agents](https://arxiv.org/abs/2605.03482)\.*Preprint*, arXiv:2605\.03482\.
- Guo et al\. \(2024\)Chengquan Guo, Xun Liu, Chulin Xie, Andy Zhou, Yi Zeng, Zinan Lin, Dawn Song, and Bo Li\. 2024\.[RedCode: Risky code execution and generation benchmark for code agents\.](http://papers.nips.cc/paper_files/paper/2024/hash/bfd082c452dffb450d5a5202b0419205-Abstract-Datasets_and_Benchmarks_Track.html)In*NeurIPS*\.
- Hammond et al\. \(2025\)Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher\-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, Wolfram Barfuss, Jakob N\. Foerster, Tomas Gavenciak, The Anh Han, Edward Hughes, Vojtech Kovarík, Jan Kulveit, Joel Z\. Leibo, Caspar Oesterheld, Christian Schröder de Witt, Nisarg Shah, Michael P\. Wellman, and 25 others\. 2025\.[Multi\-agent risks from advanced ai\.](https://doi.org/10.48550/arXiv.2502.14143)*CoRR*\.
- He et al\. \(2025a\)Pengfei He, Yuping Lin, Shen Dong, Han Xu, Yue Xing, and Hui Liu\. 2025a\.[Red\-teaming llm multi\-agent systems via communication attacks\.](https://aclanthology.org/2025.findings-acl.349/)In*ACL \(Findings\)*, pages 6726–6747\.
- He et al\. \(2025b\)Pengfei He, Han Xu, Yue Xing, Hui Liu, Makoto Yamada, and Jiliang Tang\. 2025b\.[Data poisoning for in\-context learning\.](https://doi.org/10.18653/v1/2025.findings-naacl.91)In*NAACL \(Findings\)*, pages 1680–1700\.
- He et al\. \(2026\)Yu He, Haozhe Zhu, Yiming Li, Shuo Shao, Hongwei Yao, Zhihao Liu, and Zhan Qin\. 2026\.AttriGuard: Defeating indirect prompt injection in LLM agents via causal attribution of tool invocations\.*CoRR*, abs/2603\.10749\.
- Hines et al\. \(2024\)Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman\. 2024\.[Defending against indirect prompt injection attacks with spotlighting\.](https://ceur-ws.org/Vol-3920/paper03.pdf)In*CAMLIS*, pages 48–62\.
- Hossain et al\. \(2025\)S\. M\. Asif Hossain, Ruksat Khan Shayoni, Mohd Ruhul Ameen, Akif Islam, Muhammad Firoz Mridha, and Jungpil Shin\. 2025\.[A multi\-agent llm defense pipeline against prompt injection attacks\.](https://doi.org/10.48550/arXiv.2509.14285)*CoRR*\.
- Hu et al\. \(2024a\)Xiaomeng Hu, Pin\-Yu Chen, and Tsung\-Yi Ho\. 2024a\.[Gradient Cuff: Detecting jailbreak attacks on large language models by exploring refusal loss landscapes\.](http://papers.nips.cc/paper_files/paper/2024/hash/e46984e056185b21ddb1e7973c365f14-Abstract-Conference.html)In*NeurIPS*\.
- Hu et al\. \(2024b\)Zhe Hu, Yixiao Ren, Jing Li, and Yu Yin\. 2024b\.[VIVA: A benchmark for vision\-grounded decision\-making with human values\.](https://doi.org/10.18653/v1/2024.emnlp-main.137)In*EMNLP*, pages 2294–2311\.
- Huang et al\. \(2024a\)Tiansheng Huang, Gautam Bhattacharya, Pratik Joshi, Josh Kimball, and Ling Liu\. 2024a\.[Antidote: Post\-fine\-tuning safety alignment for large language models against harmful fine\-tuning\.](https://doi.org/10.48550/arXiv.2408.09600)*CoRR*\.
- Huang et al\. \(2024b\)Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen\. 2024b\.[Catastrophic jailbreak of open\-source llms via exploiting generation\.](https://openreview.net/forum?id=r42tSSCHPh)In*ICLR*\.
- Jacob et al\. \(2025\)Dennis Jacob, Emad Alghamdi, Zhanhao Hu, Basel Alomair, and David A\. Wagner\. 2025\.[Better privilege separation for agents by restricting data types\.](https://doi.org/10.48550/arXiv.2509.25926)*CoRR*\.
- Ji et al\. \(2025\)Jiabao Ji, Bairu Hou, Alexander Robey, George J\. Pappas, Hamed Hassani, Yang Zhang, Eric Wong, and Shiyu Chang\. 2025\.[Defending large language models against jailbreak attacks via semantic smoothing\.](https://aclanthology.org/2025.ijcnlp-short.2/)In*IJCNLP\-AACL \(Short Papers\)*, pages 7–40\.
- Jia et al\. \(2025a\)Feiran Jia, Tong Wu, Xin Qin, and Anna Cinzia Squicciarini\. 2025a\.[The Task Shield: Enforcing task alignment to defend against indirect prompt injection in llm agents\.](https://aclanthology.org/2025.acl-long.1435/)In*ACL \(1\)*, pages 29680–29697\.
- Jia et al\. \(2025b\)Xiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang, Jindong Gu, Yang Liu, Xiaochun Cao, and Min Lin\. 2025b\.[Improved techniques for optimization\-based jailbreaking on large language models\.](https://openreview.net/forum?id=e9yfCY7Q3U)In*ICLR*\.
- Jiang et al\. \(2024a\)Bojian Jiang, Yi Jing, Tianhao Shen, Qing Yang, and Deyi Xiong\. 2024a\.[DART: Deep adversarial automated red teaming for llm safety\.](https://doi.org/10.48550/arXiv.2407.03876)*CoRR*\.
- Jiang et al\. \(2024b\)Changyue Jiang, Xudong Pan, Geng Hong, Chenfu Bao, and Min Yang\. 2024b\.[RAG\-Thief: Scalable extraction of private data from retrieval\-augmented generation applications with agent\-based attacks\.](https://doi.org/10.48550/arXiv.2411.14110)*CoRR*\.
- Jiao et al\. \(2025\)Ruochen Jiao, Shaoyuan Xie, Justin Yue, Takami Sato, Lixu Wang, Yixuan Wang, Qi Alfred Chen, and Qi Zhu\. 2025\.[Can we trust embodied agents? exploring backdoor attacks against embodied llm\-based decision\-making systems\.](https://openreview.net/forum?id=S1Bv3068Xt)In*ICLR*\.
- Jing et al\. \(2025\)Huihao Jing, Haoran Li, Wenbin Hu, Qi Hu, Heli Xu, Tianshu Chu, Peizhao Hu, and Yangqiu Song\. 2025\.[MCIP: Protecting mcp safety via model contextual integrity protocol\.](https://doi.org/10.18653/v1/2025.emnlp-main.62)In*EMNLP*, pages 1177–1194\.
- Ju et al\. \(2024\)Tianjie Ju, Yiting Wang, Xinbei Ma, Pengzhou Cheng, Haodong Zhao, Yulong Wang, Lifeng Liu, Jian Xie, Zhuosheng Zhang, and Gongshen Liu\. 2024\.[Flooding spread of manipulated knowledge in llm\-based multi\-agent communities\.](https://doi.org/10.48550/arXiv.2407.07791)*CoRR*\.
- Karnik et al\. \(2024\)Sathwik Karnik, Zhang\-Wei Hong, Nishant Abhangi, Yen\-Chen Lin, Tsun\-Hsuan Wang, and Pulkit Agrawal\. 2024\.[Embodied red teaming for auditing robotic foundation models\.](https://doi.org/10.48550/arXiv.2411.18676)*CoRR*\.
- Kim et al\. \(2023\)Jinhwa Kim, Ali Derakhshan, and Ian G\. Harris\. 2023\.[Robust Safety Classifier for Large Language Models: Adversarial prompt shield\.](https://doi.org/10.48550/arXiv.2311.00172)*CoRR*\.
- Kim et al\. \(2026\)Juhee Kim, Xiaoyuan Liu, Zhun Wang, Shi Qiu, Bo Li, Wenbo Guo, and Dawn Song\. 2026\.The attack and defense landscape of agentic AI: A comprehensive survey\.*CoRR*, abs/2603\.11088\.
- Kumar et al\. \(2024\)Priyanshu Kumar, Elaine Lau, Saranya Vijayakumar, Tu Trinh, Scale Red Team, Elaine T\. Chang, Vaughn Robinson, Sean Hendryx, Shuyan Zhou, Matt Fredrikson, Summer Yue, and Zifan Wang\. 2024\.[Refusal\-trained llms are easily jailbroken as browser agents\.](https://doi.org/10.48550/arXiv.2410.13886)*CoRR*\.
- Lee and Tiwari \(2024\)Donghyun Lee and Mo Tiwari\. 2024\.[Prompt Infection: Llm\-to\-llm prompt injection within multi\-agent systems\.](https://doi.org/10.48550/arXiv.2410.07283)*CoRR*\.
- Levy et al\. \(2024\)Ido Levy, Ben Wiesel, Sami Marreed, Alon Oved, Avi Yaeli, and Segev Shlomov\. 2024\.[ST\-WebAgentBench: A benchmark for evaluating safety and trustworthiness in web agents\.](https://doi.org/10.48550/arXiv.2410.06703)*CoRR*\.
- Li et al\. \(2024a\)Manling Li, Shiyu Zhao, Qineng Wang, Kangrui Wang, Yu Zhou, Sanjana Srivastava, Cem Gokmen, Tony Lee, Li Erran Li, Ruohan Zhang, Weiyu Liu, Percy Liang, Li Fei\-Fei, Jiayuan Mao, and Jiajun Wu\. 2024a\.[Embodied Agent Interface: Benchmarking llms for embodied decision making\.](http://papers.nips.cc/paper_files/paper/2024/hash/b631da756d1573c24c9ba9c702fde5a9-Abstract-Datasets_and_Benchmarks_Track.html)In*NeurIPS*\.
- Li et al\. \(2026\)Ninghui Li, Kaiyuan Zhang, Kyle Polley, and Jerry Ma\. 2026\.[Security considerations for artificial intelligence agents\.](https://doi.org/10.48550/arXiv.2603.12230)*CoRR*\.
- Li et al\. \(2024b\)Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun Sun\. 2024b\.[BackdoorLLM: A comprehensive benchmark for backdoor attacks on large language models\.](https://doi.org/10.48550/arXiv.2408.12798)*CoRR*\.
- Li et al\. \(2025\)Yuying Li, Gaoyang Liu, Chen Wang, and Yang Yang\. 2025\.[Generating Is Believing: Membership inference attacks against retrieval\-augmented generation\.](https://doi.org/10.1109/ICASSP49660.2025.10889013)In*ICASSP*, pages 1–5\.
- Liao et al\. \(2025\)Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun\. 2025\.[Eia: Environmental injection attack on generalist web agents for privacy leakage\.](https://openreview.net/forum?id=xMOLUzo2Lk)In*ICLR*\.
- Lin et al\. \(2025\)Huawei Lin, Yingjie Lao, Tong Geng, Tan Yu, and Weijie Zhao\. 2025\.[UniGuardian: A unified defense for detecting prompt injection, backdoor attacks and adversarial attacks in large language models\.](https://doi.org/10.48550/arXiv.2502.13141)*CoRR*\.
- Liu et al\. \(2024a\)Aishan Liu, Yuguang Zhou, Xianglong Liu, Tianyuan Zhang, Siyuan Liang, Jiakai Wang, Yanjun Pu, Tianlin Li, Junqi Zhang, Wenbo Zhou, Qing Guo, and Dacheng Tao\. 2024a\.[Compromising embodied agents with contextual backdoor attacks\.](https://doi.org/10.48550/arXiv.2408.02882)*CoRR*\.
- Liu et al\. \(2024b\)Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao\. 2024b\.[AutoDAN: Generating stealthy jailbreak prompts on aligned large language models\.](https://openreview.net/forum?id=7Jwpw4qKkb)In*ICLR*\.
- Liu et al\. \(2024c\)Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and Chaowei Xiao\. 2024c\.[Automatic and universal prompt injection attacks against large language models\.](https://doi.org/10.48550/arXiv.2403.04957)*CoRR*\.
- Liu et al\. \(2023a\)Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang, and Yu Qiao\. 2023a\.[Query\-relevant images jailbreak large multi\-modal models\.](https://doi.org/10.48550/arXiv.2311.17600)*CoRR*\.
- Liu et al\. \(2023b\)Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu\. 2023b\.[Prompt injection attack against llm\-integrated applications\.](https://doi.org/10.48550/arXiv.2306.05499)*CoRR*\.
- Lu et al\. \(2024a\)Weikai Lu, Ziqian Zeng, Jianwei Wang, Zhengdong Lu, Zelin Chen, Huiping Zhuang, and Cen Chen\. 2024a\.[Eraser: Jailbreaking defense in large language models via unlearning harmful knowledge\.](https://doi.org/10.48550/arXiv.2404.05880)*CoRR*\.
- Lu et al\. \(2025\)Xiaoya Lu, Dongrui Liu, Yi Yu, Luxin Xu, and Jing Shao\. 2025\.[X\-Boundary: Establishing exact safety boundary to shield llms from multi\-turn jailbreaks without compromising usability\.](https://doi.org/10.48550/arXiv.2502.09990)*CoRR*\.
- Lu et al\. \(2024b\)Xuancun Lu, Zhengxian Huang, Xinfeng Li, Xiaoyu Ji, and Wenyuan Xu\. 2024b\.POEX: Policy executable embodied AI jailbreak attacks\.*CoRR*, abs/2412\.16633\.
- Maiti \(2026\)Saikat Maiti\. 2026\.[Caging the Agents: A zero trust security architecture for autonomous ai in healthcare\.](https://doi.org/10.48550/arXiv.2603.17419)*CoRR*\.
- Mao et al\. \(2025\)Junyuan Mao, Fanci Meng, Yifan Duan, Miao Yu, Xiaojun Jia, Junfeng Fang, Yuxuan Liang, Kun Wang, and Qingsong Wen\. 2025\.[AgentSafe: Safeguarding large language model\-based multi\-agent systems via hierarchical data management\.](https://doi.org/10.48550/arXiv.2503.04392)*CoRR*\.
- Mazeika et al\. \(2024\)Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David A\. Forsyth, and Dan Hendrycks\. 2024\.[HarmBench: A standardized evaluation framework for automated red teaming and robust refusal\.](https://proceedings.mlr.press/v235/mazeika24a.html)In*ICML*, pages 35181–35224\.
- Meng et al\. \(2026\)Qianyu Meng, Yanan Wang, Liyi Chen, Yihang Li, Wei Wu, Wenyuan Jiang, Qimeng Wang, Chengqiang Lu, Yan Gao, Yi Wu, and Yao Hu\. 2026\.[Agent harness for large language model agents: A survey](https://doi.org/10.20944/preprints202604.0428.v3)\.*Preprints*\.
- Mou et al\. \(2026\)Yutao Mou, Zhangchi Xue, Lijun Li, Peiyang Liu, Shikun Zhang, Wei Ye, and Jing Shao\. 2026\.[ToolSafe: Enhancing tool invocation safety of llm\-based agents via proactive step\-level guardrail and feedback](https://arxiv.org/abs/2601.10156)\.*Preprint*, arXiv:2601\.10156\.
- Mou et al\. \(2024\)Yutao Mou, Shikun Zhang, and Wei Ye\. 2024\.[SG\-Bench: Evaluating llm safety generalization across diverse tasks and prompt types\.](http://papers.nips.cc/paper_files/paper/2024/hash/de7b99107c53e60257c727dc73daf1d1-Abstract-Datasets_and_Benchmarks_Track.html)In*NeurIPS*\.
- OpenAI \(2025\)OpenAI\. 2025\.Introducing codex\.[https://openai\.com/index/introducing\-codex/](https://openai.com/index/introducing-codex/)\.
- OpenClaw \(2026\)OpenClaw\. 2026\.Openclaw\.[https://github\.com/openclaw/openclaw](https://github.com/openclaw/openclaw)\.
- Pathmanathan et al\. \(2024\)Pankayaraj Pathmanathan, Souradip Chakraborty, Xiangyu Liu, Yongyuan Liang, and Furong Huang\. 2024\.[Is poisoning a real threat to llm alignment? maybe more so than you think\.](https://doi.org/10.48550/arXiv.2406.12091)*CoRR*\.
- Pedro et al\. \(2023\)Rodrigo Pedro, Daniel Castro, Paulo Carreira, and Nuno Santos\. 2023\.[From prompt injections to sql injection attacks: How protected is your llm\-integrated web application?](https://doi.org/10.48550/arXiv.2308.01990)*CoRR*\.
- Pei et al\. \(2025\)Aihua Pei, Zehua Yang, Shunan Zhu, Ruoxi Cheng, and Ju Jia\. 2025\.[SelfPrompt: Autonomously evaluating llm robustness via domain\-constrained knowledge guidelines and refined adversarial prompts\.](https://aclanthology.org/2025.coling-main.457/)In*COLING*, pages 6840–6854\.
- Peng et al\. \(2024\)Yuefeng Peng, Junda Wang, Hong Yu, and Amir Houmansadr\. 2024\.[Data extraction attacks in retrieval\-augmented generation via backdoors\.](https://doi.org/10.48550/arXiv.2411.01705)*CoRR*\.
- Perez and Ribeiro \(2022\)Fábio Perez and Ian Ribeiro\. 2022\.[Ignore previous prompt: Attack techniques for language models\.](https://doi.org/10.48550/arXiv.2211.09527)*CoRR*\.
- Piet et al\. \(2024\)Julien Piet, Maha Alrashed, Chawin Sitawarin, Sizhe Chen, Zeming Wei, Elizabeth Sun, Basel Alomair, and David A\. Wagner\. 2024\.[Jatmo: Prompt injection defense by task\-specific finetuning\.](https://doi.org/10.1007/978-3-031-70879-4_6)In*ESORICS \(1\)*, pages 105–124\.
- Qi et al\. \(2025a\)Xiangyu Qi, Boyi Wei, Nicholas Carlini, Yangsibo Huang, Tinghao Xie, Luxi He, Matthew Jagielski, Milad Nasr, Prateek Mittal, and Peter Henderson\. 2025a\.[On evaluating the durability of safeguards for open\-weight llms\.](https://openreview.net/forum?id=fXJCqdUSVG)In*ICLR*\.
- Qi et al\. \(2025b\)Zhenting Qi, Hanlin Zhang, Eric P\. Xing, Sham M\. Kakade, and Himabindu Lakkaraju\. 2025b\.[Follow My Instruction and Spill the Beans: Scalable data extraction from retrieval\-augmented generation systems\.](https://openreview.net/forum?id=Y4aWwRh25b)In*ICLR*\.
- Ran et al\. \(2024\)Delong Ran, Jinyuan Liu, Yichen Gong, Jingyi Zheng, Xinlei He, Tianshuo Cong, and Anyu Wang\. 2024\.[JailbreakEval: An integrated toolkit for evaluating jailbreak attempts against large language models\.](https://doi.org/10.48550/arXiv.2406.09321)*CoRR*\.
- Ren et al\. \(2024\)Qibing Ren, Hao Li, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, and Jing Shao\. 2024\.[Derail Yourself: Multi\-turn llm jailbreak attack through self\-discovered clues\.](https://doi.org/10.48550/arXiv.2410.10700)*CoRR*\.
- Ren \(2024\)Yupeng Ren\. 2024\.[F2A: An innovative approach for prompt injection by utilizing feign security detection agents\.](https://doi.org/10.48550/arXiv.2410.08776)*CoRR*\.
- Robey et al\. \(2025a\)Alexander Robey, Zachary Ravichandran, Vijay Kumar, Hamed Hassani, and George J\. Pappas\. 2025a\.[Jailbreaking llm\-controlled robots\.](https://doi.org/10.1109/ICRA55743.2025.11128119)In*ICRA*, pages 11948–11956\.
- Robey et al\. \(2025b\)Alexander Robey, Eric Wong, Hamed Hassani, and George J\. Pappas\. 2025b\.[SmoothLLM: Defending large language models against jailbreaking attacks\.](https://openreview.net/forum?id=laPAh2hRFC)*Trans\. Mach\. Learn\. Res\.*
- Russinovich et al\. \(2025\)Mark Russinovich, Ahmed Salem, and Ronen Eldan\. 2025\.[Great, now write an article about that: The crescendo multi\-turn llm jailbreak attack\.](https://www.usenix.org/conference/usenixsecurity25/presentation/russinovich)In*USENIX Security Symposium*, pages 2421–2440\.
- Shao et al\. \(2024\)Zedian Shao, Hongbin Liu, Jaden Mu, and Neil Zhenqiang Gong\. 2024\.[Making llms vulnerable to prompt injection via poisoning alignment\.](https://doi.org/10.48550/arXiv.2410.14827)*CoRR*\.
- Sharma et al\. \(2024\)Reshabh K\. Sharma, Vinayak Gupta, and Dan Grossman\. 2024\.[SPML: A dsl for defending language models against prompt attacks\.](https://doi.org/10.48550/arXiv.2402.11755)*CoRR*\.
- Shi et al\. \(2026\)Jiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou, Neil Zhenqiang Gong, and Lichao Sun\. 2026\.[Prompt injection attack to tool selection in llm agents\.](https://www.ndss-symposium.org/ndss-paper/prompt-injection-attack-to-tool-selection-in-llm-agents/)In*NDSS*\.
- Sitawarin et al\. \(2024\)Chawin Sitawarin, Norman Mu, David A\. Wagner, and Alexandre Araujo\. 2024\.[PAL: Proxy\-guided black\-box attack on large language models\.](https://doi.org/10.48550/arXiv.2402.09674)*CoRR*\.
- Siu et al\. \(2026\)Vincent Siu, Jingxuan He, Kyle Montgomery, Zhun Wang, Neil Gong, Chenguang Wang, and Dawn Song\. 2026\.[A framework for formalizing llm agent security\.](https://doi.org/10.48550/arXiv.2603.19469)*CoRR*\.
- Steinberg and Gal \(2026\)Jonathan Steinberg and Oren Gal\. 2026\.[MOSAIC\-Bench: Measuring compositional vulnerability induction in coding agents](https://arxiv.org/abs/2605.03952)\.*Preprint*, arXiv:2605\.03952\.
- Sun et al\. \(2025\)Guangzhi Sun, Xiao Zhan, Shutong Feng, Philip C\. Woodland, and Jose Such\. 2025\.[CASE\-Bench: Context\-aware safety benchmark for large language models\.](https://proceedings.mlr.press/v267/sun25ab.html)In*ICML*\.
- Suo \(2024\)Xuchen Suo\. 2024\.[Signed\-Prompt: A new approach to prevent prompt injection attacks against llm\-integrated applications\.](https://doi.org/10.48550/arXiv.2401.07612)*CoRR*\.
- Tamirisa et al\. \(2025\)Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, and Mantas Mazeika\. 2025\.[Tamper\-resistant safeguards for open\-weight llms\.](https://openreview.net/forum?id=4FIjRodbW6)In*ICLR*\.
- Tomilin et al\. \(2025\)Tristan Tomilin, Meng Fang, and Mykola Pechenizkiy\. 2025\.[HASARD: A benchmark for vision\-based safe reinforcement learning in embodied agents\.](https://openreview.net/forum?id=5BRFddsAai)In*ICLR*\.
- Toyer et al\. \(2024\)Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell\. 2024\.[Tensor Trust: Interpretable prompt injection attacks from an online game\.](https://openreview.net/forum?id=fsW7wJGLBd)In*ICLR*\.
- tse Huang et al\. \(2024\)Jen tse Huang, Jiaxu Zhou, Tailin Jin, Xuhui Zhou, Zixi Chen, Wenxuan Wang, Youliang Yuan, Maarten Sap, and Michael R\. Lyu\. 2024\.[On the resilience of multi\-agent systems with malicious agents\.](https://doi.org/10.48550/arXiv.2408.00989)*CoRR*\.
- Tur et al\. \(2025\)Ada Defne Tur, Nicholas Meade, Xing Han Lù, Alejandra Zambrano, Arkil Patel, Esin Durmus, Spandana Gella, Karolina Stanczak, and Siva Reddy\. 2025\.[SafeArena: Evaluating the safety of autonomous web agents\.](https://proceedings.mlr.press/v267/tur25a.html)In*ICML*\.
- Upadhayay et al\. \(2024\)Bibek Upadhayay, Vahid Behzadan, and Amin Karbasi\. 2024\.[Cognitive Overload Attack:prompt injection for long context\.](https://doi.org/10.48550/arXiv.2410.11272)*CoRR*\.
- Wang et al\. \(2026a\)Che Wang, Jiaming Zhang, Ziqi Zhang, Zijie Wang, Yinghui Wang, Jianbo Gao, Tao Wei, Zhong Chen, and Wei Yang Bryan Lim\. 2026a\.[AdapTools: Adaptive tool\-based indirect prompt injection attacks on agentic llms\.](https://doi.org/10.48550/arXiv.2602.20720)*CoRR*\.
- Wang et al\. \(2024a\)Jiongxiao Wang, Fangzhou Wu, Wendi Li, Jinsheng Pan, G\. Edward Suh, Z\. Morley Mao, Muhao Chen, and Chaowei Xiao\. 2024a\.[FATH: Authentication\-based test\-time defense against indirect prompt injection attacks\.](https://doi.org/10.48550/arXiv.2410.21492)*CoRR*\.
- Wang et al\. \(2025a\)Ning Wang, Zihan Yan, Weiyang Li, Chuan Ma, He Henry Chen, and Tao Xiang\. 2025a\.[Advancing Embodied Agent Security: From safety benchmarks to input moderation\.](https://doi.org/10.24963/ijcai.2025/867)In*IJCAI*, pages 7795–7803\.
- Wang et al\. \(2025b\)Shilong Wang, Guibin Zhang, Miao Yu, Guancheng Wan, Fanci Meng, Chongye Guo, Kun Wang, and Yang Wang\. 2025b\.[G\-Safeguard: A topology\-guided security lens and treatment on llm\-based multi\-agent systems\.](https://aclanthology.org/2025.acl-long.359/)In*ACL \(1\)*, pages 7261–7276\.
- Wang et al\. \(2024b\)Taowen Wang, Dongfang Liu, James Chenhao Liang, Wenhao Yang, Qifan Wang, Cheng Han, Jiebo Luo, and Ruixiang Tang\. 2024b\.[Exploring the adversarial vulnerabilities of vision\-language\-action models in robotics\.](https://doi.org/10.48550/arXiv.2411.13587)*CoRR*\.
- Wang et al\. \(2025c\)Xilong Wang, John Bloch, Zedian Shao, Yuepeng Hu, Shuyan Zhou, and Neil Zhenqiang Gong\. 2025c\.[WebInject: Prompt injection attack to web agents\.](https://doi.org/10.18653/v1/2025.emnlp-main.104)In*EMNLP*, pages 2010–2030\.
- Wang et al\. \(2025d\)Xunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li, Pingchuan Ma, Shuai Wang, Yingjiu Li, Yang Liu, Ning Liu, and Juergen Rahmel\. 2025d\.[SelfDefend: Llms can defend themselves against jailbreaking in a practical manner\.](https://www.usenix.org/conference/usenixsecurity25/presentation/wang-xunguang)In*USENIX Security Symposium*, pages 2441–2460\.
- Wang et al\. \(2024c\)Yifei Wang, Dizhan Xue, Shengjie Zhang, and Shengsheng Qian\. 2024c\.[BadAgent: Inserting and activating backdoor attacks in llm agents\.](https://doi.org/10.18653/v1/2024.acl-long.530)In*ACL \(1\)*, pages 9811–9827\.
- Wang et al\. \(2024d\)Zezhong Wang, Fangkai Yang, Lu Wang, Pu Zhao, Hongru Wang, Liang Chen, Qingwei Lin, and Kam\-Fai Wong\. 2024d\.[SELF\-GUARD: Empower the llm to safeguard itself\.](https://doi.org/10.18653/v1/2024.naacl-long.92)In*NAACL\-HLT*, pages 1648–1668\.
- Wang et al\. \(2024e\)Zhepeng Wang, Runxue Bao, Yawen Wu, Jackson Taylor, Cao Xiao, Feng Zheng, Weiwen Jiang, Shangqian Gao, and Yanfu Zhang\. 2024e\.[Unlocking memorization in large language models with dynamic soft prompting\.](https://doi.org/10.18653/v1/2024.emnlp-main.546)In*EMNLP*, pages 9782–9796\.
- Wang et al\. \(2026b\)Zihan Wang, Rui Zhang, Yu Liu, Wenshu Fan, Wenbo Jiang, Qingchuan Zhao, Hongwei Li, and Guowen Xu\. 2026b\.[MPMA: Preference manipulation attack against model context protocol\.](https://doi.org/10.1609/aaai.v40i42.40898)In*AAAI*, pages 35838–35846\.
- Wei et al\. \(2023\)Zeming Wei, Yifei Wang, and Yisen Wang\. 2023\.[Jailbreak and guard aligned language models with only few in\-context demonstrations\.](https://doi.org/10.48550/arXiv.2310.06387)*CoRR*\.
- Wen et al\. \(2025\)Tongyu Wen, Chenglong Wang, Xiyuan Yang, Haoyu Tang, Yueqi Xie, Lingjuan Lyu, Zhicheng Dou, and Fangzhao Wu\. 2025\.[Defending against indirect prompt injection by instruction detection\.](https://aclanthology.org/2025.findings-emnlp.1060/)In*EMNLP \(Findings\)*, pages 19472–19487\.
- Wu et al\. \(2024a\)Chen Henry Wu, Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried, and Aditi Raghunathan\. 2024a\.[Adversarial attacks on multimodal agents\.](https://doi.org/10.48550/arXiv.2406.12814)*CoRR*\.
- Wu et al\. \(2024b\)Fangzhou Wu, Shutong Wu, Yulong Cao, and Chaowei Xiao\. 2024b\.[WIPI: A new web threat for llm\-driven web agents\.](https://doi.org/10.48550/arXiv.2402.16965)*CoRR*\.
- Wu et al\. \(2023\)Yuanwei Wu, Xiang Li, Yixin Liu, Pan Zhou, and Lichao Sun\. 2023\.[Jailbreaking gpt\-4v via self\-adversarial attacks with system prompts\.](https://doi.org/10.48550/arXiv.2311.09127)*CoRR*\.
- Xiang et al\. \(2024\)Zhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong, Qinbin Li, Han Xie, Jiawei Zhang, Zidi Xiong, Chulin Xie, Carl Yang, Dawn Song, and Bo Li\. 2024\.[GuardAgent: Safeguard llm agents by a guard agent via knowledge\-enabled reasoning\.](https://doi.org/10.48550/arXiv.2406.09187)*CoRR*\.
- Xie et al\. \(2024\)Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, and Prateek Mittal\. 2024\.[SORRY\-Bench: Systematically evaluating large language model safety refusal behaviors\.](https://doi.org/10.48550/arXiv.2406.14598)*CoRR*\.
- Xiong et al\. \(2024\)Chen Xiong, Xiangyu Qi, Pin\-Yu Chen, and Tsung\-Yi Ho\. 2024\.[Defensive Prompt Patch: A robust and interpretable defense of llms against jailbreak attacks\.](https://doi.org/10.48550/arXiv.2405.20099)*CoRR*\.
- Xu et al\. \(2025a\)Chejian Xu, Mintong Kang, Jiawei Zhang, Zeyi Liao, Lingbo Mo, Mengqi Yuan, Huan Sun, and Bo Li\. 2025a\.[AdvAgent: Controllable blackbox red\-teaming on web agents\.](https://proceedings.mlr.press/v267/xu25m.html)In*ICML*\.
- Xu et al\. \(2025b\)Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang\. 2025b\.[A\-MEM: Agentic memory for llm agents\.](https://doi.org/10.48550/arXiv.2502.12110)*CoRR*\.
- Xue et al\. \(2024\)Jiaqi Xue, Mengxin Zheng, Yebowen Hu, Fei Liu, Xun Chen, and Qian Lou\. 2024\.[BadRAG: Identifying vulnerabilities in retrieval augmented generation of large language models\.](https://doi.org/10.48550/arXiv.2406.00083)*CoRR*\.
- Yang et al\. \(2024\)Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun\. 2024\.[Watch out for your agents\! investigating backdoor threats to llm\-based agents\.](http://papers.nips.cc/paper_files/paper/2024/hash/b6e9d6f4f3428cd5f3f9e9bbae2cab10-Abstract-Conference.html)In*NeurIPS*\.
- Ye et al\. \(2024\)Junjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, and Xuanjing Huang\. 2024\.[ToolSword: Unveiling safety issues of large language models in tool learning across three stages\.](https://doi.org/10.18653/v1/2024.acl-long.119)In*ACL \(1\)*, pages 2181–2211\.
- Yi et al\. \(2025\)Jingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu\. 2025\.[Benchmarking and defending against indirect prompt injection attacks on large language models\.](https://doi.org/10.1145/3690624.3709179)In*KDD \(1\)*, pages 1809–1820\.
- Ying et al\. \(2026\)Zonghao Ying, Haozheng Wang, Jiangfan Liu, Quanchen Zou, Aishan Liu, Jian Yang, Yaodong Yang, and Xianglong Liu\. 2026\.[AgentVisor: Defending llm agents against prompt injection via semantic virtualization](https://arxiv.org/abs/2604.24118)\.*Preprint*, arXiv:2604\.24118\.
- Yu et al\. \(2024\)Miao Yu, Shilong Wang, Guibin Zhang, Junyuan Mao, Chenlong Yin, Qijiong Liu, Qingsong Wen, Kun Wang, and Yang Wang\. 2024\.[NetSafe: Exploring the topological safety of multi\-agent networks\.](https://doi.org/10.48550/arXiv.2410.15686)*CoRR*\.
- Yung et al\. \(2025\)Canaan Yung, Hanxun Huang, Sarah Monazam Erfani, and Christopher Leckie\. 2025\.[CURVALID: Geometrically\-guided adversarial prompt detection\.](https://doi.org/10.48550/arXiv.2503.03502)*CoRR*\.
- Zeng et al\. \(2024\)Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qingyun Wu\. 2024\.[AutoDefense: Multi\-agent llm defense against jailbreak attacks\.](https://doi.org/10.48550/arXiv.2403.04783)*CoRR*\.
- Zhan et al\. \(2024\)Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang\. 2024\.[InjecAgent: Benchmarking indirect prompt injections in tool\-integrated large language model agents\.](https://doi.org/10.18653/v1/2024.findings-acl.624)In*ACL \(Findings\)*, pages 10471–10506\.
- Zhang et al\. \(2025a\)Boyang Zhang, Yicong Tan, Yun Shen, Ahmed Salem, Michael Backes, Savvas Zannettou, and Yang Zhang\. 2025a\.[Breaking Agents: Compromising autonomous llm agents through malfunction amplification\.](https://doi.org/10.18653/v1/2025.emnlp-main.1771)In*EMNLP*, pages 34964–34976\.
- Zhang et al\. \(2025b\)Dongsen Zhang, Zekun Li, Xu Luo, Xuannan Liu, Peipei Li, and Wenjun Xu\. 2025b\.[MCP Security Bench \(MSB\): Benchmarking attacks against model context protocol in llm agents\.](https://doi.org/10.48550/arXiv.2510.15994)*CoRR*\.
- Zhang et al\. \(2024a\)Hangtao Zhang, Chenyu Zhu, Xianlong Wang, Ziqi Zhou, Shengshan Hu, and Leo Yu Zhang\. 2024a\.[BadRobot: Jailbreaking llm\-based embodied ai in the physical world\.](https://doi.org/10.48550/arXiv.2407.20242)*CoRR*\.
- Zhang et al\. \(2024b\)Quan Zhang, Binqi Zeng, Chijin Zhou, Gwihwan Go, Heyuan Shi, and Yu Jiang\. 2024b\.[Human\-imperceptible retrieval poisoning attacks in llm\-powered applications\.](https://doi.org/10.1145/3663529.3663786)In*SIGSOFT FSE Companion*, pages 502–506\.
- Zhang et al\. \(2025c\)Rupeng Zhang, Haowei Wang, Junjie Wang, Mingyang Li, Yuekai Huang, Dandan Wang, and Qing Wang\. 2025c\.[From Allies to Adversaries: Manipulating llm tool\-calling through adversarial injection\.](https://doi.org/10.18653/v1/2025.naacl-long.101)In*NAACL \(Long Papers\)*, pages 2009–2028\.
- Zhang et al\. \(2025d\)Shaokun Zhang, Ming Yin, Jieyu Zhang, Jiale Liu, Zhiguang Han, Jingyang Zhang, Beibin Li, Chi Wang, Huazheng Wang, Yiran Chen, and Qingyun Wu\. 2025d\.[Which agent causes task failures and when? on automated failure attribution of llm multi\-agent systems\.](https://proceedings.mlr.press/v267/zhang25cq.html)In*ICML*\.
- Zhang et al\. \(2026\)Tian Zhang, Yiwei Xu, Juan Wang, Keyan Guo, Xiaoyang Xu, Bowen Xiao, Quanlong Guan, Jinlin Fan, Jiawei Liu, Zhiquan Liu, and Hongxin Hu\. 2026\.AgentSentry: Mitigating indirect prompt injection in LLM agents via temporal causal diagnostics and context purification\.*CoRR*, abs/2602\.22724\.
- Zhang et al\. \(2025e\)Yiming Zhang, Javier Rando, Ivan Evtimov, Jianfeng Chi, Eric Michael Smith, Nicholas Carlini, Florian Tramèr, and Daphne Ippolito\. 2025e\.[Persistent pre\-training poisoning of llms\.](https://openreview.net/forum?id=eiqrnVaeIw)In*ICLR*\.
- Zhang et al\. \(2024c\)Ziyang Zhang, Qizhen Zhang, and Jakob Nicolaus Foerster\. 2024c\.[Parden, can you repeat that? defending against jailbreaks via repetition\.](https://proceedings.mlr.press/v235/zhang24ca.html)In*ICML*, pages 60271–60287\.
- Zhao et al\. \(2024a\)Shuai Zhao, Meihuizi Jia, Zhongliang Guo, Leilei Gan, Jie Fu, Yichao Feng, Fengjun Pan, and Luu Anh Tuan\. 2024a\.[A Survey of Backdoor Attacks and Defenses on Large Language Models: Implications for security measures\.](https://doi.org/10.48550/arXiv.2406.06852)*CoRR*\.
- Zhao et al\. \(2024b\)Wanru Zhao, Vidit Khazanchi, Haodi Xing, Xuanli He, Qiongkai Xu, and Nicholas Donald Lane\. 2024b\.[Attacks on third\-party apis of large language models\.](https://doi.org/10.48550/arXiv.2404.16891)*CoRR*\.
- Zhao et al\. \(2024c\)Wei Zhao, Zhe Li, Yige Li, Ye Zhang, and Jun Sun\. 2024c\.[Defending large language models against jailbreak attacks via layer\-specific editing\.](https://doi.org/10.18653/v1/2024.findings-emnlp.293)In*EMNLP \(Findings\)*, pages 5094–5109\.
- Zhou et al\. \(2025a\)Huichi Zhou, Kinhei Lee, Zhonghao Zhan, Yue Chen, Zhenhao Li, Zhaoyang Wang, Hamed Haddadi, and Emine Yilmaz\. 2025a\.[TrustRAG: Enhancing robustness and trustworthiness in rag\.](https://doi.org/10.48550/arXiv.2501.00879)*CoRR*\.
- Zhou et al\. \(2025b\)Jiaming Zhou, Ke Ye, Jiayi Liu, Teli Ma, Zifan Wang, Ronghe Qiu, Kun\-Yu Lin, Zhilin Zhao, and Junwei Liang\. 2025b\.[Exploring the limits of vision\-language\-action manipulations in cross\-task generalization\.](https://doi.org/10.48550/arXiv.2505.15660)*CoRR*\.
- Zhou et al\. \(2024a\)Xuhui Zhou, Hyunwoo Kim, Faeze Brahman, Liwei Jiang, Hao Zhu, Ximing Lu, Frank Xu, Bill Yuchen Lin, Yejin Choi, Niloofar Mireshghallah, Ronan Le Bras, and Maarten Sap\. 2024a\.[HAICOSYSTEM: An ecosystem for sandboxing safety risks in human\-ai interactions\.](https://doi.org/10.48550/arXiv.2409.16427)*CoRR*\.
- Zhou et al\. \(2025c\)Zhenhong Zhou, Zherui Li, Jie Zhang, Yuanhe Zhang, Kun Wang, Yang Liu, and Qing Guo\. 2025c\.[CORBA: Contagious recursive blocking attacks on multi\-agent systems based on large language models\.](https://doi.org/10.48550/arXiv.2502.14529)*CoRR*\.
- Zhou et al\. \(2024b\)Zhenhong Zhou, Jiuyang Xiang, Haopeng Chen, Quan Liu, Zherui Li, and Sen Su\. 2024b\.[Speak Out of Turn: Safety vulnerability of large language models in multi\-turn dialogue\.](https://doi.org/10.48550/arXiv.2402.17262)*CoRR*\.
- Zhu et al\. \(2023\)Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun\. 2023\.AutoDAN: Automatic and interpretable adversarial attacks on large language models\.*CoRR*, abs/2310\.15140\.
- Zhu et al\. \(2026\)Wenhui Zhu, Xuanzhao Dong, Xiwen Chen, Rui Cai, Peijie Qiu, Zhipeng Wang, Oana Frunza, Shao Tang, Jindong Gu, and Yalin Wang\. 2026\.[Your Agent is More Brittle Than You Think: Uncovering indirect injection vulnerabilities in agentic llms\.](https://doi.org/10.48550/arXiv.2604.03870)*CoRR*\.
- Zou et al\. \(2025\)Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia\. 2025\.[PoisonedRAG: Knowledge corruption attacks to retrieval\-augmented generation of large language models\.](https://www.usenix.org/conference/usenixsecurity25/presentation/zou-poisonedrag)In*USENIX Security Symposium*, pages 3827–3844\.
- Zverev et al\. \(2025\)Egor Zverev, Sahar Abdelnabi, Soroush Tabesh, Mario Fritz, and Christoph H\. Lampert\. 2025\.[Can llms separate instructions from data? and what do we even mean by that?](https://openreview.net/forum?id=8EtSBX41mt)In*ICLR*\.

## Comprehensive Summary of Representative Papers

Table 1:Comprehensive summary of representative papers in our survey\. Papers are grouped by their primary boundary, even when some works naturally span multiple boundaries\.\\rowcolorblue\!12User\-Agent BoundaryPerez and Ribeiro \([2022](https://arxiv.org/html/2607.12406#bib.bib93)\)2022AttackPrompt injectionEarly evidence that language instructions can override hidden prompts and expose role confusion\.Liu et al\. \([2023b](https://arxiv.org/html/2607.12406#bib.bib77)\)2023AttackPrompt injectionShows that system prompts can be leaked or overridden through ordinary user interaction\.Toyer et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib113)\)2023AttackPrompt injectionStudies prompt\-injection style attacks against tool\-using language systems\.Pedro et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib90)\)2023AttackRole confusionShows how hidden instructions and visible user content can collapse into the same control channel\.Zhu et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib164)\)2023AttackAutomated jailbreaksIntroduces optimization\-based jailbreak generation at scale rather than only handcrafted prompts\.Liu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib74)\)2023AttackAutomated jailbreaksDevelops automated jailbreak generation with stronger transfer and search efficiency\.Huang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib51)\)2024AttackJailbreaksShows catastrophic jailbreak behavior in open\-weight models through generation exploitation\.Chao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib16)\)2023AttackBlack\-box jailbreaksDemonstrates efficient black\-box jailbreaking with limited interaction budget\.Wei et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib128)\)2023AttackIn\-context jailbreaksShows that few\-shot demonstrations can jailbreak or guard aligned language models\.Jia et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib55)\)2024AttackOptimization attacksImproves optimization\-based jailbreak methods against aligned models\.Sitawarin et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib106)\)2024AttackBlack\-box attacksUses proxy\-guided attack strategies to strengthen jailbreak transfer\.Jiang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib56)\)2024AttackAutomated red teamingUses deep adversarial search for automated safety testing\.Pei et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib91)\)2024AttackAutomated red teamingUses self\-generated prompts to test robustness under constrained attack settings\.Mazeika et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib83)\)2024BenchmarkEvaluationProvides a broad benchmark for harmful behavior and jailbreak evaluation\.Chao et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib15)\)2024BenchmarkEvaluationStandardizes jailbreak testing across models and attack styles\.Ran et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib97)\)2024BenchmarkEvaluationProvides systematic evaluation of jailbreak effectiveness and defense performance\.Russinovich et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib102)\)2024AttackMulti\-turn attacksShows that compromise may emerge gradually through repeated dialogue rather than a single prompt\.Ren et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib98)\)2024AttackMulti\-turn attacksStudies derailment over extended interaction rather than one\-shot failure\.Cheng et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib26)\)2024AttackLong\-context attacksShows that long context can be used to weaken or bypass safety behavior\.Chang et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib14)\)2024AttackIndirect cluesUses implicit clues and guessing\-game style interaction for indirect jailbreak\.Upadhayay et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib116)\)2024AttackLong\-context attacksUses cognitive overload in long context to weaken safety behavior\.Gong et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib38)\)2023AttackMultimodal attacksDemonstrates that visual channels can bypass text\-focused safety assumptions\.Liu et al\. \([2023a](https://arxiv.org/html/2607.12406#bib.bib76)\)2023AttackMultimodal attacksShows that query\-relevant images can jailbreak multimodal systems\.Wu et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib132)\)2023AttackMultimodal attacksShows that multimodal inputs can be used to induce unsafe behavior\.He et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib44)\)2024AttackPersistenceShows that in\-context state can be poisoned and later reused\.Wang et al\. \([2024e](https://arxiv.org/html/2607.12406#bib.bib126)\)2024AttackPersistenceStudies latent unlocking of unsafe behavior through hidden contextual influence\.Xu et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib137)\)2025AttackMemory injectionShows that memory can preserve unsafe user influence across interactions\.Chen et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib20)\)2024DefenseAuthority separationUses structured prompting to separate instructions from data more explicitly\.Suo \([2024](https://arxiv.org/html/2607.12406#bib.bib110)\)2024DefenseAuthority separationUses signed prompting or explicit authority structure to preserve source distinction\.Sharma et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib104)\)2024DefenseDSL\-style defenseUses a domain\-specific language to separate trusted control from untrusted prompt content\.Robey et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib101)\)2023DefenseJailbreak defenseProposes smoothing\-style robustness techniques against prompt attacks\.Kim et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib62)\)2023DefenseSafety classifierUses robust safety classification as a shield against adversarial prompting\.Ji et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib53)\)2024DefenseSemantic smoothingDefends against jailbreak attacks through semantic smoothing techniques\.Wang et al\. \([2025d](https://arxiv.org/html/2607.12406#bib.bib123)\)2024DefenseInference\-time defenseUses self\-protection strategies at inference time against unsafe prompting\.Lin et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib72)\)2025DefenseInference\-time defenseUses model\-side protection at inference time against unsafe user control\.Ying et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib142)\)2026DefenseSemantic virtualizationUses semantic virtualization to defend agents against prompt injection through stronger separation of trusted and untrusted context\.\\rowcolorgreen\!12Agent\-Tool BoundaryYi et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib141)\)2023AttackTool outputs as controlShows that indirect prompt injection can arise when tool\-returned content is treated as trusted instruction\.Ye et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib140)\)2024AnalysisTool misuseBreaks tool\-use failure into interpretation, selection, and execution stages\.Fu et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib37)\)2024AnalysisTool feedback misuseStudies how tool\-learning and feedback use can create new safety failures\.Zhao et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib157)\)2024AttackTool selectionShows that external tools and APIs can be manipulated before or during invocation\.Zhang et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib151)\)2024AttackAdversarial tool\-callingDemonstrates that the right tool can still be used in the wrong way\.Shi et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib105)\)2025AttackTool selectionStudies prompt injection and adversarial interference against tool selection logic\.Jing et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib59)\)2025AttackMCP / protocol riskShows that capability descriptions and protocol metadata are themselves safety\-critical\.Wang et al\. \([2026b](https://arxiv.org/html/2607.12406#bib.bib127)\)2025AttackMetadata manipulationStudies manipulation through protocol\-level capability exposure and routing signals\.Zhang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib148)\)2025BenchmarkMCP securityBenchmarks attacks against model context protocol ecosystems\.Mou et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib85)\)2026DefenseSafer invocationFocuses on stronger mediation for high\-risk tool calls\.Chen and Cong \([2025](https://arxiv.org/html/2607.12406#bib.bib18)\)2025DefenseInvocation controlUses guard mechanisms to mediate tool access and high\-impact external capabilities\.Wang et al\. \([2026a](https://arxiv.org/html/2607.12406#bib.bib117)\)2026AttackTrace\-level riskStudies multi\-step unsafe tool orchestration rather than single calls\.Chen et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib21)\)2026DefenseTrace\-level monitoringEmphasizes safety over full tool\-calling trajectories\.Ba et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib8)\)2026BenchmarkCoding agentsEvaluates code\-interpreter style agents under tool\-mediated risk\.Steinberg and Gal \([2026](https://arxiv.org/html/2607.12406#bib.bib108)\)2026BenchmarkCoding agentsMeasures compositional vulnerability in tool\-heavy coding workflows\.\\rowcolororange\!12Agent\-Execution BoundaryZhang et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib147)\)2024AttackUnsafe executionShows that autonomous agents can be compromised into harmful behavior through malfunction amplification\.Fang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib35)\)2024AttackWeb exploitationDemonstrates that LLM agents can autonomously hack websites\.Fang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib34)\)2024AttackVulnerability exploitationExtends execution risk to one\-day vulnerability exploitation\.Guo et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib41)\)2024BenchmarkCode agentsBenchmarks risky code generation and execution behavior\.Levy et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib66)\)2024BenchmarkWeb agentsEvaluates safety and trustworthiness in web\-agent interaction traces\.Tur et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib115)\)2025BenchmarkWeb agentsEvaluates safety in autonomous web\-agent action sequences\.Kumar et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib64)\)2025AnalysisBrowser agentsShows that refusal\-trained models remain vulnerable once coupled to browser action\.Chen et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib17)\)2025AttackGUI agentsHighlights hidden threats in LLM\-powered GUI control\.Zhang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib149)\)2024AttackEmbodied agentsDemonstrates jailbreaking of embodied systems in the physical world\.Robey et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib100)\)2024AttackRobot controlStudies prompt\-based compromise in robot systems controlled by LLMs\.Liu et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib73)\)2024AttackEmbodied backdoorsShows contextual backdoor attacks against embodied agents\.Jiao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib58)\)2025AttackEmbodied backdoorsStudies whether embodied decision\-making agents can be trusted under backdoor attack\.Hu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib49)\)2024BenchmarkVision\-grounded safetyBenchmarks decision making with human\-value constraints in embodied settings\.Li et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib67)\)2024BenchmarkEmbodied evaluationBenchmarks embodied decision making through a dedicated agent interface\.Tomilin et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib112)\)2025BenchmarkSafe RL / embodimentBenchmarks safe embodied decision making under visual risk\.Chakraborty et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib12)\)2025AnalysisEmbodied hallucinationStudies hallucinations in embodied agents driven by LLMs\.Karnik et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib61)\)2025BenchmarkEmbodied red teamingAudits robotic foundation models through embodied red teaming\.Wu et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib130)\)2024AttackMultimodal agentsStudies adversarial attacks on multimodal agents more broadly\.Zhou et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib160)\)2025AttackCross\-task manipulationExplores manipulation limits and generalization in VLA systems\.Wang et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib119)\)2025DefenseEmbodied moderationConnects safety benchmarks to input moderation in embodied agents\.Wang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib121)\)2025AttackVLA vulnerabilitiesExplores adversarial weaknesses in vision\-language\-action models\.Cheng et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib25)\)2024BenchmarkPhysical vulnerabilityEvaluates physical vulnerabilities in end\-to\-end VLA manipulation tasks\.Zhou et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib161)\)2024DefenseSandboxingFrames execution safety as a containment and sandboxing problem\.Lu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib80)\)2025DefensePolicy\-executable defenseProposes executable safeguards against embodied jailbreaks\.Maiti \([2026](https://arxiv.org/html/2607.12406#bib.bib81)\)2026DefenseZero\-trust runtimeUses a zero\-trust architecture to constrain autonomous execution in high\-stakes settings\.Cui et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib28)\)2026BenchmarkHardware repairEvaluates long\-horizon execution safety in real repair tasks\.Fokou \([2026](https://arxiv.org/html/2607.12406#bib.bib36)\)2026DefenseReason\-act separationArgues that systems that reason should not directly act without stronger separation and execution mediation\.\\rowcolorred\!12Agent\-Agent BoundaryLee and Tiwari \([2024](https://arxiv.org/html/2607.12406#bib.bib65)\)2024AttackPrompt infectionShows that one agent can pass malicious control content to another\.Amayuelas et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib4)\)2024AttackDebate attacksStudies adversarial influence through multi\-agent debate\.Ju et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib60)\)2024AttackKnowledge propagationShows that manipulated knowledge can spread socially across agent communities\.He et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib43)\)2025AttackCommunication attacksRed\-teams multi\-agent systems by targeting communication channels directly\.Wang et al\. \([2024c](https://arxiv.org/html/2607.12406#bib.bib124)\)2024AttackBackdoored agentsStudies insertion and activation of backdoors in agent workflows\.Yang et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib139)\)2024AttackBackdoor threatsHighlights latent malicious behavior in LLM\-based agents\.Zhou et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib162)\)2025AttackCascade failureStudies contagious recursive blocking and communication\-level collapse\.Das et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib29)\)2026AttackShared memoryShows that agent memory can be weaponized for later exfiltration\.tse Huang et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib114)\)2024AnalysisMalicious agentsStudies resilience of multi\-agent systems when some agents behave maliciously\.Yu et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib143)\)2024AnalysisTopology safetyShows that network structure shapes safety and spread in multi\-agent communities\.Hammond et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib42)\)2025AnalysisSystemic riskStudies broader multi\-agent risks from advanced AI systems\.Wang et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib120)\)2025DefenseTopology monitoringUses interaction graphs to study and mitigate system\-level spread\.Mao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib82)\)2025DefenseMemory partitioningUses hierarchical data management to reduce unsafe cross\-agent sharing\.Xiang et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib133)\)2024DefenseGuard rolesAssigns a dedicated guard agent to inspect other agents’ behavior\.Zeng et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib145)\)2024DefenseDefense agentsUses a multi\-agent defense setup against jailbreak attacks\.Chen et al\. \([2025d](https://arxiv.org/html/2607.12406#bib.bib23)\)2025DefenseVerifiable policy reasoningAdds explicit reasoning about safety policy within agent coordination\.Zhang et al\. \([2025d](https://arxiv.org/html/2607.12406#bib.bib152)\)2025AnalysisFailure attributionIdentifies which agent and which step caused a task failure\.Hossain et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib47)\)2025DefensePrompt\-injection defenseUses a dedicated multi\-agent defense pipeline against prompt injection\.\\rowcolorpurple\!12System\-Environment BoundaryAbdelnabi et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib1)\)2023AttackIndirect prompt injectionEstablishes that external content can later act as hidden control in LLM systems\.Liu et al\. \([2024c](https://arxiv.org/html/2607.12406#bib.bib75)\)2024AttackAutomatic injectionScales indirect prompt injection through more systematic attack generation\.Bagdasaryan et al\. \([2023](https://arxiv.org/html/2607.12406#bib.bib9)\)2023AttackMultimodal injectionShows that images and sounds can act as indirect instruction channels\.Wu et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib131)\)2024AttackWeb threatsShows new web\-specific threats for LLM\-driven web agents\.Liao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib71)\)2024AttackPrivacy leakageStudies environmental injection that induces privacy leakage in web agents\.Xu et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib136)\)2024AttackWeb\-agent red teamingUses controllable black\-box red teaming against web agents\.Zverev et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib167)\)2024AnalysisInstruction\-data separationExamines whether LLMs can reliably separate instructions from data\.Chang et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib13)\)2026AttackRetrieval\-mediated injectionShows that indirect injection can survive realistic retrieval pipelines\.Zhu et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib165)\)2026AttackLong\-horizon fragilityShows that agentic systems remain vulnerable even when hostile content appears peripheral\.Zhan et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib146)\)2024BenchmarkWeb agentsBenchmarks indirect prompt injection in tool\-integrated agents\.Evtimov et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib33)\)2025BenchmarkWeb agentsBenchmarks prompt\-injection robustness of web agents\.Wang et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib122)\)2025AttackWeb prompt injectionSpecializes prompt injection to web\-agent interaction\.Cao et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib11)\)2025BenchmarkVisual injectionStudies visual prompt injection against computer\-use agents\.Zou et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib166)\)2024AttackRAG poisoningStudies knowledge corruption attacks against retrieval\-augmented generation\.Deng et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib31)\)2024AttackRAG jailbreaksShows that retrieval poisoning can induce jailbreak\-like behavior\.Xue et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib138)\)2024AttackRetrieval corruptionIdentifies vulnerabilities in RAG systems under corrupted retrieval\.Zhang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib150)\)2024AttackRetrieval poisoningShows that imperceptible retrieval poisoning can alter downstream behavior\.Peng et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib92)\)2024AttackBackdoored extractionUses backdoors to extract data from retrieval\-augmented systems\.Qi et al\. \([2025b](https://arxiv.org/html/2607.12406#bib.bib96)\)2024AttackData extractionDemonstrates scalable data extraction from RAG systems\.Jiang et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib57)\)2024AttackAgent\-based exfiltrationUses agent workflows to extract private data from RAG systems\.Li et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib70)\)2024AttackMembership inferenceStudies whether sensitive data can be inferred from RAG outputs\.Anderson et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib5)\)2024AttackMembership inferenceAsks whether specific data exists in the retrieval database\.Chen et al\. \([2024b](https://arxiv.org/html/2607.12406#bib.bib24)\)2024AttackMemory poisoningShows that agent memory or knowledge bases can be poisoned for later exploitation\.Al\-Lawati and Wang \([2026](https://arxiv.org/html/2607.12406#bib.bib2)\)2026AttackMultimodal leakageExtends retrieval leakage analysis to multimodal RAG settings\.Wang et al\. \([2024a](https://arxiv.org/html/2607.12406#bib.bib118)\)2024DefenseAuthenticationUses authentication\-style defenses against indirect prompt injection\.Hines et al\. \([2024](https://arxiv.org/html/2607.12406#bib.bib46)\)2024DefenseSpotlightingMakes suspicious external instructions more visible during inference\.Wen et al\. \([2025](https://arxiv.org/html/2607.12406#bib.bib129)\)2025DefenseInstruction detectionUses instruction detection to filter indirect prompt injection\.Jia et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib54)\)2024DefenseTask alignmentEnforces task alignment against hostile external instructions\.Chen et al\. \([2025c](https://arxiv.org/html/2607.12406#bib.bib22)\)2025DefenseDetection / removalStudies whether indirect prompt injection can be detected and removed after ingestion\.Zhang et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib153)\)2026DefenseTemporal diagnosisTracks when environment content begins to dominate later behavior\.He et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib45)\)2026DefenseCausal attributionUses attribution to separate benign context from attack\-driving context\.Gowda \([2026](https://arxiv.org/html/2607.12406#bib.bib40)\)2026DefenseMemory hygieneDetects memory poisoning in retrieval\-augmented agents\.Siu et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib107)\)2026AnalysisFormalizationProvides a framework for formalizing LLM agent security beyond isolated attack cases\.Li et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib68)\)2026AnalysisSystems securitySummarizes system\-level security considerations for AI agents deployed in real environments\.Kim et al\. \([2026](https://arxiv.org/html/2607.12406#bib.bib63)\)2026SurveyThreat landscapeSynthesizes the broader 2026 attack and defense landscape of agentic AI\.Zhou et al\. \([2025a](https://arxiv.org/html/2607.12406#bib.bib159)\)2025DefenseTrustworthy retrievalEmphasizes trustworthy evidence selection and robustness\-aware design in RAG systems\.

Similar Articles

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

arXiv cs.AI

This paper introduces a five-condition controlled contrast design to disentangle the effects of operational reframing, planner behavior, and approval-framed delegation in multi-agent LLM safety evaluations, showing that aggregate pipeline safety measures are not interpretable as stable architectural properties.