Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities
Summary
This position paper discusses the challenges and opportunities of integrating agentic AI into safety-critical multi-drone systems, advocating for human-centered, socio-technical design approaches to ensure trust, governance, and adoption in professional contexts.
View Cached Full Text
Cached at: 08/25/26, 04:17 AM
# Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities Source: [https://arxiv.org/html/2608.21444](https://arxiv.org/html/2608.21444) \[orcid=0000\-0002\-7851\-7339, email=merritt@cs\.aau\.dk, url=https://www\.ixd\.net, \]\*1 \[orcid=0000\-0002\-3312\-7062, email=alejp@mmmi\.sdu\.dk, \] \[orcid=0000\-0002\-4046\-1528, email=juanba@mmmi\.sdu\.dk, \] \[orcid=0000\-0003\-4512\-400X, email=mtoh@cs\.aau\.dk, \] \[orcid=0000\-0002\-9994\-2908, email=andc@mmmi\.sdu\.dk, url=https://anderslyhnechristensen\.com/, \] Alejandro Jarabo\-PeñasJuan Bravo\-ArrabalMaria\-Theresa BahodiAnders Lyhne Christensen ###### Abstract Multi\-drone systems are increasingly positioned for safety\-critical missions such as search and rescue \(SAR\) and critical infrastructure monitoring\. Yet, real\-world adoption remains constrained not only by autonomy performance, but by the difficulty of integrating agentic behavior into professional work: operators must understand, trust, and govern automation under uncertainty, time pressure, and accountability\. This position paper synthesizes the ambitions and lessons from two ongoing efforts: NAMUR, which explores LLM\-supported robot control in SAR and firefighting contexts, and PERSIST, which explores persistent drone operations for monitoring and security at critical infrastructure sites\. We argue that agentic AI should be approached as a socio\-technical design problem, where interfaces, oversight mechanisms, and evaluation practices are as critical as algorithms\. We outline a human\-centered, participatory, and iterative research approach aimed at uncovering stakeholder needs, shaping agent capabilities through successive prototypes, and producing transferable proof\-of\-concept systems and evaluation strategies for other safety\-critical contexts\. ###### keywords Agentic AI ,Multi\-drone systems ,Human\-centered AI ,Participatory design ,Safety\-critical work ,Critical infrastructure ,Search and rescue ††copyrightyear:2026††copyright:Copyright for this paper by its authors\. Use permitted under Creative Commons License Attribution 4\.0 International \(CC BY 4\.0\)\.††venue:Joint Proceedings of the ACM Intelligent User Interfaces \(IUI\) Workshops 2026, March 23\-26, 2026, Paphos, Cyprus††address:Human\-centered Computing \(Aalborg University\), 300 Selma Lagerløfs Vej, Aalborg Øst, 9220, Denmark††address:The Maersk Mc\-Kinney Moller Institute \(University of Southern Denmark\), Campusvej 55, 5230, Odense, Denmark††corresp:Corresponding author\.## 1Introduction Multi\-robot and multi\-drone systems promise new capabilities for safety\-critical work, including emergency response[14](https://arxiv.org/html/2608.21444#bib.bib11)and security and inspection at critical infrastructure[15](https://arxiv.org/html/2608.21444#bib.bib19)\. At the same time, these domains impose strict demands[19](https://arxiv.org/html/2608.21444#bib.bib3);[26](https://arxiv.org/html/2608.21444#bib.bib4);[10](https://arxiv.org/html/2608.21444#bib.bib5): operations are uncertain, accountability is high, and workflows are governed by professional roles, protocols, and risk management\. In such settings, it is rarely sufficient for autonomy to “work” in a technical sense; it must also be understandable, reliable, governable, and adoptable\. Controlling multiple robots as a*single operational system*amplifies these demands[6](https://arxiv.org/html/2608.21444#bib.bib18);[2](https://arxiv.org/html/2608.21444#bib.bib22);[7](https://arxiv.org/html/2608.21444#bib.bib27)\. As fleet size grows, the traditional one\-operator/one\-vehicle model breaks down[20](https://arxiv.org/html/2608.21444#bib.bib17): coordination overhead increases, maintaining shared situation awareness becomes harder, and small uncertainties can cascade into safety issues \(e\.g\., conflicting task priorities, deconfliction problems, or ambiguous responsibility for who approved what\)\. Operators therefore need to interact at multiple levels of granularity[18](https://arxiv.org/html/2608.21444#bib.bib26)\. That might entail setting mission\-level intent and constraints for the group, while still being able to inspect, redirect, or take control of a specific drone or subset when conditions change\. Recent progress in agentic AI and large language models \(LLMs\) has renewed interest in natural\-language tasking, mixed\-initiative planning, and autonomous execution in robotics[27](https://arxiv.org/html/2608.21444#bib.bib12)\. However, these capabilities introduce new interaction challenges[23](https://arxiv.org/html/2608.21444#bib.bib20): how should agent intentions be represented; how do operators supervise, approve, and override decisions; how can correct behavior be guaranteed in the messy realities of critical field operations? This paper is organized as follows\. Section[2](https://arxiv.org/html/2608.21444#S2)introduces the two motivating project contexts—NAMUR \(SAR and firefighting\) and PERSIST \(persistent critical\-infrastructure monitoring\)—and distills shared interaction requirements for multi\-drone control\. Section[3](https://arxiv.org/html/2608.21444#S3)synthesizes lessons from our prior field engagements to frame the key opportunities and safety\-critical tensions that motivate*governable*agentic autonomy, and outlines our participatory, iterative approach for shaping such systems with stakeholders\. Building on these requirements, Section[4](https://arxiv.org/html/2608.21444#S4)presents an LLM\-based multi\-agent system \(LLM\-MAS\) that decomposes natural\-language intent into reviewable sub\-tasks, constrains execution through deterministic tools, and enforces operator preview and approval\. Finally, Section[5](https://arxiv.org/html/2608.21444#S5)concludes with implications for future agentic multi\-drone systems in safety\-critical work\. ## 2Case Contexts and Project Objectives Figure 1:NAMUR concept illustration: an LLM\-driven natural language interaction system with “closed\-loop reasoning” to support human tasking and robot control in emergency response\.Our position is shaped by two complementary projects that rely on agentic AI in different safety\-critical conditions: \(i\) time\-critical emergency response \(NAMUR\) and \(ii\) long\-horizon, persistent operations at critical infrastructure \(PERSIST\)\. We summarize each context in terms of operational setting, stakeholders, objectives, and the agentic\-AI questions that emerge\. ### 2\.1NAMUR: Agentic support for SAR and firefighting NAMUR investigates how LLM\-supported, agentic interaction can help teams coordinate a multi\-drone system during emergency response work such as search and rescue and firefighting support \(see Figure[1](https://arxiv.org/html/2608.21444#S2.F1)\)\. The operational setting is characterized by time pressure, dynamic hazards, incomplete and rapidly changing situational information, and a strong need for coordination across roles[14](https://arxiv.org/html/2608.21444#bib.bib11);[9](https://arxiv.org/html/2608.21444#bib.bib15)\. Deploying multi\-robot systems in the field also depends on robust communication and edge\-cloud infrastructure[4](https://arxiv.org/html/2608.21444#bib.bib16), and requires interfaces that scale across heterogeneous robot platforms\. Because operators with different roles and expertise tend to interact differently with each robot type, agentic AI can serve as a unifying abstraction layer—providing a more coherent, comprehensible relationship between users and diverse physical systems\. In these contexts, technology is only useful when it integrates with established command structures and communication practices, and when responsibility for high\-stakes decisions remains clear\. #### Stakeholders and work setting\. The primary stakeholders include incident commanders and coordinators, field responders operating in hazardous environments, and technology operators responsible for robotic assets\. The work is organized through role\-specialized decision\-making and structured communication, where updates and tasking must be concise, timely, and auditable[17](https://arxiv.org/html/2608.21444#bib.bib6);[28](https://arxiv.org/html/2608.21444#bib.bib7)\. This setting foregrounds interaction demands that go beyond natural language control: operators must be able to supervise, confirm, and constrain what an agentic system does under uncertainty\. #### Project objectives\. NAMUR’s objectives are to: \(1\) enable higher\-level tasking of robotic assets \(e\.g\., translating intent into feasible robot actions\); \(2\) support mixed\-initiative interaction where the system proposes actions, but humans authorize execution; \(3\) make agentic behavior interpretable enough for operators to assess risk, timing, and consequences; and \(4\) study these interactions in realistic, mission\-oriented exercises to surface failure modes, governance needs, and usability constraints that would be invisible in purely lab\-based evaluations\. #### Why agentic AI is compelling here\. Agentic AI can help transform fragmented inputs \(radio updates, map cues, observations\) into actionable suggestions and structured plans[12](https://arxiv.org/html/2608.21444#bib.bib8);[22](https://arxiv.org/html/2608.21444#bib.bib10), and can reduce coordination overhead by maintaining continuity across rapid task switches\. At the same time, emergency response makes the limits of agentic AI especially visible: uncertainty is unavoidable, and behavior designed for convenience can become harmful without explicit authorization workflows[21](https://arxiv.org/html/2608.21444#bib.bib13);[8](https://arxiv.org/html/2608.21444#bib.bib9), conservative defaults, and clear escalation boundaries\. ### 2\.2PERSIST: Persistent multi\-drone operations for critical infrastructure PERSIST investigates persistent, multi\-drone operations for monitoring, inspection, and security at critical infrastructure sites \(see Figure[2](https://arxiv.org/html/2608.21444#S2.F2)\)\. In contrast to emergency response, the dominant challenge here is not a single high\-tempo mission but sustained operations across days and shifts, where drones must repeatedly execute routine tasks, respond to anomalies, and integrate into existing organizational workflows for security and maintenance\. Recent work on large language model \(LLM\) agents in industrial automation highlights the broader relevance of agentic approaches for orchestrating complex operational systems beyond the lab[29](https://arxiv.org/html/2608.21444#bib.bib29)\. #### Stakeholders and work setting\. Key stakeholders include site security personnel, operations and maintenance staff, and organizational decision\-makers responsible for safety, compliance, and continuity of service\. The setting includes constrained physical environments \(restricted zones, sensitive assets\), operational rhythms \(shift handovers, scheduled inspections\), and an expectation that systems behave predictably and recover gracefully from routine disruptions \(weather, connectivity, access changes, false alarms\)\. #### Project objectives\. PERSIST’s objectives are to: \(1\) move from ad hoc drone flights to persistent, repeatable operations that can be delegated and scheduled; \(2\) reduce the barrier to entry for professional drone use through interface support for planning, monitoring, and handover; \(3\) enable scalable supervision where one operator can manage multiple vehicles and mission types; and \(4\) prototype and validate the approach in realistic infrastructure settings, producing proofs of concept with transfer potential to other critical sites\. #### Why agentic AI is compelling here\. Agentic AI can support long\-horizon orchestration: selecting and parameterizing routine missions, adapting schedules as conditions change, and assisting with anomaly triage and reporting\. The risk profile differs from NAMUR: failures may be less acute in the moment, but persistent operations amplify the importance of reliability, drift management, operator fatigue, and organizational trust\. Here, the central interaction challenge is sustained governability: users must be able to understand what the system has been doing over time, why exceptions occurred, and what remains unresolved at shift handover\. Figure 2:PERSIST overview illustration \(from proposal material\): examples of task types motivating persistent multi\-drone operations \(e\.g\., routine monitoring, inspection, and security workflows\)\. ### 2\.3Cross\-cutting themes and contrasts Together, NAMUR and PERSIST highlight complementary stressors for agentic multi\-drone systems\. Tempo and horizon\.NAMUR emphasizes rapid decision cycles and short operational windows; PERSIST emphasizes persistent activity, repeatability, and organizational integration across shifts\. Governance and accountability\.Both contexts demand clear authority and auditability, but in different forms: emergency response requires immediate confirmation and escalation pathways, while infrastructure operations require policy\-aligned routines, handover summaries, and post hoc accountability\. Implication for our research approach\.Across both cases, we treat agentic AI as a socio\-technical design material: we study how people coordinate, decide, and verify; we prototype agentic behaviors that remain governable; and we iteratively shape the system with stakeholders rather than presuming a final autonomy solution\. Figure 3:Latest prototype interface screenshot: used in participatory walkthroughs to elicit requirements for oversight, uncertainty communication, and operational fit\. ## 3Agentic AI for Multi\-Drone Operations in Safety\-Critical Work: Lessons, Tensions, and an Agenda Our perspective on agentic AI is grounded in prior human\-centered studies of multi\-robot systems in safety\-critical environments, including co\-design and prototype evaluations with Danish and Spanish emergency services in search\-and\-rescue \(SAR\) and firefighting contexts[14](https://arxiv.org/html/2608.21444#bib.bib11);[16](https://arxiv.org/html/2608.21444#bib.bib2), as well as co\-design work with security personnel at a power plant\. Across these contexts, practitioners want autonomy that is*self\-sustaining*enough to be operationally viable, yet*transparent*, easily learnable, and aligned with local protocols and expertise\. This motivates a cautiously optimistic stance: agentic AI can reduce the coordination burden and accelerate sensemaking, but only if paired with interaction mechanisms that keep humans able to authorize, supervise, and diagnose actions under uncertainty\. In the ideal scenario, a single operator can supervise a mission involving a fleet of drones through mixed\-initiative control: high autonomy for routine progress and rapid intervention at multiple levels of granularity when risk or uncertainty increases\. ### 3\.1Opportunities that motivate agent support Agentic AI is compelling in these domains not because it can replace professional judgment, but because it can help teams manage*scale*\(multiple vehicles, multiple data streams, multiple stakeholders\) without collapsing situation awareness[11](https://arxiv.org/html/2608.21444#bib.bib24)\. Across our cases, the most salient opportunities are: \(i\) selective attention and summarization that reduces continuous manual scanning, \(ii\) mixed\-initiative planning and replanning that maintain coherent coverage as information changes, and \(iii\) long\-horizon orchestration for persistent operations, including scheduling, anomaly triage, and reporting aligned to staffing constraints\. ### 3\.2Safety\-critical tensions that require governability The same capabilities introduce tensions that cannot be addressed by autonomy alone\. First, practitioners need systems that can operate without constant attention, yet resist black\-box behavior that cannot be explained, constrained, or overridden\. Second, strong cueing and alerting can improve short\-term performance but encourage attention tunneling, reducing holistic overview during supervision[3](https://arxiv.org/html/2608.21444#bib.bib25)\. Third, trust must be calibrated under imperfect perception and dynamic conditions[1](https://arxiv.org/html/2608.21444#bib.bib1); conservative defaults, uncertainty cues, and verification workflows are essential\. Finally, adoption depends on operational fit: alignment with roles, protocols, and accountability practices, including auditable authorization and decision trails\. ### 3\.3Research approach: participatory shaping of agentic autonomy Rather than aiming for a single, final autonomy concept, we treat agentic AI as a socio\-technical design material that must be iteratively shaped through stakeholder engagement with functional prototypes[5](https://arxiv.org/html/2608.21444#bib.bib23)\. We begin with needs discovery grounded in concrete scenarios and demonstrations to map work practices, decision points, and cognitive bottlenecks, then use successive prototypes \(see[Figure 3](https://arxiv.org/html/2608.21444#S2.F3)for an example of our functional multi\-drone control interface prototype deployed in a real\-world context[16](https://arxiv.org/html/2608.21444#bib.bib2)& companion video of live control of UGVs with LLM\-enabled voice commands\.111Video showing live control of UGVs with LLM\-enabled voice commands:[https://www\.youtube\.com/watch?v=Hna7AtDS5wo](https://www.youtube.com/watch?v=Hna7AtDS5wo)\) as boundary objects to negotiate what the agent should do, what must require explicit authorization, and what evidence and uncertainty cues the system must surface\. Across iterations, this process yields actionable interaction requirements for*governability*: configurable constraints that domain experts can tune quickly, clear intervention points for high\-stakes actions, and observability features that let both operators and developers diagnose behavior and learn from incidents\. #### Bridge to architecture\. Together, these cases and lessons motivate an architecture that decomposes intent into reviewable sub\-tasks, constrains action through deterministic tools, supports scheduling and long\-horizon memory, and enforces explicit operator preview and approval before execution\. ## 4An LLM\-MAS Architecture for Multi\-drone Control The functional prototype implements a unified agentic AI architecture that meets the requirements of the NAMUR and PERSIST projects and has been demonstrated in live control of UGVs and UAVs[16](https://arxiv.org/html/2608.21444#bib.bib2)\. The architecture is implemented as a LLM\-MAS composed of five specialized agents:*Coordinator*,*Event*,*Spatial*,*Swarm*, and*Summarizer*\. Several agents, including the Spatial Agent and the Swarm Agent, interface with external software tools that provide deterministic capabilities such as spatial computation, multi\-drone motion execution, memory storage, and task scheduling\. Collectively, these components enable the decomposition of high\-level natural language commands into well\-defined sub\-tasks[24](https://arxiv.org/html/2608.21444#bib.bib28);[25](https://arxiv.org/html/2608.21444#bib.bib14)that can be previewed, validated, and explicitly approved by an operator prior to execution \(Figure[4](https://arxiv.org/html/2608.21444#S4.F4)\)\. The specific set of tools available to the LLM\-MAS is deployment\-dependent\. For SAR scenarios, dedicated tools support functionalities such as compliant coverage path planning\. In contrast, for persistent drone operations at biomass power plants, specialized tools are provided for generating flight paths and image acquisition plans that enable accurate biomass estimation\. Figure 4:Architecture of the LLM\-MAS use in the functional prototype\.In the following, we describe the responsibilities of each agent using a representative operator command as an example\. - •Coordinator Agent:Serves as the central orchestrator of the system\. It receives the operator’s command, infers the underlying intent, and decomposes it into concrete sub\-queries and execution requests that are delegated to subordinate agents\. - •Events Agent:Manages events and periodic, delayed, and future\-oriented directives issued by the operator\. It maintains and monitors scheduled and recurring tasks \(e\.g\.,“Estimate biomass volume every weekday at 2 pm”or“Return all drones to base at 18:00”\) and ensures that they are executed or surfaced at the appropriate time and context\. - •Spatial Agent:Provides spatial reasoning and situational awareness by accessing a GeoJSON\-formatted map alongside the live positions of all active drones and tracked first responders\. It identifies relevant semantic map features \(e\.g\., the“eastern tree line”or the“biomass pile next to the water”\) and accesses tools to compute spatial relationships between entities as required\. - •Swarm Agent:Handles motion planning and task execution requests for one or more available drones\. It selects suitable drone platforms based on availability and operational context and issues coordinated control commands when multiple drones are involved[13](https://arxiv.org/html/2608.21444#bib.bib21)\. Given the real\-world consequences of these actions, all requests generated by this agent are subject to operator preview and confirmation\. Upon receiving target coordinates from the Spatial Agent, the Swarm Agent invokes an execution tool that publishes the desired actions together with the selected drone identifier\(s\) to a ROS topic, thereby triggering the downstream planning and flight control pipeline\. - •Summarizer Agent:Generates a concise, human\-readable summary of the actions executed and the resulting system state after each command\. In addition to being presented to the operator as confirmation, these summaries are persistently stored in the system’s interaction memory\. The stored summaries provide a compact, structured record of prior interactions and outcomes, which is made available as contextual input to the Coordinator Agent in subsequent dialogue turns\. This persistent context enables coherent multi\-turn interactions and reference resolution across commands, such as interpreting follow\-up instructions that rely on previously mentioned entities or actions \(e\.g\.,“Move drone 5 to the main building”followed by“Now, move it to the nearest safe landing zone”\)\. By explicitly separating responsibilities across specialized agents and constraining their behavior through structured prompts and tool interfaces, the proposed LLM\-MAS enables transparent, interpretable, and safe translation of natural language instructions into coordinated multi\-drone mission execution\. Furthermore, the integration of persistent interaction history and explicit task management supports robust multi\-turn dialogue and long\-horizon autonomy under continuous operator supervision\. ## 5Conclusion Agentic AI holds promise for scaling multi\-drone operations in SAR and critical infrastructure monitoring, but introduces challenges around oversight, trust calibration, operational fit, and evaluation\. Grounded in NAMUR and PERSIST, we framed these requirements and presented an LLM\-MAS architecture that operationalizes governable agentic autonomy through constrained tool use and explicit operator preview/approval\. This framing supports future research in autonomy and provides interactive infrastructure required to deploy agentic systems responsibly\. Looking ahead, we see three implications for future agentic multi\-drone systems in safety\-critical work\. First,governability should be treated as a core system property, not a UI add\-on: architectures should make authorization points, constraints, and intervention mechanisms explicit and auditable\. Second,agentic behavior should be built from constrained, inspectable tool userather than unconstrained free\-form action, enabling predictable failure modes, reproducible debugging, and clearer responsibility boundaries\. Third,evaluation should reflect real operational use\. This means measuring more than task success: we should assess situation awareness \(and whether alerts cause attention tunneling\), trust calibration under uncertainty, coordination overhead, and whether the system supports accountable after\-action review\. ###### Acknowledgements\. This work was supported by the Innovation Fund Denmark for DIREC project U07 \(the PERSIST project\), and the Independent Research Fund Denmark under grant 10\.46540/4264\-00105B \(the NAMUR project\)\. ## Declaration on Generative AI During the preparation of this work, the authors used Grammarly and GPT 5\.2 in order to: Grammar and spelling check\. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the publication’s content\. ## References - Ahlskoget al\.\(2024\)J\. Ahlskog, M\. Bahodi, A\. Lugmayr, and T\. MerrittFostering trust through user interface design in multi\-drone search and rescue\.InProceedings of the Second International Symposium on Trustworthy Autonomous Systems,TAS ’24,New York, NY, USA\.External Links:ISBN 9798400709890,[Link](https://doi.org/10.1145/3686038.3686052),[Document](https://dx.doi.org/10.1145/3686038.3686052)Cited by:[§3\.2](https://arxiv.org/html/2608.21444#S3.SS2.p1.1)\. - Bahodiet al\.\(2025\)M\. Bahodi, N\. Lau, N\. van Berkel, K\. A\. R\. Grøntved, M\. B\. Skov, and T\. MerrittBefore it falls: supporting drone fleet management through battery visualizations\.InIFIP Conference on Human\-Computer Interaction,pp\. 347–370\.Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p2.1)\. - Bahodiet al\.\(2024\)M\. Bahodi, N\. van Berkel, M\. Skov, and T\. MerrittShow Me What’s Wrong: Impact of Explicit Alerts on Novice Supervisors of a Multi\-Robot Monitoring System\.InProceedings of the Second International Symposium on Trustworthy Autonomous Systems,TAS ’24,New York, NY, USA,pp\. 1–17\.External Links:ISBN 979\-8\-4007\-0989\-0,[Link](https://dl.acm.org/doi/10.1145/3686038.3686069),[Document](https://dx.doi.org/10.1145/3686038.3686069)Cited by:[§3\.2](https://arxiv.org/html/2608.21444#S3.SS2.p1.1)\. - Bravo\-Arrabalet al\.\(2021\)J\. Bravo\-Arrabal, M\. Toscano\-Moreno, J\. J\. F\. Lozano, A\. Mandow, J\. A\. Gómez\-Ruiz, and A\. García\-CerezoThe Internet of Cooperative Agents Architecture \(X\-IoCA\) for Robots, Hybrid Sensor Networks, and MEC Centers in Complex Environments: A Search and Rescue Case Study\.Sensors21\(23\),pp\. 7843\.External Links:[Document](https://dx.doi.org/10.3390/S21237843)Cited by:[§2\.1](https://arxiv.org/html/2608.21444#S2.SS1.p1.1)\. - Bødker and Kyng \(2018\)S\. Bødker and M\. KyngParticipatory design that matters—facing the big issues\.ACM Trans\. Comput\.\-Hum\. Interact\.25\(1\)\.External Links:ISSN 1073\-0516,[Link](https://doi.org/10.1145/3152421),[Document](https://dx.doi.org/10.1145/3152421)Cited by:[§3\.3](https://arxiv.org/html/2608.21444#S3.SS3.p1.1)\. - Christensenet al\.\(2022\)A\. L\. Christensen, K\. A\. R\. Grøntved, M\. O\. Hoang, N\. van Berkel, M\. Skov, A\. Scovill, G\. Edwards, K\. R\. Geipel, L\. Dalgaard, U\. P\. S\. Lundquist,et al\.The HERD project: human\-multi\-robot interaction in search & rescue and in farming\.InAdjunct Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems,pp\. 1–4\.Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p2.1)\. - Chunget al\.\(2018\)S\. Chung, A\. A\. Paranjape, P\. Dames, S\. Shen, and V\. KumarA survey on aerial swarm robotics\.IEEE Transactions on Robotics34\(4\),pp\. 837–855\.External Links:[Document](https://dx.doi.org/10.1109/TRO.2018.2857475)Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p2.1)\. - Cleland\-Huanget al\.\(2025\)J\. Cleland\-Huang, P\. A\. A\. Granadeno, A\. M\. R\. Bernal, D\. Hernandez, M\. Murphy, M\. Petterson, and W\. ScheirerCognitive Guardrails for Open\-World Decision Making in Autonomous Drone Swarms\.CoRRabs/2505\.23576\.External Links:2505\.23576,[Document](https://dx.doi.org/10.48550/arXiv.2505.23576)Cited by:[§2\.1](https://arxiv.org/html/2608.21444#S2.SS1.SSS0.Px3.p1.1)\. - Delmericoet al\.\(2019\)J\. A\. Delmerico, S\. Mintchev, A\. Giusti, B\. Gromov, K\. Melo, T\. Horvat, C\. Cadena, M\. Hutter, A\. J\. Ijspeert, D\. Floreano, L\. M\. Gambardella, R\. Siegwart, and D\. ScaramuzzaThe current state and future outlook of rescue robotics\.J\. Field Robotics36\(7\),pp\. 1171–1191\.External Links:[Document](https://dx.doi.org/10.1002/ROB.21887)Cited by:[§2\.1](https://arxiv.org/html/2608.21444#S2.SS1.p1.1)\. - Drew \(2021\)D\. S\. DrewMulti\-agent systems for search and rescue applications\.Current Robotics Reports2,pp\. 189–200\.External Links:[Document](https://dx.doi.org/10.1007/s43154-021-00048-3)Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p1.1)\. - Endsley \(1995\)M\. R\. EndsleyToward a Theory of Situation Awareness in Dynamic Systems\.Human Factors37\(1\),pp\. 32–64\.External Links:ISSN 0018\-7208,[Link](https://doi.org/10.1518/001872095779049543),[Document](https://dx.doi.org/10.1518/001872095779049543)Cited by:[§3\.1](https://arxiv.org/html/2608.21444#S3.SS1.p1.1)\. - Goecks and Waytowich \(2023\)V\. G\. Goecks and N\. R\. WaytowichDisasterResponseGPT: Large Language Models for Accelerated Plan of Action Development in Disaster Response Scenarios\.InWorkshop on Challenges in Deployable Generative AI at International Conference on Machine Learning \(ICML\),External Links:2306\.17271,[Document](https://dx.doi.org/10.48550/arXiv.2306.17271)Cited by:[§2\.1](https://arxiv.org/html/2608.21444#S2.SS1.SSS0.Px3.p1.1)\. - Grøntvedet al\.\(2023\)K\. A\. Grøntved, U\. P\. Lundquist, and A\. L\. ChristensenDecentralized multi\-UAV trajectory task allocation in search and rescue applications\.In2023 21st International Conference on Advanced Robotics \(ICAR\),pp\. 35–41\.Cited by:[4th item](https://arxiv.org/html/2608.21444#S4.I1.i4.p1.1)\. - Hoanget al\.\(2023\)M\. O\. Hoang, K\. A\. R\. Grøntved, N\. van Berkel, M\. B\. Skov, A\. L\. Christensen, and T\. MerrittDrone Swarms to Support Search and Rescue Operations: Opportunities and Challenges\.InCultural Robotics: Social Robots and Their Emergent Cultural Ecologies,B\. J\. Dunstan, J\. T\. K\. V\. Koh, D\. Turnbull Tillman, and S\. A\. Brown \(Eds\.\),pp\. 163–176\.External Links:[Document](https://dx.doi.org/10.1007/978-3-031-28138-9%5F11),ISBN 978\-3\-031\-28138\-9Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.21444#S2.SS1.p1.1),[§3](https://arxiv.org/html/2608.21444#S3.p1.1)\. - Jacobsenet al\.\(2023\)R\. H\. Jacobsen, L\. Matlekovic, L\. Shi, N\. Malle, N\. Ayoub, K\. Hageman, S\. Hansen, F\. F\. Nyboe, and E\. EbeidDesign of an autonomous cooperative drone swarm for inspections of safety critical infrastructure\.Applied Sciences13\(3\),pp\. 1256\.Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p1.1)\. - Jarabo\-Peñaset al\.\(2025\)A\. Jarabo\-Peñas, J\. Bravo\-Arrabal, D\. Lin\-Yang, F\. Pastor, R\. Ladig, J\. J\. Fernández\-Lozano, A\. Christensen, and A\. García\-CerezoReal\-World Deployment of an LLM\-Enabled Voice\-Commanded UGV for Logistics in SAR Missions\.In2025 IEEE International Symposium on Safety Security Rescue Robotics \(SSRR\),Cited by:[§3\.3](https://arxiv.org/html/2608.21444#S3.SS3.p2.1),[§3](https://arxiv.org/html/2608.21444#S3.p1.1),[§4](https://arxiv.org/html/2608.21444#S4.p1.1)\. - Jensen and Thompson \(2016\)J\. Jensen and S\. ThompsonThe incident command system: a literature review\.Disasters40\(1\),pp\. 158–182\.Cited by:[§2\.1](https://arxiv.org/html/2608.21444#S2.SS1.SSS0.Px1.p1.1)\. - Kimet al\.\(2020\)L\. H\. Kim, D\. S\. Drew, V\. Domova, and S\. FollmerUser\-defined swarm robot control\.InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems,CHI ’20,New York, NY, USA,pp\. 1–13\.External Links:ISBN 9781450367080,[Document](https://dx.doi.org/10.1145/3313831.3376814)Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p2.1)\. - Murphyet al\.\(2008\)R\. R\. Murphy, S\. Tadokoro, D\. Nardi, A\. Jacoff, P\. Fiorini, H\. Choset, and A\. M\. ErkmenSearch and Rescue Robotics\.InSpringer Handbook of Robotics,B\. Siciliano and O\. Khatib \(Eds\.\),pp\. 1151–1173\.External Links:[Document](https://dx.doi.org/10.1007/978-3-540-30301-5%5F51)Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p1.1)\. - Plankeet al\.\(2020\)L\. J\. Planke, Y\. Lim, A\. Gardi, R\. Sabatini, T\. Kistan, and N\. EzerA cyber\-physical\-human system for one\-to\-many UAS operations: cognitive load analysis\.Sensors20\(19\),pp\. 5467\.Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p2.1)\. - Robeyet al\.\(2024\)A\. Robey, Z\. Ravichandran, V\. Kumar, H\. Hassani, and G\. J\. PappasJailbreaking LLM\-controlled Robots\.CoRRabs/2410\.13691\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2410.13691),2410\.13691Cited by:[§2\.1](https://arxiv.org/html/2608.21444#S2.SS1.SSS0.Px3.p1.1)\. - Sadiket al\.\(2025\)A\. R\. Sadik, M\. Ashfaq, N\. Mäkitalo, and T\. MikkonenHuman\-LLM Synergy in Context\-Aware Adaptive Architecture for Scalable Drone Swarm Operation\.CoRRabs/2509\.05355\.External Links:2509\.05355,[Document](https://dx.doi.org/10.48550/arXiv.2509.05355)Cited by:[§2\.1](https://arxiv.org/html/2608.21444#S2.SS1.SSS0.Px3.p1.1)\. - Sapkotaet al\.\(2025\)R\. Sapkota, K\. I\. Roumeliotis, and M\. KarkeeUAVs meet agentic AI: a multidomain survey of autonomous aerial intelligence and agentic UAVs\.arXiv preprint arXiv:2506\.08045\.Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p3.1)\. - Singhet al\.\(2023\)I\. Singh, V\. Blukis, A\. Mousavian, A\. Goyal, D\. Xu, J\. Tremblay, D\. Fox, J\. Thomason, and A\. GargProgPrompt: generating situated robot task plans using large language models\.In2023 IEEE International Conference on Robotics and Automation \(ICRA\),pp\. 11523–11530\.External Links:[Document](https://dx.doi.org/10.1109/ICRA48891.2023.10161317)Cited by:[§4](https://arxiv.org/html/2608.21444#S4.p1.1)\. - Tsushimaet al\.\(2025\)Y\. Tsushima, S\. Yamamoto, A\. A\. Ravankar, J\. V\. S\. Luces, and Y\. HirataTask Planning for a Factory Robot Using Large Language Model\.IEEE Robotics Autom\. Lett\.10\(3\),pp\. 2383–2390\.External Links:[Document](https://dx.doi.org/10.1109/LRA.2025.3531153)Cited by:[§4](https://arxiv.org/html/2608.21444#S4.p1.1)\. - Ventura and Lima \(2012\)R\. Ventura and P\. U\. LimaSearch and rescue robots: The civil protection teams of the future\.In2012 Third International Conference on Emerging Security Technologies,pp\. 12–19\.Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p1.1)\. - Wanget al\.\(2025\)J\. Wang, E\. Shi, H\. Hu, C\. Ma, Y\. Liu, X\. Wang, Y\. Yao, X\. Liu, B\. Ge, and S\. ZhangLarge language models for robotics: Opportunities, challenges, and perspectives\.J\. Autom\. and Intell\.4,pp\. 52–64\.Cited by:[§1](https://arxiv.org/html/2608.21444#S1.p3.1)\. - Wolbers and Boersma \(2013\)J\. Wolbers and K\. BoersmaThe common operational picture as collective sensemaking\.Journal of Contingencies and Crisis management21\(4\),pp\. 186–199\.Cited by:[§2\.1](https://arxiv.org/html/2608.21444#S2.SS1.SSS0.Px1.p1.1)\. - Xiaet al\.\(2023\)Y\. Xia, M\. Shenoy, N\. Jazdi, and M\. WeyrichTowards autonomous system: flexible modular production system enhanced with large language model agents\.In2023 IEEE International Conference on Emerging Technologies and Factory Automation \(ETFA\),External Links:[Document](https://dx.doi.org/10.1109/ETFA54631.2023.10275362)Cited by:[§2\.2](https://arxiv.org/html/2608.21444#S2.SS2.p1.1)\.
Similar Articles
Security and Privacy in Agentic AI: Grand Challenges and Future Directions
This paper presents key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise with thirty international experts. It identifies emerging risks from increased AI autonomy and permissions, including prompt injection attacks and malicious applications.
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security
This survey provides a comprehensive examination of trustworthy agentic AI, focusing on safety, robustness, privacy, and system security. It clarifies key concepts, identifies risks along the agent workflow, summarizes mitigation strategies, and consolidates evaluation metrics and benchmarks, aiming to serve as a practical reference for deploying agentic AI in high-stakes environments.
AI safety and alignment
The article discusses concerns about AI safety and alignment as AI becomes more intelligent and integrated into society, referencing Anthropic's call for a pause to address potential catastrophic risks.
Understanding Cognition-Induced Risks in Agentic AI Systems
The paper systematically analyzes risks in agentic AI systems induced by expanding cognitive capabilities across physical, social, and self-referential levels, and proposes mitigation strategies to ensure their safe development.
Engineering Trustworthy Agentic AI for Critical Systems
This survey proposes a trustworthiness model for agentic AI in critical engineering systems, covering safety, robustness, transparency, accountability, and security across domains like power systems and autonomous vehicles.