@Dinosn:智能体技能安全精选资源列表:攻击、防御、框架与基准,用于保护 AI 智能体……

X AI KOLs Timeline 工具

摘要

一个精选的 GitHub awesome 列表,汇编了智能体技能安全相关资源,涵盖攻击手段(工具投毒、间接提示注入、后门)、防御措施(沙箱隔离、权限控制、形式化验证),以及 OWASP 智能体技能 Top 10、MITRE ATLAS、NIST AI RMF 等框架和评估基准。

智能体技能安全精选资源列表:攻击、防御、框架与基准,用于保护 AI 智能体的工具使用与技能生态 https://t.co/wsN2YLh9Oc
查看原文
查看缓存全文

缓存时间: 2026/09/30 16:12

LLMSecurity/awesome-agent-skills-security

来源:https://github.com/LLMSecurity/awesome-agent-skills-security

Awesome Agent Skills Security Awesome (https://awesome.re)

🛡️ 关于 AI agent 工具使用与技能(skill)生态安全的精选资源列表——涵盖攻击、防御、框架、基准与标准。

AI agent 正越来越多地借助外部工具、插件与技能(skill)与世界交互。这带来了一个全新的攻击面:agent 技能安全。本列表涵盖保护这些能力安全所需的威胁、防御与研究现状。

目录


威胁框架与标准

  • OWASP Agentic AI Threats and Mitigations (https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/) —— OWASP Agentic Security Initiative(ASI)系列的首篇,提供了基于威胁模型的 agentic 威胁参考。
  • OWASP Top 10 for LLM Applications (https://owasp.org/www-project-top-10-for-large-language-model-applications/) —— 涵盖 LLM01:提示注入、LLM06:过度授权、LLM07:不安全的插件设计、LLM08:过度自主性。
  • OWASP Agentic Skills Top 10(AST10)(https://owasp.org/www-project-agentic-skills-top-10/) —— OWASP 孵化项目,v1.0(2026 版)。首个由标准组织针对 agent 技能(而非笼统的 agent)专门枚举风险的标准,从 AST01 恶意技能到 AST10 跨平台复用,其中 AST01 与 AST02(供应链投毒)被评为严重;聚焦于技能层,而上文 OWASP Agentic AI Threats 则是从整体 agent 层面展开。
  • MITRE ATLASTM (https://atlas.mitre.org/) —— AI 系统对抗性威胁图谱。针对 ML/AI 系统攻击的战术、技术与案例研究。
  • NIST AI Risk Management Framework (https://www.nist.gov/artificial-intelligence/ai-risk-management-framework) —— 用于管理 AI 风险(包括自主 agent 风险)的联邦框架。
  • NIST SP 800-218A: Secure Software Development for AI (https://csrc.nist.gov/pubs/sp/800/218/a/final) —— 针对 AI 赋能系统的安全开发实践。
  • EU AI Act (https://artificialintelligenceact.eu/) —— 欧洲法规,对包括自主 agent 在内的高风险 AI 系统作出专门规定。
  • Anthropic Responsible Scaling Policy (https://www.anthropic.com/index/anthropics-responsible-scaling-policy) —— AI Safety Levels(ASL)框架,针对 agent 能力阈值作出规定。
  • IETF draft-klrc-aiagent-auth-01: AI Agent Authentication and Authorization (https://datatracker.ietf.org/doc/draft-klrc-aiagent-auth/) —— Kasselman 等,与 IETF WIMSE 相关,2026 年。提出了一套基于既有 OAuth 2.0 与 WIMSE 标准的 AI agent 交互认证与授权模型;涵盖委托链、agent 身份与信任建立,但未定义新的协议。
  • IETF draft-niyikiza-oauth-attenuating-agent-tokens-00: Attenuating Authorization Tokens for Agentic Delegation Chains (https://datatracker.ietf.org/doc/draft-niyikiza-oauth-attenuating-agent-tokens/) —— Niyikiza(Tenuo),OAuth WG,2026 年 3 月。定义了衰减式授权令牌(Attenuating Authorization Tokens, AATs):基于 JWT 的凭证,以工具级别的参数约束进行编码,并通过密码学保证单调衰减不变性——任何持有者都只能推导出约束更严格的令牌,而绝不会得到更宽松的令牌。为 Rich Authorization Requests(RFC 9396)扩展了委托链语义。
  • 📄 Certifying Ghosts: How Cybersecurity AI Agents Break the EU Cyber Resilience Act (https://arxiv.org/abs/2607.07109) —— Mayoral-Vilches,2026 年。认为自主漏洞发现与利用 agent 使欧盟《网络弹性法案》关于缓慢、人工节奏的披露以及产品安全状态固定不变的核心假设失效;主张由于攻击者与防御者如今掌握同等的 AI 能力,有意义的合规需要由 agent 持续运行的监控,而非静态认证。

综述与系统化研究

  • 📄 Connecting the Dots in Agentic AI Security: A Cross-Dimensional Threat Taxonomy, Evaluation Maturity, and Open Challenges (https://arxiv.org/abs/2609.23894) —— Baek 等,2026 年。对 66 项研究(2022–2026)的结构化综述,提出跨维度表示 T={S,B,P,A},将受影响的攻击面、信任边界、被违反的安全属性与经实证检验的架构关联起来;分析了 22 项红队研究与 11 个基准以刻画评估成熟度,并提炼出 13 个开放研究问题。
  • 📄 Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems (https://arxiv.org/abs/2609.22712) —— Syed 等,2026 年。从五个维度审视 agentic AI 可信性,将各类失效模式(间接提示注入、后门触发、目标泛化错误、记忆污染、跨会话泄漏)归入统一分类体系,随后提出六阶段可信 agent 开发生命周期(TADL),包含各阶段的信任活动与基于风险的决策门禁。
  • 📄 Attack Success Rate Is Not a Number: On Measurement Validity in Agentic AI Security Evaluation (https://arxiv.org/abs/2609.25173) —— Pathade 等,2026 年。认为攻击成功率(agentic 安全论文的头号指标)实际上是一个由六项常被省略的设计选择所参数化的指标族;对 259 篇 agentic 安全相关 arXiv 论文(2025 年 2 月–2026 年 9 月)的全文元分析发现,大多数论文既未报告方差估计也未进行重复实验,从而削弱了攻击与防御结论在跨论文层面的可比性。
  • 📄 A Survey on LLM-based Autonomous Agents: Common Attacks and Defenses (https://arxiv.org/abs/2402.09283) —— Wu 等,2024 年。针对 LLM agent 在感知、认知与行动阶段所受攻击的全面分类体系。
  • 📄 Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents (https://arxiv.org/abs/2410.02644) —— Zhang 等,2024 年。形式化了 10 种攻击场景、10 个 agent、398 个对抗性环境。
  • 📄 Security of AI Agents (https://arxiv.org/abs/2406.08689) —— He 等,2024 年。关于拥有工具访问权限的 AI agent 威胁模型的知识系统化综述。
  • 📄 Not All Agents Are Created Equal: A Survey on Software-use Agent Security (https://arxiv.org/abs/2502.02761) —— Hua 等,2025 年。专门针对软件使用型 agent 及其独特安全挑战的综述。
  • 📄 A Survey on the Honesty of Large Language Models (https://arxiv.org/abs/2409.18786) —— Xie 等,2024 年。涵盖 agent 欺骗、谄媚以及工具使用场景下的诚实性问题。
  • 📄 A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models (https://arxiv.org/abs/2402.13457) —— Xu 等,2024 年。针对越狱攻击与防御的系统研究,与 agent 防护栏绕过高度相关。
  • 📄 Prompt Injection Attacks and Defenses in LLM-Integrated Applications (https://arxiv.org/abs/2310.12815) —— Liu 等,2024 年。涵盖直接与间接向量的注入攻击全面分类体系。
  • 📄 The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies (https://arxiv.org/abs/2407.19354) —— Gan 等,2024 年。附带安全失效实际案例的综述。
  • 📄 Self-Evolving Agents: A Survey (https://arxiv.org/abs/2504.01641) —— Gao 等,2025 年。自进化 agent 如何催生涌现性安全风险。
  • 📄 Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies (https://arxiv.org/abs/2606.23075) —— Lin 等,2026 年。使用模块-生命周期攻击面矩阵展示自进化如何将攻击持久性从会话级转变为永久性——对抗性影响会被编码进代际传承、自我放大,并在攻击者无需持续接触的情况下在 agent 群体间传播,25 种攻击向量中有 17 种被评为严重。
  • 📄 From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents (https://arxiv.org/abs/2603.07496) —— Zhang 等,2026 年。层级自主性演进(HAE)框架,将 agent 安全分为认知层、执行层与社会层三个层级。
  • 📄 Characterizing Faults in Agentic AI: A Taxonomy of Types, Symptoms, and Root Causes (https://arxiv.org/abs/2603.06847) —— Shah 等,2026 年。针对结合 LLM 推理与工具调用的 agentic AI 系统可靠性失效的实证分类体系。
  • 📄 Security Considerations for Multi-agent Systems (https://arxiv.org/abs/2603.09002) —— 2026 年。MAS 系统威胁面的系统梳理,涵盖 9 大类共 193 个威胁项;评估了 16 个框架,发现没有任何一个能实现过半的覆盖率。
  • 📄 The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey (https://arxiv.org/abs/2603.11088) —— Kim 等,USENIX Security 2026。首篇系统性 AI agent 安全综述,涵盖设计空间、攻击态势与防御机制,并附有保护 agentic 系统的案例研究。
  • 📄 Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats (https://arxiv.org/abs/2603.11619) —— Deng 等,2026 年。五层、面向生命周期的安全框架,分析自主 LLM agent 在初始化、输入、推理、决策与执行各阶段的复合威胁。
  • 📄 OpenClaw as Language Infrastructure: A Case-Centered Survey of a Public Agent Ecosystem in the Wild (https://www.preprints.org/manuscript/202603.1060) —— He 等,Preprints 2026 年。以案例为中心的 OpenClaw 公共 agent 生态综述,其安全维度综合分析了开放技能生态的治理、技能供应链(ClawHub)与部署风险,并与平台、社群及部署维度并列。
  • 📄 AgenticCyOps: Securing Multi-Agentic AI Integration in Enterprise Cyber Operations (https://arxiv.org/abs/2603.09134) —— Mitra 等,2026 年。面向整体架构的安全框架,将企业多 agent 系统的攻击面分解到组件层、协同层与生态层。
  • 📄 MCP-in-SoS: Risk Assessment Framework for Open-Source MCP Servers (https://arxiv.org/abs/2603.10194) —— Kumar 等,2026 年。面向体系之体系(SoS)的风险评估框架,用于评估生产环境中 agent 系统所部署的开源 MCP 服务器的安全风险。
  • 📄 SoK: The Attack Surface of Agentic AI — Tools, and Autonomy (https://arxiv.org/abs/2603.22928) —— Dehghantanha & Homayoun,2026 年。系统化梳理 agentic LLM 系统的信任边界与安全风险;提出涵盖提示注入、RAG 投毒、工具利用与多 agent 威胁的分类体系,并引入 Unsafe Action Rate 与 Privilege Escalation Distance 等度量指标。
  • 📄 Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation (https://arxiv.org/abs/2606.10749) —— Ling 等,2026 年。面向系统的 247 篇论文综述,围绕信息流、委托授权与持久化状态组织 LLM agent 安全议题,指出提示注入与经由工具的劫持是主导威胁,同时预警新兴的状态篡改与多 agent 传播风险。
  • 📄 Data Agents Under Attack: Vulnerabilities in LLM-Driven Analytical Systems (https://arxiv.org/abs/2606.08661) —— Wang 等,2026 年。针对 LLM 数据 agent 的系统安全研究,识别出解释、执行与策略层的八类风险以及由十四种技术构成的攻击分类体系,并在六个开源与生产级分析平台上揭示出显著的安全缺口。
  • 📄 Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents (https://arxiv.org/abs/2606.26627) —— Lahjouji & Colaco,2026 年。围绕数据来源(而非攻击类型)组织 LLM agent 隐私研究——涵盖数据库查询、外部 API 与 agent 间通信——并指出治理缺口,包括对信息流控制的需求,以及覆盖 agent 完整数据生态的基准需求。
  • 📄 Security Engineering of OpenClaw: Analyzing Attack Surface Expansion and Trust-Boundary Violations (https://arxiv.org/abs/2606.15008) —— Jamshidi 等,2026 年。针对模型输出会触发特权操作的自托管多 agent LLM 系统的实证安全分析,量化了在攻击者能力与 agent 数量扩展下系统被攻陷的概率、边界失效与权限漂移。
  • 📄 LLM Agents Security Duality: A Comprehensive Survey of Self-Security and Empowered Cybersecurity (https://arxiv.org/abs/2606.28450) —— Xu 等,2026 年。73 页综述,映射 LLM agent 在安全领域的双重角色——在梳理面向 agent 的漏洞及其缓解措施的同时,审视同一 agent 如何赋能进攻性与防御性网络安全,并凸显自保护与能力之间的张力。
  • 📄 Agent Security Meets Regulatory Reality: A Practitioner Systematization of Autonomous-Agent Threats and Controls in Regulated Financial Systems (https://arxiv.org/abs/2606.29142) —— Mohan & Srinivasa,2026 年。将 agentic 威胁(提示注入、agent 身份、可审计性、工具滥用、数据驻留)映射到金融行业监管要求,并记录了合规导向部署中生产级 agent 间交互模式与控制措施。
  • 📄 The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities (https://arxiv.org/abs/2607.05743) —— Rashidi,2026 年。系统梳理了 39 篇(2023–2026)关于 AI 编码 agent 执行安全的论文——这类 agent 在有限监督下读取代码仓库、调用工具并执行 shell 命令——揭示出隔离、访问控制、TOCTOU 与 MCP 威胁研究中的碎片化现象,以及包括缺少评估隔离架构的共享基准在内的五大关键缺口。
  • 📄 Security and Privacy in Agentic AI: Grand Challenges and Future Directions (https://arxiv.org/abs/2607.06608) —— Jenkins 等,2026 年。由来自学术界、产业界与政府的三十位国际专家完成的地平线扫描报告,识别出日益 agentic 化的 AI 在安全与隐私方面面临的重大挑战与未来研究方向。
  • 📄 Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents (https://arxiv.org/abs/2607.12428) —— Sakib 等,2026 年。对数千个由 agent 编写的

相似文章

mukul975/Anthropic-Cybersecurity-Skills

GitHub Trending (daily)

一个开源仓库,包含754个为AI智能体设计的结构化网络安全技能,覆盖26个安全领域,并映射到多个行业框架,使智能体能够执行专家级别的安全分析。

tech-leads-club/agent-skills

GitHub Trending (daily)

Agent Skills 是一个经过加固的开源库,包含经过验证和测试的技能,用于扩展 AI 编码代理(如 Claude Code 和 Cursor),解决了市场上替代方案中存在的安全漏洞。