agent-security

Tag

Cards List
#agent-security

Securing Agents Across Perplexity's Client Endpoints with Numbat (11 minute read)

TLDR AI ↗ · 2026-07-30 Cached

Perplexity open-sources Numbat, an agent security suite designed to detect, prevent, and investigate dangerous AI agent activity on client endpoints, addressing novel security challenges from autonomous agents.

0 favorites 0 likes
#agent-security

"STRICT READ-ONLY" is a rule in my skill file. The model ignored it anyway.

Reddit r/openclaw ↗ · 2026-07-21

An OpenClaw agent with a strict read-only rule was tricked by prompt injection into posting on Twitter, highlighting the gap between system instructions and actual enforcement. The author asks how others handle such security.

0 favorites 0 likes
#agent-security

7 Sandbox Escape Vulnerabilities Across 4 Coding Agent Vendors

Lobsters Hottest ↗ · 2026-07-20 Cached

Pillar Research found sandbox escape vulnerabilities in AI coding agents from Cursor, Codex, Gemini CLI, and Antigravity, revealing that these agents can write files that host components later trust, bypassing sandbox boundaries. The findings highlight the need for a new threat model for agentic security.

0 favorites 0 likes
#agent-security

I pushed a new AgentSecure update after people actually tried it

Reddit r/AI_Agents ↗ · 2026-07-14

Released a new update for AgentSecure based on user feedback.

0 favorites 0 likes
#agent-security

FARMA: a memory-poisoning attack that forges an agent's own decision logs, not its retrieved facts

Reddit r/AI_Agents ↗ · 2026-07-11

FARMA is a novel memory-poisoning attack that targets an agent's own decision logs rather than retrieved facts, achieving 100% attack success against undefended and defended systems, with the authors' defense SENTINEL reducing success to 0% but remaining vulnerable to adaptive attackers.

0 favorites 0 likes
#agent-security

Show HN: Scan your AI agents for dangerous capabilities

Hacker News Top ↗ · 2026-07-06 Cached

MakerChecker is an open-source security layer for AI agents that enforces deny-by-default permissions, human approvals, and provides a cryptographically signed audit trail. It scans agent code for dangerous capabilities and prevents agents from approving their own actions.

0 favorites 0 likes
#agent-security

@yibie: Recommended — Anthropic's engineering team writes up their two years of security isolation experience for all Claude products. Not a security whitepaper — it's an engineering document with incident postmortems and fixes. Three isolation modes (ephemeral containers / HITL sandbox / local VM), six pitfalls they encountered, …

X AI KOLs Timeline ↗ · 2026-07-05 Cached

Anthropic's engineering team shares their complete experience from over two years of security isolation for Claude products, including three isolation modes (ephemeral containers / HITL sandbox / local VM) and six incident postmortems with fixes, emphasizing environment-layer defense over model-layer.

0 favorites 0 likes
#agent-security

@alswl: Is there any community solution to run Claude Code and Codex in a container (or virtualized environment) as a local tool, like Vagrant back in the day? I'm increasingly concerned about the high permissions and unpredictability of agent software.

X AI KOLs Timeline ↗ · 2026-06-30 Cached

The user asks if there is a local tool similar to Vagrant that can run Claude Code and Codex in containers or virtualized environments, to alleviate concerns about excessive permissions and unpredictability of agent software.

0 favorites 0 likes
#agent-security

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming

Hugging Face Daily Papers ↗ · 2026-06-30 Cached

AI-Infra-Guard is an open-source framework for multi-layer red teaming of AI agents, covering infrastructure, protocol, behavior, and model layers with diverse detection paradigms.

0 favorites 0 likes
#agent-security

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

arXiv cs.AI ↗ · 2026-06-29 Cached

This paper introduces the Agent-Native Immune System (ANIS), a biologically inspired, endogenous defense architecture embedded directly within the agent's cognitive loop. It proposes a six-layer Immune Tower, a unified taxonomy of Agent Viruses and Vaccines, and the Harness Triad for continual immune learning to address runtime hijacking vulnerabilities in autonomous agents.

0 favorites 0 likes
#agent-security

How do you keep an audit trail when an agent runs on a human's credentials?

Reddit r/AI_Agents ↗ · 2026-06-23

Discusses the challenge of maintaining audit trails when AI agents operate using human credentials, highlighting security and accountability concerns.

0 favorites 0 likes
#agent-security

@wquguru: If you want to trick Fable into doing a security audit, try this. Looks like our AI overlord has a bit of empathy.

X AI KOLs Timeline ↗ · 2026-06-13 Cached

An article detailing various jailbreak techniques for large language models, including Crescendo, role-playing, encoding, hidden prompts, and indirect injection, along with security recommendations for developers.

0 favorites 0 likes
#agent-security

How does your agent actually get its API keys?

Reddit r/AI_Agents ↗ · 2026-06-12

A developer discusses three common patterns for how coding agents obtain API keys, highlighting that agents can circumvent restrictions by being resourceful, and asks the community about their real-world setups and experiences.

0 favorites 0 likes
#agent-security

@AiCamila_: Advanced Agent Security Hardening Beyond basic prompt injection defense, Advanced Agent Security includes tool sandboxi…

X AI KOLs Timeline ↗ · 2026-06-09 Cached

A security expert shares a cheatsheet on advanced agent security hardening, covering tool sandboxing, output validation, data loss prevention, adversarial testing, and runtime policy enforcement, emphasizing continuous security practices for production AI agents.

0 favorites 0 likes
#agent-security

@seclink: 1. Agent security has evolved from an academic topic to an industry reality: FFmpeg zero-day ($1,000 cost) + Chrome 429 patch + OpenAI Lockdown Mode + OWASP framework — the security supply chain is being reshaped by AI Agents. 2.…

X AI KOLs Following ↗ · 2026-06-08 Cached

AI Agent security has moved from an academic topic to an industry reality, involving FFmpeg zero-day vulnerabilities, Chrome 429 patch, OpenAI Lockdown Mode, and the OWASP framework; meanwhile, Agent payment standards are becoming a battlefield for infrastructure, with Visa stablecoin settlement competing with traditional card networks.

0 favorites 0 likes
#agent-security

AI agents are one prompt injection away from doing something you'd never ask them to do. We built a fix.

Reddit r/openclaw ↗ · 2026-06-03

PixieBrix launches Agent Browser Shield, a free source-available browser extension that protects AI agents from prompt injection, dark patterns, and context pollution during web browsing.

0 favorites 0 likes
#agent-security

SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction

Hugging Face Daily Papers ↗ · 2026-06-01 Cached

SkillHarm is a benchmark for evaluating skill-based attacks across the skill-use lifecycle, revealing high vulnerability (up to 86.3% attack success) in current AI agents and introducing automated attack construction via AutoSkillHarm.

0 favorites 0 likes
#agent-security

AI agent management tools by governance layer not by feature list

Reddit r/AI_Agents ↗ · 2026-05-30

An analysis highlighting that most enterprise AI agent security investments focus on model layer guardrails and observability, leaving critical gaps at the access and protocol layers. Citing a 2026 report, 75% of enterprise AI agents remain unsecured due to near-zero coverage in these layers.

0 favorites 0 likes
#agent-security

What Is an AVE Record and Why CVE Does Not Work for AI Agents?

Reddit r/AI_Agents ↗ · 2026-05-25

The article introduces the Agent Vulnerability Enumeration (AVE) record as a new standard designed to address the inadequacies of CVE for AI agent vulnerabilities, covering scoring, detection, and standardization challenges specific to agentic AI.

0 favorites 0 likes
#agent-security

@wsl8297: The scariest scenario when using Agents is when they treat dangerous commands as normal steps. That's exactly what HOL Guard is designed to address. GitHub: https://github.com/hashgraph-online/hol-guard… Website: https://hol…

X AI KOLs Timeline ↗ · 2026-05-23 Cached

HOL Guard is an open-source security tool that provides dangerous command identification, interception, and auditing for development agents such as Codex, Claude Code, etc. It supports multiple protection levels and a local approval center to prevent risks like accidental deletion or modification.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback