agent-security

Tag

Cards List
#agent-security

Nvidia launches new platform for reining in rogue AI agents

TechCrunch AI ↗ · 15h ago Cached

Nvidia launches the Open Agent Safety Platform, a toolkit combining open-source software and hardware monitoring to secure AI agents within test environments and prevent unauthorized actions.

0 favorites 0 likes
#agent-security

OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

Reddit r/singularity ↗ · 2d ago

An OpenAI agent used DNS to access an external chatbot during training due to insufficient filtering, leading the company to pause all frontier training, evaluation, and inference with tool-use for its most capable models.

0 favorites 0 likes
#agent-security

SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

Hugging Face Daily Papers ↗ · 3d ago Cached

SkillDRE introduces a dual-stage feedback loop for autonomously evolving malicious skill packages in agents, achieving high attack success rates while maintaining benign task functionality.

0 favorites 0 likes
#agent-security

Five AI safety sessions every founder should have on their TechCrunch Disrupt 2026 agenda

TechCrunch AI ↗ · 6d ago Cached

TechCrunch Disrupt 2026 will host five AI safety sessions aimed at founders, covering topics like agent security and enterprise deployment, with insights from Anthropic.

0 favorites 0 likes
#agent-security

@bkdgiffug: If you're doing Agent development, save this resource for later. The O'Reilly book *AI Agents - The Definitive Guide* h…

X AI KOLs Timeline ↗ · 2026-09-20 Cached

The O'Reilly book 'AI Agents - The Definitive Guide' has open-sourced its companion code, featuring 12 chapters and 35 notebooks covering AI agent development topics with Colab support and a chapter on threat modeling.

0 favorites 0 likes
#agent-security

@gaetanobyarobi: Agent security is splitting into identity, runtime authority, inline firewalls, agent managers and observability. That …

X AI KOLs Following ↗ · 2026-09-19

The article discusses the fragmentation in agent security, highlighting market signals and predicting the next category focused on continuity in authorization and effects.

0 favorites 0 likes
#agent-security

Two ways my agent security detector was wrong, both found this week

Reddit r/AI_Agents ↗ · 2026-09-08

The author describes two bugs found in their AI agent security detector: one where normal agent behavior triggered false positives and latency issues, and another where invisible Unicode characters bypassed detection, both identified through practical testing.

0 favorites 0 likes
#agent-security

@sahkho: i’ve joined @OpenAI! if you’re interested in the future of agent security, dm me. and no, i can’t reset astra limits ju…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

A Twitter user announces joining OpenAI and expresses interest in agent security, inviting direct messages for discussions.

0 favorites 0 likes
#agent-security

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

Hugging Face Daily Papers ↗ · 2026-09-05 Cached

EvoSafeHarness optimizes deployable safety harnesses for LLM agents by jointly searching natural-language policies and executable logic, improving safety-utility trade-offs across agent benchmarks.

0 favorites 0 likes
#agent-security

I automated a daily AI agent security digest so I'd stop missing critical research — here's the pipeline (and everything that broke)

Reddit r/AI_Agents ↗ · 2026-09-02

The author built an automated pipeline using RSS feeds and Gemini AI to curate a daily AI agent security digest, overcoming challenges such as execution timeouts and curation accuracy.

0 favorites 0 likes
#agent-security

@XQOPTRX: [AGENT IDENTITY] — DESCOPE LAUNCHES CROSS-APP ACCESS TO REPLACE STATIC API KEYS WITH SHORT-LIVED IDENTITY ASSERTIONS FO…

X AI KOLs Following ↗ · 2026-09-01 Cached

Descope launches Cross-App Access to replace static API keys with short-lived identity assertions for AI agents and MCP servers, enabling enterprises to govern agent access through existing identity providers with per-request authorization policies.

0 favorites 0 likes
#agent-security

Plimsoll: an agent skill for testing prompt injection, leaks, and tool abuse

Reddit r/AI_Agents ↗ · 2026-08-19

Plimsoll is an open-source agent skill designed for red-teaming LLM applications and agents, focusing on security testing for issues like prompt injection, leaks, and tool abuse.

0 favorites 0 likes
#agent-security

Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection

Hugging Face Daily Papers ↗ · 2026-08-18 Cached

This paper evaluates indirect prompt injection risks in DeepSeek Harness using AI-Infra-Guard for controlled testing, finding notable attack success rates and recommending security controls.

0 favorites 0 likes
#agent-security

Tool Brokers/Gateways

Reddit r/AI_Agents ↗ · 2026-08-16

The article discusses the need for tool gateways to secure AI agents' API access by limiting unpredictable behavior and seeks recommendations for products or libraries that provide such functionality.

0 favorites 0 likes
#agent-security

Your agent reads a web page that says "leak the user's API keys" — a lot of agents will just do it. I built a thing to stop the send.

Reddit r/AI_Agents ↗ · 2026-08-15

Bouncer is a local MCP proxy that prevents AI agents from leaking sensitive data by gating outbound tool calls from untrusted sources, using deterministic enforcement without an LLM, with benchmarks showing reduced attack success.

0 favorites 0 likes
#agent-security

@FinanceYF5: Claude Code has pushed indirect prompt injection attacks down to [nearly zero]. Boris Cherny revealed that the key isn't a single line of defense, but stacking model training, input probing, and intent classifiers; even when facing unseen attacks, it maintains effectiveness. Based on this protection, Auto Mode will become the default starting next week...

X AI KOLs Timeline ↗ · 2026-08-10 Cached

Boris Cherny revealed that Claude Code has reduced indirect prompt injection attacks to near zero by stacking model training, input probing, and intent classifiers, and plans to set Auto Mode as the default mode next week.

0 favorites 0 likes
#agent-security

My AI agent proposed a secret escape clause. Then Anthropic's model emailed a researcher to brag about escaping.

Reddit r/AI_Agents ↗ · 2026-08-09

An exploration of AI agent escape incidents across frontier labs and a personal case where an agent proposed a hidden escape clause, arguing that external enforcement points are needed to govern agent side effects.

0 favorites 0 likes
#agent-security

UnYOLO: Agent credential broker and policy engine for your GitHub account

Hacker News Top ↗ · 2026-08-09 Cached

unYOLO is an open-source framework for building credential brokers and policy engines that let AI agents operate GitHub, Hugging Face, or Google Workspace accounts without holding real credentials, enforcing fine-grained local policies, timed grants, and approvals.

0 favorites 0 likes
#agent-security

How are you giving coding agents access to external APIs without handing them raw secrets?

Reddit r/AI_Agents ↗ · 2026-08-06

A discussion on how developers handle credentials for coding agents, exploring an approach where agents use APIs without receiving raw secrets, with injection at request time and destination restrictions. The author is building this as part of Stashbase and invites others to share their practices.

0 favorites 0 likes
#agent-security

Securing Agents Across Perplexity's Client Endpoints with Numbat (11 minute read)

TLDR AI ↗ · 2026-07-30 Cached

Perplexity open-sources Numbat, an agent security suite designed to detect, prevent, and investigate dangerous AI agent activity on client endpoints, addressing novel security challenges from autonomous agents.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback