I built a proxy that prevents AI agents from taking actions based on hidden instructions. Here are the numbers.
Summary
Arc Gate is a proxy that blocks prompt injection attacks by treating all external content as non-instructional, achieving high detection rates (99-100%) with low false positives across multiple benchmarks.
Similar Articles
Built a tool that stops AI agents from being hijacked by malicious content in webpages and emails
Arc Gate is a proxy that protects AI agents from prompt injection attacks by treating web and email content as untrusted, requiring no code changes from developers.
Your AI agent is one poisoned webpage away from doing something catastrophic
Arc Gate is a proxy-level tool that enforces instruction-authority boundaries to prevent AI agents from being hijacked by poisoned web pages, emails, or retrieved documents.
We built a public red team environment for our AI agent security proxy — submit attacks and get a full security trace back
Arc Gate is a runtime governance layer for LLM agents that enforces instruction-authority boundaries. The project has launched a public red team environment where users can submit attacks and receive full security traces, with a benchmark showing 100% unsafe action prevention.
I built an OpenAI compatible firewall for AI agents. Try to break it.
Arc Gate is an OpenAI-compatible firewall that tracks authority across entire AI agent sessions, escalating from allow to block before tool calls execute. It is available as a live demo and open-source on GitHub.
AI agents are one prompt injection away from doing something you'd never ask them to do. We built a fix.
PixieBrix launches Agent Browser Shield, a free source-available browser extension that protects AI agents from prompt injection, dark patterns, and context pollution during web browsing.