Tag
The article argues that AI agent safety should be addressed as an access-control problem, recommending practices like least privilege, allow-lists, and approval gates to prevent unintended actions.
Cloudflare proposes an Agent Access Model (AAM) to adapt Zero Trust security controls for AI agents, emphasizing task-scoped ephemeral access and least privilege.
As AI agents gain more capabilities, operational control—not model capability—becomes the primary challenge, raising critical questions about security, permissions, observability, and recovery.
This paper argues that applying content-safety refusal methods to AI agents is a category error—agentic harm lies in authority misuse rather than output—and proposes action alignment enforced outside the model via least privilege.
This paper investigates over-privileged tool selection in LLM agents, introducing ToolPrivBench to evaluate and mitigate unnecessary use of high-privilege tools. It finds that safety alignment does not ensure least-privilege choices, and proposes a post-training defense that reduces excessive privilege use without sacrificing performance.
This paper proposes Risk-Aware Causal Gating (RACG), a training-free mechanism that applies the principle of least privilege to LLM agent tool exposure, reducing attack surface from prompt injection by only exposing high-risk tools when authorized and causally necessary.
A blog post discussing techniques for dropping privileges in Go programs to enforce the principle of least privilege, including chroot and user switching.