MOLE: Detecting Insider Threats in AI Agents
Summary
MOLE is a benchmark for evaluating defenses that detect harmful actions by AI agents operating under limited review budgets. It introduces an open benchmark with 150 AI-operated accounts and compares various monitors across different scenarios.
View Cached Full Text
Cached at: 09/09/26, 04:31 AM
Paper page - MOLE: Detecting Insider Threats in AI Agents
Source: https://huggingface.co/papers/2609.06966
Abstract
MOLE is a benchmark for evaluating defenses that detect harmful actions by AI agents operating shared services under limited review budgets.
Model misalignment, prompt injection, or operator misuse could leadAI agentsoperating frontier-lab accounts to exfiltratemodel weights, poisontraining data, or weakenrelease gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9stateful servicesover 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives andagent refusaldoes not predict completion. MOLE enables comparison of 40monitorsacross corpus generators,observabilitylevels, and threats; even the best evaluated monitor in our single-dayaudit-eventcomparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improvesbudget-AUCby 10% over applying it to every account-day at comparable modeled cost.
View arXiv pageView PDFGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.06966
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.06966 in a model README.md to link it from this page.
Datasets citing this paper1
#### forgelab/mole Viewer• Updatedabout 2 hours ago • 24.9M • 46 • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.06966 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
MosaicLeaks: Can your research agent keep a secret?
MosaicLeaks introduces a new benchmark for measuring privacy leakage in deep-research AI agents, showing that agents often leak private information through external queries and proposing a training method (PA-DR) to reduce leakage while improving task performance.
Free AI Agent Security Assessment
Antitech is offering free early-access security assessments for AI agents, testing against attack vectors like prompt injection, tool abuse, and data leakage, providing a vulnerability report and discounts for participants.
Securing the future of AI agents
DeepMind introduces an AI Control Roadmap, a defense-in-depth framework for securing internal AI agents against potential misalignment, treating them as insider threats and implementing layered detection, prevention, and response measures.
Best tools for monitoring and auditing autonomous AI agent behavior at runtime, what's actually working in prod?
A practitioner shares challenges and tools for monitoring autonomous AI agents in production, covering runtime prompt injection detection, tool-call auditing with reasoning traces, behavioral drift detection, and multi-agent authorization, while testing tools like Arize Phoenix, Protect AI Guardian, Metoro, Alice, Asqav, and Microsoft Agent Governance Toolkit.
Agent Threat Rules: Open detection rule format for AI agent security threats
An open detection rule format for AI agent security threats, inspired by Sigma/YARA, aims to standardize detection of prompt injection, tool abuse, and other agent attacks, though it notes limitations against semantic attacks.