MOLE: Detecting Insider Threats in AI Agents

Hugging Face Daily Papers Papers

Summary

MOLE is a benchmark for evaluating defenses that detect harmful actions by AI agents operating under limited review budgets. It introduces an open benchmark with 150 AI-operated accounts and compares various monitors across different scenarios.

Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.
Original Article
View Cached Full Text

Cached at: 09/09/26, 04:31 AM

Paper page - MOLE: Detecting Insider Threats in AI Agents

Source: https://huggingface.co/papers/2609.06966

Abstract

MOLE is a benchmark for evaluating defenses that detect harmful actions by AI agents operating shared services under limited review budgets.

Model misalignment, prompt injection, or operator misuse could leadAI agentsoperating frontier-lab accounts to exfiltratemodel weights, poisontraining data, or weakenrelease gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9stateful servicesover 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives andagent refusaldoes not predict completion. MOLE enables comparison of 40monitorsacross corpus generators,observabilitylevels, and threats; even the best evaluated monitor in our single-dayaudit-eventcomparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improvesbudget-AUCby 10% over applying it to every account-day at comparable modeled cost.

View arXiv pageView PDFGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.06966

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.06966 in a model README.md to link it from this page.

Datasets citing this paper1

#### forgelab/mole Viewer• Updatedabout 2 hours ago • 24.9M • 46 • 1

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.06966 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

MosaicLeaks: Can your research agent keep a secret?

Hugging Face Blog

MosaicLeaks introduces a new benchmark for measuring privacy leakage in deep-research AI agents, showing that agents often leak private information through external queries and proposing a training method (PA-DR) to reduce leakage while improving task performance.

Free AI Agent Security Assessment

Reddit r/AI_Agents

Antitech is offering free early-access security assessments for AI agents, testing against attack vectors like prompt injection, tool abuse, and data leakage, providing a vulnerability report and discounts for participants.

Securing the future of AI agents

Google DeepMind Blog

DeepMind introduces an AI Control Roadmap, a defense-in-depth framework for securing internal AI agents against potential misalignment, treating them as insider threats and implementing layered detection, prevention, and response measures.