defense-evaluation

Tag

Cards List
#defense-evaluation

MOLE: Detecting Insider Threats in AI Agents

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

MOLE is a benchmark for evaluating defenses that detect harmful actions by AI agents operating under limited review budgets. It introduces an open benchmark with 150 AI-operated accounts and compares various monitors across different scenarios.

0 favorites 0 likes
← Back to home

Submit Feedback