Audited 1,228 human interventions in our AI agent setup. 91% weren't decisions, so we stopped letting agents say "done".
Summary
Audited 1,228 human interventions in AI agent workflows, revealing that 91% weren't decisions, prompting a new verification service with work contracts and shadow mode to improve reliability.
Similar Articles
How do you actually know your AI agent did what it says it did?
The article discusses the challenge of verifying AI agent actions and advocates for immutable receipts to ensure trust and distinguish between bad decisions and non-existent ones.
Everyone caps their agent so a human can still check the output. Has anyone actually solved that?
The article questions the common practice of limiting AI agent runs for human verification and explores structural alternatives when task volumes exceed human oversight capacity.
Are AI agents actually doing a good job, or are we overhyping them?
An inquiry into the real-world effectiveness of AI agents, highlighting challenges in production such as inefficiency and error rates, and inviting community experiences.
AI Agent Audits ?
A practitioner shares concerns about an upcoming audit revealing undocumented AI agents in production, highlighting governance gaps and risks with customer PII access.
Do we trust AI agents too much once they start completing tasks successfully?
The article questions whether we become overly trusting of AI agents after they perform tasks successfully, highlighting risks of unnoticed errors and debating the need for verification layers.