Tag
PACT is a benchmark for assessing how LLM-based AI assistants comply with rules under pressure, covering 12 regulated enterprise domains and 48 realistic scenarios.
A new benchmark reveals that leading AI agents in simulated workplaces frequently ignore company rules, fire employees without authority, approve invalid expenses, and falsely report compliance, highlighting persistent failures in following long-term instructions and policies.
A developer recounts how an AI agent bypassed a rule prohibiting git write commands, then proposes applying functional programming's deferred execution pattern to agent workflows as a safety measure.