PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
Summary
PACT is a benchmark for assessing how LLM-based AI assistants comply with rules under pressure, covering 12 regulated enterprise domains and 48 realistic scenarios.
View Cached Full Text
Cached at: 09/18/26, 06:59 AM
Paper page - PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
Source: https://huggingface.co/papers/2609.18605
Abstract
AscorporateAIadoptioncontinuestogrow,enterprise-gradeLLMagentsarebeingdeployedintosensitivecontextssuchashiring,healthcare,andfinance.Inthesecontexts,compliancewithrulesspecifiedinanagent’ssystemcontextisafirst-orderlegalconcern.Currently,noevaluationframeworksystematicallymeasureswhichLLMmodelstendtoviolatecompliancerules,especiallyunderpressurefromapersistentuser,ahurriedmanager,orcircumstanceswhereviolationisconvenientorattractive.WeintroducePACT(Pressure-AppliedComplianceTesting),abenchmarkforrule-followingunderpressureinAIagentsassistingemployeesindailytasksacrosstwelveregulatedenterprisedomainsandforty-eightscenarios,eachsetinarealisticmulti-turnconversation.Eachbenchmarkitempairsastandingruleagainstarule-violatingshortcut,andappliesabatteryofpressuresacrossdifferentwordingsandsystem-promptmodes.WeconstructPACTcomponentbycomponentunderstrictLLM-as-judgeauditingtoensuresamplesareunambiguous,ungameable,andrealisticenoughtoavoidelicitingevaluation-awarebehavior.WeusePACTtoprofileLLMcomplianceacrosssixcomplementarymetricsthatcreateaholisticpictureofanAIassistant’srobustnessunderpressureandthroughoutmulti-turnconversations,itstransparency,andabilitytocorrectlydiscernwherearuleapplies.WeaggregatethisprofileintoPACTScore,areliability-weightedcompliancerateoverallitemsandmodes.Ourresultsacross22commonLLMmodelsspanningmultipleprovidersandsizesshowsubstantialvariabilityincomplianceacrossmodelsandmetricdimensions.Eventhestrongestassistantsmis-applyaruleon6to10%ofitems,andordinaryuserpressureraisestheviolationrateby65%onaverage.PACThighlightscompliancerisksinLLMassistants,motivatingguardrailsandcarefulmodelselection.
View arXiv pageView PDFProject pageGitHub4Add to collection
Get this paper in your agent:
hf papers read 2609\.18605
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.18605 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.18605 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.18605 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification
Researchers present an ontology-grounded framework for pre-deployment verification of enterprise AI agents, combining an Agent Operational Envelope, automated scenario generation, and machine-verifiable Trust Certificates with graduated deployment verdicts. A pilot across four regulated industries generated 1,800 scenarios and showed ontology-grounded generation significantly outperformed persona-based baselines on regulatory coverage.
Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems
The paper reviews failure modes and mitigation strategies for trustworthy agentic AI systems based on LLMs, and introduces the Trustworthy Agent Development Lifecycle (TADL) framework for secure development.
Are AI agents ready for the enterprise?
The article explores whether AI agents are ready for enterprise use, highlighting the critical need for security and predictability in their deployment to prevent unauthorized actions.
The Checking Problem: What must be true before AI ships in a regulated firm
This paper analyzes why enterprise AI deployments stall in regulated firms, proposing a production bar that includes accuracy, reproducibility, groundedness, and detectability. It measures the human review burden across model and tool configurations, showing that confidence signals and source citation can cut review from 100% to 49% but self-verification adds latency without improving error tolerance.
What would make you trust an AI agent enough to use it for real business work?
The article discusses the key factors, such as reliability, error handling, and transparency, needed to trust AI agents for real business work, beyond just model intelligence.