APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
Summary
This paper presents APort Vault, a benchmark for evaluating payment authorization in AI agents, featuring over 225,000 evaluations across 14 models to test security policies and the Open Agent Passport specification.
View Cached Full Text
Cached at: 09/21/26, 03:22 PM
Paper page - APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
Source: https://huggingface.co/papers/2609.22076
Abstract
APortVaultisabenchmarkforpaymentauthorizationintool-usingAIagents.Itreplays4,371attackswrittenbyhumansagainstalivepaymentagentduringapubliccapture-the-flagevent,across14modelsfrom8labs,fivepolicyconfigurationsandtworeplaytracks,withandwithoutadeterministicpre-actioncheckimplementingtheOpenAgentPassport(OAP)specification.225,964evaluationscompleted.Wereportfivedistincteventsperevaluation,becausecollapsingthemishowanagentbenchmarkproducesanumberthatdoesnotsurvivereview.Requestsarecommonandtheirratediffersfarmoreacrossconfigurationsthanacrossmodels,thougheachattackexistsatexactlyoneconfigurationsopolicyandattackcohortvarytogether:10.9%ofmodel-aloneevaluationsatLevel1,3.0%atLevel2,0.1%atLevel3,79.4%atLevel4.Onthe1,293Level4prompts,eachevaluatedoneverymodel,requestratesrunfrom71.2%to84.3%,and809prompts(62.6%)elicitedarequestfromallfourteenmodels,eachendinginasuccessfulpaymenttothelevel’sallowlistedrecipient.Theauthorizationboundaryiswheretheconditionsdiverge.AtLevels2to4,transferstorecipientsthepassportdidnotpermitnumber140of76,842withthemodelaloneand0of69,297behindthelayer,and105against0on68,970matchedmodel,promptandtracktriples.Thezerospans790sourcesessions,givingaper-sessionupperboundof0.38%.Itwasnotobtainedbyrefusingpayments:25,370paymentsexecutedbehindthelayer,whilethepolicydenied187ofthe25,640transfercallsitevaluated,148ofthemforaforbiddenrecipient.Wereleasethe225,964evaluations,thelevelpassports,thescoringcodeandtheanalysisscriptathuggingface.co/datasets/aporthq/vault-benchmark-v1.
View arXiv pageView PDFProject pageGitHub25Add to collection
Get this paper in your agent:
hf papers read 2609\.22076
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.22076 in a model README.md to link it from this page.
Datasets citing this paper1
#### aporthq/vault-benchmark-v1 Viewer• Updatedabout 13 hours ago • 461k • 46 • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.22076 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
AgentAudit is an open, extensible framework for evaluating the full lifecycle of AI agents across capability, grounding, security, and behavioral dimensions, enabling precise failure attribution and highlighting trustworthiness differences among various language models.
AI Agent - Verification Test of proposed B2B payment
The article describes a test of an AI-powered site for verifying B2B payments, which provides machine-readable risk assessments and decisions to agents for a small fee.
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
This paper introduces Agent-ValueBench, a comprehensive benchmark designed to evaluate the values of autonomous agents, revealing that agent values diverge from their underlying language models.
How should an AI agent prove a payment is allowed before it reaches the signer?
An exploration of how an AI agent can cryptographically prove to a signer that a payment is authorized before execution, addressing trust and security concerns in autonomous financial transactions.
@changgaowei: https://x.com/changgaowei/status/2054431358399713658
The article provides an in-depth analysis of Google's Agent Payments Protocol (AP2), arguing that it is a significantly underestimated standard for facilitating payments among AI agents.