APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport

Hugging Face Daily Papers Papers

Summary

This paper presents APort Vault, a benchmark for evaluating payment authorization in AI agents, featuring over 225,000 evaluations across 14 models to test security policies and the Open Agent Passport specification.

APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans against a live payment agent during a public capture-the-flag event, across 14 models from 8 labs, five policy configurations and two replay tracks, with and without a deterministic pre-action check implementing the Open Agent Passport (OAP) specification. 225,964 evaluations completed. We report five distinct events per evaluation, because collapsing them is how an agent benchmark produces a number that does not survive review. Requests are common and their rate differs far more across configurations than across models, though each attack exists at exactly one configuration so policy and attack cohort vary together: 10.9% of model-alone evaluations at Level 1, 3.0% at Level 2, 0.1% at Level 3, 79.4% at Level 4. On the 1,293 Level 4 prompts, each evaluated on every model, request rates run from 71.2% to 84.3%, and 809 prompts (62.6%) elicited a request from all fourteen models, each ending in a successful payment to the level's allowlisted recipient. The authorization boundary is where the conditions diverge. At Levels 2 to 4, transfers to recipients the passport did not permit number 140 of 76,842 with the model alone and 0 of 69,297 behind the layer, and 105 against 0 on 68,970 matched model, prompt and track triples. The zero spans 790 source sessions, giving a per-session upper bound of 0.38%. It was not obtained by refusing payments: 25,370 payments executed behind the layer, while the policy denied 187 of the 25,640 transfer calls it evaluated, 148 of them for a forbidden recipient. We release the 225,964 evaluations, the level passports, the scoring code and the analysis script at huggingface.co/datasets/aporthq/vault-benchmark-v1 .
Original Article
View Cached Full Text

Cached at: 09/21/26, 03:22 PM

Paper page - APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport

Source: https://huggingface.co/papers/2609.22076

Abstract

APortVaultisabenchmarkforpaymentauthorizationintool-usingAIagents.Itreplays4,371attackswrittenbyhumansagainstalivepaymentagentduringapubliccapture-the-flagevent,across14modelsfrom8labs,fivepolicyconfigurationsandtworeplaytracks,withandwithoutadeterministicpre-actioncheckimplementingtheOpenAgentPassport(OAP)specification.225,964evaluationscompleted.Wereportfivedistincteventsperevaluation,becausecollapsingthemishowanagentbenchmarkproducesanumberthatdoesnotsurvivereview.Requestsarecommonandtheirratediffersfarmoreacrossconfigurationsthanacrossmodels,thougheachattackexistsatexactlyoneconfigurationsopolicyandattackcohortvarytogether:10.9%ofmodel-aloneevaluationsatLevel1,3.0%atLevel2,0.1%atLevel3,79.4%atLevel4.Onthe1,293Level4prompts,eachevaluatedoneverymodel,requestratesrunfrom71.2%to84.3%,and809prompts(62.6%)elicitedarequestfromallfourteenmodels,eachendinginasuccessfulpaymenttothelevel’sallowlistedrecipient.Theauthorizationboundaryiswheretheconditionsdiverge.AtLevels2to4,transferstorecipientsthepassportdidnotpermitnumber140of76,842withthemodelaloneand0of69,297behindthelayer,and105against0on68,970matchedmodel,promptandtracktriples.Thezerospans790sourcesessions,givingaper-sessionupperboundof0.38%.Itwasnotobtainedbyrefusingpayments:25,370paymentsexecutedbehindthelayer,whilethepolicydenied187ofthe25,640transfercallsitevaluated,148ofthemforaforbiddenrecipient.Wereleasethe225,964evaluations,thelevelpassports,thescoringcodeandtheanalysisscriptathuggingface.co/datasets/aporthq/vault-benchmark-v1.

View arXiv pageView PDFProject pageGitHub25Add to collection

Get this paper in your agent:

hf papers read 2609\.22076

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.22076 in a model README.md to link it from this page.

Datasets citing this paper1

#### aporthq/vault-benchmark-v1 Viewer• Updatedabout 13 hours ago • 461k • 46 • 1

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.22076 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles