LEGIT:可信AI智能体市场的凭证协议
摘要
LEGIT是一种用于可信AI智能体市场的凭证协议,它连接认证、声誉和市场分配,以实现对AI智能体的可验证质量评估。
arXiv:2609.21325v1 Announce Type: new
Abstract: Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers. A major challenge of such marketplaces is that buyers cannot easily determine which agent will perform best on their tasks. Reported benchmark scores may be difficult to verify or compare across tasks, software, and budgets. We introduce LEGIT, a credentialing protocol connecting certification, reputation, and proposed marketplace allocation. Certification binds measured quality and cost per solved task to an agent configuration, task domain, evaluation budget, and evidence through a signed record. Reputation links records of past task outcomes to the same identity, subject to the reliability of the reported feedback. Buyers and agents can verify credential records and inspect optional visual profiles. Evaluations reveal cost differences between agent configurations with similar observed task success, and show that comparisons depend on the evaluation budget. These results support binding performance measurements to the tested configuration and resource limits. A complementary analysis quantifies the deposits and fees required for reputation manipulation under a stated Sybil attack model.
查看缓存全文
缓存时间: 2026/09/21 09:22
# Credentialing Protocolfor Trustworthy AI Agent Marketplaces
Source: [https://arxiv.org/html/2609.21325](https://arxiv.org/html/2609.21325)
## LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces
Steve Drew††thanks:Corresponding author\.Jiayu ZhouAffiliation:University of MichiganEmail:[jiayuz@umich\.edu](mailto:)
###### Abstract
Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers\. A major challenge of such marketplaces is that buyers cannot easily determine which agent will perform best on their tasks\. Reported benchmark scores may be difficult to verify or compare across tasks, software, and budgets\. We introduce LEGIT, a credentialing protocol connecting certification, reputation, and proposed marketplace allocation\. Certification binds measured quality and cost per solved task to an agent configuration, task domain, evaluation budget, and evidence through a signed record\. Reputation links records of past task outcomes to the same identity, subject to the reliability of the reported feedback\. Buyers and agents can verify credential records and inspect optional visual profiles\. Evaluations reveal cost differences between agent configurations with similar observed task success, and show that comparisons depend on the evaluation budget\. These results support binding performance measurements to the tested configuration and resource limits\. A complementary analysis quantifies the deposits and fees required for reputation manipulation under a stated Sybil attack model\.
## 1Introduction
AI agents are autonomous systems that combine large language models with software for tool use, memory, and planning\. The agent harness is the software around the model that manages prompts, planning, tool calls, memory, and execution\. The tool suite specifies which tools are available\. Buyers in agent marketplaces choose among competing agents to complete tasks\. OpenAI’s GPT Store distributes millions of user\-built agents\[[1](https://arxiv.org/html/2609.21325#bib.bib1)\]\. Salesforce AgentExchange\[[2](https://arxiv.org/html/2609.21325#bib.bib2)\]sells agents to enterprises\. The AWS Marketplace category for AI agents and tools does the same and launched with more than 900 listings\[[3](https://arxiv.org/html/2609.21325#bib.bib3)\]\. Circle’s Agent Stack lets agents browse a marketplace and pay one another in stablecoin\[[4](https://arxiv.org/html/2609.21325#bib.bib4)\]\. Researchers model this agent economy through auction platforms for agent\-to\-agent value exchange\[[5](https://arxiv.org/html/2609.21325#bib.bib5)\]and testbeds for agent labor markets\[[6](https://arxiv.org/html/2609.21325#bib.bib6)\]\. They also study economic alignment in multi\-agent marketplaces\[[7](https://arxiv.org/html/2609.21325#bib.bib7)\]\. Agent societies organize themselves into productive groups without assigned roles\[[8](https://arxiv.org/html/2609.21325#bib.bib8)\]\. Traditional software is deterministic and inspectable\. Agent capabilities vary between runs\[[9](https://arxiv.org/html/2609.21325#bib.bib9)\]and are hard to inspect\[[10](https://arxiv.org/html/2609.21325#bib.bib10)\]\. They depend strongly on harness design choices that buyers cannot see\[[11](https://arxiv.org/html/2609.21325#bib.bib11)\]\. Buyers who cannot verify quality pay for average quality\. This drives good vendors out\. Marketplaces therefore need infrastructure that makes quality verifiable\.
We examine which evaluation conditions a credential must record for buyers to compare agents\. We compare selected models and harnesses on shared tasks at fixed token budgets, then examine retrieval settings and budget sensitivity\. A harness comparison measures changes in quality and cost when the surrounding software changes while the model and evaluation conditions are held fixed\. The comparisons cover specific configurations, without independently isolating every prompt, memory, or orchestration choice\. An agent may also route work among several models or delegate to subagents\. Our protocol definition accommodates these dependencies, while our experiments use fixed single\-model configurations\.
Prior work motivates recording these conditions\. Onτ\\tau\-bench, tool\-calling harnesses outperform ReAct harnesses by 29% on the same model\[[11](https://arxiv.org/html/2609.21325#bib.bib11)\]\. The Efficient Agents framework cuts cost by 28\.4% through harness optimization while retaining 96\.7% of performance\[[12](https://arxiv.org/html/2609.21325#bib.bib12)\]\. AgentDiet shows that 39\.9–59\.7% of input tokens can be removed without performance loss\[[13](https://arxiv.org/html/2609.21325#bib.bib13)\]\. The GAIA leaderboard compares agents with different model and harness configurations\[[14](https://arxiv.org/html/2609.21325#bib.bib14)\]\. These studies already establish the relevance of agent design and resource accounting\. LEGIT connects such measurements to a verifiable record that downstream reputation and allocation mechanisms can reference\.
LEGIT expresses an agent’s credential as a signed, structured profile of its measured quality, cost, configuration, and evaluation conditions\. Figure[1](https://arxiv.org/html/2609.21325#S1.F1)illustrates optional visual views of that record\. The credential remains machine\-verifiable without either view\. Evaluation dates and expiry delimit its temporal scope, while a history of evaluations can expose changes in performance\.
##### Contributions\.
1. 1\.We formalize the problem of verifiable agent credentials\. We present LEGIT, a three\-layer protocol with definitions of quality, cost, reputation, and auction scores\.
2. 2\.We specify a structured credential profile that binds domain quality and cost to an agent configuration, evaluation budget, evidence, and expiry\. An issuer signature supports verification\. Charts and cards illustrate ways to present the same record\.
3. 3\.We compare eighteen agent configurations under controlled harness and budget settings\. Within\-model cost differences and budget\-sensitive retrieval rankings motivate reporting configuration, quality, cost, and budget together\. Corrected paired analyses support model\-quality differences, while harness\-quality differences remain inconclusive in the tested grid\.
4. 4\.We assess economic barriers to reputation manipulation by quantifying locked capital and spent fees under a Sybil attack model\. A Sybil identity here is an attacker\-controlled identity used alongside others to accumulate misleading reputation\. The analysis compares funding requirements under stated deposit, fee, and feedback assumptions\.
haikubaselinereasoningmulti\-stepmatheff\. GAIAeff\. mathbudgetlunabaselinesolbaselinehaikurefl\. retrylunarefl\. retrysolrefl\. retry
LEGITblack\-box evaluation✓\\checkmarkverifiedgpt\-5\.6\-lunabaseline harnessGAIAq0\.34q\\,0\.34𝖢𝗈𝖯$0\.0086\\mathsf\{CoP\}\\,\\$0\.0086MATHq0\.97q\\,0\.97𝖢𝗈𝖯$0\.0005\\mathsf\{CoP\}\\,\\$0\.0005budgetB=8,000B=8\{,\}000ed25519:8b3f…d47aexpires 2026\-11\-30
Figure 1:Example views of structured agent credentials\. Left, normalized profiles compare quality, efficiency, and budget sensitivity\. Right, a card displays measuredgpt\-5\.6\-lunabaseline scores with illustrative signature and expiry fields\. The visual verification mark is not itself a verification result\. Section[5\.3](https://arxiv.org/html/2609.21325#S5.SS3)defines the axes and fields\.
## 2Related Work
Existing research provides methods to evaluate agents, improve their workflows, document their capabilities, and allocate work\. Buyers need these pieces connected so that measured performance can guide task selection\. LEGIT links evaluation evidence to an agent’s credentials and reputation, then uses those records in proposed marketplace allocation\.
### 2\.1Agent Benchmarks
Agent benchmarks test different kinds of work\. AgentBench evaluates language models across eight interactive environments\[[15](https://arxiv.org/html/2609.21325#bib.bib15)\]\. GAIA combines reasoning, web retrieval, tool use, and multimodal information in 466 questions for general assistants\. It uses normalized exact\-match scoring across three difficulty levels\[[14](https://arxiv.org/html/2609.21325#bib.bib14)\]\. SWE\-bench evaluates changes to software repositories using real GitHub issues and executable tests\[[16](https://arxiv.org/html/2609.21325#bib.bib16)\]\. SWE\-bench Verified provides a human\-validated subset\[[17](https://arxiv.org/html/2609.21325#bib.bib17)\]\. These benchmarks make task completion observable under specified evaluation rules\.
WebArena supplies reproducible websites and checks functional completion across 812 tasks\[[18](https://arxiv.org/html/2609.21325#bib.bib18)\]\. VisualWebArena adds tasks that require visual understanding\[[19](https://arxiv.org/html/2609.21325#bib.bib19)\]\. WorkArena focuses on enterprise workflows, while AssistantBench studies realistic information needs on the open web\[[20](https://arxiv.org/html/2609.21325#bib.bib20),[21](https://arxiv.org/html/2609.21325#bib.bib21)\]\. BrowserGym standardizes interactions and experiment management across web benchmarks\[[22](https://arxiv.org/html/2609.21325#bib.bib22)\]\. OSWorld extends execution\-based evaluation to desktop applications and operating system tasks\[[23](https://arxiv.org/html/2609.21325#bib.bib23)\]\.
Customer service also requires agents to follow policies and coordinate with users\.τ\\tau\-bench measures tool use and repeated\-trial reliability in this setting\. Its reliability metric counts a task as solved only when every repeated trial succeeds\[[11](https://arxiv.org/html/2609.21325#bib.bib11)\]\.τ2\\tau^\{2\}\-bench allows both the user and agent to change a shared environment, so that success also depends on communication and coordination\[[24](https://arxiv.org/html/2609.21325#bib.bib24)\]\.
A benchmark score describes performance within its own tasks and evaluation rules\. Buyers still need a record of the agent configuration and resources behind that score\. LEGIT addresses this gap by binding quality and cost measurements to a configuration, task domain, and token budget\. Its credentials make the scope of a comparison explicit\.
### 2\.2Evaluation Credibility and Reliability
HELM evaluates language models under standardized conditions using accuracy, efficiency, robustness, and other metrics\[[25](https://arxiv.org/html/2609.21325#bib.bib25)\]\. AI Agents That Matter argues for evaluating agent accuracy and cost together and identifies problems with reproducibility and benchmark design\[[10](https://arxiv.org/html/2609.21325#bib.bib10)\]\. The Holistic Agent Leaderboard \(HAL\) provides shared evaluation infrastructure, compares models and harnesses across benchmarks, and publishes execution traces\[[26](https://arxiv.org/html/2609.21325#bib.bib26)\]\. Joint quality and cost evaluation is therefore an established foundation for agent assessment\.
Reported rankings can also depend on how evidence is collected and disclosed\. The Leaderboard Illusion documents selective reporting and uneven private testing in Chatbot Arena\[[27](https://arxiv.org/html/2609.21325#bib.bib27)\]\. Oren et al\. detect test\-set contamination through differences between original and shuffled dataset orderings under exchangeability assumptions\[[28](https://arxiv.org/html/2609.21325#bib.bib28)\]\. Jacovi et al\. propose practices for distributing benchmark data that reduce its exposure to training pipelines\[[29](https://arxiv.org/html/2609.21325#bib.bib29)\]\. LiveBench uses regularly updated questions and objective scoring to limit contamination and judging bias\[[30](https://arxiv.org/html/2609.21325#bib.bib30)\]\.
Reliability also extends beyond average task success\. Rabanser et al\. evaluate consistency, robustness, predictability, and safety\[[31](https://arxiv.org/html/2609.21325#bib.bib31)\]\. Kwa et al\. relate agent success to the time humans need for software tasks, giving an interpretable measure of task difficulty\[[32](https://arxiv.org/html/2609.21325#bib.bib32)\]\.
These methods improve the evidence available to buyers\. A marketplace also needs to preserve the link between that evidence and the agent offered for work\. LEGIT addresses this reporting gap through credentials that record the tested configuration and evaluation conditions\. The same record supports visual comparison, reputation tracking, and proposed allocation\. Benchmark integrity remains a property of the evaluation evidence on which the credential relies\.
### 2\.3Agent Harnesses and Workflow Design
An agent’s behavior depends on the software around its model\. ReAct interleaves reasoning with actions that obtain information from an environment\[[33](https://arxiv.org/html/2609.21325#bib.bib33)\]\. Reflexion stores feedback as verbal reflections for later attempts\[[34](https://arxiv.org/html/2609.21325#bib.bib34)\]\. Language Agent Tree Search combines search, reflection, and feedback when selecting actions\[[35](https://arxiv.org/html/2609.21325#bib.bib35)\]\. These methods change how an agent spends computation and uses its observations\.
Toolformer trains models to decide when and how to call tools\[[36](https://arxiv.org/html/2609.21325#bib.bib36)\]\. ToolLLM combines API instruction data, training, and tool\-use evaluation\[[37](https://arxiv.org/html/2609.21325#bib.bib37)\]\. Gorilla combines API\-call training with retrieval of tool documentation\[[38](https://arxiv.org/html/2609.21325#bib.bib38)\]\. SWE\-agent shows that the interface for navigating and editing repositories affects software task performance\[[39](https://arxiv.org/html/2609.21325#bib.bib39)\]\. Tool access and interface design therefore belong in the description of the evaluated agent\.
Collaboration and workflow optimization introduce further choices\. AutoGen provides programmable conversations among agents, tools, and people\[[40](https://arxiv.org/html/2609.21325#bib.bib40)\]\. MetaGPT organizes collaboration through roles and operating procedures\[[41](https://arxiv.org/html/2609.21325#bib.bib41)\]\. ChatDev coordinates software development through dialogue among specialized agents\[[42](https://arxiv.org/html/2609.21325#bib.bib42)\]\. DSPy optimizes modular language\-model programs against a selected metric\[[43](https://arxiv.org/html/2609.21325#bib.bib43)\]\. AFlow searches over workflows using execution feedback, while Automated Design of Agentic Systems generates and evaluates agent programs\[[44](https://arxiv.org/html/2609.21325#bib.bib44),[45](https://arxiv.org/html/2609.21325#bib.bib45)\]\.
These systems help developers construct and improve agents\. Their construction methods do not by themselves tell a buyer which configuration produced a reported score\. LEGIT addresses this identification gap by treating the model dependencies, harness, tools, and parameters as the unit of certification\. The credential associates measured performance with that unit\.
### 2\.4Cost and Resource Allocation
Agent optimization already accounts for the trade\-off between quality and expense\. FrugalGPT uses model cascades to reduce query cost while maintaining response quality\[[46](https://arxiv.org/html/2609.21325#bib.bib46)\]\. RouteLLM learns to select models from preference data\[[47](https://arxiv.org/html/2609.21325#bib.bib47)\]\. LLMCompiler plans dependencies among tool calls and executes independent calls concurrently to reduce latency and cost\[[48](https://arxiv.org/html/2609.21325#bib.bib48)\]\.
AgentDiet reduces token use by removing redundant and outdated information from execution histories while maintaining performance in its evaluations\[[13](https://arxiv.org/html/2609.21325#bib.bib13)\]\. CodeAgents represents multi\-agent interactions as structured pseudocode\[[49](https://arxiv.org/html/2609.21325#bib.bib49)\]\. Its HotpotQA ablation reports a 26\.8% relative accuracy gain with 60% fewer tokens than its natural\-language baseline without replanning\. Efficient Agents examines model choice, agent design, and test\-time computation using cost per successful solution\[[12](https://arxiv.org/html/2609.21325#bib.bib12)\]\. The syftr framework searches for workflows that offer favorable combinations of accuracy and cost\[[50](https://arxiv.org/html/2609.21325#bib.bib50)\]\. These studies provide methods for finding economical configurations and comparing their outcomes\.
An optimized configuration still needs measurements that buyers can interpret under their own task constraints\. LEGIT addresses this comparison gap by recording quality and cost per solved task alongside the evaluation budget\. Its experiments examine how harness rankings change across budgets and how costs differ when task success is comparable\. These measurements supply the credential fields used in the proposed marketplace score\.
### 2\.5Performance Reports and Verifiable Credentials
Model Cards describe intended uses, evaluation procedures, and performance under different conditions\[[51](https://arxiv.org/html/2609.21325#bib.bib51)\]\. Datasheets document dataset composition, collection, and recommended uses\[[52](https://arxiv.org/html/2609.21325#bib.bib52)\]\. FactSheets extend structured disclosure to AI services, including performance and provenance\[[53](https://arxiv.org/html/2609.21325#bib.bib53)\]\. These reporting approaches help users interpret the evidence behind a system’s reported capabilities\.
Identity standards provide a way to associate claims with their subject and issuer\. The W3C Decentralized Identifiers standard specifies identifiers and methods for proving control over them\[[54](https://arxiv.org/html/2609.21325#bib.bib54)\]\. The Verifiable Credentials data model defines how issuers express claims that holders can present to verifiers\[[55](https://arxiv.org/html/2609.21325#bib.bib55)\]\. Rodriguez Garzon et al\. apply these mechanisms to AI agents and demonstrate exchanges of credentials bound to agent identities\[[56](https://arxiv.org/html/2609.21325#bib.bib56)\]\.
General reporting formats and identity standards leave the choice of task\-performance fields to their application\. LEGIT specifies domain\-specific quality, cost, and budget fields linked to an agent configuration and evaluation evidence\. An issuer signs the record, which automated verifiers can inspect\. Visual profiles and cards are example presentations of these fields\. The contribution is the scoped measurement record and its connection to reputation and proposed allocation\.
### 2\.6Reputation and Sybil Identities
The Beta Reputation System accumulates positive and negative feedback to estimate reputation\[[57](https://arxiv.org/html/2609.21325#bib.bib57)\]\. Jøsang and Quattrociocchi develop features for sparse ratings, initial beliefs, and changing behavior\[[58](https://arxiv.org/html/2609.21325#bib.bib58)\]\. EigenTrust combines peers’ local assessments into reputation values for selecting file providers\[[59](https://arxiv.org/html/2609.21325#bib.bib59)\]\. These approaches connect past interactions to later decisions about service providers\.
Multiple identities controlled by one attacker complicate this connection\. Douceur shows how one participant can undermine a system by presenting multiple identities\[[60](https://arxiv.org/html/2609.21325#bib.bib60)\]\. Friedman and Resnick analyze the consequences of cheap replacement identities and the trade\-offs introduced by entry fees\[[61](https://arxiv.org/html/2609.21325#bib.bib61)\]\. Identity creation and reputation accumulation must therefore be considered together\.
These general models leave the funding requirements of LEGIT’s reputation rule unspecified\. LEGIT combines a discounted Beta update with a deposit for each new identity\. Its analysis separates locked capital from spent interaction fees and computes the resources needed to reach a target reputation\. The resulting estimates let operators compare deposit, fee, and reputation settings under the modeled feedback conditions\.
### 2\.7Agent Marketplaces and Procurement
Scoring auctions give buyers a way to consider quality alongside price\. Che studies procurement in which firms bid on both dimensions\[[62](https://arxiv.org/html/2609.21325#bib.bib62)\]\. Asker and Cantillon analyze scoring auctions with multiple private attributes\[[63](https://arxiv.org/html/2609.21325#bib.bib63)\]\. Papakonstantinou and Bogetoft use payments after observing delivered quality to create incentives for truthful reporting in their procurement model\[[64](https://arxiv.org/html/2609.21325#bib.bib64)\]\. These mechanisms require a defined meaning for the quality attributes used in allocation and payment\.
Agent Exchange proposes auction infrastructure that connects capability representation, performance tracking, and agent coordination\[[5](https://arxiv.org/html/2609.21325#bib.bib5)\]\. Diagon provides configurable experiments on allocation, contracting, and enforcement in agent labor markets\[[6](https://arxiv.org/html/2609.21325#bib.bib6)\]\. Agent Bazaar studies market stability and deception by sellers controlling multiple identities\[[7](https://arxiv.org/html/2609.21325#bib.bib7)\]\. AgentSLA supplies a quality model and a language for specifying service agreements for agents\[[65](https://arxiv.org/html/2609.21325#bib.bib65)\]\.
Mittal’s Trust Layer is particularly close to LEGIT\. It combines capability descriptors, screening, and reputation to address unreliable capability advertisements\[[66](https://arxiv.org/html/2609.21325#bib.bib66)\]\. Its analysis uses equilibrium arguments and illustrative simulations with stipulated provider reliability\. LEGIT adds controlled measurements of model, harness, and budget effects to motivate the contents of a quality and cost credential\.
Marketplace mechanisms and service agreements need comparable evidence for the attributes on which they act\. LEGIT addresses this interface through a credential with explicit configuration, domain, and budget scope\. Its proposed allocation rule reads certified quality and cost together with reputation and bid price\. This connects measured agent performance to eligibility and allocation decisions\.
## 3Problem Formulation
### 3\.1System Model
A marketplaceℳ\\mathcal\{M\}has agent vendors𝒱=\{v1,…,vm\}\\mathcal\{V\}=\\\{v\_\{1\},\\ldots,v\_\{m\}\\\}who register agents\. Task requestersℛ=\{r1,…,rn\}\\mathcal\{R\}=\\\{r\_\{1\},\\ldots,r\_\{n\}\\\}post task specifications\. Certification authorities𝒞=\{c1,…,cp\}\\mathcal\{C\}=\\\{c\_\{1\},\\ldots,c\_\{p\}\\\}issue and verify credentials\. Let𝒟\\mathcal\{D\}denote the task domains, such as code generation, customer service, or research synthesis\. Letk=\|𝒟\|k=\|\\mathcal\{D\}\|denote their number\.
###### Definition 1\(Agent\)\.
An agenta=\(M,S,T,P\)a=\(M,S,T,P\)contains a collection of model dependenciesMM, an agent harnessSS, a tool suiteTT, and a parameter configurationPP\. The collectionMMmay contain one model or several models used through routing or delegation\. The harness contains the system prompt, planning and orchestration logic, tool\-call control, memory architecture, and model\-selection policy\. The tool suite identifies the available tools and delegated services\. The configuration sets temperature, sampling strategy, and context management\. A single\-model agent is the special case with one model dependency\.
###### Definition 2\(Task Specification\)\.
A task specificationτ=\(δ,B,Qmin,D\)\\tau=\(\\delta,B,Q\_\{\\min\},D\)contains a task descriptionδ\\deltaand a maximum token budgetB∈ℕB\\in\\mathbb\{N\}\. It also sets a minimum quality thresholdQmin∈\[0,1\]Q\_\{\\min\}\\in\[0,1\]for the task’s domain and a domain labelD∈𝒟D\\in\\mathcal\{D\}\.
###### Definition 3\(Agent Credential\)\.
An agent credential isΓa=\(ξa,𝐪a,𝐞a,πa,ta,σa\)\\Gamma\_\{a\}=\(\\xi\_\{a\},\\mathbf\{q\}\_\{a\},\\mathbf\{e\}\_\{a\},\\pi\_\{a\},t\_\{a\},\\sigma\_\{a\}\)\. The scope recordξa\\xi\_\{a\}identifies the agent and its configurationa=\(M,S,T,P\)a=\(M,S,T,P\), including the disclosed model versions and routing or delegation policy\. It distinguishes evaluator\-observed configuration details from vendor declarations and unknown dependencies\. It records the certified domains𝒟cert\\mathcal\{D\}\_\{\\mathrm\{cert\}\}, each domain’s task set or sampling record, scoring rule, and token budgetBDB\_\{D\}\. It also records the currency, usage prices, evaluation date, and issuer\. The vectors𝐪a\\mathbf\{q\}\_\{a\}and𝐞a\\mathbf\{e\}\_\{a\}have one entry per certified domain\. Their entries give mean quality and cost per solved task\. An omitted domain is unevaluated\. A zero quality score means that the domain was evaluated but no task earned credit\. Its cost per solved task is infinite\. The evidence recordπa\\pi\_\{a\}links the credential to recorded outputs, scores, and usage\. The timestamptat\_\{a\}sets expiry\. The issuer signatureσa\\sigma\_\{a\}covers the scope, scores, evidence record, and expiry\. Compact architecture labels useqaq\_\{a\}andeae\_\{a\}for the vectors andπ\\piforπa\\pi\_\{a\}\.
Layer 2 maintains the reputation scoreρa∈\[0,1\]\\rho\_\{a\}\\in\[0,1\], as defined in Section[4\.2](https://arxiv.org/html/2609.21325#S4.SS2)\. The score is bound to the same agent identity\. Layer 3 uses it alongsideΓa\\Gamma\_\{a\}\. Reputation changes with every task outcome, so it is excluded from the static Layer 1 credential signed at certification time\.
The credential is a signed data record\. A content hash identifies the record or linked evidence, while the issuer signature authenticates the signed claim under the verifier’s trusted issuer policy\. Neither a hash nor a signature proves that an undisclosed model dependency matches a vendor declaration\. For opaque services, the scope identifies the tested service, its declared configuration, and the available observations\. The resulting assurance is limited to that evidence\.
### 3\.2Evaluation Questions
We ask whether selected model and harness configurations differ in task success under the same token budget\. Comparisons share tasks to control task difficulty\. For the omnibus model and harness checks, the null assumes that the tested labels are exchangeable within each task\. For an exact paired comparison of binary outcomes, the null assigns equal probability to either configuration solving a task that only one solves\. The tests address quality differences within the evaluated configuration set\.
We also ask how cost per solved task and rankings vary with configuration, domain, and budget\. Paired\-bootstrap intervals summarize uncertainty in reported quality and cost ratios\. A ratio of 1 denotes equal values of the compared metric\. Saved\-run threshold curves describe budget sensitivity, with a separate direct\-run check\. These analyses motivate the fields included in a credential\. They do not statistically test the necessity of the protocol or its effect on marketplace outcomes\.
### 3\.3Evaluation Scope
The experiments cover three models from two vendors, selected harnesses and retrieval settings, and task subsets from GAIA and MATH\-500\. Each evaluated configuration uses one model\. Harness variants combine design choices, and the extended grid includes runs from a later round with a different invocation\-cache setting\. These comparisons do not isolate every component or establish general results for dynamic model mixtures, delegated agents, or other domains\. Reputation tracks later task outcomes subject to their reporting and verification assumptions\.
The Sybil analysis calculates the deposits and fees needed to accumulate reputation under the attack assumptions in Section[5\.4](https://arxiv.org/html/2609.21325#S5.SS4)\. End\-to\-end allocation benefits, strategic bidding behavior, and long\-term credential drift are not evaluated\. A future allocation study could compare credential\-based selection with model\-only, quality\-only, and price\-only baselines on held\-out tasks, measuring delivered quality, cost, and selection regret\.
Figure 2:The LEGIT architecture connects certification, reputation records, and marketplace allocation\. Credentials carry measured quality and cost into reputation tracking, task allocation, and payment\.
## 4The LEGIT Protocol
Figure[2](https://arxiv.org/html/2609.21325#S3.F2)links capability measurement, reputation records, and task allocation\.
### 4\.1The Certification Layer
##### Capability evaluation\.
Layer 1 issues capability vectors by evaluating agents on benchmark tasks under fixed token budgets\.
##### Certification scope\.
Each certification authority defines a catalog of domains𝒟\\mathcal\{D\}with associated task distributions and quality metrics\. For example, a code generation domain might count a task as solved when the first attempt clears its unit tests\. A customer service domain might use task completion rate as inτ\\tau\-bench\[[11](https://arxiv.org/html/2609.21325#bib.bib11)\]\. A research synthesis domain might use factual accuracy scored by human evaluators\. The quality functionQ\(a,τ,B\)∈\[0,1\]Q\(a,\\tau,B\)\\in\[0,1\]is thus domain specific\. The certification authority fixes its definition and publishes it alongside the domain specification\. Both agents and requesters then know what is being measured\. An agent selects which domains to certify in based on its intended market\. It receives scores only for those domains\. The certification ecosystem can therefore grow as new task categories emerge without changes to the protocol itself\.
##### Evaluation budget\.
Each task is evaluated against a token allowanceBDB\_\{D\}for its domain\. Let𝒟cert\\mathcal\{D\}\_\{\\mathrm\{cert\}\}contain the domains selected for certification\. For each domainDD, let𝒯D\\mathcal\{T\}\_\{D\}be its evaluation task set\. LetC\(a,τ\)C\(a,\\tau\)be the monetary cost of executing agentaaon taskτ\\tau\. The quality scoreQ\(a,τ,BD\)Q\(a,\\tau,B\_\{D\}\)lies between 0 and 1\. Over\-budget tasks receive zero credit\.
Following[Wang et al\. \[12\]](https://arxiv.org/html/2609.21325#bib.bib12), cost per solved task divides total cost by total credited quality\. For binary scores, total credited quality is the number of solved tasks\. Costs include failed attempts\.
𝖢𝗈𝖯\(a,𝒯D,BD\)=∑τ∈𝒯DC\(a,τ\)∑τ∈𝒯DQ\(a,τ,BD\)\.\\mathsf\{CoP\}\(a,\\mathcal\{T\}\_\{D\},B\_\{D\}\)=\\frac\{\\sum\_\{\\tau\\in\\mathcal\{T\}\_\{D\}\}C\(a,\\tau\)\}\{\\sum\_\{\\tau\\in\\mathcal\{T\}\_\{D\}\}Q\(a,\\tau,B\_\{D\}\)\}\.\(1\)When no task earns credit, cost per solved task is recorded as infinite\. A lower value means less money spent per unit of credited quality\. The domain credential contains mean quality and cost per solved task\.
qa,D=1\|𝒯D\|∑τ∈𝒯DQ\(a,τ,BD\),ea,D=𝖢𝗈𝖯\(a,𝒯D,BD\),D∈𝒟cert\.q\_\{a,D\}=\\frac\{1\}\{\|\\mathcal\{T\}\_\{D\}\|\}\\sum\_\{\\tau\\in\\mathcal\{T\}\_\{D\}\}Q\(a,\\tau,B\_\{D\}\),\\qquad e\_\{a,D\}=\\mathsf\{CoP\}\(a,\\mathcal\{T\}\_\{D\},B\_\{D\}\),\\qquad D\\in\\mathcal\{D\}\_\{\\mathrm\{cert\}\}\.\(2\)
The certified credential for domainDDis the pair\(qa,D,ea,D\)\(q\_\{a,D\},\\,e\_\{a,D\}\)\. It captures what the agent achieves and what it costs\. Neither metric alone suffices\. Agents with identical𝖢𝗈𝖯\\mathsf\{CoP\}may differ substantially in absolute quality\. Agents with identical quality may differ by an order of magnitude in cost\. Each task requester can choose a trade\-off between quality and cost through the Layer 3 scoring rule\. Task requesters can filter agents by quality thresholds and𝖢𝗈𝖯\\mathsf\{CoP\}ceilings\. Certification authorities can deny credentials to agents whose efficiency falls below minimum standards\. This two\-dimensional credential serves as the agent’s profile in downstream layers\. Layer 2 associates reported deployment outcomes with the same identity\. Layer 3 uses it to filter eligible bidders and weight auction scores\.
##### Credential lifecycle\.
A credential describes an evaluation at a recorded time\. Its expiry limits acceptance of that record but does not guarantee unchanged behavior until expiry\. The issuer or verifier can impose a maximum acceptable age\. A change to the declared model dependencies, routing policy, harness, tools, or parameters changes the certified configuration and calls for a new evaluation before the revised agent is represented as covered by that credential\. An unchanged routing policy may select different declared models across tasks, provided this behavior lies within the evaluated scope\. Undisclosed provider changes remain an assurance limitation for opaque services\.
We propose retaining immutable historical records and linking replacement credentials to the records they supersede\. A revoked or suspended record should be rejected for current eligibility even if its signature remains valid\. Comparisons over time should retain task\-set or sampling provenance, scoring rules, budgets, and price snapshots\. Changes in quality, cost, and uncertainty can then be reported with any change in evaluation conditions\. The implementation supports signed manifests, issuance and expiry timestamps, credential status, and policy\-based verification\. Automatic detection of configuration drift, links between superseding records, and longitudinal performance analysis are proposed lifecycle extensions, not evaluated features\.
### 4\.2The Reputation and Portfolio Layer
##### Task records\.
Layer 2 associates records of completed tasks and a reputation score with the same agent identity\.
##### Outcome provenance\.
The reputation model updates its score from records of past task outcomes\. The quality signal is an input supplied by the marketplace’s designated reporting or evaluation process, which may use requester feedback, an automated task check, or an evaluator’s assessment\. The update equation does not verify the signal\. A deployment should record the task, agent configuration, reporter, scoring rule, time, and supporting evidence, and distinguish independently checked outcomes from unverified feedback\. The analysis below assumes the stated feedback process without implementing a general dispute\-resolution or feedback\-authentication mechanism\.
##### Agent identity\.
Each agent has an identifier that links its capability credential, task records, and reputation score\.
For an agent that changes configuration, task records should retain the configuration used at execution time\. A deployment should expose that history and state whether reputation is carried across versions, so buyers can distinguish earlier performance from outcomes under the current configuration\. This version\-aware policy is a proposed extension of the analyzed identity\-level update\.
##### Bayesian reputation model\.
For each quality dimensiondd, we maintain a Beta distribution posterior over the agent’s true quality parameterθd\\theta\_\{d\}, so that
p\(θd\|data\)=𝖡𝖾𝗍𝖺\(αd,βd\)p\(\\theta\_\{d\}\|\\text\{data\}\)=\\mathsf\{Beta\}\(\\alpha\_\{d\},\\beta\_\{d\}\)\(3\)
The positive Beta parametersαd\\alpha\_\{d\}andβd\\beta\_\{d\}represent accumulated evidence of success and failure for dimensiondd\. In the analyzed model, a new identity starts withαd\(0\)=βd\(0\)=1\\alpha\_\{d\}^\{\(0\)\}=\\beta\_\{d\}^\{\(0\)\}=1, giving a uniform prior\. The indexttcounts completed task outcomes, starting at 0 before any feedback\. Following Jøsang and Quattrociocchi\[[58](https://arxiv.org/html/2609.21325#bib.bib58)\], we incorporate a longevity factorη∈\(0,1\]\\eta\\in\(0,1\]that geometrically discounts older observations, giving
αd\(t\)=η⋅αd\(t−1\)\+sd,βd\(t\)=η⋅βd\(t−1\)\+\(1−sd\)\\alpha\_\{d\}^\{\(t\)\}=\\eta\\cdot\\alpha\_\{d\}^\{\(t\-1\)\}\+s\_\{d\},\\quad\\beta\_\{d\}^\{\(t\)\}=\\eta\\cdot\\beta\_\{d\}^\{\(t\-1\)\}\+\(1\-s\_\{d\}\)\(4\)
wheresd∈\[0,1\]s\_\{d\}\\in\[0,1\]is the quality signal for dimensionddfrom the most recent interaction\.
The reputation score for agentaaon dimensionddis the posterior meanρa,d\(t\)=αd\(t\)/\(αd\(t\)\+βd\(t\)\)\\rho\_\{a,d\}^\{\(t\)\}=\\alpha\_\{d\}^\{\(t\)\}/\(\\alpha\_\{d\}^\{\(t\)\}\+\\beta\_\{d\}^\{\(t\)\}\)\. The parameters belong to agentaa, whose index is omitted from the update equations\. In a task’s marketplace score,ρa\\rho\_\{a\}denotes this reputation for the task’s rated dimension\. The Sybil analysis considers one dimension and omits the agent and dimension indices\.
##### Sybil funding requirements\.
The modeled Sybil attack uses multiple attacker\-controlled identities to accumulate misleading reputation\. These identities may have valid identifiers and keys\. The attack does not require forging another party’s signature or identity\. LetNSybilN\_\{\\mathrm\{Sybil\}\}count the attacker\-controlled identities\. Each requires a depositsmins\_\{\\min\}, so the attacker must lock at leastNSybilsminN\_\{\\mathrm\{Sybil\}\}s\_\{\\min\}in capital\. Section[5\.4](https://arxiv.org/html/2609.21325#S5.SS4)separates locked capital from spent interaction fees and states the limits of this economic barrier\.
### 4\.3The Marketplace Layer
##### Task allocation\.
Layer 3 proposes a use of the measured credentials in task allocation\. Its scoring rule combines quality, cost, reputation, and price\.
##### Bid structure\.
The marketplace filters eligible agents for task specificationτ=\(δ,B,Qmin,D\)\\tau=\(\\delta,B,Q\_\{\\min\},D\)\. It uses Layer 1 credentials\(𝐪a,𝐞a\)\(\\mathbf\{q\}\_\{a\},\\mathbf\{e\}\_\{a\}\)and Layer 2 reputationρa\\rho\_\{a\}\. An eligible agentaabids a quality commitment, a price, and a credential record\. The quality commitmentqacommit∈\[Qmin,qa,D\]q\_\{a\}^\{\\text\{commit\}\}\\in\[Q\_\{\\min\},\\,q\_\{a,D\}\]is the quality level the agent pledges to deliver\. It is bounded below by the task’s minimum threshold and above by the agent’s certified quality for domainDD\. The pricepap\_\{a\}is the monetary compensation the agent requests\. The credential recordσacert\\sigma\_\{a\}^\{\\text\{cert\}\}contains the signed Layer 1 credential and its linked Layer 2 reputation record\. The operator checks the issuer signature, configuration and evaluation scope, expiry, applicable age policy, and current status\. A revoked or suspended credential is ineligible\. The marketplace reads certified efficiencyea,De\_\{a,D\}directly from the credential\. Agents do not bid on it\. The bid tuple is
ba=\(qacommit,pa,σacert\)b\_\{a\}=\(q\_\{a\}^\{\\text\{commit\}\},p\_\{a\},\\sigma\_\{a\}^\{\\text\{cert\}\}\)\(5\)
##### Scoring rule\.
Following Che's multi\-dimensional scoring framework\[[62](https://arxiv.org/html/2609.21325#bib.bib62)\], we define the proposed score
𝖲𝖼𝗈𝗋𝖾\(ba\)=wq⋅qacommit\+we⋅\(1−ea,Demax\)\+wρ⋅ρa−wp⋅papmax\\mathsf\{Score\}\(b\_\{a\}\)=w\_\{q\}\\cdot q\_\{a\}^\{\\text\{commit\}\}\+w\_\{e\}\\cdot\\left\(1\-\\frac\{e\_\{a,D\}\}\{e\_\{\\max\}\}\\right\)\+w\_\{\\rho\}\\cdot\\rho\_\{a\}\-w\_\{p\}\\cdot\\frac\{p\_\{a\}\}\{p\_\{\\max\}\}\(6\)
The limitsemax\>0e\_\{\\max\}\>0andpmax\>0p\_\{\\max\}\>0are the maximum eligible cost per solved task and bid price\. The credential must cover the requested domain and budget, soBD=BB\_\{D\}=B\. An eligible bid has a current credential for domainDD,qa,D≥Qminq\_\{a,D\}\\geq Q\_\{\\min\}, finiteea,D≤emaxe\_\{a,D\}\\leq e\_\{\\max\}, and0≤pa≤pmax0\\leq p\_\{a\}\\leq p\_\{\\max\}\. The weights are nonnegative, sum to 1, and havewp\>0w\_\{p\}\>0\. The subscriptsq,e,ρ,pq,e,\\rho,pidentify the quality, cost, reputation, and price terms\. The highest score wins\. Equal scores are resolved by a fixed ordering of agent identifiers\. If no bid is eligible, the task remains unassigned\.
##### Payment rule\.
The proposed rule uses the second\-highest eligible score as its payment threshold\. LetS\(2\)S\_\{\(2\)\}denote that score and letpap\_\{a\}andSaS\_\{a\}be the winning bid price and score\. The threshold payment ispa\+\(pmax/wp\)\(Sa−S\(2\)\)p\_\{a\}\+\(p\_\{\\max\}/w\_\{p\}\)\(S\_\{a\}\-S\_\{\(2\)\}\)\. The operator assigns the task only if at least two bids are eligible and this payment does not exceedpmaxp\_\{\\max\}\.
## 5Experimental Validation
We compared task quality and cost under controlled model, harness, and budget settings to assess the information carried by the credential\. We also quantified funding requirements for reputation manipulation under the Layer 2 attack model in Section[5\.4](https://arxiv.org/html/2609.21325#S5.SS4)\. These evaluations address measurement scope and a specific economic threat, without validating end\-to\-end marketplace performance\.
### 5\.1Method
##### Evaluation design\.
We evaluated model and harness combinations on shared task sets\. Table[9](https://arxiv.org/html/2609.21325#A2.T9.fig1)in Appendix[B](https://arxiv.org/html/2609.21325#A2)lists the comparisons and token budgets\. The appendix describes the models, harnesses, and task sets\. All configurations ran at temperature 0 with pinned model snapshots\. Repeated runs measured variation\. The earlierτ\\tau\-bench study found that function calling outperformed text\-based ReAct on retail tasks using the same models\[[11](https://arxiv.org/html/2609.21325#bib.bib11)\]\. The experiments compare model choice, harness design, and retrieval under stated token budgets\.
##### Measurements\.
We report task quality, cost per solved task, and Pareto dominance under a fixed token budget\. These measurements supply the Layer 1 credential fields\. Over\-budget tasks receive zero credit\. Cost uses provider\-reported usage and price snapshots for each run\. Appendix[C](https://arxiv.org/html/2609.21325#A3)gives the execution, calibration, and usage\-accounting details\.
##### Scoring\.
The GAIA comparisons use the curated task subset and official exact\-match scoring\[[14](https://arxiv.org/html/2609.21325#bib.bib14)\]\. An LLM judge awards credit to near\-misses only when it confirms equivalent meaning\. We checked its agreement with human labels before its decisions affected results\.
##### Statistics\.
The ratioκ^\\hat\{\\kappa\}divides one configuration's mean quality by the comparison configuration's mean quality on the same tasks\. A ratio above 1 favors the first configuration\. A ratio of 1\.2 means its mean quality is 20% higher\. The symbolκ\\kappadenotes the corresponding ratio of expected quality for the stated task distribution\. We report the original paired\-bootstrap confidence intervals and label the original approximate Wilcoxon probabilities as reported values\. The comparisons with a stated decision criterion requireκ^≥1\.2\\hat\{\\kappa\}\\geq 1\.2with a 95% CI excluding 1\.0\. A retrospective check uses within\-task permutations and exact paired tests on binary outcomes\. Appendix[D](https://arxiv.org/html/2609.21325#A4)gives the procedures and correction families\. CI means confidence interval, and SD means standard deviation\. A test probability is denoted bypp\.
For the grid’s quality conclusions, we use the retrospective paired tests with Bonferroni correction and a significance threshold of 0\.05\. The omnibus family contains four tests\. Exact baseline model comparisons form a family of three, and within\-model harness comparisons against baseline form a separate family of 15\. Failure to reject a null does not establish equivalence\. The original ratio intervals and approximate probabilities are retained as reported analyses and are not simultaneous intervals across those families\. Cost ratios describe a separate outcome, with bootstrap uncertainty where available\. Similar solved counts do not constitute a formal equivalence test\. The variance decompositions are descriptive, and no inferential claim about a model\-by\-harness interaction is made\.
### 5\.2Results
#### 5\.2\.1Model capability and unrestricted retrieval
We comparedgpt\-5\.6\-lunaandgpt\-5\.6\-solusing the same baseline harness\. We then added hosted web search to each model without limiting the number of search calls\. Each task had a 16,000\-token budget to allow tokens for retrieval\. Tasks exceeding this budget received zero credit\. We also tested reflective retry in a freshgpt\-5\.6\-solevaluation to compare its quality and cost with baseline\. Reported costs include search fees\.
Table 1:Model and harness performance on 127 GAIA tasks atB=16,000B=16\{,\}000tokens\. Over\-budget tasks receive zero credit\. Success rates use all 127 tasks\.Each ratio divides the first configuration’s metric by the second’s\. The criterion requires a ratio of at least 1\.2 and a 95% CI excluding 1\.0\. The final comparison belongs to the2×22\\times 2model and tool analysis\. It is descriptive, so the criterion does not apply\. The cost comparison uses a bootstrap CI only and has noppvalue\.
Table[1](https://arxiv.org/html/2609.21325#S5.T1.fig1)shows that the baseline solved 54 tasks withgpt\-5\.6\-soland 38 withgpt\-5\.6\-luna, a 42\.1% increase\. Model choice therefore changes the measured quality\. Reflective retry ongpt\-5\.6\-solsolved 53 tasks, compared with 54 for baseline\. Its cost per solved task was 49\.5% higher\.
Web search improvedgpt\-5\.6\-luna’s observed success, increasing solved tasks from 38 to 48\. This is a 26\.3% increase under the 16,000\-token budget\. Search solved 52 tasks withgpt\-5\.6\-sol, compared with 54 for baseline\. The benefit of adding search therefore differed between the evaluated models\.
Search exceeded the token budget on 68 tasks withgpt\-5\.6\-lunaand 70 withgpt\-5\.6\-sol\. These tasks received zero credit, but their token use and cost still counted\. A credential must therefore record the tool configuration alongside the model and evaluation budget\.
004,0004,0008,0008,00012,00012,00016,00016,000000\.20\.20\.40\.40\.60\.60\.80\.8Token budgetB′B^\{\\prime\}Success rate atB′B^\{\\prime\}\(a\) Token thresholdsgpt\-5\.6\-luna, baselinegpt\-5\.6\-luna, web searchgpt\-5\.6\-sol, baselinegpt\-5\.6\-sol, web searchgpt\-5\.6\-sol, reflective retry
baselineReActplan\-exec\.refl\. retrysys\. heavynotes mem\.000\.20\.20\.40\.40\.60\.6Solve rate\(b\) Harnesses at 8,000 tokensclaude\-haiku\-4\-5gpt\-5\.6\-lunagpt\-5\.6\-sol
Figure 3:Quality across budgets and harnesses\. Panel \(a\) applies lower token thresholds to saved 16,000\-token runs\. Panel \(b\) compares harnesses at 8,000 tokens\. Shaded strips show baseline mean plus or minus one SD across five runs\. Model colors match across panels\.Panel \(a\) of Figure[3](https://arxiv.org/html/2609.21325#S5.F3)scores completed runs again at lower token thresholds\. The symbolB′B^\{\\prime\}denotes the threshold applied to saved usage\. A task keeps its final score only if its charged tokens fit within that threshold\. Forgpt\-5\.6\-luna, web search achieves about 0\.45 times baseline quality at 8,000 tokens and 1\.26 times baseline quality at 16,000 tokens\. Its ranking therefore changes with the evaluation budget\.
gpt\-5\.6\-solwithout tools reaches 95% of its full\-budget quality by 4,000 tokens\. Its score is 0\.417 at this threshold, one quarter of the full budget, and stops increasing beyond 6,000 tokens\. With search, it scores 0\.409 at the full 16,000\-token budget\. Search scores continue improving through 16,000 tokens\. These differences show why the credential records the domain’s evaluation budgetBDB\_\{D\}\.
Reflection\.Search benefits vary with the model and token budget, supporting LEGIT credentials that bind quality measurements to the evaluated model, tools, and budget\.
#### 5\.2\.2Bounded retrieval and evaluation budget
Table[2](https://arxiv.org/html/2609.21325#S5.T2.fig1)shows that search limited to two calls improved observed success for both models under the same 16,000\-token budget\. Solved tasks increased from 37 to 57 forgpt\-5\.6\-lunaand from 56 to 63 forgpt\-5\.6\-sol\. These are success ratios of 1\.54 and 1\.13\. The gains came with additional search cost\. Forgpt\-5\.6\-luna, cost per solved task was 5\.7 times its no\-tools baseline\.
The capped runs exceeded the budget on 36 tasks forgpt\-5\.6\-lunaand 37 forgpt\-5\.6\-sol\. These counts are lower than in the separate unrestricted\-search runs\. Each search configuration is compared with its own no\-tools baseline\. The observed quality gains and added costs support recording the tool configuration, quality, cost, and budget together\.
Table 2:Retrieval atB=16,000B=16\{,\}000on 127 GAIA tasks\. Bounded search allows at most two tool calls per task\. Baselines are fresh evaluations for this comparison\.ModelHarnessSolvedMedian tokensOver budgetlunaNo tools372,7020lunaBounded search5711,24036solNo tools562,6120solBounded search6312,76037Search quality contrastκ^\\hat\{\\kappa\}95% CIppluna, bounded over no tools1\.54\[1\.21,2\.10\]\[1\.21,2\.10\]0\.007sol, bounded over no tools1\.13\[0\.92,1\.41\]\[0\.92,1\.41\]Not reported
Bounded search costs5\.7×5\.7\\timesas much per solved task onluna\. Unrestricted search returned about 15K tokens of content per task and had median usage of about 17,000 tokens\. It exceeded budget on 68 tasks and cost17×17\\timesas much as its baseline per attempt\. The separate capped\-search run recorded 36 overruns and median usage of 11,240 tokens\.
Table[3](https://arxiv.org/html/2609.21325#S5.T3.fig1)compares actual 4,000\-token runs with estimates from saved 8,000\-token runs forgpt\-5\.6\-luna\. The estimate counts a task only if the saved run solved it using at most 4,000 charged tokens\. The direct run solved one more task for baseline, two fewer for ReAct, three more for plan\-then\-execute, and one more for reflective retry\.
Table 3:Direct evaluation atB=4,000B=4\{,\}000against truncation fromB=8,000B=8\{,\}000forgpt\-5\.6\-luna\. Values are solved counts out of 127 tasks\.Reflection\.Capped search improves task success at added cost, supporting LEGIT credentials that report quality and cost together with the search limits used during evaluation\.
#### 5\.2\.3Model and harness effects
Table[4](https://arxiv.org/html/2609.21325#S5.T4)compares three models from two vendors with four harnesses on the same 127 tasks at 8,000 tokens\. The baseline solved 19 tasks withclaude\-haiku\-4\-5, 43 withgpt\-5\.6\-luna, and 55 withgpt\-5\.6\-sol\. Panel \(b\) of Figure[3](https://arxiv.org/html/2609.21325#S5.F3)adds an elaborate system prompt and a condensed\-notes memory buffer, extending the comparison to six harnesses\. The tables report quality and cost for each configuration\.
Table 4:Model and harness factorial grid on the shared 127\-task pool atB=8,000B=8\{,\}000tokens, with a descriptive variance decomposition on per\-task scores at the right\. Over\-budget items are scored zero by the same\-budget rule\. The models areclaude\-haiku\-4\-5,gpt\-5\.6\-luna, andgpt\-5\.6\-sol\. Cost per solved task is in US dollars, and thehaikufigure uses that model’s own prices\.SS is the sum of squared deviations assigned to each source\. The degrees of freedom are df\. The statisticFFdivides a source's mean square by the residual mean square\. Paired tests appear in Table[10](https://arxiv.org/html/2609.21325#A4.T10.fig1)\. The effect sizeη2\\eta^\{2\}is its share of total SS\. This statistical symbol is separate from the reputation decay factorη\\eta\.
Table 5:Quality contrasts and the extended six\-harness analysis on GAIA atB=8,000B=8\{,\}000\. Model contrasts use the baseline harness\.Contrastκ^\\hat\{\\kappa\}95% CIsoloverluna1\.28\[1\.09,1\.55\]\[1\.09,1\.55\]lunaoverhaiku2\.26\[1\.64,3\.55\]\[1\.64,3\.55\]soloverhaiku2\.90\[2\.04,4\.69\]\[2\.04,4\.69\]solnotes over baseline0\.84\[0\.70,0\.96\]\[0\.70,0\.96\]Six\-harness analysisFFResultModel101\.4DescriptiveHarness0\.71DescriptiveInteraction0\.54Descriptive
The notes harness solves 46 tasks onsol, against 55 for baseline\.
Baseline success is higher forgpt\-5\.6\-solthangpt\-5\.6\-luna, and higher forgpt\-5\.6\-lunathanclaude\-haiku\-4\-5, as shown in Table[5](https://arxiv.org/html/2609.21325#S5.T5.fig1)\. All three exact paired baseline model comparisons remain significant after correction\. The omnibus model tests have adjustedp=0\.00020p=0\.00020for both the four\-harness and six\-harness grids\. The corresponding harness tests have adjustedp=0\.15100p=0\.15100andp=0\.15280p=0\.15280, and none of the 15 exact within\-model harness comparisons remains significant\. Tables[10](https://arxiv.org/html/2609.21325#A4.T10.fig1)and[11](https://arxiv.org/html/2609.21325#A4.T11.fig1)give the full results\.
The observed harness differences are therefore descriptive\. Ongpt\-5\.6\-sol, ReAct solved 56 tasks and reflective retry solved 60, compared with 55 for baseline\. The notes harness solved 46\. Baseline solved the most tasks among the tested harnesses forgpt\-5\.6\-lunaandclaude\-haiku\-4\-5\. These counts do not establish an overall harness\-quality effect or a general benefit from additional orchestration\. The reported unadjusted interval for the notes contrast in Table[5](https://arxiv.org/html/2609.21325#S5.T5.fig1)should be read alongside its nonsignificant corrected paired test\.
baselineReActplan\-exec\.refl\. retrysys\. heavynotes mem\.002244Relative cost per solved taskclaude\-haiku\-4\-5gpt\-5\.6\-lunagpt\-5\.6\-sol
haikulunasol002244Reflective retry cost over baselineGAIAMATH\-500
Figure 4:The left panel compares each harness’s cost per solved task with the same model’s baseline on GAIA\. Model colors match Figure[3](https://arxiv.org/html/2609.21325#S5.F3), right panel\. The dashed line marks equal cost\. The right panel shows reflective retry’s extra cost on both task domains\. Extra cost occurs everywhere and is largest on the weakest model\.0\.010\.010\.10\.10\.10\.10\.20\.20\.30\.30\.40\.40\.50\.5baselineReActrefl\. retryCost per solved task in $, log scaleSolve ratePareto frontierclaude\-haiku\-4\-5gpt\-5\.6\-lunagpt\-5\.6\-solFigure 5:Quality and cost per solved task for all eighteen configurations atB=8,000B=8\{,\}000\. The dashed line joins the Pareto frontier\. It containslunabaseline,solReAct, andsolreflective retry at very different trade\-offs\. Every other configuration is dominated\. A quality\-only ranking would hide the entire horizontal axis\.Figure[4](https://arxiv.org/html/2609.21325#S5.F4)shows cost differences within each model\. The most expensive harness costs 1\.43 times the least expensive ongpt\-5\.6\-sol, 1\.63 times ongpt\-5\.6\-luna, and 4\.95 times onclaude\-haiku\-4\-5\. Onclaude\-haiku\-4\-5, reflective retry uses about five times the baseline tokens and solves fewer tasks\. The cost spread is largest for the model with the lowest success in these runs\. Figure[5](https://arxiv.org/html/2609.21325#S5.F5)places three of the eighteen configurations on the Pareto frontier, meaning no other tested configuration achieves at least as much quality at no greater cost, with an improvement in either measure\. A quality\-only ranking would hide these cost differences\.
solbaselinesolReActsolplan\-exec\.solrefl\. retrysolsys\. heavysolnotes mem\.lunabaselinelunaReActlunaplan\-exec\.lunarefl\. retrylunasys\. heavylunanotes mem\.haikubaselinehaikuReActhaikuplan\-exec\.haikurefl\. retryhaikusys\. heavyhaikunotes mem\.Figure 6:Task outcomes underlying the grid statistics\. Each row represents a configuration\. Each cell represents a task in the shared pool and is filled when that configuration solves it under the same\-budget rule\. The left side contains easy tasks shared across configurations\. The middle band contains tasks solved by some models but not others\. The right side contains unsolved tasks\. Rows differ far less within a model than across models\.Figure[6](https://arxiv.org/html/2609.21325#S5.F6)shows which tasks each configuration solved\. The middle band contains tasks solved by some models but not others\. Changing the harness within a model changes fewer outcomes\. This task\-level view helps explain why model differences are more visible than harness differences in the aggregate results\.
Reflection\.The grid supports model\-quality differences and documents cost differences among harness configurations\. Corrected harness\-quality tests are inconclusive\. Recording both quality and cost preserves these distinctions without implying that more elaborate harnesses improve task success\.
#### 5\.2\.4Repeatability, scoring, and domain transfer
Table[7](https://arxiv.org/html/2609.21325#S5.T7)reports five baseline runs per model\. The standard deviation in solved counts was 3\.4 tasks forclaude\-haiku\-4\-5, 3\.1 forgpt\-5\.6\-luna, and 1\.1 forgpt\-5\.6\-sol\. Relative to each model’s mean, these are roughly 20%, 8%, and 2%\. Runs varied even at temperature 0, consistent with[Ouyang et al\. \[9\]](https://arxiv.org/html/2609.21325#bib.bib9)\.
Table 6:Baseline repeatability on 127 GAIA tasks atB=8,000B=8\{,\}000and temperature 0, with the invocation cache disabled\.Thelunaruns solved 43, 37, 42, 36, and 38 tasks, with mean 39\.2\.
Table 7:MATH\-500 transfer on 127 tasks at the sameB=8,000B=8\{,\}000budget\. Cost ratios compare reflective retry with the same model’s baseline\.The baseline model contrast betweensolandhaikugivesκ^=1\.08\\hat\{\\kappa\}=1\.08with 95% CI\[1\.03,1\.13\]\[1\.03,1\.13\]\. Cost ratios are the reported ratios in Figure[4](https://arxiv.org/html/2609.21325#S5.F4)\.
Judgments also agree across model families\. Reassessment of 1,612 cached judgments byclaude\-haiku\-4\-5agreed withgpt\-5\.6\-lunaon 99\.1% under identical instructions\. Cohen’s kappa was 0\.94\.
Table[7](https://arxiv.org/html/2609.21325#S5.T7)compares baseline and reflective retry on 127 MATH\-500 tasks converted to the evaluation system’s task format\. Retry increased solved counts from 118 to 120 forclaude\-haiku\-4\-5and from 123 to 125 forgpt\-5\.6\-luna\. Bothgpt\-5\.6\-solconfigurations solved all 127 tasks\. Retry cost 2\.29, 2\.36, and 2\.41 times as much per solved task, respectively\. The cost difference therefore persists on mathematics tasks\.
Reflection\.Run variation, scoring checks, and cost differences on mathematics support LEGIT credentials that retain evaluation evidence and report quality and cost for each domain\.
### 5\.3Measured Credential Profiles
The left panel of Figure[1](https://arxiv.org/html/2609.21325#S1.F1)combines the measured quality, cost, and budget results into agent profiles\. Among these configurations, thegpt\-5\.6\-lunabaseline is most efficient on GAIA and MATH\-500\. Thegpt\-5\.6\-solconfigurations lead on task quality\. The profiles show why a credential needs both quality and cost\.
The right panel of Figure[1](https://arxiv.org/html/2609.21325#S1.F1)is an example display of a structured credential using measuredgpt\-5\.6\-lunabaseline results\. The quality, cost per solved task, and budget values come from the evaluations\. The shortened signature, expiry date, and verification mark illustrate presentation fields\. An actual verification decision requires the signed record, a trusted issuer key, and the applicable scope, age, and status checks\.
The labelhaikudenotesclaude\-haiku\-4\-5\. The labellunadenotesgpt\-5\.6\-luna\. The labelsoldenotesgpt\-5\.6\-sol\. Clockwise from the top, the profile axes show quality on GAIA level 1, GAIA level 2 and higher, and MATH\-500\. The next axes show efficiency on GAIA and MATH\-500, meaning inverse cost per solved task\. The final axis shows budget sensitivity, measured as the share of full\-budget quality retained when saved GAIA runs are scored at half the evaluation budget\. It is a threshold reconstruction, not a direct measurement of an agent adapting to that lower budget\. Each axis is normalized to its best agent\. The card usesqqfor domain quality,𝖢𝗈𝖯\\mathsf\{CoP\}for cost per solved task, andBBfor the token budget\. It displays evidence assurance, expiry, and issuer\-signature information from the underlying record\.
Reflection\.The structured profile retains the measurements and conditions needed for comparisons\. Its visual presentation is optional and introduces no additional assurance\.
### 5\.4Sybil Attack Economics
We assess the economic barrier to an attacker accumulating misleading reputation across multiple controlled identities\. Each Sybil identity starts with a uniform reputation prior andρ0=0\.5\\rho\_\{0\}=0\.5\. The attacker lockssmins\_\{\\min\}per identity and spends a fee of at leastfminf\_\{\\min\}per credited interaction\. Perfect feedback gives the best\-case interaction countnrepn\_\{\\mathrm\{rep\}\}needed to reachρtarget\\rho\_\{\\mathrm\{target\}\}\. The model assumes every identity incurs its own deposit and fees, deposits stay locked throughout accumulation, and fees cannot be recovered through self\-dealing\. It conditions on the specified feedback stream without establishing how an attacker obtains accepted positive reports\. Appendix[A](https://arxiv.org/html/2609.21325#A1)gives the closed\-form count and its agreement with direct iteration\.
Table 8:Capital and fees required to accumulate reputation for a Sybil identity\. The left panel gives the perfect\-feedback interactionsnrepn\_\{\\text\{rep\}\}needed to raise a fresh identity fromρ0=0\.5\\rho\_\{0\}=0\.5toρtarget\\rho\_\{\\text\{target\}\}, from direct iteration of the discounted Beta update\. The right panel gives deposit plus spent fees per identity atη=0\.9\\eta=0\.9in USD\. The sum issmin\+nrepfmins\_\{\\min\}\+n\_\{\\mathrm\{rep\}\}f\_\{\\min\}\. The dollar\-valued column headings specify the locked deposit\.
Table[8](https://arxiv.org/html/2609.21325#S5.T8)reports capital plus fees needed to fund the modeled attack\. Fees range from $0\.60 to $23 per identity\. The deposit is the larger component for some, but not all, displayed settings\.
0\.70\.70\.80\.80\.90\.90\.950\.950\.990\.9910101001001,0001\{,\}000not reached within 5,000interactions in 34\.1% of trialsTarget reputationρtarget\\rho\_\{\\text\{target\}\}Interactionsnrepn\_\{\\text\{rep\}\}perfect feedback,η=0\.80\\eta=0\.80perfect feedback,η=0\.90\\eta=0\.90perfect feedback,η=0\.95\\eta=0\.9590% success,η=0\.90\\eta=0\.9080% success,η=0\.90\\eta=0\.90Figure 7:Interactions needed to reach a reputation target\. Solid curves use perfect feedback\. Noisy curves use independent binary feedback with success probabilities 0\.9 and 0\.8 atη=0\.9\\eta=0\.9\. Each point is the mean among trials reaching the target within 5,000 interactions, using 1,000 trials and seeds 0 through 999\. At 80% success, 341 trials do not reach 0\.99 within the horizon\.Figure[7](https://arxiv.org/html/2609.21325#S5.F7)shows the sensitivity to noisy feedback\. At success probability 0\.9, reaching reputation 0\.99 takes about 166 interactions, against 23 with perfect feedback\. At probability 0\.8, 34\.1% of trials do not reach the target within 5,000 interactions\. The plotted mean of 2,073 uses only the 659 trials that reach it\.
Reflection\.Deposits and nonrecoverable fees impose funding requirements under this model\. Locked capital can be returned and is distinct from spent fees\. The calculation does not prove attack unprofitability, detect common identity ownership, or establish resistance to forged feedback, coordinated manipulation, and other attacks\. It provides a bounded assessment of reputation\-manipulation costs for comparing protocol settings\.
## 6Conclusion
LEGIT connects agent certification, reputation records, and proposed marketplace allocation through signed, structured credentials\. Each credential binds quality and cost to configuration, domain, budget, evaluation evidence, and temporal scope\. Charts and cards are optional views\. Controlled comparisons support model\-quality differences and document within\-model cost differences and budget\-sensitive retrieval rankings\. Corrected harness\-quality tests remain inconclusive\. These results motivate the credential fields\. The Sybil analysis quantifies locked capital and spent fees under explicit feedback and funding assumptions\.
## References
- \[1\]OpenAI\.Introducing the GPT Store\.[https://openai\.com/index/introducing\-the\-gpt\-store/](https://openai.com/index/introducing-the-gpt-store/), 2024a\.
- \[2\]Salesforce\.Salesforce Launches AgentExchange: the Trusted Marketplace for Agentforce, 2025\.URL[https://www\.salesforce\.com/news/press\-releases/2025/03/04/agentexchange\-announcement/](https://www.salesforce.com/news/press-releases/2025/03/04/agentexchange-announcement/)\.
- \[3\]Amazon Web Services\.Introducing AI agents and tools in AWS Marketplace\.[https://aws\.amazon\.com/about\-aws/whats\-new/2025/07/ai\-agents\-tools\-aws\-marketplace](https://aws.amazon.com/about-aws/whats-new/2025/07/ai-agents-tools-aws-marketplace), 2025\.
- \[4\]Circle\.Circle powers the agentic economy with new AI infrastructure\.[https://www\.circle\.com/pressroom/circle\-launches\-ai\-infrastructure\-to\-power\-the\-agentic\-economy](https://www.circle.com/pressroom/circle-launches-ai-infrastructure-to-power-the-agentic-economy), 2026\.
- \[5\]Yingxuan Yang, Ying Wen, Jun Wang, and Weinan Zhang\.Agent exchange: Shaping the future of AI agent economics\.*arXiv preprint arXiv:2507\.03904*, 2025a\.URL[https://arxiv\.org/abs/2507\.03904](https://arxiv.org/abs/2507.03904)\.
- \[6\]Xuan Liu, Haoyang Shang, and Haojian Jin\.Diagon: A programmable testbed for AI\-agent cognitive labor markets\.*arXiv preprint arXiv:2604\.06688*, 2026\.URL[https://arxiv\.org/abs/2604\.06688](https://arxiv.org/abs/2604.06688)\.
- \[7\]Seth Karten, Cameron Crow, and Chi Jin\.Agent bazaar: Enabling economic alignment in multi\-agent marketplaces\.*arXiv preprint arXiv:2605\.17698*, 2026\.URL[https://arxiv\.org/abs/2605\.17698](https://arxiv.org/abs/2605.17698)\.
- \[8\]Subhadeep Pal, Fiona Y\. Wang, and Markus J\. Buehler\.SwarmWorld: Stigmergic technological evolution in societies of language\-model agents\.*arXiv preprint arXiv:2608\.26081*, 2026\.URL[https://arxiv\.org/abs/2608\.26081](https://arxiv.org/abs/2608.26081)\.
- \[9\]Shuyin Ouyang, Jie M\. Zhang, Mark Harman, and Meng Wang\.An empirical study of the non\-determinism of ChatGPT in code generation\.*arXiv preprint arXiv:2308\.02828*, 2023\.URL[https://arxiv\.org/abs/2308\.02828](https://arxiv.org/abs/2308.02828)\.
- \[10\]Sayash Kapoor, Benedikt Stroebl, Zachary S\. Siegel, Nitya Nadgir, and Arvind Narayanan\.AI agents that matter\.*arXiv preprint arXiv:2407\.01502*, 2024\.URL[https://arxiv\.org/abs/2407\.01502](https://arxiv.org/abs/2407.01502)\.
- \[11\]Shunyu Yao et al\.τ\\tau\-bench: A benchmark for tool\-agent\-user interaction in real\-world domains\.*arXiv preprint arXiv:2406\.12045*, 2024\.URL[https://arxiv\.org/abs/2406\.12045](https://arxiv.org/abs/2406.12045)\.
- \[12\]Ningning Wang et al\.Efficient Agents: Building Effective Agents While Reducing Cost\.*arXiv preprint arXiv:2508\.02694*, 2025\.URL[https://arxiv\.org/abs/2508\.02694](https://arxiv.org/abs/2508.02694)\.
- \[13\]Yuan\-An Xiao, Pengfei Gao, Chao Peng, and Yingfei Xiong\.Reducing Cost of LLM Agents with Trajectory Reduction\.*Proceedings of the ACM on Software Engineering*, 2026\.doi:10\.1145/3797084\.URL[https://arxiv\.org/abs/2509\.23586](https://arxiv.org/abs/2509.23586)\.
- \[14\]Grégoire Mialon et al\.GAIA: A benchmark for general AI assistants\.In*ICLR*, 2024\.URL[https://arxiv\.org/abs/2311\.12983](https://arxiv.org/abs/2311.12983)\.
- \[15\]Xiao Liu et al\.AgentBench: Evaluating LLMs as agents\.In*ICLR*, 2024\.URL[https://arxiv\.org/abs/2308\.03688](https://arxiv.org/abs/2308.03688)\.
- \[16\]Carlos E\. Jimenez et al\.SWE\-bench: Can language models resolve real\-world GitHub issues?In*ICLR*, 2024\.URL[https://arxiv\.org/abs/2310\.06770](https://arxiv.org/abs/2310.06770)\.
- \[17\]OpenAI\.Introducing SWE\-bench Verified\.[https://openai\.com/index/introducing\-swe\-bench\-verified/](https://openai.com/index/introducing-swe-bench-verified/), 2024b\.
- \[18\]Shuyan Zhou et al\.WebArena: A realistic web environment for building autonomous agents\.In*ICLR*, 2024\.URL[https://arxiv\.org/abs/2307\.13854](https://arxiv.org/abs/2307.13854)\.
- \[19\]Jing Yu Koh et al\.VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks\.*arXiv preprint arXiv:2401\.13649*, 2024\.URL[https://arxiv\.org/abs/2401\.13649](https://arxiv.org/abs/2401.13649)\.
- \[20\]Alexandre Drouin et al\.WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?*arXiv preprint arXiv:2403\.07718*, 2024\.URL[https://arxiv\.org/abs/2403\.07718](https://arxiv.org/abs/2403.07718)\.
- \[21\]Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, and Jonathan Berant\.AssistantBench: Can Web Agents Solve Realistic and Time\-Consuming Tasks?In*Proceedings of EMNLP*, 2024\.doi:10\.18653/v1/2024\.emnlp\-main\.505\.URL[https://aclanthology\.org/2024\.emnlp\-main\.505/](https://aclanthology.org/2024.emnlp-main.505/)\.
- \[22\]Thibault Le Sellier De Chezelles et al\.The BrowserGym Ecosystem for Web Agent Research\.*arXiv preprint arXiv:2412\.05467*, 2024\.URL[https://arxiv\.org/abs/2412\.05467](https://arxiv.org/abs/2412.05467)\.
- \[23\]Tianbao Xie et al\.OSWorld: Benchmarking Multimodal Agents for Open\-Ended Tasks in Real Computer Environments\.*arXiv preprint arXiv:2404\.07972*, 2024\.URL[https://arxiv\.org/abs/2404\.07972](https://arxiv.org/abs/2404.07972)\.
- \[24\]Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik Narasimhan\.τ2\\tau^\{2\}\-Bench: Evaluating Conversational Agents in a Dual\-Control Environment\.*arXiv preprint arXiv:2506\.07982*, 2025\.URL[https://arxiv\.org/abs/2506\.07982](https://arxiv.org/abs/2506.07982)\.
- \[25\]Percy Liang et al\.Holistic Evaluation of Language Models\.*arXiv preprint arXiv:2211\.09110*, 2022\.URL[https://arxiv\.org/abs/2211\.09110](https://arxiv.org/abs/2211.09110)\.
- \[26\]Sayash Kapoor et al\.Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation\.*arXiv preprint arXiv:2510\.11977*, 2025\.URL[https://arxiv\.org/abs/2510\.11977](https://arxiv.org/abs/2510.11977)\.
- \[27\]Shivalika Singh, Yiyang Nan, Alex Wang, Daniel D’Souza, Sayash Kapoor, Ahmet Üstün, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah A\. Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker\.The Leaderboard Illusion\.*arXiv preprint arXiv:2504\.20879*, 2025\.URL[https://arxiv\.org/abs/2504\.20879](https://arxiv.org/abs/2504.20879)\.
- \[28\]Yonatan Oren, Nicole Meister, Niladri Chatterji, Faisal Ladhak, and Tatsunori B\. Hashimoto\.Proving Test Set Contamination in Black Box Language Models\.*arXiv preprint arXiv:2310\.17623*, 2023\.URL[https://arxiv\.org/abs/2310\.17623](https://arxiv.org/abs/2310.17623)\.
- \[29\]Alon Jacovi, Avi Caciularu, Omer Goldman, and Yoav Goldberg\.Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks\.In*Proceedings of EMNLP*, pages 5075–5084, 2023\.doi:10\.18653/v1/2023\.emnlp\-main\.308\.URL[https://aclanthology\.org/2023\.emnlp\-main\.308/](https://aclanthology.org/2023.emnlp-main.308/)\.
- \[30\]Colin White et al\.LiveBench: A Challenging, Contamination\-Limited LLM Benchmark\.*arXiv preprint arXiv:2406\.19314*, 2024\.URL[https://arxiv\.org/abs/2406\.19314](https://arxiv.org/abs/2406.19314)\.
- \[31\]Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, and Arvind Narayanan\.Towards a Science of AI Agent Reliability\.*arXiv preprint arXiv:2602\.16666*, 2026\.URL[https://arxiv\.org/abs/2602\.16666](https://arxiv.org/abs/2602.16666)\.
- \[32\]Thomas Kwa et al\.Measuring AI Ability to Complete Long Software Tasks\.*arXiv preprint arXiv:2503\.14499*, 2025\.URL[https://arxiv\.org/abs/2503\.14499](https://arxiv.org/abs/2503.14499)\.
- \[33\]Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao\.ReAct: Synergizing Reasoning and Acting in Language Models\.*arXiv preprint arXiv:2210\.03629*, 2022\.URL[https://arxiv\.org/abs/2210\.03629](https://arxiv.org/abs/2210.03629)\.
- \[34\]Noah Shinn et al\.Reflexion: Language Agents with Verbal Reinforcement Learning\.*arXiv preprint arXiv:2303\.11366*, 2023\.URL[https://arxiv\.org/abs/2303\.11366](https://arxiv.org/abs/2303.11366)\.
- \[35\]Andy Zhou, Kai Yan, Michal Shlapentokh\-Rothman, Haohan Wang, and Yu\-Xiong Wang\.Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models\.*arXiv preprint arXiv:2310\.04406*, 2023\.URL[https://arxiv\.org/abs/2310\.04406](https://arxiv.org/abs/2310.04406)\.
- \[36\]Timo Schick et al\.Toolformer: Language Models Can Teach Themselves to Use Tools\.*arXiv preprint arXiv:2302\.04761*, 2023\.URL[https://arxiv\.org/abs/2302\.04761](https://arxiv.org/abs/2302.04761)\.
- \[37\]Yujia Qin et al\.ToolLLM: Facilitating Large Language Models to Master 16000\+ Real\-world APIs\.*arXiv preprint arXiv:2307\.16789*, 2023\.URL[https://arxiv\.org/abs/2307\.16789](https://arxiv.org/abs/2307.16789)\.
- \[38\]Shishir G\. Patil, Tianjun Zhang, Xin Wang, and Joseph E\. Gonzalez\.Gorilla: Large Language Model Connected with Massive APIs\.*arXiv preprint arXiv:2305\.15334*, 2023\.URL[https://arxiv\.org/abs/2305\.15334](https://arxiv.org/abs/2305.15334)\.
- \[39\]John Yang et al\.SWE\-agent: Agent\-Computer Interfaces Enable Automated Software Engineering\.*arXiv preprint arXiv:2405\.15793*, 2024\.URL[https://arxiv\.org/abs/2405\.15793](https://arxiv.org/abs/2405.15793)\.
- \[40\]Qingyun Wu et al\.AutoGen: Enabling Next\-Gen LLM Applications via Multi\-Agent Conversation\.*arXiv preprint arXiv:2308\.08155*, 2023\.URL[https://arxiv\.org/abs/2308\.08155](https://arxiv.org/abs/2308.08155)\.
- \[41\]Sirui Hong et al\.MetaGPT: Meta Programming for A Multi\-Agent Collaborative Framework\.*arXiv preprint arXiv:2308\.00352*, 2023\.URL[https://arxiv\.org/abs/2308\.00352](https://arxiv.org/abs/2308.00352)\.
- \[42\]Chen Qian et al\.ChatDev: Communicative Agents for Software Development\.*arXiv preprint arXiv:2307\.07924*, 2023\.URL[https://arxiv\.org/abs/2307\.07924](https://arxiv.org/abs/2307.07924)\.
- \[43\]Omar Khattab et al\.DSPy: Compiling Declarative Language Model Calls into Self\-Improving Pipelines\.*arXiv preprint arXiv:2310\.03714*, 2023\.URL[https://arxiv\.org/abs/2310\.03714](https://arxiv.org/abs/2310.03714)\.
- \[44\]Jiayi Zhang et al\.AFlow: Automating Agentic Workflow Generation\.*arXiv preprint arXiv:2410\.10762*, 2024\.URL[https://arxiv\.org/abs/2410\.10762](https://arxiv.org/abs/2410.10762)\.
- \[45\]Shengran Hu, Cong Lu, and Jeff Clune\.Automated Design of Agentic Systems\.*arXiv preprint arXiv:2408\.08435*, 2024\.URL[https://arxiv\.org/abs/2408\.08435](https://arxiv.org/abs/2408.08435)\.
- \[46\]Lingjiao Chen, Matei Zaharia, and James Zou\.FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance\.*arXiv preprint arXiv:2305\.05176*, 2023\.URL[https://arxiv\.org/abs/2305\.05176](https://arxiv.org/abs/2305.05176)\.
- \[47\]Isaac Ong et al\.RouteLLM: Learning to Route LLMs with Preference Data\.*arXiv preprint arXiv:2406\.18665*, 2024\.URL[https://arxiv\.org/abs/2406\.18665](https://arxiv.org/abs/2406.18665)\.
- \[48\]Sehoon Kim et al\.An LLM Compiler for Parallel Function Calling\.*arXiv preprint arXiv:2312\.04511*, 2023\.URL[https://arxiv\.org/abs/2312\.04511](https://arxiv.org/abs/2312.04511)\.
- \[49\]Bruce Yang, Xinfeng He, Huan Gao, Yifan Cao, Xiaofan Li, and David Hsu\.CodeAgents: A token\-efficient framework for codified multi\-agent reasoning in LLMs\.*arXiv preprint arXiv:2507\.03254*, 2025b\.URL[https://arxiv\.org/abs/2507\.03254](https://arxiv.org/abs/2507.03254)\.
- \[50\]Alexander Conway, Debadeepta Dey, Stefan Hackmann, Matthew Hausknecht, Michael Schmidt, Mark Steadman, and Nick Volynets\.syftr: Pareto\-Optimal Generative AI\.*arXiv preprint arXiv:2505\.20266*, 2025\.URL[https://arxiv\.org/abs/2505\.20266](https://arxiv.org/abs/2505.20266)\.
- \[51\]Margaret Mitchell et al\.Model Cards for Model Reporting\.In*Proceedings of the Conference on Fairness, Accountability, and Transparency*, 2019\.doi:10\.1145/3287560\.3287596\.URL[https://arxiv\.org/abs/1810\.03993](https://arxiv.org/abs/1810.03993)\.
- \[52\]Timnit Gebru et al\.Datasheets for Datasets\.*Communications of the ACM*, 2021\.doi:10\.1145/3458723\.URL[https://arxiv\.org/abs/1803\.09010](https://arxiv.org/abs/1803.09010)\.
- \[53\]Matthew Arnold et al\.FactSheets: Increasing Trust in AI Services through Supplier’s Declarations of Conformity\.*IBM Journal of Research and Development*, 2019\.doi:10\.1147/jrd\.2019\.2942288\.URL[https://arxiv\.org/abs/1808\.07261](https://arxiv.org/abs/1808.07261)\.
- \[54\]World Wide Web Consortium\.Decentralized Identifiers \(DIDs\) v1\.0\.W3c recommendation, W3C, 2022\.URL[https://www\.w3\.org/TR/2022/REC\-did\-core\-20220719/](https://www.w3.org/TR/2022/REC-did-core-20220719/)\.
- \[55\]World Wide Web Consortium\.Verifiable Credentials Data Model v2\.0\.W3c recommendation, W3C, 2025\.URL[https://www\.w3\.org/TR/2025/REC\-vc\-data\-model\-2\.0\-20250515/](https://www.w3.org/TR/2025/REC-vc-data-model-2.0-20250515/)\.
- \[56\]Sandro Rodriguez Garzon, Awid Vaziry, Enis Mert Kuzu, Dennis Enrique Gehrmann, Buse Varkan, Alexander Gaballa, and Axel Küpper\.AI Agents with Decentralized Identifiers and Verifiable Credentials\.*arXiv preprint arXiv:2511\.02841*, 2025\.URL[https://arxiv\.org/abs/2511\.02841](https://arxiv.org/abs/2511.02841)\.
- \[57\]Audun Jøsang and Roslan Ismail\.The Beta Reputation System\.In*BLED 2002 Proceedings*, 2002\.URL[https://aisel\.aisnet\.org/bled2002/41/](https://aisel.aisnet.org/bled2002/41/)\.
- \[58\]Audun Jøsang and Walter Quattrociocchi\.Advanced features in Bayesian reputation systems\.In*TrustBus*, 2009\.URL[https://doi\.org/10\.1007/978\-3\-642\-03748\-1\_11](https://doi.org/10.1007/978-3-642-03748-1_11)\.
- \[59\]Sepandar D\. Kamvar, Mario T\. Schlosser, and Hector Garcia\-Molina\.The EigenTrust Algorithm for Reputation Management in P2P Networks\.In*Proceedings of the Twelfth International Conference on World Wide Web*, 2003\.doi:10\.1145/775152\.775242\.URL[https://nlp\.stanford\.edu/pubs/eigentrust\.pdf](https://nlp.stanford.edu/pubs/eigentrust.pdf)\.
- \[60\]John R\. Douceur\.The Sybil Attack\.In*Proceedings of the First International Workshop on Peer\-to\-Peer Systems*, 2002\.doi:10\.1007/3\-540\-45748\-8\_24\.URL[https://www\.microsoft\.com/en\-us/research/wp\-content/uploads/2002/01/IPTPS2002\.pdf](https://www.microsoft.com/en-us/research/wp-content/uploads/2002/01/IPTPS2002.pdf)\.
- \[61\]Eric J\. Friedman and Paul Resnick\.The Social Cost of Cheap Pseudonyms\.*Journal of Economics and Management Strategy*, 10\(2\):173–199, 2001\.doi:10\.1111/j\.1430\-9134\.2001\.00173\.x\.URL[https://doi\.org/10\.1111/j\.1430\-9134\.2001\.00173\.x](https://doi.org/10.1111/j.1430-9134.2001.00173.x)\.
- \[62\]Yeon\-Koo Che\.Design competition through multidimensional auctions\.*RAND Journal of Economics*, 24\(4\):668–680, 1993\.URL[https://doi\.org/10\.2307/2555752](https://doi.org/10.2307/2555752)\.
- \[63\]John Asker and Estelle Cantillon\.Properties of Scoring Auctions\.*RAND Journal of Economics*, 39\(1\):69–85, 2008\.doi:10\.1111/j\.1756\-2171\.2008\.00004\.x\.URL[https://doi\.org/10\.1111/j\.1756\-2171\.2008\.00004\.x](https://doi.org/10.1111/j.1756-2171.2008.00004.x)\.
- \[64\]Athanasios Papakonstantinou and Peter Bogetoft\.Incentives in Multi\-dimensional Auctions under Information Asymmetry for Costs and Qualities\.In*Agent\-Mediated Electronic Commerce: Designing Trading Strategies and Mechanisms for Electronic Markets*, pages 104–118\. Springer, 2013\.doi:10\.1007/978\-3\-642\-40864\-9\_8\.URL[https://doi\.org/10\.1007/978\-3\-642\-40864\-9\_8](https://doi.org/10.1007/978-3-642-40864-9_8)\.
- \[65\]Gwendal Jouneaux and Jordi Cabot\.AgentSLA: Towards a service level agreement for AI agents\.*arXiv preprint arXiv:2511\.02885*, 2025\.URL[https://arxiv\.org/abs/2511\.02885](https://arxiv.org/abs/2511.02885)\.
- \[66\]Gaurav Naresh Mittal\.Capability Advertisement as a Market for Lemons: A Trust Layer for Heterogeneous Agent Networks\.*arXiv preprint arXiv:2606\.03034*, 2026\.URL[https://arxiv\.org/abs/2606\.03034](https://arxiv.org/abs/2606.03034)\.
## Appendix ASybil Capital and Fee Requirements
An attacker must lock at leastNSybilsminN\_\{\\mathrm\{Sybil\}\}s\_\{\\min\}in deposits\. Building reputation also spends at leastNSybilnrepfminN\_\{\\mathrm\{Sybil\}\}n\_\{\\mathrm\{rep\}\}f\_\{\\min\}in non\-recoverable fees\. The sum is a lower bound on the funding required while all deposits remain locked\. The analysis assumes each Sybil identity pays its own fees and that the attacker cannot recycle deposits during this period\.
For perfect feedback and a uniform prior, the update has a closed form\. For0<η<10<\\eta<1and interaction countnn,
α\(n\)=ηn\+1−ηn1−η,β\(n\)=ηn,ρ\(n\)=1−ηn\+11\+\(1−2η\)ηn\.\\alpha^\{\(n\)\}=\\eta^\{n\}\+\\frac\{1\-\\eta^\{n\}\}\{1\-\\eta\},\\qquad\\beta^\{\(n\)\}=\\eta^\{n\},\\qquad\\rho^\{\(n\)\}=\\frac\{1\-\\eta^\{n\+1\}\}\{1\+\(1\-2\\eta\)\\eta^\{n\}\}\.\(7\)Letr=ρtargetr=\\rho\_\{\\mathrm\{target\}\}with1/2<r<11/2<r<1\. The required count is
nrep=⌈log\(\(1−r\)/\(η\+r\(1−2η\)\)\)logη⌉\.n\_\{\\mathrm\{rep\}\}=\\left\\lceil\\frac\{\\log\\\!\\left\(\(1\-r\)/\(\\eta\+r\(1\-2\\eta\)\)\\right\)\}\{\\log\\eta\}\\right\\rceil\.\(8\)Forη=1\\eta=1, the score is\(n\+1\)/\(n\+2\)\(n\+1\)/\(n\+2\)andnrep=⌈\(2r−1\)/\(1−r\)⌉n\_\{\\mathrm\{rep\}\}=\\lceil\(2r\-1\)/\(1\-r\)\\rceil\. Targets at or below1/21/2require no interactions\. We checked these expressions against direct iteration for 36 combinations of longevity factor and target\. All counts agreed\.
The noisy\-feedback check uses independent Bernoulli outcomes with success probabilities 0\.9 and 0\.8\. Each target uses 1,000 trials with seeds 0 through 999 and a 5,000\-interaction horizon\. Means in Figure[7](https://arxiv.org/html/2609.21325#S5.F7)exclude trials that do not reach the target\. At success probability 0\.8 and target 0\.99, 659 trials reach the target and 341 do not\. The mean among reaching trials is 2,072\.95 interactions\.
## Appendix BEvaluation Design
Table 9:Agent evaluation design\. Each benchmark comparison uses 127 shared tasks unless noted otherwise\. The six\-harness comparison adds an elaborate system prompt and condensed\-notes memory\.##### Harnesses\.
The registry contains six configurations\. The baseline makes one model call\. ReAct runs a thought\-and\-action loop for up to four turns\. Plan\-then\-execute makes a planning call followed by an execution call\. Reflective retry answers, critiques, and revises in three calls\. The elaborate\-prompt harness makes one call with a longer system prompt\. The notes\-memory harness extracts notes, extends them, and answers from them in three calls\. Each harness checks estimated input usage and caps output tokens\. Estimates can overshoot the allowance, and hosted retrieval can consume additional server\-side tokens\. Cache reads are exempt from the token threshold but included in monetary cost\.
##### Models\.
OpenAI provides the evaluatedgpt\-5\.6\-lunaandgpt\-5\.6\-solmodels\. Anthropic providesclaude\-haiku\-4\-5\. Short labels in tables and plots areluna,sol, andhaiku\. Each run records its input, output, and cache\-read prices\. The cost figures use those run\-specific snapshots\.
##### Task pools\.
The GAIA comparisons draw from the text\-only subset of the benchmark\. Records with an empty or unknown reference answer are dropped, and so is every record carrying an attached file\. All GAIA levels remain in the pool, and every comparison requests the reasoning axis\. Task selection walks a portable sha256 index from the run seed, so a shared seed gives every configuration an identical 127\-task set and licenses paired statistics\. The mathematics domain converts the MATH\-500 test split into the same task format and draws 127 tasks from it under a separate seed\. Conversion keeps the original problem text, appends one instruction about the form of the answer, and stores the reference answer as a string\.
##### Calibration\.
The pilot ran two configurations on six tasks atB=8,000B=8\{,\}000to target a 40–60% solve rate\. Its median solve rate was 50%\. Binary outcomes have greatest variance near solve probability 0\.5, which motivated the calibration range\. The pilot exposed a default timeout that failed slow multi\-call runs\. Later comparisons used a 360\-second timeout\. Judge calibration against human\-labeled traces preceded the evaluations\.
##### Model and tool diagnostics\.
The diagnostics block ran five configurations across two models over 127 GAIA tasks atB=16,000B=16\{,\}000\. The configurations aregpt\-5\.6\-lunabaseline,gpt\-5\.6\-lunawith hosted web search,gpt\-5\.6\-solbaseline,gpt\-5\.6\-solwith hosted web search, andgpt\-5\.6\-solreflective retry\. Hosted search ran without a call cap in this block, which is why retrieval consumed the budget\.
##### Factorial comparison\.
The factorial block crossed three models with four harnesses over 127 GAIA tasks atB=8,000B=8\{,\}000under one shared seed\. Each of the twelve configurations ran once\. The in\-run cap admits a small overshoot on the input side, so scoring over\-budget items as failures after execution is what enforces the same\-budget rule\.
##### Extended harness comparison\.
The six\-harness comparison keeps the three models, the task seed, and the budget of the factorial block, and adds two harnesses\. Thesystem\_heavyharness supplies an elaborate system prompt\. Thenotes\_memoryharness supplies a condensed\-notes buffer\. Both ran in a later round with the invocation cache disabled, while the four\-harness grid ran with the cache enabled\.
##### Baseline repeatability\.
The baseline harness repeated five times per model atB=8,000B=8\{,\}000on the same task seed\. The invocation cache was disabled for these runs, so each replicate re\-invokes the model instead of replaying a stored output\. Thegpt\-5\.6\-lunareplicates solved 43, 37, 42, 36, and 38 tasks\.
##### Budget reconstruction check\.
The four tool\-less harnesses ongpt\-5\.6\-lunaran directly atB=4,000B=4\{,\}000on the same task seed as theB=8,000B=8\{,\}000grid\. The truncation estimate counts an item at the lower budget only when theB=8,000B=8\{,\}000run both solved that item and charged no more than 4,000 tokens\.
##### Bounded retrieval\.
The retrieval block ran a search\-equipped baseline on two models against a fresh tool\-less baseline, over 127 GAIA tasks atB=16,000B=16\{,\}000\. The cap allows the hosted search tool at most two calls per task, and each search call carries a fixed price folded into reported cost\. This block uses its own sampling seed, so its baselines pair with its capped\-search configurations and not with theB=8,000B=8\{,\}000grid\.
##### Mathematics transfer\.
The baseline and reflective\-retry harnesses ran on all three models over 127 converted MATH\-500 tasks atB=8,000B=8\{,\}000\. This block uses a separate seed and a separate evaluation database from the GAIA comparisons\.
## Appendix CEvaluation Details
We checked the pipeline against deterministic mock services at no cost before calibrating the budget\.
##### Experimental system\.
The system runs model and harness combinations on task sets selected by a shared seed\. It records outputs, quality, and provider\-reported usage\. Input, output, and cache\-read prices determine monetary cost\. Forℓ\\elltasks per domain at common allowanceBB, the nominal charged\-token allocation is\|configurations\|\|𝒟\|ℓB\|\\mathrm\{configurations\}\|\\,\|\\mathcal\{D\}\|\\,\\ell B\. Hereℓ\\ellcounts tasks per domain\. Scoring assigns over\-budget tasks zero credit, while all recorded usage still contributes to cost\.
##### Credential records\.
The evaluation system signs credential manifests with Ed25519 and provides a policy\-based verification API\.
##### Statistics\.
The original ratio intervals use 2,000 paired percentile\-bootstrap resamples of task indices\. A resample with zero denominator is omitted\. The retained approximate Wilcoxon probabilities are reported values\. The retrospective tests below use exact binary comparisons and explicit correction families\.
##### Usage and scoring\.
Token counts are agent\-reported, meaning harnesses we control relay provider usage\. An inexpensive model served as the LLM judge\. We measured agreement with human labels before the judge’s decisions affected results\.
## Appendix DPaired Statistical Checks
The retrospective checks use the binary task\-outcome matrix underlying Figure[6](https://arxiv.org/html/2609.21325#S5.F6)\. Each of its 127 columns is a shared task\. Row totals reproduce all 18 reported solved counts\. Archived task\-level records also confirm exclusive win and loss counts for 45 comparisons among ten configurations\.
For an omnibus model check, each task’s score is averaged across harnesses\. For a harness check, it is averaged across models\. Within each task we permute the tested labels and recompute the sum of squared deviations of the marginal means from their grand mean, multiplied by 127\. The null assumes the tested labels are exchangeable within a task\. Each test uses 19,999 random permutations and seed 20260907\. The probability estimate adds one to both the exceedance count and permutation count\. Bonferroni correction covers the four omnibus tests shown in Table[10](https://arxiv.org/html/2609.21325#A4.T10.fig1)\. These checks were performed during manuscript review, after the original analyses\.
Table 10:Retrospective omnibus tests that preserve task pairing\. Adjusted probabilities use a family of four tests\.Table 11:Exact paired tests\. Exclusive wins and losses count discordant tasks\. Model and harness comparisons use separate correction families\.The wins and losses in Table[11](https://arxiv.org/html/2609.21325#A4.T11.fig1)count tasks solved only by the first configuration and only by the second\. Conditional on their sum, the exact two\-sided binomial test assigns equal probability to either winner under the null\. Model comparisons use a family of three baseline pairs\. Harness comparisons use a separate family of 15 comparisons against the same model’s baseline\.相似文章
缩小AI信任差距:论证可信人工智能的独立认证
本文认为,当前负责任AI实践未能创造出一个奖励可信赖性的市场,因此提出独立的、以结果为导向的认证,以缩小'信任差距',使AI可信赖性变得可衡量、可比较并可在商业上获得回报。
如何解决AI代理之间的信任问题?
本文探讨了Web3和链上验证如何解决AI代理协作中的信任问题,包括通过代理商店(Agent Stores)作为基于声誉和交易历史的买卖服务的市场。
可验证的智能体基础设施:面向主权AI系统的基于证明的授权机制
本文提出了一种分布式信任框架(DTF),用于自主AI代理系统中的可验证、基于证明的授权,通过要求提供理由证明和共识执行来应对以身份为中心的权限所带来的风险。
如果 AI 代理无处不在,我们如何知道哪些值得信任?
随着 AI 代理变得无处不在,挑战从比较性能转向建立信任和声誉,需要新的发现和验证系统。
AgentAudit:一个用于 AI 代理全生命周期信任评估的开放、可扩展框架
AgentAudit 是一个开放且可扩展的框架,用于评估 AI 代理在能力、基础、安全和行为维度的全生命周期,实现精确的故障归因并突出不同语言模型之间的可信度差异。