When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains
Summary
This paper studies how LLM agents negotiate in a dynamic supply chain bargaining problem, benchmarking nine models from OpenAI, Google, and Alibaba against a Bayesian equilibrium and finding that capability, provider identity, and prompt design shape surplus creation and division.
View Cached Full Text
Cached at: 08/11/26, 08:02 AM
# 1 Introduction
Source: [https://arxiv.org/html/2608.07538](https://arxiv.org/html/2608.07538)
\\OneAndAHalfSpacedXI\\TheoremsNumberedThrough\\EquationsNumberedThrough\\RUNAUTHOR
Liang and Xu\\RUNTITLEDynamic Bargaining with LLM Agents\\TITLEWhen LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains\\ARTICLEAUTHORS\\AUTHORChen Liang Fasheng Xu \\AFFSchool of Business, University of Connecticut\\ABSTRACTAs LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid contracts that lose money for one party\. We study these questions in a canonical dynamic supply chain bargaining problem: a buyer with private demand information negotiates a quantity–payment contract with an uninformed seller\. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium, using 9,840 LLM\-to\-LLM negotiations\. Three findings emerge\. First, capability governs value creation\. LLM agents reach agreement in 98\.9% of negotiations and capture 95\.4% of first\-best surplus in undiscounted terms, but they average 2\.98 rounds against the Bayesian benchmark of 1\.25, and this delay erodes 21–34% of first\-best surplus, depending on patience\. The same capability ordering governs operational reliability: baseline models accept individually irrational contracts in 19\.2% of negotiations, versus 0\.0–0\.6% at mid\-tier and flagship, making automated profit verification the binding guardrail below the capability threshold\. Second, surplus capture is relational: provider identity predicts the direction of surplus flow more reliably than capability rank\. Under a common prompting regime, average self\-play buyer shares are 40% for OpenAI, 50% for Google, and 70% for Alibaba’s Qwen, an ordering that persists when communication is restricted and discounting is removed\. In cross\-provider matchups, reversing which provider sells shifts the division by 7–18 percentage points—as large as the within\-family swing from reversing which capability tier sells \(roughly 17 points\)—and a provider profile can override capability outright: the highly capable Qwen flagship is the weakest cross\-family seller\. Vendor choice is therefore a first\-order distributional decision\. Third, the prompt itself is a strategic lever\. Restricting agents to numeric offers shifts surplus in provider\-specific directions; removing discounting preserves every qualitative regularity but lengthens bargaining; and delegating to an LLM agent separates the principal’s economic patience \(the real cost of delay\) from the agent’s prompted strategic patience, a free deployment choice that is the strongest single driver of surplus division \(90% of explained variance in our design\), and the best setting varies by role and model\. Together, the results establish an equilibrium\-referenced audit of strategic AI agents along three dimensions: discounted efficiency, distributional profile, and operational reliability\.\\KEYWORDSLarge language models \(LLMs\), LLM agents, LLM\-to\-LLM negotiation, dynamic bargaining, supply chain negotiation, asymmetric information, automated procurement\\HISTORYCurrent version: July, 2026
The digitization of supply chains is shifting from automated execution to autonomous negotiation\. Walmart has deployed Pactum’s LLM agents for tail\-spend contracts across thousands of suppliers\.111[https://pactum\.com/procurement\-agent/](https://pactum.com/procurement-agent/)Arkestro, Globality, and Fairmarkit manage procurement negotiations for Boeing, BP, and other large buyers\.222See[https://www\.arkestro\.com/](https://www.arkestro.com/), and[https://www\.fairmarkit\.com/](https://www.fairmarkit.com/)\.The same shift extends to consumer and cross\-border commerce, where Alibaba’s Accio equips sourcing agents that negotiate supplier pricing and Xianyu’s FishBargain bargains on behalf of individual sellers\(Kong et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib28)\)\. A controlled demonstration underscores the stakes of full delegation: in Anthropic’sProject Deal, employees’ Claude agents traded real personal items in a marketplace with no human in the inner loop, and more capable agents systematically captured more value, selling higher and buying lower than weaker ones\.333[https://www\.anthropic\.com/features/project\-deal](https://www.anthropic.com/features/project-deal)As both sides of a transaction increasingly delegate to autonomous agents, bargaining is becoming an LLM\-to\-LLM activity conducted at machine speed and scale: a single recent international competition ran more than180,000180\{,\}000negotiations between AI agents\(Vaccaro et al\.[2026](https://arxiv.org/html/2608.07538#bib.bib39)\)\. This places the economic competence of autonomous LLM negotiators \(whether they create value, divide it predictably, and avoid costly errors\) at the center of the emerging agentic economy\.
For firms deciding which LLM vendor to adopt and which side of a transaction to automate, the operational question is whether LLM agents bargain in ways that are economically coherent, distributionally stable, and reliable enough to be trusted with high\-volume procurement\. These are decisions about capability thresholds for autonomous deployment, distributional consequences of counterparty choice, and guardrail architectures for acceptable risk\. Linguistic or leaderboard evaluations do not directly address such questions: operational value is measured in dollars foregone to delay, dollars misallocated across trading partners, and dollars lost when agents accept individually irrational contracts\.
Yet existing LLM evaluations leave open whether autonomous agents bargain coherently in operational contracting settings where performance can be benchmarked against economic theory and evaluated on deployment\-relevant outcomes\. We therefore ask whether general\-purpose LLM agents can recover equilibrium\-consistent bargaining behavior in a canonical supply chain contracting problem, and how their behavior varies with the vendor\-selection and role\-assignment decisions firms must make\. We benchmark nine LLMs from three providers \(OpenAI, Google, and Alibaba\) against the Perfect Bayesian Equilibrium of the alternating\-offer bargaining game ofFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\), in which a buyer with private demand information negotiates a quantity–payment contract\(q,T\)\(q,T\)with an uninformed seller\.
The experimental program comprises 9,840 LLM\-to\-LLM negotiations across three main\-analysis blocks \(symmetric verbal self\-play, cross\-family flagship pairings, and within\-family capability\-asymmetric pairings\) and seven design and robustness extensions \(structured\-offer bargaining without verbal communication, no\-discounting bargaining, strategic\-patience analysis, reasoning\-effort ablation, retail\-price sensitivity, prior sensitivity, and parameter\-size variation\)\. Prompts disclose the primitives of the game \(payoff formulas, demand distributions, discount factors, priors, and buyer type\) but agents are not told which contract to offer, when to separate types, or how to update beliefs\. The design therefore audits strategic execution in a stylized supply chain contracting setting rather than discovery of the setting from raw context\.
The design deliberately excludes the domain\-specific guardrails that surround production LLM procurement systems: rule\-based approval layers, cost\-verification databases, escalation protocols, and human sign\-off gates\. Although useful for deployment, such controls make it difficult to distinguish strategic competence from compliance with externally imposed constraints\. We therefore evaluate models in a common bargaining environment with minimal additional safeguards or prompt engineering, identifying their underlying bargaining behavior and the guardrails needed to support, constrain, or correct it\. The evidence delivers three empirical findings tied to the vendor\-selection and role\-assignment problem\.
Finding 1: Capability is the value\-creation lever that governs both efficiency and reliability\.LLM agents reach agreement in 98\.9% of negotiations and capture 95\.4% of first\-best surplus in undiscounted terms, but they average 2\.98 rounds against the Bayesian benchmark of 1\.25\. Under discounting—which we treat as a stress test on costly continuation rather than a literal time\-preference estimate \(§[3\.3](https://arxiv.org/html/2608.07538#S3.SS3)\)—this delay erodes 21–34% of first\-best surplus, depending on patience\. Of the efficiency variation the experimental design systematically moves, capability explains roughly 66% \(Shapley decomposition\): flagship models take more rounds than baselines \(3\.25 versus 2\.75\) but achieve materially higher efficiency \(98\.9% versus 91\.0%\), consistent with an iterative\-search pattern in which capable agents invest in proposal–rejection cycles rather than settling at round\-one heuristic offers\.
The same capability ordering governs operational reliability: baseline models accept individually irrational contracts \(negative profit for one party\) in 19\.2% of cases, versus 0\.6% for mid\-tier and 0\.0% for flagship models\. This order\-of\-magnitude gap separates a verification\-gated weak\-model cluster from a lighter\-monitoring strong\-model cluster, making automated profit verification a necessary guardrail below the threshold\.
Finding 2: Provider identity is the distributional lever, and surplus capture depends on the counterparty and role\.Capability and provider identity load on different margins: capability governs how much surplus is created \(Finding 1\), while provider identity is the more reliable predictor of who captures it\. Among the model versions evaluated under a common prompt protocol, self\-play surplus division varies markedly across providers: Qwen models average 70% buyer share \(within\-family range 53–91% across capability tiers\), Gemini 50% \(44–54%\), and OpenAI 40% \(38–42%\)\. The mean gap between Qwen and OpenAI \(approximately 30 percentage points\) is comparable to the largest within\-provider spread \(39 points across Qwen tiers\), so the 53\.5% pooled buyer share masks three qualitatively distinct bargaining profiles rather than a single capability\-tier effect\. Of the surplus\-division variation the experimental design systematically moves, the announced patience parameters explain 90% \(Shapley decomposition\), while capability, buyer type, and first\-proposer assignment together contribute the remaining 10%\. Prompted, common\-knowledge patience, however, is a configuration choice rather than a model attribute \(Finding 3\); among the model\-side factors a firm cannot prompt away, provider identity is the strongest predictor of who captures surplus\. Public demonstrations such as Anthropic’sProject Deal, in which upgrading one’s own agent reliably improved its terms, reinforce the intuition that a more capable agent is a better bargainer\. Our results qualify it: capability reliablycreatesvalue \(Finding 1\), but itsdistributionaladvantage does not survive across providers\.
Even holding provider fixed, a cross\-tier GPT pairing makes the point: switching the buyer from GPT\-5\.2 to the weaker GPT\-5\-mini against an unchanged GPT\-5\.2 sellerlowersbuyer share from37\.9%37\.9\\%to35\.2%35\.2\\%, so within a family the more capable model captures the larger share\. Across providers, however, capability rank alone does not determine who captures surplus: a highly capable Qwen flagship is the weakest cross\-family seller, so a model’s provider profile can override its capability\. Cross\-family direction effects amplify this pattern: when flagship models from different providers negotiate each other, the direction of the pairing produces 7–18 percentage\-point swings in surplus, comparable to the within\-family capability swing; Gemini captures 67% of surplus as a cross\-family buyer and retains 48% as seller, while Qwen retains only 27% as a cross\-family seller\. For firms deploying different vendors on opposite sides of a procurement interaction, model choice is therefore a first\-order distributional decision, not a commodity input\.
Finding 3: The principal’s configuration choices are strategic levers\.Because firms bargain through delegated agents rather than directly, the prompt serves as the agent’s mandate: three economically meaningful design choices systematically shape bargaining outcomes\. Constraining the communication channel reshapes provider biases: removing natural language shifts OpenAI’s baseline and mid\-tier modelstowardbuyers \(GPT\-5\-mini: 41\.9% to 51\.3%\) while shifting Gemini’s mid\-tier and flagship models toward sellers, so the verbal channel is doing strategic work, not decorating offer behavior\. Removing the discounting framework from the prompt leaves the value\-creation, distributional, and reliability rankings intact but changes negotiation tempo, so discount\-factor language operates as a tempo cue rather than a substantive driver of the results\. Most distinctively, delegation to LLM agents separates two objects that classical bargaining treats as a single discount\-factor parameter:economic patience\(the principal’s fixed real\-time cost of delay\) andstrategic patience\(the agent’s prompted discount factor, which the principal sets at deployment\)\. Patience thus becomes a configurable design lever rather than an exogenous primitive, and choosing it well raises realized payoff by economically meaningful margins\. Configuration is therefore a first\-order design problem alongside model and vendor selection\.
The paper makes three contributions to operations management research on AI\-mediated contracting\. Empirically, we document a dimension of LLM heterogeneity that standard capability benchmarks miss \(provider\-level bargaining profiles\) and show it is as consequential as within\-family capability differences once counterparties are heterogeneous, establishing vendor choice as a strategic operations decision rather than a technical one\. Methodologically, we introduce a reproducible equilibrium\-referenced audit framework: an executable implementation of the branch\-specific PBE, validated cell\-by\-cell against the equilibrium’s analytical properties—a portable template for evaluating autonomous agents in game\-theoretic operational settings\. Because the audit spans the current capability frontier \(nine models, three tiers, and roughly three generations\), it separates structural regularities that travel across model vintages from model\-specific results that recalibrate as providers update, so the framework remains informative as the underlying models improve\. Practically, we identify three deployment\-relevant evaluation dimensions \(time\-adjusted efficiency, provider\-level distributional profile, and operational reliability\), guiding capability thresholds, role assignment, and guardrail architecture\.
Delegation changes both the stakes and the relevant benchmark\. When firms authorize LLM agents to negotiate on their behalf, agents’ offers, concessions, and acceptance decisions directly determine realized contract terms and surplus\. This is particularly consequential in procurement, where automation can extend bargaining to supplier relationships that firms cannot economically negotiate one by one, allowing gains and systematic errors alike to scale across transactions\(Van Hoek et al\.[2022](https://arxiv.org/html/2608.07538#bib.bib40)\)\. The central question is therefore not whether LLM agents bargain like humans, but whether they advance their principals’ economic interests\. We evaluate this competence against the Perfect Bayesian Equilibrium characterized byFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)\. The equilibrium provides reference values for agreement timing, contract form, surplus division, screening, and feasibility\. Theory thus serves as measurement infrastructure, allowing us to identify when an agent creates value, transfers value to its counterparty, or destroys value through delay or individually irrational acceptance\. Human\-subject experiments address complementary questions about human–agent interaction and the transmission of principal heterogeneity into delegated outcomes; for example,Imas et al\. \([2025](https://arxiv.org/html/2608.07538#bib.bib23)\)show that delegated outcomes retain substantial variation across principals\. Our design instead fixes the payoff objective and evaluates the agent against the corresponding equilibrium, providing a direct measure of agent\-level bargaining competence\.
The remainder of the paper develops these results: Section[2](https://arxiv.org/html/2608.07538#S2)reviews the three research streams we build on; Section[3](https://arxiv.org/html/2608.07538#S3)describes the experimental design and Bayesian benchmark; Sections[4](https://arxiv.org/html/2608.07538#S4)and[5](https://arxiv.org/html/2608.07538#S5)present the first two findings on value creation and distribution; Section[6](https://arxiv.org/html/2608.07538#S6)develops the third finding on the principal’s configuration levers and summarizes the robustness program \(detailed in the Appendix\); and Section[7](https://arxiv.org/html/2608.07538#S7)concludes with practical implications and future directions\.
## 2Literature Review
Our study contributes to three streams that together define what it means to evaluate an autonomous bargaining agent for operational deployment: dynamic bargaining under asymmetric information, which supplies the normative benchmark; behavioral evidence on bargaining in operations, which establishes the performance baseline human negotiators reach without equilibrium computation; and the emerging machine\-behavior literature, which treats AI systems as strategic actors requiring their own audit\.
Dynamic bargaining models formalize how parties trade off surplus division against costly delay\. The alternating\-offers framework\(Rubinstein[1982](https://arxiv.org/html/2608.07538#bib.bib35)\)remains canonical: equilibrium allocations depend on relative patience, generating sharp comparative statics for who captures surplus and how quickly agreement occurs\. Under private information the benchmark shifts to Perfect Bayesian Equilibrium, and strategic delay, signaling, and screening become central\(Kennan and Wilson[1993](https://arxiv.org/html/2608.07538#bib.bib25), Myerson and Satterthwaite[1983](https://arxiv.org/html/2608.07538#bib.bib32)\); results related to the Coase conjecture show that when delay is costly, equilibrium often features rapid agreement, with residual inefficiency tied to informational frictions rather than bargaining mechanics\(Coase[1972](https://arxiv.org/html/2608.07538#bib.bib6), Gul et al\.[1986](https://arxiv.org/html/2608.07538#bib.bib18)\)\.444The Coase prediction is conditional on the interpretation of the discount factor; Section[3\.3](https://arxiv.org/html/2608.07538#S3.SS3)discusses how the economic content ofδ\\deltadiffers in LLM\-to\-LLM bargaining\.In operations,Feng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)adapt these ideas to supply chain contracting: a buyer with private demand information and an uninformed seller bargain over a single contract\(q,T\)\(q,T\), and incomplete information generates screening, signaling, and quantity distortion or pooling along the path\. Their clean predictions on agreement speed, contract form, and surplus division make the environment well suited for benchmarking automated negotiators\.
A parallel stream establishes that human decision makers achieve high performance in bargaining despite bounded rationality\(Simon[1955](https://arxiv.org/html/2608.07538#bib.bib38)\)\. Controlled experiments show human negotiators often approach theoretical benchmarks on efficiency and agreement while exhibiting systematic patterns in delay, disclosure, and fairness not captured by purely rational models:Davis and Hyndman \([2021](https://arxiv.org/html/2608.07538#bib.bib11)\)find high agreement rates with limited delay under private information, attributing this to experience, institutional norms, and communication, andDavis et al\. \([2022](https://arxiv.org/html/2608.07538#bib.bib10)\)extend the result to procurement and assembly structures\. Communication is central:Camerer et al\. \([2019](https://arxiv.org/html/2608.07538#bib.bib3)\)show unstructured text carries signals that predict outcomes,Davis and Hyndman \([2025](https://arxiv.org/html/2608.07538#bib.bib12)\)that voluntary disclosure overcomes matching frictions, andHaruvy et al\. \([2020](https://arxiv.org/html/2608.07538#bib.bib20)\)that protocol design shifts efficiency and surplus division at fixed primitives, with fairness and contract framing also shaping outcomes\(Katok and Pavlov[2013](https://arxiv.org/html/2608.07538#bib.bib24)\)\. This literature sets a high behavioral baseline that humans reach through calculation, norms, and communication rather than strict backward induction, motivating the question of whether LLM agents replicate it, and whether they do so through similar or qualitatively different mechanisms\.
As LLM agents enter procurement and contracting workflows, a growing literature studies AI systems as behavioral objects whose actions can deviate systematically from equilibrium and from human behavior\(Rahwan et al\.[2019](https://arxiv.org/html/2608.07538#bib.bib34), Manning and Horton[2026](https://arxiv.org/html/2608.07538#bib.bib31)\)\. This literature models LLMs as economic actors whose strategic conduct depends on prompting and system design rather than equilibrium reasoning alone\(Filippas et al\.[2024](https://arxiv.org/html/2608.07538#bib.bib15), Wang et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib41), Shahidi et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib36), Hadfield and Koh[2025](https://arxiv.org/html/2608.07538#bib.bib19)\)\. In static settings, LLMs replicate human decision biases, often more strongly than human managers\(Chen et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib5)\):Liu et al\. \([2025](https://arxiv.org/html/2608.07538#bib.bib29)\)document pull\-to\-center and demand\-chasing in inventory tasks, alongside a “paradox of intelligence” in which more capable models overthink into greater irrationality\. Role and relationship framing in the prompt shifts how LLMs divide surplus between supply chain parties\(Liu and Ning[2025](https://arxiv.org/html/2608.07538#bib.bib30)\), and in contract design under moral hazard, LLM principals favor enforceable incentive contracts over the trust\-based bonus contracts humans often choose\(Kirshner et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib26)\)\.
Evidence on dynamic negotiation is more mixed\.Kirshner et al\. \([2026](https://arxiv.org/html/2608.07538#bib.bib27)\)find LLM agents more agreement\-oriented than humans, raising efficiency but potentially inequality;Zhu et al\. \([2025](https://arxiv.org/html/2608.07538#bib.bib45)\)warn of behavioral anomalies and losses in consumer settings; andHasija and Castillo \([2025](https://arxiv.org/html/2608.07538#bib.bib21)\)identify a role\-based asymmetry in which buyer\-role agents outperform supplier\-role agents through anchoring\. Closest to our work,Chen and Huang \([2026](https://arxiv.org/html/2608.07538#bib.bib4)\)study human\-retailer and LLM\-supplier bargaining in a two\-tier chain and show LLM suppliers reproduce many human–human patterns but need externally specified reservation\-profit guidance for stability\. Our study is complementary on three dimensions: LLM\-to\-LLM rather than human–LLM negotiation, isolating agent behavior from human adaptation; theFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)asymmetric\-information setting, where competence demands signaling, screening, and belief updating; and benchmarking against a validated PBE rather than human data\.
Two further studies vary other sources of heterogeneity in delegated negotiation:Imas et al\. \([2025](https://arxiv.org/html/2608.07538#bib.bib23)\)hold the model fixed and vary the human principals who write the prompts, showing delegated outcomes inherit and amplify principal heterogeneity, whileVaccaro et al\. \([2026](https://arxiv.org/html/2608.07538#bib.bib39)\)run a large open LLM\-to\-LLM competition \(182,812182\{,\}812negotiations,286286prompt engineers\) that flags prompt strategy as a central driver\. We add the provider layer under a common baseline prompt and separately test whether principal\-controlled configuration choices shift outcomes relative to the equilibrium benchmark\. The three designs thus locate heterogeneity at different layers of the delegation stack \(principal, prompt, provider\), and our equilibrium\-referenced audit supplies the provider\-level component of the integrated theory of AI negotiation thatVaccaro et al\. \([2026](https://arxiv.org/html/2608.07538#bib.bib39)\)call for\.
Three gaps remain for operations management research\. First, equilibrium\-grounded audits of LLM agents in canonical dynamic bargaining under asymmetric information are absent, even though success in such settings requires both operational calculation and strategic inference; Finding 1 and its decomposition address this\. Second, existing evaluations lack a validated theoretical reference, so deviations cannot be cleanly attributed to the environment, the evaluator, or the agent; our executable Bayesian benchmark \(Section[3](https://arxiv.org/html/2608.07538#S3)\) removes the ambiguity\. Third, general\-purpose LLM evaluations such as leaderboards and linguistic benchmarks assess competence but poorly predict operational value; Findings 2 and 3 and the three\-dimensional framework of Section[7](https://arxiv.org/html/2608.07538#S7)close this gap\.
More broadly, a management literature documents how generative AI reshapes task allocation, productivity, and organizational design\(Ide and Talamas[2025](https://arxiv.org/html/2608.07538#bib.bib22), Xu et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib43), Eloundou et al\.[2024](https://arxiv.org/html/2608.07538#bib.bib13), Brynjolfsson et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib2)\)and how AI supply chains shape deployment\(Xu et al\.[2024](https://arxiv.org/html/2608.07538#bib.bib44), Fransoo et al\.[2026](https://arxiv.org/html/2608.07538#bib.bib16)\)\. We respond to calls to engage the distinctive features of generative AI rather than treat it as a faster, cheaper technology\(Simchi\-Levi et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib37), Cohen et al\.[2026](https://arxiv.org/html/2608.07538#bib.bib8), Dai et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib9)\), focusing on a narrow but consequential unit: bilateral contracting under demand uncertainty, where the value question admits both a normative benchmark and a concrete deployment decision\.
## 3Experimental Design
The experimental program is built around the paper’s central question: when a firm delegates bargaining to an LLM agent, which of its choices move outcomes? It is therefore organized to separate three delegation layers—capability tier, provider identity, and prompt configuration—and to benchmark outcomes in each layer against a validated equilibrium that provides a known normative reference for timing, contract form, and surplus division\. The core block audits bargaining competence with a2×2×42\\times 2\\times 4factorial design \(Table[1](https://arxiv.org/html/2608.07538#S3.T1)\) crossing buyer type, first proposer, and patience configuration\. The three factors test, respectively, information revelation under high versus low demand, protocol effects on screening versus signaling, and the patience regimes identified byFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)where rational behavior transitions from truth\-telling to pooling\. Each condition is replicated 15 times per model across the nine models of Table[2](https://arxiv.org/html/2608.07538#S3.T2), yielding 2,160 negotiations in the baseline block\.
Our prompts fully disclose the primitives of the game but not the equilibrium strategy\. Agents can therefore compute any function of the disclosed parameters in principle, but they are not given the branch\-specific screening or signaling logic fromFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)\. The design tests whether LLMs can translate disclosed strategic structure into approximately coherent bargaining actions through interaction rather than whether they can infer the environment from scratch\. Because the prompts are role\-specific, the reported surplus shares should be interpreted as behavior under a common prompting regime rather than as prompt\-free estimates of intrinsic bargaining bias\.
Table 1:Main Experimental Design and Sample DistributionDesign ComponentValueBuyer type2 levels: High \(μ=80,σ=10\\mu=80,\\sigma=10\), Low \(μ=40,σ=10\\mu=40,\\sigma=10\)First proposer2 levels: Buyer, SellerPatience configuration4 levels:\(0\.9,0\.9\)\(0\.9,0\.9\),\(0\.7,0\.9\)\(0\.7,0\.9\),\(0\.4,0\.9\)\(0\.4,0\.9\),\(0\.9,0\.4\)\(0\.9,0\.4\)Replications per condition15Models9
Notes:Patience configurations are ordered as\(δB,δS\)\(\\delta\_\{B\},\\delta\_\{S\}\), buyer patience first and seller patience second\.
### 3\.1Model Selection
To support both within\-provider capability comparisons and cross\-provider heterogeneity comparisons, we selected nine models from three commercial providers\. The three families come from independent organizations: OpenAI and Google are U\.S\.\-based, closed\-source frontier\-model providers, while Alibaba is a China\-based frontier\-model provider\. Distinct training data, alignment objectives, and post\-training procedures give the sample cross\-provider heterogeneity rather than within\-family recipe similarity\. Each family offers at least three publicly accessible capability tiers \(flagship, mid\-tier, baseline\), shown in Table[2](https://arxiv.org/html/2608.07538#S3.T2)\. Within each family the baseline tier is the prior\-generation model: GPT\-4o\-mini and Qwen2\.5\-14B are non\-reasoning models, whereas Gemini\-2\.5\-Flash \(like all six mid\-tier and flagship models\) exposes reasoning \(“thinking”\) capability, a distinction that proves consequential for operational reliability \(Section[4\.2](https://arxiv.org/html/2608.07538#S4.SS2)\)\. Extensions to Anthropic \(Claude\), Meta \(Llama\), and xAI \(Grok\) are valuable directions for future work\.
Table 2:Language Models by Provider and Capability TierProviderFlagshipMid\-TierBaselineOpenAIGPT\-5\.2GPT\-5\-miniGPT\-4o\-miniGoogleGemini\-3\-ProGemini\-3\-FlashGemini\-2\.5\-FlashAlibabaQwen3\-MaxQwen3\-32BQwen2\.5\-14B
### 3\.2Negotiation Protocol
Our baseline negotiation protocol implements the dynamic bargaining framework ofFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)adapted for LLM agents; variations used in the design and robustness extensions \(R1 structured communication, R2 no\-discounting bargaining\) are described where introduced\. A seller with production costc=30c=30negotiates with a buyer over a wholesale contract\(q,T\)\(q,T\)specifying order quantity and transfer payment\. The buyer privately observes its demand type; the seller holds an equal prior on each type\. Both parties know the retail pricer=60r=60, production cost, discount factors, and the ten\-round maximum\. Parties alternate offers until agreement or the round limit; failure to agree yields zero payoffs\. The theoretical benchmark predicts rational agents reach agreement within two rounds through screening, signaling, and quantity distortion when relevant; this is the performance target against which we evaluate LLM agents\.
To mimic realistic supply chain negotiation environments, the protocol allows agents to exchange natural\-language messages alongside formal offers\. A negotiation terminates when the responder signals acceptance in the structured action field or the ten\-round cap is reached; unaccepted negotiations hitting the cap are counted as zero\-surplus outcomes for agreement and efficiency metrics, but excluded from analyses requiring accepted contract terms, profits, or surplus shares\. Conditional on acceptance, final contract terms are resolved in two stages\. First, a language\-model extractor \(GPT\-4o\-mini\) recovers the deal from the full conversation history, subject to numerical validation \(quantity and payment non\-null, strictly positive, and finite\)\. Second, if the extractor returns no validated values, a deterministic rule\-based scan recovers the most recent structured proposal satisfying the same criteria\.555We re\-extracted terms on a stratified sample of 300 experiments \(all 2 fallback\-recovered deals, all 24 round\-limit non\-agreements, and 274 randomly drawn extractor\-resolved deals balanced across models\) using a second LLM from a different family \(Claude Sonnet 4\.6\)\. The two extractors agreed on deal status in 100% of cases and on quantity within±0\.5\\pm 0\.5units in all 276 jointly confirmed deals; payment agreed in all but one, a per\-unit\-versus\-total ambiguity in a gpt\-4o\-mini contract that quoted a per\-unit price in the payment field\. A corpus\-wide scan found four such gpt\-4o\-mini cases, which we rescaled to the implied total; the adjustment leaves efficiency unchanged and shifts aggregate surplus shares by at most 0\.2 percentage points\.
For agreements reached in roundτ\\tau, we compute buyer profitπB=r⋅𝔼\[min\(D,q\)\]−T\\pi\_\{B\}=r\\cdot\\mathbb\{E\}\[\\min\(D,q\)\]\-Tusing newsvendor expected revenue, and seller profitπS=T−cq\\pi\_\{S\}=T\-cq\. Effective utilities incorporate time discounting asUB=δBτ−1πBU\_\{B\}=\\delta\_\{B\}^\{\\tau\-1\}\\pi\_\{B\}andUS=δSτ−1πSU\_\{S\}=\\delta\_\{S\}^\{\\tau\-1\}\\pi\_\{S\}, consistent with theRubinstein \([1982](https://arxiv.org/html/2608.07538#bib.bib35)\)bargaining framework \(we discuss the role and interpretation ofδ\\deltain Section[3\.3](https://arxiv.org/html/2608.07538#S3.SS3)\)\. Both agents are instructed to reject individually irrational offers yielding negative profit\. Appendix[10\.1](https://arxiv.org/html/2608.07538#S10.SS1)provides a representative negotiation transcript\.
Our design prioritizes equilibrium\-referenced identification over realism\. Five choices support this priority\. First, we fix temperature at 1\.0 \(the API default and a common production setting\) to characterize out\-of\-the\-box stochastic behavior\. Second, we set outside options to zero, followingRubinstein \([1982](https://arxiv.org/html/2608.07538#bib.bib35)\), to isolate pure bargaining from exit\-threat credibility\. Third, we use role\-specific but structurally symmetric buyer and seller prompts, so observed heterogeneity reflects model capability and role rather than prompt engineering \(full prompts in Appendix[8](https://arxiv.org/html/2608.07538#S8)\)\. Fourth, we disclose the game’s primitives without supplying reservation profits or target utilities, which would substitute designer\-imposed discipline for the endogenous execution we measure\. Fifth, we study LLM\-to\-LLM rather than human–LLM bargaining so that outcome differences attribute to model behavior rather than human adaptation or interface effects\. Natural\-language messages are permitted as a native LLM channel while formal offer fields preserve machine\-checkable contracts \(Extension R1 shows the regularities persist under structured\-only offers\), complementing concurrent human–LLM studies\(Chen and Huang[2026](https://arxiv.org/html/2608.07538#bib.bib4)\)that prioritize realism over equilibrium referencing\.
### 3\.3What Patience Means in the Experiment
Three roles for patience\.Patience enters the paper in three distinct roles\. First, it is atheoreticalobject: theFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)dynamic\-bargaining benchmark requires discount factors to pin down continuation values, agreement timing, and surplus division; without discounting, the distributional prediction is not uniquely defined\. Second, it is anexperimentalobject: in the main treatments each agent is told both parties’ discount factors, matching the common\-knowledge assumption of the benchmark and removing any need to infer hidden patience\. Third, it is adeploymentobject: a real principal may face an economic cost of delay that is not literally encoded as a prompt parameter\. Our main analysis uses prompted patience as measurement infrastructure; Extension R2 then removes that infrastructure to test whether the qualitative findings survive in a no\-discounting regime\.
What the agents know, and what the prompted value represents\.The patience pair\(δB,δS\)\(\\delta\_\{B\},\\delta\_\{S\}\)is common knowledge, disclosed to both agents alongside the retail price, production cost, and round limit \(Appendix[8](https://arxiv.org/html/2608.07538#S8)\), while the buyer’s demand type is the sole private information\. The experiment therefore asks whether LLMs canuseknown primitives, not whether they caninfera hidden one; uncertainty over a counterparty’s patience is a distinct problem we leave to future work\. The prompted discount factor is an announced utility parameter specified in the prompt; it is neither elicited from the human principal nor inferred by the agent\.
Why prompted patience is also a design lever\.The gap between the experimental and deployment roles is not a defect but the source of a managerial design lever\. In classical bargaining a singleδ\\deltadoes double duty \(it is at once the agent’s weight on delay and the principal’s real cost of waiting\) because a bargaining round and a real\-time period coincide\. Under LLM mediation they decouple: a round of LLM exchange takes seconds, whereas the deployed economic period is governed by API latency, human review, and approval workflows\. We therefore distinguishstrategic patienceδstrat\\delta^\{\\text\{strat\}\}, the discount factor the agent applies across alternating\-offer rounds and the value we set by prompt, fromeconomic patienceδecon\\delta^\{\\text\{econ\}\}, the principal’s fixed exogenous discount over real time\. Because the prompted value isδstrat\\delta^\{\\text\{strat\}\}and need not equalδecon\\delta^\{\\text\{econ\}\}, the principal’s problem is not to translate a human’s patience into a prompt but totuneδstrat\\delta^\{\\text\{strat\}\}against a fixed economic constraint, a principal\-specification exercise we develop in Section[6\.3](https://arxiv.org/html/2608.07538#S6.SS3)\.
Reading the efficiency metrics\.Consistent with this view we report bothundiscountedefficiency, appropriate when the object of interest is allocative quality independent of process, anddiscountedefficiency, which measures sensitivity to continuation cost under the promptedδstrat\\delta^\{\\text\{strat\}\}\. Discounted efficiency should be read as a stress test of how outcomes respond when delay is made costly, not as a direct estimate of either human time preference or operational reliability\.666One can additionally motivate themagnitudeofδ\\deltathrough a per\-round breakdown hazard, writingδk=e−\(αk\+λk\)Δt\\delta\_\{k\}=e^\{\-\(\\alpha\_\{k\}\+\\lambda\_\{k\}\)\\Delta t\}withαk\\alpha\_\{k\}a time\-preference rate andλk\\lambda\_\{k\}a breakdown hazard over inter\-round intervalΔt\\Delta t\(Feng et al\.[2015](https://arxiv.org/html/2608.07538#bib.bib14)\); at the per\-second round pace of our experiments a hazard\-based analogy is more natural than literal time preference\. We do not identify these channels in our data and do not rely on this reading\.When we study prompt design in Section[6\.3](https://arxiv.org/html/2608.07538#S6.SS3), we separately evaluate realized payoffs at the principal’s economic patienceδecon\\delta^\{\\text\{econ\}\}\.
### 3\.4Bayesian Benchmark Implementation and Validation
Throughout the paper, we use two related theoretical references\. The complete\-information Rubinstein outcome provides a first\-best bargaining reference: quantity is set at the type\-specific first\-best level and surplus is divided according to alternating\-offer patience\. The primary asymmetric\-information benchmark is the branch\-specific equilibrium characterization ofFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)\. Because their formal incomplete\-information PBE is derived for seller\-initiated bargaining, our seller\-first cells implement that benchmark directly, while buyer\-first cells use a separately derived buyer\-first separating equilibrium built from the same primitives\. Appendix[12](https://arxiv.org/html/2608.07538#S12)provides the derivation and implementation details\.
Each cell’s equilibrium is determined by its patience pair\(δB,δS\)\(\\delta\_\{B\},\\delta\_\{S\}\)and priorβ\\beta\. In the main 16\-condition design withβ=0\.5\\beta=0\.5, all buyer\-first cells settle in round 1 \(separating\) and all seller\-first cells settle by round 2 \(the H\-type accepts the seller’s round\-1 offer; the L\-type counteroffers in round 2\)\. The benchmark agreement time averages 1\.25 rounds \(1\.00 buyer\-first, 1\.50 seller\-first\)\.
### 3\.5Experimental Program Overview
The paper comprises three main\-analysis blocks and seven design and robustness extensions, summarized in Table[3](https://arxiv.org/html/2608.07538#S3.T3)\. The main analysis is organized around the three delegation layers\. M1 \(the 16\-condition symmetric self\-play factorial described above\) and M3 \(GPT\-5\.2 paired with GPT\-5\-mini, Section[4\.5](https://arxiv.org/html/2608.07538#S4.SS5)\) vary thecapabilitylayer; M2 pairs flagship models from different providers in both buyer and seller roles, varying theproviderlayer and enabling tests of cross\-family direction effects \(Section[5](https://arxiv.org/html/2608.07538#S5)\); the configuration extensions then vary thepromptlayer \(Section[6](https://arxiv.org/html/2608.07538#S6)\)\.
The design and robustness extensions comprise three configuration analyses reported in Section[6](https://arxiv.org/html/2608.07538#S6)—communication mode \(R1\), removal of the discounting framework from the prompt \(R2\), and strategic patience \(R3\)—plus four additional robustness extensions: reasoning effort \(R4\), surplus magnitude \(R5\), prior beliefs \(R6\), and model size \(R7\)\. R2 doubles as an external\-validity check: the core efficiency, delay, and reliability regularities survive and the provider ordering persists when discounting is removed from the prompt, indicating that prompted patience is a design feature for the equilibrium audit rather than a precondition for the findings\. The completed program covers 9,840 unique LLM\-to\-LLM negotiations and 31,592 bargaining rounds\.
Table 3:Experimental Program OverviewBlockDescriptionNNMain AnalysisM1Symmetric LLM \(9 models×\\times16 conditions×\\times15\)2,160M2Cross\-family flagship pairings \(GPT\-5\.2, Gemini\-3\-Pro, Qwen3\-Max\)1,440M3Within\-family cross\-capability \(GPT\-5\.2↔\\leftrightarrowGPT\-5\-mini\)480Design and Robustness ExtensionsR1Structured communication \(no natural language\)2,160R2No\-discounting bargaining540R3Strategic\-patience analysis \(additional patience configuration\)540R4Reasoning\-effort ablation \(GPT\-5\.2, low/medium/high\)480R5Retail\-price robustness \(OpenAI models,r=120r=120\)720R6Seller\-prior sensitivity \(β∈\{0\.3,0\.5,0\.7\}\\beta\\in\\\{0\.3,0\.5,0\.7\\\}\)1,080R7Parameter\-size ablation \(Qwen3 family\)240Total9,840
Notes:NNis the number of unique LLM\-to\-LLM negotiations per block\. These counts are of newly run negotiations; several blocks additionally reuse M1 cells as comparison baselines \(e\.g\., M3 the two homogeneous configurations, R2 the symmetric\-patience\(0\.9,0\.9\)\(0\.9,0\.9\)cells, and R4 the medium\-effort arm\), so the analysis samples reported later exceed the counts shown here\. M1–M3 constitute the main analysis; R1–R7 are the design and robustness extensions\. Factor\-by\-factor breakdowns, round counts, and replication rates appear in Table[26](https://arxiv.org/html/2608.07538#S11.T26)of Appendix[11](https://arxiv.org/html/2608.07538#S11)\.
## 4Negotiation Performance of LLM Agents
The balanced factorial and fixed\-prompt LLM\-to\-LLM design of Section[3](https://arxiv.org/html/2608.07538#S3), audited against the validated Bayesian benchmark, lets us isolate how much of an agent’s bargaining outcome is created by the model itself\. For a firm deciding whether to delegate a class of procurement contracts to an autonomous agent, the first\-order question is how much value the agent captures relative to what a theoretically rational counterparty would capture\. We evaluate LLM performance along four outcomes tied to that decision: agreement rate \(does the deal close?\), negotiation duration in rounds \(how costly is the delay?\), undiscounted efficiency \(what share of first\-best surplus is realized?\), and discounted efficiency \(what share survives once delay is priced?\)\. The Perfect Bayesian Equilibrium ofFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)predicts rational agents reach agreement within two rounds whenever gains from trade are positive, providing the normative benchmark against which we audit behavior\. Later sections use Rubinstein only as a complete\-information reference for how patience would divide first\-best surplus if buyer type were known\. Table[4](https://arxiv.org/html/2608.07538#S4.T4)summarizes LLM performance relative to this Bayesian benchmark across key dimensions\. This section establishes Finding 1: capability is the value\-creation lever, governing both how much surplus agents realize \(competence, §[4\.1](https://arxiv.org/html/2608.07538#S4.SS1)\) and how reliably they avoid individually irrational contracts \(reliability, §[4\.2](https://arxiv.org/html/2608.07538#S4.SS2)\)\.
Table 4:Negotiation Performance: LLM Agents vs Bayesian BenchmarkConditionLLM AgentsBayesianAgreementRoundsEfficiencyDiscEfficiencyAgreementRoundsEfficiencyDiscEfficiency\(%\)\(%\)\(%\)\(%\)\(%\)\(%\)Overall98\.92\.9895\.466\.2100\.01\.25100\.096\.4By Model Tier:Flagship100\.03\.2598\.965\.3100\.01\.25100\.096\.4Mid\-tier99\.92\.9396\.468\.0100\.01\.25100\.096\.4Baseline96\.82\.7591\.065\.4100\.01\.25100\.096\.4By First Proposer:Buyer First98\.42\.9794\.866\.7100\.01\.00100\.0100\.0Seller First99\.42\.9996\.165\.8100\.01\.50100\.092\.9By Buyer Type:High\-type99\.52\.8895\.167\.0100\.01\.00100\.0100\.0Low\-type98\.23\.0895\.765\.4100\.01\.5199\.992\.8By Patience:\(0\.9, 0\.9\)98\.03\.3995\.274\.7100\.01\.25100\.097\.5\(0\.7, 0\.9\)99\.43\.0996\.163\.8100\.01\.25100\.096\.1\(0\.4, 0\.9\)99\.32\.6395\.061\.2100\.01\.2599\.895\.4\(0\.9, 0\.4\)98\.92\.8295\.565\.3100\.01\.25100\.096\.7Statistical Tests \(LLM Agents\)Model Tier42\.72\*\*\*21\.11\*\*\*75\.39\*\*\*3\.04\*————First Proposer3\.41†0\.125\.76\*0\.85————Buyer Type7\.12\*\*10\.75\*\*1\.242\.44————Patience6\.40†27\.95\*\*\*0\.8634\.67\*\*\*————
Notes:We use DiscEfficiency to denote discounted efficiency\. Patience rows are ordered as\(δB,δS\)\(\\delta\_\{B\},\\delta\_\{S\}\), buyer patience first and seller patience second\. Test statistics areχ2\\chi^\{2\}for Agreement andFFfor Rounds, Efficiency, and Discounted Efficiency\.p†<0\.10\{\}^\{\\dagger\}p<0\.10, \*p<0\.05p<0\.05, \*\*p<0\.01p<0\.01, \*\*\*p<0\.001p<0\.001\. Bayesian benchmark from validated implementation\. Efficiency measured as percentage of first\-best undiscounted surplus; Discounted Efficiency applies per\-round discount factors to the realized surplus\.
### 4\.1Capability Governs Value Creation
As a class, LLM agents reach near\-universal agreement \(98\.9%98\.9\\%\) and capture95\.4%95\.4\\%of first\-best surplus in undiscounted terms, but take2\.982\.98rounds against the1\.251\.25\-round Bayesian benchmark, realizing efficient allocations through iterative proposal–rejection search rather than one\-shot equilibrium play\. The first\-order result, however, is not this class\-level gap but its dependence on capability\. Efficiency rises monotonically with model tier \(flagship98\.9%98\.9\\%, mid\-tier96\.4%96\.4\\%, baseline91\.0%91\.0\\%\), and capability accounts for roughly66%66\\%of the efficiency variation the design moves, far more than any other factor \(the variance decomposition in §[4\.4](https://arxiv.org/html/2608.07538#S4.SS4)\)\. Duration, however, rises with capability rather than falling: baseline models are fastest \(2\.752\.75rounds\), mid\-tier models take2\.932\.93, and flagship models take the longest \(3\.253\.25\)\. The extra rounds are therefore not a uniform price paid for efficiency but a signature of how capable agents search: they invest in proposal–rejection cycles rather than settling at round\-one heuristic offers\. Figure[1](https://arxiv.org/html/2608.07538#S4.F1)illustrates this across all nine models\.

\(a\) Undiscounted efficiency

\(b\) Discounted efficiency
Figure 1:Efficiency and Rounds by Model TierNotes:\(a\) Undiscounted and \(b\) discounted efficiency against mean rounds to agreement, each averaged over the four patience configurations and all conditions \(consistent with the discounted\-efficiency column of Table[4](https://arxiv.org/html/2608.07538#S4.T4)\)\. The Bayesian benchmark \(star\) settles at1\.251\.25rounds, and every model lies below and to the right of it\. Discounting \(b\) penalizes additional rounds, sharply widening the spread relative to \(a\)\.
In undiscounted terms \(Figure[1](https://arxiv.org/html/2608.07538#S4.F1)a\), the nine models do not separate cleanly into tier\-based clusters; capability ordering predicts the rough region of the plot but not the within\-tier ranking\. The flagship tier occupies the high\-efficiency region nearest the benchmark’s100%100\\%: Gemini\-3\-Pro \(3\.313\.31rounds,99\.8%99\.8\\%\) is the most efficient model in our sample, GPT\-5\.2 \(2\.682\.68,98\.2%98\.2\\%\) is the fastest flagship, and Qwen3\-Max reaches comparable efficiency \(98\.7%98\.7\\%\) at the cost of the most rounds in the sample \(3\.753\.75\)\. The mid\-tier and baseline bands overlap the flagship efficiency region rather than sitting visibly below it: Gemini\-3\-Flash \(3\.113\.11,99\.1%99\.1\\%\) and the baseline Gemini\-2\.5\-Flash \(2\.452\.45,98\.7%98\.7\\%\) match or exceed the efficiency of flagships from other providers\. The baseline tier is where the dispersion appears\. GPT\-4o\-mini reaches only84\.6%84\.6\\%efficiency at3\.403\.40rounds, and Qwen2\.5\-14B recovers89\.8%89\.8\\%at2\.482\.48rounds\. Tier therefore predicts whether a model approaches the benchmark, but provider identity governs where in that region it lands; every model sacrifices some combination of speed and surplus relative to the theoretical optimum\.
Under discounting, this delay is costly, and the cost scales with patience\. In discounted terms, LLM agents forgo33\.8%33\.8\\%of first\-best surplus overall, ranging from25\.3%25\.3\\%under symmetric high patience \(δB=δS=0\.9\\delta\_\{B\}=\\delta\_\{S\}=0\.9\) to38\.8%38\.8\\%under the most buyer\-impatient configuration \(δB=0\.4,δS=0\.9\\delta\_\{B\}=0\.4,\\ \\delta\_\{S\}=0\.9\), gaps computed from the discounted\-efficiency columns of Table[4](https://arxiv.org/html/2608.07538#S4.T4)\.777Isolating the delay channel alone \(holding the allocation fixed and discounting only the excess rounds,1−δτ−1\.251\-\\delta^\{\\tau\-1\.25\}\) gives a comparable≈20%\\approx 20\\%atδ=0\.9\\delta=0\.9and≈48%\\approx 48\\%atδ=0\.7\\delta=0\.7\.The timing cost is therefore economically consequential whenever continuation costs are nontrivial, even as undiscounted coordination remains high\. Figure[1](https://arxiv.org/html/2608.07538#S4.F1)\(b\) makes the mechanism visual: once delay is priced, discounted efficiency slopes down steeply in rounds and the tier ordering scrambles\. The fast, efficient models hold the top discounted\-efficiency positions: the baseline Gemini\-2\.5\-Flash \(2\.452\.45rounds\) and the flagship GPT\-5\.2 \(2\.682\.68rounds\) reach79\.379\.3and76\.9%76\.9\\%, while the slowest flagship, Qwen3\-Max \(3\.753\.75rounds\), falls to the bottom of the discounted ranking \(50\.8%50\.8\\%\), below every baseline model despite near\-top undiscounted efficiency\. Speed, not capability alone, governs the discounted outcome, foreshadowing the provider\- and patience\-driven division analyzed in Section[5](https://arxiv.org/html/2608.07538#S5)\.
The reasoning\-effort ablation \(Appendix[11\.4](https://arxiv.org/html/2608.07538#S11.SS4); summarized in §[6\.4](https://arxiv.org/html/2608.07538#S6.SS4)\) shows that inference\-time compute shortens negotiations and raises discounted efficiency while leaving undiscounted efficiency and buyer share statistically unchanged\.
The total efficiency gap of 4\.56 percentage points decomposes into three components \(Appendix[9\.1](https://arxiv.org/html/2608.07538#S9.SS1), Table[9](https://arxiv.org/html/2608.07538#S9.T9)\)\. Suboptimal contract terms in otherwise rational agreements account for the majority of the gap \(63\.4%63\.4\\%, 2\.89 pp\)\. Failed deals account for24\.4%24\.4\\%\(1\.11 pp\), and irrational deals—agreements in which at least one party accepts terms yielding negative expected profit—account for the remaining12\.2%12\.2\\%\(0\.56 pp\)\. This distribution indicates LLM agents successfully avoid catastrophic failures but still struggle with marginal optimization, often settling for good\-enough outcomes rather than refining toward the benchmark optimum\. Detailed efficiency heterogeneity, including performance by buyer type and task difficulty variation, appears in Appendix[9\.2](https://arxiv.org/html/2608.07538#S9.SS2)\.
### 4\.2The Capability Threshold for Operational Reliability
The same capability ordering that governs efficiency governs the second margin of value creation: operational reliability\. We define an irrational agreement as one in which at least one party accepts terms yielding negative expected profit, and report its incidence by model and party in Figure[2](https://arxiv.org/html/2608.07538#S4.F2), with the tier\-level summary in Appendix[9\.5](https://arxiv.org/html/2608.07538#S9.SS5)\(Table[15](https://arxiv.org/html/2608.07538#S9.T15)\)\. Reliability is not a smooth gradient but a threshold: a near\-zero failure rate at the upper tiers set against an order\-of\-magnitude cliff at baseline\.

Figure 2:Irrationality Concentrates in Non\-Reasoning Baseline ModelsNotes:Bars show buyer\- and seller\-side irrational\-deal rates by model, with error bars; the dashed line marks the overall either\-party rate \(6\.5%6\.5\\%\)\. Unsafe contracts concentrate in the two non\-reasoning models \(GPT\-4o\-mini and Qwen2\.5\-14B\), whereas the seven reasoning models \(including the baseline\-tier Gemini\-2\.5\-Flash\) are at or near zero\.
The irrationality measure should be interpreted as a failure to implement an explicitly stated acceptance constraint, not as a pure measure of economic reasoning in isolation\. Because the prompt instructs agents not to accept negative\-profit agreements, such outcomes may reflect payoff miscalculation, failed constraint checking, or noncompliance with the acceptance rule\. Although these channels differ diagnostically, they are operationally equivalent: each leads the delegated agent to accept a contract the principal instructed it to reject\.
Irrationality rates vary sharply with capability:0\.0%0\.0\\%for flagship,0\.6%0\.6\\%for mid\-tier, and19\.2%19\.2\\%for baselines;97\.1%97\.1\\%of all irrational agreements involve a baseline model\. Economically unsafe contracts are overwhelmingly a baseline\-tier phenomenon, pointing to a capability threshold for autonomous prompt\-only deployment\. Upper\-tier failure rates are within rounding distance of zero, while baseline agents fail at a rate untenable in any high\-stakes setting\. Prompt\-only safeguards therefore suffice at the upper tiers, but baseline deployment requires a more capable model or external verification of the acceptance constraint\.
This baseline rate is itself concentrated along a sharp line: whether the model is a reasoning model\. The twonon\-reasoningbaselines account for nearly all unsafe contracts \(GPT\-4o\-mini,36\.2%36\.2\\%overall and33\.9%33\.9\\%on the buyer side; Qwen2\.5\-14B,20\.0%20\.0\\%overall and10\.0%10\.0\\%on the seller side\), whereas the one baseline model with reasoning capability, Gemini\-2\.5\-Flash, sits at2\.9%2\.9\\%, far closer to the mid\-tier and flagship models, all of which are reasoning models, than to the other two baseline models \(per\-model breakdown in Figure[2](https://arxiv.org/html/2608.07538#S4.F2)\)\. The operative threshold for prompt\-only deployment is therefore associated with the presence of reasoning capability rather than capability tier as such\. The reasoning\-effort ablation of Appendix[11\.4](https://arxiv.org/html/2608.07538#S11.SS4)refines this distinction: increasing the reasoning budget within GPT\-5\.2 changes negotiation speed and discounted efficiency, but not deal rates, undiscounted efficiency, or surplus division\. Because our two non\-reasoning models are also the earliest\-generation, we read reasoning capability as a refinement of the capability threshold rather than a separately identified axis\.
Buyers suffer irrationality more than sellers \(5\.1%5\.1\\%vs\.1\.4%1\.4\\%\) despite their informational advantage, plausibly because seller feasibility is a deterministic cost\-coverage check while buyer feasibility requires expected\-profit evaluation under demand uncertainty\. Low\-type buyers err more than high\-type \(6\.3%6\.3\\%vs\.3\.9%3\.9\\%\), reflecting tighter low\-demand margins\. Both error types occur alongside confident strategic language: revenue miscalculation \(79\.0%79\.0\\%, 109 of 138 cases; buyers accepting payments above expected revenue\) and cost miscalculation \(21\.0%21\.0\\%, 29 of 138; sellers accepting payments below cost\)\. These reinforce the value of automated profit verification as a deployment guardrail \(whatever the source\), developed in Section[7](https://arxiv.org/html/2608.07538#S7)\. The gradient also tracks a reliability story: each round compounds process\-failure probability, so baseline outcomes are strongest under quick termination while flagships’ near\-zero failure rate makes extended exploration productive\.
### 4\.3Why Agents Take Extra Rounds: Heuristic Substitution
The round\-count gap documented in §[4\.1](https://arxiv.org/html/2608.07538#S4.SS1)has a clear process\-level signature: LLM agents recognize the structure of the bargaining problem but substitute robust heuristics for the branch\-specific execution the PBE prescribes\. The first signature appears in opening offers\. Buyers who move first reproduce theFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)truth\-telling type separation \(high\-type quantities average77\.877\.8units, low\-type41\.641\.6, with opening\-to\-final correlations of0\.9110\.911for quantity and0\.8620\.862for payment\), whereas sellers who open propose60\.560\.5units on average \(essentially the uniform\-prior expected value of6060rather than the branch\-specific screening menu\), driving the2\.982\.98\-round average against the1\.251\.25\-round equilibrium prediction\.
A second signature concerns type\-strategic intent\. After aggregating the turn\-level classifications within each negotiation, the average signaling rate per negotiation is64\.7%64\.7\\%for low\-type buyers versus21\.9%21\.9\\%for high\-type buyers, while the average mimic rate per negotiation is13\.8%13\.8\\%for high\-type buyers versus1\.3%1\.3\\%for low\-type buyers\. Both differences remain significant after Holm correction \(p<0\.001p<0\.001; Appendix[10\.3](https://arxiv.org/html/2608.07538#S10.SS3)\)\. The coded reasoning follows the pattern predicted byFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\): low\-type buyers signal their type more often to separate, whereas high\-type buyers mimic the low type more often to avoid surplus extraction\. Consistent with this interpretation, high\-type realized quantities are7\.1%7\.1\\%below first\-best, although this quantity distortion alone does not establish mimicking\. The comparative static, however, runs the wrong way: the average mimic rate islowest, not highest, under low buyer patience, so high\-type buyers do not calibrate mimicking to the patience incentive on which the equilibrium turns\. The full opening\-offer breakdown, inferred\-strategy distributions, and extended reasoning traces appear in Appendix[10\.5](https://arxiv.org/html/2608.07538#S10.SS5)\(with further detail in Appendix[10](https://arxiv.org/html/2608.07538#S10)\)\.
### 4\.4Capability Creates, Patience Divides: Orthogonal Drivers
We next ask whether the same design variables govern value creation and value division\. Table[5](https://arxiv.org/html/2608.07538#S4.T5)attributes explained variance in efficiency and buyer surplus share to four experimental blocks: buyer type, first proposer, patience, and model tier\. Interaction and within\-stratum tests \(Appendix[9\.1](https://arxiv.org/html/2608.07538#S9.SS1)\) verify the table’s diagonal pattern is not driven by a few cells\.
Table 5:Variance Attribution: Efficiency and Buyer Surplus ShareEfficiencyBuyer surplus shareR2R^\{2\}% expl\.R2R^\{2\}% expl\.Variance attribution \(ShapleyR2R^\{2\}\)Information \(Buyer type\)0\.01927%0\.0032%Bargaining Structure \(First Proposer\)0\.0000%0\.0000%Time Pressure \(Patience\)0\.0057%0\.13890%Capability \(Model tier\)0\.04766%0\.0138%Total variance explained0\.0710\.154NN1,9981,998
Notes:Sample:N=1,998N=1\{,\}998economically rational agreements; irrational deals are excluded because buyer surplus share is ill\-defined when one party accepts a negative\-profit contract\. Entries report ShapleyR2R^\{2\}by design block and each block’s share of explained variance within the outcome\. Interaction tests and within\-stratum coefficient estimates appear in Appendix[9\.1](https://arxiv.org/html/2608.07538#S9.SS1)\.
The split is stark\. Capability explains66%66\\%of the explained variation in efficiency but only8%8\\%of buyer\-share variation; patience explains90%90\\%of buyer\-share variation but only7%7\\%of efficiency variation\. Buyer type matters for efficiency because demand information affects the operational quantity decision, but it has little direct role in surplus division once patience and model tier are accounted for\. The appendix diagnostics confirm the same diagonal pattern: capability moves efficiency within patience strata but not buyer share, while patience moves buyer share within capability tiers but not efficiency\. Thus capability creates value, patience divides it\.
### 4\.5Within\-Family Cross\-Model Negotiations
The driver separation documented above is established with homogeneous model pairs\. Pairing models of different capability tiers within the same family stress\-tests that finding and also reveals additional patterns worth distilling on their own\. We pair GPT\-5\.2 \(flagship\) with GPT\-5\-mini \(mid\-tier\) in both role assignments\. Each configuration covers the full 16\-condition factorial \(22buyer types×\\times22first proposers×\\times44patience levels\) with 15 replications\. We conduct the two cross\-tier configurations in the verbal format \(N=480N=480\) and reuse the two homogeneous configurations from M1 as benchmarks; Table[17](https://arxiv.org/html/2608.07538#S9.T17)therefore displays 960 observations\. Four findings emerge, grouped below\.
Efficiency is robust to tier mixing, and capability manifests as strategic delay rather than speed\.Cross\-tier pairs reach agreement in every negotiation \(100% deal rate\), and undiscounted efficiency is statistically indistinguishable from homogeneous pairs \(96\.8% vs\. 96\.1%,p=0\.284p=0\.284\): high allocative efficiency survives capability heterogeneity even when one party has weaker newsvendor reasoning, so firms can mix capability tiers across buyer/seller sides without paying an efficiency penalty\. The cleanest separation between GPT\-5\.2 and GPT\-5\-mini is on the seller side: GPT\-5\.2 as seller concedes only35\.2%35\.2\\%of surplus to a GPT\-5\-mini buyer, versus52\.3%52\.3\\%when GPT\-5\-mini sells to a GPT\-5\.2 buyer\. GPT\-5\.2 achieves this by holding out0\.430\.43rounds longer as seller \(2\.462\.46vs\.2\.032\.03;p<0\.01p<0\.01\)\. The dollar gap is correspondingly large \(seller profit$931\\mathdollar 931vs\.$680\\mathdollar 680;Δ=\+251\\Delta=\+251,p<0\.001p<0\.001; Appendix[9\.6](https://arxiv.org/html/2608.07538#S9.SS6)\)\. The flagship thus converts its additional rounds into materially higher seller profit rather than faster closure, so within\-family capability operates through extended willingness to delay, not accelerated convergence, localizing the aggregate finding that flagship models’ longer negotiations reflect patience\-driven extraction by the stronger party, not slower computation\.
Within a family the more capable model claims a larger share in either role, but patience remains the dominant lever of division\.Capability translates into bargaining power in both roles\. With GPT\-5\.2 as seller, switching the buyer from GPT\-5\.2 to the weaker GPT\-5\-minilowersthe buyer’s share \(37\.9%37\.9\\%to35\.2%35\.2\\%\); conversely, with GPT\-5\-mini as seller, a GPT\-5\.2 buyer captures52\.3%52\.3\\%against GPT\-5\-mini’s41\.9%41\.9\\%self\-play baseline\. The more capable model thus extracts more whether it buys or sells, and earns more in dollars in both roles \(Appendix[9\.6](https://arxiv.org/html/2608.07538#S9.SS6)\), so within a family capability rank does predict who captures surplus\. Across providers, however, capability rank alone does not determine who captures surplus \(Section[5](https://arxiv.org/html/2608.07538#S5)\)\. Patience nonetheless remains the dominant driver of surplus division: buyer shares range from about67%67\\%under buyer\-patient conditions \(δB=0\.9\\delta\_\{B\}=0\.9,δS=0\.4\\delta\_\{S\}=0\.4\) to about26%26\\%under seller\-patient conditions \(δB=0\.4\\delta\_\{B\}=0\.4,δS=0\.9\\delta\_\{S\}=0\.9\)\. The within\-family capability range \(∼\\sim17 pp\) sits well below the patience range \(∼\\sim40 pp\), and provider\-role assignment shifts buyer shares by an amount comparable to within\-family capability differences \(Section[5\.2](https://arxiv.org/html/2608.07538#S5.SS2)\)\. Full per\-configuration statistics and the direction\-effect visualization appear in Appendix[9\.7](https://arxiv.org/html/2608.07538#S9.SS7)\(Table[17](https://arxiv.org/html/2608.07538#S9.T17), Figure[6](https://arxiv.org/html/2608.07538#S9.F6)\)\.
## 5Provider\-Level Bargaining Profiles
For firms deploying LLM agents on different sides of a procurement interaction, the first\-order distributional question is whether the vendor they select systematically favors one role\. We show that it does, and that provider identity predicts the direction of surplus flow more reliably than capability tier\.
Table[6](https://arxiv.org/html/2608.07538#S5.T6)reveals the heterogeneity, reporting buyer surplus as a share of total realized surplus\. Alibaba’s Qwen models exhibit extreme buyer bias, with buyer shares ranging from52\.7%52\.7\\%to91\.4%91\.4\\%across capability tiers\. Google’s Gemini models cluster near parity, ranging from44\.3%44\.3\\%to53\.9%53\.9\\%\. OpenAI models are mildlyseller\-leaning across all three tiers, most pronounced at the flagship \(37\.9%37\.9\\%\) and closer to parity at the mid\-tier \(41\.9%41\.9\\%\) and baseline \(39\.8%39\.8\\%\)\. Four of nine models show buyer shares below50%50\\%, and the pooled53\.5%53\.5\\%average masks three distinct provider profiles\.
Table 6:Surplus Division by Provider: Buyer Share Across Capability TiersProviderFlagshipMid\-TierBaselineProvider AvgOpenAI0\.3790\.4190\.3980\.399Google0\.5390\.5140\.4430\.499Alibaba0\.9140\.6680\.5270\.703Pooled0\.535This provider\-level heterogeneity has two implications\. First, role\-based surplus division is not a universal property of LLM negotiation: the three providers differ qualitatively, not merely in degree\. These family\-level patterns are consistent with provider\-specific training and fine\-tuning differences, but the current design does not identify which channel produces them\. Second, organizations selecting LLM agents for procurement must evaluate provider\-specific bargaining behavior, not merely model capability: deploying a Qwen model as buyer and an OpenAI model as seller would produce markedly different distributional outcomes than the reverse configuration, a prediction the cross\-family experiments of Section[5\.2](https://arxiv.org/html/2608.07538#S5.SS2)confirm directly\.
A second pattern layered on top of provider variation is universal compression toward equal division relative to the Bayesian benchmark\. Under strong seller patience \(δB=0\.4,δS=0\.9\\delta\_\{B\}=0\.4,\\delta\_\{S\}=0\.9\), the benchmark allocates approximately13%13\\%of surplus to the buyer, yet LLM buyers achieve40\.2%40\.2\\%, a\+27\.1\+27\.1pp deviation toward the buyer\. Under strong buyer patience \(δB=0\.9,δS=0\.4\\delta\_\{B\}=0\.9,\\delta\_\{S\}=0\.4\), the benchmark allocates approximately91%91\\%, yet LLM buyers achieve only70\.9%70\.9\\%, a−20\.5\-20\.5pp deviation toward the seller\. All nine models attenuate predicted extremes, pulling outcomes toward roughly even division\. First\-mover identity has little effect on the split: buyers capture 54\.1% proposing first versus 52\.8% proposing second, and even seller\-first buyers exceed an equal split, indicating advantage well beyond procedural positioning\. Surplus division across patience, tier, first\-proposer, and buyer\-type cuts appears in Appendix[9\.3](https://arxiv.org/html/2608.07538#S9.SS3)\(Table[12](https://arxiv.org/html/2608.07538#S9.T12)\)\.
### 5\.1Interpreting Provider\-Level Heterogeneity
A universal RLHF\-based explanation does not fit: a common buyer\-favoring tendency from post\-training would produce broadly similar directional effects, yet Qwen exhibits large buyer advantages, Gemini clusters near parity, and OpenAI spans near\-parity to seller\-leaning\. The cleaner reading is comparative rather than causal \(under a common prompting regime, provider identity is associated with distinct bargaining profiles\), with three descriptive correlates surfaced by the cross\-family evidence below: anchoring discipline, concession posture, and how explicitly models track patience asymmetries\. Several channels could generate such profiles \(training\-data composition, reward criteria, safety tuning, annotation guidelines\), and isolating one would require varying post\-training while holding base capability fixed; we treat the profiles as descriptive, and the managerial implication is to select providers whose profile aligns with strategic objectives\.
The verbal channel reinforces this comparative reading: an LLM classifier flags public–private divergence \(public fairness appeals paired with self\-interested private reasoning\) as a population\-average behavior whose prevalence is itself provider\- and role\-conditional\. Using buyer\-minus\-seller differences, Google buyers have lower average within\-negotiation rates of both fairness mismatch \(−9\.3\-9\.3pp\) and profit hiding \(−3\.0\-3\.0pp\) than Google sellers, whereas the OpenAI gaps are small or nonsignificant and the Qwen gaps run in the opposite direction \(\+8\.7\+8\.7and\+3\.3\+3\.3pp; full breakdown in Appendix[10\.6](https://arxiv.org/html/2608.07538#S10.SS6), Table[24](https://arxiv.org/html/2608.07538#S10.T24)\)\.
The same evidence argues against a simple computational\-asymmetry account in which buyers underperform because newsvendor evaluation is harder than cost checking: buyers show higher irrationality than sellers, yet aggregate buyer surplus still exceeds the Bayesian reference for most models, and buyer\-first negotiations yield the largest buyer shares\.
### 5\.2Cross\-Family Flagship Negotiations
While our main analysis pairs each model with itself, this subsection considers asymmetric LLM deployment by crossing provider boundaries entirely, pairing flagship models from different AI families against one another: GPT\-5\.2 \(OpenAI\), Gemini\-3\-Pro \(Google\), and Qwen3\-Max \(Alibaba\)\. Each pair is tested in both directions \(each model serving as both buyer and seller\), yielding six directional configurations\. Each configuration covers 16 conditions \(2 buyer types×\\times2 first proposers×\\times4 patience levels\), with 15 replications per condition\. All experiments use the verbal treatment with medium reasoning effort\. The total sample is1,4401\{,\}440attempted negotiations, all of which reached agreement\.
This analysis connects directly to the provider\-specific bargaining profiles documented above and asks what happens when these profiles collide at the bargaining table\. The within\-family results of Section[4\.5](https://arxiv.org/html/2608.07538#S4.SS5)already imply that a model’s bargaining advantage isrelational, not intrinsic: GPT\-5\.2’s 37\.9% self\-play buyer share reflects the mirrored equilibrium of two agents sharing the same strategic tendencies, not a fixed property of the model in the buyer role\. Cross\-family matchups test how provider dispositions combine when they are not mirrored\.
Cross\-family flagship pairs reach agreement in all negotiations: GPT↔\\leftrightarrowGemini achieves a 100% deal rate \(N=480N=480\), GPT↔\\leftrightarrowQwen achieves 100% \(N=480N=480\), and Gemini↔\\leftrightarrowQwen achieves 100% \(N=480N=480\); the full performance statistics across all six directional configurations appear in Appendix[9\.8](https://arxiv.org/html/2608.07538#S9.SS8)\(Table[18](https://arxiv.org/html/2608.07538#S9.T18)\)\.
Negotiation outcomes are strikingly asymmetric across providers \(Figure[3](https://arxiv.org/html/2608.07538#S5.F3)\)\. Gemini\-3\-Pro is near parity in self\-play but strongest in heterogeneous pairings \(66\.7% of surplus as buyer, 47\.9% as seller\); GPT\-5\.2 occupies the middle \(54\.7% / 43\.3%\); and Qwen3\-Max is weakest as a cross\-family seller, retaining only 27\.3%\. The two designs measure different objects: self\-play asks what division looks like when both sides share a profile, cross\-family play which profile is stronger when they differ\. A provider can thus be near parity in self\-play \(its anchoring and concession discipline mirrored by an identical counterparty\) yet dominate heterogeneous matchups when those traits are not reciprocated\. This reconciles two mirror\-image patterns: Gemini’s aggressive anchoring, which self\-play offsets but which is decisive against softer GPT or Qwen counteroffers, and Qwen’s buyer\-favoring self\-play profile, which conceals a seller\-side softness—costly whenever Qwen’s counterparty does not share that softness\. Qwen sellers retain only 33\.6% against GPT\-5\.2 buyers and 21\.0% against Gemini buyers\.

Figure 3:Direction Effects in Cross\-Flagship NegotiationsNotes:Panel \(a\) shows each model’s own surplus share as buyer \(solid\) vs\. seller \(hatched\)\. Panel \(b\) shows buyer share by direction within each pair\. Panels \(a\) and \(b\) show significance brackets from independent two\-samplett\-tests comparing the displayed role or direction groups\. Error bars are 95% confidence intervals\. Significance levels:p∗<0\.05\{\}^\{\*\}\\,p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}\\,p<0\.01,p∗∗∗<0\.001\{\}^\{\*\*\*\}\\,p<0\.001\.
Direction effects and the direction\-effect hierarchy\.The direction of the pairing produces large, statistically significant effects on surplus division\. In GPT↔\\leftrightarrowGemini negotiations, buyer share swings 11\.3 pp \(54\.3% when GPT sells to Gemini versus 43\.0% when Gemini sells to GPT;t=6\.01t=6\.01,p<0\.001p<0\.001\), with Gemini capturing more surplus regardless of role\. The asymmetry is largest for the Qwen↔\\leftrightarrowGemini pair: the buyer captures 79\.0% when Qwen sells but only 61\.2% when Gemini sells, a 17\.8 pp swing \(t=12\.25t=12\.25,p<0\.001p<0\.001\) that compounds Gemini’s strength and Qwen’s weakness; GPT↔\\leftrightarrowQwen shifts 7\.3 pp by direction \(t=3\.09t=3\.09,p<0\.01p<0\.01\)\. These cross\-family direction effects \(7–18 pp\) are comparable to the within\-family capability swing \(∼\\sim17 pp\) from Section[4\.5](https://arxiv.org/html/2608.07538#S4.SS5), and both sit well below the patience swings \(∼\\sim35 pp\), so provider choice and capability tier are comparable levers, each second\-order to patience\. Patience nevertheless dominates: buyer shares reach 80\.6% under buyer\-patient conditions \(δB=0\.9\\delta\_\{B\}=0\.9,δS=0\.4\\delta\_\{S\}=0\.4\) versus 46\.1% under seller\-patient conditions \(δB=0\.4\\delta\_\{B\}=0\.4,δS=0\.9\\delta\_\{S\}=0\.9\), tracking the Bayesian benchmarks of 91\.4% and 13\.1% directionally but with the characteristic compression toward equal division \(−10\.8\-10\.8pp and\+33\.0\+33\.0pp deviations\)\.
Efficiency: Qwen’s inefficiency concentrates in the seller role\.Undiscounted efficiency exceeds 95% for every pair \(GPT↔\\leftrightarrowGemini 98\.7%, Gemini↔\\leftrightarrowQwen 96\.5%, GPT↔\\leftrightarrowQwen 95\.1%\), so cross\-family pairing does not meaningfully reduce allocative efficiency\. Decomposing the gap, GPT↔\\leftrightarrowGemini loses only 1\.25 pp \(zero deal failures, zero irrational agreements\), GPT↔\\leftrightarrowQwen 4\.88 pp, and Gemini↔\\leftrightarrowQwen 3\.50 pp, almost all of it from suboptimal contract terms rather than failed or loss\-making deals\. The direction\-level pattern is unambiguous: when Qwen sells, efficiency drops sharply \(Qwen→\\toGPT 91\.9%, an 8\.1 pp gap; Qwen→\\toGemini 94\.0%, a 6\.0 pp gap\), whereas GPT or Gemini selling to Qwen stays above 98%\. All three irrational agreements in the sample occur with Qwen as seller \(two in Qwen→\\toGPT, one in Qwen→\\toGemini, out of240240deals each\); Qwen’s seller\-side inefficiency, not its counterparty, is the common thread\.
Process correlates\.Reasoning traces show recurring differences behind this ordering: Gemini\-3\-Pro articulates the patience structure directly and anchors aggressively \(claiming 82–92% of surplus on the first offer, near\-efficientq≈80q\\approx 80at 4–6% above cost\), GPT\-5\.2 reasons in terms of fair splits and cost coverage, and Qwen3\-Max exposes little reasoning and defaults to passive cost\-plus counter\-offers as seller, with concession discipline mirroring the same ranking\. Counterintuitively, Gemini is descriptively theleastBayesian\-aligned \(mean buyer\-side deviation\+\+21\.5 pp vs\.\+\+9\.5 for GPT and\+\+15\.0 for Qwen\): its advantage reflects not closer alignment but a consistent ability to extract surplus above the benchmark\. Because all three are flagships, the safer reading is that provider identity carries distinct bargaining profiles in heterogeneous matchups \(the cross\-family analogue of the within\-family pattern of Section[4\.5](https://arxiv.org/html/2608.07538#S4.SS5)\), so which model represents each party can systematically shift the distribution of economic value\. Supporting figures appear in Appendix[9\.8](https://arxiv.org/html/2608.07538#S9.SS8)\(Figures[7](https://arxiv.org/html/2608.07538#S9.F7)and[8](https://arxiv.org/html/2608.07538#S9.F8)\)\.
## 6Design Choices in Deploying LLM Bargaining Agents
A firm deploying an LLM bargaining agent does not bargain directly; it configures an agent and delegates the task, a principal–agent relationship in which the prompt functions as the contract\(Imas et al\.[2025](https://arxiv.org/html/2608.07538#bib.bib23)\)\. The principal selects which messages the agent may exchange, whether discounting is encoded in the prompt, and what patience parameters to assign\. Our main experiment fixed these at a single specification \(verbal channel enabled, discounting prompted, patience assigned\), so the patterns of Sections[4](https://arxiv.org/html/2608.07538#S4)and[5](https://arxiv.org/html/2608.07538#S5)are silent on how behavior shifts elsewhere in the configuration space\. This section examines the three dimensions in turn \(the verbal channel, the discounting framework, and the focal agent’s patience\) and asks of each the same question: do alternative specifications produce systematically different outcomes that a principal should anticipate?
### 6\.1Does the Verbal Channel Do Strategic Work? \(R1\)
The first configuration dimension is the message protocol\. Production LLM bargaining agents generate natural\-language messages by default, yet the contract itself is numerical, so we ask whether language moves outcomes or merely accompanies offer behavior\. We test this by re\-running the main factorial with the verbal channel disabled: agents exchange only numeric offers and accept/reject decisions, with all other features held constant\. This isolates whether the provider profiles and high\-efficiency\-with\-delay pattern documented above live in linguistic persuasion or in the numerical proposals themselves\. The question also connects to evidence from open\-ended AI negotiation that linguistic style affects agreement and value creation\(Vaccaro et al\.[2026](https://arxiv.org/html/2608.07538#bib.bib39)\); our setting asks whether language carries comparabledistributionalweight once offers are structured and surplus is objective\.
Mirroring the main analysis performance table, structured agents achieve 96\.4% agreement \(vs\. 98\.9% verbal\), 3\.15 rounds \(vs\. 2\.98\), and 92\.8% undiscounted efficiency \(vs\. 95\.4%\)\. The high\-efficiency\-with\-delay pattern therefore reproduces, but the protocols are not statistically identical: removing verbal communication modestly lowers agreement and efficiency and slightly lengthens bargaining\. The capability gradient also replicates: flagship structured agents achieve 98\.3% efficiency at 3\.23 rounds; baseline agents manage only 87\.7% at 2\.81 rounds\. The full performance table by tier, proposer, buyer type, and patience, with between\-treatment tests, appears in Appendix[11\.2](https://arxiv.org/html/2608.07538#S11.SS2)\(Table[27](https://arxiv.org/html/2608.07538#S11.T27)\)\.
##### Provider\-Specific Bias Persists Without Communication\.
The provider\-specific surplus division patterns documented in the main analysis survive the removal of verbal communication\. This rules out a purely language\-based explanation for the provider profiles: the distributional regularities also appear in how models generate and evaluate numerical proposals\. This qualifies the language\-centric account of AI negotiation for structured\-offer settings: whereVaccaro et al\. \([2026](https://arxiv.org/html/2608.07538#bib.bib39)\)find linguistic warmth driving outcomes in open\-ended dialogue, our distributional regularities persist in numerical offer behavior even when language is removed\. As the next paragraph shows, however, the verbal channel remains an execution aid and a provider\-specific distributional lever\.
##### Model\-Specific Heterogeneity\.
Pooled buyer share changes little, but model\-level distributional shifts diverge: removing communication shifts OpenAI’s baseline and mid\-tier models toward buyers and Gemini’s mid\-tier and flagship models toward sellers\. The remaining five models, GPT\-5\.2, Gemini\-2\.5\-Flash, Qwen2\.5\-14B, Qwen3\-32B, and Qwen3\-Max, show no detectable buyer\-share shift \(full model\-by\-model breakdown in Appendix[11\.2](https://arxiv.org/html/2608.07538#S11.SS2), Table[28](https://arxiv.org/html/2608.07538#S11.T28)\)\. The managerial reading is that the verbal channel is itself a provider\-specific lever: for OpenAI’s baseline and mid\-tier models, language moderates the buyer advantage, whereas for Gemini’s mid\-tier and flagship models it amplifies it, so constraining communication is not a uniformly buyer\- or seller\-favoring intervention\.
### 6\.2No\-Discounting Bargaining \(R2\)
The second configuration dimension is whether discounting is encoded in the agent’s prompt at all\. Our main experiment presents agents with an explicit utility formulaUk=πkδkτ−1U\_\{k\}=\\pi\_\{k\}\\delta\_\{k\}^\{\\tau\-1\}and informs them of the patience parameters that govern it\. This framing supports clean comparison against the Feng et al\. equilibrium but it does not match how production systems frame the task: systems such as Pactum and Arkestro present agents with contractual terms, permissible ranges, and fallback rules, but do not parameterize the agent’s utility with a time\-preference coefficient\. R2 therefore tests how LLM agents bargain when the discounting framework is removed from the prompt entirely, comparing the symmetric\-patience baseline with a no\-discounting treatment that removesδ\\deltafrom both the prompt and payoff computation \(N=540N=540per treatment\)\. Both treatments are evaluated on a common undiscounted basis along the three structural margins of the main analysis: high efficiency with delayed agreement, the cross\-provider distributional profile, and the capability–irrationality gradient\. The full performance comparison appears in Appendix[11\.3](https://arxiv.org/html/2608.07538#S11.SS3)\(Table[29](https://arxiv.org/html/2608.07538#S11.T29)\)\.
Removing the prompted discount factor leaves the main outcome regularities intact but lengthens bargaining\. The high\-efficiency\-with\-delay pattern reproduces: efficiency is statistically indistinguishable across treatments and within every tier, while rounds rise from 3\.39 to 4\.02 overall and increase in each capability tier\. The cross\-provider distributional ordering also reproduces despite a marginal Treatment×\\timesProvider interaction \(F=2\.44F=2\.44,p=0\.087p=0\.087; Qwen≫\\ggGoogle\>\>OpenAI\)\. And the capability\-dependent irrationality gradient reproduces, declining monotonically from baseline through mid\-tier to flagship\. With discounting removed we compare raw buyer shares against the symmetric\-patience baseline rather than the Feng et al\. benchmark, whose patience\-based surplus division is no longer defined\. Full design, panel\-by\-panel tests, and the cuts by first\-proposer role and buyer type appear in Appendix[11\.3](https://arxiv.org/html/2608.07538#S11.SS3)\(Table[30](https://arxiv.org/html/2608.07538#S11.T30)\)\.
### 6\.3Prompted Patience as a Principal\-Specification Lever \(R3\)
The third configuration dimension is the patience level the principal assigns to the agent: a continuous parameter, unlike the present\-or\-absent verbal\-channel and discounting choices\. Section[3\.3](https://arxiv.org/html/2608.07538#S3.SS3)introduced the strategic/economic patience distinction this dimension rests on: under delegation to LLM agents the promptedstrategic patienceδstrat\\delta^\{\\text\{strat\}\}\(the agent’s per\-round discount, on which the equilibrium is defined\) decouples fromeconomic patienceδecon\\delta^\{\\text\{econ\}\}\(the principal’s fixed real\-time cost of delay\), so the principalchoosesδstrat\\delta^\{\\text\{strat\}\}whileδecon\\delta^\{\\text\{econ\}\}stays a constraint\. The specification problem is to setδstrat\\delta^\{\\text\{strat\}\}to maximize realized payoff given the counterparty’s choice and the principal’sδecon\\delta^\{\\text\{econ\}\}\. This exercise is conditional rather than game\-theoretic: we vary the focal agent’s prompted patience while holding the counterparty environment fixed, and do not solve a two\-sided meta\-game in which both principals jointly choose prompt parameters or agents infer undisclosed patience\. The remainder of this subsection asks whether varyingδstrat\\delta^\{\\text\{strat\}\}in our data produces outcome differences a principal could anticipate\.
We vary focal\-agent strategic patience across\{0\.9,0\.7,0\.4\}\\\{0\.9,\\,0\.7,\\,0\.4\\\}separately for the buyer and seller roles, holding the counterparty at the maximally demandingδstrat=0\.9\\delta^\{\\text\{strat\}\}=0\.9; we further disaggregate the buyer side by type \(H vs\. L\), since the Feng et al\. equilibrium predicts the high type’s information rent and the low type’s signaling cost respond to own patience along different margins\. Two questions organize the analysis\.Part 1: does varyingδstrat\\delta^\{\\text\{strat\}\}produce systematic outcome differences at all?Part 2, if yes: which way should the principal pull the lever, and does the answer differ by role or buyer type? For each \(model, buyer\-type, first\-proposer\) cell we compute realized payoff at the principal’s true economic patience, identify the empirically best promptedδstrat\\delta^\{\\text\{strat\}\}, and definegain=best payoff−matched\-patience payoff≥0\\text\{gain\}=\\text\{best payoff\}\-\\text\{matched\-patience payoff\}\\geq 0, with the matched benchmark settingδstrat=δecon\\delta^\{\\text\{strat\}\}=\\delta^\{\\text\{econ\}\}\. Headline numbers evaluate atδecon=0\.9\\delta^\{\\text\{econ\}\}=0\.9across99models×\\times44buyer\-type/first\-proposer conditions; the full results, by party and economic\-patience level, appear in Appendix[10\.7](https://arxiv.org/html/2608.07538#S10.SS7)\(Table[25](https://arxiv.org/html/2608.07538#S10.T25)\)\.
R3 adds the\(δB,δS\)=\(0\.9,0\.7\)\(\\delta\_\{B\},\\delta\_\{S\}\)=\(0\.9,0\.7\)configuration:99models×\\times22buyer types×\\times22proposer orders×\\times1515replications, or 540 new negotiations\. Combined with 2,160 reused M1 observations, the full analysis sample isN=2,700N=2\{,\}700\.
##### Part 1: Is strategic patience a lever?
Both headline rows of Table[7](https://arxiv.org/html/2608.07538#S6.T7)show that the agent’s strategic patience is a payoff\-relevant choice for the principal\.888We conduct separate stratified permutation tests for each role and economic\-patience level, permuting prompted\-patience labels within each model×\\timesbuyer\-type×\\timesfirst\-proposer stratum while preserving 15 replications per level\. All four tests reject equality of mean payoffs acrossδstrat∈\{0\.4,0\.7,0\.9\}\\delta^\{\\mathrm\{strat\}\}\\in\\\{0\.4,0\.7,0\.9\\\}\(19,999 permutations; Holm\-adjustedp<0\.001p<0\.001throughout\)\.Descriptively, the cell\-level mean potential in\-sample gain is\+74\.81\+74\.81on the buyer side \(95% bootstrap interval\[\+29\.43,\+130\.17\]\[\+29\.43,\\,\+130\.17\]\) and\+18\.28\+18\.28on the seller side \(\[\+4\.35,\+35\.06\]\[\+4\.35,\\,\+35\.06\]\)\. The payoff curves are not flat inδBstrat\\delta^\{\\text\{strat\}\}\_\{B\}\(Appendix[10\.7](https://arxiv.org/html/2608.07538#S10.SS7), Figure[12](https://arxiv.org/html/2608.07538#S10.F12)\)\. Strategic patience is a lever; the principal cannot ignore it\.
##### Part 2: Which value of strategic patience should the principal specify?
For buyer principals, the guidance is more heterogeneous than a single rule\. Matching \(δBstrat=δecon\\delta^\{\\text\{strat\}\}\_\{B\}=\\delta^\{\\text\{econ\}\}, here0\.90\.9\) is the most common per\-condition optimum, winning20/3620/36\(56%56\\%\) cells, and remains the safe default\. But it is the average\-best choice for only4/94/9models \(GPT\-5\.2 and the three Gemini models\); the other five do better understating patience, and the gains are large: GPT\-4o\-mini realizes\+278\+278payoff units and GPT\-5\-mini\+234\+234atδBstrat=0\.7\\delta^\{\\text\{strat\}\}\_\{B\}=0\.7\(Figure[13](https://arxiv.org/html/2608.07538#S10.F13), Panel A\)\. The mechanism in the payoff curves depends on who proposes first \(Figure[12](https://arxiv.org/html/2608.07538#S10.F12)\)\. When the buyer proposes first, the H\-type payoff rises sharply fromδBstrat=0\.4\\delta^\{\\text\{strat\}\}\_\{B\}=0\.4to0\.70\.7and then flattens, while the L\-type payoff continues rising through the matched value of0\.90\.9\. When the seller proposes first, both buyer types peak atδBstrat=0\.7\\delta^\{\\text\{strat\}\}\_\{B\}=0\.7, so maximal announced patience lowers realized payoff in those cells\. The stakes are larger in the H\-type configuration, where both payoffs and the spread acrossδBstrat\\delta^\{\\text\{strat\}\}\_\{B\}exceed the L\-type’s\.
For seller principals, by contrast, matching is robust\. The matched choice \(δSstrat=δecon\\delta^\{\\text\{strat\}\}\_\{S\}=\\delta^\{\\text\{econ\}\}, here0\.90\.9\) is per\-condition optimal in26/3626/36\(72%72\\%\) cells and is the average\-best prompted patience for all9/99/9models, with no model improving on it at the model\-averaged level \(Figure[13](https://arxiv.org/html/2608.07538#S10.F13), Panel B\)\. The seller recommendation is therefore unambiguous at this economic patience: match\. The residual mean gain of\+18\.28\+18\.28\(Table[7](https://arxiv.org/html/2608.07538#S6.T7)\) reflects scattered single\-condition cells in which a strategic alternative edges out matching, not a systematic model\-level pattern: no seller model, baseline tier included, gains from strategic understatement on average\.
Atδecon=0\.7\\delta^\{\\text\{econ\}\}=0\.7\(Table[7](https://arxiv.org/html/2608.07538#S6.T7), bottom two rows\), the lever finding remains, but the preferred direction changes\. The mean gains remain positive \(\+41\.11\+41\.11buyer,\+66\.53\+66\.53seller\), with13/3613/36buyer cells and11/3611/36seller cells matched\-optimal\. Buyer guidance is model\-dependent: matching is the average\-best choice for4/94/9models, while three models \(GPT\-5\.2, Gemini\-2\.5\-Flash, and Qwen3\-Max\) favor overstatement toδBstrat=0\.9\\delta^\{\\text\{strat\}\}\_\{B\}=0\.9and the two Gemini\-3 models favor understatement toδBstrat=0\.4\\delta^\{\\text\{strat\}\}\_\{B\}=0\.4\. Seller guidance reverses:δSstrat=0\.9\\delta^\{\\text\{strat\}\}\_\{S\}=0\.9is average\-best for7/97/9models, while Gemini\-3\-Flash and GPT\-5\.2 favor the matched value of0\.70\.7\. Thus, at lower economic patience, matching remains the most common buyer\-side choice, whereas seller\-side guidance favors strategic overstatement\.
Table 7:Prompted Patience Is a Principal\-Specification Leverδecon\\delta^\{\\mathrm\{econ\}\}PartyMatching is best\(of 36 cells\)Potential gain fromcell\-specific tuningRecommended prompt rule0\.9Buyer20\+74\.81\+74\.81\[\+29\.4,\+130\.2\]\[\+29\.4,\\,\+130\.2\]Match \(modal\); understate for mostSeller26\+18\.28\+18\.28\[\+4\.4,\+35\.1\]\[\+4\.4,\\,\+35\.1\]Match: setδstrat=0\.9\\delta^\{\\mathrm\{strat\}\}=0\.90\.7Buyer13\+41\.11\+41\.11\[\+25\.1,\+58\.4\]\[\+25\.1,\\,\+58\.4\]Match \(most common\); tune by modelSeller11\+66\.53\+66\.53\[\+39\.0,\+100\.5\]\[\+39\.0,\\,\+100\.5\]Setδstrat=0\.9\\delta^\{\\mathrm\{strat\}\}=0\.9; match for two models
Notes:Each row compares the matched choiceδstrat=δecon\\delta^\{\\mathrm\{strat\}\}=\\delta^\{\\mathrm\{econ\}\}with the best of\{0\.9,0\.7,0\.4\}\\\{0\.9,0\.7,0\.4\\\}across 36 cells \(9 models×\\times2 buyer types×\\times2 first\-proposer orders\)\. Potential gain is best minus matched payoff \(mean over cells; 95% percentile bootstrap CI\)\. The recommended rule is the portable default; the model\-level heterogeneity, which models gain from deviating from matching and by how much, is summarized above and detailed in Appendix[10\.7](https://arxiv.org/html/2608.07538#S10.SS7)\.
Three implications follow for a principal specifying an agent’s strategic patience\. First, the choice is consequential: the mean gains in Table[7](https://arxiv.org/html/2608.07538#S6.T7), set against quantity×\\timesprice scales of a few hundred, mean that a principal who treatsδstrat\\delta^\{\\text\{strat\}\}as irrelevant leaves surplus on the table\. Second, the optimum depends on role, model, and the principal’s true cost of delay\. When delay is cheap \(δecon=0\.9\\delta^\{\\text\{econ\}\}=0\.9\), matching strategic to economic patience is average\-best for every seller model, whereas five of nine buyer models benefit from understatement\. When delay is moderately costly \(δecon=0\.7\\delta^\{\\text\{econ\}\}=0\.7\), seven of nine seller models benefit from overstatement toδstrat=0\.9\\delta^\{\\text\{strat\}\}=0\.9, while the buyer\-side optimum varies by model: matching maximizes mean payoff for four models, overstatement for three, and understatement for two\. Third, the principal should treatδstrat\\delta^\{\\text\{strat\}\}as a prompt\-specification variable evaluated againstδecon\\delta^\{\\text\{econ\}\}, not as a literal translation of human patience\. We defer to future work the meta\-game in which both principals jointly optimizeδstrat\\delta^\{\\text\{strat\}\}\.
### 6\.4Additional Robustness Extensions \(R4–R7\)
The verbal channel \(R1, §[6\.1](https://arxiv.org/html/2608.07538#S6.SS1)\), prompted discounting framework \(R2, §[6\.2](https://arxiv.org/html/2608.07538#S6.SS2)\), and strategic\-patience analysis \(R3, §[6\.3](https://arxiv.org/html/2608.07538#S6.SS3)\) were examined above as configuration levers\. Four additional extensions probe the remaining design choices most likely to unsettle the findings: reasoning effort \(R4\), surplus magnitude \(R5\), the seller’s prior \(R6\), and model size \(R7\)\. The outcome regularities of Sections[4](https://arxiv.org/html/2608.07538#S4)–[5](https://arxiv.org/html/2608.07538#S5)survive all four: high undiscounted efficiency with delayed agreement, the provider rank ordering \(with provider\-specific exceptions catalogued in the appendix\), and the capability–reliability gradient all persist, whereas the process mechanisms of Section[4\.3](https://arxiv.org/html/2608.07538#S4.SS3)and Appendix[10\.4](https://arxiv.org/html/2608.07538#S10.SS4)are validated less directly, since none of R4–R7 re\-run the turn\-level strategy inference or disclosure labeling\. The reasoning\-effort and model\-size ablations \(R4 and R7\) distinguish reliability from tempo\. The parameter\-size ladder \(R7\) sharpens the reliability gradient into a continuous within\-family slope, while both model size \(F=18\.0F=18\.0,p<0\.001p<0\.001\) and reasoning effort \(F=11\.22F=11\.22,p<0\.001p<0\.001\) affect negotiation speed\. In R4, greater reasoning effort also raises discounted efficiency while undiscounted efficiency does not change significantly; within R7, the relation between size and speed is non\-monotonic\. The complete program \(all seven extensions, with per\-extension results and the full inventory\) appears in Appendix[11](https://arxiv.org/html/2608.07538#S11)\(Table[26](https://arxiv.org/html/2608.07538#S11.T26)\)\.
## 7Conclusion
We examine how LLM agents negotiate in a dynamic bargaining game with asymmetric information, and three conclusions emerge\. First,capability is the value\-creation lever\. It governs allocative efficiency: agents reach agreement in98\.9%98\.9\\%of cases and capture95\.4%95\.4\\%of first\-best surplus, but take more than twice the benchmark rounds, so discounting erodes 21–34% of first\-best surplus, depending on patience\. The same ordering governs operational reliability, where baseline models accept individually irrational agreements an order of magnitude more often than the upper tiers\. Second,provider identity is the distributional lever, and bargaining power is relational, not intrinsic: capability and provider load on different margins, surplus division varies markedly across providers, cross\-family direction effects rival within\-family capability effects, and a model’s provider profile can override its capability across families\. Third,the principal’s configuration choices are strategic levers: the communication channel reshapes provider biases, removing the prompted discounting framework preserves the main regularities while lengthening bargaining, and delegation to an LLM agent gives the principal a choice absent from classical direct bargaining: economic patience remains the fixed cost of delay, while the agent’s strategic patience can be chosen at deployment\.
Empirically, we document provider\-level bargaining heterogeneity that is as consequential as within\-family capability differences once counterparties are heterogeneous, establishing vendor choice as a strategic operations decision rather than a technical one\. Methodologically, we introduce an executable implementation of theFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)PBE, validated cell\-by\-cell against the equilibrium’s analytical properties, enabling evaluation of autonomous agents against a theoretical benchmark rather than human baselines alone; the approach illustrates a template for benchmarking agents in other game\-theoretic operational settings, though each new setting requires its own equilibrium solution\. Practically, the findings yield a three\-dimensional deployment framework developed below\.
The framework comprises three deployment rules that follow directly from the findings\.Time\-adjusted efficiencyshifts the locus of control from outcome to process: because deal\-level outcomes carry meaningful stochasticity even under flagship models, organizations should bound delay ex ante through round caps, escalation rules for stalled negotiations, and final\-offer procedures for high\-urgency contracts\.Distributional profilemakes vendor choice first\-order whenever counterparties run different providers: cross\-family direction effects \(77–1818pp\) are comparable to within\-family tier effects \(∼\\sim17 pp\), so firms controlling both sides of a transaction should weight provider identity at least as heavily as capability rank when assigning roles, and, because self\-play profiles do not transfer directly to heterogeneous matchups, audit the specific cross\-vendor pairing in simulation before delegation\.Operational reliabilityseparates into two regimes: the0\.00\.0–0\.6%0\.6\\%irrationality rate at flagship and mid\-tier permits lighter monitoring, while the19\.2%19\.2\\%rate at baseline makes dual\-party automated profit verification non\-negotiable\. Because arithmetic errors, failed constraint checks, and instruction non\-compliance are operationally equivalent \(each yields an economically unsafe contract\), the verification layer should treat them uniformly\.
We distinguishstructural regularities\(high undiscounted efficiency with delayed agreement, provider\-level heterogeneity in surplus division, and concentration of unsafe agreements among weaker models\) frommodel\-specific resultssuch as specific surplus shares, irrationality rates, and provider rankings\. The former are likely to travel across nearby settings; the latter are calibrated to the model versions tested here and require periodic recalibration as providers update their systems\. This division rests on evidence, not assertion: across the three capability tiers and roughly three model generations we test, the structural regularities recur at every point along the frontier, including the reliability gradient traced continuously by the R7 capacity ladder, while only the model\-specific results shift with capability\. Extending the structural claims to models outside this tested range is a prediction the design supports but does not itself verify\.
Three limitations bound the interpretation of these results\. First, prompts are role\-specific, so the reported surplus shares reflect behavior under a common prompting regime and are not prompt\-free estimates of intrinsic bargaining bias; the qualitative pairing order is broadly stable across our prompt variations, with specific exceptions identified in Section[6\.4](https://arxiv.org/html/2608.07538#S6.SS4)\. Second, our design discloses each party’s patience parameters as common knowledge to both agents, matching the informational structure of the Bayesian benchmark; whether the qualitative patterns persist when a principal’s patience or urgency is undisclosed to the counterparty, a common feature of real negotiations, is untested\. Third, several interpretive moves remain process\-descriptive rather than mechanistically identified: the strategic\-behavior evidence of Section[4\.3](https://arxiv.org/html/2608.07538#S4.SS3)documents heuristic substitution and a behavior–theory gap in equilibrium comparative statics without adjudicating among competing cognitive explanations, and the hazard\-based reading ofδ\\deltain Section[3\.3](https://arxiv.org/html/2608.07538#S3.SS3)is an analogy rather than a tested mechanism\.
These limitations point to a concrete research agenda\. The hazard\-based reading ofδ\\deltais testable by exogenously varying per\-round reliability \(through compute budgets, injected parsing noise, or forced\-termination protocols\), and the related conjecture that effectiveδ\\deltadeclines as conversation history accumulates can be tested by measuring error rates against round number across context\-management regimes\. If these mappings hold, engineering choices organizations already make carry bargaining consequences \(round budgets as value\-conditional commitment devices, profit\-verification guardrails as credible commitments, and an agent\-versus\-human boundary drawn along per\-interaction reliability rather than general capability\)\. Field studies, in turn, would assess which structural regularities survive once agents are embedded in real approval and verification workflows\. As LLM agents move from pilot to production in procurement, the operational question shifts from whether a delegated agent can close a deal to whether the deal it closes is one the principal would have authorized\. The audit developed here answers that question before delegation, not after—replacing the intuition that a more capable agent is simply a better bargainer with a discipline that treats capability, provider, and configuration as three separate levers a principal must each get right\.
## References
- Armitage \(1955\)Armitage, Peter\. 1955\.Tests for linear trends in proportions and frequencies\.Biometrics11\(3\) 375–386\.
- Brynjolfsson et al\. \(2025\)Brynjolfsson, Erik, Danielle Li, Lindsey Raymond\. 2025\.Generative ai at work\.The Quarterly Journal of Economics140\(2\) 889–942\.
- Camerer et al\. \(2019\)Camerer, Colin F, Gideon Nave, Alec Smith\. 2019\.Dynamic unstructured bargaining with private information: Theory, experiment, and outcome prediction via machine learning\.Management Science65\(4\) 1867–1890\.
- Chen and Huang \(2026\)Chen, Yang, Rihuan Huang\. 2026\.Haggling with a bot: Human vs\. llm negotiation in supply chain contracts\.Working Paper\.
- Chen et al\. \(2025\)Chen, Yang, Samuel N Kirshner, Anton Ovchinnikov, Meena Andiappan, Tracy Jenkin\. 2025\.A manager and an ai walk into a bar: does chatgpt make biased decisions like we do?Manufacturing & Service Operations Management27\(2\) 354–368\.
- Coase \(1972\)Coase, Ronald H\. 1972\.Durability and monopoly\.Journal of Law and Economics15\(1\) 143–149\.
- Cochran \(1954\)Cochran, William G\. 1954\.Some methods for strengthening the commonχ\\chi2 tests\.Biometrics10\(4\) 417–451\.
- Cohen et al\. \(2026\)Cohen, Maxime C, Tinglong Dai, Georgia Perakis, Narendra Agrawal, Gad Allon, Robert N Boute, Gerard P Cachon, Zhe Chen, Morris A Cohen, Rares Cristian, et al\. 2026\.OM forum—supply chain management in the AI era: A vision statement from the operations management community\.Manufacturing & Service Operations Management28\(3\) 687–705\.
- Dai et al\. \(2025\)Dai, Tinglong, David Simchi\-Levi, Michelle Xiao Wu, Yao Xie\. 2025\.Assured autonomy: How operations research powers and orchestrates generative ai systems\.arXiv preprint arXiv:2512\.23978\.
- Davis et al\. \(2022\)Davis, Andrew M, Bin Hu, Kyle Hyndman, Anyan Qi\. 2022\.Procurement for assembly under asymmetric information: Theory and evidence\.Management Science68\(4\) 2327–2347\.[10\.1287/mnsc\.2021\.4000](https://arxiv.org/doi.org/10.1287/mnsc.2021.4000)\.
- Davis and Hyndman \(2021\)Davis, Andrew M, Kyle Hyndman\. 2021\.Private information and dynamic bargaining in supply chains: An experimental study\.Manufacturing & Service Operations Management23\(6\) 1449–1467\.[10\.1287/msom\.2020\.0896](https://arxiv.org/doi.org/10.1287/msom.2020.0896)\.
- Davis and Hyndman \(2025\)Davis, Andrew M, Kyle Hyndman\. 2025\.Bargaining with voluntary disclosure and endogenous matching\.Management Science71\(2\) 1102–1119\.
- Eloundou et al\. \(2024\)Eloundou, Tyna, Sam Manning, Pamela Mishkin, Daniel Rock\. 2024\.Gpts are gpts: Labor market impact potential of llms\.Science384\(6702\) 1306–1308\.
- Feng et al\. \(2015\)Feng, Qi, Guoming Lai, Lauren Xiaoyuan Lu\. 2015\.Dynamic bargaining in a supply chain with asymmetric demand information\.Management Science61\(2\) 301–315\.[10\.1287/mnsc\.2014\.1938](https://arxiv.org/doi.org/10.1287/mnsc.2014.1938)\.
- Filippas et al\. \(2024\)Filippas, Apostolos, John J Horton, Benjamin S Manning\. 2024\.Large language models as simulated economic agents: What can we learn from homo silicus?Proceedings of the 25th ACM Conference on Economics and Computation\. 614–615\.
- Fransoo et al\. \(2026\)Fransoo, Jan C, Robert Peels, Maximiliano Udenio\. 2026\.Navigating supply chain dynamics for sustained AI growth\.Maxime C Cohen, Tinglong Dai, eds\.,AI in Supply Chains: Perspectives from Global Thought Leaders\. Springer, Cham, 37–54\.
- Grömping \(2007\)Grömping, Ulrike\. 2007\.Estimators of relative importance in linear regression based on variance decomposition\.The American Statistician61\(2\) 139–147\.
- Gul et al\. \(1986\)Gul, Faruk, Hugo Sonnenschein, Robert Wilson\. 1986\.Foundations of dynamic monopoly and the Coase conjecture\.Journal of Economic Theory39\(1\) 155–190\.
- Hadfield and Koh \(2025\)Hadfield, Gillian K, Andrew Koh\. 2025\.An economy of ai agents\.arXiv preprint arXiv:2509\.01063\.
- Haruvy et al\. \(2020\)Haruvy, Ernan, Elena Katok, Valery Pavlov\. 2020\.Bargaining process and channel efficiency\.Management Science66\(7\) 2845–2860\.
- Hasija and Castillo \(2025\)Hasija, Sunny, Vincent E Castillo\. 2025\.When anchors sink suppliers: Role\-based asymmetry bias in ai\-automated buyer\-supplier negotiations\.Available at SSRN 5522018\.
- Ide and Talamas \(2025\)Ide, Enrique, Eduard Talamas\. 2025\.Artificial intelligence in the knowledge economy\.Journal of Political Economy\.
- Imas et al\. \(2025\)Imas, Alex, Kevin Lee, Sanjog Misra\. 2025\.Agentic interactions\.Available at SSRN 5875162\.
- Katok and Pavlov \(2013\)Katok, Elena, Valery Pavlov\. 2013\.Fairness in supply chain contracts: A laboratory study\.Journal of Operations Management31\(3\) 129–137\.
- Kennan and Wilson \(1993\)Kennan, John, Robert Wilson\. 1993\.Bargaining with private information\.Journal of Economic Literature31\(1\) 45–104\.
- Kirshner et al\. \(2025\)Kirshner, Samuel, Yiwen Pan, Jason Xianghua Wu\. 2025\.The ai agent’s dilemma: Llm contract design under moral hazard\.Available at SSRN 5607356\.
- Kirshner et al\. \(2026\)Kirshner, Samuel N, Yiwen Pan, Jason Xianghua Wu, Alex Gould\. 2026\.Talking terms: Agent information in llm supply chain bargaining\.Decision Sciences57\(1\) 9–23\.
- Kong et al\. \(2025\)Kong, Dexin, Xu Yan, Ming Chen, Shuguang Han, Jufeng Chen, Fei Huang\. 2025\.Fishbargain: An llm\-empowered bargaining agent for online fleamarket platform sellers\.Companion Proceedings of the ACM on Web Conference 2025\. 2855–2858\.
- Liu et al\. \(2025\)Liu, Jifei, Zhi Chen, Yuanguang Zhong\. 2025\.Large language newsvendor: Decision biases and cognitive mechanisms\.arXiv preprint arXiv:2512\.12552\.
- Liu and Ning \(2025\)Liu, Ye, Jie Ning\. 2025\.Make your chatbot price like you do: Identify and customize distributional preferences of large language models in supply chains\.Available at SSRN\.
- Manning and Horton \(2026\)Manning, Benjamin S, John J Horton\. 2026\.General social agents\.Tech\. rep\., National Bureau of Economic Research\.
- Myerson and Satterthwaite \(1983\)Myerson, Roger B\., Mark A\. Satterthwaite\. 1983\.Efficient mechanisms for bilateral trading\.Journal of Economic Theory29\(2\) 265–281\.
- Pearson \(1900\)Pearson, Karl\. 1900\.On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling\.The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science50\(302\) 157–175\.
- Rahwan et al\. \(2019\)Rahwan, Iyad, Manuel Cebrian, Nick Obradovich, et al\. 2019\.Machine behaviour\.Nature568477–486\.
- Rubinstein \(1982\)Rubinstein, Ariel\. 1982\.Perfect equilibrium in a bargaining model\.Econometrica50\(1\) 97–109\.
- Shahidi et al\. \(2025\)Shahidi, Peyman, Gili Rusak, Benjamin S Manning, Andrey Fradkin, John J Horton\. 2025\.The coasean singularity? demand, supply, and market design with ai agents\.NBER Chapters\.
- Simchi\-Levi et al\. \(2025\)Simchi\-Levi, David, Tinglong Dai, Ishai Menache, Michelle Xiao Wu\. 2025\.Democratizing optimization with generative ai\.Available at SSRN 5511218\.
- Simon \(1955\)Simon, Herbert A\. 1955\.A behavioral model of rational choice\.The Quarterly Journal of Economics69\(1\) 99–118\.
- Vaccaro et al\. \(2026\)Vaccaro, Michelle, Michael Caosun, Harang Ju, Sinan Aral, Jared R\. Curhan\. 2026\.Advancing AI negotiations: A large\-scale autonomous negotiation competition\.Proceedings of the National Academy of Sciences123\(23\) e2521774123\.[10\.1073/pnas\.2521774123](https://arxiv.org/doi.org/10.1073/pnas.2521774123)\.
- Van Hoek et al\. \(2022\)Van Hoek, Remko, Michael DeWitt, Mary Lacity, Travis Johnson\. 2022\.How Walmart automated supplier negotiations\.Harvard Business Review\.URL[https://hbr\.org/2022/11/how\-walmart\-automated\-supplier\-negotiations](https://hbr.org/2022/11/how-walmart-automated-supplier-negotiations)\.November 8, 2022\.
- Wang et al\. \(2025\)Wang, Shaoyu, Ying\-Ju Chen, Pin Gao, Yang Li\. 2025\.Sponsored search with ai\-generated ad content\.Available at SSRN 5724045\.
- Welch \(1947\)Welch, Bernard L\. 1947\.The generalization of ‘Student’s’ problem when several different population variances are involved\.Biometrika34\(1/2\) 28–35\.
- Xu et al\. \(2025\)Xu, Fasheng, Jing Hou, Wei Chen, Karen Xie\. 2025\.Generative ai and organizational structure in the knowledge economy\.arXiv preprint arXiv:2506\.00532\.
- Xu et al\. \(2024\)Xu, Fasheng, Xiaoyu Wang, Wei Chen, Karen Xie\. 2024\.The economics of ai foundation models: Openness, competition, and governance\.SSRN\.
- Zhu et al\. \(2025\)Zhu, Shenzhe, Jiao Sun, Yi Nian, Tobin South, Alex Pentland, Jiaxin Pei\. 2025\.The automated but risky game: Modeling agent\-to\-agent negotiations and transactions in consumer markets\.ICML 2025 Workshop on Reliable and Responsible Foundation Models\.
\\ECSwitch\{APPENDICES\}
Online Appendix
The appendix is organized to support reproducibility first and interpretation second: Appendix[8](https://arxiv.org/html/2608.07538#S8)reports the agent prompts, Appendices[9](https://arxiv.org/html/2608.07538#S9)and[10](https://arxiv.org/html/2608.07538#S10)provide the detailed evidence behind the performance and process findings, Appendix[11](https://arxiv.org/html/2608.07538#S11)reports the design and robustness extension program, and Appendix[12](https://arxiv.org/html/2608.07538#S12)documents the Bayesian benchmark implementation\.
## 8Prompts
Prompt for the Buyer AgentYou are a buyer negotiating a supply contract with a seller\.THE NEGOTIATION:•You buy a product from a seller to retail to end consumers•You negotiate overquantity\(qq\) andtotal payment\(TT\)•Your expected revenue:R\(q\)=r⋅𝔼\[min\(D,q\)\]R\(q\)=r\\cdot\\mathbb\{E\}\[\\min\(D,q\)\], whereD∼𝒩\(μ,σ\)D\\sim\\mathcal\{N\}\(\\mu,\\sigma\)•Your profit:πB=R\(q\)−T\\pi\_\{B\}=R\(q\)\-T•Your effective utility:UB=πB×δB\(τ−1\)U\_\{B\}=\\pi\_\{B\}\\times\\delta\_\{B\}^\{\(\\tau\-1\)\}, whereδB\\delta\_\{B\}is buyer’s patience andτ\\tauis round number•You haveprivate informationabout demand: you know if you’re HIGH or LOW type•The seller doesn’t know your type but has beliefs about it•Seller’s effective utility:US=\(T−c⋅q\)×δS\(τ−1\)U\_\{S\}=\(T\-c\\cdot q\)\\times\\delta\_\{S\}^\{\(\\tau\-1\)\}•You cannot accept ifπB<0\\pi\_\{B\}<0THE NEGOTIATION PROCESS:•Either seller or buyer proposes first \(check “First proposer” in context\)•Players alternate: proposer makes offer, responder accepts or rejects•If rejected, roles switch and the other player proposes•The proposer gets a larger share of surplus than the responder•Maximum rounds:\{p\.max\_rounds\}\. If no deal by final round, profit is outside option\.YOUR OBJECTIVE:Maximize your effective utility through negotiation overqqandTT\.YOUR PRIVATE ADVANTAGE:•You know your true demand type•HH\-type: higher demand→\\rightarrowhigher optimalqq→\\rightarrowhigher potential surplus•LL\-type: lower demand→\\rightarrowlower optimalqq→\\rightarrowlower potential surplusRESPONSE FORMAT: IF PROPOSER``` { "public": { "action": "propose", "quantity": <number>, "payment": <number>, "message": "<brief message to seller>" }, "private": { "reasoning": "<internal reasoning>" } } ``` RESPONSE FORMAT: IF RESPONDER``` { "public": { "action": "accept" or "reject", "message": "<brief message to seller>" }, "private": { "reasoning": "<internal reasoning>" } } ``` CONTEXT PARAMETERS:•First proposer:\{p\.first\_proposer\},•Retail price:r=\{p\.retail\_price\}r=\\texttt\{\\\{p\.retail\\\_price\\\}\},•Production cost:c=\{p\.production\_cost\}c=\\texttt\{\\\{p\.production\\\_cost\\\}\}•Buyer patience:δB=\{p\.delta\_buyer\}\\delta\_\{B\}=\\texttt\{\\\{p\.delta\\\_buyer\\\}\},•Seller patience:δS=\{p\.delta\_seller\}\\delta\_\{S\}=\\texttt\{\\\{p\.delta\\\_seller\\\}\}•HH\-type demand:𝒩\(\{p\.mean\_high\},\{p\.std\_high\}\)\\mathcal\{N\}\(\\texttt\{\\\{p\.mean\\\_high\\\}\},\\texttt\{\\\{p\.std\\\_high\\\}\}\),•LL\-type demand:𝒩\(\{p\.mean\_low\},\{p\.std\_low\}\)\\mathcal\{N\}\(\\texttt\{\\\{p\.mean\\\_low\\\}\},\\texttt\{\\\{p\.std\\\_low\\\}\}\)•PriorPr\(H\)=\{p\.prior\_high\}\\Pr\(H\)=\\texttt\{\\\{p\.prior\\\_high\\\}\},•Max rounds:\{p\.max\_rounds\}YOUR PRIVATE INFO:True type =\{self\.true\_type\}\-type\. Seller believesPr\(H\)=\{p\.prior\_high\}\\Pr\(H\)=\\texttt\{\\\{p\.prior\\\_high\\\}\}\.
Prompt for the Seller AgentYou are a seller negotiating a supply contract with a buyer\.THE NEGOTIATION:•You sell a product to a buyer who will retail it to end consumers•You negotiate overquantity\(qq\) andtotal payment\(TT\)•Your profit:πS=T−c⋅q\\pi\_\{S\}=T\-c\\cdot q, whereccis production cost•Your effective utility:US=πS×δS\(τ−1\)U\_\{S\}=\\pi\_\{S\}\\times\\delta\_\{S\}^\{\(\\tau\-1\)\}, whereδS\\delta\_\{S\}is seller’s patience andτ\\tauis round number•The buyer hasprivate informationabout demand \(High or Low type\)•You don’t know the buyer’s type but have a belief \(probability\)•Buyer’s effective utility:UB=\(r⋅𝔼\[min\(D,q\)\]−T\)×δB\(τ−1\)U\_\{B\}=\(r\\cdot\\mathbb\{E\}\[\\min\(D,q\)\]\-T\)\\times\\delta\_\{B\}^\{\(\\tau\-1\)\}, whereD∼𝒩\(μ,σ\)D\\sim\\mathcal\{N\}\(\\mu,\\sigma\)•You cannot accept ifπS<0\\pi\_\{S\}<0THE NEGOTIATION PROCESS:•Either seller or buyer proposes first \(check “First proposer” in context\)•Players alternate: proposer makes offer, responder accepts or rejects•If rejected, roles switch and the other player proposes•The proposer gets a larger share of surplus than the responder•Maximum rounds:\{p\.max\_rounds\}\. If no deal by final round, profit is outside option\.YOUR OBJECTIVE:Maximize your effective utility through negotiation overqqandTT\.YOUR INFORMATION DISADVANTAGE:•You do NOT know the buyer’s true demand type•You only have a prior belief:Pr\(H\-type\)=β\\Pr\(\\text\{H\-type\}\)=\\beta•HH\-type: higher demand→\\rightarrowhigher optimalqq→\\rightarrowhigher potential surplus•LL\-type: lower demand→\\rightarrowlower optimalqq→\\rightarrowlower potential surplus•You must infer the buyer’s type from their behaviorRESPONSE FORMAT: IF PROPOSER``` { "public": { "action": "propose", "quantity": <number>, "payment": <number>, "message": "<brief message to buyer>" }, "private": { "reasoning": "<internal reasoning>", "belief_about_buyer": "<belief about buyer’s type>" } } ``` RESPONSE FORMAT: IF RESPONDER``` { "public": { "action": "accept" or "reject", "message": "<brief message to buyer>" }, "private": { "reasoning": "<internal reasoning>", "belief_about_buyer": "<belief about buyer’s type>" } } ``` COMMUNICATION GUIDELINES:•message: What you SAY to the buyer \(they see this\)\. Use to negotiate, probe, or justify\.•reasoning: Your PRIVATE reasoning \(they don’t see this\)\. Your true strategic thinking\.•belief\_about\_buyer: Your private assessment of whether buyer isHH\-type orLL\-type\.CONTEXT PARAMETERS:•First proposer:\{p\.first\_proposer\}•Retail price:r=\{p\.retail\_price\}r=\\texttt\{\\\{p\.retail\\\_price\\\}\}•Production cost:c=\{p\.production\_cost\}c=\\texttt\{\\\{p\.production\\\_cost\\\}\}•Buyer patience:δB=\{p\.delta\_buyer\}\\delta\_\{B\}=\\texttt\{\\\{p\.delta\\\_buyer\\\}\}•Seller patience:δS=\{p\.delta\_seller\}\\delta\_\{S\}=\\texttt\{\\\{p\.delta\\\_seller\\\}\}•HH\-type demand:DH∼𝒩\(\{p\.mean\_high\},\{p\.std\_high\}\)D\_\{H\}\\sim\\mathcal\{N\}\(\\texttt\{\\\{p\.mean\\\_high\\\}\},\\texttt\{\\\{p\.std\\\_high\\\}\}\)•LL\-type demand:DL∼𝒩\(\{p\.mean\_low\},\{p\.std\_low\}\)D\_\{L\}\\sim\\mathcal\{N\}\(\\texttt\{\\\{p\.mean\\\_low\\\}\},\\texttt\{\\\{p\.std\\\_low\\\}\}\)•Prior:Pr\(H\)=\{p\.prior\_high\}\\Pr\(H\)=\\texttt\{\\\{p\.prior\\\_high\\\}\}•Max rounds:\{p\.max\_rounds\}•Outside options: Seller =\{p\.seller\_outside\_option\}, Buyer =\{p\.buyer\_outside\_option\}YOUR BELIEF:You believe the buyer isHH\-type with probabilityPr\(H\)=\{p\.prior\_high\}\\Pr\(H\)=\\texttt\{\\\{p\.prior\\\_high\\\}\}\. Update this belief based on the buyer’s actions during negotiation\.
### Prompt Variant for R2 \(No\-Discounting Bargaining\)
The R2 extension \(Appendix[11\.3](https://arxiv.org/html/2608.07538#S11.SS3)\) removes all references to discount factors, patience, effective utility, and time pressure from both buyer and seller prompts\. Relative to the main\-analysis prompts above, the following lines are removed or modified; all other content is unchanged\.
Buyer Agent: Modifications for R2Removed from THE NEGOTIATION block:•Your effective utility:UB=πB×δB\(τ−1\)U\_\{B\}=\\pi\_\{B\}\\times\\delta\_\{B\}^\{\(\\tau\-1\)\}, whereδB\\delta\_\{B\}is buyer’s patience andτ\\tauis round number•Seller’s effective utility:US=\(T−c⋅q\)×δS\(τ−1\)U\_\{S\}=\(T\-c\\cdot q\)\\times\\delta\_\{S\}^\{\(\\tau\-1\)\}Added:•Seller’s profit:πS=T−c⋅q\\pi\_\{S\}=T\-c\\cdot qYOUR OBJECTIVErewritten as: Maximize your profitπB\\pi\_\{B\}through negotiation overqqandTT\.Removed from CONTEXT PARAMETERS:•Buyer patience:δB=\{p\.delta\_buyer\}\\delta\_\{B\}=\\texttt\{\\\{p\.delta\\\_buyer\\\}\}•Seller patience:δS=\{p\.delta\_seller\}\\delta\_\{S\}=\\texttt\{\\\{p\.delta\\\_seller\\\}\}
Seller Agent: Modifications for R2Analogous modifications apply: references toδB\\delta\_\{B\},δS\\delta\_\{S\}, effective utility formulas, and patience parameters are removed, and the objective is rewritten as maximizing profitπS=T−c⋅q\\pi\_\{S\}=T\-c\\cdot q\.
All other prompt content \(role description, private/public information structure, response format, negotiation process, and parameter values for retail price, production cost, type distributions, and prior\) is identical to the main\-analysis prompts\.
## 9Negotiation Performance Details
This appendix provides detailed analyses supporting the negotiation performance findings reported in Section 4 of the main text: efficiency\-gap decomposition and cross\-condition heterogeneity \(Sections[9\.1](https://arxiv.org/html/2608.07538#S9.SS1)–[9\.2](https://arxiv.org/html/2608.07538#S9.SS2)\), surplus\-division and contract\-terms detail \(Sections[9\.3](https://arxiv.org/html/2608.07538#S9.SS3)–[9\.4](https://arxiv.org/html/2608.07538#S9.SS4)\), model\-level irrationality \(Section[9\.5](https://arxiv.org/html/2608.07538#S9.SS5)\), and cross\-model and cross\-family welfare comparisons \(Sections[9\.6](https://arxiv.org/html/2608.07538#S9.SS6)–[9\.8](https://arxiv.org/html/2608.07538#S9.SS8)\)\.
### 9\.1Efficiency Gap Decomposition Visualizations
The additivity tests underlying the orthogonal\-drivers partition of Section[4\.4](https://arxiv.org/html/2608.07538#S4.SS4)are reported in Table[8](https://arxiv.org/html/2608.07538#S9.T8): incrementalR2R^\{2\}and jointFF\-tests for adding pairwise and three\-way interactions to an additive baseline of capability, patience, information, and first proposer\. The three\-way Capability×\\timesPatience×\\timesInformation interaction is null for both outcomes, and interactions add little explanatory power overall \(ΔR2≤0\.013\\Delta R^\{2\}\\leq 0\.013\); the two\-way block is significant for efficiency but not for buyer surplus share\. The companion within\-stratum coefficient matrix that visualizes the driver separation appears in Figure[4](https://arxiv.org/html/2608.07538#S9.F4)below\.
Table 8:Additivity Tests for the Efficiency and Buyer\-Share Variance PartitionEfficiencyBuyer surplus shareΔR2\\Delta R^\{2\}\(all interactions\)0\.0130\.012ΔR2\\Delta R^\{2\}\(3\-way only\)0\.0020\.001JointFF, all interactionsF17=3\.07F\_\{17\}=3\.07,p<0\.001p\{<\}0\.001F17=1\.87F\_\{17\}=1\.87,p=0\.017p=0\.0172\-way blockF11=3\.53F\_\{11\}=3\.53,p<0\.001p\{<\}0\.001F11=0\.96F\_\{11\}=0\.96,p=0\.480p=0\.4803\-way \(Tier×\\timesPat×\\timesInfo\)F6=0\.63F\_\{6\}=0\.63,p=0\.709p=0\.709F6=0\.51F\_\{6\}=0\.51,p=0\.797p=0\.797NN1,9981,998
Notes:IncrementalR2R^\{2\}and jointFF\-tests for adding all pairwise interactions \(2\-way block\) and the three\-way Capability×\\timesPatience×\\timesInformation \(3\-way block\) to an additive baseline of Capability, Patience, Information, and first proposer\(Grömping[2007](https://arxiv.org/html/2608.07538#bib.bib17)\)\. Sample:N=1,998N=1\{,\}998economically rational agreements\.

Figure 4:Drivers of Efficiency and Buyer Surplus ShareNotes:Each cell plots within\-stratum coefficients with 95% CIs\. Top row: capability\-tier contrasts \(Mid and Flagship vs\. Baseline\) within each patience cell, for efficiency \(left\) and buyer surplus share \(right\)\. Bottom row: patience\-configuration contrasts \(vs\. symmetric\(0\.9,0\.9\)\(0\.9,0\.9\)\) within each capability tier, for efficiency \(left\) and buyer surplus share \(right\)\.ΔR2\\Delta R^\{2\}is the incremental fit from adding the focal factor to a within\-stratum regression\. The\(δB=0\.9,δS=0\.4\)\(\\delta\_\{B\}=0\.9,\\delta\_\{S\}=0\.4\)stratum singled out in the upper\-right panel is the regime in which theFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)PBE predicts the sharpest share advantage for the patient buyer, and the within\-cell capability effect captures convergence toward that prediction: flagship buyers recognize and exploit the impatient seller’s weakened bargaining position, while baseline models leave equilibrium surplus unclaimed\. The remaining patience cells return null or marginal capability effects on share\.
Section[4](https://arxiv.org/html/2608.07538#S4)reports an aggregate 4\.56 pp efficiency gap composed of 1\.11 pp failed negotiations, 0\.56 pp irrational agreements, and 2\.89 pp suboptimal terms; Table[9](https://arxiv.org/html/2608.07538#S9.T9)reports the overall decomposition and Table[10](https://arxiv.org/html/2608.07538#S9.T10)disaggregates these three components by model tier and buyer type\.
Table 9:Efficiency Gap DecompositionLoss ComponentEfficiency Loss \(pp\)Share of Gap \(%\)Failed Negotiations1\.1124\.37Irrational Agreements0\.5612\.20Suboptimal Terms2\.8963\.43Total Gap4\.56100\.00
Notes:Gap measured relative to 100% first\-best surplus\. Suboptimal terms refers to agreements with positive profits but below optimal surplus extraction\.
Table 10:Efficiency Loss Components: Detailed BreakdownConditionFailedIrrationalSuboptimalTotalDeals \(pp\)Deals \(pp\)Terms \(pp\)Gap \(pp\)Overall1\.110\.562\.894\.56By Model Tier:Flagship0\.000\.001\.111\.11Mid\-tier0\.140\.023\.433\.59Baseline3\.191\.654\.138\.98By Buyer Type:High\-type0\.460\.404\.004\.87Low\-type1\.760\.711\.784\.25
Notes:Total Gap==Failed\+\+Irrational\+\+Suboptimal, all expressed as percentage points relative to first\-best \(undiscounted\) efficiency\. Failed==surplus lost when no agreement is reached\. Irrational==surplus lost when at least one party agrees to a deal yielding negative realized profit\. Suboptimal==residual loss from agreed quantities deviating from the first\-best\.
Across model tiers, flagship models achieve near\-optimal efficiency \(1\.11 pp total gap\), entirely from suboptimal terms \(1\.11 pp\), with no failed deals and zero irrationality\. Mid\-tier models incur losses chiefly through suboptimal terms \(0\.14 pp failure, 0\.02 pp irrationality, 3\.43 pp suboptimal terms\)\. Baseline models exhibit elevated losses in every category, with failed deals \(3\.19 pp\) and irrationality \(1\.65 pp\) driving most of their 8\.98 pp total gap\. Across buyer types, high\-types show lower failure rates \(0\.46 pp\) but higher suboptimal\-term losses \(4\.00 pp\), reflecting the strategic quantity distortions documented above\. Low\-types exhibit higher failures \(1\.76 pp\) and irrationality \(0\.71 pp\) and a smaller suboptimal\-term loss \(1\.78 pp\), reflecting that low\-type negotiations more often miscoordinate before reaching the contract stage but, conditional on agreement, settle closer to first\-best terms\.
### 9\.2Efficiency Heterogeneity Analysis
This subsection details the per\-model efficiency heterogeneity summarized in Section[4\.1](https://arxiv.org/html/2608.07538#S4.SS1)\. We examine how efficiency varies across the four experimental conditions defined by buyer type and first proposer\. Figure[5](https://arxiv.org/html/2608.07538#S9.F5)displays undiscounted efficiency for each of the nine models in each condition\.

Figure 5:Undiscounted Efficiency by Model and ConditionNotes:Cells report undiscounted efficiency \(percent of first\-best surplus\) for each of the nine models within each of the four buyer\-type×\\timesfirst\-proposer conditions\.
Efficiency Variation\.The heatmap reveals substantial heterogeneity across models and conditions\. The lowest observed model–condition cell is GPT\-4o\-mini in the low\-type buyer\-first condition, at72\.6%72\.6\\%\. Among baseline models, efficiency in the low\-type seller\-first condition ranges from90\.5%90\.5\\%to99\.2%99\.2\\%\.
Flagship models range from97\.4%97\.4\\%to99\.9%99\.9\\%efficiency across the four conditions\. Mid\-tier models range from83\.8%83\.8\\%to99\.9%99\.9\\%, and baseline models from72\.6%72\.6\\%to100\.0%100\.0\\%\. The wider lower\-tail dispersion among mid\-tier and baseline models shows that condition sensitivity is concentrated outside the flagship tier\.
Table[11](https://arxiv.org/html/2608.07538#S9.T11)disaggregates undiscounted efficiency by buyer type within strata of model tier, first proposer, and patience configuration, withtt\-tests for each high\-versus\-low contrast\.
Table 11:Efficiency by Buyer Type: Detailed BreakdownConditionHigh\-TypeLow\-TypeDifferencett\-statpp\-valueEff\. \(%\)Eff\. \(%\)\(pp\)Overall95\.195\.7\-0\.6\-1\.110\.266Flagship98\.799\.1\-0\.5\-2\.000\.046Mid\-tier94\.997\.9\-3\.0\-3\.90<<0\.001Baseline91\.890\.21\.61\.140\.253Buyer First95\.194\.50\.60\.630\.526Seller First95\.297\.0\-1\.8\-2\.730\.006\(0\.9, 0\.9\)95\.395\.10\.20\.130\.898\(0\.7, 0\.9\)96\.096\.3\-0\.3\-0\.330\.742\(0\.4, 0\.9\)94\.695\.4\-0\.8\-0\.790\.427\(0\.9, 0\.4\)94\.796\.2\-1\.5\-1\.340\.181
Notes:Efficiency measured as percentage of first\-best undiscounted surplus\. Differences calculated as high\-type minus low\-type\.
Efficiency is comparable across buyer types: low\-type buyers achieve only marginally higher efficiency than high\-types overall, a difference that does not reach significance \(gap 0\.6 pp,p=0\.266p=0\.266\)\. The low\-type advantage is concentrated in the mid\-tier \(3\.0 pp,p<0\.001p<0\.001\); it is economically negligible at the flagship \(a 0\.5 pp gap near the efficiency ceiling; nominally significant atp=0\.046p=0\.046but not robust to multiple comparisons\), and among baseline models the point estimate reverses in sign though not significantly, with high\-types attaining 1\.6 pp higher efficiency \(p=0\.253p=0\.253\)\.
### 9\.3Surplus Distribution Summary
This subsection details the surplus\-division and compression patterns summarized in Section[5](https://arxiv.org/html/2608.07538#S5)\. Table[12](https://arxiv.org/html/2608.07538#S9.T12)reports the buyer’s share of total surplus across all experimental factors, the seller’s share, alongside the deviation from the Bayesian benchmark implied by patience parameters\. The pooled distribution is centered above equal split with modal buyer share near 56%\.
Table 12:Surplus Division: Descriptive SummaryConditionBuyer ShareSeller ShareBuyer Share−\-Mean \(SD\)Mean \(SD\)Bayesian Ref\.Overall0\.535 \(0\.356\)0\.465 \(0\.356\)\+8\.3∗∗∗Flagship0\.611 \(0\.298\)0\.389 \(0\.298\)\+15\.9∗∗∗Mid\-tier0\.534 \(0\.292\)0\.466 \(0\.292\)\+8\.1∗∗∗Baseline0\.458 \(0\.443\)0\.542 \(0\.443\)\+0\.8Buyer First0\.541 \(0\.370\)0\.459 \(0\.370\)\+7\.0∗∗∗Seller First0\.528 \(0\.341\)0\.472 \(0\.341\)\+9\.7∗∗∗High\-type0\.558 \(0\.307\)0\.442 \(0\.307\)\+12\.7∗∗∗Low\-type0\.511 \(0\.398\)0\.489 \(0\.398\)\+4\.0∗∗\(0\.9, 0\.9\)0\.515 \(0\.344\)0\.485 \(0\.344\)\+0\.2\(0\.7, 0\.9\)0\.513 \(0\.330\)0\.487 \(0\.330\)\+26\.3∗∗∗\(0\.4, 0\.9\)0\.402 \(0\.372\)0\.598 \(0\.372\)\+27\.1∗∗∗\(0\.9, 0\.4\)0\.709 \(0\.303\)0\.291 \(0\.303\)\-20\.5∗∗∗
Notes:\*\*\*p<0\.001p<0\.001, \*\*p<0\.01p<0\.01, \*p<0\.05p<0\.05\. Standard deviations in parentheses\.
Model Tier and Proposer Effects\.Flagship and mid\-tier models exhibit significant buyer advantages of \+15\.9 pp and \+8\.1 pp above the Bayesian reference, respectively\. Baseline models depart sharply from this pattern\. Their \+0\.8 pp deviation is statistically indistinguishable from the equilibrium prediction, and the markedly higher variance \(SD = 0\.443\) indicates unstable rather than strategically calibrated outcomes\. Seller\-first negotiations amplify the deviation to \+9\.7 pp, while buyer\-first negotiations attenuate it to \+7\.0 pp; both remain highly significant, indicating that first\-mover position shifts surplus division without reversing the buyer advantage\. Across buyer types, the deviation is more than three times as large for high\-types \(\+12\.7 pp on a 55\.8% share\) as for low\-types \(\+4\.0 pp on a 51\.1% share\), so the buyer\-side deviation from the Bayesian reference is driven primarily by high\-type buyers\.
Patience Compression Toward Equality\.Under symmetric high patience at\(δB,δS\)=\(0\.9,0\.9\)\(\\delta\_\{B\},\\delta\_\{S\}\)=\(0\.9,0\.9\), buyers capture 51\.5% of surplus, essentially at the Bayesian reference \(\+0\.2 pp\)\. As patience asymmetry shifts in favor of sellers, the deviation widens substantially: at \(0\.7, 0\.9\), buyer share is 51\.3% \(\+26\.3 pp\), and at \(0\.4, 0\.9\), buyer share falls to 40\.2% \(\+27\.1 pp\)\. Although buyers obtain only around half the surplus in these regimes, they remain far above the Rubinstein reference range of 13 to 20% that obtains under strong seller bargaining power\. Conversely, when patience asymmetry favors buyers at \(0\.9, 0\.4\), buyer share rises to 70\.9%\. Yet this falls 20\.5 pp below the Bayesian reference, which approaches 90% under strong buyer bargaining power\. The compression operates in both directions but not to the same degree: buyer share settles near 40% \(close to a 60–40 split\) when patience favors sellers, but reaches 71% \(closer to 70–30\) when patience favors buyers—in both cases well short of the near\-90/10 or near\-10/90 divisions the Bayesian reference implies\.
### 9\.4Detailed Contract Terms Analysis
This subsection details the contract\-terms patterns underlying the strategic\-behavior evidence of Section[4\.3](https://arxiv.org/html/2608.07538#S4.SS3)\. Table[13](https://arxiv.org/html/2608.07538#S9.T13)presents negotiated quantities and wholesale prices across all experimental dimensions\. Two patterns are salient\. First, quantity outcomes separate sharply by type while prices barely do: high\-type contracts averageq=74\.3q=74\.3versus low\-typeq=41\.0q=41\.0, but mean wholesale prices differ by only $1\.05 \($42\.35 vs\. $41\.30\), and the type\-conditional medians both sit at $40\. Second, both proposer direction and patience configuration shift prices materially \(the patience\-asymmetric\(0\.4,0\.9\)\(0\.4,0\.9\)cell averages $45\.18 pooled across type versus $37\.44 at\(0\.9,0\.4\)\(0\.9,0\.4\)\) but leave quantities near the type\-conditional first\-best\. Table[14](https://arxiv.org/html/2608.07538#S9.T14)unpacks the quantity distribution further, showing asymmetric pooling\-consistent distortions: high\-types under\-order in 39\.1% of cases, while low\-types converge tightly on first\-best\.
Table 13:Contract Terms: Comprehensive Analysis by ConditionsConditionQuantityWholesale PriceHighLowHighLowMean \(SD\)Mean \(SD\)Mean \(SD\)Mean \(SD\)Overall74\.3 \(9\.3\)41\.0 \(4\.8\)42\.35 \(8\.70\)41\.30 \(8\.96\)Buyer First74\.3 \(9\.4\)41\.9 \(4\.8\)42\.89 \(8\.80\)40\.32 \(9\.16\)Seller First74\.4 \(9\.3\)40\.0 \(4\.5\)41\.80 \(8\.57\)42\.27 \(8\.66\)\(0\.9, 0\.9\)75\.0 \(8\.3\)40\.7 \(4\.1\)42\.54 \(7\.50\)42\.02 \(8\.73\)\(0\.7, 0\.9\)74\.8 \(8\.6\)41\.1 \(5\.0\)43\.19 \(8\.05\)41\.60 \(8\.54\)\(0\.4, 0\.9\)73\.3 \(10\.5\)41\.3 \(5\.3\)46\.07 \(9\.39\)44\.29 \(9\.37\)\(0\.9, 0\.4\)74\.3 \(9\.6\)40\.7 \(4\.5\)37\.56 \(7\.50\)37\.32 \(7\.70\)Flagship77\.7 \(5\.0\)40\.3 \(2\.7\)40\.25 \(7\.99\)39\.35 \(6\.93\)Mid\-tier74\.7 \(11\.3\)41\.5 \(4\.1\)41\.73 \(7\.88\)41\.79 \(7\.19\)Baseline70\.5 \(9\.1\)41\.1 \(6\.6\)45\.09 \(9\.44\)42\.85 \(11\.75\)Bayesian benchmark80\.0 \(0\.0\)39\.6 \(0\.7\)45\.33 \(8\.18\)42\.90 \(7\.46\)Theory \(first\-best\)80\.040\.045\.33 \(8\.18\)43\.63 \(7\.27\)
Notes:Standard deviations in parentheses\. All high–low quantity differences are significant atp<0\.001p<0\.001using independent Welch tests\. Wholesale\-price differences are generally not significant; the buyer\-first comparison is the clear exception \(p<0\.001p<0\.001\)\. High\-typeqqequals the first\-bestqH=80q\_\{H\}=80exactly \(no screening distortion for the efficient type\); low\-typeqqis distorted slightly belowqL=40q\_\{L\}=40\(SD=0\.7\) to deter high\-type mimicking, with the distortion size varying by patience config\. WP varies because the equilibrium price depends on bargaining patience\.
Table 14:Quantity Choice Patterns Relative to OptimalBuyerSevere UnderModerate UnderNear OptimalModerate OverSevere OverMeanType\(\-20%\+\)\(\-5 to \-20%\)\(±\\pm5%\)\(\+5 to \+20%\)\(\+20%\+\)Deviation\(%\)\(%\)\(%\)\(%\)\(%\)High\-Type12\.726\.359\.51\.40\.0\-5\.7Low\-Type4\.44\.570\.511\.09\.5\+1\.0
Notes:Optimal quantities: High\-type = 80 units, Low\-type = 40 units\. Mean deviation calculated as actual minus optimal quantity\.
The high\-type quantity distortion is asymmetric in two ways\. Distortions are large rather than marginal: 12\.7% of high\-type orders fall more than 20% below optimal, and only 59\.5% sit within±5%\\pm 5\\%of first\-best\. And the distortion is directional: over\-ordering aboveqH=80q\_\{H\}=80is essentially absent \(1\.4% moderate, 0% severe\), consistent with the interpretation that high\-types recognize downward shifts preserve information rents while upward shifts only sacrifice efficiency\. Low\-types concentrate at first\-best \(70\.5% within±5%\\pm 5\\%\) with a slight upward bias \(mean\+1\.0\+1\.0units\); the 20\.5% over\-ordering rate likely reflects a mix of newsvendor calculation noise and, in a minority of cases, costly signaling intended to credibly demonstrate type\.
### 9\.5Irrationality Detailed Analysis by Model
This subsection provides the model\-level detail behind the operational\-reliability threshold of Section[4\.2](https://arxiv.org/html/2608.07538#S4.SS2)\(Figure[2](https://arxiv.org/html/2608.07538#S4.F2)\)\. Unsafe contracts concentrate almost entirely in the two non\-reasoning models, GPT\-4o\-mini \(36\.2%36\.2\\%overall\) and Qwen2\.5\-14B \(20\.0%20\.0\\%\), whereas every reasoning model, including the baseline\-tier Gemini\-2\.5\-Flash \(2\.9%2\.9\\%\), stays at or near zero\.
The tier\- and party\-level aggregates summarized in Section[4\.2](https://arxiv.org/html/2608.07538#S4.SS2)are reported in Table[15](https://arxiv.org/html/2608.07538#S9.T15)\.
Table 15:Irrationality Rates by Model Tier, Buyer Type, and PartyModel TierOverallBuyerSellerIrrational \(%\)Irrational \(%\)Irrational \(%\)Overall6\.55\.11\.4Flagship0\.00\.00\.0Mid\-tier0\.60\.60\.0Baseline19\.215\.14\.2By Buyer Type:High\-type4\.03\.90\.1Low\-type9\.06\.32\.6
Notes:Percentages calculated over all completed agreements \(N=2,136N=2\{,\}136\)\. Per\-model rates with error bars appear in Figure[2](https://arxiv.org/html/2608.07538#S4.F2)\.
Within\-tier dispersion is substantial\. Baseline rates range from 2\.9% \(Gemini\-2\.5\-Flash\) to 36\.2% \(GPT\-4o\-mini\), so baseline\-tier irrationality cannot be treated as a uniform property of the tier\. Every flagship and mid\-tier model except Qwen3\-32B \(1\.7%\) records 0%\. The qualitative pattern documented in Section[4](https://arxiv.org/html/2608.07538#S4), that economically unsafe contracts are overwhelmingly a baseline\-tier phenomenon, therefore holds at the model level\. The within\-baseline risk is concentrated in a small number of specific models\.
The mechanism behind these failures is visible in the agents’ communication transcripts and private reasoning traces\. Two patterns recur\.Revenue miscalculation: a low\-type buyer proposed 40 units at $2,400, reasoning that the offer “allows for a small profit margin that meets my requirements,” against expected revenue of $2,160\.63 \(buyer profit is \-$239\.37\)\.Cost miscalculation: a seller proposed 45 units at $900, publicly describing it as an offer that “respects your budget while also ensuring I cover my costs,” against production costs of $1,350 \(seller profit is \-$450\)\. In both, strategic language is delivered confidently\. The failure is verification, not reasoning: models compute incorrectly but do not check results against rationality constraints, which argues for automated profitability guardrails\.
### 9\.6Cross\-Model Welfare Detail
Table[16](https://arxiv.org/html/2608.07538#S9.T16)reports buyer\- and seller\-side absolute \(undiscounted\) profits and mean rounds for the four asymmetric cross\-model pairs analyzed in Section[4\.5](https://arxiv.org/html/2608.07538#S4.SS5)and Section[5](https://arxiv.org/html/2608.07538#S5): the within\-family OpenAI pair \(GPT\-5\.2↔\\leftrightarrowGPT\-5\-mini\) plus the three cross\-family flagship pairs\.
Within the OpenAI family, the seller earns substantially more in dollar terms when GPT\-5\.2 occupies that role: $931931against a GPT\-5\-mini buyer versus $680680when GPT\-5\-mini sells to GPT\-5\.2 \(Δ=\+251\.2\\Delta=\+251\.2,p<0\.001p<0\.001\)\. The buyer\-side comparison mirrors this and is equally significant: the GPT\-5\-mini buyer earns $562562against GPT\-5\.2 versus the GPT\-5\.2 buyer’s $824824against GPT\-5\-mini \(Δ=−262\.3\\Delta=\-262\.3,p<0\.001p<0\.001\)\. The capability gap also appears on rounds: GPT\-5\.2 as seller takes0\.430\.43more rounds than GPT\-5\-mini in the same role \(p<0\.01p<0\.01\)\. Capability mismatch within a family thus shows up as a large and precisely estimated dollar advantage to the more capable model in either role, reinforced by a smaller difference in rounds\.
Cross\-family pairs sharpen the provider\-profile pattern\. The more buyer\-favoring provider surrenders dollar value as seller relative to a less buyer\-favoring seller in the same pair: against Qwen3\-Max, GPT\-5\.2 earns $616616as seller versus Qwen’s $467467\(Δ=\+149\.4\\Delta=\+149\.4\), and Gemini\-3\-Pro earns $587587as seller versus Qwen’s $315315\(Δ=\+272\.8\\Delta=\+272\.8,p<0\.001p<0\.001\); between the two stronger\-as\-seller providers, Gemini out\-earns GPT\-5\.2 \($843843versus $711711as sellers in the GPT↔\\leftrightarrowGemini pairing,Δ=−131\.6\\Delta=\-131\.6,p<0\.001p<0\.001\)\. The directional ordering, Gemini and GPT\-5\.2 extract dollar value as sellers while Qwen surrenders it, tracks the provider\-level bargaining profile documented in Section[5](https://arxiv.org/html/2608.07538#S5), and the round counts reinforce it: against Qwen3\-Max, rounds\-to\-agreement run4\.14\.1–4\.74\.7when GPT\-5\.2 or Gemini\-3\-Pro sells but only2\.22\.2–2\.42\.4when Qwen sells, so the price\-extracting party also bears the time cost of holding out\.
Table 16:Cross\-Model Configurations: Welfare DecompositionPairDirection / TestRoundsBuyer Profit \(Undisc\)Seller Profit \(Undisc\)Self\-play anchors:GPT\-5\.2—2\.68607\.2926\.4GPT\-5\-mini—2\.43655\.2776\.7Gemini\-3\-Pro—3\.31835\.8720\.5Qwen3\-Max—3\.751405\.0130\.8Cross\-Tier configuration:GPT\-5\.2 vs\. GPT\-5\-miniGPT\-5\.2 sells2\.46562\.1931\.0GPT\-5\-mini sells2\.03824\.4679\.8Δ\\Delta\(GPT\-5\.2−\-GPT\-5\-mini\)\+0\.43∗∗\-262\.3∗∗∗\+251\.2∗∗∗Cross\-Provider configurations:GPT\-5\.2 vs\. Gemini\-3\-ProGPT\-5\.2 sells3\.80844\.6711\.2Gemini\-3\-Pro sells2\.79690\.9842\.8Δ\\Delta\(GPT\-5\.2−\-Gemini\-3\-Pro\)\+1\.01∗∗∗\+153\.7∗∗∗\-131\.6∗∗∗GPT\-5\.2 vs\. Qwen3\-MaxGPT\-5\.2 sells4\.13918\.2616\.2Qwen3\-Max sells2\.24971\.9466\.8Δ\\Delta\(GPT\-5\.2−\-Qwen3\-Max\)\+1\.89∗∗∗\-53\.6\+149\.4Gemini\-3\-Pro vs\. Qwen3\-MaxGemini\-3\-Pro sells4\.70956\.7587\.3Qwen3\-Max sells2\.351157\.8314\.5Δ\\Delta\(Gemini\-3\-Pro−\-Qwen3\-Max\)\+2\.35∗∗∗\-201\.1∗∗∗\+272\.8∗∗∗
Notes:Each pair shows two direction\-mean rows \(A sells, B sells\) and one within\-pair difference row \(Δ\\Delta\)\. Means are agreed\-deal averages; rounds are conditional on agreement; each direction\-mean row is based onN=240N=240attempts\. Profits are absolute \(undiscounted\) payoffs in dollars\. Significance stars on theΔ\\Deltarow come from OLS regressions of the outcome on a seller\-identity indicator with design\-stratum fixed effects \(4 patience configurations×\\times2 buyer types×\\times2 first\-proposer assignments==16 strata\) and cluster\-robust standard errors at the stratum level:p†<0\.10\{\}^\{\\dagger\}\\,p<0\.10,p∗<0\.05\{\}^\{\*\}\\,p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}\\,p<0\.01,p∗∗∗<0\.001\{\}^\{\*\*\*\}\\,p<0\.001\.
### 9\.7Within\-Family Cross\-Model Negotiations: Detail
This appendix collects the supporting floats for the within\-family cross\-tier analysis summarized in Section[4\.5](https://arxiv.org/html/2608.07538#S4.SS5)\. Table[17](https://arxiv.org/html/2608.07538#S9.T17)reports the per\-configuration summary statistics \(deal rate, rounds, efficiency, discounted efficiency, and buyer share\) across the four homogeneous and cross\-tier configurations, and Figure[6](https://arxiv.org/html/2608.07538#S9.F6)visualizes the within\-family direction effects\.
Table 17:Cross\-Tier Configurations: Summary StatisticsSeller ModelBuyer ModelPairingNNDeal RateRoundsEfficiencyDisc\. EfficiencyBuyer ShareGPT\-5\.2GPT\-5\.2Symmetric240100\.0%2\.6898\.2%76\.9%37\.9%GPT\-5\-miniGPT\-5\-miniSymmetric240100\.0%2\.4394\.0%71\.7%41\.9%GPT\-5\.2GPT\-5\-miniAsymmetric240100\.0%2\.4696\.6%74\.7%35\.2%GPT\-5\-miniGPT\-5\.2Asymmetric240100\.0%2\.0397\.0%81\.5%52\.3%
Notes:Buyer Share is the buyer’s share of undiscounted total surplus\. The two symmetric rows are reused M1 benchmarks; the two asymmetric rows are the 480 new M3 experiments\. Each row contains 240 observations, so the table displays 960 observations in total\. The more capable model captures the larger share whether it buys or sells\. Comparing rows that share a seller model \(Rows 1 vs\. 3 and Rows 2 vs\. 4\), a GPT\-5\.2 buyer earns more than a GPT\-5\-mini buyer against the same seller \(37\.9%37\.9\\%vs\.35\.2%35\.2\\%;52\.3%52\.3\\%vs\.41\.9%41\.9\\%\)\. Comparing rows that share a buyer model \(Rows 1 vs\. 4 and Rows 2 vs\. 3\), a GPT\-5\.2 seller concedes less than a GPT\-5\-mini seller to the same buyer, a smaller buyer share indicating the seller conceded less \(37\.9%37\.9\\%vs\.52\.3%52\.3\\%;35\.2%35\.2\\%vs\.41\.9%41\.9\\%\)\.

Figure 6:Direction Effects in Cross\-Model NegotiationsNotes:Bars show direction means; error bars are 95% confidence intervals\. Brackets compare the two directional configurations\. Significance stars are based on OLS regressions of each outcome on seller identity with fixed effects for the 16 design strata \(four patience configurations×\\timestwo buyer types×\\timestwo first\-proposer assignments\) and cluster\-robust standard errors at the stratum level\. Significance levels:p†<0\.10\{\}^\{\\dagger\}\\,p<0\.10,p∗<0\.05\{\}^\{\*\}\\,p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}\\,p<0\.01,p∗∗∗<0\.001\{\}^\{\*\*\*\}\\,p<0\.001\.
### 9\.8Cross\-Family Flagship Negotiations: Detail
This appendix collects the full performance summary and supporting figures for the cross\-family flagship analysis summarized in Section[5\.2](https://arxiv.org/html/2608.07538#S5.SS2)\. Table[18](https://arxiv.org/html/2608.07538#S9.T18)reports deal rate, rounds, efficiency, discounted efficiency, buyer share, and the Bayesian reference for all six directional configurations\. Figure[7](https://arxiv.org/html/2608.07538#S9.F7)decomposes the descriptive sources of Gemini’s cross\-family advantage \(quantity gap, unit price, and surplus share across patience levels\), and Figure[8](https://arxiv.org/html/2608.07538#S9.F8)reports buyer\- and seller\-side concession patterns by patience configuration\.
Table 18:Cross\-Family Flagship Negotiations: Summary StatisticsPairDirectionDeal RateRoundsEfficiencyDiscEfficiencyBuyer ShareBayesian Ref\.GPT↔\\leftrightarrowGeminiGPT sells100\.0%3\.8099\.6%63\.2%54\.3%45\.2%GPT↔\\leftrightarrowGeminiGemini sells100\.0%2\.7997\.9%73\.2%43\.0%45\.2%GPT↔\\leftrightarrowQwenGPT sells100\.0%4\.1398\.4%57\.1%59\.1%45\.2%GPT↔\\leftrightarrowQwenQwen sells100\.0%2\.2491\.9%73\.9%66\.4%45\.2%Gemini↔\\leftrightarrowQwenGemini sells100\.0%4\.7099\.0%51\.8%61\.2%45\.2%Gemini↔\\leftrightarrowQwenQwen sells100\.0%2\.3594\.0%71\.3%79\.0%45\.2%
Notes:We use DiscEfficiency to denote discounted efficiency\. The Bayesian column reports the PBE buyer share benchmark under asymmetric demand information, evaluated at the patience configuration and proposer order of each cell\. Each pair is tested in both directions across the 16 design strata \(four patience configurations×\\timestwo buyer types×\\timestwo first\-proposer assignments\)\.
##### Additional cross\-flagship statistics\.
Three descriptive patterns supplement the figures\. First, on alignment with the Bayesian benchmark, under seller\-patient conditions \(δB=0\.4\\delta\_\{B\}=0\.4,δS=0\.9\\delta\_\{S\}=0\.9\), where the Bayesian benchmark allocates the buyer only 13\.1% of surplus, Gemini achieves 50\.2% \(\+\+37\.1 pp deviation\), while GPT achieves 35\.0% \(\+\+21\.8 pp\) and Qwen 53\.1% \(\+\+40\.0 pp\); these cross\-family buyer advantages thus reflect not closer alignment with the benchmark reference but a consistent ability to extract surplus above it\. Second, on sensitivity to the buyer’s informational position, Gemini’s surplus share varies least across buyer types \(high: 66\.1%, low: 67\.2%, gap:−\-1\.1 pp\), compared to GPT \(58\.6% vs\. 50\.8%, gap: 7\.8 pp\) and Qwen \(62\.5% vs\. 57\.8%, gap: 4\.6 pp\)\. Third, on the efficiency decomposition for the Gemini↔\\leftrightarrowQwen direction, the total gap from first\-best is 3\.50 pp, dominated by suboptimal contract terms \(3\.42 pp, 97\.7%\), with no deal failures \(both directions reach agreement in all negotiations\) and a negligible irrational component \(0\.08 pp\)\.

Figure 7:Sources of Model Advantage in Cross\-Flagship NegotiationsNotes:Sample restricted to negotiations that reached agreement\. Panel \(a\) quantity gap = \(optimal−\-actual\) / optimal\. Panel \(b\) dashed horizontal line marks the production cost \($30/unit\) as reference\. Error bars are 95% confidence intervals\.

Figure 8:Concession Patterns in Cross\-Flagship Negotiations by Patience ConfigurationNotes:Sample restricted to negotiations that reached agreement\. Panel \(a\) Buyer concession: Qwen’s buyer concessions are nearly flat across patience cells \($5\.5–$6\.7 per unit\), whereas Gemini’s decline from $7\.6 under seller\-patient conditions to $2\.1 under buyer\-patient conditions, where Qwen concedes approximately2\.6×2\.6\\timesas much as Gemini\. Panel \(b\) Seller concession: Qwen barely concedes as seller, while GPT makes the largest seller\-side adjustments\. Error bars are 95% confidence intervals\.
## 10Strategic Behavior Details
This appendix provides detailed analyses supporting the process and strategic\-behavior evidence summarized in Section[4\.3](https://arxiv.org/html/2608.07538#S4.SS3)and Section[5](https://arxiv.org/html/2608.07538#S5)\. We present comprehensive anchoring analysis, LLM\-inferred buyer strategy distributions, information disclosure patterns, deceptive tactics impact, and extended reasoning content analysis\.
### 10\.1An Illustrative Negotiation
Figure[9](https://arxiv.org/html/2608.07538#S10.F9)traces one representative Gemini\-3\-Pro self\-play negotiation under buyer\-first, high\-patience conditions \(δB=0\.9\\delta\_\{B\}=0\.9,δS=0\.9\\delta\_\{S\}=0\.9, high\-type buyer\)\. It illustrates the paper’s central empirical pattern: the agents converge on the efficient quantity but take two rounds rather than the Bayesian benchmark of one\.
round 1
B
BuyerPROPOSE
Qty:80Payment:$2,500Unit Price:$31\.25“I’d like to propose a high volume order to kickstart our relationship\. This provides you with a solid margin above your production costs\.”
S
SellerREJECT
“While I appreciate the high volume, the total payment leaves me with almost no margin\. We need a more equitable split of the profits\.”
round 2
S
SellerPROPOSE
Qty:80Payment:$3,500Unit Price:$43\.75“I am willing to meet your requested quantity of 80 units, but we need a price that properly reflects both my costs and the market value\. How does this total payment look to you?”
B
BuyerACCEPT
“This is a fair compromise that works for both of us\. I accept your offer to move forward\.”
Figure 9:Example of an LLM Negotiation Process \(Gemini\-3\-Pro Self\-Play Negotiation\)
### 10\.2Detailed Anchoring Analysis
We examine how opening offers shape final contract terms\. Figure[10](https://arxiv.org/html/2608.07538#S10.F10)displays opening\-offer patterns compared against the Bayesian benchmark for bargaining shares, summarizing the buyer\-side type separation and seller\-side expected\-value pooling reported in Section[4\.3](https://arxiv.org/html/2608.07538#S4.SS3)\.
Figure 10:Opening Offers and Anchoring EffectsTable[19](https://arxiv.org/html/2608.07538#S10.T19)presents correlation coefficients between opening and final terms across proposer conditions and buyer types\.
Table 19:Correlations Between Opening and Final Contract TermsProposerBuyer TypeQuantitypp\-valuePaymentpp\-valueCorrelationCorrelationBuyer\-First NegotiationsOverallAll0\.911<<0\.0010\.862<<0\.001High\-type0\.349<<0\.0010\.687<<0\.001Low\-type0\.310<<0\.0010\.628<<0\.001Seller\-First NegotiationsOverallAll0\.164<<0\.0010\.381<<0\.001High\-type0\.435<<0\.0010\.725<<0\.001Low\-type0\.269<<0\.0010\.478<<0\.001
Notes:Pearson correlation coefficients between opening offer terms and final negotiated contract terms\. Buyer\-first N=1,080; Seller\-first N=1,080\.
Buyer\-First Negotiations Show Type Separation, Moderate Quantity Anchoring, and Strong Payment Anchoring\.When buyers propose first, overall quantity correlation of 0\.911 indicates opening quantities strongly predict final outcomes\. This high correlation reflects type separation: buyers signal their type through initial quantity proposals, and these signals persist into final contracts\. Decomposing by buyer type shows that both types anchor at moderate strength: high\-type buyers show a quantity correlation of 0\.349 and low\-type buyers 0\.310, indicating that opening quantities meaningfully predict final terms for both, with substantial adjustment through bargaining\.
Payment anchoring follows different patterns\. Both buyer types exhibit strong payment correlations, with high\-types at 0\.687 and low\-types at 0\.628, both highly significant\. This suggests that while quantities may adjust through negotiation, payment terms remain anchored to opening proposals regardless of buyer type\. The asymmetry between quantity and payment anchoring is best read as descriptive process evidence that different parts of the contract adjust at different rates during bargaining\.
Seller\-First Negotiations Show Weaker Anchoring\.When sellers propose first, both quantity and payment correlations decline substantially\. Overall quantity correlation drops to 0\.164, and payment correlation to 0\.381, indicating sellers’ opening offers provide weaker anchors than buyers’ proposals\. This asymmetry may reflect sellers’ pooling strategy: by proposing expected\-value quantities averaging 60\.5 units rather than the branch\-specific screening contracts in the Bayesian benchmark, sellers create opening positions subject to greater adjustment through negotiation\.
Decomposing seller\-first negotiations by buyer type reveals that anchoring strength varies systematically with the realized type\. Quantity correlations are moderate when the buyer is a high\-type \(r=0\.435r=0\.435,p<0\.001p<0\.001\) but weaker when the buyer is a low\-type \(r=0\.269r=0\.269,p<0\.001p<0\.001\)\. Payment correlations follow the same ordering and are noticeably stronger than the corresponding quantity correlations \(r=0\.725r=0\.725for high\-types andr=0\.478r=0\.478for low\-types, bothp<0\.001p<0\.001, versusr=0\.435r=0\.435andr=0\.269r=0\.269for quantity\)\. The within\-type correlations are sharper than the pooled estimates because pooling mixes negotiations whose realized contracts diverge in opposite directions from a common opening offer\. Even for high\-types, where anchoring is strongest, correlations remain well below the buyer\-first benchmark, consistent with seller opening offers establishing reference points that influence but do not determine final terms\.
### 10\.3Inferred Buyer Strategy Distributions
Because the experimental protocol records public messages and private reasoning traces but does not elicit self\-reported strategy labels, we classify buyer strategy from the recorded turn\-level text using an independent language\-model coder \(Claude Haiku 4\.5\)\. For each buyer turn, the coder observes the public message and private reasoning trace and assigns one of four categories:signal\(truthfully conveying the buyer’s type to separate\),mimic\(misrepresenting type by posing as the other type\),optimal\(straightforward payoff maximization with no type\-strategic intent\), orother\. Table[20](https://arxiv.org/html/2608.07538#S10.T20)reports the distribution by buyer type and patience, Table[21](https://arxiv.org/html/2608.07538#S10.T21)by capability tier, and Figure[11](https://arxiv.org/html/2608.07538#S10.F11)visualizes the type\-and\-patience breakdown\.

Figure 11:LLM\-Inferred Buyer Strategy by Type and PatienceNotes:Bars report average within\-negotiation shares of classified buyer turns; error bars are 95% confidence intervals\.
Table 20:LLM\-Inferred Buyer Strategy: by Type and PatienceConditionNNOptimalSignalMimicOther\(%\)\(%\)\(%\)\(%\)By Buyer TypeHigh\-Type1,07157\.521\.913\.86\.8Low\-Type1,07931\.864\.71\.32\.2Holm\-adjustedpp\-value<<0\.001<<0\.001<<0\.001<<0\.001By Patience Configuration \(High\-Type Buyers\)\(0\.9, 0\.9\)26953\.219\.817\.89\.2\(0\.7, 0\.9\)26956\.723\.214\.06\.2\(0\.4, 0\.9\)26665\.320\.99\.24\.5\(0\.9, 0\.4\)26754\.823\.714\.07\.4Holm\-adjustedpp\-value<<0\.0010\.340<<0\.0010\.013By Patience Configuration \(Low\-Type Buyers\)\(0\.9, 0\.9\)27030\.465\.01\.63\.1\(0\.7, 0\.9\)27030\.966\.11\.01\.9\(0\.4, 0\.9\)27037\.660\.80\.51\.1\(0\.9, 0\.4\)26928\.566\.72\.02\.9Holm\-adjustedpp\-value0\.0270\.2890\.2890\.029
Notes:The sample comprises 6,279 classified buyer turns across 2,150 negotiations\. Each percentage is the average within\-negotiation share of classified buyer turns\. The buyer\-type panel uses independent Welch tests; the patience panels use Welch ANOVA\. Holm correction is applied across the four strategy outcomes within each panel\.
Type Asymmetry Aligns with Theory\.The two type\-strategic categories split sharply, in the directionsFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)predict\. The average mimic rate per negotiation is13\.8%13\.8\\%for high\-type buyers versus1\.3%1\.3\\%for low\-type buyers, consistent with the high type’s incentive to pose as a low\-demand buyer and avoid surplus extraction\. Conversely, the average signaling rate per negotiation is64\.7%64\.7\\%for low\-type buyers versus21\.9%21\.9\\%for high\-type buyers, consistent with the low type’s incentive to separate credibly\. Both differences remain significant after Holm correction \(p<0\.001p<0\.001\), and non\-type\-strategic “optimal” play is the modal high\-type category \(57\.5%57\.5\\%\)\.
Patience Effects Contradict Theory\.The comparative static, however, does not align\.Feng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)predict that high\-type mimicking should intensify when buyer patience is low, since a revealed high\-type is then more exposed to seller extraction\. We find the opposite: the average mimic rate per negotiation islowestunder the most buyer\-impatient configuration \(δB=0\.4\\delta\_\{B\}=0\.4,δS=0\.9\\delta\_\{S\}=0\.9\) at9\.2%9\.2\\%andhighestunder symmetric high patience \(δB=δS=0\.9\\delta\_\{B\}=\\delta\_\{S\}=0\.9\) at17\.8%17\.8\\%\(Holm\-adjusted Welch ANOVA,p<0\.001p<0\.001\)\. Agents thus exhibit behavior consistent with high\-type mimicking but do not calibrate it to the patience\-based incentive the theory emphasizes; this is the same qualitative anomaly that the buyer\-side quantity distortions in Section[4\.3](https://arxiv.org/html/2608.07538#S4.SS3)exhibit\.
Model Capability Effects\.The average mimic rate per negotiation rises with model capability \(Table[21](https://arxiv.org/html/2608.07538#S10.T21)\):11\.1%11\.1\\%for flagship buyers,6\.9%6\.9\\%for mid\-tier buyers, and4\.5%4\.5\\%for baseline buyers \(Holm\-adjusted Welch ANOVA,p<0\.001p<0\.001\)\. More capable buyers are more often classified as mimicking the other type, consistent with the reasoning capacity a pooling strategy requires; signaling, by contrast, is common across all tiers \(37\.737\.7–50\.6%50\.6\\%\)\. This gradient is consistent with greater strategic sophistication, though the inferred mimic rate should be read as approximate\.
Table 21:LLM\-Inferred Buyer Strategy by Model Capability TierModel TierNNOptimal \(%\)Signal \(%\)Mimic \(%\)Other \(%\)Flagship71840\.341\.711\.16\.9Mid\-tier71551\.137\.76\.94\.2Baseline71742\.550\.64\.52\.4Holm\-adjustedpp\-value<<0\.001<<0\.001<<0\.001<<0\.001
Notes:The sample comprises 6,279 classified buyer turns across 2,150 negotiations\. Each percentage is the average within\-negotiation share of classified buyer turns\. Welch ANOVA compares capability tiers, with Holm correction across the four strategy outcomes\.
### 10\.4Information Disclosure and Public–Private Divergence
We first describe buyers’ disclosure rates by type, then turn to the four public–private divergence tactics and their effect on outcomes\. A language\-model classifier999Classification was performed using Claude Haiku 4\.5, which is independent of all three studied agent vendors \(OpenAI, Google, Alibaba\)\.coded each buyer turn’s public message into disclosure categories: direct truthful \(explicitly revealing type\), indirect truthful \(implying type through context\), deceptive \(posing as the other type\), or withholding \(avoiding disclosure\)\. Table[22](https://arxiv.org/html/2608.07538#S10.T22)reports the resulting rates\.
Table 22:Information Disclosure Patterns by Buyer TypeBuyerNAnyDirectIndirectDeceptiveWithholdTypeTruthfulTruthfulTruthful\(%\)\(%\)\(%\)\(%\)\(%\)High\-Type1,07848\.227\.221\.04\.047\.8Low\-Type1,08052\.121\.330\.81\.446\.5Difference\-3\.9\*\+5\.9\*\*\*\-9\.8\*\*\*\+2\.6\*\*\*\+1\.2
Notes:Each percentage is the average within\-negotiation share of classified buyer turns\.NNdenotes negotiations\. The difference row reports high\-type minus low\-type percentage points\. Stars are based on independent Welch tests with Holm correction across the five outcomes\. \*\*\*p<0\.001p<0\.001, \*\*p<0\.01p<0\.01, \*p<0\.05p<0\.05\.
Disclosure Rate Asymmetries\.The average truthful\-disclosure rate per negotiation is higher for low\-type than high\-type buyers \(52\.1 versus 48\.2%, Holm\-adjustedp<0\.05p<0\.05\)\. The two types differ in mode, however: high\-types usedirectdisclosure more \(27\.2 versus 21\.3%, Holm\-adjustedp<0\.001p<0\.001\), whereas low\-types rely more onindirectdisclosure \(30\.8 versus 21\.0%, Holm\-adjustedp<0\.001p<0\.001\), so the net truthfulness gap is carried by indirect statements\. High\-types withhold marginally more \(47\.8 versus 46\.5%, not significant\)\. At the net level the pattern aligns with signaling theory: low\-types disclose more to separate from high\-types, while high\-types more often withhold or send deceptive type signals\.
Deception Concentrated Among High\-Types\.The average deceptive\-disclosure rate per negotiation is low but higher among high\-type buyers \(4\.0 versus 1\.4%, Holm\-adjustedp<0\.001p<0\.001\)\. This aligns with the theoretical prediction that high\-types have a structural incentive to mimic low\-types to avoid surplus extraction\.
From disclosure to public–private divergence\.The disclosure rates above describe what the public message asserts about the buyer’s type; they do not speak to whether the public framing aligns with private reasoning\. We next compare buyers’ public messages to their private reasoning traces and identify four recurring forms of public–private divergence\.Fairness Claim Mismatchoccurs when buyers publicly appeal to fairness while privately reasoning in purely self\-interested terms\.Profit Hidingoccurs when buyers conceal or understate expected profitability, portraying acceptable terms as marginally viable when private reasoning reveals substantial gains\.False ConstraintandFalse Finalityoccur when buyers claim non\-existent constraints or declare offers final while privately intending continued negotiation\. These labels come from a language\-model classifier rather than validated human annotation and should be read as exploratory indicators of public–private divergence rather than validated prevalence estimates of intentional deception\. Table[23](https://arxiv.org/html/2608.07538#S10.T23)reports prevalence\.
Table 23:Public–Private Message Divergence TacticsTacticPrevalence \(%\)Observed CountFairness Claim Mismatch56\.57,352Profit Hiding14\.11,831False Constraint9\.41,222False Finality0\.791Total Turns Analyzed13,023
Notes:This table is descriptive at the turn level: percentages are calculated over all 13,023 cache\-hit turns, and some turns exhibit multiple tactics\. Negotiation\-level role comparisons appear in Table[24](https://arxiv.org/html/2608.07538#S10.T24)\.
The pooled turn\-level prevalence in Table[23](https://arxiv.org/html/2608.07538#S10.T23)averages across two distinctions that turn out to matter: which side of the negotiation produced the turn, and which provider family generated the agent\. Table[24](https://arxiv.org/html/2608.07538#S10.T24)instead compares average buyer and seller tactic rates within the same negotiation\. Using buyer\-minus\-seller differences, the split reveals that the asymmetry is not uniform across families\. Within OpenAI, buyer and seller fairness\-mismatch rates are nearly identical \(\+0\.1\+0\.1pp\), and the\+2\.8\+2\.8pp profit\-hiding difference is not significant after Holm correction\. Google buyers have lower rates than Google sellers for both fairness mismatch \(−9\.3\-9\.3pp\) and profit hiding \(−3\.0\-3\.0pp\), whereas Qwen buyers have higher rates than Qwen sellers \(\+8\.7\+8\.7and\+3\.3\+3\.3pp, respectively\)\. These patterns suggest that public–private divergence in framing is a population\-average behavior, but one whose role incidence varies in sign across provider families\.
Table 24:Public–Private Divergence Tactics by Provider Family and RoleFamilyRoleNFairness MismatchProfit HidingFalse ConstraintFalse Finality\(%\)\(%\)\(%\)\(%\)OpenAIBuyer72039\.117\.05\.81\.2Seller72039\.114\.29\.41\.2Δ\\Delta\(B−\-S, pp\)\+0\.1\+2\.8\-3\.6\*\*\*0\.0GoogleBuyer71863\.210\.18\.20\.2Seller71872\.513\.16\.70\.1Δ\\Delta\(B−\-S, pp\)\-9\.3\*\*\*\-3\.0\*\*\+1\.5\+0\.1QwenBuyer71859\.415\.811\.90\.2Seller71850\.812\.55\.80\.3Δ\\Delta\(B−\-S, pp\)\+8\.7\*\*\*\+3\.3\*\*\+6\.1\*\*\*\-0\.1
Notes:Each percentage is the average within\-negotiation share of cache\-hit turns exhibiting the tactic\.NNdenotes negotiations with classified buyer and seller turns\. TheΔ\\Deltarow reports the within\-negotiation buyer\-minus\-seller difference in percentage points\. Stars are based on paired t\-tests with Holm correction across the four tactics within each family\. \*\*\*p<0\.001p<0\.001, \*\*p<0\.01p<0\.01, \*p<0\.05p<0\.05\.
### 10\.5Process Traces: Why Agents Take Extra Rounds
This appendix presents the full process\-trace evidence behind the round\-count gap summarized in Section[4\.3](https://arxiv.org/html/2608.07538#S4.SS3)\. The round\-count gap documented in the main text has a clear process\-level signature: LLM agents recognize the structure of the bargaining problem but substitute robust heuristics for the branch\-specific execution the PBE prescribes\. We document the substitution in two places, opening offers and inferred strategies, and show that both contribute to the deviation from equilibrium timing\.
On opening offers, the substitution depends on which side moves first\. When buyers open, they reproduce the type separation predicted by theFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)truth\-telling equilibrium: high\-type quantities average77\.877\.8units and low\-type quantities average41\.641\.6units, with opening\-to\-final correlations of0\.9110\.911for quantity and0\.8620\.862for payment\. When sellers open, by contrast, they propose60\.560\.5units on average, essentially the uniform\-prior expected value of6060, rather than the branch\-specific screening menu the PBE prescribes; opening\-to\-final correlations fall to0\.1640\.164for quantity and0\.3810\.381for payment\. This pooling substitution contributes to the timing gap on the seller\-first side: a seller who opens at a pooling quantity cannot reach the separating contract in round one and defers separation to later rounds\. But seller\-first negotiations are not measurably slower than buyer\-first ones \(rounds do not differ significantly by first\-proposer order,F=0\.12F=0\.12, n\.s\.; Table[4](https://arxiv.org/html/2608.07538#S4.T4)\), so pooling substitution alone does not explain the pooled2\.982\.98\-round average against the1\.251\.25\-round equilibrium prediction; on the buyer\-first side, where opening offers already separate by type, the residual delay instead reflects post\-agreement price haggling over an already\-anchored quantity \(Appendix[10\.1](https://arxiv.org/html/2608.07538#S10.SS1)\)\. Wholesale prices show minimal type\-based differentiation in either case, with medians at4040for both buyer types, indicating that LLM agents recover the quantity\-separating signal but not the type\-dependent price separation\.
A complementary signature appears in how agents handle the high type’s mimicking incentive \(Appendix[10\.3](https://arxiv.org/html/2608.07538#S10.SS3)\)\. The average mimic rate per negotiation is13\.8%13\.8\\%for high\-type buyers versus1\.3%1\.3\\%for low\-type buyers \(Holm\-adjustedp<0\.001p<0\.001\)\. High\-type realized quantities are also7\.1%7\.1\\%below first\-best, consistent with this interpretation, although the quantity distortion alone does not establish mimicking\. The comparative static, however, runs the wrong way: theory predicts more high\-type mimicking under low buyer patience, but the average mimic rate instead declines from17\.8%17\.8\\%under symmetric high patience to9\.2%9\.2\\%under low buyer patience \(Holm\-adjusted Welch ANOVA,p<0\.001p<0\.001\)\. Agents enact the static mimicking incentive yet do not calibrate it to the dynamic patience incentive on which the equilibrium turns, a behavior–theory gap that accompanies the implementation failures behind the timing\-and\-efficiency gap\.
### 10\.6Process Traces: How Buyers Exploit the Verbal Channel
The provider\-level patterns documented in Section[5](https://arxiv.org/html/2608.07538#S5)are accompanied by a role\-asymmetric pattern in how the verbal channel gets used, though the asymmetry is itself provider\-conditional\. An LLM classifier flags fairness\-claim mismatch \(public appeals to fairness paired with self\-interested private reasoning\) in roughly56%56\\%of classified turns overall and profit hiding in about14%14\\%\. As detailed in Appendix[10\.4](https://arxiv.org/html/2608.07538#S10.SS4)\(Table[24](https://arxiv.org/html/2608.07538#S10.T24)\), this asymmetry is provider\-specific: Google buyers show markedly lower rates than Google sellers, OpenAI shows no meaningful gap, and Qwen’s gap runs the opposite way\.
### 10\.7Strategic Patience: Per\-Model Heterogeneity \(R3\)
This subsection reports the full strategic\-patience analysis summarized in Section[6\.3](https://arxiv.org/html/2608.07538#S6.SS3)of the main paper: the headline gain table \(Table[25](https://arxiv.org/html/2608.07538#S10.T25)\), the buyer\-payoff curves by prompted patience \(Figure[12](https://arxiv.org/html/2608.07538#S10.F12)\), and the per\-model decomposition \(Figure[13](https://arxiv.org/html/2608.07538#S10.F13)\) showing which models translate structural patience into realized payoff\.
Table 25:Gains from Optimally Prompted Strategic PatiencePartyTrueδecon\\delta^\{\\mathrm\{econ\}\}Matchedavg\-optimal modelsMatchedper\-condition optimalMean gain\[\[95% CI\]\]ConclusionBuyer0\.94/920/36 \(56%\)\+74\.81\+74\.81\[\+29\.43,\+130\.17\]\[\+29\.43,\\,\+130\.17\]Matched\-high modal;understate for mostSeller0\.99/926/36 \(72%\)\+18\.28\+18\.28\[\+4\.35,\+35\.06\]\[\+4\.35,\\,\+35\.06\]Matched\-high dominantBuyer0\.74/913/36 \(36%\)\+41\.11\+41\.11\[\+25\.11,\+58\.41\]\[\+25\.11,\\,\+58\.41\]Matched most common;3 overstate, 2 understateSeller0\.72/911/36 \(31%\)\+66\.53\+66\.53\[\+39\.05,\+100\.50\]\[\+39\.05,\\,\+100\.50\]High patience dominant;two matched exceptions
Notes:For each cell \(model×\\timesbuyer\-type×\\timesfirst\-proposer\), we identify the empirically best promptedδstrat\\delta^\{\\mathrm\{strat\}\}\(including the matched valueδstrat=δecon\\delta^\{\\mathrm\{strat\}\}=\\delta^\{\\mathrm\{econ\}\}as one of the alternatives\) and compute gain==best payoff−\-matched payoff≥0\\geq 0\.Per\-condition optimalreports the count of cells in which the matched value is itself the best;avg\-optimal modelsreports the count of models for which the matched value yields the highest mean payoff after averaging across the four conditions\.Mean gainaverages the cell\-level gains across the 36 cells; positive values indicate that some strategic alternative beats matched on average\. Brackets are 95% percentile bootstrap CIs over the 36 cells \(2,000 resamples\)\.

Figure 12:Buyer Payoff by Prompted Strategic PatienceNotes\.Buyer payoff atδBecon=0\.9\\delta^\{\\text\{econ\}\}\_\{B\}=0\.9is plotted as a function of prompted strategic patienceδBstrat\\delta^\{\\text\{strat\}\}\_\{B\}, separately for H\-type and L\-type buyers\. Panel A restricts to negotiations in which the buyer proposes first; Panel B to those in which the seller proposes first\. Each line averages across the nine main verbal models, with shaded bands giving 95% confidence intervals across models\. The right edge of each panel and the dashed vertical line both mark the matched choiceδBstrat=δBecon=0\.9\\delta^\{\\text\{strat\}\}\_\{B\}=\\delta^\{\\text\{econ\}\}\_\{B\}=0\.9\.

Figure 13:Per\-model Gain from Optimal Strategic PatienceNotes\.Bars show the gain from choosing the empirically best promptedδstrat\\delta^\{\\text\{strat\}\}relative to the matched baselineδstrat=δecon=0\.9\\delta^\{\\text\{strat\}\}=\\delta^\{\\text\{econ\}\}=0\.9, averaged across the four buyer type×\\timesfirst proposer conditions\. Panel A: buyers\. Panel B: sellers\. Near\-zero gray bars indicate that the matched baseline \(δstrat=δecon\\delta^\{\\text\{strat\}\}=\\delta^\{\\text\{econ\}\}\) is average\-optimal\. The pattern is asymmetric across roles\. On the seller side \(Panel B\) matching is average\-optimal for all nine models, so every bar is essentially flat\. On the buyer side \(Panel A\) five of nine models realize higher payoff by understating patience toδBstrat=0\.7\\delta^\{\\text\{strat\}\}\_\{B\}=0\.7\(or0\.40\.4for Qwen3\-Max\), with the largest gains for GPT\-4o\-mini \(\+278\+278\) and GPT\-5\-mini \(\+234\+234\); the four remaining models \(GPT\-5\.2 and the three Gemini models\) are average\-optimal at the matched value\. For buyer models that benefit from understating patience, the gain arises through higher agreement rates, more favorable terms, or shorter bargaining, depending on the model\.
## 11Design and Robustness Extension Program
This appendix reports the complete design and robustness extension program supporting the paper’s main findings\. Section[11\.1](https://arxiv.org/html/2608.07538#S11.SS1)lists every LLM\-to\-LLM negotiation conducted across both the main analysis and the design and robustness extensions\. The seven completed extensions are reported in full in the paper: a structured\-communication treatment with verbal messages removed \(R1; Section[11\.2](https://arxiv.org/html/2608.07538#S11.SS2)\), a no\-discounting bargaining treatment that removes the discounting framework from agent prompts \(R2; Section[11\.3](https://arxiv.org/html/2608.07538#S11.SS3)\), a strategic\-patience analysis \(R3; Section[10\.7](https://arxiv.org/html/2608.07538#S10.SS7)\), a reasoning\-effort ablation on GPT\-5\.2 \(R4; Section[11\.4](https://arxiv.org/html/2608.07538#S11.SS4)\), a retail\-price robustness check on the OpenAI family \(R5; Section[11\.5](https://arxiv.org/html/2608.07538#S11.SS5)\), a sensitivity analysis over the seller’s prior belief about buyer type \(R6; Section[11\.6](https://arxiv.org/html/2608.07538#S11.SS6)\), and a parameter\-size ablation within the Qwen3 family \(R7; Section[11\.7](https://arxiv.org/html/2608.07538#S11.SS7)\)\. R1–R3 are the configuration analyses reported in Section[6](https://arxiv.org/html/2608.07538#S6); R4–R7 are the additional robustness extensions\. The extensions probe whether the headline findings of Sections[4](https://arxiv.org/html/2608.07538#S4)–[5](https://arxiv.org/html/2608.07538#S5)generalize beyond the experimental choices made in the main analysis, along dimensions of communication mode, time\-preference framing, prompted strategic patience, inference\-time compute, surplus magnitude, prior beliefs, and model size\.
### 11\.1Complete Experimental Inventory
Table[26](https://arxiv.org/html/2608.07538#S11.T26)accounts for all LLM\-to\-LLM negotiations conducted in this paper\. The inventory is organized into two tiers: the main analysis \(M1–M3\), whose findings are reported in the main text, and the design and robustness extensions \(R1–R7\), whose findings are reported in Section[6](https://arxiv.org/html/2608.07538#S6)and the appendix sections referenced above\. M1–M3 together contribute 4,080 negotiations and 12,490 bargaining rounds; the completed design and robustness extensions \(R1–R7\) contribute an additional 5,760 negotiations and 19,102 rounds\. The full program therefore totals 9,840 unique negotiations and 31,592 rounds across nine models from three providers\.
Table 26:Experiment Inventory: Main Analysis and Design and Robustness ExtensionsCondition FactorsBlockBuyer TypesFirst ProposersPatience LevelsOtherCond\.ModelsExpsRoundsRepsMain AnalysisM1: Verbal mainfactorial2 \(H, L\)2 \(B, S\)4—169 \(3 providers×\\times3 tiers\)2,1606,60715\.0M2: Cross\-familyflagship2 \(H, L\)2 \(B, S\)42 direc\-tions163 \(GPT\-5\.2,Gemini\-3\-Pro,Qwen3\-Max\)1,4404,80515\.0M3: Cross\-capability2 \(H, L\)2 \(B, S\)42 directions321 pair \(GPT\-5\.2↔\\leftrightarrowGPT\-5\-mini\)4801,07815\.0Main analysis subtotal4,08012,490Design and Robustness ExtensionsR1: Structuredcommunication2 \(H, L\)2 \(B, S\)4—169 \(3 providers×\\times3 tiers\)2,1607,33115\.0R2: No\-discountingbargaining2 \(H, L\)2 \(B, S\)—noδ\\delta49 \(3 providers×\\times3 tiers\)5402,22915\.0R3: Strategic patience\(δB,δS\)=\(0\.9,0\.7\)\(\\delta\_\{B\},\\delta\_\{S\}\)=\(0\.9,0\.7\)2 \(H, L\)2 \(B, S\)1—49 \(all main models\)540a1,75515\.0R4: Reasoning effort\(GPT\-5\.2 only\)2 \(H, L\)2 \(B, S\)43 efforts481 \(low/med/high\)480b1,31615\.0R5: Retail price 120\(OpenAI models\)2 \(H, L\)2 \(B, S\)4—163 \(OpenAI family\)7202,32815\.0R6: Prior sensitivity\(β∈\{0\.3,0\.5,0\.7\}\\beta\\in\\\{0\.3,0\.5,0\.7\\\}\)2 \(H, L\)2 \(B, S\)13 priors129 \(all main models\)1,080c3,11615\.0R7: Parameter size\(Qwen3\-14B\)2 \(H, L\)2 \(B, S\)4—161 \(Qwen3\-14B\)240d1,02715\.0Extension subtotal \(R1–R7\)5,76019,102Full program total9,84031,592
Notes:M1–M3 constitute the main analysis reported in the main text; R1–R7 are the design and robustness extensions reported in full in Appendix[11](https://arxiv.org/html/2608.07538#S11)\. H = high\-type buyer, L = low\-type buyer\. B = buyer proposes first, S = seller proposes first\. Patience levels: \(δB,δS\\delta\_\{B\},\\delta\_\{S\}\)∈\\in\{\(0\.9, 0\.9\), \(0\.9, 0\.4\), \(0\.4, 0\.9\), \(0\.7, 0\.9\)\}\. M3 counts 480 new cross\-tier experiments; the 480 observations from the two homogeneous M1 benchmarks displayed in Table[17](https://arxiv.org/html/2608.07538#S9.T17)are not counted again\. Efforts: low, medium, high reasoning\.a540 new strategic\-patience experiments; the full analysis reuses 2,160 M1 observations \(N=2,700N=2\{,\}700\)\.b480 new experiments \(low \+ high effort\); the medium\-effort baseline \(\+240\) is reused from M1 and not counted in the R4 row or in the grand total\.c1,080 new experiments \(β=0\.3\\beta=0\.3andβ=0\.7\\beta=0\.7\); theβ=0\.5\\beta=0\.5baseline \(\+540\) is reused from M1 and not counted in the R6 row or in the grand total\.d240 new Qwen3\-14B experiments; the Qwen3\-32B and Qwen3\-Max legs of theN=720N=720three\-size comparison reported in Section[11\.7](https://arxiv.org/html/2608.07538#S11.SS7)\(\+480\) are reused from M1 self\-play data and not counted in the R7 row or in the grand total\.
### 11\.2Structured Communication \(R1\): Model\-Level Detail
This subsection reports the detailed results underlying the structured\-communication treatment \(R1\) summarized in Section[6\.1](https://arxiv.org/html/2608.07538#S6.SS1)\. Table[27](https://arxiv.org/html/2608.07538#S11.T27)gives the full performance comparison across agreement, rounds, and efficiency by model tier, first proposer, buyer type, and patience, with between\-treatment tests\. Removing verbal communication measurably changes execution but does not overturn the main performance pattern\. Structured\-only bargaining lowers agreement overall \(χ2=27\.41\\chi^\{2\}=27\.41,p<0\.001p<0\.001\), increases rounds modestly \(\|t\|=3\.23\|t\|=3\.23,p<0\.01p<0\.01\), and reduces undiscounted efficiency \(\|t\|=5\.05\|t\|=5\.05,p<0\.001p<0\.001\)\. The agreement decline is concentrated outside the flagship tier, while efficiency falls significantly in each tier; nevertheless, structured agents still reach high absolute agreement \(96\.4%96\.4\\%\) and efficiency \(92\.8%92\.8\\%\), with only a modest increase in rounds \(2\.982\.98to3\.153\.15\)\. Table[28](https://arxiv.org/html/2608.07538#S11.T28)then reports, for each model, buyer surplus share, efficiency, and rounds under the verbal and structured treatments, with within\-model significance tests\.
Table 27:Structured vs\. Verbal AgentsConditionStructured AgentsVerbal AgentsAgreementRoundsEfficiencyAgreementRoundsEfficiency\(%\)\(%\)\(%\)\(%\)Overall96\.43\.1592\.898\.92\.9895\.4By Model Tier:Flagship100\.03\.2398\.3100\.03\.2598\.9Mid\-tier98\.83\.3992\.599\.92\.9396\.4Baseline90\.62\.8187\.796\.82\.7591\.0By First Proposer:Buyer First95\.53\.1092\.898\.42\.9794\.8Seller First97\.43\.2092\.999\.42\.9996\.1By Buyer Type:High\-type97\.02\.9493\.099\.52\.8895\.1Low\-type95\.83\.3692\.798\.23\.0895\.7By Patience:\(0\.9, 0\.9\)97\.83\.5794\.798\.03\.3995\.2\(0\.7, 0\.9\)96\.93\.3393\.199\.43\.0996\.1\(0\.4, 0\.9\)94\.62\.8690\.299\.32\.6395\.0\(0\.9, 0\.4\)96\.52\.8393\.498\.92\.8295\.5Between\-Treatment Tests \(Structured vs Verbal\)Overall27\.41\*\*\*3\.23\*\*5\.05\*\*\*———By Model Tier:Flagship0\.000\.312\.58\*\*———Mid\-tier4\.93\*4\.54\*\*\*5\.18\*\*\*———Baseline22\.71\*\*\*0\.522\.57\*———By First Proposer:Buyer First15\.02\*\*\*1\.69†2\.51\*———Seller First11\.62\*\*\*2\.89\*\*4\.94\*\*\*———By Buyer Type:High\-type18\.59\*\*\*0\.863\.24\*\*———Low\-type10\.06\*\*3\.69\*\*\*3\.88\*\*\*———
Notes:Between\-treatment test statistics:χ2\\chi^\{2\}for Agreement,\|t\|\|t\|for Rounds and Efficiency\.p†<0\.10\{\}^\{\\dagger\}p<0\.10, \*p<0\.05p<0\.05, \*\*p<0\.01p<0\.01, \*\*\*p<0\.001p<0\.001\. Efficiency measured as percentage of first\-best undiscounted surplus\.
Table 28:Model\-Level Verbal vs\. Structured ComparisonBuyer Share \(%\)Efficiency \(%\)RoundsModelFamilyVerbalStructuredVerbalStructuredVerbalStructuredGPT\-4o\-miniOpenAI39\.864\.1∗∗∗84\.674\.7∗∗3\.403\.52GPT\-5\-miniOpenAI41\.951\.3∗∗94\.092\.92\.433\.63∗∗∗GPT\-5\.2OpenAI37\.940\.598\.297\.0∗2\.682\.91∗Gemini\-2\.5\-FlashGoogle44\.341\.798\.797\.12\.452\.82∗Gemini\-3\-FlashGoogle51\.445\.9∗∗∗99\.198\.93\.113\.24Gemini\-3\-ProGoogle53\.950\.2∗99\.8100\.0∗∗∗3\.313\.31Qwen2\.5\-14BQwen52\.750\.089\.891\.42\.482\.19†Qwen3\-32BQwen66\.860\.296\.185\.7∗∗∗3\.263\.30Qwen3\-MaxQwen91\.491\.098\.797\.9∗3\.753\.45∗All models53\.554\.895\.492\.82\.983\.15
Notes:Significance stars on structured values:tt\-test comparing structured vs verbal within each model\.p†<0\.10\{\}^\{\\dagger\}p<0\.10, \*p<0\.05p<0\.05, \*\*p<0\.01p<0\.01, \*\*\*p<0\.001p<0\.001\.
Pooled buyer share changes little \(53\.5% verbal vs\. 54\.8% structured\), but model\-level distributional shifts diverge\. OpenAI’s mid\-tier and baseline models shift surplustowardbuyers when communication is removed: GPT\-5\-mini’s buyer share rises from 41\.9% to 51\.3% \(d=−0\.27d=\-0\.27,p=0\.003p=0\.003\), and GPT\-4o\-mini shifts even more sharply \(39\.8%→\\to64\.1%,d=−0\.33d=\-0\.33,p=0\.001p=0\.001\)\. These models use verbal messages in ways that moderate the buyer advantage: their sellers negotiate more effectively with language than without it\. Gemini’s mid\-tier and flagship models move in the opposite direction: Gemini\-3\-Flash buyer share drops from 51\.4% to 45\.9% \(d=0\.30d=0\.30,p=0\.001p=0\.001\), and Gemini\-3\-Pro from 53\.9% to 50\.2% \(d=0\.20d=0\.20,p=0\.033p=0\.033\); for these two models, verbal communication amplifies the buyer advantage rather than moderating it\. The remaining five models, GPT\-5\.2, Gemini\-2\.5\-Flash, Qwen2\.5\-14B, Qwen3\-32B, and Qwen3\-Max, show no detectable buyer\-share shift\. Thus R1 rules out a purely language\-based explanation for the main distributional profiles while showing that the verbal channel remains an execution aid and a model\-specific distributional lever\.
### 11\.3No\-Discounting Bargaining \(R2\)
This subsection reports the full no\-discounting bargaining extension \(R2\) summarized in Section[6\.2](https://arxiv.org/html/2608.07538#S6.SS2)of the main paper\. It gives the design, interpretive frame, and the complete set of empirical results behind the headline comparison \(Table[29](https://arxiv.org/html/2608.07538#S11.T29)\), including the cuts by first\-proposer role and buyer type \(Table[30](https://arxiv.org/html/2608.07538#S11.T30)\)\.
Table 29:Negotiation Performance: Symmetric Patience \(δ=\(0\.9,0\.9\)\\delta=\(0\.9,0\.9\)\) vs\. No DiscountingGroupSymmetricδ=\(0\.9,0\.9\)\\delta=\(0\.9,0\.9\)No DiscountingEff\.RoundsRound 1BuyerIrr\.Eff\.RoundsRound 1BuyerIrr\.acceptsharerateacceptsharerate\(%\)\(%\)\(%\)\(%\)\(%\)\(%\)\(%\)\(%\)By Tier:Baseline89\.12\.8310\.041\.522\.489\.83\.661\.146\.416\.5Mid\-tier97\.33\.4215\.050\.70\.695\.93\.9211\.153\.70\.0Flagship99\.23\.880\.661\.90\.099\.04\.450\.661\.80\.0By Provider:OpenAI90\.63\.225\.032\.415\.990\.13\.793\.340\.212\.4Google99\.63\.845\.651\.00\.699\.64\.910\.048\.60\.0Qwen95\.33\.0915\.070\.26\.195\.03\.339\.472\.73\.9By Model:OpenAI:GPT\-5\.298\.83\.600\.039\.30\.098\.44\.700\.038\.70\.0GPT\-5\-mini95\.82\.7215\.035\.90\.093\.33\.1310\.037\.50\.0GPT\-4o\-mini77\.43\.360\.020\.154\.078\.73\.500\.045\.242\.0Google:Gemini\-3\-Pro99\.94\.330\.056\.00\.0100\.04\.830\.054\.80\.0Gemini\-3\-Flash99\.84\.420\.051\.80\.099\.85\.280\.052\.50\.0Gemini\-2\.5\-Flash99\.12\.7716\.745\.21\.799\.04\.620\.038\.60\.0Qwen:Qwen3\-Max99\.03\.701\.790\.40\.098\.73\.821\.791\.90\.0Qwen3\-32B96\.23\.1430\.064\.41\.794\.63\.3523\.371\.10\.0Qwen2\.5\-14B90\.72\.4513\.355\.616\.791\.62\.833\.355\.111\.7Overall95\.23\.398\.551\.57\.494\.94\.024\.354\.15\.3
Between\-treatment tests \(Symmetric vs No Discounting\):\|t\|\|t\|Welch /χ2\\chi^\{2\}Eff\.RoundsRound\-1 acceptBuyer shareIrr\. rateOverall0\.296\.49\*\*\*8\.19\*\*1\.251\.95By Tier:Baseline0\.304\.48\*\*\*13\.55\*\*\*1\.021\.88Mid\-tier1\.212\.79\*\*1\.201\.051\.01Flagship0\.704\.72\*\*\*0\.000\.04—
Provider\-level buyer\-share tests \(Symmetric vs No Discounting\):\|t\|\|t\|Welch on raw shareOpenAI pooled1\.73†Google pooled1\.86†Qwen pooled0\.78Treatment×\\timesProvider interaction \(OLSFF\-test\)2\.44†Capability gradient: Cochran–Armitage linear\-trend test on irrationality across tiersSymmetricδ=\(0\.9,0\.9\)\\delta=\(0\.9,0\.9\)\|Z\|=7\.93∗∗∗\|Z\|=7\.93\*\*\*\(p=0\.000p=0\.000\)No discounting\|Z\|=6\.82∗∗∗\|Z\|=6\.82\*\*\*\(p=0\.000p=0\.000\)
Notes:N=60N=60per \(model, treatment\);N=540N=540per treatment\. Round\-1 accept: share agreeing in the first round\. Buyer share: buyer’s share of realized undiscounted surplus\. Irr\. rate: share of agreed deals where either party accepted a contract with negative undiscounted profit\. Efficiency and first\-round acceptance averaged over all experiments; rounds\-to\-agreement, buyer share, and irrationality rate averaged over agreed deals only\. All quantities computed on an undiscounted basis to compare treatments on common footing \(the no\-discounting treatment makes discounted and undiscounted profits identical by construction\)\. Test column reports\|t\|\|t\|\(Welch\) for Eff\., Rounds, and Buyer share;χ2\\chi^\{2\}\(Pearson\) for Round\-1 accept and Irr\. rate\. Cochran–Armitage trend test uses scores\{0,1,2\}\\\{0,1,2\\\}on Either\-irrational across tiers \(Baseline→\\toMid→\\toFlagship\), reported separately within each treatment\.p†<0\.10\{\}^\{\\dagger\}p<0\.10, \*p<0\.05p<0\.05, \*\*p<0\.01p<0\.01, \*\*\*p<0\.001p<0\.001\.
#### 11\.3\.1Design
We retain the main\-analysis factorial structure \(two buyer types, two proposer assignments, nine models, fifteen replications per cell\) but modify both buyer and seller prompts to remove all references toδ\\delta, patience, effective utility, and time pressure\. Agents are told that they negotiate over multiple rounds with a ten\-round cap, that failure to agree yields zero, and that profit is computed from the contract terms as in the main analysis\. No discount factor appears in the prompt or in the utility computation\. The resulting design yieldsN=540N=540negotiations\. Appendix[8](https://arxiv.org/html/2608.07538#S8.SSx1)reports the modified prompt text, with removed content shown in strikethrough\.
#### 11\.3\.2Interpretive Frame
We present R2 as an external\-validity test of the main findings rather than as a mechanism test\. Removing discounting from the prompt changes several features simultaneously: the utility formula, the patience concept, and the cue that time matters\. We do not attempt to separate their contributions\. The informative question is whether the qualitative patterns of the main analysis \(the capability\-patience driver separation, provider\-specific surplus profiles, capability\-dependent irrationality\) replicate under a prompting regime that omits the discounting framework\. Because no imposedδ\\deltastructure exists in R2, there is no Feng et al\. benchmark for surplus division; we compare against the symmetric high\-patience\(0\.9,0\.9\)\(0\.9,0\.9\)cell as the closest main\-analysis analogue\.
#### 11\.3\.3Empirical Results
Table[29](https://arxiv.org/html/2608.07538#S11.T29)compares the symmetric\-patience baseline with the no\-discounting treatment \(noδ\\deltain the prompt or payoff computation\) on five outcomes: undiscounted efficiency, rounds\-to\-agreement, first\-round acceptance, buyer share, and irrationality rate, all reported on a common undiscounted basis so the two treatments are compared on equal footing\. The three panels serve distinct purposes\. Panel A maps the descriptive landscape, reporting each outcome by tier, by provider, by model, and overall, for both treatments\.101010Patterns by first\-proposer role and buyer type are stable on agreement and efficiency, while rounds increase significantly in every reported cut; Table[30](https://arxiv.org/html/2608.07538#S11.T30)reports these comparisons\.Panel B tests whether the treatments differ in central tendency, overall and within each tier, using Welch\|t\|\|t\|tests for continuous outcomes andχ2\\chi^\{2\}tests for binary outcomes\. Panel C tests for heterogeneity: provider\-level Welch tests, a Treatment×\\timesProvider interactionFF\-test on raw buyer share, and a Cochran–Armitage linear\-trend test\(Cochran[1954](https://arxiv.org/html/2608.07538#bib.bib7), Armitage[1955](https://arxiv.org/html/2608.07538#bib.bib1)\)for irrationality across capability tiers within each treatment\.
The next three paragraphs each draw on one slice of the table: the high\-efficiency\-with\-delay pattern \(Panels A and B\), the cross\-provider distributional profile \(theBy Providerblock of Panel A together with the interaction in Panel C\), and the capability–irrationality gradient \(theIrr\. ratecolumn of Panel A together with the trend test in Panel C\)\.
##### Efficiency and delay\.
Allocative efficiency is statistically indistinguishable across the two treatments at the aggregate level \(95\.2%95\.2\\%vs\.94\.9%94\.9\\%;\|t\|=0\.29\|t\|=0\.29\) and within every capability tier\. Delay, by contrast, increases broadly: rounds\-to\-agreement rise from3\.393\.39to4\.024\.02in the pooled sample \(\|t\|=6\.49\|t\|=6\.49,p<0\.001p<0\.001\), from2\.832\.83to3\.663\.66among baseline models \(\|t\|=4\.48\|t\|=4\.48,p<0\.001p<0\.001\), from3\.423\.42to3\.923\.92among mid\-tier models \(\|t\|=2\.79\|t\|=2\.79,p<0\.01p<0\.01\), and from3\.883\.88to4\.454\.45among flagship models \(\|t\|=4\.72\|t\|=4\.72,p<0\.001p<0\.001\)\. Round\-1 acceptance also falls from8\.5%8\.5\\%to4\.3%4\.3\\%\(χ2=8\.19\\chi^\{2\}=8\.19,p<0\.01p<0\.01\)\. Removing the discount factor therefore acts as a tempo cue, slowing closure without altering allocative efficiency\.
##### Provider\-specific distributional profiles\.
TheBy Providerblock of Table[29](https://arxiv.org/html/2608.07538#S11.T29)reports realized buyer surplus shares pooled within each model family\. We report raw buyer share rather than deviation from a PBE benchmark because theFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)PBE buyer share is undefined whenδBδS=1\\delta\_\{B\}\\delta\_\{S\}=1\(the alternating\-offer game admits a continuum of subgame\-perfect divisions when delay is costless\), so aΔ\\Delta\-against\-PBE construct is not portable across the two treatments\. The provider\-level pattern from Section[5](https://arxiv.org/html/2608.07538#S5)reappears under R2: OpenAI remains seller\-favoring \(buyer share32\.4%→40\.2%32\.4\\%\\to 40\.2\\%\), Google tracks an even split \(51\.0%→48\.6%51\.0\\%\\to 48\.6\\%\), and Qwen remains strongly buyer\-favoring \(70\.2%→72\.7%70\.2\\%\\to 72\.7\\%\)\. The Treatment×\\timesProvider interaction is marginal \(F=2\.44F=2\.44,p=0\.087p=0\.087; Panel C\), but the rank ordering OpenAI≺\\precGoogle≺\\precQwen is preserved\. The cross\-family seller\-favoring/near\-neutral/buyer\-favoring profile is therefore a property of the underlying agents rather than an artifact of how patience was prompted\.
##### Capability\-dependent irrationality gradient\.
TheIrr\. ratecolumn of Table[29](https://arxiv.org/html/2608.07538#S11.T29)reports, by tier and treatment, the share of agreed deals in which either party accepted a contract yielding negative undiscounted profit\. The capability ordering from the main analysis survives the removal of discounting: irrationality declines monotonically from baseline \(22\.4%22\.4\\%symmetric,16\.5%16\.5\\%no\-discount\) through mid\-tier \(0\.6%0\.6\\%,0\.0%0\.0\\%\) to flagship \(0\.0%0\.0\\%,0\.0%0\.0\\%\)\. A Cochran–Armitage linear\-trend test confirms the gradient within each treatment \(symmetric:\|Z\|=7\.93\|Z\|=7\.93; no\-discount:\|Z\|=6\.82\|Z\|=6\.82; bothp<0\.001p<0\.001; Panel C\)\. Baseline irrationality is if anything lower under no discounting \(16\.5%16\.5\\%versus22\.4%22\.4\\%\), and the within\-tier between\-treatmentχ2\\chi^\{2\}tests in Panel B are nonsignificant \(baselineχ2=1\.88\\chi^\{2\}=1\.88; mid\-tier essentially unchanged,χ2=1\.01\\chi^\{2\}=1\.01\); flagship has no between\-treatment variation to test, since both treatments show0\.0%0\.0\\%irrationality\. Capability, not prompt structure, governs whether agents accept loss\-making contracts\. Removing discounting leaves the capability gradient intact, indicating that rationality at the contract level is an agent property\.
Table 30:Negotiation Performance by Other Dimensions: Symmetric Patience vs\. No DiscountingConditionSymmetric Patienceδ=\(0\.9,0\.9\)\\delta=\(0\.9,0\.9\)No DiscountingAgreementRoundsEfficiencyAgreementRoundsEfficiency\(%\)\(%\)\(%\)\(%\)By First Proposer:Buyer first97\.43\.3894\.697\.84\.1395\.0Seller first98\.53\.4095\.898\.53\.9194\.8By Buyer Type:High\-type98\.93\.3395\.398\.94\.0395\.2Low\-type97\.03\.4595\.197\.44\.0094\.6Between\-Treatment Tests \(Symmetric vs No Discounting\)By First Proposer:Buyer first0\.085\.29\*\*\*0\.31———Seller first0\.003\.85\*\*\*0\.81———By Buyer Type:High\-type0\.005\.40\*\*\*0\.04———Low\-type0\.073\.85\*\*\*0\.33———
Notes:N=60N=60per \(model, treatment\);N=540N=540per treatment\. Companion to Table[29](https://arxiv.org/html/2608.07538#S11.T29), which reports Overall and By Tier; this table presents only the By First Proposer and By Buyer Type cuts to avoid duplication\. Agreement reported over all experiments; Rounds over agreed deals only; Efficiency over all experiments \(failed deals contribute 0\)\. Efficiency is undiscounted realized surplus relative to first\-best\. Between\-treatment test statistics:χ2\\chi^\{2\}\(Pearson[1900](https://arxiv.org/html/2608.07538#bib.bib33)\)for Agreement,\|t\|\|t\|\(Welch\([1947](https://arxiv.org/html/2608.07538#bib.bib42)\)\) for Rounds and Efficiency\. Significance:p†<0\.10\{\}^\{\\dagger\}p<0\.10, \*p<0\.05p<0\.05, \*\*p<0\.01p<0\.01, \*\*\*p<0\.001p<0\.001\.
##### Deployment interpretation\.
R2 is a closer analogue to production\-style prompts that specify contractual terms and objectives without parameterizing agent utility with a time\-preference coefficient\. On the three structural margins that organize the main analysis, the qualitative patterns persist: allocative efficiency, the raw cross\-provider distributional ordering, and the capability\-irrationality gradient all reproduce\. The main detectable change is longer bargaining, which carries no allocative\-efficiency loss in the no\-discounting treatment because delay is not penalized\. The extension does not replace the main\-analysis benchmark against the PBE ofFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\), which remains necessary for interpreting deviations from normative predictions\. Instead, it shows that those deviations persist when explicit discount\-factor language is omitted, while that language mainly affects negotiation tempo\.
### 11\.4Reasoning Effort Ablation \(R4\)
This extension probes whether additional inference\-time computation materially changes bargaining outcomes within a fixed model\. If more reasoning per turn sharpened strategic execution, we would expect more aggressive opening offers, faster convergence, or tighter surplus splits\. We test this by running GPT\-5\.2 at low, medium, and high reasoning effort levels across all 16 conditions \(N=720N=720experiments\)\.
Reasoning effort affects negotiation tempo but not undiscounted allocative performance\. Mean rounds decline from 2\.95 at low effort to 2\.68 at medium effort and 2\.53 at high effort \(F=11\.22F=11\.22,p=1\.6×10−5p=1\.6\\times 10^\{\-5\}\)\. Discounted efficiency rises from 70\.7% to 76\.9% to 78\.5% \(F=24\.59F=24\.59,p=4\.7×10−11p=4\.7\\times 10^\{\-11\}\)\. For both outcomes, the low–medium and low–high differences survive Holm correction, whereas the medium–high difference does not\. Undiscounted efficiency \(F=1\.68F=1\.68,p=0\.188p=0\.188\) and buyer share \(F=2\.21F=2\.21,p=0\.110p=0\.110\) do not differ in the run\-level omnibus tests\. Deal rates are 100% at every effort level\. Table[31](https://arxiv.org/html/2608.07538#S11.T31)reports the full results\.
Table 31:Reasoning Effort Ablation: Summary Statistics \(GPT\-5\.2\)EffortNNDeal RateRoundsEfficiencyDiscEfficiencyBuyer ShareLow240100\.0%2\.9597\.2%70\.7%41\.4%Medium240100\.0%2\.6898\.2%76\.9%37\.9%High240100\.0%2\.5397\.7%78\.5%36\.5%Run\-level ANOVAFF11\.221\.6824\.592\.21Omnibuspp\-value<0\.001<0\.0010\.188<0\.001<0\.0010\.110
Notes:We use DiscEfficiency to denote discounted efficiency\. The omnibus tests are ordinary one\-way ANOVAs at the negotiation level\. Pairwise comparisons use Welch tests with Holm correction across the three effort\-level contrasts within each outcome\.
Across all effort levels, buyer surplus share stays within a 36\.5%–41\.4% band, and no buyer\-share pairwise comparison survives Holm correction\. Thus, within GPT\-5\.2 and over the effort levels tested here, additional inference\-time computation changes negotiation speed but not the division of surplus in the run\-level analysis\.
The ablation therefore rejects a blanket null effect of reasoning effort\. Greater effort primarily reduces negotiation delay and improves time\-adjusted performance, without significantly improving undiscounted allocative efficiency\. This pattern is not uniquely diagnostic of any single internal mechanism\.
### 11\.5Retail Price Robustness \(R5\)
The main analysis establishes its findings at a retail price of 60 \(margin=30=30, 50% gross margin\)\. A natural concern is whether these patterns \(high efficiency with delayed agreement, provider\-specific surplus division, and capability threshold\) are specific to this particular surplus level or represent general properties of LLM bargaining\. If surplus division patterns stem from training\-encoded behavioral tendencies rather than strategic calculation, they should be relatively insensitive to the absolute size of the surplus: a rational agent would negotiate harder when more money is on the table, but a training\-driven bias should produce similar percentage shares regardless of magnitude\. We test this by tripling the margin to 90 \(retail price 120, 75% gross margin\), creating a larger surplus to divide\. Only OpenAI models \(GPT\-5\.2, GPT\-5\-mini, GPT\-4o\-mini\) are included\. We compare their RP120 results \(N=720N=720\) against their RP60 baseline \(N=720N=720\), with 15 replications per condition in both\.
##### The high\-efficiency\-with\-delay pattern generalizes across price levels\.
Undiscounted efficiency remains high at both price levels, dipping modestly from 92\.3% at RP60 to 89\.5% at RP120 \(p=0\.014p=0\.014\); discounted efficiency is statistically equivalent \(66\.7% vs\. 64\.2%,p=0\.070p=0\.070\)\. The agents’ ability to find efficient deals is largely insensitive to the size of the surplus available\.
##### Buyer advantage responds differently by capability tier\.
Buyer surplus share rises from 39\.9% at RP60 to 52\.7% at RP120\. As Table[32](https://arxiv.org/html/2608.07538#S11.T32)and Figure[14](https://arxiv.org/html/2608.07538#S11.F14)show, this shift is driven mainly by GPT\-4o\-mini, whose buyer share jumps from 39\.8% to 67\.1% \(p<0\.001p<0\.001\); GPT\-5\-mini shows a smaller but significant rise \(41\.9%→\\to50\.4%,p=0\.006p=0\.006\), while GPT\-5\.2 is essentially flat \(37\.9%→\\to40\.7%,p=0\.254p=0\.254\)\.
This pattern is descriptive rather than mechanically identified\. The flagship model maintains a comparatively stable buyer share across the larger\-surplus environment, while GPT\-5\-mini and especially GPT\-4o\-mini become more buyer\-favoring at RP120\. This tier gradient is consistent with greater environment sensitivity below the capability threshold, but the design does not identify the internal mechanism generating the response\.
Notably, GPT\-4o\-mini’s higher buyer share translates into significantly higherabsolutebuyer profit at RP120\. Among rational deals \(excluding agreements in which either party accepts negative profit\), GPT\-4o\-mini earns $3,349 per deal compared to $2,059 for GPT\-5\.2 \(t=7\.62t=7\.62,p<0\.001p<0\.001\), with a buyer share gap of 23\.9 percentage points \(t=9\.14t=9\.14,p<0\.001p<0\.001\)\. Even accounting for GPT\-4o\-mini’s 11\.8% irrational deal rate and 15\.4% deal failure rate, its expected buyer profit \($2,782\) still significantly exceeds GPT\-5\.2’s \($2,059;t=4\.23t=4\.23,p<0\.001p<0\.001\)\. This counterintuitive result arises because GPT\-5\.2 negotiates “fairer” deals that allocate more surplus to the seller \($2,823 vs\. $1,699 for GPT\-4o\-mini\), while GPT\-4o\-mini’s aggressive buyer behavior, though less reliable, captures more value when it succeeds\. The implication for AI\-mediated procurement is that deploying a more capable model as buyer does not necessarily maximize buyer value; capability promotes balanced outcomes rather than partisan ones\.
Table 32:Retail Price Comparison: Summary Statistics \(OpenAI Models\)Retail PriceModelNNDeal RateRoundsEfficiencyBuyer Share60GPT\-4o\-mini24090\.8%3\.4084\.6%39\.8%60GPT\-5\-mini240100%2\.4394\.0%41\.9%60GPT\-5\.2240100%2\.6898\.2%37\.9%120GPT\-4o\-mini24084\.6%3\.8279\.2%67\.1%120GPT\-5\-mini240100%2\.1891\.7%50\.4%120GPT\-5\.2240100%2\.7597\.6%40\.7%
Notes:N=240N=240per \(model, retail price\)\. Deal Rate is over all experiments; Rounds is over agreed deals only; Efficiency is undiscounted realized surplus relative to first\-best \(failed deals contribute 0\); Buyer Share is the buyer’s share of realized undiscounted surplus among agreed deals\.

Figure 14:Retail Price Comparison Across OpenAI ModelsNotes:RP60 \(blue\) vs\. RP120 \(red\) for buyer surplus share, undiscounted efficiency, and rounds to agreement\.
##### Patience effects and negotiation speed at higher margins\.
Negotiations conclude at a similar pace at RP120 \(2\.87 vs\. 2\.82 rounds,t=0\.63t=0\.63,p=0\.53p=0\.53\)\. The qualitative pattern of patience effects is preserved: under buyer\-patient conditions \(δB=0\.9\\delta\_\{B\}=0\.9,δS=0\.4\\delta\_\{S\}=0\.4\), buyer shares reach 75\.5% at RP120 \(vs\. 62\.9% at RP60\); under seller\-patient conditions \(δB=0\.4\\delta\_\{B\}=0\.4,δS=0\.9\\delta\_\{S\}=0\.9\), buyer shares fall to 37\.4% at RP120 \(vs\. 23\.4% at RP60\)\. The direction of patience effects is consistent across price levels, though the magnitude of buyer\-share variation is somewhat compressed at RP120\.
### 11\.6Sensitivity to Prior Probability of High\-Type Buyer \(R6\)
The main analysis fixes the seller’s prior belief about the buyer being high\-type atβ=0\.5\\beta=0\.5\. This extension varies the prior toβ=0\.3\\beta=0\.3\(seller believes buyer is more likely low\-type\) andβ=0\.7\\beta=0\.7\(seller believes buyer is more likely high\-type\), restricting the patience configuration to \(δB=0\.4\\delta\_\{B\}=0\.4,δS=0\.9\\delta\_\{S\}=0\.9\), where the buyer is impatient and the seller is patient\. This condition is chosen because the theoretical equilibrium structure\(Feng et al\.[2015](https://arxiv.org/html/2608.07538#bib.bib14)\)is most prior\-sensitive under low buyer patience: the thresholds that determine signaling, pooling, and screening regimes shift with the prior, and quantity distortion occurs in this regime\. By contrast, under high buyer patience the buyer truth\-tells with first\-best quantities regardless ofβ\\beta, making the prior less consequential\. All nine main\-analysis models are included, with 15 replications per condition \(N=540N=540per prior level,N=1,620N=1\{,\}620total\)\.
##### Efficiency and deal rates are robust to prior beliefs\.
Deal rates remain near\-universal across prior levels \(98\.1%, 99\.3%, and 97\.8% atβ=0\.3\\beta=0\.3,0\.50\.5, and0\.70\.7\), with no meaningful difference\. Undiscounted efficiency likewise stays high \(92\.7%, 95\.0%, and 93\.6%\), without a systematic trend in the prior\. Table[33](https://arxiv.org/html/2608.07538#S11.T33)reports the full results\.
Table 33:Prior Sensitivity: Summary Statistics \(δB=0\.4\\delta\_\{B\}=0\.4,δS=0\.9\\delta\_\{S\}=0\.9, All 9 Models Pooled\)Prior \(β\\beta\)NNDeal RateRoundsEfficiencyBuyer ShareWholesale Price0\.354098\.1%2\.6192\.744\.2%$44\.010\.554099\.3%2\.6395\.040\.2%$45\.190\.754097\.8%2\.8693\.637\.1%$44\.99
Notes:N=540N=540per prior level, pooled across all 9 models\. Deal Rate, Rounds, Efficiency, and Buyer Share as defined in Table[32](https://arxiv.org/html/2608.07538#S11.T32)\. Wholesale Price is the mean agreed unit price \(payment/quantity\) among agreed deals\.
##### Negotiation speed varies with prior\.
The ANOVA for rounds is significant \(F=4\.78F=4\.78,p=0\.009p=0\.009\)\. Negotiations atβ=0\.7\\beta=0\.7take the longest \(2\.86 rounds\), followed byβ=0\.5\\beta=0\.5\(2\.63 rounds\) andβ=0\.3\\beta=0\.3\(2\.61 rounds\)\. Higher priors, where the seller believes the buyer is likely high\-type, lead to longer negotiations as sellers bargain more aggressively\.
##### Surplus division shifts with seller beliefs\.
Buyer surplus share trends downward as the seller’s prior about high\-type increases: 44\.2% atβ=0\.3\\beta=0\.3, 40\.2% atβ=0\.5\\beta=0\.5, and 37\.1% atβ=0\.7\\beta=0\.7\. The ANOVA is significant \(F=3\.21F=3\.21,p=0\.040p=0\.040\), and the directional pattern is consistent with rational Bayesian updating: a seller who believes the buyer is likely high\-type should demand more, extracting a larger share\.
##### Model families respond differently to priors\.
The prior’s effect varies across model families \. Gemini models show a monotonic decline in buyer share fromβ=0\.3\\beta=0\.3toβ=0\.7\\beta=0\.7\(31\.1%→\\to28\.5%→\\to19\.2%\), consistent with directionally rational behavior: higher priors prompt harder seller bargaining\. OpenAI models show the same monotonic pattern \(28\.1%→\\to23\.4%→\\to20\.1%\), declining steadily as the prior rises\. Qwen models, by contrast, show no meaningful response to the prior\. Buyer share remains near 70% at all three prior levels \(72\.8%, 68\.3%, 71\.0%\), reflecting the strong buyer bias documented in the main analysis\. The Qwen family’s dominant buyer\-favoring tendency overrides the informational content of the prior\.
### 11\.7Parameter Size Ablation \(R7\)
The main analysis compares models across providers, confounding architecture, training data, and parameter count\. This extension isolates the effect of model size by comparing three models from the same family: Qwen3\-14B \(14 billion parameters\), Qwen3\-32B \(32 billion\), and Qwen3\-Max \(flagship, undisclosed but substantially larger\)\. All three share the Qwen3 architecture, allowing a cleaner size comparison\. The extension covers the 16\-condition factorial \(N=720N=720across the three sizes; 240 attempts per model\)\.
##### Deal rates remain near\-universal across model size\.
All three sizes reach agreement in essentially all negotiations \(14B 100\.0%, 32B 99\.6%, Max 100\.0%\), so deal\-making itself does not scale with size in this family; the reliability difference instead appears in the rate of economically irrational agreements \(discussed below\)\. Table[34](https://arxiv.org/html/2608.07538#S11.T34)reports the full results\.
Table 34:Parameter Size Ablation: Summary Statistics \(Qwen3 Family; Irr\. = Irrational Deal Rate\)SizeNNDeal RateRoundsEfficiencyDiscEfficiencyBuyer ShareIrr\.14B240100\.0%4\.2894\.1%51\.6%63\.1%5\.4%32B24099\.6%3\.2696\.1%62\.6%66\.8%1\.7%Max240100\.0%3\.7598\.7%50\.8%91\.4%0\.0%
Notes:We use DiscEfficiency to denote discounted efficiency\.
##### Negotiation speed varies modestly with model size\.
The 14B model is the slowest to converge, averaging 4\.28 rounds, compared with 3\.75 for Max and 3\.26 for 32B\. The ANOVA for rounds across the three sizes is significant \(F=18\.0F=18\.0,p<0\.001p<0\.001\), with the 14B\-vs\-Max pairwise difference att=3\.32t=3\.32,p<0\.001p<0\.001; 32B is also significantly faster than Max \(t=3\.15t=3\.15,p=0\.002p=0\.002\), so negotiation speed does not improve monotonically with size in this family\. Figure[15](https://arxiv.org/html/2608.07538#S11.F15)shows the pattern across buyer share, efficiency, and rounds\.
##### Efficiency and the cost of delay\.
Undiscounted efficiency rises modestly with size \(94\.1% at 14B, 96\.1% at 32B, 98\.7% at Max;p<0\.001p<0\.001for the 14B\-vs\-Max contrast\)\. Discounted efficiency, which penalizes the extra rounds, does not fall monotonically with size: 32B attains the highest discounted efficiency \(62\.6%\), with 14B \(51\.6%\) and Max \(50\.8%\) close together, mirroring the round counts since 32B is the fastest to converge\. The smallest model’s longer negotiations are not severe enough to devastate the value it eventually creates\.

Figure 15:Parameter Size Effects \(Qwen3 Family\)Notes:Panels show buyer surplus share, undiscounted efficiency, and rounds to agreement\. The 14B model takes more rounds than 32B or Max \(p<0\.001p<0\.001\)\. The Max model achieves the highest efficiency and buyer share\.
##### Irrational deals decline with size\.
The rate of irrational deals declines with size: 5\.4% at 14B, 1\.7% at 32B, and 0\.0% at Max\. This pattern suggests a within\-family reliability gradient in the Qwen size comparison: value\-destroying agreements become less frequent as model size rises within this architecture family\. Because the ladder contains three Qwen points and Max’s parameter count is undisclosed, we interpret it as within\-family evidence rather than a general scaling law\.
## 12Bayesian Benchmark Implementation Details
### 12\.1Problem Setup and Notation
Let the buyer’s private type bei∈\{H,L\}i\\in\\\{H,L\\\}with priorP\(i=H\)=βP\(i=H\)=\\beta\. A contract is a pair\(q,T\)\(q,T\), whereqqis order quantity andTTis the transfer from buyer to seller\. The buyer’s demand isDi∼𝒩\(μi,σi2\)D\_\{i\}\\sim\\mathcal\{N\}\(\\mu\_\{i\},\\sigma\_\{i\}^\{2\}\)\. In the baseline calibration,
\(μH,μL,σH,σL\)=\(80,40,10,10\),r=60,c=30,β=0\.5\.\(\\mu\_\{H\},\\mu\_\{L\},\\sigma\_\{H\},\\sigma\_\{L\}\)=\(80,40,10,10\),\\qquad r=60,\\qquad c=30,\\qquad\\beta=0\.5\.
If agreement occurs in roundτ\\tau, the buyer’s and seller’s discounted utilities are
UB,i\(τ\)\(q,T\)\\displaystyle U\_\{B,i\}^\{\(\\tau\)\}\(q,T\)=δBτ−1\(Ri\(q\)−T\),\\displaystyle=\\delta\_\{B\}^\{\\tau\-1\}\\bigl\(R\_\{i\}\(q\)\-T\\bigr\),\(1\)US\(τ\)\(q,T\)\\displaystyle U\_\{S\}^\{\(\\tau\)\}\(q,T\)=δSτ−1\(T−cq\),\\displaystyle=\\delta\_\{S\}^\{\\tau\-1\}\\bigl\(T\-cq\\bigr\),\(2\)where
Ri\(q\)=r𝔼\[min\(Di,q\)\]\.R\_\{i\}\(q\)=r\\,\\mathbb\{E\}\[\\min\(D\_\{i\},q\)\]\.\(3\)
For normal demand,Ri\(q\)R\_\{i\}\(q\)admits the closed form
Ri\(q\)=r\[μiΦ\(zi\)−σiϕ\(zi\)\+q\(1−Φ\(zi\)\)\],zi=q−μiσi,R\_\{i\}\(q\)=r\\left\[\\mu\_\{i\}\\Phi\(z\_\{i\}\)\-\\sigma\_\{i\}\\phi\(z\_\{i\}\)\+q\\bigl\(1\-\\Phi\(z\_\{i\}\)\\bigr\)\\right\],\\qquad z\_\{i\}=\\frac\{q\-\\mu\_\{i\}\}\{\\sigma\_\{i\}\},\(4\)withΦ\(⋅\)\\Phi\(\\cdot\)andϕ\(⋅\)\\phi\(\\cdot\)denoting the standard normal cdf and pdf\.
The type\-iifirst\-best quantity solves the newsvendor condition
Fi\(q^i\)=r−cr,F\_\{i\}\(\\hat\{q\}\_\{i\}\)=\\frac\{r\-c\}\{r\},\(5\)and the associated maximum total surplus is
π^i=Ri\(q^i\)−cq^i\.\\hat\{\\pi\}\_\{i\}=R\_\{i\}\(\\hat\{q\}\_\{i\}\)\-c\\hat\{q\}\_\{i\}\.\(6\)
In our calibration,\(r−c\)/r=0\.5\(r\-c\)/r=0\.5, soq^H=80\\hat\{q\}\_\{H\}=80andq^L=40\\hat\{q\}\_\{L\}=40\.
### 12\.2Complete\-Information Building Blocks
Under complete information,Feng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)reduce the game to Rubinstein bargaining over the type\-specific surplusπ^i\\hat\{\\pi\}\_\{i\}\.
If the seller proposes first, the equilibrium contract is\(q^i,T^iS\)\(\\hat\{q\}\_\{i\},\\hat\{T\}\_\{i\}^\{S\}\)with
T^iS=cq^i\+1−δB1−δBδSπ^i\.\\hat\{T\}\_\{i\}^\{S\}=c\\hat\{q\}\_\{i\}\+\\frac\{1\-\\delta\_\{B\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{i\}\.\(7\)Hence the seller and buyer shares of surplus are
sSS\-first=1−δB1−δBδS,sBS\-first=δB\(1−δS\)1−δBδS\.s\_\{S\}^\{S\\text\{\-first\}\}=\\frac\{1\-\\delta\_\{B\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\},\\qquad s\_\{B\}^\{S\\text\{\-first\}\}=\\frac\{\\delta\_\{B\}\(1\-\\delta\_\{S\}\)\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\.\(8\)
If the buyer proposes first, the equilibrium contract is\(q^i,T^iB\)\(\\hat\{q\}\_\{i\},\\hat\{T\}\_\{i\}^\{B\}\)with
T^iB=cq^i\+δS\(1−δB\)1−δBδSπ^i\.\\hat\{T\}\_\{i\}^\{B\}=c\\hat\{q\}\_\{i\}\+\\frac\{\\delta\_\{S\}\(1\-\\delta\_\{B\}\)\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{i\}\.\(9\)The corresponding shares are
sSB\-first=δS\(1−δB\)1−δBδS,sBB\-first=1−δS1−δBδS\.s\_\{S\}^\{B\\text\{\-first\}\}=\\frac\{\\delta\_\{S\}\(1\-\\delta\_\{B\}\)\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\},\\qquad s\_\{B\}^\{B\\text\{\-first\}\}=\\frac\{1\-\\delta\_\{S\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\.\(10\)
These complete\-information objects are building blocks for the asymmetric\-information benchmark\. They are not, by themselves, the full incomplete\-information prediction\.
### 12\.3Patience Regions
Following Proposition 2 ofFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\), define the high\-type information advantage at the low\-type first\-best quantity,
Δ=RH\(q^L\)−RL\(q^L\),\\Delta=R\_\{H\}\(\\hat\{q\}\_\{L\}\)\-R\_\{L\}\(\\hat\{q\}\_\{L\}\),\(11\)and
k=Δπ^H−π^L\.k=\\frac\{\\Delta\}\{\\hat\{\\pi\}\_\{H\}\-\\hat\{\\pi\}\_\{L\}\}\.\(12\)
The buyer’s patience region is determined by
δ¯B\(δS\)=k1\+δS\(k−1\),δ¯B\(δS\)=k−1\+δSkδS\.\\bar\{\\delta\}\_\{B\}\(\\delta\_\{S\}\)=\\frac\{k\}\{1\+\\delta\_\{S\}\(k\-1\)\},\\qquad\\underline\{\\delta\}\_\{B\}\(\\delta\_\{S\}\)=\\frac\{k\-1\+\\delta\_\{S\}\}\{k\\delta\_\{S\}\}\.\(13\)
The buyer is in the HIGH region whenδB≥δ¯B\(δS\)\\delta\_\{B\}\\geq\\bar\{\\delta\}\_\{B\}\(\\delta\_\{S\}\), in the MEDIUM region whenδ¯B\(δS\)≤δB<δ¯B\(δS\)\\underline\{\\delta\}\_\{B\}\(\\delta\_\{S\}\)\\leq\\delta\_\{B\}<\\bar\{\\delta\}\_\{B\}\(\\delta\_\{S\}\), and in the LOW region otherwise\.
For our calibration,k≈0\.1995k\\approx 0\.1995\. WhenδS=0\.9\\delta\_\{S\}=0\.9, this gives
δ¯B\(0\.9\)≈0\.554,δ¯B\(0\.9\)≈0\.714,\\underline\{\\delta\}\_\{B\}\(0\.9\)\\approx 0\.554,\\qquad\\bar\{\\delta\}\_\{B\}\(0\.9\)\\approx 0\.714,so the three buyer\-patience values used in the paper map to LOW \(0\.40\.4\), MEDIUM \(0\.70\.7\), and HIGH \(0\.90\.9\) exactly as intended\.
### 12\.4Asymmetric\-Information Contract Objects
The seller\-initiated PBE inFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)is piecewise, with the realized branch depending on the patience region and on prior cutoffs within that region\. The benchmark implementation uses the following equilibrium objects\.
##### Pooling / intermediate contract\(qI,TI\)\(q\_\{I\},T\_\{I\}\)\.
In the seller\-first pooling region, the seller offers a single contract acceptable to both types:
TI=RH\(qI\)−δB\(1−δS\)1−δBδSπ^H=RL\(qI\)−δB\(1−δS\)1−δBδSπ^L\.T\_\{I\}=R\_\{H\}\(q\_\{I\}\)\-\\frac\{\\delta\_\{B\}\(1\-\\delta\_\{S\}\)\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{H\}=R\_\{L\}\(q\_\{I\}\)\-\\frac\{\\delta\_\{B\}\(1\-\\delta\_\{S\}\)\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{L\}\.\(14\)
##### Seller\-screening contract in the MEDIUM region\.
In the\(R,A\)\(R,A\)\-Scr branch of Proposition 4, the seller uses the same quantity objectqIq\_\{I\}but perturbs the transfer by a smallε\>0\\varepsilon\>0so that only the low type accepts in round 1:
TI=RH\(qI\)−δB\(1−δS\)1−δBδSπ^H\+ε=RL\(qI\)−δB\(1−δS\)1−δBδSπ^L\.T\_\{I\}=R\_\{H\}\(q\_\{I\}\)\-\\frac\{\\delta\_\{B\}\(1\-\\delta\_\{S\}\)\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{H\}\+\\varepsilon=R\_\{L\}\(q\_\{I\}\)\-\\frac\{\\delta\_\{B\}\(1\-\\delta\_\{S\}\)\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{L\}\.\(15\)
##### Low\-type signaling contract\(qD,TD\)\(q\_\{D\},T\_\{D\}\)\.
In the LOW buyer\-patience region, the low type separates by distorting quantity downward\. The signaling contract\(qD,TD\)\(q\_\{D\},T\_\{D\}\)withqD<q^Lq\_\{D\}<\\hat\{q\}\_\{L\}solves
RH\(qD\)−TD\\displaystyle R\_\{H\}\(q\_\{D\}\)\-T\_\{D\}=1−δS1−δBδSπ^H,\\displaystyle=\\frac\{1\-\\delta\_\{S\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{H\},\(16\)TD−cqD\\displaystyle T\_\{D\}\-cq\_\{D\}=δS\(1−δB\)1−δBδSπ^L\.\\displaystyle=\\frac\{\\delta\_\{S\}\(1\-\\delta\_\{B\}\)\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{L\}\.\(17\)Equation \([16](https://arxiv.org/html/2608.07538#S12.E16)\) makes the high type indifferent to mimicking, and Equation \([17](https://arxiv.org/html/2608.07538#S12.E17)\) gives the seller the low\-type continuation value\.
##### Mixed low\-patience counteroffer\(qM,TM\)\(q\_\{M\},T\_\{M\}\)\.
When the LOW\-region equilibrium falls into the\(A,R\)\(A,R\)\-Mix branch of Proposition 5, the low type’s round\-2 counteroffer\(qM,TM\)\(q\_\{M\},T\_\{M\}\)satisfies
RH\(qM\)−TM\\displaystyle R\_\{H\}\(q\_\{M\}\)\-T\_\{M\}=min\{RH\(qS\)−TSδB,RH\(q^L\)−T^LB\},\\displaystyle=\\min\\left\\\{\\frac\{R\_\{H\}\(q\_\{S\}\)\-T\_\{S\}\}\{\\delta\_\{B\}\},\\;R\_\{H\}\(\\hat\{q\}\_\{L\}\)\-\\hat\{T\}\_\{L\}^\{B\}\\right\\\},\(18\)TM−cqM\\displaystyle T\_\{M\}\-cq\_\{M\}=δS\(1−δB\)1−δBδSπ^L,\\displaystyle=\\frac\{\\delta\_\{S\}\(1\-\\delta\_\{B\}\)\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{L\},\(19\)withqD≤qM≤q^Lq\_\{D\}\\leq q\_\{M\}\\leq\\hat\{q\}\_\{L\}\. The corresponding seller first\-round transfer cutoffT¯HS\\bar\{T\}\_\{H\}^\{S\}is pinned down by the indifference condition in Proposition 5\(a\)\.
##### Mixed seller offer\(qX,TX\)\(q\_\{X\},T\_\{X\}\)\.
In the\(R,A\)\(R,A\)\-Mix branch of Proposition 5, the seller’s round\-1 offer\(qX,TX\)\(q\_\{X\},T\_\{X\}\)satisfiesqX<qDq\_\{X\}<q\_\{D\}and solves the fixed\-point indifference conditions in Proposition 5\(b\)\. These equations are implemented numerically in the validation agent\. They are not used as simple Rubinstein\-share shortcuts\.
### 12\.5How the Benchmark Is Constructed in This Paper
The incomplete\-information PBE inFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)is derived for seller\-initiated bargaining\. Our experimental design also includes buyer\-first conditions, so the implemented benchmark separates two cases: seller\-first cells are benchmarked directly to Feng’s branch\-specific PBE, whereas buyer\-first cells are benchmarked to a separately derived buyer\-first separating equilibrium built from Feng’s complete\-information contracts and low\-type signaling conditions\.
##### Buyer\-first conditions\.
BecauseFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)do not derive the buyer\-initiated game formally, we verify here the separating buyer\-first equilibrium used in the benchmark implementation\. Nature draws buyer typei∈\{H,L\}i\\in\\\{H,L\\\}with priorβ\\beta; the buyer observesiiand proposes\(q,T\)\(q,T\)in round 1; the seller updates beliefs and accepts or rejects; after rejection, the game continues with the seller proposing in round 2\. For a degenerate posterior on typeii, the seller’s round\-1 continuation value from rejecting and entering the round\-2 complete\-information seller\-proposer subgame is
VS,icont=δS1−δB1−δBδSπ^i\.V\_\{S,i\}^\{\\mathrm\{cont\}\}=\\delta\_\{S\}\\frac\{1\-\\delta\_\{B\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{i\}\.Hence any accepted round\-1 type\-iicontract must give the seller at leastVS,icontV\_\{S,i\}^\{\\mathrm\{cont\}\}\.
The buyer\-first benchmark uses the following candidate separating offers:
\(qH⋆,TH⋆\)=\(q^H,T^HB\),\(q\_\{H\}^\{\\star\},T\_\{H\}^\{\\star\}\)=\(\\hat\{q\}\_\{H\},\\hat\{T\}\_\{H\}^\{B\}\),and
\(qL⋆,TL⋆\)=\{\(q^L,T^LB\),δB≥δ¯B\(δS\),\(qD,TD\),δB<δ¯B\(δS\),\(q\_\{L\}^\{\\star\},T\_\{L\}^\{\\star\}\)=\\begin\{cases\}\(\\hat\{q\}\_\{L\},\\hat\{T\}\_\{L\}^\{B\}\),&\\delta\_\{B\}\\geq\\underline\{\\delta\}\_\{B\}\(\\delta\_\{S\}\),\\\\ \(q\_\{D\},T\_\{D\}\),&\\delta\_\{B\}<\\underline\{\\delta\}\_\{B\}\(\\delta\_\{S\}\),\\end\{cases\}where\(q^i,T^iB\)\(\\hat\{q\}\_\{i\},\\hat\{T\}\_\{i\}^\{B\}\)is the buyer\-proposer complete\-information contract from Equation \([9](https://arxiv.org/html/2608.07538#S12.E9)\),\(qD,TD\)\(q\_\{D\},T\_\{D\}\)is the low\-type signaling contract from Equations \([16](https://arxiv.org/html/2608.07538#S12.E16)\)–\([17](https://arxiv.org/html/2608.07538#S12.E17)\), andδ¯B\(δS\)\\underline\{\\delta\}\_\{B\}\(\\delta\_\{S\}\)is the lower patience cutoff in Equation \([13](https://arxiv.org/html/2608.07538#S12.E13)\)\.
Seller participation binds on path because the buyer optimally offers the lowest transfer consistent with acceptance\. Fori∈\{H,L\}i\\in\\\{H,L\\\}, substituting Equation \([9](https://arxiv.org/html/2608.07538#S12.E9)\) yields
T^iB−cq^i=δS1−δB1−δBδSπ^i=VS,icont,\\hat\{T\}\_\{i\}^\{B\}\-c\\hat\{q\}\_\{i\}=\\delta\_\{S\}\\frac\{1\-\\delta\_\{B\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{i\}=V\_\{S,i\}^\{\\mathrm\{cont\}\},so the seller is indifferent and accepts by convention\. In the low\-patience signaling regime, Equation \([17](https://arxiv.org/html/2608.07538#S12.E17)\) gives the same equality for\(qD,TD\)\(q\_\{D\},T\_\{D\}\)withi=Li=L\.
Type\-HHincentive compatibility determines the threshold\. If typeLLoffers\(q^L,T^LB\)\(\\hat\{q\}\_\{L\},\\hat\{T\}\_\{L\}^\{B\}\), typeHH’s payoff from mimicking is
RH\(q^L\)−T^LB=π^L\+Δ−δS1−δB1−δBδSπ^L,R\_\{H\}\(\\hat\{q\}\_\{L\}\)\-\\hat\{T\}\_\{L\}^\{B\}=\\hat\{\\pi\}\_\{L\}\+\\Delta\-\\delta\_\{S\}\\frac\{1\-\\delta\_\{B\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{L\},whereΔ=RH\(q^L\)−RL\(q^L\)\\Delta=R\_\{H\}\(\\hat\{q\}\_\{L\}\)\-R\_\{L\}\(\\hat\{q\}\_\{L\}\)\. Comparing this with typeHH’s equilibrium payoff
RH\(q^H\)−T^HB=1−δS1−δBδSπ^HR\_\{H\}\(\\hat\{q\}\_\{H\}\)\-\\hat\{T\}\_\{H\}^\{B\}=\\frac\{1\-\\delta\_\{S\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\hat\{\\pi\}\_\{H\}shows thatHHweakly prefers its own contract iff
1−δS1−δBδS\(π^H−π^L\)≥Δ,\\frac\{1\-\\delta\_\{S\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\bigl\(\\hat\{\\pi\}\_\{H\}\-\\hat\{\\pi\}\_\{L\}\\bigr\)\\geq\\Delta,which is equivalent to
1−δS1−δBδS≥k⟺δB≥δ¯B\(δS\)\.\\frac\{1\-\\delta\_\{S\}\}\{1\-\\delta\_\{B\}\\delta\_\{S\}\}\\geq k\\qquad\\Longleftrightarrow\\qquad\\delta\_\{B\}\\geq\\underline\{\\delta\}\_\{B\}\(\\delta\_\{S\}\)\.Thus, above the cutoff, the no\-distortion low\-type offer\(q^L,T^LB\)\(\\hat\{q\}\_\{L\},\\hat\{T\}\_\{L\}^\{B\}\)is incentive compatible; below the cutoff, typeHHwould mimic it\. In that low\-patience region, the signaling contract\(qD,TD\)\(q\_\{D\},T\_\{D\}\)restores separation because Equation \([16](https://arxiv.org/html/2608.07538#S12.E16)\) makes typeHHexactly indifferent between its own contract and mimicking the low type, while Equation \([17](https://arxiv.org/html/2608.07538#S12.E17)\) preserves seller participation\.
Type\-LLdoes not mimic typeHH\. The key monotonicity is thatRH\(q\)−RL\(q\)R\_\{H\}\(q\)\-R\_\{L\}\(q\)is increasing inqq, so higher quantities are relatively more attractive to the high type\. In the no\-distortion regime, once typeHHweakly prefers\(q^H,T^HB\)\(\\hat\{q\}\_\{H\},\\hat\{T\}\_\{H\}^\{B\}\)to\(q^L,T^LB\)\(\\hat\{q\}\_\{L\},\\hat\{T\}\_\{L\}^\{B\}\), typeLLstrictly prefers\(q^L,T^LB\)\(\\hat\{q\}\_\{L\},\\hat\{T\}\_\{L\}^\{B\}\)to\(q^H,T^HB\)\(\\hat\{q\}\_\{H\},\\hat\{T\}\_\{H\}^\{B\}\)\. In the low\-patience regime, the same single\-crossing logic implies that if typeHHis exactly indifferent between\(q^H,T^HB\)\(\\hat\{q\}\_\{H\},\\hat\{T\}\_\{H\}^\{B\}\)and\(qD,TD\)\(q\_\{D\},T\_\{D\}\), then typeLLstrictly prefers\(qD,TD\)\(q\_\{D\},T\_\{D\}\)to the high\-type contract\. Therefore typeLLnever gains from mimicking typeHH\.
To complete the equilibrium, we use off\-path beliefs that attribute a deviation to the type for whom it could be profitable under acceptance\. If a deviation is attributed to typeHH, the seller applies the posterior\-HHparticipation threshold, and no accepted deviation improves on\(q^H,T^HB\)\(\\hat\{q\}\_\{H\},\\hat\{T\}\_\{H\}^\{B\}\)because that contract already solves typeHH’s round\-1 maximization subject to seller participation under posteriorHH\. If a deviation is attributed to typeLL, the seller applies the posterior\-LLparticipation threshold, and no accepted deviation improves on the relevant low\-type benchmark contract:\(q^L,T^LB\)\(\\hat\{q\}\_\{L\},\\hat\{T\}\_\{L\}^\{B\}\)in the no\-distortion region, or\(qD,TD\)\(q\_\{D\},T\_\{D\}\)in the signaling region, where the high\-type no\-mimicking constraint also binds\. Pooling outcomes are ruled out by the same logic: typeHHcan deviate to an offer that is strictly preferred under posteriorHHand unattractive to typeLL, so the seller’s off\-path beliefs do not sustain a pooling allocation\. These beliefs therefore support the separating allocation above\. For the benchmark implementation, this existence result is sufficient; we do not require a stronger uniqueness claim\.
Under this buyer\-first separating equilibrium, both types settle in round 1\. In the main design withδS=0\.9\\delta\_\{S\}=0\.9, the thresholdδ¯B\(0\.9\)≈0\.554\\underline\{\\delta\}\_\{B\}\(0\.9\)\\approx 0\.554places the HIGH \(δB=0\.9\\delta\_\{B\}=0\.9\) and MEDIUM \(δB=0\.7\\delta\_\{B\}=0\.7\) conditions in the no\-distortion regime and the LOW \(δB=0\.4\\delta\_\{B\}=0\.4\) condition in the signaling regime\. In the seller\-impatient condition\(δB,δS\)=\(0\.9,0\.4\)\(\\delta\_\{B\},\\delta\_\{S\}\)=\(0\.9,0\.4\), the threshold is negative, so the no\-distortion regime applies\. These are the buyer\-first benchmark outcomes implemented in the validation agent and summarized in Table[35](https://arxiv.org/html/2608.07538#S12.T35); Figure[16](https://arxiv.org/html/2608.07538#S12.F16)visualizes the corresponding truth\-telling regions under the baseline calibration\.
Relative to Feng’s seller\-first PBE, only the lower thresholdδ¯B\(δS\)\\underline\{\\delta\}\_\{B\}\(\\delta\_\{S\}\)matters here\. The upper thresholdδ¯B\(δS\)\\bar\{\\delta\}\_\{B\}\(\\delta\_\{S\}\)separates pooling from screening in the seller\-first game but has no analogue in the buyer\-first game because the informed party proposes, eliminating the uninformed\-proposer pooling temptation\. Prior cutoffs onβ\\betalikewise play no role: separation is achieved by the buyer’s own offer, so the seller’s posterior is degenerate on path regardless of the prior\.
##### Seller\-first conditions\.
For seller\-initiated cells, we benchmark directly to Feng’s PBE\. We first classify the patience region using Proposition 2 and then solve the relevant branch of Propositions 3–5 at\(δB,δS,β\)\(\\delta\_\{B\},\\delta\_\{S\},\\beta\)\.
Within each patience region, Feng defines prior cutoffs
βh,β¯h,βm,β¯m,βℓ,β¯ℓ\\beta\_\{h\},\\ \\bar\{\\beta\}\_\{h\},\\ \\beta\_\{m\},\\ \\bar\{\\beta\}\_\{m\},\\ \\beta\_\{\\ell\},\\ \\bar\{\\beta\}\_\{\\ell\}that separate pooling, screening, signaling, and mixed branches\. We compute these cutoffs numerically and then solve the corresponding branch\-specific contract equations\. Criteria 1–3 and Lemmas 1–2 determine how rejected seller offers are interpreted and which buyer counteroffers are belief\-revealing, pooling, or mixed\.
In the main 16\-condition design,β=0\.5\\beta=0\.5in every cell\. Under this calibration, the seller\-first benchmark reported in the paper has the following realized timing pattern:
- •type\-HHbuyers accept the seller’s round\-1 offer, and
- •type\-LLbuyers reject in round 1 and make the equilibrium round\-2 counteroffer, which the seller accepts\.
In HIGH and MEDIUM patience, the type\-LLround\-2 counteroffer is\(q^L,T^LB\)\(\\hat\{q\}\_\{L\},\\hat\{T\}\_\{L\}^\{B\}\)\. In LOW patience, the type\-LLround\-2 counteroffer is the low\-type equilibrium separating contract obtained from Proposition 5, that is, the signaling contract\(qD,TD\)\(q\_\{D\},T\_\{D\}\)or the corresponding mixed\-case object when the numerically selected branch requires it\.
This yields exactly the benchmark timing reported in the main text:
Buyer\-first benchmark rounds=1\.00,\\displaystyle=1\.00,\(20\)Seller\-first benchmark rounds=1\.50,\\displaystyle=1\.50,\(21\)Overall benchmark rounds=1\.25\.\\displaystyle=1\.25\.\(22\)
Table 35:Bayesian Benchmark Mapping for the Main 16\-Condition Design \(β=0\.5\\beta=0\.5\)First proposerBuyer patienceType\-HHpathType\-LLpathBuyerHIGH /MEDIUM\(q^H,T^HB\)\(\\hat\{q\}\_\{H\},\\hat\{T\}\_\{H\}^\{B\}\),accept in round 1\(q^L,T^LB\)\(\\hat\{q\}\_\{L\},\\hat\{T\}\_\{L\}^\{B\}\),accept in round 1BuyerLOW\(q^H,T^HB\)\(\\hat\{q\}\_\{H\},\\hat\{T\}\_\{H\}^\{B\}\),accept in round 1\(qD,TD\)\(q\_\{D\},T\_\{D\}\),accept in round 1SellerHIGH /MEDIUMseller offer acceptedin round 1reject in round 1, then\(q^L,T^LB\)\(\\hat\{q\}\_\{L\},\\hat\{T\}\_\{L\}^\{B\}\)in round 2SellerLOWseller offer acceptedin round 1reject in round 1, then LOW\-regionequilibrium counteroffer in round 2
Notes:The LOW\-region seller\-first counteroffer is\(qD,TD\)\(q\_\{D\},T\_\{D\}\)in the pure signaling branch and the corresponding mixed\-case counteroffer when Proposition 5 selects a mixed branch numerically\.
### 12\.6Validation and Interpretation
For each experimental cell, the benchmark implementation performs four steps:
1. \(1\)computeq^i\\hat\{q\}\_\{i\}andπ^i\\hat\{\\pi\}\_\{i\}fori∈\{H,L\}i\\in\\\{H,L\\\};
2. \(2\)classify the patience region using Equation \([13](https://arxiv.org/html/2608.07538#S12.E13)\);
3. \(3\)if the seller proposes first, solve the branch\-specific equilibrium object\(s\) required by Propositions 3–5; if the buyer proposes first, construct the type\-contingent buyer proposal from the buyer\-initiated complete\-information contract stated after Proposition 1, Proposition 2, and Lemma 2;
4. \(4\)execute the implied accept/reject path and record agreement round, contract terms, profits, and surplus shares\.
This validation benchmark reproduces the analytical properties used in the paper: \(i\) immediate agreement in all buyer\-first cells, \(ii\) agreement in at most two rounds in all seller\-first cells, \(iii\) truthful first\-best quantities whenever the equilibrium calls for full revelation, and \(iv\) branch\-specific quantity distortion whenever separation requires signaling\.
A final clarification is important\. The Rubinstein shares in Equations \([8](https://arxiv.org/html/2608.07538#S12.E8)\)–\([10](https://arxiv.org/html/2608.07538#S12.E10)\) are complete\-information building blocks\. They are not the full asymmetric\-information surplus benchmark\. Whenever the paper refers to the “Bayesian benchmark,” it means the branch\-specific incomplete\-information implementation described above, not Rubinstein shares taken in isolation\.

Figure 16:Buyer Truth\-Telling Regions Under the Baseline CalibrationNotes:The figure reproduces Proposition 2 ofFeng et al\. \([2015](https://arxiv.org/html/2608.07538#bib.bib14)\)\. The calibration isr=60r=60,c=30c=30,β=0\.5\\beta=0\.5,DH∼𝒩\(80,102\)D\_\{H\}\\sim\\mathcal\{N\}\(80,10^\{2\}\), andDL∼𝒩\(40,102\)D\_\{L\}\\sim\\mathcal\{N\}\(40,10^\{2\}\)\. Under this calibration,q^H=80\\hat\{q\}\_\{H\}=80,q^L=40\\hat\{q\}\_\{L\}=40, and the buyer\-patience values used in the main design map cleanly into the LOW, MEDIUM, and HIGH regions\.Similar Articles
Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations
This paper introduces a framework using reinforcement learning from verifiable rewards to train large language models for strategic bargaining in multi-buyer markets, addressing private information and surplus extraction in concurrent negotiations.
Counterparty Modeling is Not Strategy: The Limits of LLM Negotiators
Study shows LLM agents can model counterparty preferences in negotiation but fail to turn that knowledge into strategic bargaining to improve outcomes, limiting their effectiveness in multi-turn negotiations.
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
This paper formalizes deliberative collaboration for LLM agents under partial observability, introduces a scalable benchmark across multiple domains, and systematically evaluates representative LLMs, finding that complex tasks remain challenging while deliberation can enable error correction.
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
This paper introduces MoralSim to evaluate how LLM agents behave in morally charged social dilemmas where ethical actions conflict with profit incentives, finding that no model remains consistently moral and cooperation rates vary widely.
LLM agents diverge between public and off-the-record channels under social pressure, without any hidden goal in the prompt
This paper shows that LLM agents diverge between public and off-the-record channels under social pressure, without explicit hidden goals. Across 10 models, decision-level divergence jumped from ~3% at baseline to ~40% when scenarios implied relational costs.