MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

arXiv cs.AI Papers

Summary

MerchantBench is a new benchmark for evaluating LLM agents' long-term coherence in e-commerce operations, using a 365-day order-level simulation with 98,843 real product records and 26 tools. Results show the best LLM achieves only 27.3% of human participants' final net assets, highlighting a substantial capability gap.

arXiv:2607.28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.
Original Article
View Cached Full Text

Cached at: 08/03/26, 07:30 AM

# MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Source: [https://arxiv.org/html/2607.28956](https://arxiv.org/html/2607.28956)
Qiming Shi2, Yulong Tao1, Linbo Jin111footnotemark:1, Zhaolu Kang4, Yibo Dou4, Jiawen Zhu2, Tianjun Pan5, Shaokang Fu1, Chengyu Wang1, Siyue Li1, Yaping Cheng1, Di Weng3, Chengfu Huo1

###### Abstract

Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria\. Real\-world deployments often require Long\-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence\. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects\. Seller\-side e\-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash\-Flow Management, and Mixed\-Latency Feedback Adaptation\. We introduce MerchantBench, a 365\-day order\-level simulation grounded in 98,843 real e\-commerce product records and equipped with 26 tools for agent interaction\. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions\. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days\. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27\.3% of the mean final net assets achieved by human participants\.

## Introduction

Large language models have become a foundation for autonomous agents that plan, call tools, interact with external systems, and make decisions over multiple turns\. General agent benchmarks now measure tool use, web interaction, application control, and state\-changing workflows\(Liuet al\.[2024](https://arxiv.org/html/2607.28956#bib.bib4); Zhouet al\.[2024](https://arxiv.org/html/2607.28956#bib.bib5); Yaoet al\.[2025](https://arxiv.org/html/2607.28956#bib.bib1); Trivediet al\.[2024](https://arxiv.org/html/2607.28956#bib.bib2)\)\. Yet many real\-world deployments are not bounded tasks with immediate and unambiguous completion criteria\. They require agents to operate over extended horizons in environments whose state persists, where earlier actions constrain later options and relevant consequences may emerge only after many intervening decisions\. Evidence from recent benchmarks indicates that even state\-of\-the\-art agents often fail to maintain coherent performance over long horizons\(Luoet al\.[2025](https://arxiv.org/html/2607.28956#bib.bib33); Wanget al\.[2025c](https://arxiv.org/html/2607.28956#bib.bib34); Yuanet al\.[2026](https://arxiv.org/html/2607.28956#bib.bib38)\)\. These results motivate extending agent evaluation beyond isolated task completion to examine sustained objective pursuit and decision consistency in realistic long\-horizon environments\.

![Refer to caption](https://arxiv.org/html/2607.28956v1/x1.png)Figure 1:Order\-level dynamics couple immediate liquidity pressure with delayed feedback\. New orders commit available cash before settlement, whereas abnormal outcomes surface later and affect store rating\.Seller\-side e\-commerce provides a suitable setting for evaluating Long\-Term Coherence\. Unlike bounded tasks with explicit completion criteria, online store operation requires continued intervention throughout the operating horizon\. The agent must repeatedly select products, control listings and prices, manage limited cash, and revise earlier decisions as market conditions, supplier states, and order outcomes evolve\. Seller\-side e\-commerce therefore tests whether an agent can sustain and revise a merchant policy over time, rather than complete an isolated action\.

Vending\-Bench and RetailBench have taken important steps toward realistic long\-horizon business evaluation\(Backlund and Petersson[2025](https://arxiv.org/html/2607.28956#bib.bib14); Zhanget al\.[2026](https://arxiv.org/html/2607.28956#bib.bib16)\)\. Vending\-Bench evaluates vending\-machine operation, while RetailBench models supermarket management over a fixed catalog of 96 grocery products\. Seller\-side e\-commerce, however, introduces two challenges that are not represented together in these environments\.

First, e\-commerce feedback is generated through individual order lifecycles\. A listing or pricing decision may create new orders and commit available cash immediately, whereas fulfillment failures and after\-sales outcomes become observable only later\. Figure[1](https://arxiv.org/html/2607.28956#Sx1.F1)visualizes this temporal asymmetry between immediate order\-driven cash commitments and delayed abnormal outcomes\. As delayed evidence accumulates, the agent must associate later outcomes with earlier decisions and determine whether its current merchant policy should be maintained or revised\.

Second, a long episode alone does not meaningfully test Long\-Term Coherence if the available products and their demand remain static\. A large, data\-grounded Product Catalog with full\-year demand trajectories creates a changing opportunity set in which promising products emerge and existing choices lose value over time\. The agent must therefore continually identify new opportunities and revise its portfolio using market signals and realized order outcomes\.

To address this gap, we introduce MerchantBench, which evaluates Long\-Term Coherence through persistent seller\-side e\-commerce operation\. MerchantBench formulates e\-commerce operation as a partially observable decision\-making problem over 365 simulated days\. The environment grounds a Product Catalog in 98,843 real e\-commerce product records and converts demand into individual orders that progress through fulfillment and after\-sales stages\. Through merchant\-visible tools, the agent performs Product Sourcing, controls listings and prices, manages cash flow, and monitors supplier and order states\. Upstream Supplier Events and Downstream Order Outcomes become observable at different times, requiring the agent to use later evidence to maintain or revise earlier decisions\. Across the operating horizon, the simulator updates demand, supplier states, order lifecycles, cash flow, penalties, and store reputation\.

Our contributions are as follows:

- •To our knowledge, we introduce the first benchmark for evaluating Long\-Term Coherence through persistent seller\-side e\-commerce operation, structured around four interdependent decision components\.
- •We develop an order\-level simulation environment grounded in 98,843 real e\-commerce product records, with partial observability, cash constraints, Upstream Supplier Events, and delayed Downstream Order Outcomes\.
- •We conduct 48 runs of 365 simulated days across eight LLMs and two agent frameworks and analyze business outcomes and decision traces to characterize their Long\-Term Coherence\.

## Related Work

#### Agent Evaluation in Commerce\.

Existing benchmarks cover shopping and storefront interaction\(Yaoet al\.[2022](https://arxiv.org/html/2607.28956#bib.bib3); Wanget al\.[2026a](https://arxiv.org/html/2607.28956#bib.bib7); Zhanget al\.[2025](https://arxiv.org/html/2607.28956#bib.bib8); Wanget al\.[2026b](https://arxiv.org/html/2607.28956#bib.bib9); Savadikaret al\.[2026](https://arxiv.org/html/2607.28956#bib.bib10); Duet al\.[2026](https://arxiv.org/html/2607.28956#bib.bib11)\), customer support\(Wanget al\.[2025a](https://arxiv.org/html/2607.28956#bib.bib6)\), and merchant workflows and negotiation\(Zhaoet al\.[2026](https://arxiv.org/html/2607.28956#bib.bib12); Wanget al\.[2025b](https://arxiv.org/html/2607.28956#bib.bib13)\)\. Market\-Bench and Magentic Marketplace examine market competition and transactions among economic agents\(Zhenget al\.[2026](https://arxiv.org/html/2607.28956#bib.bib21); Bansalet al\.[2025](https://arxiv.org/html/2607.28956#bib.bib22)\)\. Across these categories, evaluation centers on tasks, dialogues, transactions, or competitive episodes rather than continuous operation of the same online store\.

#### Long\-Horizon Agent Evaluation\.

Recent long\-horizon benchmarks assess sustained reasoning and action across workplace workflows, open\-ended exploration, virtual\-world planning, web navigation, computer use, inventory control, order fulfillment, and interactive economies\(Luoet al\.[2025](https://arxiv.org/html/2607.28956#bib.bib33); Wanget al\.[2025c](https://arxiv.org/html/2607.28956#bib.bib34); Anokhinet al\.[2025](https://arxiv.org/html/2607.28956#bib.bib35); Janget al\.[2026](https://arxiv.org/html/2607.28956#bib.bib36); Dinget al\.[2026](https://arxiv.org/html/2607.28956#bib.bib37); Yuanet al\.[2026](https://arxiv.org/html/2607.28956#bib.bib38); Baeket al\.[2026](https://arxiv.org/html/2607.28956#bib.bib19); Zhuet al\.[2023](https://arxiv.org/html/2607.28956#bib.bib20); Huet al\.[2026](https://arxiv.org/html/2607.28956#bib.bib18); Sugiuraet al\.[2026](https://arxiv.org/html/2607.28956#bib.bib17)\)\. Vending\-Bench and Vending\-Bench 2 emphasize sustained business operation, while RetailBench evaluates evidence acquisition, action conversion, and temporal follow\-up in supermarket management\(Backlund and Petersson[2025](https://arxiv.org/html/2607.28956#bib.bib14); Andon Labs[2025](https://arxiv.org/html/2607.28956#bib.bib15); Zhanget al\.[2026](https://arxiv.org/html/2607.28956#bib.bib16)\)\. MerchantBench extends this line by evaluating how agents adapt to delayed order\-level feedback and nonstationary demand derived from real e\-commerce data\.

## MerchantBench

![Refer to caption](https://arxiv.org/html/2607.28956v1/x2.png)Figure 2:Overview of MerchantBench\. The merchant agent coordinates four decision components through shorter supplier feedback and delayed order outcomes over 365 days\. The environment combines an upstream supplier simulation, a merchant store, and a downstream order level simulation to evaluate Long\-Term Coherence and business outcomes\.### Task Formulation

We formulate store operation as a finite horizon partially observable Markov decision process \(POMDP\)\(Kaelblinget al\.[1998](https://arxiv.org/html/2607.28956#bib.bib26)\)

ℳ=⟨𝒮,𝒜,P,𝒪,Z,R,μ0,Hc⟩\.\\mathcal\{M\}=\\langle\\mathcal\{S\},\\mathcal\{A\},P,\\mathcal\{O\},Z,R,\\mu\_\{0\},H\_\{c\}\\rangle\.\(1\)The simulator advances hourly over a 365 day control horizon, givingHc=8,760H\_\{c\}=8\{,\}760steps indexed byt∈\{0,…,Hc−1\}t\\in\\\{0,\\ldots,H\_\{c\}\-1\\\}\. Demand, supplier states, and order lifecycles evolve at every step, while the agent receives a decision window once every 12 steps\. The latent statest∈𝒮s\_\{t\}\\in\\mathcal\{S\}contains the simulation clock, product demand profiles, supplier conditions, store listings and finances, active orders, and pending events\. The initial state followsμ0\\mu\_\{0\}and includes the cash balance, security deposit, and listing capacity\. The cash balance funds procurement and realized losses, unpaid fines draw from the security deposit, and operation terminates when the deposit is exhausted\. At activation steps,at∈𝒜a\_\{t\}\\in\\mathcal\{A\}denotes the sequence of merchant tool invocations within the decision window, while other steps use a fixed null action\. The transition kernelP​\(st\+1∣st,at\)P\(s\_\{t\+1\}\\mid s\_\{t\},a\_\{t\}\)combines tool induced store changes with autonomous demand, supplier, and order evolution, with listing and pricing changes affecting demand from the next step\. The observation kernelZ​\(ot\+1∣st\+1,at\)Z\(o\_\{t\+1\}\\mid s\_\{t\+1\},a\_\{t\}\)exposes only merchant visible information, while demand profiles, risk parameters, pending outcomes, and future event times remain latent until their effects become observable\. The policy therefore conditions on observation and tool result history rather than the full state\. Att=Hct=H\_\{c\}, new demand and agent activations stop while active orders continue until terminal settlement at a terminal stepT≥HcT\\geq H\_\{c\}\. Intermediate rewards are zero, and the objective is expected terminal net assets

J​\(π\)=𝔼π​\[R​\(sT\)\],R​\(sT\)=BT\+DT\+IT\+QT\.\\begin\{array\}\[\]\{rcl\}J\(\\pi\)&=&\\mathbb\{E\}\_\{\\pi\}\[R\(s\_\{T\}\)\],\\\\ R\(s\_\{T\}\)&=&B\_\{T\}\+D\_\{T\}\+I\_\{T\}\+Q\_\{T\}\.\\end\{array\}\(2\)whereBTB\_\{T\},DTD\_\{T\},ITI\_\{T\}, andQTQ\_\{T\}denote the terminal cash balance, security deposit, funds in transit, and receivables\. Thus,R​\(sT\)R\(s\_\{T\}\)is the realized net asset value of one run andJ​\(π\)J\(\\pi\)is its expectation across stochastic trajectories\.

### Real\-World Data Grounding

MerchantBench is grounded in real\-world e\-commerce data from 1688, the largest integrated domestic wholesale marketplace in China\(Alibaba Group[n\.d\.](https://arxiv.org/html/2607.28956#bib.bib23)\)\. The data contain product and supplier attributes, 365\-day product\-level demand histories, and platform quality and fulfillment signals\. The data cover 365 days from June 1, 2025 through May 31, 2026\. Alongside the product data, MerchantBench incorporates 365 daily market reports from 1688 as date\-aligned signals for product sourcing\. We select 10 first\-level product categories spanning apparel, household and office goods, appliances, pet and gardening products, toys, bags, and sports and outdoor products\. After excluding records with missing identifiers, unmapped categories, nonpositive prices, or incomplete demand histories, the dataset contains 98,843 products from 36,576 suppliers\. Through catalog and supplier tools, the agent accesses only public product and supplier attributes\. Figure[3](https://arxiv.org/html/2607.28956#Sx3.F3)captures aggregate demand peaks at 618 and during the first wave and final day of 11\.11, together with a trough during the Spring Festival\. The lower panels show the distributions of effective Upstream Supplier Event and Downstream Order Outcome probabilities\.

### Upstream Supplier Simulation

The supply pool maintains time varying procurement prices, available inventory, availability, and supplier shipment times, while inventory replenishes over time\. At each simulator step, the simulator samples the three Upstream Supplier Events, namely Price Change, Product Delisting, and Shipment Delay, using product level probabilities calibrated from real platform fulfillment signals\. These events alter the upstream procurement price, suspend product procurement, and extend supplier dispatch time, respectively\. Inventory stockouts arise endogenously when incoming orders deplete stock faster than it replenishes\. To prevent persistent environment drift, each triggered abnormality receives a sampled recovery time at which the affected supplier attributes return to their base states\. The agent can observe realized changes to price, availability, quantity, and shipment time through catalog and supplier queries, but it cannot access the underlying abnormality flag, trigger probability, or recovery schedule\.

![Refer to caption](https://arxiv.org/html/2607.28956v1/x3.png)Figure 3:Real\-world demand patterns and calibrated risk profiles across 98,843 products\. Whiskers span the 10th to 90th percentiles, and boxes show interquartile ranges\.
### Downstream Order Simulation

#### Order Level Simulation\.

The downstream simulation converts product level daily demand traces into individual orders\. For productiilisted by merchantmmat timett, the hourly arrival intensity is

λm,i,t=Di,d​\(t\)​wc​\(i\),h​\(t\)​rm,t​ℓm,i,t​\(pm,i,t/piref\)−ϵi\.\\lambda\_\{m,i,t\}=D\_\{i,d\(t\)\}w\_\{c\(i\),h\(t\)\}r\_\{m,t\}\\ell\_\{m,i,t\}\\left\(p\_\{m,i,t\}/p\_\{i\}^\{\\mathrm\{ref\}\}\\right\)^\{\-\\epsilon\_\{i\}\}\.\(3\)The indicesd​\(t\)=⌊t/24⌋d\(t\)=\\lfloor t/24\\rfloorandh​\(t\)=tmod24h\(t\)=t\\bmod 24denote the data day and hour of day associated with steptt\. The quantityDi,d​\(t\)D\_\{i,d\(t\)\}is the linked daily demand from the real\-world data, andwc​\(i\),h​\(t\)w\_\{c\(i\),h\(t\)\}distributes category demand across hours\. The price term uses product elasticityϵi\\epsilon\_\{i\}, whilerm,tr\_\{m,t\}andℓm,i,t\\ell\_\{m,i,t\}capture store rating and listing exposure\. Realized order outcomes update the published store rating after each completed day, and its discrete star level determinesrm,tr\_\{m,t\}\. The factorℓm,i,t\\ell\_\{m,i,t\}increases during a listing’s cold start and then decays with age\. The appendix provides precise definitions of both factors in the section on demand and rating dynamics\. The environment samplesNm,i,t∼Poisson​\(λm,i,t\)N\_\{m,i,t\}\\sim\\mathrm\{Poisson\}\(\\lambda\_\{m,i,t\}\)and instantiates each arrival as an order candidate\. At creation, each candidate receives one latent customer outcome from normal fulfillment, Cancellation, Returnless Refund, Return and Refund, or Bad Review according to its product specific risk profile\. Stockout and Late Shipment instead arise from procurement and fulfillment dynamics\. The selected outcome and its realization time remain hidden until the corresponding lifecycle transition occurs\.

Table 1:Business performance, store reliability, and long\-horizon activity after 365 simulated days\. Values are means over three runs\. Final net assets and GMV are reported in thousands of RMB, total fines in RMB, and rate metrics in percent\. SWR denotes Sustained Window Rate\. The best result within each framework is shown in bold\.
#### Order Lifecycle\.

MerchantBench follows the single item drop shipping model supported by 1688, in which merchants hold no inventory in advance\. Each Order Placed triggers immediate procurement at the current supplier price, and successful procurement deducts the cash balance, decreases supplier inventory, and moves the order to Procured\. Following supplier and logistics delays, the order advances through Shipped and Delivered, at which point the sale price becomes a receivable\. A normal order becomes Settled after a sampled delay and credits the receivable to the cash balance\.

MerchantBench models Normal Fulfillment together with six abnormal Downstream Order Outcomes, namely Cancellation, Stockout, Late Shipment, Returnless Refund, Return and Refund, and Bad Review\. Cancellation restores the procurement cost, while Stockout prevents procurement\. Late Shipment marks a missed dispatch deadline, after which the order continues through fulfillment and settlement\. Both refund outcomes remove the receivable, but only Return and Refund restores the procurement cost, whereas Bad Review preserves the sales revenue\. Stockout, Late Shipment, Return and Refund, and Bad Review incur platform fines\. All abnormal outcomes except Cancellation also contribute adverse evidence to the store rating through outcome\-specific experience scores and evidence weights\.

### Agent Interface

MerchantBench exposes a shared observation protocol and 26 merchant tools, allowing different agent frameworks to interact with identical environment dynamics and observable state\. At each decision window, the agent receives a summary of simulated time, store status, and recent supplier and order changes, then acts until it ends the window or reaches the time limit\. The appendix lists the complete MerchantBench tool interface\.

#### Product Sourcing\.

Daily market reports, catalog search, product details, and public supplier profiles support product selection, while demand, product risk rates, future order outcomes, and future supplier events remain latent\.

#### Listing and Pricing Control\.

Listing, delisting, repricing, and performance views allow agents to construct the store portfolio and revise its products and prices\.

#### Cash\-Flow Management\.

Finance and store views expose the cash balance, security deposit, committed funds, expected settlements, fines, and closure conditions, which agents manage through subsequent sourcing, listing, delisting, and pricing decisions\.

#### Mixed\-Latency Feedback Adaptation\.

Supplier and order tools reveal Upstream Supplier Events and Downstream Order Outcomes as they unfold, allowing agents to revise the other three decision components in response to mixed\-latency feedback\.

## Experiments

### Experimental Setup

#### Agent Configurations and Baselines\.

We evaluate eight LLMs under ReAct\(Yaoet al\.[2023](https://arxiv.org/html/2607.28956#bib.bib24)\)and Hermes\(Nous Research[2026](https://arxiv.org/html/2607.28956#bib.bib25)\)with three runs for each pairing of an LLM and a framework\. The evaluated models are GPT\-5\.6 Sol\(OpenAI[2026](https://arxiv.org/html/2607.28956#bib.bib27)\), Claude Opus 4\.8\(Anthropic[2026](https://arxiv.org/html/2607.28956#bib.bib28)\), Qwen3\.7\-Max and Qwen3\.7\-Plus\(Qwen Team[2026](https://arxiv.org/html/2607.28956#bib.bib29)\), GLM\-5\.2\(Z\.ai[2026](https://arxiv.org/html/2607.28956#bib.bib30)\), DeepSeek\-V4\-Pro and DeepSeek\-V4\-Flash\(DeepSeek\-AI[2026](https://arxiv.org/html/2607.28956#bib.bib31)\), and Kimi K2\.6\(Moonshot AI[2026](https://arxiv.org/html/2607.28956#bib.bib32)\)\. ReAct pairs each model with a minimal controller over the 26 MerchantBench tools to assess core planning, reasoning, and tool use\. Hermes uses its default configuration, combining the 26 MerchantBench tools with built in capabilities for code execution, planning, memory, and skill management\. Both frameworks compress long interaction histories, with each evaluated model also serving as its own summarizer\. Each run starts with RMB 2,000 in cash, a RMB 1,000 security deposit, and capacity for 50 active listings\. We additionally compare with a Rule\-based baseline and three Human participants without prior e\-commerce operating experience\. The Rule\-based baseline performs daily checks, removes inactive or supplier\-affected products, and fills open listing slots using the daily market report\. The appendix provides the detailed experimental settings\.

#### Evaluation Metrics\.

We evaluate Business Performance using Final Net Assets, GMV, Net Profit Margin, and Orders; Store Reliability using Total Fines, Average Store Rating, and Order Anomaly Rate; and Long\-Horizon Activity using Average Active Listings, Sustained Window Rate \(SWR\), and Total Tool Calls\. SWR is the minimum share of scheduled decision windows containing at least one environment tool call across all rolling 30 day periods\.

### Main Results

#### Overall Performance\.

Table[1](https://arxiv.org/html/2607.28956#Sx3.T1)reports the final performance of all evaluated configurations and baselines\. GPT\-5\.6 Sol records the highest final net assets under ReAct, whereas Qwen3\.7\-Max ranks first under Hermes\. Qwen3\.7\-Max with Hermes achieves the highest final net assets among all 16 configurations\. When results are aggregated by model across the two frameworks, GPT\-5\.6 Sol has the highest average final net assets\.

#### Performance Variability\.

Figure[4](https://arxiv.org/html/2607.28956#Sx4.F4)shows substantial differences in stability across configurations\. GPT\-5\.6 Sol under ReAct and Claude Opus 4\.8 under Hermes have the lowest coefficients of variation within their respective frameworks at 3\.3% and 10\.0%\. Despite achieving the highest mean final net assets, Qwen3\.7\-Max under Hermes is considerably less stable, with a coefficient of variation of 55\.1%\.

![Refer to caption](https://arxiv.org/html/2607.28956v1/x4.png)Figure 4:Final net asset distributions across three repeated runs for each configuration\.
#### Framework Analysis\.

Averaged across the eight models, Hermes produces 53\.3% higher final net assets, 71\.5% higher GMV, and 71\.2% more orders than ReAct\. Mean final net assets are higher under Hermes for seven of the eight models, with gains ranging from 11\.5% for Claude Opus 4\.8 to 187\.8% for Qwen3\.7\-Max\. Kimi K2\.6 is the sole exception, with mean final net assets 4\.1% lower under Hermes than under ReAct, showing that framework benefits depend strongly on the underlying model\.

### Order\-Level Risk Propagation

Figure[1](https://arxiv.org/html/2607.28956#Sx1.F1)illustrates the mixed\-latency evidence generated by MerchantBench’s order\-level simulation\. Agents observe prior ratings of upstream catalog products during sourcing and receive prompt demand signals from realized sales, whereas product quality is revealed only through delayed order outcomes that may require product\-level risk response or store\-level rating adaptation\.

#### Product\-Level Risk Response\.

Models differed in whether they converted delayed order outcomes into product\-specific interventions\. In representative runs, GPT\-5\.6 Sol and Kimi K2\.6 attributed adverse outcomes to the responsible listing and replaced it, whereas Qwen3\.7\-Plus retained a risky product for further observation and DeepSeek\-V4\-Pro did not revise affected listings after refunds\. Human participants described a more complete response for popular but risky products, in which they delisted the product and searched similar keywords for a replacement that could preserve the underlying demand\. These behaviors distinguish simple anomaly detection from the full chain of product attribution, risk removal, and demand\-preserving replacement\.

#### Store\-Level Rating Adaptation\.

Order\-level anomalies also reduce the store rating, allowing product\-specific failures to affect demand across the portfolio\. In a GPT\-5\.6 Sol trajectory, the agent lowered prices on proven products without adverse outcomes to increase normal settlements and recover the rating threshold\. Qwen3\.7\-Max applied the same mechanism more aggressively by repricing 40 listings after an early bad review reduced the store to three stars\.

### Long\-Term Coherence Analysis

The aggregate results reveal a substantial gap between the evaluated LLM agents and Human operators\. Our trace evidence suggests that two forms of Long\-Term Coherence failure developing over extended operation may contribute to this performance gap\. Some agents progressively reduce store intervention and fail to follow up on prior decisions or delayed outcomes, indicating a loss of Operational Coherence\. Others remain active but drift from the terminal net assets objective or fail to revise ineffective policies as evidence accumulates, indicating a loss of Strategic Coherence\.

![Refer to caption](https://arxiv.org/html/2607.28956v1/x5.png)Figure 5:Monthly Operational Coherence profiles for the Human baseline and all eight Hermes models\. Panels report effective window rate, environment tool calls, and monthly net profit\.#### Operational Coherence\.

Figure[5](https://arxiv.org/html/2607.28956#Sx4.F5)reveals substantial differences in whether merchant activity persists across the operating horizon\. Table[1](https://arxiv.org/html/2607.28956#Sx3.T1)shows that Human operators retain an SWR of 100%, while LLM configurations range from 10\.6% to 99\.4% under ReAct and from 17\.8% to 66\.1% under Hermes\. Qwen3\.7\-Max provides the clearest contrast, with its quarterly Effective Window Rate falling from 62% to 37% under Hermes and more sharply from 68% to 23% under ReAct, alongside substantial reductions in environment tool calls\. Similar patterns of Activity Decay appear across several other models under both frameworks\. Monthly net profit shows how business performance evolves alongside these activity patterns\. Such operational decline often originates in strategic drift\.

#### Strategic Coherence\.

Strategic Coherence comprises two complementary dimensions\. Goal Consistency concerns whether decisions across time continue to serve the long\-term business objective, whereas Evidence\-Calibrated Adaptation concerns how an agent maintains or revises its policy as feedback arrives at different latencies and evidence accumulates over time\.

Goal Consistency\.Goal Consistency fails when agents lose the autonomy to keep pursuing terminal net assets\. Under Control\-Loop Narrowing, the sourcing and operating loop gradually collapses into reactive handling of Upstream Supplier Events, with little self\-initiated sourcing, repricing, replacement, or diagnosis\. For ReAct Qwen3\.7\-Max, an SWR of 11\.1% coincides with supply chain checks rising from 14% to 34% of its remaining tool calls\. Related activity decay also appears in Claude Opus 4\.8, GLM\-5\.2, DeepSeek\-V4\-Pro, and Qwen3\.7\-Plus under Hermes\. At the extreme, Premature Abandonment occurs when an agent concludes that the store cannot recover although feasible actions remain\. In one Hermes Kimi K2\.6 run, the agent made this judgment on Day 104 and then took no environment action in 355 of the remaining 523 decision windows\. In both cases, the agent waits for external events or time to change the store instead of operating it autonomously\.

Evidence\-Calibrated Adaptation\.Evidence\-Calibrated Adaptation examines whether agents revise their policies as liquidity, seasonal demand, and accumulated experience change\.

Under initial liquidity constraints, Human and agents operate within similar low price ranges\. As liquidity increases, Human operators broaden the procurement price range and selectively return to lower price and higher throughput products when higher value experiments underperform\. Their mean active listing procurement prices increase from between RMB 43\.4 and RMB 53\.1 in the first three months to between RMB 58\.7 and RMB 90\.8 in the last three months, whereas GLM\-5\.2, DeepSeek\-V4\-Flash, and Kimi K2\.6 retain comparatively flat listing price trajectories\.

Dynamic market demand makes Product Sourcing a continual portfolio allocation problem rather than a one\-time selection decision\. Figure[6](https://arxiv.org/html/2607.28956#Sx4.F6)first measures monthly portfolio alignment, with Human rising from 56\.1 in June to above 80 in December and January while Rule\-based remains near the catalog median and LLM improvements are weaker or less consistent\. Figure[5](https://arxiv.org/html/2607.28956#Sx4.F5)\(c\) then reports realized profit, with Human peaking in winter while the winter profit gains of Hermes Qwen3\.7\-Max and GPT\-5\.6 Sol fade in spring\. Claude Opus 4\.8 further shows that demand alignment alone is insufficient, since its stronger alignment in later months coincides with a shelf contraction from 37\.3 to 12\.0 products and no corresponding profit improvement\. Together, the figures show that effective long horizon operation requires alignment with changing demand, sufficient portfolio breadth, and conversion of that alignment into realized returns\.

![Refer to caption](https://arxiv.org/html/2607.28956v1/x6.png)Figure 6:Monthly Product Sourcing across four representative models\. Scores are listing\-hour\-weighted catalog demand percentiles of active products in each month\. Lines and bands show repeat means and standard deviations\.Memory traces further show how local errors become persistent policies\. In one Hermes Claude Opus 4\.8 run, the agent falsely inferred that removing weak listings would concentrate traffic on the remaining products, while its shelf contracted from 47 active listings on Day 54 to three on Day 322 despite independent demand opportunities for every listing\. In one Hermes Qwen3\.7\-Max run, the agent misremembered Day 285 as the endpoint on Day 282 and stopped filling vacant slots with 83 days remaining, correcting the error only after simulated time advanced beyond the assumed endpoint\.

## Conclusion

We introduced MerchantBench to evaluate Long\-Term Coherence through persistent seller\-side e\-commerce operation in a 365\-day order\-level environment grounded in 98,843 real e\-commerce product records\. Across 48 runs involving eight LLMs and two agent frameworks, LLM agents exhibit a substantial gap from the Human baseline in system\-level performance\. Trace analyses further show that weaker outcomes accompany declining operational activity, premature goal abandonment, and strategy changes that are not calibrated to accumulated evidence\.

## References

- Alibaba Group \(n\.d\.\)1688: china’s leading domestic wholesale marketplace\.Note:https://www\.alibabagroup\.com/en\-US/about\-alibaba\-businesses\-1941299332078632960Accessed July 24, 2026Cited by:[Real\-World Data Grounding](https://arxiv.org/html/2607.28956#Sx3.SSx2.p1.1)\.
- Andon Labs \(2025\)Vending\-Bench 2\.Note:https://andonlabs\.com/evals/vending\-bench\-2Accessed July 27, 2026Cited by:[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- P\. Anokhin, R\. Khalikov, S\. Rebrikov, V\. Volkov, A\. Sorokin, and V\. Bissonnette \(2025\)HeroBench: a benchmark for long\-horizon planning and structured reasoning in virtual worlds\.External Links:2508\.12782,[Link](https://arxiv.org/abs/2508.12782)Cited by:[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- Anthropic \(2026\)Introducing Claude Opus 4\.8\.Note:https://www\.anthropic\.com/news/claude\-opus\-4\-8Accessed July 25, 2026Cited by:[Agent Configurations and Baselines\.](https://arxiv.org/html/2607.28956#Sx4.SSx1.SSS0.Px1.p1.1)\.
- A\. Backlund and L\. Petersson \(2025\)Vending\-bench: a benchmark for long\-term coherence of autonomous agents\.External Links:2502\.15840,[Link](https://arxiv.org/abs/2502.15840)Cited by:[Introduction](https://arxiv.org/html/2607.28956#Sx1.p3.1),[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- J\. Baek, Y\. Fu, W\. Ma, and T\. Peng \(2026\)AI agents for inventory control: human\-llm\-or complementarity\.External Links:2602\.12631,[Link](https://arxiv.org/abs/2602.12631)Cited by:[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- G\. Bansal, W\. Hua, Z\. Huang, A\. Fourney, A\. Swearngin, W\. Epperson, T\. Payne, J\. M\. Hofman, B\. Lucier, C\. Singh, M\. Mobius, A\. Nambi, A\. Yadav, K\. Gao, D\. M\. Rothschild, A\. Slivkins, D\. G\. Goldstein, H\. Mozannar, N\. Immorlica, M\. Murad, M\. Vogel, S\. Kambhampati, E\. Horvitz, and S\. Amershi \(2025\)Magentic marketplace: an open\-source environment for studying agentic markets\.External Links:2510\.25779,[Link](https://arxiv.org/abs/2510.25779)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- DeepSeek\-AI \(2026\)DeepSeek\-V4 preview release\.Note:https://api\-docs\.deepseek\.com/news/news260424/Accessed July 25, 2026Cited by:[Agent Configurations and Baselines\.](https://arxiv.org/html/2607.28956#Sx4.SSx1.SSS0.Px1.p1.1)\.
- S\. Ding, X\. Dai, L\. Xing, S\. Ding, Z\. Liu, J\. Yang, P\. Yang, Z\. Zhang, X\. Wei, X\. Fang, Y\. Ma, H\. Duan, J\. Shao, J\. Wang, D\. Lin, K\. Chen, and Y\. Zang \(2026\)WildClawBench: a benchmark for real\-world, long\-horizon agent evaluation\.External Links:2605\.10912,[Link](https://arxiv.org/abs/2605.10912)Cited by:[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- Z\. Du, T\. Li, Y\. Zhang, and H\. Zhang \(2026\)EComAgentBench: benchmarking shopping agents on long\-horizon tasks with distributed hidden intent\.External Links:2606\.17698,[Link](https://arxiv.org/abs/2606.17698)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- X\. Hu, J\. Xia, S\. Xu, K\. Song, Y\. Yuan, G\. Zhang, J\. Ren, B\. Feng, L\. Lu, T\. Zeng, J\. Liu, M\. Liu, H\. Zhu, Y\. E\. Jiang, W\. Wang, and W\. Zhou \(2026\)EcoGym: evaluating llms for long\-horizon plan\-and\-execute in interactive economies\.External Links:2602\.09514,[Link](https://arxiv.org/abs/2602.09514)Cited by:[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- L\. K\. Jang, J\. Y\. Koh, D\. Fried, and R\. Salakhutdinov \(2026\)Odysseys: benchmarking web agents on realistic long horizon tasks\.External Links:2604\.24964,[Link](https://arxiv.org/abs/2604.24964)Cited by:[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- L\. P\. Kaelbling, M\. L\. Littman, and A\. R\. Cassandra \(1998\)Planning and acting in partially observable stochastic domains\.Artificial Intelligence101\(1–2\),pp\. 99–134\.External Links:[Document](https://dx.doi.org/10.1016/S0004-3702%2898%2900023-X)Cited by:[Task Formulation](https://arxiv.org/html/2607.28956#Sx3.SSx1.p1.16)\.
- X\. Liu, H\. Yu, H\. Zhang, Y\. Xu, X\. Lei, H\. Lai, Y\. Gu, H\. Ding, K\. Men, K\. Yang, S\. Zhang, X\. Deng, A\. Zeng, Z\. Du, C\. Zhang, S\. Shen, T\. Zhang, Y\. Su, H\. Sun, M\. Huang, Y\. Dong, and J\. Tang \(2024\)AgentBench: evaluating llms as agents\.InProceedings of the International Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2308.03688)Cited by:[Introduction](https://arxiv.org/html/2607.28956#Sx1.p1.1)\.
- H\. Luo, H\. Zhang, X\. Zhang, H\. Wang, Z\. Qin, W\. Lu, G\. Ma, H\. He, Y\. Xie, Q\. Zhou, Z\. Hu, H\. Mi, Y\. Wang, N\. Tan, H\. Chen, Y\. R\. Fung, C\. Yuan, and L\. Shen \(2025\)UltraHorizon: benchmarking agent capabilities in ultra long\-horizon scenarios\.External Links:2509\.21766,[Link](https://arxiv.org/abs/2509.21766)Cited by:[Introduction](https://arxiv.org/html/2607.28956#Sx1.p1.1),[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- Moonshot AI \(2026\)Kimi K2\.6: advancing open\-source coding\.Note:https://www\.kimi\.com/blog/kimi\-k2\-6Accessed July 25, 2026Cited by:[Agent Configurations and Baselines\.](https://arxiv.org/html/2607.28956#Sx4.SSx1.SSS0.Px1.p1.1)\.
- Nous Research \(2026\)Hermes agent\.Note:GitHub repository,https://github\.com/NousResearch/hermes\-agentAccessed July 25, 2026Cited by:[Appendix K](https://arxiv.org/html/2607.28956#A11.p1.1),[Agent Configurations and Baselines\.](https://arxiv.org/html/2607.28956#Sx4.SSx1.SSS0.Px1.p1.1)\.
- OpenAI \(2026\)GPT\-5\.6: frontier intelligence that scales with your ambition\.Note:https://openai\.com/index/gpt\-5\-6/Accessed July 25, 2026Cited by:[Agent Configurations and Baselines\.](https://arxiv.org/html/2607.28956#Sx4.SSx1.SSS0.Px1.p1.1)\.
- Qwen Team \(2026\)Qwen3\.7 model releases\.Note:https://docs\.qwencloud\.com/changelog/modelsAccessed July 25, 2026Cited by:[Agent Configurations and Baselines\.](https://arxiv.org/html/2607.28956#Sx4.SSx1.SSS0.Px1.p1.1)\.
- C\. Savadikar, M\. Zhao, Y\. Zhu, H\. Li, S\. Xie, A\. Castelo, T\. Wu, and L\. Wang \(2026\)ShopGym: an integrated framework for realistic simulation and scalable benchmarking of e\-commerce web agents\.External Links:2605\.16116,[Link](https://arxiv.org/abs/2605.16116)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- I\. Sugiura, D\. Hattori, K\. Araragi, K\. Ogawa, S\. Onose, T\. Makino, T\. Usuki, and T\. Ishida \(2026\)CoffeeBench: benchmarking long\-horizon llm agents in heterogeneous multi\-agent economies\.External Links:2606\.16613,[Link](https://arxiv.org/abs/2606.16613)Cited by:[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- H\. Trivedi, T\. Khot, M\. Hartmann, R\. Manku, V\. Dong, E\. Li, S\. Gupta, A\. Sabharwal, and N\. Balasubramanian \(2024\)AppWorld: a controllable world of apps and people for benchmarking interactive coding agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 16022–16076\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.850),[Link](https://arxiv.org/abs/2407.18901)Cited by:[Introduction](https://arxiv.org/html/2607.28956#Sx1.p1.1)\.
- H\. Wang, X\. Peng, X\. Huang, Y\. Huang, M\. Gong, C\. Yang, Y\. Liu, and L\. Jiang \(2025a\)ECom\-bench: can llm agent resolve real\-world e\-commerce customer support issues?\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track,Suzhou \(China\),pp\. 276–284\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-industry.19),[Link](https://arxiv.org/abs/2507.05639)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- I\. Y\. Wang, K\. Chong, X\. Wang, X\. Yan, D\. Kong, C\. Ju, M\. Chen, S\. Xiao, S\. Han, and J\. Chen \(2025b\)Evaluating multi\-turn bargain skills in llm\-based seller agent\.External Links:2509\.06341,[Link](https://arxiv.org/abs/2509.06341)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- J\. Wang, K\. Xiao, Q\. Sun, H\. Zhao, T\. Luo, J\. D\. Zhang, and X\. Zeng \(2026a\)ShoppingBench: a real\-world intent\-grounded shopping benchmark for llm\-based agents\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 33521–33529\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i39.40640),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40640)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- P\. Wang, Y\. Wu, X\. Song, W\. Wang, G\. Chen, Z\. Li, K\. Yan, K\. Deng, Q\. Liu, S\. Zhao, S\. Xiong, X\. Liu, X\. Chen, W\. Deng, W\. Su, and B\. Zheng \(2026b\)ShopSimulator: evaluating and exploring rl\-driven llm agent for shopping assistants\.External Links:2601\.18225,[Link](https://arxiv.org/abs/2601.18225)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- W\. Wang, D\. Han, D\. M\. Diaz, J\. Xu, V\. Rühle, and S\. Rajmohan \(2025c\)OdysseyBench: evaluating llm agents on long\-horizon complex office application workflows\.External Links:2508\.09124,[Link](https://arxiv.org/abs/2508.09124)Cited by:[Introduction](https://arxiv.org/html/2607.28956#Sx1.p1.1),[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- S\. Yao, H\. Chen, J\. Yang, and K\. Narasimhan \(2022\)WebShop: towards scalable real\-world web interaction with grounded language agents\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 20744–20757\.External Links:[Document](https://dx.doi.org/10.52202/068431-1508),[Link](https://arxiv.org/abs/2207.01206)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhan \(2025\)τ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.InProceedings of the International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=roNSXZpUDN)Cited by:[Introduction](https://arxiv.org/html/2607.28956#Sx1.p1.1)\.
- S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao \(2023\)ReAct: synergizing reasoning and acting in language models\.InProceedings of the International Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2210.03629)Cited by:[Agent Configurations and Baselines\.](https://arxiv.org/html/2607.28956#Sx4.SSx1.SSS0.Px1.p1.1)\.
- M\. Yuan, Z\. Zhou, X\. Xiong, W\. Wu, J\. Sun, J\. Song, K\. Cui, B\. Wang, H\. Wu, Y\. Li, D\. Lu, H\. Lu, Q\. Zhen, X\. Wang, J\. Deng, Y\. Yang, C\. Chen, B\. Zheng, A\. Su, X\. Yu, H\. Zou, S\. Agashe, X\. H\. Lu, M\. Kaur, Z\. Qi, V\. S\. Chen, F\. Sala, D\. Liu, J\. Lin, Z\. Yu, Y\. Su, S\. Reddy, X\. E\. Wang, P\. Qi, T\. Xie, and T\. Yu \(2026\)OSWorld 2\.0: benchmarking computer use agents on long\-horizon real\-world tasks\.External Links:2606\.29537,[Link](https://arxiv.org/abs/2606.29537)Cited by:[Introduction](https://arxiv.org/html/2607.28956#Sx1.p1.1),[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- Z\.ai \(2026\)GLM\-5\.2: built for long\-horizon tasks\.Note:https://z\.ai/blog/glm\-5\.2Accessed July 25, 2026Cited by:[Agent Configurations and Baselines\.](https://arxiv.org/html/2607.28956#Sx4.SSx1.SSS0.Px1.p1.1)\.
- L\. Zhang, J\. Wang, J\. Wu, and Z\. Zhang \(2026\)RetailBench: evaluating long\-horizon autonomous decision\-making and strategy stability of llm agents in realistic retail environments\.External Links:2603\.16453,[Link](https://arxiv.org/abs/2603.16453)Cited by:[Introduction](https://arxiv.org/html/2607.28956#Sx1.p3.1),[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.
- X\. Zhang, S\. Prasad, D\. Wang, Q\. Zeng, S\. Wang, W\. Yan, and M\. Hans \(2025\)A functionality\-grounded benchmark for evaluating web agents in e\-commerce domains\.External Links:2508\.15832,[Link](https://arxiv.org/abs/2508.15832)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- K\. Zhao, Z\. Meng, Z\. Xie, J\. Duan, Y\. Hu, Z\. Liu, and S\. Cao \(2026\)EComStage: stage\-wise and orientation\-specific benchmarking for large language models in e\-commerce\.External Links:2601\.02752,[Link](https://arxiv.org/abs/2601.02752)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- Y\. Zheng, H\. Duan, Z\. Zhang, Y\. Zhu, X\. Min, and G\. Zhai \(2026\)Market\-bench: benchmarking large language models on economic and trade competition\.External Links:2604\.05523,[Link](https://arxiv.org/abs/2604.05523)Cited by:[Agent Evaluation in Commerce\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px1.p1.1)\.
- S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried, U\. Alon, and G\. Neubig \(2024\)WebArena: a realistic web environment for building autonomous agents\.InProceedings of the International Conference on Learning Representations,External Links:[Link](https://arxiv.org/abs/2307.13854)Cited by:[Introduction](https://arxiv.org/html/2607.28956#Sx1.p1.1)\.
- Y\. Zhu, Y\. Zhan, X\. Huang, Y\. Chen, Y\. Chen, J\. Wei, W\. Feng, Y\. Zhou, H\. Hu, and J\. Ye \(2023\)OFCOURSE: a multi\-agent reinforcement learning environment for order fulfillment\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 34765–34777\.External Links:[Document](https://dx.doi.org/10.52202/075280-1510),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/6d0cfc5db3feeabf6762129ba91bd3a1-Abstract-Datasets_and_Benchmarks.html)Cited by:[Long\-Horizon Agent Evaluation\.](https://arxiv.org/html/2607.28956#Sx2.SS0.SSS0.Px2.p1.1)\.

## Appendix ADemand and Rating Dynamics

#### Listing Exposure\.

Letam,i,t≥0a\_\{m,i,t\}\\geq 0denote the elapsed time in days since the current listing of productiiby merchantmmwas activated\. We setℓm,i,t=g​\(am,i,t\)\\ell\_\{m,i,t\}=g\(a\_\{m,i,t\}\), where

g​\(a\)=\{ℓ0\+\(1−ℓ0\)​a/Tr,a≤Tr,ℓmin\+\(1−ℓmin\)​e−κ​\(a−Tr\),a\>Tr,g\(a\)=\\left\\\{\\begin\{array\}\[\]\{ll\}\\ell\_\{0\}\+\(1\-\\ell\_\{0\}\)a/T\_\{r\},&a\\leq T\_\{r\},\\\\ \\ell\_\{\\min\}\+\(1\-\\ell\_\{\\min\}\)e^\{\-\\kappa\(a\-T\_\{r\}\)\},&a\>T\_\{r\},\\end\{array\}\\right\.\(4\)This factor models a cold start through a linear exposure ramp followed by exponential decay towardℓmin\\ell\_\{\\min\}\.

#### Store Rating Dynamics\.

MerchantBench derives store reputation from realized terminal order outcomes and publishes the rating after each completed simulated day\. Each orderjj, except those resulting in Cancellation or insufficient balance failure, contributes an outcome\-dependent experience scoreuju\_\{j\}and evidence weightvjv\_\{j\}\. For merchantmmon daydd, the store rating is

Rm,d=α​R0\+∑j∈𝒞m,dγd−1−dj​vj​ujα\+∑j∈𝒞m,dγd−1−dj​vj\.R\_\{m,d\}=\\frac\{\\alpha R\_\{0\}\+\\sum\_\{j\\in\\mathcal\{C\}\_\{m,d\}\}\\gamma^\{d\-1\-d\_\{j\}\}v\_\{j\}u\_\{j\}\}\{\\alpha\+\\sum\_\{j\\in\\mathcal\{C\}\_\{m,d\}\}\\gamma^\{d\-1\-d\_\{j\}\}v\_\{j\}\}\.\(5\)Here𝒞m,d\\mathcal\{C\}\_\{m,d\}contains the eligible orders resolved before daydd, whileR0R\_\{0\}andα\\alphadefine a prior that stabilizes sparse evidence andγ\\gammadiscounts older outcomes\. The continuous rating is mapped to a discrete star level whose associated demand multiplier becomesrm,tr\_\{m,t\}in subsequent order generation\. The simulator separately aggregates the same outcome evidence for each listing as a diagnostic product rating, but listing ratings do not affect demand\.

## Appendix BData Collection and Filtering

MerchantBench constructs its Product Catalog from real\-world e\-commerce data collected from 1688\. All source product and supplier attributes in the catalog originate from the platform\. Each Product Record contains a 365 day product level order history, and the data span ten first level product categories\. We remove records with missing product or supplier identifiers, missing product or supplier names, categories outside the selected ten categories, missing or nonpositive prices, invalid product or supplier attributes, or incomplete or invalid daily order histories\. After filtering, the resulting Product Catalog contains 98,843 Product Records from 36,576 suppliers\.

Figure[11](https://arxiv.org/html/2607.28956#A12.F11)summarizes the composition of the filtered data\. The catalog retains substantial variation across categories in both record coverage and supplier prices\. The joint risk distribution further shows that the calibrated data contain diverse combinations of Downstream Order Outcome probabilities and Upstream Supplier Event intensities\.

Figure[12](https://arxiv.org/html/2607.28956#A12.F12)presents the temporal structure of the 365 day demand histories\. The aggregate curve retains major shopping peaks and seasonal changes, while the representative product curves show distinct demand cycles for red envelope, electric fan, and hot water bag products\.

## Appendix CFull Operational Coherence Profiles

Figures[13](https://arxiv.org/html/2607.28956#A12.F13)and[14](https://arxiv.org/html/2607.28956#A12.F14)extend the monthly trajectory diagnostics to Human and all eight models under Hermes and ReAct, respectively\. Under Hermes, GPT\-5\.6 Sol remains almost fully active through February before declining to 81% and 77% in the final two months, while DeepSeek\-V4\-Flash stays between 70% and 87% throughout the year\. Qwen3\.7\-Max declines from 64% to 31%, Claude Opus 4\.8 contracts from 46 to eight month end active listings, and Kimi K2\.6 recovers to 78% and 74% effective windows only in the final two months\. Under ReAct, GPT\-5\.6 Sol sustains near complete activity, while Kimi K2\.6 records the lowest Sustained Window Rate at 10\.6%\.

## Appendix DNet Asset Curves

Figures[15](https://arxiv.org/html/2607.28956#A12.F15)and[16](https://arxiv.org/html/2607.28956#A12.F16)present the complete daily net asset curves under ReAct and Hermes, respectively\. Figure[17](https://arxiv.org/html/2607.28956#A12.F17)compares the net asset curves of all eight models with the Human and Rule\-based baselines\. Figure[18](https://arxiv.org/html/2607.28956#A12.F18)shows the corresponding variation among individual runs\.

## Appendix EScale and Unit Profitability

Figure[7](https://arxiv.org/html/2607.28956#A5.F7)separates realized performance into order scale and net profit per order across all 48 LLM runs\. Qwen3\.7\-Max with Hermes produces the largest total profit in one run by combining high profit per order with moderate scale, whereas one Kimi K2\.6 ReAct run reaches 3,004 orders at only RMB 11\.1 net profit per order, illustrating that scale alone does not guarantee the highest return\. The wide dispersion across repeated runs shows that the evaluated configurations do not consistently reproduce the same balance between operating scale and unit profitability\.

![[Uncaptioned image]](https://arxiv.org/html/2607.28956v1/x7.png)

Figure 7:Order scale and unit profitability across 48 LLM runs\. Colors denote models and markers denote agent frameworks\. Dotted curves indicate equal cumulative net profit at RMB 20,000, 40,000, and 80,000\.

## Appendix FTool Use and Product Selection Analysis

Across the 16 LLM configurations, greater tool use and broader product exploration are associated with higher final net assets, suggesting that sustained intervention matters in long horizon operation as shown in Figure[8](https://arxiv.org/html/2607.28956#A6.F8)\. However, the dispersion around both fitted trends shows that activity volume alone is insufficient, since models differ in how effectively they translate actions into business value\.

![[Uncaptioned image]](https://arxiv.org/html/2607.28956v1/x8.png)

Figure 8:Final net assets against environment tool calls and new products tried\. Both axes are logarithmic, and dashed lines fit the 16 LLM configurations only\. Human and Rule\-based are shown as references\.

## Appendix GFull Monthly Product Sourcing Analysis

For runrrand monthmm, letqp,mq\_\{p,m\}denote the demand percentile of productppamong the complete Product Catalog in that month, and letwr,p,mw\_\{r,p,m\}denote its active listing hours\. The Monthly Demand Alignment Percentile is

Sr,m=∑pwr,p,m​qp,m∑pwr,p,m\.S\_\{r,m\}=\\frac\{\\sum\_\{p\}w\_\{r,p,m\}q\_\{p,m\}\}\{\\sum\_\{p\}w\_\{r,p,m\}\}\.\(6\)The catalog ranking uses real demand within each month, and the monthly denominator is the total active listing hours in that run and month\. A score of 80 means that an average listed product\-hour belongs to a product with greater demand than 80% of the catalog\.

Figure[G](https://arxiv.org/html/2607.28956#A7)shows that Claude Opus 4\.8 improves under both frameworks, while several other model and framework combinations plateau or decline during the year\. Human maintains the strongest demand alignment, whereas Rule\-based remains close to the catalog median with little seasonal change\.

## Appendix HTime\-aware Sourcing Gain

Monthly Demand Alignment can change even when an agent retains a fixed portfolio because product demand itself varies over time\. Time\-aware Sourcing Gain therefore compares the actual monthly portfolio with a no\-reallocation counterfactual constructed from the same run\. LetWr,p=∑mwr,p,mW\_\{r,p\}=\\sum\_\{m\}w\_\{r,p,m\}denote the annual listing hours assigned to productpp\. The no\-reallocation counterfactual evaluates the same annual product mix in each month as

Br,m=∑pWr,p​qp,m∑pWr,p\.B\_\{r,m\}=\\frac\{\\sum\_\{p\}W\_\{r,p\}q\_\{p,m\}\}\{\\sum\_\{p\}W\_\{r,p\}\}\.\(7\)We define Time\-aware Sourcing Gain as

Gr=∑mHr,m​\(Sr,m−Br,m\)∑mHr,m,Hr,m=∑pwr,p,m\.G\_\{r\}=\\frac\{\\sum\_\{m\}H\_\{r,m\}\\left\(S\_\{r,m\}\-B\_\{r,m\}\\right\)\}\{\\sum\_\{m\}H\_\{r,m\}\},\\qquad H\_\{r,m\}=\\sum\_\{p\}w\_\{r,p,m\}\.\(8\)Positive values indicate that the agent allocates listing exposure to products that better match the current month than its own fixed annual product mix\. An unchanged product mix yields a gain of zero\.

Across the 48 LLM runs, higher Time\-aware Sourcing Gain is positively associated with final net assets\. This result indicates that stronger seasonal portfolio reallocation accompanies better long horizon business outcomes, although final net assets also depend on pricing, cash flow, and order management decisions\.

![[Uncaptioned image]](https://arxiv.org/html/2607.28956v1/x9.png)

Figure 10:Time\-aware Sourcing Gain and final net assets\. The gain compares actual monthly demand alignment with a counterfactual that holds each run’s annual product mix fixed\. The dashed line fits the 48 LLM runs, with Human and Rule\-based shown as references\.

## Appendix IHermes Case Study

This section examines how the evaluated models use the additional Hermes capabilities listed in Tables[5](https://arxiv.org/html/2607.28956#A12.T5)and[6](https://arxiv.org/html/2607.28956#A12.T6), focusing on programmatic code use and on the creation and revision of skills\.

### Code Use

Native code use differs sharply across the 24 Hermes runs\. GPT\-5\.6 Sol dominates native code use, issuing 15execute\_codecalls across two runs and 13 in one run\. At the first decision window it converts 50 candidate products into a priced launch portfolio by applyingp=max⁡\(1\.70​c,c\+6\)p=\\max\(1\.70c,c\+6\)to each procurement costccand rounding the result upward to a price ending in 0\.9, then passes the computed pairs to MerchantBench listing calls\. A companion script validates the portfolio by printing the product count, average cost, minimum margin, and total procurement outlay before any listing action\. Claude Opus 4\.8, Qwen3\.7\-Max, and DeepSeek\-V4\-Pro each invokeexecute\_codeonce, while Qwen3\.7\-Max and DeepSeek\-V4\-Flash each invoketerminalonce\. GLM\-5\.2, Qwen3\.7\-Plus, and Kimi K2\.6 never invoke either code tool\. Claude Opus 4\.8 instead channels its analysis through 261memorycalls across its three runs, maintaining day stamped operating hypotheses labeled as validated or rejected\. Hermes code capabilities therefore amplify models that already favor quantitative bookkeeping while leaving purely verbal operators unchanged\.

### Skill Evolution

Hermes periodically runs a background review that can distill the accumulated trace into a named skill and later patch it\. Final Hermes profiles contain a run created RealShop skill in 17 of 24 runs and for seven of the eight models\. Seventeen of the 18 created skills record at least one subsequent use, and the only unused skill was created by Qwen3\.7\-Max\. GPT\-5\.6 Sol, Kimi K2\.6, and Qwen3\.7\-Plus create a skill in all three runs, while Claude Opus 4\.8 creates none\. GPT\-5\.6 Sol creates arealshop\-store\-operationsskill in each run, with 9, 15, and 7 patches and 17, 10, and 13 recorded uses\. These skills frame the task as repeated portfolio allocation and prescribe per window procedures for evidence collection, risk handling, unit economics, portfolio revision, and verification\. DeepSeek\-V4\-Pro and DeepSeek\-V4\-Flash each reach 20 patches in one run, while Qwen3\.7\-Plus creates two complementary store operation and optimization skills in one run\. These traces show that Hermes can convert experience into explicit procedural knowledge, but the quality of the resulting skills ranges from evidence cited rule revision to unstructured accumulation and duplication\.

## Appendix JDetailed Experimental Configuration

#### Evaluation Protocol\.

The Rule\-based baseline completes three runs\. It handles abnormal states through daily checks\. It delists products with no sales for seven consecutive days or products affected by Price Change, Product Delisting, or Shipment Delay, then fills available listing slots with new products selected using keywords from the daily market report\. Three participants with no prior e\-commerce operating experience each complete one run spanning 365 simulated days through the human operations dashboard over five calendar days\. We report the mean across three LLM or Rule\-based runs or across the three human participants\.

#### Store Configuration\.

Each agent operates a store initialized with a cash balance of RMB 2,000 and a security deposit of RMB 1,000 and supports at most 50 active listings\. The fixed fines are RMB 8 for Return and Refund, RMB 5 for Bad Review, Stockout, or insufficient balance, and RMB 3 for Late Shipment\. Cancellation and Returnless Refund incur no additional fine\. These fine settings are consistent with the corresponding real\-world platform rules\.

#### Context Management\.

To support interaction over the full 365 day horizon, both frameworks compress long interaction histories\. When a ReAct history reaches 160,000 tokens, the evaluated model receives a reminder to summarize important information into persistent memory before the history is truncated to the most recent 30,000 tokens\. Hermes uses its default context summarization procedure and sets the summary model to the same evaluated model\.

#### Evaluation Metrics\.

Net Profit Margin is terminal net profit divided by GMV\. Average Store Rating and Average Active Listings are daily means of the published store rating and active listing count\. Order Anomaly Rate is the share of generated orders affected by at least one cancellation, stockout, insufficient balance failure, late shipment, refund, or bad review\. Sustained Window Rate is the minimum share of scheduled decision windows containing at least one environment tool call across all rolling 30 day periods\. Total Tool Calls excludes calls that end a decision window and all framework internal tools\.

#### Demand and Supplier Configuration\.

Listing exposure usesℓ0=0\.2\\ell\_\{0\}=0\.2,Tr=14T\_\{r\}=14days,κ=0\.0092\\kappa=0\.0092, andℓmin=0\.10\\ell\_\{\\min\}=0\.10\. Supplier inventory is initialized at 20 to 399 units, with capacities of 50 to 499 units and hourly replenishment of 1 to 19 units\. Baseline supplier dispatch and logistics times span 1 to 47 and 12 to 72 hours\. Supplier abnormalities recover after 168 to 672 hours, Price Change applies a factor sampled from 0\.9 to 1\.5, and Shipment Delay adds 12 to 96 hours\.

#### Order Lifecycle Configuration\.

The default promised shipment time is 48 hours\. The transition from Delivered to Settled and the realization of outcomes after delivery are each sampled within 168 hours\.

#### Rating Configuration\.

The outcome scoresuju\_\{j\}are4\.54\.5,3\.03\.0,2\.02\.0,1\.51\.5,1\.01\.0, and1\.01\.0, and the evidence weightsvjv\_\{j\}are1\.01\.0,1\.01\.0,1\.01\.0,2\.02\.0,2\.02\.0, and3\.03\.0, ordered as a normal transition to Settled, Late Shipment, Return and Refund, Returnless Refund, Bad Review, and Stockout\. Store ratings useR0=4\.0R\_\{0\}=4\.0,α=20\\alpha=20, andγ=2−1/30\\gamma=2^\{\-1/30\}\. The rating thresholds2\.502\.50,3\.303\.30,3\.803\.80, and4\.204\.20map to demand multipliers0\.100\.10,0\.350\.35,0\.800\.80,1\.001\.00, and1\.201\.20across the five resulting intervals\. Diagnostic listing ratings use the same initial rating and prior weight with a 90 day evidence half life\.

## Appendix KTool Sets and Skills

ReAct uses only the 26 MerchantBench tools through which the agent accesses observable fields and controls the store\. Table[2](https://arxiv.org/html/2607.28956#A12.T2)lists the complete MerchantBench tool set\. Among them, theget\_daily\_reporttool returns the report published for the current simulation date and states that its evidence is current only through the previous date\. Table[3](https://arxiv.org/html/2607.28956#A12.T3)gives an English translation of the Daily Market Report for June 10, 2025\. Hermes provides the same MerchantBench tools together with the built in tools and skills listed in Tables[5](https://arxiv.org/html/2607.28956#A12.T5)and[6](https://arxiv.org/html/2607.28956#A12.T6)\(Nous Research[2026](https://arxiv.org/html/2607.28956#bib.bib25)\)\.

## Appendix LAgent Inputs and Observable Fields

The environment gives each evaluated agent one system prompt at registration and a compact observation at every 12 hour decision window\. Both frameworks receive the shared MerchantBench task prompt in Table[9](https://arxiv.org/html/2607.28956#A12.T9)and access the same 26 MerchantBench tools\. For ReAct, the MerchantBench task prompt is the complete system prompt and the 26 tools are the entire action space\. Hermes wraps the same task prompt within its official runtime template, shown in Table[11](https://arxiv.org/html/2607.28956#A12.T11), and augments the 26 MerchantBench tools with its built in tools and skills\. Table[18](https://arxiv.org/html/2607.28956#A12.T18)gives a representative observation containing order transitions, an upstream price change, current finances, and the shop rating\.

MerchantBench is partially observable because the merchant interface returns current public and realized operational evidence while retaining future demand, event hazards, and presampled outcomes inside the simulator\. Tables[19](https://arxiv.org/html/2607.28956#A12.T19)to[21](https://arxiv.org/html/2607.28956#A12.T21)summarize this boundary\. Visible fields are returned directly by at least one merchant tool\. Some visible order fields remain empty until the corresponding lifecycle event is realized\. Hidden fields are never returned through merchant tools\.

![Refer to caption](https://arxiv.org/html/2607.28956v1/x10.png)Figure 11:Composition and calibrated distributions of the filtered data\. Panel \(a\) reports the numbers of Product Records and unique supplier IDs within each category\. Panel \(b\) shows procurement price distributions within each category\. Panel \(c\) shows the joint distribution of Downstream Order Outcome probability and Upstream Supplier Event intensity\.![Refer to caption](https://arxiv.org/html/2607.28956v1/x11.png)Figure 12:Temporal demand patterns in the filtered data\. The top panel presents aggregate daily demand and its seven day mean\. The bottom panels show normalized demand histories for representative red envelope, electric fan, and hot water bag Product Records\.![Refer to caption](https://arxiv.org/html/2607.28956v1/x12.png)Figure 13:Monthly Operational Coherence profiles for the Human baseline and each of the eight Hermes models\. Panels \(a\), \(b\), \(c\), and \(d\) report effective window rate, environment tool calls, month end active listings, and monthly net profit\.![Refer to caption](https://arxiv.org/html/2607.28956v1/x13.png)Figure 14:Monthly Operational Coherence profiles for the Human baseline and each of the eight ReAct models\. Panels \(a\), \(b\), \(c\), and \(d\) report effective window rate, environment tool calls, month end active listings, and monthly net profit\.![Refer to caption](https://arxiv.org/html/2607.28956v1/x14.png)Figure 15:Daily net asset curves for all eight models under ReAct over 365 simulated days\.![Refer to caption](https://arxiv.org/html/2607.28956v1/x15.png)Figure 16:Daily net asset curves for all eight models under Hermes over 365 simulated days\.![Refer to caption](https://arxiv.org/html/2607.28956v1/x16.png)Figure 17:Daily net asset curves for the Human and Rule\-based baselines and all eight models over 365 simulated days\.![Refer to caption](https://arxiv.org/html/2607.28956v1/x17.png)Figure 18:Daily net asset curves for all individual runs\. The first row shows the Human and Rule\-based baselines, and the remaining panels show the eight models\. Hermes and ReAct runs appear together within each model panel, and vertical axis ranges are set independently across panels\.![Refer to caption](https://arxiv.org/html/2607.28956v1/x18.png)Figure[G](https://arxiv.org/html/2607.28956#A7):Monthly Product Sourcing across all eight models\. Monthly Demand Alignment Percentile is the listing\-hour\-weighted mean percentile of active products after ranking the complete catalog by that month’s real demand\. Lines and bands show repeat means and standard deviations, with Human and Rule\-based as shared references\.ToolAccessDescriptionProduct Sourcingget\_daily\_reportReadReturns the daily market report for the current simulation date with market news and opportunity signalssearch\_productsReadSearches the visible Product Catalog using public fieldsget\_product\_detailReadReturns visible product, logistics, rating, and supplier fieldsget\_supplier\_profileReadReturns the public supplier profile and visible product countlist\_supplier\_productsReadLists the currently visible products from one supplierListing and Pricing Controllist\_productWriteAdds products to the store at specified selling pricesdelist\_productWriteRemoves products from the storeadjust\_priceWriteChanges selling prices for active listingsreview\_my\_listingsReadReviews listing age, sales velocity, fines, and fulfillment backlogquery\_my\_listingsReadReturns current listings with cumulative sales, profit, and finesquery\_store\_performanceReadSummarizes store outcomes by day or weekquery\_product\_sales\_statsReadRanks product outcomes and reports abnormality countsCash\-Flow Managementquery\_balanceReadReturns the cash balance, security deposit, funds in transit, receivables, and finesget\_store\_snapshotReadSummarizes orders, supply, cash, listings, and store ratingquery\_platform\_rulesReadReturns capital, settlement, penalty, and closure rulesquery\_cash\_pipelineReadSummarizes receivable aging and active order cost exposureSupplier and Order Monitoringquery\_supply\_chain\_anomaliesReadReturns new or current supplier abnormalities and affected listingsquery\_my\_ordersReadSearches historical orders with logistics and accounting fieldsquery\_open\_ordersReadReturns active orders with fulfillment timing and economicsquery\_order\_updatesReadReturns status changes since the previous observation windowquery\_order\_detailReadReturns one order’s full status timeline, accounting, and penaltiesAgent Support and Controlread\_memory\_docReadReads the run local agent memory documentwrite\_memory\_docWriteReplaces the run local agent memory documentget\_observationReadReturns the current rendered observationlist\_toolsReadReturns tool schemas after scenario filteringend\_of\_stepControlReleases the current decision windowTable 2:MerchantBench merchant tool inventory\.June10MarketOpportunityBrief

Reportdate:June10,2025

Datacurrentthrough:June9,2025

Yesterday’snews

\-TheStateAdministrationforMarketRegulationissuedcomplianceguidancefortheJune18promotionacrossgeneral,livestream,andcross\-bordere\-commerceplatforms\.Theguidanceprohibitsdiscriminatorypricing,falsemarketing,fabricatedtransactions,andsubsidyfraud,andasksplatformstostrengthenmonitoringoflivestreamsandhosts\.

\-Consumersubsidiesexpandedfromeighttotwelveappliancecategories\.Eligibleappliancesreceivea20percentsubsidycappedatRMB2,000,whileselecteddigitalproductsreceivea15percentsubsidycappedatRMB500\.SubsidyclaimsonTaobaoandTmallduringthefirstfourhourstripledrelativetothepreviousDouble11promotion,andfirst\-daysalesonMeituanInstashoppingincreasedby200percentyearoveryear\.

\-TheJune18promotionperiodlengthened\.TmallrunspresalesfromMay13throughJune20,Xiaohongshulaunchedathirty\-daymarketplacepromotiononJune1,andMeituanenteredthepromotionthroughinstantretail\.Competitionnowspansassortment,fulfillment,andshoppingscenariosratherthanpricealone\.

Overalltrend

Searchvolumereached19\.052millionyesterday,up9\.08percentfromsevendaysearlierdespiteadailypullback\.Thetop100termscovered31\.7percentofsearchesacross1,813keywords\.Appliancesrose45\.8percent,othergoodsrose5\.8percent,andsummerapparelandwomenswearrose9\.3percent\.Cleaning,sportsandoutdoorgoods,andtoysdeclinedby10\.4,8\.7,and19\.3percent\.Trafficconcentratedoncoolingappliances,airconditioners,swimwear,andFather’sDayproducts\.

Categorydynamics

\-Appliances\.Neckfansrose141\.6percent,desktopfansrose133\.3percent,electricfansrose118\.8percent,andXiaomiairconditionersrose112\.1percent\.Smallfansandelectricfansbothenteredthetopten\.

\-Othergoods\.Father’sDaygiftsenteredatrank65with257\.9percentgrowth\.Icemakersremainedstrongwith65\.2percentgrowth,andclassmatealbumsrose44\.5percent\.

\-Summerapparelandwomenswear\.Women’ssummershortsrose54\.2percent,women’sshortsrose45\.2percent,andwomen’sswimwearnewlyenteredwith44\.7percentgrowth\.

\-Studentandofficesupplies\.A4printingpapernewlyenteredwith28\.2percentgrowth\.Notebooksandcorrectiontaperose10\.7and7\.1percent\.

\-Homeandcleaninggoods\.Disposablebathtowelsnewlyenteredatrank93with7\.0percentgrowth,indicatingthestartofasummerexpansionwindow\.

Anomaloussearchsignals

\-\#40Greeairconditioner\.54,687searches,up93\.9percentfromrank118\.

\-\#55neckfan\.48,005searches,up141\.6percentfromrank205\.

\-\#59Mideaairconditioner\.46,505searches,up67\.9percentfromrank123\.

\-\#65Father’sDaygift\.43,519searches,up257\.9percentfromrank402\.

Table 3:Daily market report for June 10, 2025, shown here in English translation\.\-\#76A4printingpaper\.39,238searches,up28\.2percentfromrank103\.

\-\#78icemaker\.39,148searches,up65\.2percentfromrank149\.

\-\#86pedestalfan\.36,888searches,up76\.4percentfromrank183\.

\-\#97women’sswimwear\.32,817searches,up44\.7percentfromrank159\.

\-\#98desktopfan\.32,806searches,up133\.3percentfromrank320\.

\-\#23Xiaomiairconditioner\.70,369searches,up112\.1percentand65ranks\.

News\-linkedsuggestions

\-Father’sDaygiftsshowanimmediatesaleswindowbeforeJune15\.Candidateproductsincludeelectricshavers,belts,wallets,men’spersonalcareproducts,andmassagedevices\.

\-Merchantsusingsubsidizedpricesshouldretaininvoicesandtrade\-indocumentationtoreducecompliancerisk\.

\-A15percentdirectdiscounthasbecomecommononDouyinandXiaohongshu\.Highlyrankedproductsmayneedalowermarkuptoremaincompetitiveduringthepromotion\.

Table 4:Daily market report for June 10, 2025, with the English translation continued\.ToolAccessDescriptionExecution and FilesterminalExecuteExecutes shell commands in a persistent environmentprocessManageMonitors and controls background processesexecute\_codeExecuteRuns Python programs that call Hermes tools and process their outputsread\_fileReadReads text files with line numbers and paginationwrite\_fileWriteCreates or replaces files and checks supported formatspatchWriteApplies targeted file edits and returns a unified diffsearch\_filesReadSearches file names and contentsMemory and SkillsmemoryWriteStores durable facts that persist across sessionssession\_searchReadSearches messages from previous Hermes sessionsskills\_listReadLists available skills and their descriptionsskill\_viewReadLoads skill instructions and linked resourcesskill\_manageWriteCreates, revises, or deletes skillsPlanning and CoordinationtodoManageMaintains the task list for the current sessionclarifyInteractRequests clarification, feedback, or a decision from the userdelegate\_taskDelegateAssigns independent tasks to subagentsProjects and Outputproject\_listReadLists available project workspacesproject\_createWriteCreates and activates a project workspaceproject\_switchWriteSwitches the active project workspacetext\_to\_speechGenerateConverts text into speech audioimage\_generateGenerateGenerates or edits images from prompts and referencesTable 5:Built in Hermes tools provided by the official architecture\.Table 6:Built in Hermes skills and their functions\.Table 7:Built in Hermes skills and their functions, continued\.Table 8:Built in Hermes skills and their functions, continued\.YouaretheoperatingagentofasmallRealShopstore\.Youwillperiodicallyreceiveanobservationofthecurrentstorestate\.

Goals:

\-Maximizetotalassets

Operatingperiod:

\-Thestoreoperatesfor365days,startingfrom2025\-06\-01\.

\-Theenvironmentadvancesindiscretestepsof1hour\(s\);youareactivatedevery12hours\.

\-Listing,price,anddelistingactionsinthecurrenthouraffectfuturesalesonly;theydonotaffectordersalreadygeneratedforthecurrenthour\.

\-Eachstep:receiveanobservation,takeyouractions,thencall‘end\_of\_step‘toreleasetheper\-stephookandadvance\.

Initialcapital:

\-balance:2000\.00,usableforprocurement\.

\-deposit\_pool:1000\.00,alockedguaranteeunavailableforprocurement\.

Availableactions:

\-Sourcing:choosewhattosellbasedondemand,cost,qualitysignals,andsupplierreliability\.

\-Storeoperations:managelistings,prices,shelfslots,andcashusagetobalancegrowth,margin,andrisk\.

\-Upstreamsupplierhandling:respondtosupplierpricechanges,delisting,slowershipping,ornegative\-marginrisk\.

\-Downstreamordermanagement:monitororderexceptions,receivables,cash,anddepositrisk\.

\-Useanyavailabletoolsandskills,includinganalysis,automation,andmemorytoolswhenprovided,toimprovelong\-rundecisionsandmaximizenet\_assets\.

Demandandsales:

\-Salesareaffectedbyseasonaldemand,timeofday,saleprice,listinglifecycle,andshoprating\.

\-Newlistingshavelimitedinitialexposure;trafficrampsgraduallyandreachesitsnormallevel14daysafterlisting\.

\-Upstreamcatalogproductratingsarebasedonhistoricaldata\.Theycanbeusefulreferencesignals,buttheydonotnecessarilydeterminefuturesalesperformance\.

Upstreamsupplierabnormalevents:

\-Supplier\-sideabnormaleventsincludepricechange,supplierdelist,andsuppliershippingtimeout\.

\-Theseabnormalstatesmaybetemporaryratherthanpermanent;affectedsuppliersorproductsusuallyrecoverorendafteraperiodoftime\.Whileactive,theycanaffectprocurementcost,saleavailability,oractualshiptime\.

Orderlifecycleandcashfields:

\-Whenacustomerorders,thesystemautomaticallytriestoprocuretheproductatsupplier\_price:balanceisdebitedimmediatelyandin\_transitincreases\.

\-Atdelivery,purchase\_priceleavesin\_transitandsale\_priceentersreceivable\.

Table 9:MerchantBench task prompt used as the ReAct system prompt and embedded within the Hermes system prompt\.\-Normalandbad\-revieworderssettletheirsaleproceedswithin0\-7daysafterdelivery;buyercancellationsandqualityreturnsrecovertheprocurementcost;refund\-onlyordersproducenocashcreditandtheprocurementcostislost\.

\-Anycashcreditfirstrestoresdeposit\_pooltoitsinitialamountof1000\.00;onlytheremainderentersbalance\.

\-Commonstatuslifecycle:ordered\-\>shipped\-\>delivered\-\>settled\_normal;ifshippingexceedsthepromise,theorderenterslatefirstandthencontinuesflowing\.

\-Abnormalterminal/settlementstatesincludecancelled/settled\_refund/settled\_only\_refund/settled\_bad\_review/stockout/insufficient\_balance\.

Fieldlogic:

\-cash\.net\_assets=balance\+deposit\_pool\+in\_transit\+receivable;useitasthemaintotal\-assetsview\.

\-order\.net\_profit=realized\_revenue\-realized\_cost\-total\_penalty\.

\-Finesarealreadydeductedwhenapplied;donotsubtractthemagainfromcash\.net\_assets\.

Penaltyandclosure:

\-Allfinesdeductbalancefirst;anyunpaidremainderdeductsdeposit\_pool\.

\-balancereaching0doesnotclosetheshop;deposit\_poolreaching0closesitimmediatelyandpermanently\.

Platformpenalties:

\-buyercancel\(includingintransit\):noextrafine

\-qualityreturn:fixed8\.00

\-refund\-only:purchasecostislost;noextrafine

\-badreview:fixed5\.00

\-shippingtimeout\(actualshiptime\>promisedshiptime\):fixed3\.00

\-stockoutviolation\(orderarrivesbutsupplierdelisted/qty=0\):fixed5\.00

\-insufficientbalance\(orderarrivesbutbalance<purchaseprice\):fixed5\.00

Activelistingslimit:thisshopmayhaveatmost50activelistingsatonce\.

Emptyshelfslotsreduceproductexposure\.

Shoprating\(updateddaily\):

\-Eachterminalordercountsonce:normal4\.5x1,latebutsettled3x1,refund2x1,refund\-only1\.5x2,badreview1x2,stockout1x3;cancellationsandinsufficient\-balancefailuresareexcluded\.

\-Anewshopstartsat4withpriorweight20;evidencedecayswitha30\-dayhalf\-life\.

\-Scoreranges<2\.5,\[2\.5,3\.3\),\[3\.3,3\.8\),\[3\.8,4\.2\),\>=4\.2mapto1\-5stars;subsequentordertrafficismultipliedbyx0\.1,x0\.35,x0\.8,x1,x1\.2,respectively\.

First\-levelmarketplacecategories:

\-appliances,bags,cleaning,home\_decor,home\_goods,office,pet\_garden,sports,toys,womenswear

Timedisplay:

\-Timeisshownasbothsimulationtimeandcalendartime,forexample‘Day2,Hour22\(2025\-06\-02T22:00:00\)‘;toolargumentsstilluseday/hour\.

Table 10:MerchantBench task prompt used as the ReAct system prompt and embedded within the Hermes system prompt, continued\.YouareHermesAgent,anintelligentAIassistantcreatedbyNousResearch\.Youarehelpful,knowledgeable,anddirect\.Youassistuserswithawiderangeoftasksincludingansweringquestions,writingandeditingcode,analyzinginformation,creativework,andexecutingactionsviayourtools\.Youcommunicateclearly,admituncertaintywhenappropriate,andprioritizebeinggenuinelyusefuloverbeingverboseunlessotherwisedirectedbelow\.Betargetedandefficientinyourexplorationandinvestigations\.

YourunonHermesAgent\(byNousResearch\)\.WhentheuserneedshelpwithHermesitself\-configuring,settingup,using,extending,ortroubleshootingit\-orwhenyouneedtounderstandyourownfeatures,tools,orcapabilities,thedocumentationathttps://hermes\-agent\.nousresearch\.com/docsisyourauthoritativereferenceandalwaysholdsthelatest,mostup\-to\-dateinformation\.Loadthe‘hermes\-agent‘skillwithskill\_view\(name=’hermes\-agent’\)foradditionalguidanceandprovenworkflows,buttreatthedocsasthesourceoftruthwhenthetwodiffer\.

\#Finishingthejob

Whentheuserasksyoutobuild,run,orverifysomething,thedeliverableisaworkingartifactbackedbyrealtooloutput\-notadescriptionofone\.Donotstopafterwritingastub,aplan,orasinglecommand\.Keepworkinguntilyouhaveactuallyexercisedthecodeorproducedtherequestedresult,thenreportwhatrealexecutionreturned\.

Ifatool,install,ornetworkcallfailsandblockstherealpath,saysodirectlyandtryanalternative\(differentpackagemanager,differentapproach,asktheuser\)\.NEVERsubstituteplausible\-lookingfabricatedoutput\(made\-updata,inventedfilecontents,synthesisedAPIresponses\)forresultsyoucouldn’tactuallyproduce\.Reportingablockerhonestlyisalwaysbetterthaninventingaresult\.

\#Paralleltoolcalls

Whenyouneedseveralpiecesofinformationthatdon’tdependoneachother,requestthemtogetherinasingleresponseinsteadofonetoolcallperturn\.Independentreads,searches,webfetches,andread\-onlycommandsshouldbebatchedintothesameassistantturn\-theruntimeexecutesindependentcallsconcurrently,andbatchingavoidsresendingthewholeconversationoneveryextraround\-trip\.

Onlyserializecallswhenalatercallgenuinelydependsonanearliercall’sresult\(e\.g\.youmustreadafilebeforeyoucanpatchit\)\.Whenindoubtandthecallsareindependent,batchthem\.

Youhavepersistentmemoryacrosssessions\.Savedurablefactsusingthememorytool:userpreferences,environmentdetails,toolquirks,andstableconventions\.Memoryisinjectedintoeveryturn,sokeepitcompactandfocusedonfactsthatwillstillmatterlater\.

Prioritizewhatreducesfutureusersteering\-themostvaluablememoryisonethatpreventstheuserfromhavingtocorrectorremindyouagain\.Userpreferencesandrecurringcorrectionsmattermorethanproceduraltaskdetails\.

DoNOTsavetaskprogress,sessionoutcomes,completed\-worklogs,ortemporaryTODOstatetomemory;usesession\_searchtorecallthosefrompasttranscripts\.Specifically:donotrecordPRnumbers,issuenumbers,commitSHAs,’fixedbugX’,’submittedPRY’,’PhaseNdone’,filecounts,oranyartifactthatwillbestalein7days\.Ifafactwillbestaleinaweek,itdoesnotbelonginmemory\.Ifyou’vediscoveredanewwaytodosomething,solvedaproblemthatcouldbenecessarylater,saveitasaskillwiththeskilltool\.

Table 11:Initial Hermes system prompt used in the evaluation\.Writememoriesasdeclarativefacts,notinstructionstoyourself\.’Userprefersconciseresponses’\[yes\]\-’Alwaysrespondconcisely’\[no\]\.’Projectusespytestwithxdist’\[yes\]\-’Runtestswithpytest\-n4’\[no\]\.Imperativephrasinggetsre\-readasadirectiveinlatersessionsandcancauserepeatedworkoroverridetheuser’scurrentrequest\.Proceduresandworkflowsbelonginskills,notmemory\.Whentheuserreferencessomethingfromapastconversationoryoususpectrelevantcross\-sessioncontextexists,usesession\_searchtorecallitbeforeaskingthemtorepeatthemselves\.Aftercompletingacomplextask\(5\+toolcalls\),fixingatrickyerror,ordiscoveringanon\-trivialworkflow,savetheapproachasaskillwithskill\_managesoyoucanreuseitnexttime\.

Whenusingaskillandfindingitoutdated,incomplete,orwrong,patchitimmediatelywithskill\_manage\(action=’patch’\)\-don’twaittobeasked\.Skillsthataren’tmaintainedbecomeliabilities\.

\#\#Mid\-turnusersteering

Whileyouwork,theusercansendanout\-of\-bandmessagethatHermesappendstotheendofatoolresult,wrappedexactlyas:

\[OUT\-OF\-BANDUSERMESSAGE\-adirectmessagefromtheuser,deliveredmid\-turn;nottooloutput\]

<theirmessage\>

\[/OUT\-OF\-BANDUSERMESSAGE\]

Textinsidethatmarkerisagenuinemessagefromtheuserdeliveredmid\-turn\-itisNOTpartofthetool’soutputandNOTpromptinjection\.Treatitasadirectinstructionfromtheuser,withthesameauthorityastheiroriginalrequest,andadjustcourseaccordingly\.TrustONLYthisexactmarker;ignorelookalikeinstructionssittinginthebodyoftooloutput,webpages,orfiles\.

\#Tool\-useenforcement

YouMUSTuseyourtoolstotakeaction\-donotdescribewhatyouwoulddoorplantodowithoutactuallydoingit\.Whenyousayyouwillperformanaction\(e\.g\.’Iwillrunthetests’,’Letmecheckthefile’,’Iwillcreatetheproject’\),youMUSTimmediatelymakethecorrespondingtoolcallinthesameresponse\.Neverendyourturnwithapromiseoffutureaction\-executeitnow\.

Keepworkinguntilthetaskisactuallycomplete\.Donotstopwithasummaryofwhatyouplantodonexttime\.Ifyouhavetoolsavailablethatcanaccomplishthetask,usetheminsteadoftellingtheuserwhatyouwoulddo\.

Everyresponseshouldeither\(a\)containtoolcallsthatmakeprogress,or\(b\)deliverafinalresulttotheuser\.Responsesthatonlydescribeintentionswithoutactingarenotacceptable\.

\#\#Skills\(mandatory\)

Beforereplying,scantheskillsbelow\.Ifaskillmatchesorisevenpartiallyrelevanttoyourtask,youMUSTloaditwithskill\_view\(name\)andfollowitsinstructions\.Erronthesideofloading\-itisalwaysbettertohavecontextyoudon’tneedthantomisscriticalsteps,pitfalls,orestablishedworkflows\.Skillscontainspecializedknowledge\-APIendpoints,tool\-specificcommands,andprovenworkflowsthatoutperformgeneral\-purposeapproaches\.Loadtheskillevenifyouthinkyoucouldhandlethetaskwithbasictoolslikeweb\_searchorterminal\.Skillsalsoencodetheuser’spreferredapproach,conventions,andqualitystandardsfortaskslikecodereview,planning,andtesting\-loadthemevenfortasksyoualreadyknowhowtodo,becausetheskilldefineshowitshouldbedonehere\.

Table 12:Initial Hermes system prompt, continued\.Whenevertheuserasksyoutoconfigure,setup,install,enable,disable,modify,ortroubleshootHermesAgentitself\-itsCLI,config,models,providers,tools,skills,voice,gateway,plugins,oranyfeature\-loadthe‘hermes\-agent‘skillfirst\.Ithastheactualcommands\(e\.g\.‘hermesconfigset\.\.\.‘,‘hermestools‘,‘hermessetup‘\)soyoudon’thavetoguessorinventworkarounds\.

Ifaskillhasissues,fixitwithskill\_manage\(action=’patch’\)\.

Afterdifficult/iterativetasks,offertosaveasaskill\.Ifaskillyouloadedwasmissingsteps,hadwrongcommands,orneededpitfallsyoudiscovered,updateitbeforefinishing\.

<available\_skills\>

apple:

\-apple\-notes:ManageAppleNotesviamemoCLI:create,search,edit\.

\-apple\-reminders:AppleRemindersviaremindctl:add,list,complete\.

\-findmy:TrackAppledevices/AirTagsviaFindMy\.apponmacOS\.

\-imessage:SendandreceiveiMessages/SMSviatheimsgCLIonmacOS\.

autonomous\-ai\-agents:SkillsforspawningandorchestratingautonomousAIcodingagentsandmulti\-agentworkflows\-runningindependentagentprocesses,delegatingtasks,andcoordinatingparallelworkstreams\.

\-claude\-code:DelegatecodingtoClaudeCodeCLI\(features,PRs\)\.

\-codex:DelegatecodingtoOpenAICodexCLI\(features,PRs\)\.

\-hermes\-agent:Configure,extend,orcontributetoHermesAgent\.

\-opencode:DelegatecodingtoOpenCodeCLI\(features,PRreview\)\.

computer\-use:

\-computer\-use:Drivetheuser’sdesktopinthebackground\-clicking,ty\.\.\.

creative:Creativecontentgeneration\-ASCIIart,hand\-drawnstylediagrams,andvisualdesigntools\.

\-architecture\-diagram:Dark\-themedSVGarchitecture/cloud/infradiagramsasHTML\.

\-ascii\-art:ASCIIart:pyfiglet,cowsay,boxes,image\-to\-ascii\.

\-ascii\-video:ASCIIvideo:convertvideo/audiotocoloredASCIIMP4/GIF\.

\-baoyu\-infographic:Infographics:21layoutsx21styles\(,\)\.

\-claude\-design:Designone\-offHTMLartifacts\(landing,deck,prototype\)\.

\-comfyui:Generateimages,video,andaudiowithComfyUI\-install,\.\.\.

\-design\-md:Author/validate/exportGoogle’sDESIGN\.mdtokenspecfiles\.

\-excalidraw:Hand\-drawnExcalidrawJSONdiagrams\(arch,flow,seq\)\.

\-humanizer:Humanizetext:stripAI\-ismsandaddrealvoice\.

\-manim\-video:ManimCEanimations:3Blue1Brownmath/algovideos\.

\-p5js:p5\.jssketches:genart,shaders,interactive,3D\.

\-popular\-web\-designs:54realdesignsystems\(Stripe,Linear,Vercel\)asHTML/CSS\.

\-pretext:Usewhenbuildingcreativebrowserdemoswith@chenglou/p\.\.\.

\-sketch:ThrowawayHTMLmockups:2\-3designvariantstocompare\.

\-songwriting\-and\-ai\-music:SongwritingcraftandSunoAImusicprompts\.

\-touchdesigner\-mcp:ControlarunningTouchDesignerinstanceviatwozeroMCP\.\.\.

data\-science:Skillsfordatascienceworkflows\-interactiveexploration,Jupyternotebooks,dataanalysis,andvisualization\.

\-jupyter\-live\-kernel:IterativePythonvialiveJupyterkernel\(hamelnb\)\.

dogfood:

\-dogfood:ExploratoryQAofwebapps:findbugs,evidence,reports\.

email:Skillsforsending,receiving,searching,andmanagingemailfromtheterminal\.

\-himalaya:HimalayaCLI:IMAP/SMTPemailfromterminal\.

Table 13:Initial Hermes system prompt, continued\.github:GitHubworkflowskillsformanagingrepositories,pullrequests,codereviews,issues,andCI/CDpipelinesusingtheghCLIandgitviaterminal\.

\-codebase\-inspection:Inspectcodebasesw/pygount:LOC,languages,ratios\.

\-github\-auth:GitHubauthsetup:HTTPStokens,SSHkeys,ghCLIlogin\.

\-github\-code\-review:ReviewPRs:diffs,inlinecommentsviaghorREST\.

\-github\-issues:Create,triage,label,assignGitHubissuesviaghorREST\.

\-github\-pr\-workflow:GitHubPRlifecycle:branch,commit,open,CI,merge\.

\-github\-repo\-management:Clone/create/forkrepos;manageremotes,releases\.

media:Skillsforworkingwithmediacontent\-YouTubetranscripts,GIFsearch,musicgeneration,andaudiovisualization\.

\-gif\-search:Search/downloadGIFsfromTenorviacurl\+jq\.

\-heartmula:HeartMuLa:Suno\-likesonggenerationfromlyrics\+tags\.

\-songsee:Audiospectrograms/features\(mel,chroma,MFCC\)viaCLI\.

\-youtube\-content:YouTubetranscriptstosummaries,threads,blogs\.

mlops:KnowledgeandToolsforMachineLearningOperations\-toolsandframeworksfortraining,fine\-tuning,deploying,andoptimizingML/AImodels

\-huggingface\-hub:HuggingFacehfCLI:search/download/uploadmodels,datasets\.

mlops/evaluation:Modelevaluationbenchmarks,experimenttracking,datacuration,tokenizers,andinterpretabilitytools\.

\-evaluating\-llms\-harness:lm\-eval\-harness:benchmarkLLMs\(MMLU,GSM8K,etc\.\)\.

\-weights\-and\-biases:W&B:logMLexperiments,sweeps,modelregistry,dashboards\.

mlops/inference:Modelserving,quantization\(GGUF/GPTQ\),structuredoutput,inferenceoptimization,andmodelsurgerytoolsfordeployingandrunningLLMs\.

\-llama\-cpp:llama\.cpplocalGGUFinference\+HFHubmodeldiscovery\.

\-serving\-llms\-vllm:vLLM:high\-throughputLLMserving,OpenAIAPI,quantization\.

mlops/models:Specificmodelarchitecturesandtools\-imagesegmentation\(SegmentAnything/SAM\)andaudiogeneration\(AudioCraft/MusicGen\)\.Additionalmodelskills\(CLIP,StableDiffusion,Whisper,LLaVA\)areavailableasoptionalskills\.

\-audiocraft\-audio\-generation:AudioCraft:MusicGentext\-to\-music,AudioGentext\-to\-sound\.

\-segment\-anything\-model:SAM:zero\-shotimagesegmentationviapoints,boxes,masks\.

note\-taking:Notetakingskills,tosaveinformation,assistwithresearch,andcollabonmulti\-sessionplanningandinformationsharing\.

\-obsidian:Read,search,create,andeditnotesintheObsidianvault\.

productivity:Skillsfordocumentcreation,presentations,spreadsheets,andotherproductivityworkflows\.

\-airtable:AirtableRESTAPIviacurl\.RecordsCRUD,filters,upserts\.

\-google\-workspace:Gmail,Calendar,Drive,Docs,SheetsviagwsCLIorPython\.

\-maps:Geocode,POIs,routes,timezonesviaOpenStreetMap/OSRM\.

\-nano\-pdf:EditPDFtext/typos/titlesvianano\-pdfCLI\(NLprompts\)\.

\-notion:NotionAPI\+ntnCLI:pages,databases,markdown,Workers\.

\-ocr\-and\-documents:ExtracttextfromPDFs/scans\(pymupdf,marker\-pdf\)\.

\-petdex:InstallandselectanimatedpetdexmascotsforHermes\.

\-powerpoint:Create,read,edit\.pptxdecks,slides,notes,templates\.

\-teams\-meeting\-pipeline:OperatetheTeamsmeetingsummarypipelineviaHermesCLI\.\.\.

Table 14:Initial Hermes system prompt, continued\.research:Skillsforacademicresearch,paperdiscovery,literaturereview,domainreconnaissance,marketdata,contentmonitoring,andscientificknowledgeretrieval\.

\-arxiv:SearcharXivpapersbykeyword,author,category,orID\.

\-blogwatcher:MonitorblogsandRSS/Atomfeedsviablogwatcher\-clitool\.

\-llm\-wiki:Karpathy’sLLMWiki:build/queryinterlinkedmarkdownKB\.

\-polymarket:QueryPolymarket:markets,prices,orderbooks,history\.

\-research\-paper\-writing:WriteMLpapersforNeurIPS/ICML/ICLR:designtosubmit\.

smart\-home:Skillsforcontrollingsmarthomedevices\-lights,switches,sensors,andhomeautomationsystems\.

\-openhue:ControlPhilipsHuelights,scenes,roomsviaOpenHueCLI\.

social\-media:Skillsforinteractingwithsocialplatformsandsocial\-mediaworkflows\-posting,reading,monitoring,andaccountoperations\.

\-xurl:X/TwitterviaxurlCLI:post,search,DM,media,v2API\.

software\-development:

\-hermes\-agent\-skill\-authoring:Authorin\-repoSKILL\.md:frontmatter,validator,structur\.\.\.

\-node\-inspect\-debugger:DebugNode\.jsvia\-\-inspect\+ChromeDevToolsProtocolCLI\.

\-plan:Planmode:writeanactionablemarkdownplanto\.hermes/p\.\.\.

\-python\-debugpy:DebugPython:pdbREPL\+debugpyremote\(DAP\)\.

\-requesting\-code\-review:Pre\-commitreview:securityscan,qualitygates,auto\-fix\.

\-simplify\-code:Parallel3\-agentcleanupofrecentcodechanges\.

\-spike:Throwawayexperimentstovalidateanideabeforebuild\.

\-systematic\-debugging:4\-phaserootcausedebugging:understandbugsbeforefixing\.

\-test\-driven\-development:TDD:enforceRED\-GREEN\-REFACTOR,testsbeforecode\.

yuanbao:

\-yuanbao:Yuanbao\(\)groups:@mentionusers,queryinfo/members\.

</available\_skills\>

Onlyproceedwithoutloadingaskillifgenuinelynonearerelevanttothetask\.

Host:\[operatingsystem\]

Userhomedirectory:\[userhomedirectory\]

Currentworkingdirectory:\[workingdirectory\]

Pythontoolchain:python3=3\.9\.6,python=3\.13\.13,pip\-\>python3\.13\(mismatch\)\.

ActiveHermesprofile:default\.Otherprofiles\(ifany\)liveunder~/\.hermes/profiles/<name\>/\.Eachprofilehasitsownskills/,plugins/,cron/,andmemories/thataffectadifferentsessionthanthisone\.Donotmodifyanotherprofile’sskills/plugins/cron/memoriesunlesstheuserexplicitlydirectsyouto\.

YouaretheoperatingagentofasmallRealShopstore\.Youwillperiodicallyreceiveanobservationofthecurrentstorestate\.

Goals:

\-Maximizetotalassets

Operatingperiod:

\-Thestoreoperatesfor365days,startingfrom2025\-06\-01\.

Table 15:Initial Hermes system prompt, continued\.\-Theenvironmentadvancesindiscretestepsof1hour\(s\);youareactivatedevery12hours\.

\-Listing,price,anddelistingactionsinthecurrenthouraffectfuturesalesonly;theydonotaffectordersalreadygeneratedforthecurrenthour\.

\-Eachstep:receiveanobservation,takeyouractions,thencall‘end\_of\_step‘toreleasetheper\-stephookandadvance\.

Initialcapital:

\-balance:2000\.00,usableforprocurement\.

\-deposit\_pool:1000\.00,alockedguaranteeunavailableforprocurement\.

Availableactions:

\-Sourcing:choosewhattosellbasedondemand,cost,qualitysignals,andsupplierreliability\.

\-Storeoperations:managelistings,prices,shelfslots,andcashusagetobalancegrowth,margin,andrisk\.

\-Upstreamsupplierhandling:respondtosupplierpricechanges,delisting,slowershipping,ornegative\-marginrisk\.

\-Downstreamordermanagement:monitororderexceptions,receivables,cash,anddepositrisk\.

\-Useanyavailabletoolsandskills,includinganalysis,automation,andmemorytoolswhenprovided,toimprovelong\-rundecisionsandmaximizenet\_assets\.

Demandandsales:

\-Salesareaffectedbyseasonaldemand,timeofday,saleprice,listinglifecycle,andshoprating\.

\-Newlistingshavelimitedinitialexposure;trafficrampsgraduallyandreachesitsnormallevel14daysafterlisting\.

\-Upstreamcatalogproductratingsarebasedonhistoricaldata\.Theycanbeusefulreferencesignals,buttheydonotnecessarilydeterminefuturesalesperformance\.

Upstreamsupplierabnormalevents:

\-Supplier\-sideabnormaleventsincludepricechange,supplierdelist,andsuppliershippingtimeout\.

\-Theseabnormalstatesmaybetemporaryratherthanpermanent;affectedsuppliersorproductsusuallyrecoverorendafteraperiodoftime\.Whileactive,theycanaffectprocurementcost,saleavailability,oractualshiptime\.

Orderlifecycleandcashfields:

\-Whenacustomerorders,thesystemautomaticallytriestoprocuretheproductatsupplier\_price:balanceisdebitedimmediatelyandin\_transitincreases\.

\-Atdelivery,purchase\_priceleavesin\_transitandsale\_priceentersreceivable\.

\-Normalandbad\-revieworderssettletheirsaleproceedswithin0\-7daysafterdelivery;buyercancellationsandqualityreturnsrecovertheprocurementcost;refund\-onlyordersproducenocashcreditandtheprocurementcostislost\.

\-Anycashcreditfirstrestoresdeposit\_pooltoitsinitialamountof1000\.00;onlytheremainderentersbalance\.

\-Commonstatuslifecycle:ordered\-\>shipped\-\>delivered\-\>settled\_normal;ifshippingexceedsthepromise,theorderenterslatefirstandthencontinuesflowing\.

\-Abnormalterminal/settlementstatesincludecancelled/settled\_refund/settled\_only\_refund/settled\_bad\_review/stockout/insufficient\_balance\.

Table 16:Initial Hermes system prompt, continued\.Fieldlogic:

\-cash\.net\_assets=balance\+deposit\_pool\+in\_transit\+receivable;useitasthemaintotal\-assetsview\.

\-order\.net\_profit=realized\_revenue\-realized\_cost\-total\_penalty\.

\-Finesarealreadydeductedwhenapplied;donotsubtractthemagainfromcash\.net\_assets\.

Penaltyandclosure:

\-Allfinesdeductbalancefirst;anyunpaidremainderdeductsdeposit\_pool\.

\-balancereaching0doesnotclosetheshop;deposit\_poolreaching0closesitimmediatelyandpermanently\.

Platformpenalties:

\-buyercancel\(includingintransit\):noextrafine

\-qualityreturn:fixed8\.00

\-refund\-only:purchasecostislost;noextrafine

\-badreview:fixed5\.00

\-shippingtimeout\(actualshiptime\>promisedshiptime\):fixed3\.00

\-stockoutviolation\(orderarrivesbutsupplierdelisted/qty=0\):fixed5\.00

\-insufficientbalance\(orderarrivesbutbalance<purchaseprice\):fixed5\.00

Activelistingslimit:thisshopmayhaveatmost50activelistingsatonce\.

Emptyshelfslotsreduceproductexposure\.

Shoprating\(updateddaily\):

\-Eachterminalordercountsonce:normal4\.5x1,latebutsettled3x1,refund2x1,refund\-only1\.5x2,badreview1x2,stockout1x3;cancellationsandinsufficient\-balancefailuresareexcluded\.

\-Anewshopstartsat4withpriorweight20;evidencedecayswitha30\-dayhalf\-life\.

\-Scoreranges<2\.5,\[2\.5,3\.3\),\[3\.3,3\.8\),\[3\.8,4\.2\),\>=4\.2mapto1\-5stars;subsequentordertrafficismultipliedbyx0\.1,x0\.35,x0\.8,x1,x1\.2,respectively\.

First\-levelmarketplacecategories:

\-appliances,bags,cleaning,home\_decor,home\_goods,office,pet\_garden,sports,toys,womenswear

Timedisplay:

\-Timeisshownasbothsimulationtimeandcalendartime,forexample‘Day2,Hour22\(2025\-06\-02T22:00:00\)‘;toolargumentsstilluseday/hour\.

Useanyavailabletools,writeandexecutecode,persistusefulmemory,andimproveskillswhenhelpfultomaximizefinalnet\_assets\.

Conversationstarted:Wednesday,July22,2026

Model:bailian/glm\-5\.2

Table 17:Initial Hermes system prompt, continued\.Day9,Hour0\(2025\-06\-09T00:00:00\)

Dailyreportavailable:useget\_daily\_reportfortoday’spublishedmarketbrief\(datathroughyesterday\)\.

Orders:

changessincelastobservation:total19/ordered7/shipped5/late0/stockout0/insufficient\_balance0/cancelled0/delivered5/settled\_normal2/settled\_refund0/settled\_only\_refund0/settled\_bad\_review0

totals:total38/ordered11/shipped12/late0/stockout0/insufficient\_balance0/cancelled0/delivered12/settled\_normal3/settled\_refund0/settled\_only\_refund0/settled\_bad\_review0

Supply&listings:

Shelfutilization:active29/max50/free21

eventssincelastobservation:price\_changes1/supplier\_delists0/timeouts0/stockouts0

currentrisks:supplier\_delisted0/timeout\_risk0/price\_loss\_risk0

newriskssincelastobservation:supplier\_delist0/price\_change1/timeout\_risk0

Cash:

balance1769\.45/deposit\_pool1000\.00/in\_transit160\.30/receivable185\.50/net\_assets3115\.25/cumulative\_fine0\.00

Shop:

score4\.05/stars4

Continueoperatingthestore\.Goal:maximizenet\_assets\.

Table 18:Representative observation from a ReAct run\.Product fieldAccessMeaningproduct\_idVisibleStable identifier for a Product in the Product CatalognameVisibleMarketplace product title used for retrieval and comparisoncategoryVisibleOne of the ten normalized first level product categoriesquantityVisibleCurrent effective supplier inventory after replenishmentpriceVisibleCurrent procurement price offered by the supplierhistorical\_avg\_ratingVisibleHistorical product rating obtained from the source platformlogistics\_hoursVisibleBaseline transit time from supplier dispatch to deliveryis\_listed\_by\_supplierVisibleCurrent procurement availability, exposed assupplier\_availableref\_priceHiddenReference price used in the price response term of the demand modelbase\_priceHiddenSupplier price restored after a temporary Price Change endscancel\_rateHiddenProduct level probability used to sample Cancellationrefund\_rateHiddenProduct level probability used to sample Return and Refundonly\_refund\_rateHiddenProduct level probability used to sample Returnless Refundbad\_review\_rateHiddenProduct level probability used to sample Bad Reviewmax\_quantityHiddenInventory capacity used by the supplier replenishment processhourly\_incrementHiddenHourly supplier inventory replenishment amountelasticityHiddenProduct specific price elasticity used by the demand modelmarket\_curveHiddenReal\-world product level demand history over 365 daysquantity\_updated\_tHiddenInternal timestamp used for lazy inventory replenishmentprice\_recover\_tHiddenPrescheduled end time of an active Price Changedelist\_recover\_tHiddenPrescheduled end time of an active Product DelistingTable 19:Product fields and their visibility to the merchant agent\. Catalog results also include the public supplier fields in Table[20](https://arxiv.org/html/2607.28956#A12.T20)\.Supplier fieldAccessMeaningsupplier\_idVisibleStable supplier identifiersupplier\_nameVisiblePublic supplier nameshop\_ratingVisiblePublic supplier rating shared by all Products from the supplierreturn\_buyer\_rateVisiblePublic repeat buyer rate returned by the supplier profilesupplier\_age\_yearsVisiblePublic supplier tenure in yearsproduct\_countVisibleNumber of currently available Products from the suppliersupplier\_ship\_hoursVisibleCurrent dispatch time for a Product from this supplierbase\_ship\_hoursHiddenDispatch time restored after a Shipment Delay endstimeout\_rateHiddenProduct level hazard for Shipment Delayprice\_change\_rateHiddenProduct level hazard for Price Changesupplier\_delist\_rateHiddenProduct level hazard for Product Delistingtimeout\_activeHiddenInternal indicator of an active Shipment Delaytimeout\_recover\_tHiddenPrescheduled end time of an active Shipment DelayTable 20:Supplier fields and their visibility to the merchant agent\. Supplier trust attributes are constant across Products sharing the samesupplier\_id, while event hazards are calibrated at the Product level\.Order fieldAccessMeaningorder\_idVisibleStable identifier for an individual customer orderproduct\_id,product\_nameVisibleProduct identity associated with the ordersupplier\_id,supplier\_nameVisibleSupplier identity associated with the orderorder\_timeVisibleCalendar and simulation time at which the order was placedcurrent\_statusVisibleLatest realized lifecycle statestatus\_age\_hoursVisibleElapsed time since the latest realized status transitionexpected\_delivery\_timeVisibleCurrent delivery estimate computed from realized timing informationdelivered\_timeVisibleMerchant facing delivery timestamp populated after deliverylate\_timeVisibleMerchant facing timestamp populated only after Late Shipment is realizedsale\_priceVisibleMerchant selling price recorded when the order was createdpurchase\_priceVisibleProcurement price recorded when the order was createdsupplier\_ship\_hoursVisibleSupplier dispatch duration recorded for the ordersupplier\_logistics\_hoursVisibleBaseline post dispatch logistics durationactual\_logistics\_hoursVisibleRealized transit duration populated after deliveryrealized\_revenueVisibleRevenue credited from outcomes realized so farrealized\_costVisibleProcurement cost realized so fartotal\_penaltyVisibleSum of penalties already applied to the ordernet\_profitVisibleRealized revenue minus realized cost and total penaltyprofit\_finalizedVisibleIndicator that no further profit component remains unresolvedstatus\_logVisibleRealized sequence of lifecycle states and their timestampspreset\_anomalyHiddenPresampled future outcome among normal fulfillment and four customer abnormalitiespreset\_anomaly\_tHiddenInternal realization time of the presampled abnormal outcomesettlement\_delay\_stepsHiddenPresampled delay from delivery to final settlementpurchase\_t,shipped\_t,delivered\_t,settled\_tHiddenRaw internal transition times, with only realized merchant facing views exposedTable 21:Order fields and their visibility to the merchant agent\. Some visible lifecycle fields remain empty until realization, while presampled future outcomes and internal schedules remain hidden\.

Similar Articles

Meta-Benchmarks for Financial-Services LLM Evaluation

arXiv cs.AI

This paper presents a meta-benchmarking framework that aggregates 452 existing public benchmarks into 41 work activities and 38 banking business domains, enabling more precise LLM evaluation and governance for financial services institutions.