SimTrace:面向在线用户建模的有据多模态用户轨迹生成
摘要
SimTrace 是一个开源框架,利用基于匿名真实轨迹和模拟网络环境的 computer-use agent,生成忠实、细粒度的合成多模态用户点击流。在 8 项保真度指标中有 7 项超过基线,并且在增强真实数据时将下一动作预测效果提升了 11.0%。
arXiv:2609.38397v1 Announce Type: new
Abstract: Virtual clients offer a cost-effective approach to support applications such as A/B testing, recommender system development, and interface evaluation. However, building them requires access to large-scale, semantically faithful, fine-grained online user trajectories. These data are difficult to obtain because proprietary logs are subject to privacy restrictions and small businesses often lack sufficient traffic. Consequently, existing public datasets either abstract away fine-grained user interaction details or preserve rich context but remain platform-specific and small-scale. To address this gap, we propose SimTrace, a framework that generates faithful, fine-grained synthetic multimodal clickstreams through a computer-use client agent that is grounded in real user trajectories and the given web environment. SimTrace anonymizes real interactions and constructs a simulated twin of the given web environment, then uses both to generate synthetic interaction trajectories. Each action is paired with its corresponding web observations and user context, yielding a shareable alternative to confidential logs for developing computer-use agent-style virtual clients. We apply SimTrace to an e-commerce setting and evaluate both its fidelity and downstream utility. SimTrace outperforms competing baselines on 7 out of 8 fidelity metrics. Models trained on synthetic data achieve performance comparable to those trained on real data on downstream tasks such as purchase prediction and recommendation. For next action prediction task, augmenting real data with synthetic data further improves accuracy by 11.0% relative to training on real data alone. We release SimTrace as an open-source package to facilitate research on online user behavior modeling.
查看缓存全文
缓存时间: 2026/10/01 09:40
# SimTrace: Grounded Multimodal User Trajectories Generation for Online User Modeling
Source: [https://arxiv.org/html/2609.38397](https://arxiv.org/html/2609.38397)
###### Abstract
Virtual clients offer a cost\-effective approach to support applications such as A/B testing, recommender system development, and interface evaluation\. However, building them requires access to large\-scale, semantically faithful, fine\-grained online user trajectories\. These data are difficult to obtain because proprietary logs are subject to privacy restrictions and small businesses often lack sufficient traffic\. Consequently, existing public datasets either abstract away fine\-grained user interaction details or preserve rich context but remain platform\-specific and small\-scale\. To address this gap, we proposeSimTrace, a framework that generates faithful, fine\-grained synthetic multimodal clickstreams through a computer\-use client agent that is grounded in real user trajectories and the given web environment\.SimTraceanonymizes real interactions and constructs a simulated twin of the given web environment, then uses both to generate synthetic interaction trajectories\. Each action is paired with its corresponding web observations and user context, yielding a shareable alternative to confidential logs for developing computer\-use agent\-style virtual clients\. We applySimTraceto an e\-commerce setting and evaluate both its fidelity and downstream utility\.SimTraceoutperforms competing baselines on 7 out of 8 fidelity metrics\. Models trained on synthetic data achieve performance comparable to those trained on real data on downstream tasks such as purchase prediction and recommendation\. For next action prediction task, augmenting real data with synthetic data further improves accuracy by 11\.0% relative to training on real data alone\. We releaseSimTraceas an open\-source package to facilitate research on online user behavior modeling\.
## 1Introduction
LLMs are increasingly used to model online user behavior in web environments to support a broad range of applications, including A/B testing, recommender systems, usability studies, market research, and the evaluation of interactive LLM\-based systems\([Li et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib2);[Lu et al\., 2026a](https://arxiv.org/html/2609.38397#bib.bib1);[Chen et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib3);[Sun et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib5);[Wang et al\., 2026a](https://arxiv.org/html/2609.38397#bib.bib4)\)\. Many of these applications require models not only to predict final outcomes but also to reproduce the step\-by\-step interactions leading to those outcomes\. Recording these interactions as trajectories allows researchers and practitioners to inspect how user decisions unfold\. This provides an interpretable record of the decision\-making process\. Therefore, fine\-grained trajectories that faithfully pair each action with its corresponding web observation are essential for developing computer\-user virtual clients for these applications\.
However, existing public datasets generally either abstract away fine\-grained user interaction details or preserve rich context but remain limited in scale\. Large\-scale datasets\([Requena et al\., 2020](https://arxiv.org/html/2609.38397#bib.bib7);[Ben\-Shimon et al\., 2015](https://arxiv.org/html/2609.38397#bib.bib8);[Zhang et al\., 2015](https://arxiv.org/html/2609.38397#bib.bib9)\), commonly used for recommendation and user behavior prediction, typically represent sessions as sequences of item identifiers or event types, omitting the interface observations and semantic context needed to train computer\-use agents to reproduce users’ interactions\. Context\-rich datasets\([Wang et al\., 2026b](https://arxiv.org/html/2609.38397#bib.bib10);[Sun et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib5)\)collected through user studies, by contrast, preserve detailed interactions, but they are costly to construct, limited in scale, and typically confined to a single platform or environment, restricting their reuse in other settings\. Organizations can instead collect interaction logs from their own environments, but this requires sufficient user traffic, and the resulting data are often proprietary and constrained by privacy regulations\([Voigt and Bussche, 2017](https://arxiv.org/html/2609.38397#bib.bib6)\)\. Consequently, existing data sources rarely provide semantically faithful, environment\-grounded trajectories at scale\.
LLM\-based user simulation provides a promising solution for generating context\-rich trajectories at scale\. Existing studies infer personas and intents from user activity logs or surveys and use persona\-based LLM simulators to generate synthetic trajectories\([Sun et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib5);[Li et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib2)\)\. However, these methods typically prioritize plausible task completion\. Their generated step\-by\-step actions still differ substantially from those observed in real user sessions, as reported by[Zhang et al\. \(2026\)](https://arxiv.org/html/2609.38397#bib.bib11)\.
To address these limitations, we introduceSimTrace, a framework that uses a computer\-use client agent to generate synthetic trajectories conditioned on real user trajectories and a given environment to improve the fidelity of step\-by\-step interactions in user modeling tasks\. Through empirical evaluation, we demonstrate that synthetic data generated bySimTraceoutperform competing baselines\([Li et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib2);[Lu et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib12)\)on 7 out of 8 fidelity metrics\. Moreover, models trained on synthetic data perform comparably to those trained on real data across multiple downstream tasks\. Augmenting real data with synthetic data further improves next action prediction accuracy by 11%\. Together, these results show thatSimTracecan alleviate data scarcity and privacy constraints without compromising behavioral fidelity or downstream utility\. We releaseSimTraceas an open\-source package with modules for data collection, synthetic data generation, and user\-model training, together with a synthetic multimodal computer\-use trajectory dataset spanning five e\-commerce environments\. The package enables researchers to construct synthetic datasets for additional environments, facilitating the growth of shared resources for user behavior research\.111The code is available at[https://github\.com/agentic\-foundation\-modeling\-research/SimTrace](https://github.com/agentic-foundation-modeling-research/SimTrace)\.
## 2SimTrace
Figure 1:Overview ofSimTraceand the resulting dataset\.SimTracetransforms real activity logs into representative, anonymized behavior trajectories, constructs a twin sandbox environment of the given website, and uses LLM\-based teacher agent to generate synthetic interaction sessions\. Each resulting session contains user context and a sequence of timestamped steps, where each step pairs a structured action record with the corresponding webpage HTML and screenshot\. See Appendix[C\.2](https://arxiv.org/html/2609.38397#A3.SS2)for a complete data example\.We introduceSimTraceas a scalable framework for generating semantically faithful multimodal clickstreams to support user behavior modeling\. As illustrated in Figure[1](https://arxiv.org/html/2609.38397#S2.F1),SimTracecomprises three steps\. First, it transforms noisy and private activity logs into abstract, anonymized behavioral trajectories \(Section[2\.1](https://arxiv.org/html/2609.38397#S2.SS1)\)\. However, these abstract trajectories alone lack the fine\-grained web observations required for user behavior modeling\. To recover this interaction context in a controllable environment,SimTraceconstructs a simulated twin grounded in the given web environment using a website\-agnostic approach \(Section[2\.2](https://arxiv.org/html/2609.38397#S2.SS2)\)\. Finally, a teacher agent uses the anonymized trajectories as behavioral guidance and navigates the twin environment to generate fine\-grained interactions paired with corresponding web observations and user context \(Section[2\.3](https://arxiv.org/html/2609.38397#S2.SS3)\)\.
### 2\.1Transform Activity Logs
Activity logs record fine\-grained interactions, including clicks, keystrokes, navigation paths, and time spent on each page\([Bucklin and Sismeiro, 2009](https://arxiv.org/html/2609.38397#bib.bib15)\)\. For environments without an existing logging pipeline, our package integrates with commonly used platforms, including Microsoft Clarity\([Microsoft Clarity,](https://arxiv.org/html/2609.38397#bib.bib16)\)and PostHog\([PostHog,](https://arxiv.org/html/2609.38397#bib.bib17)\), to facilitate activity\-log collection\.
##### Select Representative Sessions\.
To obtain a compact yet representative set of user interactions, we select sessions that maximize the diversity of the given environment context and user behavior\. In the e\-commerce setting, the environment context is represented by the range of products encountered, while user behavior is captured by three conversion\-funnel outcomes\([Mittal, 2025](https://arxiv.org/html/2609.38397#bib.bib18)\): Browse\-only, Cart Abandonment, and Checkout\. These outcomes are defined as follows:*Browse\-only*means that a user views products without adding any product to the cart or proceeding to checkout;*Cart Abandonment*means that a user adds at least one product to the cart but does not proceed to checkout; and*Checkout*means that a user proceeds to checkout\.
Based on these representations, our objective is to select a subset of sessions that maximizes product coverage while maintaining a balanced distribution across the three outcomes\. To solve this problem, we use a greedy algorithm as an efficient method that directly captures our selection criteria\. The algorithm cycles through the three outcomes and selects the eligible session with the largest marginal coverage of previously unseen products\. This iterative procedure terminates when it reaches the specified sample\-size budget or a configured product\-coverage threshold, such as 95%\. Appendix[A](https://arxiv.org/html/2609.38397#A1)provides implementation details and the pseudocode\.
##### Standardize and Anonymize Records\.
To provide a common interface between heterogeneous logging tools and downstream models, we standardize raw events using a semantic\-action taxonomy\. This standardization is necessary because logging tools differ in their event schemas and levels of granularity\. Some tools record only high\-level actions that trigger page transitions, whereas others capture low\-level events such as individual mouse movements\. These tool\-specific and overly granular events can introduce noise and make activity logs difficult to compare across tools or use for behavior modeling\. We therefore design a taxonomy that maps raw events to semantic actions in two categories:*domain\-specific actions*, which capture user goals and task\-state changes specific to a particular domain, and*domain\-independent actions*, which capture navigation behavior that transfers across environments\. For our e\-commerce instantiation, we build on and extend the taxonomy introduced by[Requena et al\. \(2020\)](https://arxiv.org/html/2609.38397#bib.bib7), following the mutually exclusive and collectively exhaustive \(MECE\) principle\([Minto, Barbara, 1987](https://arxiv.org/html/2609.38397#bib.bib20)\)\. Table[1](https://arxiv.org/html/2609.38397#S2.T1)presents the resulting taxonomy\.
After standardization, we further anonymize user, store, session, and product identifiers using keyed HMAC\-SHA\-256 hashes\([National Institute of Standards and Technology, 2008](https://arxiv.org/html/2609.38397#bib.bib19)\)\. We also abstract product information into a product category and a price bucket\. Each transformed record containsstore\_id,session\_id,user\_id,timestamp,semantic\_action,product\_hash,product\_category, andprice\_bucket\.
##### Enrich Sessions\.
Standardization and anonymization operate at the record level, but session\-wide patterns can reveal latent user intent and preferences valuable for behavior modeling\. We represent this session\-level user context throughintentandpersonainferred from interaction history\. For our e\-commerce instantiation, we construct the persona following SimGym\([Li et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib2)\)\. The persona comprises five dimensions:*price sensitivity*, reflecting preferences across budget, mid\-range, and premium products;*exploration depth*, reflecting the extent of search and browsing; and*premium focus*,*performance focus*, and*ethics focus*, capturing attention to luxury, durability, and ethical sourcing, respectively\. See Appendix[C\.1](https://arxiv.org/html/2609.38397#A3.SS1)\. The final transformed log represents each session through its semantic\-action trajectory, inferred intent, and persona, as illustrated in Figure[1](https://arxiv.org/html/2609.38397#S2.F1)\.
Table 1:Taxonomy for transforming raw clickstream events in e\-commerce domain\.CategorySemantic ActionDefinitionsearchThe user enters a query in the search bar\.Domain\-detailThe user visits a product page or interacts with product attributes\.specificaddThe user adds a product to the cart\.ActionsremoveThe user removes a product from the cart\.checkoutThe user proceeds to purchase using*Buy Now*or*Checkout*\.gotoThe user navigates to a new non\-product\-detail page, e\.g\. applying a filter, changing the sort order, or following a navigation link\.Domain\-backThe user returns to the previous page\.independent ActionsstayThe user remains on the same page while a new event is recorded, typically because of a same\-page interaction, such as scrolling\.terminateThe label is appended as the final action of every session\.
### 2\.2Simulate Twin Environment
To provide web observations of corresponding actions, fine\-grained trajectories need to be generated through interactions with an environment\. To support controllable generation when direct interaction with the original environment is infeasible, such as when sensitive content must be masked, we provide a website\-agnostic approach for constructing a simulated sandbox of a given environment\.
Inspired by ShopGym\([Savadikar et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib14)\), we autonomously construct the simulated sandbox through three modules:exploration,generation, andverification\. During exploration, a powerful coding agent \(e\.g\.Piharness\([Earendil Inc\.& Contributors,](https://arxiv.org/html/2609.38397#bib.bib36)\)withGPT\-5\.5\) follows the task lists we designed to inspect the original environment and collect browser evidence, including page screenshots\. These observations are consolidated into specification capturing the visual theme, navigation structure, and supported interactions\. The generation module, which is driven by the same agent, then transforms this specification into source code for a runnable sandbox while replacing sensitive content with synthetic alternatives\. See Appendix[B](https://arxiv.org/html/2609.38397#A2)for implementation and cost details\. To ensure that the simulated environment preserves similar user experience, including user\-observed content and user actions, we introduce two verifiers: a*visual\-fidelity verifier*and an*action\-replay verifier*\. We modify the simulated environment iteratively until all the verifier feedback is resolved\.
The*visual\-fidelity verifier*compares screenshots of the generated sandbox against reference screenshots collected during exploration\. An LLM\-based evaluator scores each screenshot from 0 to 10 along five dimensions: layout, color theme, component correspondence, content density, and language consistency\. The sandbox passes visual verification only if the average score across these dimensions meets a configured threshold\. Otherwise, the verifier returns concrete discrepancies and repair instructions for the next code\-generation iteration\.
The*action\-replay verifier*evaluates whether observed user trajectories remain executable in the simulated environment\. It samples trajectories from real clickstreams and replays them in the simulated environment using Playwright browser automation\. For example, adetailaction passes if the associated product path returns a non\-error response and renders a product page, whereas anaddaction passes if the verifier can successfully activate an add\-to\-cart control\. A trajectory passes only if every action is reproduced successfully\. Each failure report identifies the corresponding trajectory, action, and reason and is supplied to the next code\-generation iteration\.
### 2\.3Generate Synthetic Data
To reproduce fine\-grained multimodal interaction trajectories based on the transformed sessions, we use a LLM\-based teacher agent to navigate the simulated environment\. This teacher agent needs to plan step\-by\-step actions while maintaining consistency with the user’s high\-level intent and persona\. It also needs to ensure that each action is valid and realistic before execution, since an incorrect action can irreversibly alter the environment state and cause the trajectory to deviate from the reference behavior\. We therefore develop a nested reflection–verification architecture: the outer loop maintains task\-level coherence across the trajectory, while the inner loop validates each proposed action before execution\. See Figure[1](https://arxiv.org/html/2609.38397#S2.F1)Section III\.
To maintain coherence over multiple interaction steps, the outer loop repeatedly plans, acts, and reflects\. Theplan moduledetermines the next step using the current web observation and guidance from prior reflections\. Theact moduletranslates this plan into a browser action with a corresponding rationale\. After execution, thereflect moduleexamines the resulting observation and provides high\-level guidance for subsequent planning\.
To prevent an incorrect action from distorting the remaining trajectory, theverify modulein the inner loop evaluates each proposal against three criteria: \(1\) executability, whether the action can be performed on the current page; \(2\) trajectory consistency, whether the action aligns with the reference trajectory; and \(3\) realism, whether a real user would plausibly take the action in the current context\. For each criterion, the verifier returns a rationale, score, and confidence value\. If an action fails verification, the action module receives the failure reason and generates a revised action\.
## 3E\-commerce Synthetic Dataset
UsingSimTrace, we generate an e\-commerce synthetic dataset across five simulated stores covering apparel, food, cookware, toys, and accessories, and four languages: English, French, Turkish, and Hindi\.222The dataset is available at[https://huggingface\.co/datasets/luyunan/SimTrace](https://huggingface.co/datasets/luyunan/SimTrace)\.Following OPeRA\([Wang et al\., 2026b](https://arxiv.org/html/2609.38397#bib.bib10)\), each session data contains three components: action records, web observations, and user context\.
Action Records\.Each session contains an ordered sequence of actions, each with an action type, a natural\-language rationale, and action\-specific metadata\. Table[3](https://arxiv.org/html/2609.38397#S3.T3)shows the action type distribution\. Each action type defines its own metadata fields\. For example, aclickaction includes atargetfield that identifies the selected interface element in the web observation, such as"search\_icon", whereas atypeaction includes aninputfield specifying the text entered by the agent\.
Web Observations\.At each step, the dataset records a web observation comprising an HTML representation of the current page and a screenshot\. Together, they provide complementary structural and visual context for interpreting the action and its rationale\. Following common practice in prior work\([Lu et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib12);[Sun et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib5);[Wang et al\., 2026b](https://arxiv.org/html/2609.38397#bib.bib10)\), we simplify the raw HTML into a semantic representation that preserves meaningful interface elements\.
User Context\.Each session is paired with an inferred persona and intent derived from the raw activity logs following SimGym\([Li et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib2)\)\. The persona summarizes the user’s behavior and preferences, while the intent specifies the session\-level goal\. See Appendix[C\.1](https://arxiv.org/html/2609.38397#A3.SS1)for an example\.
Table[2](https://arxiv.org/html/2609.38397#S3.T2)compares our dataset with existing public datasets for modeling online user behavior\. Large recommendation\-oriented datasets\([Requena et al\., 2020](https://arxiv.org/html/2609.38397#bib.bib7);[Ben\-Shimon et al\., 2015](https://arxiv.org/html/2609.38397#bib.bib8);[Zhang et al\., 2015](https://arxiv.org/html/2609.38397#bib.bib9)\)provide substantial scale but primarily record product\-level interactions with limited context\. Another line of datasets record real user interactions[Lu et al\. \(2026b\)](https://arxiv.org/html/2609.38397#bib.bib29);[Wang et al\. \(2026b\)](https://arxiv.org/html/2609.38397#bib.bib10)\. They preserve finer\-grained user interactions but are typically tied to a specific data\-collection platform and do not provide an interactive environment\. In contrast, our dataset combines fine\-grained interaction context with controllable sandbox environments across multiple platforms, and can be further expanded or customized to new environments\. Appendix[C\.2](https://arxiv.org/html/2609.38397#A3.SS2)provides a complete example session\.
DatasetSizeTasksActionSpaceMulti\-modalMulti\-platformSandboxCoveo203kPurchase Prediction6 Semantic Actions✗✗✗YOOCHOOSE9MRecommendationProduct Identifiers✗✗✗Tmall8MRecommendationProduct Identifiers✗✗✗SHOPCART31,865User Behavior Simulationclick, type, terminate✗✗✗OPeRA692All AboveRich GUI Actions✓✗✗SimTrace1,513All AboveRich GUI Actions✓✓✓Table 2:Comparison ofSimTracewith existing datasets\.*Multi\-platform*indicates whether a dataset spans multiple environment;*Sandbox*indicates whether it is paired with controllable environments for interaction\.
ActionCount\(%\)Click8,195 \(79\.40\)Terminate901 \(8\.73\)Scroll551 \(5\.34\)Type295 \(2\.86\)Back277 \(2\.68\)Select94 \(0\.91\)Clear8 \(0\.08\)Total10,321Table 3:Action distribution forSimTrace\.
## 4Experiments
We evaluateSimTracealong two complementary perspectives:fidelityanddownstream task utility\. Fidelity assesses how closely the synthetic data reproduce the characteristics of real data, whereas downstream utility assesses whether they provide effective supervision for user behavior modeling\. Specifically, our experiments address the following research questions:
- •RQ1 \- Fidelity: How faithfully doesSimTracereproduce real user behavior compared with existing methods?
- •RQ2 \- Downstream Task Utility: Can synthetic data generated bySimTraceeffectively support downstream user behavior modeling tasks?
- •RQ3 \- Low\-resource Setting: How effectively canSimTracesupport user behavior modeling in low\-resource settings, such as next\-action prediction?
### 4\.1Fidelity \(RQ1\)
##### Experimental setup\.
To evaluate the fidelity ofSimTrace, we compare it with prompting and agentic baselines\.*Persona*uses the persona representation from SimGym\([Li et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib2)\);*Trajectory*uses a reference trajectory as a one\-shot demonstration; and*Trajectory \+ Persona*combines both\. We use UXAgent\([Lu et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib12)\)as the agentic baseline\. Its planning–reflection loop conceptually corresponds to the outer loop inSimTrace, providing a practical comparison for the added benefit of action\-level verification\. For each method, we generate 200 sessions on bothgemini\-3\-flash\([Google, 2025](https://arxiv.org/html/2609.38397#bib.bib33)\)andgpt\-5\.6\-sol\([OpenAI, 2026](https://arxiv.org/html/2609.38397#bib.bib35)\)models\.
##### Evaluation metrics\.
Behavioral fidelity is multifaceted and cannot be adequately characterized by a single metric\. We therefore evaluate it at three complementary levels\. Outcome\-level fidelity uses Jensen–Shannon divergence \(JSD\) to assess whether synthetic data preserve the distribution of three session outcomes: browsing only, cart abandonment, and checkout\. Sequence\-level fidelity captures finer\-grained trajectory similarity using action\-frequency JSD, action transition matrix distance, and trajectory\-level Levenshtein distance\. In addition to these structural measures, semantic\-level fidelity assesses whether generated trajectories reflect user preference similar to that observed in real sessions\. Specifically, it evaluates product coherence, diversity, alignment and rationale alignment\. Together, these metrics capture population\-level statistical similarity and session\-level behavioral plausibility\. Appendix[D](https://arxiv.org/html/2609.38397#A4)provides detailed definitions of each metric\.
##### Analysis\.
Table[4](https://arxiv.org/html/2609.38397#S4.T4)shows thatSimTraceachieves the best fidelity on 7 out of 8 metrics, with consistent gains across all three evaluation levels\. We also test with thegpt\-5\.6\-solmodel, it largely preserves the method ranking with a slight absolute gain\. See Table[6](https://arxiv.org/html/2609.38397#A1.T6)for results\.
Beyond the quantitative results, our qualitative analysis reveals two practical issues when generating synthetic trajectories from activity logs\. The first issue, which we refer to asgranularity mismatch, occurs when activity logs omit the interface\-level steps required to reproduce an event\. For example, asearchevent may record only the query, whereas an agent must locate the search field, enter the query, and submit it through multiple actions\. The prompt\-based*Trajectory*method often advances to the next logged event before completing these intermediate steps, producing discontinuous trajectories\. Our outer reflection–planning loop addresses this issue by tracking task progress and expanding each coarse event into the required actions while treating the reference trajectory as high\-level guidance\. The second issue, which we refer to asinterface drift, occurs when actions recorded in an earlier environment are no longer directly executable in the current interface and require intermediate actions to reach the intended state\. Without action\-level verification, an agent may repeatedly attempt infeasible actions until exhausting its action budget\. Our inner verifier instead checks proposed actions before execution, identifies appropriate intermediate steps, and redirects local deviations toward the reference behavior\. Together, these mechanisms allowSimTraceto use logged trajectories as flexible behavioral guidance rather than brittle execution scripts\.
Table 4:Fidelity comparison between different generation methods ongemini\-3\-flashmodel\. The best results are shown inbold\. Statistically significant improvements \(p<0\.05p<0\.05\) over all other baselines are marked with∗\.Outcome\-levelSequence\-levelSemantic\-levelMethodOutcomeJSD↓\\downarrow\[0,1\]\[0,1\]Action Freq\.JSD↓\\downarrow\[0,1\]\[0,1\]Transition Matrix𝑳𝟏\\bm\{L\_\{1\}\}Distance↓\\downarrow\[0,1\]\[0,1\]TrajectoryLevenshtein↓\\downarrow\[0,1\]\[0,1\]ProductCoherence Gap↓\\downarrow\[0,1\]\[0,1\]ProductDiversity Ratio↓\\downarrow\[0,∞\)\[0,\\infty\)ProductAlignment↑\\uparrow\[−1,1\]\[\-1,1\]RationaleAlignment↑\\uparrow\[−1,1\]\[\-1,1\]Prompting\-based methodsPersona2\.94×10−12\.94\{\\times\}10^\{\-1\}5\.40×10−25\.40\{\\times\}10^\{\-2\}4\.30×10−14\.30\{\\times\}10^\{\-1\}6\.74×10−16\.74\{\\times\}10^\{\-1\}3\.92×10−23\.92\{\\times\}10^\{\-2\}2\.84×10−12\.84\{\\times\}10^\{\-1\}5\.75×10−15\.75\{\\times\}10^\{\-1\}6\.10×10−16\.10\{\\times\}10^\{\-1\}Trajectory8\.56×10−38\.56\{\\times\}10^\{\-3\}1\.10×10−21\.10\{\\times\}10^\{\-2\}2\.34×10−12\.34\{\\times\}10^\{\-1\}2\.18×10−12\.18\{\\times\}10^\{\-1\}6\.86×10−26\.86\{\\times\}10^\{\-2\}7\.74×𝟏𝟎−𝟐\\mathbf\{7\.74\{\\times\}10^\{\-2\}\}5\.32×10−15\.32\{\\times\}10^\{\-1\}6\.52×10−16\.52\{\\times\}10^\{\-1\}Trajectory \+ Persona3\.36×10−33\.36\{\\times\}10^\{\-3\}9\.73×10−39\.73\{\\times\}10^\{\-3\}2\.39×10−12\.39\{\\times\}10^\{\-1\}2\.85×10−12\.85\{\\times\}10^\{\-1\}6\.85×10−26\.85\{\\times\}10^\{\-2\}9\.15×10−29\.15\{\\times\}10^\{\-2\}5\.43×10−15\.43\{\\times\}10^\{\-1\}6\.54×10−16\.54\{\\times\}10^\{\-1\}Agentic methodsUXAgent1\.86×10−31\.86\{\\times\}10^\{\-3\}9\.58×10−39\.58\{\\times\}10^\{\-3\}2\.30×10−12\.30\{\\times\}10^\{\-1\}3\.50×10−13\.50\{\\times\}10^\{\-1\}3\.81×10−23\.81\{\\times\}10^\{\-2\}9\.68×10−29\.68\{\\times\}10^\{\-2\}5\.52×10−15\.52\{\\times\}10^\{\-1\}6\.59×10−16\.59\{\\times\}10^\{\-1\}SimTrace7\.98×𝟏𝟎−𝟓∗\\mathbf\{7\.98\{\\times\}10^\{\-5\}\{\}^\{\*\}\}4\.96×𝟏𝟎−𝟑∗\\mathbf\{4\.96\{\\times\}10^\{\-3\}\}^\{\*\}2\.27×𝟏𝟎−𝟏\\mathbf\{2\.27\{\\times\}10^\{\-1\}\}1\.80×𝟏𝟎−𝟏\\mathbf\{1\.80\{\\times\}10^\{\-1\}\}3\.75×𝟏𝟎−𝟐\\mathbf\{3\.75\{\\times\}10^\{\-2\}\}8\.50×10−28\.50\{\\times\}10^\{\-2\}5\.92×𝟏𝟎−𝟏\\mathbf\{5\.92\{\\times\}10^\{\-1\}\}6\.73×𝟏𝟎−𝟏\\mathbf\{6\.73\{\\times\}10^\{\-1\}\}
Table 5:Downstream performance of models trained on real \(TRTR\) and synthetic \(TSTR\) sessions and evaluated on the same held\-out real test set\.Δ\\Deltais TSTR\-TRTR\. All 95% CIs forΔ\\Deltainclude zero, indicating no statistically significant differences between two training conditions\.MethodMetricTRTRTSTRΔ\\Delta95% CIPurchase predictionSFT \(Qwen3\.5\-4B\)F1↑\\uparrow64\.3665\.52\+1\.16\+1\.16\[−3\.00,\+5\.00\]\[\-3\.00,\\,\+5\.00\]Session\-based recommendationRAIN \(Graph\-based\)MRR@5↑\\uparrow12\.4712\.58\+0\.11\+0\.11\[−1\.59,\+1\.70\]\[\-1\.59,\\,\+1\.70\]HR@5↑\\uparrow21\.2520\.85−0\.40\-0\.40\[−3\.20,\+2\.27\]\[\-3\.20,\\,\+2\.27\]MRR@10↑\\uparrow13\.2013\.64\+0\.44\+0\.44\[−1\.24,\+2\.03\]\[\-1\.24,\\,\+2\.03\]HR@10↑\\uparrow26\.7129\.12\+2\.41\+2\.41\[−0\.71,\+5\.40\]\[\-0\.71,\\,\+5\.40\]DIMO \(Context\-aware\)MRR@5↑\\uparrow17\.8416\.49−1\.35\-1\.35\[−3\.67,\+1\.16\]\[\-3\.67,\\,\+1\.16\]HR@5↑\\uparrow31\.8330\.27−1\.56\-1\.56\[−5\.05,\+2\.28\]\[\-5\.05,\\,\+2\.28\]MRR@10↑\\uparrow18\.9118\.34−0\.57\-0\.57\[−2\.85,\+2\.41\]\[\-2\.85,\\,\+2\.41\]HR@10↑\\uparrow39\.8640\.46\+0\.60\+0\.60\[−2\.83,\+3\.95\]\[\-2\.83,\\,\+3\.95\]
### 4\.2Downstream Task Utility \(RQ2\)
To assess whether the synthetic sessions generated by our pipeline provide useful training signals across multiple downstream tasks\. We evaluateSimTraceon two tasks: purchase prediction and session\-based recommendation\. For each task, we compare models trained and tested on real sessions \(TRTR\) with models trained on synthetic sessions and tested on real sessions \(TSTR\)\. Both conditions use the same held\-out real test set, isolating the effect of replacing real training trajectories with their synthetic counterparts\.
##### Purchase Prediction
Given the interaction history before a session’s terminal event, the model predicts whether the session ends in a purchase\. We evaluate this binary classification task using the F1 score\. For training, we use all available matched real–synthetic session pairs and balance the purchase and non\-purchase classes, yielding 739 examples in each training condition\. To construct the held\-out test set, following prior work\([Requena et al\., 2020](https://arxiv.org/html/2609.38397#bib.bib7)\), we exclude real sessions with lengthL<5L<5to ensure sufficient interaction history\. We then randomly sample 500 of the remaining sessions to form the final held\-out test set\.
We finetuneQwen3\.5\-4B\([Qwen Team, 2026](https://arxiv.org/html/2609.38397#bib.bib34)\)with QLoRA for five epochs with a learning rate of1×10−51\\times 10^\{\-5\}\. We report preliminary results averaged over three random seeds, together with the TSTR–TRTR difference and its 95% confidence interval\.
##### Session\-Based Recommendation
Let𝒱\\mathcal\{V\}denote the item set, and letS=\(v1,v2,…,vL\)S=\(v\_\{1\},v\_\{2\},\\ldots,v\_\{L\}\)denote a sequence of item interactions within a user session\. Given an observed sessionSS, session\-based recommendation aims to predict the next itemvL\+1∈𝒱v\_\{L\+1\}\\in\\mathcal\{V\}by ranking all candidate items and returning the top\-KKitems\. Following prior work\([Zhang et al\., 2023](https://arxiv.org/html/2609.38397#bib.bib23)\), we remove sessions containing only one interaction and items appearing fewer than five times\. Applying these filters to the matched real and synthetic data yields 357 training sessions covering 110 items\. Since ID\-based recommendation models cannot score items unseen during training, we restrict the held\-out real test set to the shared training\-item vocabulary\. The resulting test set contains 434 sessions involving 40 items\.
We apply it to two complementary recommendation methods\. RAIN\([Zeng et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib24)\)is a recent graph\-based session recommendation method with publicly available code\. DIMO\([Zhang et al\., 2024](https://arxiv.org/html/2609.38397#bib.bib25)\)is a context\-aware method that additionally incorporates product titles, vendors, product types, and images into its recommendation model\. DIMO therefore allows us to assess whether the rich multimodal content generated by our pipeline provides additional utility for downstream recommendation\. We evaluate both models using Hit Rate atKK\(HR@KK\) and Mean Reciprocal Rank atKK\(MRR@KK\), whereK∈\{5,10\}K\\in\\\{5,10\\\}\. We repeat each experiment with three random seeds and use bootstrap resampling to estimate the TSTR–TRTR differences and 95% confidence intervals\.
##### Findings
Table[5](https://arxiv.org/html/2609.38397#S4.T5)reports downstream performance for the models trained on real and synthetic sessions using the same held\-out real test set\. It shows that every 95% confidence interval includes zero, indicating no statistically significant difference between the two training conditions under our experimental setting\. These results suggest that the synthetic sessions preserve task\-relevant training signals for both purchase prediction and session\-based recommendation\.
Figure 2:Next action prediction performance ofQwen3\.5\-9Bon the OPeRA test set\. Augmenting real with synthetic data improves all metrics, with RL achieving the best overall performance\.
### 4\.3Low\-Resource Setting \(RQ3\)
For small businesses, limited traffic hinders the training of user simulation models that can perform fine\-grained navigation behaviors and enable practitioners to analyze how users interact with a website step by step\. Therefore, we assess how effectively synthetic data generated bySimTracesupports next\-action prediction in low\-resource settings\. Specifically, we sample 50 sessions from OPeRA\([Wang et al\., 2026b](https://arxiv.org/html/2609.38397#bib.bib10)\)as the real training data and compare real\-only training \(Real\-only\) with synthetic\-only \(Synth\-only\) and synthetic augmentation \(Real \+ Synth\) training\. All these trainings useQwen3\.5\-9B\([Qwen Team, 2026](https://arxiv.org/html/2609.38397#bib.bib34)\)as the base model and are evaluated on the held\-out OPeRA test set using exact match, action\-type F1, and action\-type accuracy\. A prediction counts as an exact match only when both itsactionandtargetfields exactly match the reference\.
##### SFT Training and Analysis\.
We formulate each session as a multi\-turn conversation, where each user turn contains the current web observation and each assistant turn contains a rationale followed by a structured action\. The system prompt specifies the task instruction, persona, and shopping intent\. See Appendix[E](https://arxiv.org/html/2609.38397#A5)for the complete template and Appendix[F\.1](https://arxiv.org/html/2609.38397#A6.SS1)for training details\.
Figure[2](https://arxiv.org/html/2609.38397#S4.F2)shows that SFT on synthetic data improves the base model without finetuning by 11\.9–20\.6 percentage points across the three metrics\. Training on the 50 real sessions yields stronger performance, likely because of distributional differences between synthetic and real supervision, as analyzed in Appendix[F\.2](https://arxiv.org/html/2609.38397#A6.SS2)\. Despite this gap, synthetic data provides complementary supervision: jointly training on real and synthetic sessions produces further gains of 1\.5–8\.5 percentage points over real\-only\. The improvement mainly comes from broadertargetfield coverage, which reduces overprediction of frequent values and helps the model recognize infrequent or previously unseen interface elements\. See Appendix[F\.3](https://arxiv.org/html/2609.38397#A6.SS3)for details\.
However,targetprediction remains the main bottleneck for exact\-match performance\. Target field values are long and highly sparse: 77\.4% of unique targets occur only once in OPeRA, and 14\.9% contain at least 200 characters\. Nevertheless, their hierarchical structure encodes behavioral similarity\. For example,"customers\_also\_bought\.<product1\>"and"customers\_also\_bought\.<product2\>"both indicate exploration of recommended alternatives, whereas a target beginning with"review"indicates further investigation of the current product\. Exact match treats these predictions as equally different, despite the first two reflecting similar behavioral directions\. We therefore introduce an RL objective that assigns partial credit according to hierarchical similarity oftargetfield\.
##### RL Training and Analysis\.
Following the approach of\([DeepSeek\-AI et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib21)\), we initialize the training with SFT on synthetic data to learn the dependencies among web context, action type, and target\. We then apply RL on the real sessions to align the learned policy with the behavioral patterns and target distribution of the real environment\. Each real session is converted into⟨context,action⟩\\langle\\text\{context\},\\text\{action\}\\ranglepairs, where the context contains all preceding web observations and executed actions and the prediction is the structured next action\.
Based on our observations from the SFT results, we design a reward function for thetargetfield that assigns partial credit based on the prediction\. The reward comprises three terms: a string\-similarity reward, a binary validity reward indicating whether the predicted target exists in the current UI context, and an exact\-match bonus\. Lett^\\hat\{t\}andttdenote the predicted and reference targets:
Rtarget=wsubssub\(t^,t\)\+wvalid𝕀\[t^∈𝒯\(xt\)\]\+wexact𝕀\[t^=t\],R\_\{\\mathrm\{target\}\}=w\_\{sub\}\{s\_\{\\mathrm\{sub\}\}\}\(\\hat\{t\},t\)\+w\_\{valid\}\\mathbb\{I\}\[\\hat\{t\}\\in\\mathcal\{T\}\(x\_\{t\}\)\]\+w\_\{exact\}\\mathbb\{I\}\[\\hat\{t\}=t\],wheressub\(t^,t\)s\_\{\\mathrm\{sub\}\}\(\\hat\{t\},t\)is the symmetric longest\-common\-substring ratio between predicted and ground\-truth target, and𝒯\(xt\)\\mathcal\{T\}\(x\_\{t\}\)is the set of semantic identifiers in the current web observation\.
Target accuracy alone, however, does not ensure the prediction accuracy\. Following[Zhang et al\. \(2026\)](https://arxiv.org/html/2609.38397#bib.bib11), the full reward therefore combinesRtargetR\_\{\\mathrm\{target\}\}withRactionR\_\{\\mathrm\{action\}\}, which rewards correct action\-type predictions, andRformatR\_\{\\mathrm\{format\}\}, which rewards valid structured responses\. We optimize this multi\-reward objective using Group Reward\-Decoupled Normalization Policy Optimization\([Liu et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib22)\); implementation details are provided in Appendix[F\.1](https://arxiv.org/html/2609.38397#A6.SS1)\. Relative to mixed\-data SFT, the RL stage further improves the three metrics by 1\.7%, 1\.8%, and 5\.0%, respectively\. The improvements mainly come from better output validity and target selection\. RL eliminates malformed JSON predictions and more often redirects predictions from an unrelated UI component to the correct hierarchical target path\. For example, it correctsreviews\.popover\.review\_images\.nextto the ground\-truthbuybox\.purchase\_form\.add\_to\_cart\. It also improves predictions when the correct parent path is identified, but the final target segment is wrong\. See Appendix[F\.4](https://arxiv.org/html/2609.38397#A6.SS4)for details\.
## 5Related Work
##### LLM\-Based User Simulation
Recent LLM\-based user simulation can be categorized into two paradigms: prompting LLMs and training dedicated user models\. Among prompt\-based methods, a large amount of work focuses on steering LLM behavior through personas\. PAARS\([Mansour et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib13)\)induces shopper personas from anonymized sessions, while SimGym\([Li et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib2)\)constructs buyer profiles for offline A/B testing\. Other work enriches simulators with domain knowledge\. SAGE grounds conversations in business principles\([Shea et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib26)\), whereas VISTA enables interaction through both UI and API actions\([Lu et al\., 2026a](https://arxiv.org/html/2609.38397#bib.bib1)\)\. Although flexible, prompting\-based methods rely on powerful LLMs, making long multimodal simulations costly\. Dedicated user models reduce this cost by learning human behavior into model parameters\. UserLM\([Naous et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib27)\)and Socratic\([Kong et al\., 2024](https://arxiv.org/html/2609.38397#bib.bib28)\)learn from human\-authored turns to reproduce realistic questions, while Shop\-R1\([Zhang et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib11)\)and Customer\-R1\([Wang et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib32)\)fine\-tune models to create virtual shoppers for e\-commerce domain\. However, these models remain tied to their training domains and interaction protocols\. In contrast,SimTracegenerates environment\-specific supervision for training efficient user models tailored to the given web environment\.
##### Datasets for User Behavior Modeling
User behavior modeling requires data that captures both user actions and the context in which they occur\. Existing datasets generally fall into two categories\. Large\-scale datasets contain millions of shopping or recommendation sessions but typically represent them as sequences of item identifiers or coarse event types, omitting the interface observations and semantic context needed to reconstruct user decisions\([Requena et al\., 2020](https://arxiv.org/html/2609.38397#bib.bib7);[Ben\-Shimon et al\., 2015](https://arxiv.org/html/2609.38397#bib.bib8);[Zhang et al\., 2015](https://arxiv.org/html/2609.38397#bib.bib9)\)\. Context\-rich datasets preserve semantic actions, webpage content, but are costly to collect, limited in scale, and usually tied to a single platform\([Wang et al\., 2026b](https://arxiv.org/html/2609.38397#bib.bib10);[Sun et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib5)\)\.SimTracebridges this gap by transforming private logs into standardized, anonymized trajectories and generating scalable multimodal clickstreams of the given environment\.
## 6Conclusion
We introduceSimTrace, a scalable framework for generating semantically faithful multimodal computer\-use client\-agent trajectories for user modeling\.SimTracetransforms real activity logs into standardized and anonymized trajectories, constructs a simulated twin environment grounded on the given web, and uses a teacher agent to generate behaviorally grounded interactions\. Our evaluation shows thatSimTracereproduces real user behavior more faithfully than the baselines\. Besides, models trained on synthetic data perform comparably to those trained on real data for multiple downstream tasks\. Augmenting real data with synthetic data further improves next\-action prediction accuracy by 11\.0%\. These results demonstrate the potential ofSimTraceto alleviate real\-data scarcity and reduce reliance on sensitive user logs when training downstream models\.
## 7Limitations
The proposed data generation method is designed to be domain\-adaptive, while our empirical evaluation focuses on an e\-commerce scenario\. Further validation is needed in other domains, where user semantic action taxonomy, privacy requirements may differ\. Future work should evaluateSimTracein domains such as finance, education, and healthcare\. Secondly,SimTracedoes not fully solve the cold\-start problem, as trajectory\-conditioned generation still requires initial activity logs\. Although synthetic data generated from a similar environment can provide useful supervision when target\-domain data are scarce, our results show that such data remain less effective than real target\-domain interactions for next\-action prediction\. Future work should investigate cross\-environment schema alignment and minimal\-data adaptation to bootstrap generation from related environments while reducing reliance on target\-domain logs\.
#### Acknowledgments
This research was supported by grant funding from Shopify and Toloka\. We gratefully acknowledge both organizations for their generous support and thank our collaborators for their constructive feedback and valuable discussions throughout the project, which helped shape the research direction and strengthen this work\.
## References
- Ben\-Shimonet al\.\(2015\)D\. Ben\-Shimon, A\. Tsikinovsky, M\. Friedmann, B\. Shapira, L\. Rokach, and J\. HoerleRecSys Challenge 2015 and the YOOCHOOSE Dataset\.InProceedings of the 9th ACM Conference on Recommender Systems,RecSys ’15,New York, NY, USA,pp\. 357–358\.External Links:ISBN 978\-1\-4503\-3692\-5,[Link](https://dl.acm.org/doi/10.1145/2792838.2798723),[Document](https://dx.doi.org/10.1145/2792838.2798723)Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p2.1),[§3](https://arxiv.org/html/2609.38397#S3.p5.1),[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px2.p1.1)\.
- Bucklin and Sismeiro \(2009\)R\. E\. Bucklin and C\. SismeiroClick here for internet insight: advances in clickstream data analysis in marketing\.Journal of Interactive Marketing23\(1\),pp\. 35–48\.External Links:[Document](https://dx.doi.org/10.1016/j.intmar.2008.10.004)Cited by:[§2\.1](https://arxiv.org/html/2609.38397#S2.SS1.p1.1)\.
- Chenet al\.\(2025\)L\. Chen, Q\. Dai, Z\. Zhang, X\. Feng, M\. Zhang, P\. Tang, X\. Chen, Y\. Zhu, and Z\. DongRecUserSim: A Realistic and Diverse User Simulator for Evaluating Conversational Recommender Systems\.InCompanion Proceedings of the ACM on Web Conference 2025,pp\. 133–142\.Note:arXiv:2507\.22897 \[cs\]External Links:[Link](http://arxiv.org/abs/2507.22897),[Document](https://dx.doi.org/10.1145/3701716.3715258)Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p1.1)\.
- DeepSeek\-AIet al\.\(2025\)DeepSeek\-AI, D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi, X\. Zhang, X\. Yu, Y\. Wu, Z\. F\. Wu, Z\. Gou, Z\. Shao, Z\. Li, Z\. Gao, A\. Liu, B\. Xue, B\. Wang, B\. Wu, B\. Feng, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan, D\. Dai, D\. Chen, D\. Ji, E\. Li, F\. Lin, F\. Dai, F\. Luo, G\. Hao, G\. Chen, G\. Li, H\. Zhang, H\. Bao, H\. Xu, H\. Wang, H\. Ding, H\. Xin, H\. Gao, H\. Qu, H\. Li, J\. Guo, J\. Li, J\. Wang, J\. Chen, J\. Yuan, J\. Qiu, J\. Li, J\. L\. Cai, J\. Ni, J\. Liang, J\. Chen, K\. Dong, K\. Hu, K\. Gao, K\. Guan, K\. Huang, K\. Yu, L\. Wang, L\. Zhang, L\. Zhao, L\. Wang, L\. Zhang, L\. Xu, L\. Xia, M\. Zhang, M\. Zhang, M\. Tang, M\. Li, M\. Wang, M\. Li, N\. Tian, P\. Huang, P\. Zhang, Q\. Wang, Q\. Chen, Q\. Du, R\. Ge, R\. Zhang, R\. Pan, R\. Wang, R\. J\. Chen, R\. L\. Jin, R\. Chen, S\. Lu, S\. Zhou, S\. Chen, S\. Ye, S\. Wang, S\. Yu, S\. Zhou, S\. Pan, S\. S\. Li, S\. Zhou, S\. Wu, S\. Ye, T\. Yun, T\. Pei, T\. Sun, T\. Wang, W\. Zeng, W\. Zhao, W\. Liu, W\. Liang, W\. Gao, W\. Yu, W\. Zhang, W\. L\. Xiao, W\. An, X\. Liu, X\. Wang, X\. Chen, X\. Nie, X\. Cheng, X\. Liu, X\. Xie, X\. Liu, X\. Yang, X\. Li, X\. Su, X\. Lin, X\. Q\. Li, X\. Jin, X\. Shen, X\. Chen, X\. Sun, X\. Wang, X\. Song, X\. Zhou, X\. Wang, X\. Shan, Y\. K\. Li, Y\. Q\. Wang, Y\. X\. Wei, Y\. Zhang, Y\. Xu, Y\. Li, Y\. Zhao, Y\. Sun, Y\. Wang, Y\. Yu, Y\. Zhang, Y\. Shi, Y\. Xiong, Y\. He, Y\. Piao, Y\. Wang, Y\. Tan, Y\. Ma, Y\. Liu, Y\. Guo, Y\. Ou, Y\. Wang, Y\. Gong, Y\. Zou, Y\. He, Y\. Xiong, Y\. Luo, Y\. You, Y\. Liu, Y\. Zhou, Y\. X\. Zhu, Y\. Xu, Y\. Huang, Y\. Li, Y\. Zheng, Y\. Zhu, Y\. Ma, Y\. Tang, Y\. Zha, Y\. Yan, Z\. Z\. Ren, Z\. Ren, Z\. Sha, Z\. Fu, Z\. Xu, Z\. Xie, Z\. Zhang, Z\. Hao, Z\. Ma, Z\. Yan, Z\. Wu, Z\. Gu, Z\. Zhu, Z\. Liu, Z\. Li, Z\. Xie, Z\. Song, Z\. Pan, Z\. Huang, Z\. Xu, Z\. Zhang, and Z\. ZhangDeepSeek\-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning\.Nature645\(8081\),pp\. 633–638\.Note:arXiv:2501\.12948 \[cs\.CL\]External Links:ISSN 0028\-0836, 1476\-4687,[Link](http://arxiv.org/abs/2501.12948),[Document](https://dx.doi.org/10.1038/s41586-025-09422-z)Cited by:[§4\.3](https://arxiv.org/html/2609.38397#S4.SS3.SSS0.Px2.p1.1)\.
- \[5\]Earendil Inc\.& ContributorsPi\.Note:[https://pi\.dev/](https://pi.dev/)Accessed: 2026\-07\-25Cited by:[§2\.2](https://arxiv.org/html/2609.38397#S2.SS2.p2.1)\.
- Google \(2025\)GoogleGemini 3 flash\.Note:Accessed: August 28, 2026External Links:[Link](https://deepmind.google/models/gemini/flash/)Cited by:[§C\.2](https://arxiv.org/html/2609.38397#A3.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.38397#S4.SS1.SSS0.Px1.p1.1)\.
- Konget al\.\(2024\)C\. Kong, Y\. Fan, X\. Wan, F\. Jiang, and B\. WangPlatoLM: Teaching LLMs in Multi\-Round Dialogue via a User Simulator\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 7841–7863\.External Links:[Link](https://aclanthology.org/2024.acl-long.424/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.424)Cited by:[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px1.p1.1)\.
- Liet al\.\(2026\)H\. Li, V\. Malik, Z\. Z\. Foumani, A\. Castelo, S\. Xie, A\. Fan, K\. Y\. Koay, Y\. Zhu, M\. Feghhi, R\. Uliana, Z\. Zhang, A\. O\. Martins, M\. Zhao, F\. Pelland, J\. Faerman, N\. LeBlanc, A\. Glazer, A\. McNamara, Z\. Wu, and L\. WangSimGym: A Framework for A/B Test Simulation in E\-Commerce with Traffic\-Grounded VLM Agents\.arXiv\(en\)\.Note:Version Number: 1External Links:[Link](https://arxiv.org/abs/2605.19219),[Document](https://dx.doi.org/10.48550/ARXIV.2605.19219)Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p1.1),[§1](https://arxiv.org/html/2609.38397#S1.p3.1),[§1](https://arxiv.org/html/2609.38397#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.38397#S2.SS1.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2609.38397#S3.p4.1),[§4\.1](https://arxiv.org/html/2609.38397#S4.SS1.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px1.p1.1)\.
- Lin \(1991\)J\. LinDivergence measures based on the Shannon entropy\.IEEE Transactions on Information Theory37\(1\),pp\. 145–151\.External Links:[Document](https://dx.doi.org/10.1109/18.61115)Cited by:[Appendix D](https://arxiv.org/html/2609.38397#A4.p2.1)\.
- Liuet al\.\(2026\)S\. Liu, X\. Dong, X\. Lu, S\. Diao, P\. Belcak, M\. Liu, M\. Chen, H\. Yin, Y\. F\. Wang, K\. Cheng, Y\. Choi, J\. Kautz, and P\. MolchanovGDPO: Group reward\-Decoupled Normalization Policy Optimization for Multi\-reward RL Optimization\.arXiv\.Note:arXiv:2601\.05242 \[cs\.CL\]External Links:[Link](http://arxiv.org/abs/2601.05242),[Document](https://dx.doi.org/10.48550/arXiv.2601.05242)Cited by:[§F\.1](https://arxiv.org/html/2609.38397#A6.SS1.SSS0.Px2.p1.1),[§4\.3](https://arxiv.org/html/2609.38397#S4.SS3.SSS0.Px2.p3.1)\.
- Luet al\.\(2026a\)Y\. Lu, R\. Shea, Y\. Zhang, and Z\. YuVISTA: A Versatile Interactive User Simulation Toolkit for Agent Evaluation\.arXiv\.Note:arXiv:2606\.11079 \[cs\.CL\]External Links:[Link](http://arxiv.org/abs/2606.11079),[Document](https://dx.doi.org/10.48550/arXiv.2606.11079)Cited by:[Appendix D](https://arxiv.org/html/2609.38397#A4.p4.1),[§1](https://arxiv.org/html/2609.38397#S1.p1.1),[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px1.p1.1)\.
- Luet al\.\(2026b\)Y\. Lu, J\. Huang, Y\. Han, B\. Yao, S\. Bei, J\. Gesi, Y\. Xie, Y\. Sang, Zheshen, Wang, Q\. He, and D\. WangCan LLM Agents Simulate Multi\-Turn Human Behavior? Evidence from Real Online Customer Behavior Data\.arXiv\.Note:arXiv:2503\.20749 \[cs\.CL\]External Links:[Link](http://arxiv.org/abs/2503.20749),[Document](https://dx.doi.org/10.48550/arXiv.2503.20749)Cited by:[§3](https://arxiv.org/html/2609.38397#S3.p5.1)\.
- Luet al\.\(2025\)Y\. Lu, B\. Yao, H\. Gu, J\. Huang, J\. Wang, Y\. Li, J\. Gesi, Q\. He, T\. J\. Li, and D\. WangUXAgent: an LLM agent\-based usability testing framework for web design\.InExtended Abstracts of the CHI Conference on Human Factors in Computing Systems,External Links:[Document](https://dx.doi.org/10.1145/3706599.3719729)Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p4.1),[§3](https://arxiv.org/html/2609.38397#S3.p3.1),[§4\.1](https://arxiv.org/html/2609.38397#S4.SS1.SSS0.Px1.p1.1)\.
- Mansouret al\.\(2025\)S\. Mansour, L\. Perelli, L\. Mainetti, G\. Davidson, and S\. D’AmatoPAARS: Persona Aligned Agentic Retail Shoppers\.arXiv\.Note:arXiv:2503\.24228 \[cs\]External Links:[Link](http://arxiv.org/abs/2503.24228),[Document](https://dx.doi.org/10.48550/arXiv.2503.24228)Cited by:[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px1.p1.1)\.
- \[15\]Microsoft ClarityMicrosoft clarity\.Note:[https://clarity\.microsoft\.com/](https://clarity.microsoft.com/)Accessed: 2026\-08\-16Cited by:[§2\.1](https://arxiv.org/html/2609.38397#S2.SS1.p1.1)\.
- Minto, Barbara \(1987\)Minto, BarbaraThe minto pyramid principle: logic in writing, thinking, and problem solving\.Minto International\.Note:Introduced the MECE \(Mutually Exclusive, Collectively Exhaustive\) frameworkCited by:[§2\.1](https://arxiv.org/html/2609.38397#S2.SS1.SSS0.Px2.p1.1)\.
- Mittal \(2025\)B\. MittalThe psychology of online shopping cart abandonment: a scrutiny of the current research framework and building an improved model of the online shopper journey\.Electronic Commerce Research25\(2\),pp\. 777–803\.External Links:[Document](https://dx.doi.org/10.1007/s10660-022-09667-0),[Link](https://ideas.repec.org/a/spr/elcore/v25y2025i2d10.1007_s10660-022-09667-0.html)Cited by:[§2\.1](https://arxiv.org/html/2609.38397#S2.SS1.SSS0.Px1.p1.1)\.
- Naouset al\.\(2025\)T\. Naous, P\. Laban, W\. Xu, and J\. NevilleFlipping the Dialogue: Training and Evaluating User Language Models\.arXiv\.Note:arXiv:2510\.06552 \[cs\]External Links:[Link](http://arxiv.org/abs/2510.06552),[Document](https://dx.doi.org/10.48550/arXiv.2510.06552)Cited by:[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px1.p1.1)\.
- National Institute of Standards and Technology \(2008\)National Institute of Standards and TechnologyThe keyed\-hash message authentication code \(hmac\)\.Technical reportTechnical ReportFIPS PUB 198\-1,U\.S\. Department of Commerce\.Cited by:[§2\.1](https://arxiv.org/html/2609.38397#S2.SS1.SSS0.Px2.p2.1)\.
- OpenAI \(2026\)OpenAIOpenAI\.Note:Accessed: August 28, 2026External Links:[Link](https://developers.openai.com/api/docs/models/gpt-5.6-sol)Cited by:[§4\.1](https://arxiv.org/html/2609.38397#S4.SS1.SSS0.Px1.p1.1)\.
- \[21\]PostHogPostHog\.Note:[https://posthog\.com/](https://posthog.com/)Accessed: 2026\-08\-16Cited by:[§2\.1](https://arxiv.org/html/2609.38397#S2.SS1.p1.1)\.
- Qwen Team \(2026\)Qwen TeamQwen3\.5: towards native multimodal agents\.External Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[§4\.2](https://arxiv.org/html/2609.38397#S4.SS2.SSS0.Px1.p2.1),[§4\.3](https://arxiv.org/html/2609.38397#S4.SS3.p1.1)\.
- Requenaet al\.\(2020\)B\. Requena, G\. Cassani, J\. Tagliabue, C\. Greco, and L\. LacasaShopper intent prediction from clickstream e\-commerce data with minimal browsing information\.Scientific Reports10\(1\),pp\. 16983\(en\)\.External Links:ISSN 2045\-2322,[Link](https://www.nature.com/articles/s41598-020-73622-y),[Document](https://dx.doi.org/10.1038/s41598-020-73622-y)Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.38397#S2.SS1.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.38397#S3.p5.1),[§4\.2](https://arxiv.org/html/2609.38397#S4.SS2.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px2.p1.1)\.
- Savadikaret al\.\(2026\)C\. Savadikar, M\. Zhao, Y\. Zhu, H\. Li, S\. Xie, A\. Castelo, T\. Wu, and L\. WangShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E\-Commerce Web Agents\.arXiv\.Note:arXiv:2605\.16116 \[cs\.AI\]External Links:[Link](http://arxiv.org/abs/2605.16116),[Document](https://dx.doi.org/10.48550/arXiv.2605.16116)Cited by:[Appendix B](https://arxiv.org/html/2609.38397#A2.p1.1),[§2\.2](https://arxiv.org/html/2609.38397#S2.SS2.p2.1)\.
- Sheaet al\.\(2025\)R\. Shea, Y\. Lu, L\. Qiu, and Z\. YuSAGE: A Top\-Down Bottom\-Up Knowledge\-Grounded User Simulator for Multi\-turn AGent Evaluation\.arXiv\.Note:arXiv:2510\.11997 \[cs\]External Links:[Link](http://arxiv.org/abs/2510.11997),[Document](https://dx.doi.org/10.48550/arXiv.2510.11997)Cited by:[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px1.p1.1)\.
- Sunet al\.\(2025\)L\. Sun, S\. Fu, B\. Yao, Y\. Lu, W\. Li, H\. Gu, J\. Gesi, J\. Huang, C\. Luo, and D\. WangLLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic\-AI\-based Shopping Assistants?\.arXiv\.Note:arXiv:2509\.21501 \[cs\] version: 1External Links:[Link](http://arxiv.org/abs/2509.21501),[Document](https://dx.doi.org/10.48550/arXiv.2509.21501)Cited by:[Appendix D](https://arxiv.org/html/2609.38397#A4.p5.1),[§1](https://arxiv.org/html/2609.38397#S1.p1.1),[§1](https://arxiv.org/html/2609.38397#S1.p2.1),[§1](https://arxiv.org/html/2609.38397#S1.p3.1),[§3](https://arxiv.org/html/2609.38397#S3.p3.1),[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px2.p1.1)\.
- Voigt and Bussche \(2017\)P\. Voigt and A\. v\. d\. BusscheThe EU General Data Protection Regulation \(GDPR\): A Practical Guide\.1st edition,Springer Publishing Company, Incorporated\.External Links:ISBN 978\-3\-319\-57958\-0Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p2.1)\.
- Wanget al\.\(2026a\)M\. Wang, D\. J\. Zhang, and H\. ZhangLarge Language Models for Market Research: A Data\-augmentation Approach\.arXiv\(en\)\.Note:arXiv:2412\.19363 \[cs\.AI\]External Links:[Link](http://arxiv.org/abs/2412.19363),[Document](https://dx.doi.org/10.48550/arXiv.2412.19363)Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p1.1)\.
- Wanget al\.\(2026b\)Z\. Wang, Y\. Lu, W\. Li, A\. Amini, B\. Sun, Y\. Bart, W\. Lyu, J\. Gesi, T\. Wang, J\. Huang, Y\. Su, U\. Ehsan, M\. Alikhani, T\. J\. Li, L\. Chilton, and D\. WangOPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation\.arXiv\.Note:arXiv:2506\.05606 \[cs\]External Links:[Link](http://arxiv.org/abs/2506.05606),[Document](https://dx.doi.org/10.48550/arXiv.2506.05606)Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p2.1),[§3](https://arxiv.org/html/2609.38397#S3.p1.1),[§3](https://arxiv.org/html/2609.38397#S3.p3.1),[§3](https://arxiv.org/html/2609.38397#S3.p5.1),[§4\.3](https://arxiv.org/html/2609.38397#S4.SS3.p1.1),[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px2.p1.1)\.
- Wanget al\.\(2025\)Z\. Wang, Y\. Lu, Y\. Zhang, J\. Huang, and D\. WangCustomer\-R1: Personalized Simulation of Human Behaviors via RL\-based LLM Agent in Online Shopping\.arXiv\.Note:arXiv:2510\.07230 \[cs\.CL\]External Links:[Link](http://arxiv.org/abs/2510.07230),[Document](https://dx.doi.org/10.48550/arXiv.2510.07230)Cited by:[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px1.p1.1)\.
- Xiaet al\.\(2024\)Y\. Xia, C\. Wang, J\. Mabry, and G\. ChengAdvancing Retail Data Science: Comprehensive Evaluation of Synthetic Data\.arXiv\.Note:arXiv:2406\.13130 \[cs\.LG\]External Links:[Link](http://arxiv.org/abs/2406.13130),[Document](https://dx.doi.org/10.48550/arXiv.2406.13130)Cited by:[Appendix D](https://arxiv.org/html/2609.38397#A4.p2.1)\.
- Zenget al\.\(2025\)X\. Zeng, S\. Li, Z\. Zhang, L\. Jin, Z\. Guo, and K\. WeiRAIN: reconstructed\-aware in\-context enhancement with graph denoising for session\-based recommendation\.Neural Networks184,pp\. 107056\.Cited by:[§4\.2](https://arxiv.org/html/2609.38397#S4.SS2.SSS0.Px2.p2.1)\.
- Zhanget al\.\(2023\)X\. Zhang, B\. Xu, F\. Ma, C\. Li, L\. Yang, and H\. LinBeyond Co\-occurrence: Multi\-modal Session\-based Recommendation\.arXiv\(en\)\.Note:arXiv:2309\.17037 \[cs\.IR\]External Links:[Link](http://arxiv.org/abs/2309.17037),[Document](https://dx.doi.org/10.48550/arXiv.2309.17037)Cited by:[§4\.2](https://arxiv.org/html/2609.38397#S4.SS2.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2024\)X\. Zhang, B\. Xu, Z\. Ren, X\. Wang, H\. Lin, and F\. MaDisentangling ID and Modality Effects for Session\-based Recommendation\.InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 1883–1892\.Note:arXiv:2404\.12969 \[cs\.IR\]External Links:[Link](http://arxiv.org/abs/2404.12969),[Document](https://dx.doi.org/10.1145/3626772.3657748)Cited by:[§4\.2](https://arxiv.org/html/2609.38397#S4.SS2.SSS0.Px2.p2.1)\.
- Zhanget al\.\(2026\)Y\. Zhang, T\. Wang, J\. Gesi, Z\. Wang, Y\. Lu, J\. Lin, S\. Zhan, V\. Gao, R\. Jiao, J\. Liu, K\. Qian, Y\. Tang, R\. Xue, H\. Zhang, Q\. Cui, Y\. Guo, and D\. WangShop\-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning\.arXiv\.Note:arXiv:2507\.17842 \[cs\]External Links:[Link](http://arxiv.org/abs/2507.17842),[Document](https://dx.doi.org/10.48550/arXiv.2507.17842)Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p3.1),[§4\.3](https://arxiv.org/html/2609.38397#S4.SS3.SSS0.Px2.p3.1),[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2015\)Y\. Zhang, L\. Pang, L\. Shi, and B\. WangLarge Scale Purchase Prediction with Historical User Actions on B2C Online Retail Platform\.arXiv\.Note:arXiv:1408\.6515 \[cs\.LG\]External Links:[Link](http://arxiv.org/abs/1408.6515),[Document](https://dx.doi.org/10.48550/arXiv.1408.6515)Cited by:[§1](https://arxiv.org/html/2609.38397#S1.p2.1),[§3](https://arxiv.org/html/2609.38397#S3.p5.1),[§5](https://arxiv.org/html/2609.38397#S5.SS0.SSS0.Px2.p1.1)\.
## Appendix ARepresentative Session Selection
LetSSdenote the candidate sessions\. Each sessionssis associated with an outcomeysy\_\{s\}, an interaction\-sequence lengthlsl\_\{s\}, and a set of encountered productspsp\_\{s\}\. Our objective is to select a subset of sessions that maximizes product coverage subject to two constraints: \(1\) the selected sessions are evenly distributed across the outcomes, and \(2\) thelsl\_\{s\}are larger than 3 for meaningful interactions and smaller than the empirical 95th\-percentile length constraint to avoid unusually long sessions\.
We solve this selection problem using a greedy algorithm\. We first filter the candidate sessions according to the sequence\-length constraint \(e\.g\.3≤ls≤203\\leq l\_\{s\}\\leq 20\)\. At each iteration, we consider each outcome and select the eligible session with the largest marginal coverage of previously unseen products\. The procedure terminates when it reaches the specified sample\-size budget or a configured product\-coverage threshold, such as 95%\. See Algorithm[1](https://arxiv.org/html/2609.38397#alg1)for pseudocode\.
Algorithm 1Session Selection1:Candidate sessions
SS; optional budget
BB; product coverage threshold
τ\\tau; minimum length
lminl\_\{\\min\}; maximum length
lmaxl\_\{\\max\}
2:Selected sessions
XX
3:
S′←\{s∈S:lmin≤ls≤lmax\}S^\{\\prime\}\\leftarrow\\\{s\\in S:l\_\{\\min\}\\leq l\_\{s\}\\leq l\_\{\\max\}\\\}
4:
U←⋃s∈S′PsU\\leftarrow\\bigcup\_\{s\\in S^\{\\prime\}\}P\_\{s\}⊳\\trianglerightUUis product union
5:
𝒴←\{Browser,Cart Abandoner,Checkout\}\\mathcal\{Y\}\\leftarrow\\\{\\text\{Browser\},\\text\{Cart Abandoner\},\\text\{Checkout\}\\\}
6:
X←∅,C←∅X\\leftarrow\\emptyset,\\ C\\leftarrow\\emptyset
7:while
\(B≠∅and\|X\|<B\)or\(B=∅and\|C\|/\|U\|<τ\)\(B\\neq\\emptyset\\textbf\{ and \}\|X\|<B\)\\textbf\{ or \}\(B=\\emptyset\\textbf\{ and \}\|C\|/\|U\|<\\tau\)do⊳\\trianglerightCCis the covered product list
8:for
y∈𝒴y\\in\\mathcal\{Y\}do
9:
Ay←\{s∈S′∖X:ys=y\}A\_\{y\}\\leftarrow\\\{s\\in S^\{\\prime\}\\setminus X:y\_\{s\}=y\\\}⊳\\trianglerightAyA\_\{y\}is available sessions for outcomeyy
10:if
Ay=∅A\_\{y\}=\\emptysetthen
11:return
XXwith an infeasibility warning
12:endif
13:
sy∗←argmaxs∈Ay\|Ps∖C\|s\_\{y\}^\{\*\}\\leftarrow\\operatorname\*\{arg\\,max\}\_\{s\\in A\_\{y\}\}\|P\_\{s\}\\setminus C\|
14:
X←X∪\{sy∗\}X\\leftarrow X\\cup\\\{s\_\{y\}^\{\*\}\\\}
15:
C←C∪Psy∗C\\leftarrow C\\cup P\_\{s\_\{y\}^\{\*\}\}
16:endfor
17:endwhile
18:return
XX
Table 6:Fidelity comparison between different generation methods on thegpt\-5\.6\-solmodel\. The best results are shown inbold\. Statistically significant improvements \(p<0\.05p<0\.05\) over all other baselines are marked with∗\.Outcome\-levelSequence\-levelSemantic\-levelMethodOutcomeJSD↓\\downarrow\[0,1\]\[0,1\]Action Freq\.JSD↓\\downarrow\[0,1\]\[0,1\]Transition Matrix𝑳𝟏\\bm\{L\_\{1\}\}Distance↓\\downarrow\[0,1\]\[0,1\]TrajectoryLevenshtein↓\\downarrow\[0,1\]\[0,1\]ProductCoherence Gap↓\\downarrow\[0,1\]\[0,1\]ProductDiversity Ratio↓\\downarrow\[0,∞\)\[0,\\infty\)ProductAlignment↑\\uparrow\[−1,1\]\[\-1,1\]RationaleAlignment↑\\uparrow\[−1,1\]\[\-1,1\]Prompting\-based methodsPersona1\.94×10−11\.94\{\\times\}10^\{\-1\}4\.39×10−24\.39\{\\times\}10^\{\-2\}3\.82×10−13\.82\{\\times\}10^\{\-1\}6\.08×10−16\.08\{\\times\}10^\{\-1\}6\.06×10−26\.06\{\\times\}10^\{\-2\}1\.03×10−11\.03\{\\times\}10^\{\-1\}5\.97×10−15\.97\{\\times\}10^\{\-1\}6\.47×𝟏𝟎−𝟏\\mathbf\{6\.47\{\\times\}10^\{\-1\}\}Trajectory3\.98×10−33\.98\{\\times\}10^\{\-3\}1\.65×10−21\.65\{\\times\}10^\{\-2\}2\.28×10−12\.28\{\\times\}10^\{\-1\}2\.59×10−12\.59\{\\times\}10^\{\-1\}7\.80×10−27\.80\{\\times\}10^\{\-2\}6\.45×10−26\.45\{\\times\}10^\{\-2\}5\.41×10−15\.41\{\\times\}10^\{\-1\}6\.38×10−16\.38\{\\times\}10^\{\-1\}Trajectory \+ Persona1\.20×10−31\.20\{\\times\}10^\{\-3\}2\.03×10−22\.03\{\\times\}10^\{\-2\}2\.38×10−12\.38\{\\times\}10^\{\-1\}2\.80×10−12\.80\{\\times\}10^\{\-1\}8\.00×10−28\.00\{\\times\}10^\{\-2\}7\.14×10−27\.14\{\\times\}10^\{\-2\}5\.89×10−15\.89\{\\times\}10^\{\-1\}6\.34×10−16\.34\{\\times\}10^\{\-1\}Agentic methodsUXAgent7\.69×10−47\.69\{\\times\}10^\{\-4\}1\.25×10−21\.25\{\\times\}10^\{\-2\}2\.53×10−12\.53\{\\times\}10^\{\-1\}3\.71×10−13\.71\{\\times\}10^\{\-1\}1\.00×10−11\.00\{\\times\}10^\{\-1\}2\.58×𝟏𝟎−𝟐\\mathbf\{2\.58\{\\times\}10^\{\-2\}\}5\.92×10−15\.92\{\\times\}10^\{\-1\}6\.07×10−16\.07\{\\times\}10^\{\-1\}SimTrace0\.00∗\\mathbf\{0\.00^\{\*\}\}5\.59×𝟏𝟎−𝟑\\mathbf\{5\.59\{\\times\}10^\{\-3\}\}2\.27×𝟏𝟎−𝟏\\mathbf\{2\.27\{\\times\}10^\{\-1\}\}2\.39×𝟏𝟎−𝟏\\mathbf\{2\.39\{\\times\}10^\{\-1\}\}5\.79×𝟏𝟎−𝟐\\mathbf\{5\.79\{\\times\}10^\{\-2\}\}3\.90×10−23\.90\{\\times\}10^\{\-2\}6\.18×𝟏𝟎−𝟏\\mathbf\{6\.18\{\\times\}10^\{\-1\}\}6\.41×10−16\.41\{\\times\}10^\{\-1\}
## Appendix BSimulate Twin Environment
We construct a simulated twin of each original website using an autonomous exploration and generation pipeline adapted from ShopGym\([Savadikar et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib14)\)\. The objective is to reproduce the visual organization and interaction patterns that shape the user’s navigation experience, while allowing identifying content to be replaced with synthetic alternatives due to sensitivity information\. Exploration produces specifications supported by browser evidence\. Then the generation module uses these specifications and the captured or synthetic static data to implement a runnable sandbox\. Below, we describe the exploration components, the specification representation, and the generation procedure\.
### B\.1Exploration
##### Domain\-specific interface\-area Configuration\.
To reconstruct a website faithfully, the agent must first identify both its visual structure and its interactive behavior\. We therefore formulate exploration as a configurable inspection process that specifies which interface areas the agent should examine and what evidence it should collect\. This configuration serves as a coverage checklist for the coding agent\. While some interface areas are common across domains, such as the homepage, header, and navigation, others depend on the target website\. The exploration configuration can therefore be adapted to different website domains\.
In our e\-commerce setting, we configure the following nine interface areas:
1. 1\.Homepage:the hero region and subsequent sections, including their ordering, layouts, content, and interactive elements\.
2. 2\.Header and navigation:header structure, announcement bars\.
3. 3\.Cart:its presentation as a drawer, popup, or page, together with empty and populated states, quantity changes, and item removal\.
4. 4\.Product pages:image galleries, product information, variant selectors, quantity controls, availability signals, and recommendations\.
5. 5\.Collections:product\-card layouts, filtering, sorting, and pagination\.
6. 6\.Search:the search entry point, predictive suggestions, and the presentation of results\.
7. 7\.Footer:link groups, policy links, payment indicators, and social links\.
8. 8\.Information pages:shipping, returns, privacy, terms, frequently asked questions, and contact and about pages when present\.
9. 9\.Floating interfaces:cookie banners, newsletter popups, chat widgets, and age gates, including their appearance and dismissal behavior\.
The exploration plan includes all configured areas\. When a feature is absent or an interaction cannot be exercised, the agent explicitly records this outcome rather than inferring its behavior\. The agent may also append new exploration tasks when it encounters functionality not covered by the predefined configuration\.
##### Interactive inspection\.
The coding agent executes the exploration plan using Playwright browser automation\. Rather than inspecting only static page layouts, it also exercises relevant interactions to capture state transitions\. For example, cart inspection includes adding an item, changing its quantity, and removing it, while search inspection includes entering query prefixes and observing the suggestion panel\. For each distinct interface state, the agent records a full\-page screenshot together with the corresponding accessibility snapshot\. We organize this evidence by exploration task so that each observation remains associated with the interface area, interaction, and state that produced it\.
##### Structured properties recording\.
Screenshots alone do not fully specify how an interface behaves\. We therefore complement visual evidence with structured descriptions of each observed element\. The agent records six properties:*type*,*layout*,*content*,*behavior*,*styling*, and*UX notes*\. These fields capture the element’s functional role, spatial arrangement, displayed information, response to user interaction, visual appearance, and relevant edge cases\. For example, the description of a navigation panel records its opening trigger, screen position, entry hierarchy, and dismissal behavior\. This representation makes interactive functionality and state transitions explicit while retaining screenshots as visual references\.
### B\.2Specifications
The exploration outputs are consolidated into a specification manual with three complementary components\. First, a natural\-language design manual summarizes the overall website content, structure and visual style\. Second, a structured element\-level specification records interface\-specific properties for each area defined in the exploration configuration, including layout, navigation\-menu design, filtering dimensions, pagination behavior, product\-option controls, etc\. Third, we retain prefetched static data files, such as product and collection JSON files, which provide the structured content required to populate the reconstructed interface\.
Together, these components summarize the website’s interface structure and behavior from its underlying content and provide the coding agent with sufficient information to reproduce the observed environment\.
### B\.3Generation
Given the resulting specification, the generation stage constructs a simulated twin that preserves the structure and interaction patterns of the original environment while replacing sensitive vendor\-specific content\.
##### Synthetic content generation\.
In our e\-commerce setting, we treat vendor\-specific product information as sensitive content, including product titles, descriptions, vendor identities, and images\. We therefore transform the prefetched static product data before using it in the reconstructed website\.
For each product, we first identify its original vendor information and generate a synthetic vendor identity\. We then regenerate textual product content, including titles and descriptions, conditioned on the synthetic vendor identity while preserving the product’s underlying semantic attributes\. Product images require additional care because directly transforming the original images may retain identifying or copyrighted visual content\. We therefore use a two\-stage generation pipeline\. The first model produces a textual description of the original product image\. A second image\-generation model then receives only this description and the synthetic vendor information and generates a new image without access to the original image\. This separation preserves high\-level product semantics while reducing direct visual reproduction of the source content\.
##### Stepwise source\-code generation\.
Generating the entire website in a single step makes it difficult to maintain consistency and localize errors\. Instead, the coding agent implements the configured interface areas incrementally\. For each generation task, it receives the relevant portion of the design manual, the associated structured specification and static data, and the current source code\. It then implements the corresponding layouts, content, and interactions\.
The generation plan, source code, and intermediate artifacts persist across iterations\. Consequently, later tasks can build directly on previously implemented components, while errors can be corrected locally without regenerating the complete website\. The final application, together with its synthetic content and supporting data, constitutes the simulated twin used for downstream user simulation\.
### B\.4Cost
Constructing a simulated twin requires approximately 28–30M tokens and 2\.9–4\.2 hours end\-to\-end using the Pi harness with GPT\-5\.5, with an estimated API cost of $60–64 per website\. Token usage includes cached input, uncached input, and output tokens\.
## Appendix CDataset
### C\.1Persona Example
Persona Example``` { "behavioral": { "exploration_depth": 0.45, "price_sensitivity": "mid-range" }, "values": { "ethics": 0.1, "performance": 0.5, "premium": 0.2 }, "reasoning": { "ethics": "No indicators of environmental or ethical considerations (e.g., non-toxic plastics, sustainable packaging, or local manufacturing) were found in the product data. Score: 0.1", "exploration_depth": "The user spent a significant amount of time (avg 347.5s) relative to the number of products viewed (2), suggesting they were carefully reading descriptions rather than browsing broadly. Score: 0.45", "performance": "100% of browsed products focus on the tactile performance and functional outcomes (’slow-rise experience’, ’groovy tactile experience’, ’stress relief’). Focus is on how the product functions as a sensory tool. Score: 0.5 (max for browsing only)", "premium": "Description mentions ’velvety texture’ and ’perfected the formula’, signaling a preference for established brand quality over generic toys, though it stops short of luxury or exclusive keywords. Score: 0.2", "price_sensitivity": "Browsed price bucket is standard for name-brand sensory fidget toys. There is no evidence of extreme budget-seeking or high-end collector-grade pricing. Classification: ’mid-range’" }, "confidence": { "behavioral": 0.4, "values": 0.2 } } ```
### C\.2Dataset Details
We usegemini\-3\-flash\([Google, 2025](https://arxiv.org/html/2609.38397#bib.bib33)\)to generate the e\-commerce synthetic dataset\. We release the dataset under the Creative Commons Attribution\-NonCommercial 4\.0 International license \(CC BY\-NC 4\.0\)\. The following example illustrates the structure of a complete session in the generated dataset\. For readability, thesimplified\_domandpersonafields are truncated\.
\{
"store\_id":"68e1563a\-676a\-694b\-ce31\-820f79080bad",
"session\_id":"d79f58cc\-30fc\-046f\-bc52\-38014c1642a7",
"user\_id":"4c5381ad\-ab85\-6e0d\-ca9f\-71ce393443f1",
"timestamp":"2026\-07\-01T21:56:48\.861744",
"action\_type":"click",
"target":"needoh",
"rationale":"I’llstartbycheckingouttheNeeDohbrandcollection,asI’mlookingfor
high\-qualitysensorytoysformycollection\.",
"input\_text":null,
"simplified\_dom":"<html\><bodyparser\-is\-focused=\\"true\\"\>
\.\.\.
<ahref=\\"/collections/novelty\-toys\\"
parser\-semantic\-id=\\"needoh\\"
parser\-clickable=\\"true\\"\>
<span\>NeeDoh</span\>
</a\>
\.\.\.
</body\></html\>",
"action\_json":\{
"action":"click",
"description":"ClickingontheNeeDohbrandcollection\.",
"rationale":"I’llstartbycheckingouttheNeeDohbrand
collection,asI’mlookingforhigh\-qualitysensory
toysformycollection\.",
"target":"needoh",
"url":"/collections/novelty\-toys"
\},
"image":"41fd9aa7f6e8c0e5f260c3da\.png",
"intent":"I’mlookingforSensoryToys,greetingcards,sensorytoys,StuffedAnimals\."
"persona":\{
"behavioral":\{
"exploration\_depth":0\.45,
"price\_sensitivity":"mid\-range"
\},
"values":\{
"ethics":0\.1,
"performance":0\.5,
"premium":0\.2
\},
"reasoning":\{\.\.\.\},
"confidence":\{\.\.\.\}
\}
\}
## Appendix DFidelity Metrics
This section defines the eight fidelity metrics reported in Table[4](https://arxiv.org/html/2609.38397#S4.T4)and Table[6](https://arxiv.org/html/2609.38397#A1.T6)\. LetAAdenote the semantic action set defined in Table[1](https://arxiv.org/html/2609.38397#S2.T1), with\|A\|=9\|A\|=9\. This standardized taxonomy abstracts away website\-specific actions and enables the calculation of distributional metrics, such as Jensen–Shannon divergence \(JSD\) and Transition Matrix L1 Distance\.
Outcome JSDmeasures the divergence between the session outcome distributions of real and synthetic data\. We categorize each session as*Browser*,*Cart Abandoner*, or*Checkout*\. Jensen–Shannon divergence \(JSD\)\([Lin, 1991](https://arxiv.org/html/2609.38397#bib.bib30)\)is widely used to evaluate user simulators by quantifying the discrepancy between simulated and real user behavior\([Xia et al\., 2024](https://arxiv.org/html/2609.38397#bib.bib31)\)\. Using base\-2 logarithms, the score ranges from 0 to 1, with lower values indicating more similar outcome distributions\.
Action Frequency JSDmeasures the divergence between the marginal action distributions of real and synthetic data\. We pool all actions in each corpus into a normalized frequency vector overAAand compute the JSD between the two vectors\. Unlike Outcome JSD, which considers only terminal outcomes, this metric captures how frequently each action occurs throughout a session\. For example, a generator may reproduce the real checkout rate while omitting thesearchanddetailactions that typically precede a purchase\. The score ranges from 0 to 1, with lower values indicating more similar action distributions\.
Transition MatrixL1L\_\{1\}Distancemeasures the similarity of local action\-to\-action dynamics following prior work\([Lu et al\., 2026a](https://arxiv.org/html/2609.38397#bib.bib1)\)\. For each corpus, we estimate a row\-normalized first\-order transition matrixP∈\[0,1\]\|A\|×\|A\|P\\in\[0,1\]^\{\|A\|\\times\|A\|\}, wherePij=Pr\(at\+1=j∣at=i\)P\_\{ij\}=\\Pr\(a\_\{t\+1\}=j\\mid a\_\{t\}=i\)\. We compute the normalizedL1L\_\{1\}distance as1\|A\|∑i12∑j\|Pijreal−Pijsyn\|\\frac\{1\}\{\|A\|\}\\sum\_\{i\}\\frac\{1\}\{2\}\\sum\_\{j\}\\lvert P^\{\\mathrm\{real\}\}\_\{ij\}\-P^\{\\mathrm\{syn\}\}\_\{ij\}\\rvert\. This metric captures sequential structure that marginal action frequencies cannot distinguish\. For example, replacing everydetail→\\rightarrowaddtransition withadd→\\rightarrowdetailpreserves action frequencies but increases the transition distance\. The score ranges from 0 to 1, with lower values indicating more similar transition patterns\.
Trajectory Levenshteinmeasures the normalized Levenshtein edit distance between the semantic action sequences of real and synthetic sessions, following prior work\([Sun et al\., 2025](https://arxiv.org/html/2609.38397#bib.bib5)\)\. A score of 0 indicates that the agent replayed its reference trajectory exactly\.
Product Coherence Gapmeasures whether synthetic sessions maintain the same degree of topical consistency as real sessions\. For each session containing at least two products, we compute the mean pairwise cosine similarity between the embeddings of the viewed products\. We rescale the value to\[0,1\]\[0,1\]\. We then report the absolute difference between the mean coherence scores of the real and synthetic corpora\. The lower values indicating more similar within\-session coherence\. A large gap in either direction suggests that the agent either browses across unrelated product categories or focuses more narrowly than real users\.
Product Diversity Ratiomeasures the relative difference in overall product coverage between the real and synthetic corpora\. Letnrealn^\{\\mathrm\{real\}\}andnsynn^\{\\mathrm\{syn\}\}denote the numbers of unique products browsed across all real and synthetic sessions, respectively\. We compute\|nsyn/nreal−1\|\\lvert n^\{\\mathrm\{syn\}\}/n^\{\\mathrm\{real\}\}\-1\\rvert\. A score of 0 indicates that the synthetic corpus covers the same number of unique products as the real corpus, with lower values indicating more similar product diversity\. Although this metric is unbounded in principle, the synthetic corpora in our experiments consistently cover fewer products than the real data\. This result highlights the need to increase product diversity in future synthetic data generation\.
Product Alignmentmeasures whether the real and synthetic sessions involve semantically similar products\. For each product in a synthetic session, we retrieve the most similar product from the matched real session using inner\-product search overL2L\_\{2\}\-normalized embeddings and average the resulting values within each session across all matched pairs\. The score ranges from \-1 to 1, with higher values indicating stronger product alignment\.
Rationale Alignmentmeasures whether matched real and synthetic sessions reflect similar underlying shopping intents\. Given the action sequence and product context of each session, an LLM generates a natural\-language description of the inferred intent\. We embed these descriptions and report the mean cosine similarity between the descriptions of each matched real–synthetic pair\. This metric captures semantic agreement despite differences in specific actions or products; for example, two sessions may follow different trajectories while both reflecting comparison shopping for a budget\-friendly gift\.
## Appendix EUser Model System Prompt
The template below is the system message used for training and evaluation\. The\{persona\}and\{intent\}fields are filled per session\. The user turn supplies the current observation as\# contextfollowed by the HTML of the page, and assistant turn is the single JSON object the template requires\.
User Model System Prompt Template``` You pretend to be a user browsing the store website and do shopping based on your intent. Your task is to predict the next action and provide rationale for the action based on your persona, intent, previous actions and context. The history action (with details described below), rationale, context and the user persona will be provided to you. # Action Space Each action object must include an ‘action‘ key specifying one of the following types. Include required fields exactly as shown. ## Click { "action": "click", "target": "<element_semantic_id>", "description": "Clicking ..." } ## Type (with optional Enter submit) { "action": "type", "target": "<input_semantic_id>", "text": "<text>", "enter": true, "description": "Typing and submitting ..." } ## Select (e.g., dropdowns) { "action": "select", "target": "<select_semantic_id>", "value": "<option_value>", "description": "Selecting ..." } ## Clear (clear an input field) { "action": "clear", "target": "<input_semantic_id>", "description": "Clearing ..." } ## Scroll (scroll the chat window or current page) { "action": "scroll", "target": "<optional_element_semantic_id>", "direction": "up", "amount": 300, "description": "Scrolling up ..." } { "action": "scroll", "target": "<optional_element_semantic_id>", "direction": "down", "amount": 300, "description": "Scrolling down ..." } ## Navigation (use when the chatbot provides a URL to open) { "action": "goto_url", "url": "https://example.com", "description": "Navigating ..." } { "action": "back", "description": "Going back ..." } { "action": "forward", "description": "Going forward ..." } { "action": "refresh", "description": "Refreshing ..." } ## Terminate (only if explicitly instructed by the step) { "action": "terminate", "description": "Terminating ..." } # Rationale The rationale is a first-person sentence (<=25 words) explaining why you are taking the action. Do not mention HTML, tag ids, or the simulation. # Context Your context will be a HTML of the webpage you are looking at. # Persona The user persona reflects the user’s price sensitivity, exploration and preference. Here is your persona: {persona} # Intent Here is your intent: {intent} # Output Format You need to predict the next action and provide rationale for the action. Your output should be a single, flat JSON object with ‘rationale‘, ‘action‘, and the type-specific fields shown above. For example: {"rationale": "I want to search for a necklace.", "action": "type", "target": "search", "text": "necklace", "enter": true, "description": "Typing and submitting necklace in the search bar."} <IMPORTANT> OUTPUT A SINGLE JSON OBJECT, NOTHING ELSE. </IMPORTANT> ```
## Appendix FDiscussion for Next\-Action Prediction
### F\.1Training Details
We useQwen3\.5\-9Bas the base model for all next\-action prediction experiments\. Both SFT and RL use full\-parameter fine\-tuning with BF16 mixed precision and gradient checkpointing\. All training is conducted on NVIDIA H200 GPUs\.
##### SFT\.
We formulate each session as a multi\-turn conversation, as described in Section[4\.3](https://arxiv.org/html/2609.38397#S4.SS3)\. To accommodate long web observations and interaction trajectories, we segment each session using a 15\-turn sliding window and truncate each example to 32K tokens\. We fine\-tune the model for up to three epochs with a learning rate of2×10−52\\times 10^\{\-5\}\.
##### RL\.
We initialize RL training from the model obtained through SFT on synthetic data and optimize it using Group Reward\-Decoupled Normalization Policy Optimization \(GDPO\)\([Liu et al\., 2026](https://arxiv.org/html/2609.38397#bib.bib22)\)\. We uses eight GPUs in total: four for full\-parameter policy optimization with DeepSpeed ZeRO\-3 and four for response generation with vLLM using tensor parallelism\. We use a group size of 4, a global batch size of 32, and a learning rate of1×10−61\\times 10^\{\-6\}\. Based on empirical tuning, we setwsub=0\.6w\_\{\\mathrm\{sub\}\}=0\.6,wvalid=0\.1w\_\{\\mathrm\{valid\}\}=0\.1, andwexact=0\.3w\_\{\\mathrm\{exact\}\}=0\.3forRtargetR\_\{\\mathrm\{target\}\}\. The RL stage takes approximately 24 hours per run\.
### F\.2Performance Gap Between Synthetic and Real Supervision
Figure[2](https://arxiv.org/html/2609.38397#S4.F2)shows that synthetic supervision fromSimTraceimproves the base model substantially but still trails supervision from real data\. To understand this gap, we examined predictions from the synthetic\-only model on the OPeRA test set\. The residual gap concentrates in three mismatches between the source platform used for synthetic data generation and the target platform used for evaluation\.
Interaction granularity\.The two platforms record user behavior at different levels of granularity\. The source logs primarily capture actions that trigger page transitions, whereas the target platform also records interactions that modify the current view without navigating to a new page\. Examples include cycling through a product image carousel, expanding a review panel, and scrolling through a long product page\. Because these fine\-grained interactions are absent from the synthetic corpus, the model receives no supervision for predicting them\.
Target vocabulary\.The platforms expose different interface elements, leaving several target\-domain identifiers without source\-domain counterparts\. For example, targets rooted atcustomers\_also\_boughtorreviewcorrespond to recommendation and review modules that rarely appear in the source data\. This mismatch further compounds the target sparsity discussed in Section[4\.3](https://arxiv.org/html/2609.38397#S4.SS3), where77\.4%77\.4\\%of unique targets occur only once\. By assigning partial credit based on target similarity, RL mitigates this problem but cannot fully recover targets that receive little or no supervision\.
Operation semantics\.Visually similar elements may also behave differently across platforms\. For example, a search button opens the search interface on the source platform but submits a typed query on the target platform\. The correct interaction order is therefore click\-then\-type in the source environment but type\-then\-click in the target environment\. A model trained only on source\-domain data learns the former ordering and transfers it to the target platform, producing action sequences that the target interface cannot execute\.
The gap may narrow when synthetic interactions are generated within a simulated twin of the target interface, or within the original interface when direct agent interaction is feasible and does not expose sensitive information\. Such settings would reduce the interface and functionality mismatches identified above\. In our experiments, however, the original activity logs do not contain the web observations associated with each action and therefore cannot directly supervise next\-action prediction\.SimTraceaddresses this limitation by using a powerful teacher agent to replay the logged behavioral trajectories and reconstruct multimodal web observation–action pairs\. From this perspective,SimTraceserves not only as an anonymized synthetic data generation framework but also as an effective data enrichment method that transforms otherwise incomplete activity logs into training examples suitable for downstream user\-modeling tasks\.
### F\.3Benefits of Synthetic Data
Our error analysis indicates that more than half of the exact\-match improvement occurs in cases where synthetic data expands the model’s effective coverage of the target space\. With limited real training data, the model tends to overpredict a small set of frequent targets rather than distinguish among less common interface elements\. Synthetic data exposes the model to a broader range of interactions, including suggested\-item exploration and product\-option selection, thereby improving its ability to identify diverse targets\.
Search suggestions provide a representative example\. Among the 40 test cases whose ground\-truth target was a specific suggested term \(e\.g\.,nav\_bar\.suggested\_terms\.bike\_bottle\_hold\), the model trained only on real data failed to predict any target correctly\. Instead, it typically predictedsearch\_inputorsearch\_button, two frequent targets that together account for 24\.03% of the real training data\. After adding synthetic data, the model correctly identified 22 of the 40 suggested terms, achieving 55\.00% accuracy on these cases, while preserving its accuracy onsearch\_input\.
We observe a similar pattern for product\-option attributes\. The test set contains four option types: color, scent, size, and style\. Because scent does not appear in the real training data, the model trained only on real data never predicts it\. Although the synthetic data also contains no scent examples, it introduces related attributes ”flavor”\. After training with these examples, the model correctly predicts scent in some test cases, likely because flavor and scent are semantically related and occupy similar structural roles in the target schema\. This result suggests that synthetic data can improve not only direct target coverage but also generalization to structurally and semantically related targets\.
### F\.4Benefits of Reinforcement Learning
RL improves both output validity and target selection\. Under the RL objective, which includes a format reward, the rate of malformed JSON predictions decreases from 1\.99% to 0%\. The target reward also assigns partial credit based on structural similarity, encouraging the model to identify the correct interface region even when it does not recover the exact target\.
Our error analysis shows that more than half of the exact\-match improvement comes from cases in which RL redirects a prediction from an unrelated top\-level UI component to the correct region\. For example, given the ground\-truth targetbuybox\.purchase\_form\.add\_to\_cart, the SFT model predicts the unrelated targetreviews\.popover\.review\_images\.next, whereas the RL\-trained model recovers the exact target\. The remaining improvements primarily occur when both models identify the correct parent path but RL corrects the final leaf segment\. Consistent with this pattern, after removing the final leaf segment from each target, the RL\-trained model achieves 44\.26% parent\-path accuracy, compared with 40\.71% for the SFT model, an improvement of 3\.55 percentage points\.相似文章
超越表面风格:基于行为一致性的多轮用户模拟器对齐研究
本文提出TRACER多轮用户模拟器,该模拟器通过强化学习将模拟行为与真实用户轨迹对齐,并引入动态营销基准测试以评估大语言模型在说服力和响应质量方面的表现。
RealUserSim:通过真实用户模拟弥合智能体基准测试中的现实差距
本文介绍了RealUserSim,一个将基于LLM的用户模拟扎根于来自14,000+真实对话的人类行为数据中的框架,旨在弥合智能体基准测试中的现实差距。研究表明,基于真实数据的模拟将行为匹配率从24.2%提升至45.3%,并揭示了协作型模拟器无法发现的失效机制。
信号:用于智能体交互的轨迹采样与分流
本文提出了一种轻量级的、基于信号的框架,通过计算低成本指标来高效分流智能体交互轨迹,这些指标能识别出信息丰富的样本,且不影响在线智能体行为,在基准测试上实现了82%的信息率。
TrajGenAgent:一种用于人类移动轨迹生成的分层LLM智能体
TrajGenAgent提出了一种分层LLM智能体框架,将宏观活动规划与微观时空实例化解耦,用于无需微调即可生成逼真的人类移动轨迹。它还引入了一种基于异常检测的评估方法,用于行为保真度。
使用Grounded Theory进行大规模智能体行为分析
AutoTraceGT自动化了在智能体轨迹上进行的扎根理论分析,以构建行为分类体系,有效恢复人类失败模式注释并改进下游预测任务。