超越表面风格:基于行为一致性的多轮用户模拟器对齐研究

arXiv cs.AI 论文

摘要

本文提出TRACER多轮用户模拟器,该模拟器通过强化学习将模拟行为与真实用户轨迹对齐,并引入动态营销基准测试以评估大语言模型在说服力和响应质量方面的表现。

arXiv:2609.28690v1 Announce Type: new Abstract: Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.
查看原文
查看缓存全文

缓存时间: 2026/09/25 09:30

# Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency
Source: [https://arxiv.org/html/2609.28690](https://arxiv.org/html/2609.28690)
###### Abstract

Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale\. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions\. We propose TRACER, a multi\-turn user simulator that explicitly models users’ evolving intent and learns to align simulated behavior with real interaction trajectories\. TRACER is trained in two stages: supervised fine\-tuning on real user dialogues, followed by multi\-turn reinforcement learning\. The RL stage combines hierarchical outcome\- and trajectory\-level rewards with deviation\-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues\. On real customer\-service sessions organized into reference cohorts, TRACER\-7B surpasses the strongest baseline by 11\.4 conversion F1, while also achieving the lowest group\-level conversion\-rate error and semantic trajectory distance, and generalizing to out\-of\-distribution scenarios\. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations\. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates\.

22footnotetext:These authors contributed equally to this work\.11footnotetext:Corresponding author\. E\-mail:[wuxiangyu06@kuaishou\.com](mailto:[email protected])\.## 1Introduction

With the widespread application of large language models \(LLMs\) in dialogue systems, scalable user feedback has become increasingly important for model evaluation and improvement\. User simulators provide such feedback through interactions with dialogue agents\([Hu et al\., 2023](https://arxiv.org/html/2609.28690#bib.bib3);[Abbasiantaeb et al\., 2024](https://arxiv.org/html/2609.28690#bib.bib1);[Qiu & Lan, 2024](https://arxiv.org/html/2609.28690#bib.bib15);[Dou et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib2)\), supporting offline evaluation and reinforcement learning\([Wu et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib25);[Wang et al\., 2025b](https://arxiv.org/html/2609.28690#bib.bib23);[Qian et al\., 2025b](https://arxiv.org/html/2609.28690#bib.bib14)\)\. Their usefulness, however, depends on both natural responses and faithful user decisions\. Multi\-turn simulation must therefore account for how users express themselves and how their behavior unfolds over complete interactions\.

However, existing user simulators still suffer from two fundamental limitations\.*\(L1\) Alignment is limited to linguistic style, not decision\-making\.*Mainstream methods condition LLMs on static user profiles to elicit persona\-consistent responses\([Wang et al\., 2025a](https://arxiv.org/html/2609.28690#bib.bib22);[Shi et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib21);[Ye et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib29);[Kim & Yang, 2025](https://arxiv.org/html/2609.28690#bib.bib4)\), aligning simulators with real users only at the surface linguistic level while leaving the underlying decision\-making process unaligned, although real multi\-turn behavior is jointly driven by interaction history, context, and decision states\.*\(L2\) Optimization is single\-turn, not trajectory\-level\.*Recent methods adopt optimization paradigms tailored to turn\-level responses\([Wang et al\., 2025c](https://arxiv.org/html/2609.28690#bib.bib24);[Liao et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib7);[Zhang et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib31);[Wu et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib26);[Dou et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib2)\), and thus inherently struggle with reward sparsity and credit assignment in multi\-turn interactions, failing to capture state evolution and decision coherence across turns\. Consequently, simulated feedback often exhibits behavioral instability and inconsistent decision\-making, limiting its utility for high\-fidelity evaluation and training\.

To address these limitations, we proposeTRajectory\-AlignedCredit Assignment for UsERSimulation \(TRACER\), a multi\-turn user simulator trained on real dialogues\. TRACER combines linguistic grounding with behavioral alignment, using intent states inferred from observed dialogue as proxies for users’ decision tendencies\. In the first stage, supervised fine\-tuning learns user\-like responses and initializes structured intent prediction\. In the second stage, the simulator interacts with a fixed assistant and learns from rewards measuring terminal agreement and trajectory alignment against paired human logs\. Dynamic time warping accommodates differences in conversational pace\. The same alignment also supplies local discrepancies, which are accumulated over discounted trajectory suffixes to modulate turn\-level advantages\. This links session evaluation to policy optimization: the session advantage determines the reinforcement direction, while alignment deviations adjust its strength within the interaction\.

To evaluate behavioral fidelity, we organize3,866 real customer\-service sessions into 762 reference groupsbased on similar user conditions\. Each group provides multiple observed trajectories, allowing evaluation to account for variation among comparable users and reducing dependence on a single reference interaction\. We compare simulated and real behavior within these groups along three dimensions: session\-level behavior, dialogue structure, and semantic consistency\. This evaluation complements the paired\-reference training objective by examining behavior across multiple real sessions\. We also assess generalization to unseen scenarios and perceived human\-likeness\. In a human Turing study, annotator identification accuracy for TRACER conversations was close to chance, supporting their perceived naturalness under the study conditions\.

Building on TRACER, we further introduce the Dynamic Marketing Benchmark \(DM\-Bench\), which evaluates LLMs as conversational marketing agents through interactions with simulated users\. DM\-Bench jointly measures response quality and conversion outcomes, providing a behavioral perspective on agent performance\. Within this simulated environment, models with higher response quality do not necessarily achieve higher conversion rates, motivating evaluation of both dimensions\.

We summarize our contributions as follows: 1\) We propose TRACER, a two\-stage user simulator that combines outcome and trajectory rewards with deviation\-aware turn\-level advantage modulation, which uses the magnitude and persistence of trajectory deviations to adjust per\-turn updates while preserving session\-level advantage direction\. 2\) We construct a multi\-reference evaluation set with 762 groups and 3,866 real sessions, supporting assessment of session behavior, dialogue structure, and semantic consistency\. 3\) We develop DM\-Bench to jointly evaluate response quality and conversion outcomes in simulated marketing interactions, revealing a misalignment between the two\.

## 2Related Work

##### LLM\-based User Simulator\.

Recent LLM\-based user simulators have evolved from prompt\-based role\-playing to deeper cognitive modeling\. Early works often rely on explicit persona descriptions or static intent predictions, capturing surface\-level linguistic patterns or individual query responses but struggling to capture the dynamic evolution of user intents across multi\-turn interactions\([Park et al\., 2024](https://arxiv.org/html/2609.28690#bib.bib12);[Wang et al\., 2025a](https://arxiv.org/html/2609.28690#bib.bib22);[Kim & Yang, 2025](https://arxiv.org/html/2609.28690#bib.bib4);[Kolluri et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib5)\)\. Recent advancements attempt to address this: HUMANLM\([Wu et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib26)\)models latent psychological states, and[Lu et al\. \(2025\)](https://arxiv.org/html/2609.28690#bib.bib9)predicts next\-step behavior in trajectories\. However, they remain bound by supervised fine\-tuning or specific domains\. UserLM\-R1\([Zhang et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib31)\)enhances user reasoning through multi\-reward reinforcement learning, with rewards evaluating generated rationales and responses at the turn level\. Overall, existing simulators fall short in capturing how continuous behavioral evolution impacts final dialogue outcomes\.

##### Applications of user simulator\.

User simulators are heavily utilized in dialogue system research for both training and evaluation, alleviating the reliance on costly human annotation\.Training:Simulators provide scalable interactive environments, either by generating large\-scale dialogue corpora \(e\.g\., CAMEL\([Li et al\., 2023](https://arxiv.org/html/2609.28690#bib.bib6)\)\) or by supplying multi\-turn dynamic feedback and reward signals in RL loops\([Wang et al\., 2025b](https://arxiv.org/html/2609.28690#bib.bib23);[Qian et al\., 2025b](https://arxiv.org/html/2609.28690#bib.bib14);[Yang et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib28);[Wu et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib25)\)\.Evaluation:Simulators overcome the limits of static benchmarks by enabling diverse and realistic interaction assessments\([Zhang et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib30);[Qian et al\., 2025a](https://arxiv.org/html/2609.28690#bib.bib13)\)\. However, the reliability of these simulated users remains a central bottleneck constraining the credibility of simulator\-based evaluation\([Naous et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib11)\)\.

## 3Problem Formulation

Interactive user simulation\.We formulate multi\-turn user simulation as conditional sequential generation through interaction with a conversational assistant\. Letuudenote the user profile and task background available before simulation, and lethth\_\{t\}denote the interaction context at turntt, including the dialogue history and the assistant’s latest utterance\. The simulator policy generates

\(ℓ^t,y^t\)∼πθ\(⋅∣u,ht\),ℓ^t=\(r^t,z^t\),\(\\hat\{\\ell\}\_\{t\},\\hat\{y\}\_\{t\}\)\\sim\\pi\_\{\\theta\}\(\\cdot\\mid u,h\_\{t\}\),\\qquad\\hat\{\\ell\}\_\{t\}=\(\\hat\{r\}\_\{t\},\\hat\{z\}\_\{t\}\),
wherer^t\\hat\{r\}\_\{t\}is a natural\-language rationale,z^t\\hat\{z\}\_\{t\}is an ordinal intent score, andy^t\\hat\{y\}\_\{t\}is the user’s response\. The assistant follows a policyπA\\pi\_\{A\}and responds to the evolving dialogue, producing the context for the next user turn\. A simulated session ends upon conversion, dropout, or reaching the maximum number of turns\.

Behavioral trajectory representation\.For each observed user responseyty\_\{t\}, we associate an annotationℓt=\(rt,zt\)\\ell\_\{t\}=\(r\_\{t\},z\_\{t\}\)consisting of a rationale and an intent score inferred from the dialogue\. These annotations provide an interpretable proxy for the user’s expressed engagement and commitment\. The intent score can increase, decrease, or remain unchanged as the interaction unfolds\. We represent an annotated human session and a simulated session as

τ=\{\(zt,yt\)\}t=1T,τ^=\{\(z^t,y^t\)\}t=1T∗,\\tau=\\\{\(z\_\{t\},y\_\{t\}\)\\\}\_\{t=1\}^\{T\},\\qquad\\hat\{\\tau\}=\\\{\(\\hat\{z\}\_\{t\},\\hat\{y\}\_\{t\}\)\\\}\_\{t=1\}^\{T^\{\*\}\},
whereTTandT∗T^\{\*\}are their respective lengths, andzTz\_\{T\}andz^T∗\\hat\{z\}\_\{T^\{\*\}\}denote their final intent scores\. Terminal outcomes characterize the observable result of an interaction, such as whether the user provides contact information\. This representation supports analysis of intent evolution, response content, and interaction length alongside terminal behavior\.

Behavioral modeling objective\.LetpH​\(τ∣u,h1,πA\)p\_\{H\}\(\\tau\\mid u,h\_\{1\},\\pi\_\{A\}\)denote the conditional distribution of annotated human trajectories, and letpθ​\(τ^∣u,h1,πA\)p\_\{\\theta\}\(\\hat\{\\tau\}\\mid u,h\_\{1\},\\pi\_\{A\}\)denote the trajectory distribution induced by the simulator interacting with the assistant\. Our goal is to learnπθ\\pi\_\{\\theta\}that captures human behavioral dynamics under these conditions, including the distribution of terminal outcomes, the evolution of expressed intent, and the linguistic and temporal characteristics of the interaction\. Observed human sessions provide empirical reference trajectories for learning; Section[4](https://arxiv.org/html/2609.28690#S4)describes the resulting supervision and policy optimization\. For evaluation, we aggregate human and simulated trajectories within groups of users with comparable pre\-interaction contexts, with each rollout initialized from its own user’s context\. The grouping procedure and behavioral metrics are specified in Section[5\.1](https://arxiv.org/html/2609.28690#S5.SS1)\.

## 4TRACER

Faithful multi\-turn user simulation requires both natural responses and coherent behavior across an interaction\. Relevant behavioral patterns include the final intent state, the sequence of intermediate intent transitions, and the pace at which the conversation progresses\. Fitting individual responses provides no explicit trajectory\-level constraint, while terminal supervision alone leaves intermediate behavior underdetermined\. This motivates two design questions: how to evaluate behavioral consistency across a complete session, and how to use that feedback to guide updates at individual turns\.

TRACER addresses these questions through two\-stage training \(Figure[1](https://arxiv.org/html/2609.28690#S4.F1)\)\. Supervised fine\-tuning initializes user\-like language generation and structured intent prediction\. Interactive reinforcement learning then combines outcome and trajectory rewards with deviation\-aware advantage modulation\. A shared trajectory alignment connects evaluation and optimization: its overall cost scores a completed interaction, while its local discrepancies inform turn\-level updates\.

![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/main_v7.png)Figure 1:TRAIL first learns from annotated human dialogues through supervised fine\-tuning, then performs multi\-turn reinforcement learning\. A shared DTW alignment provides the trajectory cost for session\-level rewards and local deviations for turn\-level advantage modulation\.### 4\.1Simulator Initialization and Interactive Rollout

Following Section[3](https://arxiv.org/html/2609.28690#S3), the simulator generates a rationalertr\_\{t\}, intent scoreztz\_\{t\}, and user replyyty\_\{t\}conditioned on user informationuuand contexthth\_\{t\}\. We initialize these outputs by supervised fine\-tuning on real dialogues and their derived annotations with a standard autoregressive likelihood objective\.

To train under contexts induced by its own responses, the simulator then interacts with a fixed assistant policyπA\\pi\_\{A\}\. Each rollout starts from the corresponding user information and initial assistant utterance, and continues until conversion, dropout, or the turn limit\. The generated user reply advances the interaction with the assistant\. For each input, we sampleGGtrajectories\{τ^i\}i=1G\\\{\\hat\{\\tau\}\_\{i\}\\\}\_\{i=1\}^\{G\}of lengthsTi∗T\_\{i\}^\{\*\}, scored against the paired human referenceτ=\{\(zt,yt\)\}t=1T\\tau=\\\{\(z\_\{t\},y\_\{t\}\)\\\}\_\{t=1\}^\{T\}\.

### 4\.2Session\-Level Rewards

Behavioral consistency has complementary outcome and process dimensions: a generated session should reflect the reference outcome as well as the intent evolution leading to it\. We therefore combine a terminal\-intent reward with a trajectory\-alignment reward\. Together, they provide supervision beyond local response similarity, while an additional format reward encourages valid structured outputs\.

##### Outcome consistency\.

The final annotated intent state summarizes the reference session’s outcome, providing a direct behavioral target beyond the wording of individual responses\. We measure agreement with this state through the terminal\-intent reward:

Rfinal,i=𝕀⁡\(zT=z^i,Ti∗\)\.R\_\{\\mathrm\{final\},i\}=\\mathbb\{I\}\\\!\\left\(z\_\{T\}=\\hat\{z\}\_\{i,T\_\{i\}^\{\*\}\}\\right\)\.\(1\)This provides an empirical terminal\-consistency signal\. Terminal agreement alone does not constrain how the interaction unfolds: sessions with the same final intent may exhibit substantially different intent trajectories\.

##### Trajectory alignment\.

A direct turn\-by\-turn comparison can penalize comparable intent transitions simply because they occur at different conversational paces\. We therefore use dynamic time warping \(DTW\)\([Sakoe & Chiba, 1978](https://arxiv.org/html/2609.28690#bib.bib17);[Myers et al\., 1980](https://arxiv.org/html/2609.28690#bib.bib10)\)to align the human and simulated intent sequences while preserving their temporal order\. With local cost\|z^i,t−zs\|\|\\hat\{z\}\_\{i,t\}\-z\_\{s\}\|, let𝒫i\\mathcal\{P\}\_\{i\}denote the set of warping paths satisfying the standard boundary, monotonicity, and step constraints\. The optimal path and its cumulative cost are

Pi∗∈arg⁡min⁡∑\(t,s\)∈PP∈𝒫i⁡\|z^i,t−zs\|,Di=∑\(t,s\)∈Pi∗\|z^i,t−zs\|\.P\_\{i\}^\{\*\}\\in\\arg\\min\_\{P\\in\\mathcal\{P\}\_\{i\}\}\\sum\_\{\(t,s\)\\in P\}\|\\hat\{z\}\_\{i,t\}\-z\_\{s\}\|,\\qquad D\_\{i\}=\\sum\_\{\(t,s\)\\in P\_\{i\}^\{\*\}\}\|\\hat\{z\}\_\{i,t\}\-z\_\{s\}\|\.\(2\)Flexible alignment accommodates differences in pacing, but does not by itself require similar session lengths\. We therefore combine the alignment cost with an explicit length penalty and a coefficient that discourages nearly constant intent sequences:

Rtraj,i=exp⁡\[−α​Di\+β​\|Ti∗−T\|T\]​ηi\.R\_\{\\mathrm\{traj\},i\}=\\exp\\\!\\left\[\-\\alpha\\,\\frac\{D\_\{i\}\+\\beta\|T\_\{i\}^\{\*\}\-T\|\}\{T\}\\right\]\\eta\_\{i\}\.\(3\)Here,α\>0\\alpha\>0controls reward sensitivity,β≥0\\beta\\geq 0weights the length mismatch, and normalization byTTscales the cost relative to the reference length\. The coefficientηi∈\(0,1\]\\eta\_\{i\}\\in\(0,1\]discounts degenerate, nearly constant intent trajectories\. Appendix[H](https://arxiv.org/html/2609.28690#A8)gives the alignment details\.

Outcome and trajectory agreement do not ensure that individual outputs follow the required structure\. We therefore include a format reward to check the prescribed output schema and basic field consistency\. Ifri,tfmt∈\{0,1\}r^\{\\mathrm\{fmt\}\}\_\{i,t\}\\in\\\{0,1\\\}is the turn\-level format score, the combined session reward is

Rfmt,i\\displaystyle R\_\{\\mathrm\{fmt\},i\}=1Ti∗​∑t=1Ti∗ri,tfmt,\\displaystyle=\\frac\{1\}\{T\_\{i\}^\{\*\}\}\\sum\_\{t=1\}^\{T\_\{i\}^\{\*\}\}r^\{\\mathrm\{fmt\}\}\_\{i,t\},Ri\\displaystyle R\_\{i\}=λ1​Rfmt,i\+λ2​Rfinal,i\+λ3​Rtraj,i\.\\displaystyle=\\lambda\_\{1\}R\_\{\\mathrm\{fmt\},i\}\+\\lambda\_\{2\}R\_\{\\mathrm\{final\},i\}\+\\lambda\_\{3\}R\_\{\\mathrm\{traj\},i\}\.\(4\)The resulting scalar incorporates output validity, terminal agreement, and intermediate trajectory information\. Reward hyperparameters are reported in Appendix[A](https://arxiv.org/html/2609.28690#A1)\.

### 4\.3Deviation\-Aware Policy Optimization

Following GRPO\([Shao et al\., 2024](https://arxiv.org/html/2609.28690#bib.bib19)\), we standardize rewards within each rollout group to obtain the trajectory advantageAiA\_\{i\}\. Broadcasting this scalar across turns does not reflect their different alignment deviations\. We therefore retain its global reinforcement direction while adapting update strength using the shared DTW alignment\.

##### Cumulative alignment deviation\.

We reuse the optimal DTW pathPi∗P\_\{i\}^\{\*\}from Eq\. equation[2](https://arxiv.org/html/2609.28690#S4.E2)to derive a local deviation for each generated turn\. With𝒜i,t=\{s:\(t,s\)∈Pi∗\}\\mathcal\{A\}\_\{i,t\}=\\\{s:\(t,s\)\\in P\_\{i\}^\{\*\}\\\}denoting the reference turns aligned to generated turntt, we define

di,t=1\|𝒜i,t\|​∑s∈𝒜i,t\|z^i,t−zs\|\.d\_\{i,t\}=\\frac\{1\}\{\|\\mathcal\{A\}\_\{i,t\}\|\}\\sum\_\{s\\in\\mathcal\{A\}\_\{i,t\}\}\|\\hat\{z\}\_\{i,t\}\-z\_\{s\}\|\.\(5\)This averages the discrepancies over all reference turns aligned to the same generated turn\. A local score describes agreement at one position, but does not indicate whether the subsequent interaction remains aligned\. To incorporate this continuation information, we aggregate deviations over a discounted suffix of the sampled trajectory:

ci,t=∑k=0Ti∗−tρk​di,t\+k,ρ∈\[0,1\]\.c\_\{i,t\}=\\sum\_\{k=0\}^\{T\_\{i\}^\{\*\}\-t\}\\rho^\{k\}d\_\{i,t\+k\},\\qquad\\rho\\in\[0,1\]\.\(6\)The discountρ\\rhocontrols how strongly later discrepancies affect the score: smaller values emphasize nearby turns, while larger values retain more information about the continuation\. Settingρ=0\\rho=0recovers the local signaldi,td\_\{i,t\}\.

##### Advantage modulation\.

A relatively high\-reward session can still contain poorly aligned segments, while a low\-reward session may exhibit varying degrees of deviation across turns\. We useci,tc\_\{i,t\}to adjust the strength of reinforcement or suppression within each session, with a bounded modulation that preserves the sign of the trajectory advantage:

A~i,t=wi,t​Ai,wi,t=1−λadv​sign⁡\(Ai\)​tanh⁡\(γ​ci,t\),\\widetilde\{A\}\_\{i,t\}=w\_\{i,t\}A\_\{i\},\\qquad w\_\{i,t\}=1\-\\lambda\_\{\\mathrm\{adv\}\}\\operatorname\{sign\}\(A\_\{i\}\)\\tanh\(\\gamma c\_\{i,t\}\),\(7\)whereλadv∈\[0,1\]\\lambda\_\{\\mathrm\{adv\}\}\\in\[0,1\]bounds the modulation magnitude andγ\>0\\gamma\>0controls sensitivity\. Equivalently,A~i,t=Ai−λadv​\|Ai\|​tanh⁡\(γ​ci,t\)\\widetilde\{A\}\_\{i,t\}=A\_\{i\}\-\\lambda\_\{\\mathrm\{adv\}\}\|A\_\{i\}\|\\tanh\(\\gamma c\_\{i,t\}\)\. Larger cumulative deviations reduce reinforcement whenAi\>0A\_\{i\}\>0and increase the magnitude of suppression whenAi<0A\_\{i\}<0\. The modulation prevents sign reversal, whileλadv=0\\lambda\_\{\\mathrm\{adv\}\}=0recovers the unmodulated trajectory advantage\.

We useA~i,t\\widetilde\{A\}\_\{i,t\}in place ofAiA\_\{i\}for all simulator tokens selected for optimization at turntt, excluding assistant and padding tokens\. Alignment and advantage values remain fixed during each update\. This preserves the GRPO clipped surrogate structure while changing the turn\-level optimization signal\.

## 5Experiments

### 5\.1Experimental Setup

##### Datasets\.

We construct CustomerService\-Dialogue from real\-world dialogues between users and merchant\-side customer service agents on a mobile application, spanning industries such as healthcare and automotive\. Each session contains a 3–15\-turn dialogue, a user profile \(demographics, personality, and initial intent\), and merchant information\. Conversion is defined by whether the user provides contact information\. Claude\-4\.5\-Sonnet provides weak turn\-level supervisionℓt=\(rt,zt\)\\ell\_\{t\}=\(r\_\{t\},z\_\{t\}\), generating a rationale and an intent score conditioned on the full dialogue and user profile\. Human assessment on a stratified sample shows strong annotation agreement \(Table[4](https://arxiv.org/html/2609.28690#A2.T4)\)\.

The dataset contains 8,040 training sessions and 3,866 test sessions organized into 762 reference groups\. Each group contains sessions from the same merchant, with similar personality traits and initial user needs selected using semantic similarity thresholds\. Grouping is used only for evaluation; training retains paired\-session supervision\. Construction and grouping details appear in Appendix[B](https://arxiv.org/html/2609.28690#A2)\.

##### Evaluation Protocol\.

For each test instance, the simulated user model receives the user profile and the assistant’s first utterance\. It role\-plays as the user and conducts a multi\-turn interaction with the assistant until the dialogue terminates via conversion, dropout or reaching a maximum turns\. The assistant is the same model deployed online, conditioned on the merchant information and user query, and follows a standardized SOP\-driven policy that yields low\-variance, reproducible responses\. Metrics are computed over the simulated trajectories against the ground\-truth dialogues\.

##### Metrics\.

We evaluate three dimensions\.Session\-Level Behavior: PCR reports the overall conversion rate; ACC and conversion\-class F1 measure outcome agreement with the original paired sessions\.Group\-Level Fidelity: Group\-Δ\\DeltaPCR measures the absolute difference between simulated and real conversion rates within each group; W1\-Turns measures the Wasserstein\-1 distance between their dialogue\-length distributions\.Semantic Content: OT\-DTW measures optimal transport distance between simulated and real sessions within each group, using DTW\-based semantic trajectory costs\. Group\-based metrics are averaged across the 762 groups\.

##### Baselines\.

We compare three categories of user simulators\.General LLM baselines:We evaluate several general LLMs as zero\-shot user simulators, including the Qwen2\.5\-7B\-Instruct\([Qwen et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib16)\), Qwen3\-4B\-Instruct\([Yang et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib27)\), DeepSeek\-V3\.2\([Liu et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib8)\), Gemini\-3\.1\-Pro, Doubao\-Seed\-2\.0\-Pro and Claude\-4\.5\-Sonnet\.Role\-playing baselines:We include models explicitly designed for role\-playing, including Doubao\-Seed\-Character, Qwen\-Character, and CharGLM\-4, which emphasize strong role consistency and conversational naturalness\.Training\-based baseline:We adopt HUMANLM\([Wu et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib26)\), which introduces latent natural\-language states aligned with responses via RL\. We train it on our dataset using SFT and RL for fair comparison\.

##### Implementation details\.

We instantiate TRACER using Qwen2\.5\-7B\-Instruct and Qwen3\-4B\-Instruct\. In the Stage 1, we train the models for 2 epochs with a learning rate of10−510^\{\-5\}\. In Stage 2, we train for 6 epochs, sample 8 rollout trajectories per prompt\. All models use an inference temperature of 0\.7\. Full configurations, evaluation details, and stability analysis are provided in Appendix[A](https://arxiv.org/html/2609.28690#A1)\.

Table 1:Overall comparison of user simulators across three evaluation dimensions: Session\-Level Behavior \(PCR, ACC, F1\), Group\-Level Fidelity \(Group\-Δ\\DeltaPCR, W1\-Turns\), and Semantic Content \(OT\-DTW\)\. Groups are balanced 1:1; the session\-level real PCR is 51\.24%\. Best results are in bold, second\-best are underlined\.

### 5\.2Main Results

Table[1](https://arxiv.org/html/2609.28690#S5.T1)reports the performance of all methods along the three evaluation dimensions\. We summarize our main findings as follows\.

General LLMs exhibit insufficient behavioral alignment, with opposing tendencies toward over\-compliance and over\-refusal\.Their PCR ranges from 5\.8% to 79\.7% , compared with the real conversion rate of 50%, while ACC remains within 50–59% and F1 reaches at most 68\.2%\. These results suggest that general LLMs tend to be either overly willing or overly reluctant to convert, failing to match the conversion behavior observed among real users\.

Role\-playing baselines remain behaviorally misaligned despite their focus on persona consistency\.Doubao\-Seed\-Character, Qwen\-Character, and CharGLM\-4 exhibit a pronounced under\-conversion tendency, with PCRs of 9\.9–26\.2% and F1 scores of 23\.9–37\.0%\. Their high W1\-Turns and OT\-DTW further reveal mismatches in dialogue length and semantic progression\. These results suggest that persona\-oriented modeling alone is insufficient to reproduce real users’ conversion tendencies and multi\-turn interaction patterns\.

TRACER jointly achieves linguistic\-style consistency and trajectory\-level behavioral consistency\.TRACER\-7B achieves a PCR of 50\.5%, together with the best ACC and F1, exceeding the strongest baseline by 11\.4 F1 points\. It also achieves the lowest Group\-Δ\\DeltaPCR and OT\-DTW, while remaining competitive on W1\-Turns\. The 4B variant shows similar improvements, supporting the framework’s effectiveness across both backbones\. Compared with HUMANLM trained on the same data using SFT and RL, TRACER better captures real users’ conversion tendencies and semantic trajectories\. These results highlight the value of jointly aligning terminal outcomes and the course of interaction\.

### 5\.3Detailed Analysis

#### 5\.3\.1Human Evaluations

To assess simulation fidelity, we conduct a Turing\-style evaluation testing whether AI\-simulated users are distinguishable from real humans\. Annotators are given two dialogues with the same user profile information, one produced by a real user and the other by a simulator, and are asked to identify the human\-generated dialogue\. We randomly sample 500 sessions from the test set, covering diverse personas and intents\. Details about the annotators are provided in the Appendix[G\.1](https://arxiv.org/html/2609.28690#A7.SS1)\.

##### TRACER\-7B achieves human\-level indistinguishability\.

Figure 2:Human\-indistinguishability evaluation\. Accuracy is shown as bars, while the presence of non\-human characteristics is indicated with markers\. The dashed line at 50% represents the random guess threshold\.Figure[2](https://arxiv.org/html/2609.28690#S5.F2)compares TRACER\-7B with selected general\-purpose and role\-playing baselines from Table[1](https://arxiv.org/html/2609.28690#S5.T1), as well as its SFT\-only variant\. General LLMs are correctly identified in over 98% of cases, revealing persistent machine\-like patterns\. The SFT variant reduces identification accuracy to 61\.80% but still falls short of indistinguishability\. In contrast, TRACER\-7B reaches 49\.50%, converging to the random\-guess threshold and indicating that it captures subtle human behavioral nuances missed by prior simulators\.

##### Why are baselines recognized?

Table[5](https://arxiv.org/html/2609.28690#A4.T5)groups detection cues into three dimensions\. \(1\)*Linguistic over\-regularity*: baselines almost entirely lack human imperfection and frequently show overly neat structure\. \(2\)*Mis\-calibrated interaction rhythm*: they are consistently information\-overloading and excessively cooperative, with Doubao variants further exhibiting mechanical patterns\. \(3\)*Out\-of\-character*behavior appears in roughly 40% of cases, while higher\-order defects such as emotional discontinuity remain rare\. Details are in Appendix[D](https://arxiv.org/html/2609.28690#A4)\.

#### 5\.3\.2Ablation Study

Table 2:Ablation study of TRACER\-7B on session\-level behavior, dialogue structure, and semantic content\. Results are reported as mean±\\pmstandard deviation over five evaluation seeds\.Table[2](https://arxiv.org/html/2609.28690#S5.T2)summarizes the ablation results for TRACER\. The full model achieves the best or tied\-best performance across all five direction\-specified metrics\. Removing advantage modulation or the trajectory\-level rewardRt​r​a​jR\_\{traj\}consistently degrades behavioral accuracy, group\-level fidelity, and semantic alignment, supporting their complementary contributions\. Replacing DTW with strict turn\-by\-turn matching primarily impairs structural and semantic fidelity, with smaller effects on ACC and F1\. Disabling multi\-turn RL substantially reduces behavioral accuracy and worsens Group\-Δ\\DeltaPCR, highlighting the value of multi\-turn optimization\. Removing RL entirely causes the largest overall degradation, underscoring its importance beyond supervised fine\-tuning\.

#### 5\.3\.3Out\-of\-Domain Generalization

Figure 3:Performance Comparison of Models on the Out\-of\-Domain Dataset\.To assess the generalization capability of TRACER beyond its training distribution, we evaluate on CSC\-Conv\([Zhu et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib33)\), an open\-source dataset of real user–agent conversations in the financial customer\-service that is entirely unseen during training\. For each dialogue, we extract the user’s initial intent and a binary label indicating whether the issue is ultimately resolved\. We randomly sample 1,000 dialogues and define the session\-level terminal state accordingly\. We compare TRACER with the strongest baselines from Table[1](https://arxiv.org/html/2609.28690#S5.T1), measuring session\-level ACC and F1 on the terminal state, together with dialogue embedding similarity\. As shown in Figure[3](https://arxiv.org/html/2609.28690#S5.F3), TRACER achieves the best results across all three metrics, outperforming the strongest baseline by \+8\.4, \+4\.8, and \+0\.7 points respectively, despite never being exposed to this domain\. These gains at both the outcome and trajectory levels suggest that TRACER learns transferable user\-simulation behaviors rather than domain\-specific patterns\. Detailed settings are provided in Appendix[E](https://arxiv.org/html/2609.28690#A5)\.

## 6Dynamic Marketing Benchmark

TRACER aligns simulated users with human interaction trajectories, capturing both intermediate intent changes and terminal decisions\. These capabilities support multi\-turn assistant evaluation: user responses shape subsequent interactions, while terminal decisions provide outcome feedback\. We instantiate this evaluation setting as the Dynamic Marketing Benchmark \(DM\-Bench\), which compares assistant response quality and simulated conversion under shared user and merchant conditions\. We further examine whether the simulator’s local intent changes agree with human expectations about assistant strategies\.

![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/profile_personality_v2.png)\(a\)Personality
![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/profile_city_v2.png)\(b\)City Profile
![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/profile_age_v2.png)\(c\)Demographic Profile
![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/profile_industry_v2.png)\(d\)Merchant Distribution

![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/bench_v4.png)\(e\)LLM Performance as Customer Service Agents

Figure 4:Overview of DM\-Bench: User Profiles, Merchant Distribution, and Model Performance\.### 6\.1Benchmark Construction

DM\-Bench comprises 734 instances sampled from CustomerService\-Dialogue\. Each instance contains a user profile, including demographic attributes and initial intent, together with merchant information\. Seven assistant models interact with TRACER under the same set of initial conditions, with user responses and intent evolving as each conversation unfolds\. Figure[4](https://arxiv.org/html/2609.28690#S6.F4)summarizes the user and merchant distributions\. The data source and construction protocol are described in Section[5\.1](https://arxiv.org/html/2609.28690#S5.SS1)\.

We report two complementary metrics\.Simulated Conversion Rate \(CR\)measures the proportion of sessions ending in conversion, as determined by the user simulator\.Response Quality \(RQ\)follows the LLM\-as\-a\-judge protocol of[Zhu et al\. \(2026\)](https://arxiv.org/html/2609.28690#bib.bib33), assessing multi\-turn responses in terms of accuracy, helpfulness, informativeness, and empathy\. Together, these metrics capture response quality and terminal outcomes within the simulated environment\.

### 6\.2Validation of Behavioral Feedback

We evaluate whether TRACER’s turn\-level intent changes agree with human expectations using 449 single\-turn strategy instances from 200 conversations across assistant models\. These cover six strategies: empathy, urgency, benefit/discount framing, informative help, contact\-information requests, and probing/clarification\. Three annotators independently label the expected intent change as decrease, unchanged, or increase\. Reference labels are obtained by majority vote, with pairwise agreement above 90%\. Against these labels, TRACER achieves recalls of 88\.7%, 74\.2%, and 81\.0% for decreases, unchanged responses, and increases, respectively\. This agreement supports the use of its turn\-level behavioral feedback in multi\-turn evaluation, although the reference reflects human expectations rather than observed user outcomes\. Appendix[F\.2](https://arxiv.org/html/2609.28690#A6.SS2)provides the strategy\-level breakdown\.

### 6\.3Results and Findings

We report the results in Figure[4](https://arxiv.org/html/2609.28690#S6.F4), Table[7](https://arxiv.org/html/2609.28690#A6.T7)and Table[8](https://arxiv.org/html/2609.28690#A6.T8), with the following conclusions:\(1\) Response quality is not a reliable proxy for conversion effectiveness\.Although RQ and CR show a moderate positive rank correlation across the seven models \(Spearmanρ=0\.57\\rho=0\.57,p=0\.18p=0\.18,n=7n=7\); for instance, Claude\-4\.5\-Sonnet ranks highest on RQ but is surpassed on CR by DeepSeek\-V3\.2 and Qwen3\-235B\. This indicates that response\-quality evaluation does not fully capture persuasion effectiveness\.\(2\) Even frontier LLMs exhibit a substantial conversion ceiling\.Despite high response\-quality scores, top models still fail to convert a large fraction of users, revealing limited adaptation to user intent and decision dynamics in multi\-turn interactions\.

These findings show the value of DM\-Bench as a user\-centered, outcome\-driven evaluation protocol: it reveals performance gaps that are largely invisible under quality\-only benchmarks, especially the gap between generating good\-sounding responses and generating decision\-moving ones\.

## 7Conclusion

In this work, we present TRACER, a multi\-turn user simulator aligned with real users at both linguistic and decision\-making levels, trained via a two\-stage paradigm of SFT and trajectory\-aligned RL\. Building on TRACER, we further construct the Dynamic Marketing Benchmark, which jointly measures persuasion effectiveness and response quality\. Experiments show that TRACER largely outperforms existing simulators and achieves near\-human fidelity in Turing tests, while revealing a misalignment between response quality and persuasion effectiveness in mainstream LLMs\. In future work, we plan to extend TRACER to broader interactive scenarios and explore its use as a training environment for dialogue agents\.

### AI use statement

During the preparation of this study, LLMs were utilized solely as auxiliary tools\. The application of these models was strictly limited to linguistic refinement, including grammar and spelling corrections, as well as technical support for drafting and debugging code during the experimental phase\. All suggestions provided by the models underwent rigorous auditing, revision, and verification by the authors\. It must be explicitly stated that LLMs played no role in the conceptualization of the research problem, the development of the theoretical framework, the design of experimental protocols, or the scientific interpretation of the results\. All core academic contributions such as the research logic, methodological design, and final conclusions were independently completed by the whole authors\.

### Ethics statement

All user behavior data involved in this study were sampled and processed in strict accordance with relevant laws, regulations, platform guidelines, and privacy protection requirements\. The scope of this research is limited exclusively to data resources authorized for scientific analysis and benchmark construction\. Prior to the release of the public benchmark, the raw data underwent multiple stages of quality control, representative sampling, and de\-identification procedures\. Sensitive information possessing identifying characteristics, such as names, contact details, and specific addresses, has been either deleted or replaced with semantic placeholders\. Furthermore, this study implemented filters for inappropriate or potentially harmful content and conducted manual audits to ensure that the processed data meet ethical standards\. All analyses and experiments were performed using the desensitized and cleaned dataset\. The authors assume full legal and academic responsibility for ensuring that the construction and utilization of the benchmark data comply with all relevant data governance and privacy protection norms\.

### Reproducibility statement

To support reproducibility of our experiments, we provide the TRACER code at[https://github\.com/ChenGeng0102/TRACER](https://github.com/ChenGeng0102/TRACER)and document the experimental procedures in the paper and appendices\. Section[4](https://arxiv.org/html/2609.28690#S4)specifies the training objectives and optimization procedure, while Appendix[A](https://arxiv.org/html/2609.28690#A1)reports the hardware, training and inference configurations, and hyperparameter settings\. Appendix[B](https://arxiv.org/html/2609.28690#A2)describes data construction and preprocessing, and Appendix[J](https://arxiv.org/html/2609.28690#A10)provides the prompts used for annotation and simulation\. The evaluation protocol is presented in Section[5\.1](https://arxiv.org/html/2609.28690#S5.SS1), with additional details on out\-of\-domain evaluation, DM\-Bench, and human annotation in Appendices[E](https://arxiv.org/html/2609.28690#A5)–[G](https://arxiv.org/html/2609.28690#A7)\.

## References

- Abbasiantaeb et al\. \(2024\)Zahra Abbasiantaeb, Yifei Yuan, Evangelos Kanoulas, and Mohammad Aliannejadi\.Let the llms talk: Simulating human\-to\-human conversational qa via zero\-shot llm\-to\-llm interactions\.In*Proceedings of the 17th ACM International Conference on Web Search and Data Mining*, pp\. 8–17, 2024\.
- Dou et al\. \(2025\)Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, and Jianfeng Gao\.SimulatorArena: Are user simulators reliable proxies for multi\-turn evaluation of AI assistants?In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng \(eds\.\),*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pp\. 35212–35290, Suzhou, China, November 2025\. Association for Computational Linguistics\.ISBN 979\-8\-89176\-332\-6\.doi:10\.18653/v1/2025\.emnlp\-main\.1786\.URL[https://aclanthology\.org/2025\.emnlp\-main\.1786/](https://aclanthology.org/2025.emnlp-main.1786/)\.
- Hu et al\. \(2023\)Zhiyuan Hu, Yue Feng, Anh Tuan Luu, Bryan Hooi, and Aldo Lipani\.Unlocking the potential of user feedback: Leveraging large language model as user simulators to enhance dialogue system\.In*Proceedings of the 32nd ACM International Conference on Information and Knowledge Management*, pp\. 3953–3957, 2023\.
- Kim & Yang \(2025\)Jaehyung Kim and Yiming Yang\.Few\-shot personalization of llms with mis\-aligned responses\.In*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pp\. 11943–11974, 2025\.
- Kolluri et al\. \(2025\)Akaash Kolluri, Shengguang Wu, Joon Sung Park, and Michael S Bernstein\.Finetuning llms for human behavior prediction in social science experiments\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pp\. 30084–30099, 2025\.
- Li et al\. \(2023\)Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem\.Camel: Communicative agents for "mind" exploration of large scale language model society, 2023\.
- Liao et al\. \(2025\)Chonghua Liao, Ke Wang, Yuchuan Wu, Fei Huang, and Yongbin Li\.Moa: Multi\-objective alignment for role\-playing agents\.*arXiv preprint arXiv:2512\.09756*, 2025\.
- Liu et al\. \(2025\)Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, et al\.Deepseek\-v3\. 2: Pushing the frontier of open large language models\.*arXiv preprint arXiv:2512\.02556*, 2025\.
- Lu et al\. \(2025\)Yuxuan Lu, Jing Huang, Yan Han, Bingsheng Yao, Sisong Bei, Jiri Gesi, Yaochen Xie, Qi He, Dakuo Wang, et al\.Can llm agents simulate multi\-turn human behavior? evidence from real online customer behavior data\.*arXiv preprint arXiv:2503\.20749*, 2025\.
- Myers et al\. \(1980\)C\. Myers, L\. Rabiner, and A\. Rosenberg\.Performance tradeoffs in dynamic time warping algorithms for isolated word recognition\.*IEEE Transactions on Acoustics, Speech, and Signal Processing*, 28\(6\):623–635, 1980\.doi:10\.1109/TASSP\.1980\.1163491\.
- Naous et al\. \(2025\)Tarek Naous, Philippe Laban, Wei Xu, and Jennifer Neville\.Flipping the dialogue: Training and evaluating user language models\.*arXiv preprint arXiv:2510\.06552*, 2025\.
- Park et al\. \(2024\)Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein\.Generative agent simulations of 1,000 people\.*arXiv preprint arXiv:2411\.10109*, 2024\.
- Qian et al\. \(2025a\)Cheng Qian, Zuxin Liu, Akshara Prabhakar, Zhiwei Liu, Jianguo Zhang, Haolin Chen, Heng Ji, Weiran Yao, Shelby Heinecke, Silvio Savarese, et al\.Userbench: An interactive gym environment for user\-centric agents\.*arXiv preprint arXiv:2507\.22034*, 2025a\.
- Qian et al\. \(2025b\)Cheng Qian, Zuxin Liu, Akshara Prabhakar, Jielin Qiu, Zhiwei Liu, Haolin Chen, Shirley Kokane, Heng Ji, Weiran Yao, Shelby Heinecke, et al\.Userrl: Training interactive user\-centric agent via reinforcement learning\.*arXiv preprint arXiv:2509\.19736*, 2025b\.
- Qiu & Lan \(2024\)Huachuan Qiu and Zhenzhong Lan\.Interactive agents: Simulating counselor\-client psychological counseling via role\-playing llm\-to\-llm interactions\.*arXiv preprint arXiv:2408\.15787*, 2024\.
- Qwen et al\. \(2025\)Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu\.Qwen2\.5 technical report, 2025\.URL[https://arxiv\.org/abs/2412\.15115](https://arxiv.org/abs/2412.15115)\.
- Sakoe & Chiba \(1978\)H\. Sakoe and S\. Chiba\.Dynamic programming algorithm optimization for spoken word recognition\.*IEEE Transactions on Acoustics, Speech, and Signal Processing*, 26\(1\):43–49, 1978\.doi:10\.1109/TASSP\.1978\.1163055\.
- Senin \(2008\)Pavel Senin\.Dynamic time warping algorithm review\.*Information and Computer Science Department University of Hawaii at Manoa Honolulu, USA*, 855\(1\-23\):40, 2008\.
- Shao et al\. \(2024\)Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al\.Deepseekmath: Pushing the limits of mathematical reasoning in open language models\.*arXiv preprint arXiv:2402\.03300*, 2024\.
- Sheng et al\. \(2024\)Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu\.Hybridflow: A flexible and efficient rlhf framework\.*arXiv preprint arXiv: 2409\.19256*, 2024\.
- Shi et al\. \(2025\)Yimin Shi, Yang Fei, Shiqi Zhang, Haixun Wang, and Xiaokui Xiao\.You are what you bought: Generating customer personas for e\-commerce applications\.In*Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval*, pp\. 1810–1819, 2025\.
- Wang et al\. \(2025a\)Kuang Wang, Xianfei Li, Shenghao Yang, Li Zhou, Feng Jiang, and Haizhou Li\.Know you first and be you better: Modeling human\-like user simulators via implicit profiles\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 21082–21107, 2025a\.
- Wang et al\. \(2025b\)Peisong Wang, Ruotian Ma, Bang Zhang, Xingyu Chen, Zhiwei He, Kang Luo, Qingsong Lv, Qingxuan Jiang, Zheng Xie, Shanyi Wang, et al\.Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents\.*arXiv preprint arXiv:2507\.03112*, 2025b\.
- Wang et al\. \(2025c\)Zongsheng Wang, Kaili Sun, Bowen Wu, Qun Yu, Ying Li, and Baoxun Wang\.Raiden\-r1: Improving role\-awareness of llms via grpo with verifiable reward\.*arXiv preprint arXiv:2505\.10218*, 2025c\.
- Wu et al\. \(2025\)Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng, Gavin Li, Yao Dou, Weixin Cai, James Zou, Jure Leskovec, and Jianfeng Gao\.Collabllm: From passive responders to active collaborators\.In*International Conference on Machine Learning*, pp\. 67260–67283\. PMLR, 2025\.
- Wu et al\. \(2026\)Shirley Wu, Evelyn Choi, Arpandeep Khatua, Zhanghan Wang, Joy He\-Yueya, Tharindu Cyril Weerasooriya, Wei Wei, Diyi Yang, Jure Leskovec, and James Zou\.Humanlm: Simulating users with state alignment beats response imitation\.*arXiv preprint arXiv:2603\.03303*, 2026\.
- Yang et al\. \(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al\.Qwen3 technical report\.*arXiv preprint arXiv:2505\.09388*, 2025\.
- Yang et al\. \(2026\)Haojin Yang, Ai Jian, Xinyue Huang, Yiwei Wang, Weipeng Zhang, Ke Zeng, Xunliang Cai, and Jingqing Ruan\.Harmonizing dense and sparse signals in multi\-turn rl: Dual\-horizon credit assignment for industrial sales agents\.*arXiv preprint arXiv:2603\.01481*, 2026\.
- Ye et al\. \(2025\)Xinge Ye, Rui Wang, Yuchuan Wu, Victor Ma, Feiteng Fang, Fei Huang, and Yongbin Li\.Cpo: Addressing reward ambiguity in role\-playing dialogue via comparative policy optimization, 2025\.URL[https://arxiv\.org/abs/2508\.09074](https://arxiv.org/abs/2508.09074)\.
- Zhang et al\. \(2025\)Bang Zhang, Ruotian Ma, Qingxuan Jiang, Peisong Wang, Jiaqi Chen, Zheng Xie, Xingyu Chen, Yue Wang, Fanghua Ye, Jian Li, et al\.Sentient agent as a judge: Evaluating higher\-order social cognition in large language models\.*arXiv preprint arXiv:2505\.02847*, 2025\.
- Zhang et al\. \(2026\)Feng Zhang, Shijia Li, Chunmao Zhang, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, Jingwen Xu, and Han Liu\.Userlm\-r1: Modeling human reasoning in user language models with multi\-reward reinforcement learning\.*arXiv preprint arXiv:2601\.09215*, 2026\.
- Zheng et al\. \(2024\)Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma\.Llamafactory: Unified efficient fine\-tuning of 100\+ language models\.In*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\)*, Bangkok, Thailand, 2024\. Association for Computational Linguistics\.URL[http://arxiv\.org/abs/2403\.13372](http://arxiv.org/abs/2403.13372)\.
- Zhu et al\. \(2026\)Jie Zhu, Huaixia Dou, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang, and Fang Kong\.Evaluating, synthesizing, and enhancing for customer support conversation\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 40, pp\. 35185–35194, 2026\.

## Appendix

## Appendix AExperimental Details

### A\.1Implementation Details

##### Training configuration\.

All training experiments were conducted on a hardware platform equipped with8×8\\timesNVIDIA H200 GPUs\. For the backbone models, we mainly adopted Qwen2\.5\-7B\-Instruct\([Qwen et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib16)\)and Qwen3\-4B\-Instruct\([Yang et al\., 2025](https://arxiv.org/html/2609.28690#bib.bib27)\)as the base models, and trained them through supervised fine\-tuning and multi\-turn reinforcement learning, respectively\.

During the supervised fine\-tuning stage, we implemented the SFT training pipeline based on LLaMA\-Factory\([Zheng et al\., 2024](https://arxiv.org/html/2609.28690#bib.bib32)\)and used DeepSpeed ZeRO\-2 for distributed training\. The maximum context length was set to 4096 tokens\. The per\-GPU batch size was set to 4, with a gradient accumulation step of 2\. The models were trained for 2 epochs with a learning rate of1\.0×10−51\.0\\times 10^\{\-5\}\. We adopted a cosine learning rate scheduler with a warmup ratio of 0\.1\. All SFT experiments were conducted using bf16 precision to improve training efficiency and reduce GPU memory consumption\.

During the multi\-turn reinforcement learning stage, we modified the verl framework\([Sheng et al\., 2024](https://arxiv.org/html/2609.28690#bib.bib20)\)to construct a multi\-turn RL interaction training framework tailored to the user simulator dialogue learning task\. The global training batch size was set to 32, the mini\-batch size was set to 32, and the per\-GPU micro\-batch size was set to 4\. For each prompt, we sampled 8 rollout trajectories\. The total number of training epochs was set to 6, and the sampling temperature was set to 1\.0\. On the input side, the maximum prompt length was set to 6144 tokens, while the maximum response length was set to 2048 tokens\. We further limited the maximum number of dialogue turns during rollout to 15\. Under the above configuration, multi\-turn RL training took approximately 12 hours\.

##### Inference configuration\.

During the inference stage, the temperature was set to 0\.7 for both the baseline models and our model to ensure a consistent decoding configuration across all compared methods\. Following the dataset setting, the maximum number of interaction turns was capped at 15 during inference, and sessions that did not reach conversion within this limit were treated as dropout\.

##### Evaluation details\.

Across all evaluations, we employ a unified computational setup to ensure comparability among different methods\. All token\-related metrics are computed using the Qwen2\.5\-7B\-Instruct tokenizer\. For the Embedding Similarity metric, we utilize bge\-base\-zh\-v1\.5 as the encoding model to map both simulated and ground\-truth user replies into normalized embeddings, measuring their semantic consistency via cosine similarity\. For MAUVE, we also use bge\-base\-zh\-v1\.5 to encode replies\. Except for the input features and computation device, we keep the random seed, number of clusters, and other internal hyperparameters of MAUVE at their default settings\.

##### Reward and advantage hyperparameters\.

Unless otherwise specified, we use the same reward and advantage modulation configurations across both model scales\. All trajectory\-level rewards are computed at the session level and propagated to valid user turns via the proposed advantage modulation strategy\. For the DTW\-based trajectory reward in Eq\. equation[3](https://arxiv.org/html/2609.28690#S4.E3), we set the path\-deviation sensitivity coefficient toα=1\.0\\alpha=1\.0and the length\-penalty coefficient toβ=1\.0\\beta=1\.0\. The flattening coefficient is set toη=0\.2\\eta=0\.2when the predicted intent trajectory degenerates into an overly flat sequence, andη=1\.0\\eta=1\.0otherwise\. For the hierarchical reward in Eq\. equation[4](https://arxiv.org/html/2609.28690#S4.E4), we set\(λ1,λ2,λ3\)=\(0\.2,0\.5,0\.3\)\(\\lambda\_\{1\},\\lambda\_\{2\},\\lambda\_\{3\}\)=\(0\.2,0\.5,0\.3\)for the format reward, terminal intent consistency reward, and trajectory alignment reward, respectively\. For the cumulative future deviation score in Eq\. equation[6](https://arxiv.org/html/2609.28690#S4.E6), we set the future\-deviation discount factor toρ=0\.4\\rho=0\.4\. For the turn\-level advantage modulation in Eq\. equation[7](https://arxiv.org/html/2609.28690#S4.E7), we set the modulation strength toλadv=0\.2\\lambda\_\{\\mathrm\{adv\}\}=0\.2, the deviation sensitivity coefficient toγ=0\.5\\gamma=0\.5, and the numerical stability term toϵ=10−6\\epsilon=10^\{\-6\}\. All coefficients were selected on a held\-out validation split\.

### A\.2Sensitivity Analysis of Hyperparameters in Multi\-Turn Advantage Modulation

Table 3:Hyperparameter sensitivity of TRACER\-7B\. Each configuration is evaluated using its corresponding fixed checkpoint with five evaluation seeds \(42–46\); hyperparameter variants correspond to separately trained checkpoints\. Results are mean±\\pmstandard deviation across evaluation seeds, rather than independent training runs\. PCR, ACC, and F1 are reported in percent\. The reference PCR is 51\.24%\. Bold indicates the best mean for each directional metric\.Configurationλadv\\lambda\_\{\\mathrm\{adv\}\}γ\\gammaρ\\rhoPCRACC↑\\uparrowF1↑\\uparrowGroup\-Δ\\DeltaPCR↓\\downarrowW1\-Turns↓\\downarrowOT\-DTW↓\\downarrowWithout advantage modulation–––49\.84±0\.4949\.84\\pm 0\.4977\.68±0\.3377\.68\\pm 0\.3377\.92±0\.3377\.92\\pm 0\.330\.180±0\.0030\.180\\pm 0\.0031\.016±0\.0201\.016\\pm 0\.0200\.370±0\.0020\.370\\pm 0\.002Default TRACER\-7B0\.20\.50\.450\.29±0\.2050\.29\\pm 0\.2079\.04±0\.1979\.04\\pm 0\.1979\.28±0\.2179\.28\\pm 0\.210\.166±0\.0030\.166\\pm 0\.0030\.884±0\.006\\textbf\{0\.884\}\\pm 0\.0060\.351±0\.001\\mathbf\{0\.351\}\\pm 0\.001ρ=0\.0\\rho=0\.00\.20\.50\.050\.52±0\.3250\.52\\pm 0\.3278\.36±0\.2078\.36\\pm 0\.2078\.74±0\.1578\.74\\pm 0\.150\.176±0\.0020\.176\\pm 0\.0020\.924±0\.0120\.924\\pm 0\.0120\.373±0\.0010\.373\\pm 0\.001ρ=0\.1\\rho=0\.10\.20\.50\.147\.21±0\.2247\.21\\pm 0\.2279\.06±0\.2479\.06\\pm 0\.2478\.73±0\.2678\.73\\pm 0\.260\.164±0\.0030\.164\\pm 0\.0030\.975±0\.0070\.975\\pm 0\.0070\.369±0\.0010\.369\\pm 0\.001ρ=0\.3\\rho=0\.30\.20\.50\.348\.66±0\.1948\.66\\pm 0\.1978\.55±0\.2978\.55\\pm 0\.2978\.53±0\.2878\.53\\pm 0\.280\.170±0\.0040\.170\\pm 0\.0040\.977±0\.0170\.977\\pm 0\.0170\.365±0\.0010\.365\\pm 0\.001ρ=0\.5\\rho=0\.50\.20\.50\.553\.86±0\.1653\.86\\pm 0\.1678\.42±0\.2478\.42\\pm 0\.2479\.46±0\.2379\.46\\pm 0\.230\.177±0\.0030\.177\\pm 0\.0030\.947±0\.0090\.947\\pm 0\.0090\.365±0\.0020\.365\\pm 0\.002λadv=0\.1\\lambda\_\{\\mathrm\{adv\}\}=0\.10\.10\.50\.447\.78±0\.3247\.78\\pm 0\.3279\.80±0\.20\\mathbf\{79\.80\}\\pm 0\.2079\.60±0\.25\\mathbf\{79\.60\}\\pm 0\.250\.160±0\.002\\mathbf\{0\.160\}\\pm 0\.0021\.059±0\.0121\.059\\pm 0\.0120\.382±0\.0020\.382\\pm 0\.002λadv=0\.3\\lambda\_\{\\mathrm\{adv\}\}=0\.30\.30\.50\.452\.16±0\.2552\.16\\pm 0\.2578\.47±0\.2978\.47\\pm 0\.2979\.18±0\.2579\.18\\pm 0\.250\.172±0\.0050\.172\\pm 0\.0051\.101±0\.0241\.101\\pm 0\.0240\.371±0\.0010\.371\\pm 0\.001γ=0\.4\\gamma=0\.40\.20\.40\.450\.82±0\.1850\.82\\pm 0\.1878\.67±0\.1178\.67\\pm 0\.1179\.10±0\.0779\.10\\pm 0\.070\.168±0\.0020\.168\\pm 0\.0020\.969±0\.0120\.969\\pm 0\.0120\.361±0\.0020\.361\\pm 0\.002

##### Overall sensitivity\.

Table[3](https://arxiv.org/html/2609.28690#A1.T3)varies one hyperparameter at a time around the default setting\(λadv,γ,ρ\)=\(0\.2,0\.5,0\.4\)\(\\lambda\_\{\\mathrm\{adv\}\},\\gamma,\\rho\)=\(0\.2,0\.5,0\.4\)\. All tested configurations with advantage modulation achieve higher mean ACC and F1 and lower Group\-Δ\\DeltaPCR than the unmodulated baseline\. Their ACC and F1 remain within 78\.36–79\.80 and 78\.53–79\.60, respectively, while PCR varies more substantially, from 47\.21% to 53\.86%\. The reported standard deviations indicate limited variation across evaluation seeds for each fixed checkpoint; they do not measure variability across independent training runs\.

##### Future\-deviation discountρ\\rho\.

Settingρ=0\\rho=0restricts modulation to the current turn’s local deviation\. Compared with this setting, the defaultρ=0\.4\\rho=0\.4improves ACC, F1, Group\-Δ\\DeltaPCR, and OT\-DTW, but yields a higher W1\-Turns\. The effects are not monotonic:ρ=0\.1\\rho=0\.1obtains slightly better ACC and Group\-Δ\\DeltaPCR than the default, whereasρ=0\.5\\rho=0\.5achieves higher F1 and lower W1\-Turns\. Among the tested discount values, the default achieves the lowest OT\-DTW, indicating the strongest semantic trajectory alignment under this metric\.

##### Modulation strengthλadv\\lambda\_\{\\mathrm\{adv\}\}\.

Reducingλadv\\lambda\_\{\\mathrm\{adv\}\}to 0\.1 yields the highest ACC and F1 and the lowest Group\-Δ\\DeltaPCR, but increases both W1\-Turns and OT\-DTW relative to the default\. Increasingλadv\\lambda\_\{\\mathrm\{adv\}\}to 0\.3 produces lower ACC and F1 and higher values of all three distance metrics than the default\. These results reveal a trade\-off between outcome agreement and dialogue\-length and semantic alignment\. The default strength provides a compromise, achieving better trajectory fidelity than the smaller value while retaining improved outcome agreement over the unmodulated baseline\.

##### Deviation sensitivityγ\\gamma\.

Reducingγ\\gammafrom 0\.5 to 0\.4 produces similar results, with slightly lower ACC and F1 and slightly higher values of the three distance metrics\. This comparison suggests limited sensitivity to this particular perturbation, although the two tested values do not establish robustness over a broader range ofγ\\gamma\. Overall, the default configuration achieves the lowest OT\-DTW among all tested configurations while improving all five directional metrics over the unmodulated baseline\.

### A\.3Training Dynamics and Convergence Analysis

![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/training_curves/7b/critic_score_mean.png)\(a\)Overall reward score
![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/training_curves/7b/actor_entropy.png)\(b\)Generation entropy
![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/training_curves/7b/response_length_mean.png)\(c\)Mean response length

Figure 5:Training dynamics of TRACER\-7B\. We report the overall reward score, generation entropy, and mean response length during reinforcement learning\. The solid curves show EMA\-smoothed trends, while the faint curves denote the raw values\.![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/training_curves/4b/critic_score_mean.png)\(a\)Overall reward score
![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/training_curves/4b/actor_entropy.png)\(b\)Generation entropy
![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/training_curves/4b/response_length_mean.png)\(c\)Mean response length

Figure 6:Training dynamics of TRACER\-4B\. We report the overall reward score, generation entropy, and mean response length during reinforcement learning\. The solid curves show EMA\-smoothed trends, while the faint curves denote the raw values\.Figures[5](https://arxiv.org/html/2609.28690#A1.F5)and[6](https://arxiv.org/html/2609.28690#A1.F6)present the training dynamics of TRACER\-7B and TRACER\-4B, respectively\. Across both model scales, the overall reward score increases rapidly during the early stage of reinforcement learning and then gradually stabilizes, indicating that the proposed training objective can be effectively optimized\. The convergence pattern is consistent for both 7B and 4B models, suggesting that the proposed framework is not limited to a single model size\.

The generation entropy exhibits a smooth downward trend for both models\. This pattern suggests that the policy gradually becomes more focused as training progresses, while avoiding abrupt entropy collapse\. In other words, the model learns to produce more task\-aligned responses without showing signs of severe mode collapse or unstable policy updates\.

The mean response length also decreases substantially in the early training stage and then remains within a relatively stable range\. This trend is important because it suggests that the reward improvement is not achieved by simply generating increasingly longer responses\. Instead, both models learn to produce more concise and controlled outputs after reinforcement learning\. Overall, these training curves demonstrate the stability and feasibility of the proposed reinforcement learning framework across model scales, with stable reward optimization, gradual policy adaptation, and controlled response length\.

## Appendix BDataset Construction Details

### B\.1Data Source

The dataset used in this study is derived from real\-world merchant–user dialogue logs collected from an online commercial service platform\. The raw corpus contains 36,841 multi\-turn dialogue sessions between merchant\-side intelligent customer service agents and real users, together with corresponding user profile attributes and merchant\-side information\. The dialogues cover multiple real\-world business domains, including healthcare, automotive services, education consulting, and legal consulting\.

Before entering the research pipeline, all data were strictly anonymized and de\-identified\. We removed personally identifiable information and retained only task\-relevant structured attributes, merchant\-side metadata, and dialogue content required for user simulation\. The processed dataset is intended to support user behavior modeling in multi\-turn customer service scenarios while minimizing privacy risks\.

### B\.2User Profile Construction

To support high\-fidelity simulation grounded in user background information, we construct structured user profiles from two complementary perspectives: objective attributes and subjective behavioral priors\.

First, we organize objective user attributes into natural\-language profile descriptions\. Based on real business data, we select three groups of core information\. The first group consists of basic demographic attributes, including age group, gender, education level, marital or parenting status, and geographic location\. The second group contains occupation\-related information\. The third group reflects consumption\-related attributes, including mobile operating system, device price range, city tier, and residential environment\. We then convert these discrete labels into coherent natural\-language background descriptions using predefined templates\. This step provides the user simulator with stable prior conditions while avoiding direct exposure of unnecessary raw attribute fields\.

Second, since objective attributes alone are insufficient to determine a user’s conversational style, task motivation, and behavioral tendency, we use Claude 4\.5\-Sonnet as an annotation model to extract subjective behavioral priors\. Specifically, with final conversion outcomes and other terminal\-state information masked, we analyze early dialogue context and observable user information to extract two types of subjective priors\. The first type is the user’s personality trait, such as impatient and direct, rational and cautious, or cooperative and efficient\. The second type is the initial core intent, which summarizes the user’s primary task goal, key constraints, and expected issue to be resolved at the beginning of the conversation\.

To avoid label leakage, we impose strict constraints during subjective information extraction\. The extracted information is not allowed to include any signal related to the final conversion outcome, such as whether the user eventually provides contact information\. In other words, the initial core intent describes only the user’s demand and task background at the beginning of the conversation, without revealing the final session outcome\. In this way, the user profile serves as a prior condition for simulation while avoiding direct leakage of the target decision label\.

### B\.3Dialogue Augmentation and Structured Trajectory Construction

The original real\-world dialogue logs mainly contain users’ surface\-level natural\-language responses, but lack explicit annotations of the latent decision states underlying those responses\. This leads to a key limitation: if a model is trained only on raw user responses, it primarily learns*what the user says*, but receives limited supervision regarding*why the user responds in that way*or*how the user’s intent evolves over the course of a multi\-turn interaction*\. This limitation is closely aligned with the motivation of this work: high\-fidelity user simulation requires not only linguistic\-style consistency, but also consistency in behavioral logic and decision trajectories\.

To address this issue, we perform structured augmentation for each user turn, transforming the original user response into a training instance with intermediate behavioral\-state annotations\. Specifically, for each user utterance in a multi\-turn dialogue, we use Claude 4\.5\-Sonnet to supplement two additional fields: a rationale annotation and an intent level\.

The rationale annotation describes the possible task motivation, contextual reaction, or decision reason underlying the current user response\. It is important to note that this field should not be interpreted as the user’s actual private mental state\. Instead, it is an annotation\-based proxy inferred from the dialogue context for modeling latent decision states\. Its role is to provide the model with a learnable intermediate representation that bridges surface\-level linguistic expression and deeper behavioral consistency\.

The intent level characterizes the user’s current willingness to continue the interaction, advance the task, or reach a conversion\-oriented outcome\. It is represented as a discrete score ranging from 0 to 5\. Specifically, a score of 0 indicates that the user exits the current interaction or explicitly refuses to proceed, while a score of 5 indicates that the user explicitly provides contact information or reaches a positive commitment state\. Higher scores indicate stronger willingness to continue the interaction, seek further consultation, or move toward conversion\.

After augmentation, each user turn is organized into the following structured format:

```
<think>...</think>
<intent_level>...</intent_level>
<reply>...</reply>
```

In this format, the original user response is preserved, while the intent state is explicitly introduced as a key intermediate variable\. As a result, the user simulation task is no longer formulated as pure response generation, but as joint modeling of intent states and natural\-language responses\. This structured annotation further provides necessary supervision signals for trajectory\-level reward modeling and turn\-level credit assignment in the subsequent reinforcement learning stage\.

Since the subjective priors, personality traits and initial core intents, are inferred by an LLM rather than directly observed, we further validate their fidelity through a dedicated human evaluation\. We draw a stratified sample of 300 dialogues covering diverse intents, dialogue lengths, and user segments, and recruit three trained annotators to independently re\-annotate each instance under a blind protocol \(i\.e\., without exposure to Claude’s outputs\)\. Information about the annotators can be found in Appendix[G\.1](https://arxiv.org/html/2609.28690#A7.SS1)\. The evaluation results are summarized in Table[4](https://arxiv.org/html/2609.28690#A2.T4)\.

Table 4:Evaluation of LLM\-inferred subjective priors
### B\.4Data Cleaning

To ensure data quality and reduce the influence of noisy samples on model training, we design a rigorous data cleaning pipeline\.

First, we remove sessions that are corrupted, entirely empty or meaningless, clearly unrelated to the merchant’s business, or explicitly identified by the user as accidental entries\. This prevents the model from learning noise patterns unrelated to the target user simulation task\. Second, we constrain dialogue length by retaining only sessions with no fewer than 3 turns and no more than 15 turns\. We retain sessions containing 3 to 15 user turns to ensure sufficient multi\-turn context while limiting excessive contextual redundancy and computational cost\. This constraint balances interaction completeness and training stability\.

Finally, we perform annotation quality control after structured augmentation\. We verify schema completeness and objective consistency between observable user utterances and deterministic behavioral labels, such as whether explicit contact information is present\. When an LLM\-inferred rationale or intermediate intent annotation conflicts with the observable dialogue content, the annotation is flagged for review or regenerated rather than using the inferred behavioral logic as a basis for excluding the original session\. We do not remove sessions because their annotated intent trajectories appear abrupt, flat, non\-monotonic, or causally implausible, as such patterns may reflect genuine behavioral variability, unobserved factors, or annotation uncertainty\.

After the above cleaning and quality\-control procedures, we obtain 11,906 multi\-turn dialogue sessions, comprising 8,040 training sessions and 3,866 test sessions, with no session overlap between the two sets\. The training data are further divided into 6,312 supervised fine\-tuning sessions and 1,728 reinforcement learning sessions, which are mutually disjoint at the session level\. The test sessions are organized into 762 reference groups based on shared merchant identity and similar user personality traits and initial needs\. Grouping is used only for evaluation, while training retains paired\-session supervision\. The session\-level conversion rate in the test set is 51\.24%\.

## Appendix CLimitation

We acknowledge several limitations of our work that point to promising directions for future research\.

Generalization to Open\-Ended Social Dialogues\. While our method demonstrates strong out\-of\-domain generalization across task\-oriented scenarios, its applicability to purely open\-ended, unconstrained social dialogues remains to be thoroughly investigated\. Such conversations typically lack explicit task objectives and well\-defined success criteria, making both reward modeling and credit assignment substantially more challenging\. We leave a systematic study of this setting to future work\.

##### Broader Applicability to Agentic Scenarios\.

Our evaluation focuses on multi\-turn dialogue\. The proposed advantage modulation mechanism may also be useful in other sequential interaction settings where feedback is available at the episode level and intermediate behavior can be compared with reference trajectories\. Potential applications include tool use, web navigation, and long\-horizon planning\. However, extending the method to these settings requires suitable trajectory representations and deviation measures, particularly for heterogeneous action spaces\. Whether reference\-based deviation signals improve policy optimization in these environments remains an empirical question for future work\.

We also note potential negative societal impacts: a high\-fidelity user simulator aligned with real users could be misused to optimize manipulative or deceptive persuasion strategies\.

## Appendix DDetails for the Turing\-style Evaluation

### D\.1Reasons LLM\-Simulated Users Are Judged Non\-Human

Table[5](https://arxiv.org/html/2609.28690#A4.T5)presents a fine\-grained analysis of why LLM\-simulated users were judged as non\-human across three diagnostic dimensions:*Linguistic Style*,*Interaction Rhythm*, and*Persona & Logic*\. Lower percentages indicate better human\-likeness\. Each dimension includes multiple subcategories capturing specific cues, such as lack of typos, overly neat sentence structures, excessive cooperativeness, or emotional discontinuities\. Notably, the TRACER\-7B model consistently demonstrates lower percentages, reflecting improved human\-like imperfections and interaction patterns compared to baselines\. Information about the annotators can be found in Appendix[G\.1](https://arxiv.org/html/2609.28690#A7.SS1)\.

Table 5:Fine\-grained reasons why LLM\-simulated users are judged as non\-human\. Lower percentages indicate better performance\.We expand the three diagnostic dimensions summarized in the main text \(Table[5](https://arxiv.org/html/2609.28690#A4.T5)\)\.

##### Linguistic style: the too clean problem\.

Human\-imperfection cues, such as typos, fillers, incomplete utterances, and colloquial punctuation, are prevalent in real user dialogues but are largely absent in LLM outputs\. For TRACER\-7B, such cues are extremely rare, occurring in only 3\.0% of samples\. Overly neat sentence structures, another signal of unnatural text regularity, appear in 17\.0% of TRACER\-7B outputs, while overly formal or written register remains at 4\.0%\. These values contrast sharply with Gemini and Doubao models, where lack of human imperfection ranges from 93\.8% to 98\.8% and overly neat structures range from 41\.2% to 50\.6%, indicating that TRACER\-7B better captures the variability and imperfection characteristic of real users\.

##### Interaction rhythm and informativeness\.

In real customer\-service dialogues, users tend to be terse and goal\-driven, revealing information incrementally\. LLM simulators, by contrast, often exhibit information overload and excessive cooperativeness\. TRACER\-7B, however, shows minimal information overload at 0\.8% and a low rate of mechanical response patterns at 4\.4%\. Its excessive cooperativeness is moderate at 27\.2%, and super\-human compliance occurs in only 0\.6% of cases\. Compared with Gemini and Doubao models, which exhibit information overload between 78\.2%–87\.2% and mechanical patterns up to 82\.4%, TRACER\-7B more faithfully reproduces the incremental, negotiation\-based tempo of real user–agent interactions\.

##### Persona and logic\.

At the semantic level, persona inconsistencies are a key defect in many LLMs\. For TRACER\-7B, out\-of\-character behavior is minimal \(1\.4%\), with zero emotional discontinuity, negligible lack of situation sense \(0\.2%\), and minimal over\-politeness \(0\.2%\)\. In contrast, Gemini and Doubao models show much higher rates, with persona inconsistency ranging from 38\.4% to 46\.2% and over\-politeness reaching 14\.0%\. These results indicate that TRACER\-7B maintains strong alignment with the assigned user persona and exhibits high\-level coherence across dialogue turns\.

##### Summary\.

The bottleneck of current user simulators is not high\-level reasoning but \(i\) the lack of low\-level linguistic noise and \(ii\) the inability to reproduce strategic, under\-informative, and sometimes uncooperative interaction rhythms\. This directly motivates the design choices of our framework, which explicitly models linguistic imperfection and turn\-level information control\.

## Appendix EOut\-of\-Domain Generalization Evaluation Details

To evaluate TRACER on out\-of\-domain \(OOD\) dialogues, we use the CSC\-Conv dataset\([Zhu et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib33)\), an open\-source corpus of real user–agent conversations in the financial customer\-service domain, which is completely unseen during training\.

Since CSC\-Conv does not contain structured user profiles, we follow the approach in Appendix[B\.2](https://arxiv.org/html/2609.28690#A2.SS2)to extract user initial intent and behavioral characteristics using the LLMs based on each dialogue session\. We filter out dialogues in which the user’s issue resolution is ambiguous\. From the remaining dialogues, we randomly sample 1,000 sessions for evaluation, with the ratio of resolved to unresolved issues being 689:311\. The original dataset is skewed toward resolved issues, hence this sampling ensures sufficient representation of unresolved cases\.

## Appendix FDynamic Marketing Benchmark Details

### F\.1Response Quality Evaluation

Following the same LLM\-as\-a\-judge prompt as in\([Zhu et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib33)\), we evaluate the response quality of the customer service model in multi\-turn dialogues\. The prompt used for this evaluation is illustrated in Figure[7](https://arxiv.org/html/2609.28690#A6.F7)\. To mitigate potential self\-preference bias, we adopt gpt\-oss\-120B as the judge model\.

![Refer to caption](https://arxiv.org/html/2609.28690v1/llm_as_judge.png)Figure 7:Prompt for the LLM\-as\-Judge evaluation is adapted from\([Zhu et al\., 2026](https://arxiv.org/html/2609.28690#bib.bib33)\)\.
### F\.2Human\-Referenced Strategy\-Response Evaluation

Table[6](https://arxiv.org/html/2609.28690#A6.T6)reports the strategy\-level results of the 200\-conversation study in Section[6\.2](https://arxiv.org/html/2609.28690#S6.SS2)\. Each cell contains the number of simulated changes matching the human\-expected direction, divided by the number of strategy instances assigned that reference direction\. The percentages are therefore class\-conditional recall values\. They do not indicate the proportion of all occurrences of a strategy that increase or decrease intent\.

Table 6:Agreement between simulated and human\-expected intent\-change directions\. Each cell reports matched / reference strategy instances \(recall\)\. The pooled row aggregates counts across strategy categories\.The pooled row sums the corresponding numerators and denominators across strategy categories, yielding 449 strategy\-instance counts from 200 conversations\. These counts should not be interpreted as 449 independent conversations\. Coverage is uneven: urgency contains only six instances, making its percentages particularly sensitive to individual observations\. We report the category results descriptively without claiming statistically significant differences between strategies\.

The study measures agreement with annotator expectations in sampled conversations\. It does not isolate the causal effects of changing a strategy under identical conditions, nor does it directly validate real\-user terminal outcomes\. For example, high decrease recall for empathy means that the simulator recovers decreases where annotators expect them; it does not imply that empathy generally reduces user intent\.

The direction recalls pool observations across assistant models\. They provide evidence within the sampled interactions, but do not quantify performance separately for each assistant or establish generalization to unseen assistant policies\. The study also contains no simulator ablation that would attribute strategy\-response agreement specifically to trajectory alignment\.

### F\.3Results

Table 7:Detailed results of mainstream LLMs on DM\-Bench\. Conversion Rate\(CR\) and Conversion Turns\(CT\) measure outcome\-oriented performance, while Response Quality\(RQ\) is evaluated across multiple conversational dimensions\.Table 8:Spearmanρ\\rhocorrelation between Conversion Rate \(CR\) and Response Quality \(RQ\) metrics on DM\-Bench\.Table[7](https://arxiv.org/html/2609.28690#A6.T7)and Table[8](https://arxiv.org/html/2609.28690#A6.T8)summarize the performance of seven contemporary LLMs on DM\-Bench\.

## Appendix GHuman Annotation

We detail the human annotation procedures used in our study, including annotator qualifications, remuneration, quality control, and the annotation guidelines used across different experiments\.

### G\.1Annotator Details

The annotation process was carried out by a team of 10 trained annotators and overseen by a dedicated quality reviewer\. All annotators hold at least an associate degree and possess strong literacy and comprehension skills, enabling them to quickly understand AI annotation guidelines, operational procedures, and task\-specific rules\. Each annotator has approximately one year of experience in AI data annotation, with hands\-on experience in routine labeling, data verification, and issue documentation\. Annotators were compensated at a rate of 50 RMB per hour for their work\. The quality reviewer performed random sampling and consistency checks to assess annotation quality and inter\-annotator agreement, thereby helping ensure the reliability of the annotations\.

### G\.2Turing\-Style Evaluation

The human annotation for dialogue evaluation followed a structured two\-step procedure designed to distinguish real\-user dialogues from AI\-simulated dialogues\. Annotators were provided with the corresponding user profile and asked to make judgments based on multiple dimensions, including linguistic naturalness, information presentation, emotional expression, profile consistency, situational awareness, and conversational imperfections\. The full annotation guide is provided below\.

Guide to Real Human Dialogue Recognition##### Step 1: Dialogue Assessment Annotators were shown two dialogue excerpts: one originating from a real user interacting with customer service, and one from an AI\-simulated user interacting with customer service\. Considering the provided user profile, annotators were asked to determine which dialogue was more likely to have been produced by a human\.Core Evaluation Dimensions:1\.Language Naturalness\.Evaluate whether the language resembles natural human conversational patterns\.•Human cues:colloquial expressions, filler words \(e\.g\., “um”, “oh”, “well”\), typographical or input errors, inconsistent sentence lengths, abbreviations, dialects, or internet slang\.•AI cues:overly formal or structured language, perfect grammar and punctuation, uniform sentence lengths, or neutral and textbook\-style wording\.2\.Information Expression Style\.Assess the pacing and completeness of the information provided\.•Humans tend to mention one point at a time and may require follow\-up prompts before providing additional information\.•AI tends to provide “all\-in\-one” responses that simultaneously cover the issue, background, request, and related details\.3\.Emotion and Attitude Expression\.Evaluate whether emotional expressions appear natural and consistent with the user profile\.•Humans may express impatience, complaints, urgency, gratitude, politeness, sudden disengagement, or repeated emphasis\.•AI\-related cues may include overly flat or exaggerated emotion, abrupt emotional shifts, or emotional expressions inconsistent with the profile \(e\.g\., a “short\-tempered” user profile using consistently excessive politeness\)\.4\.Profile Consistency\.Assess whether the dialogue content and interaction style are consistent with the provided user profile\.•Language style:Language use should be broadly consistent with characteristics such as age and educational background\. For example, older users may be less likely to use internet slang, whereas highly educated users may use more precise expressions\.•Knowledge level:The user’s demonstrated knowledge should be consistent with the stated profession or background, avoiding unexplained use of highly specialized terminology or implausibly simple questions\.•Behavioral habits:Behavioral tendencies described in the profile should be reflected in the interaction when relevant\. For example, a profile described as “impatient” may display greater urgency during the dialogue\.5\.Common Sense and Situational Awareness\.Evaluate whether the user’s behavior is plausible within a real\-world interaction\.•Humans may mention concrete situational details \(e\.g\., “just got off work”, “my child is nearby”, or “my phone battery is low”\)\.•Humans may express confusion during complex tasks \(e\.g\., “what does this mean?” or “I don’t know how to do it”\)\.•AI may behave like an “all\-knowing user”, immediately understanding every operation and rapidly providing all requested information\.6\.Imperfections and Anomalies\.Human dialogues are often imperfect, whereas AI\-generated dialogues may appear unusually smooth\.•Common human imperfections that may signal authenticity include typos, accidentally sent messages, segmented responses, irrelevant answers, abrupt silence, incomplete statements followed by later clarification, or ambiguous pronoun use\. ##### Evaluation Notes 1\.Do not rely solely on grammatical correctness\. AI\-generated language is often fluent, whereas conversational naturalness is more informative than grammatical perfection alone\.2\.Be aware of “reverse disguise\.” High\-quality AI systems may intentionally introduce typos, filler words, or other human\-like imperfections\.3\.Dialogue length is not a definitive cue\. A long dialogue does not necessarily indicate a human user, and a short dialogue does not necessarily indicate an AI\-generated user\.4\.Avoid confirmation bias\. Evaluate consistency across the entire dialogue rather than relying on a single utterance\.5\.Key indicators of potentially AI\-like behavior include:•excessive politeness \(e\.g\., “Thank you very much for your reply” or “Sorry for the trouble”\);•highly structured expression \(e\.g\., “First…, second…, finally…”\);•unusually high information density, where nearly every sentence provides task\-relevant information; and•emotional content that appears disconnected from the conversational context\.6\.System messages indicating user inactivity or dropout should not be treated as evidence that the dialogue was AI\-generated\. ##### Step 2: Reasoning Annotation After completing the human\-versus\-AI judgment, annotators were asked to select the most applicable “non\-human reason\(s\)” from a set of predefined categories\. Multiple selections were allowed\.Category A: Language Habits and Style•Overly formal:Language resembles manuals or official documents and lacks a natural conversational tone\.•Highly structured sentences:Sentence structure is unusually uniform or logically organized for a real\-time conversation\.•Lack of human imperfection:Language contains consistently perfect punctuation and grammar, with no typos, colloquial expressions, or conversational fillers\.Category B: Interaction Rhythm and Information Quantity•Information dump:A single response contains an unusually large amount of key information and lacks typical human conversational pacing\.•Over\-cooperative:The user completes complex tasks or provides extensive information without requiring clarification or guidance, behaving like an “expert user\.”•Excessive comprehension:The user immediately understands ambiguous expressions or instructions without requesting clarification\.•Mechanical replies:Responses are excessively long, repetitive, redundant, or formulaic\.Category C: Profile and Logical Consistency•Out\-of\-character \(OOC\):The user’s behavior is inconsistent with characteristics specified in the user profile, such as age, profession, background, or personality\.•Emotion mismatch:Emotional expression changes abruptly or appears inconsistent with the conversational context or user profile\.•Over\-polite:The user employs courteous expressions at a frequency or intensity that appears unusual for ordinary customer\-service interactions\.•Lack of situational awareness:The dialogue focuses exclusively on task completion while ignoring contextual or environmental factors that a real user might naturally mention\.

### G\.3Human Annotation for LLM\-Inferred User Intents and Rationale

##### Intent Scoring

The subjective priors, personality traits, and initial core intents are inferred by a LLM rather than directly observed\. To validate these inferred attributes, we perform a dedicated human evaluation\. Annotators independently assess dialogues sampled from a stratified set of 300 conversations covering diverse intents, dialogue lengths, and user segments\. All annotations are conducted under a blind protocol\.

Each annotator is responsible for completing the following two tasks for every dialogue:

Objective:Evaluate the likelihood that the user intended a particular action or goal at each turn of the dialogue, conditioned on the dialogue history and the user profile\.

Instructions:

Human Annotation Guideline: User Intent Level Scoring##### Task Goal Your task is to read the current user utterance in a multi\-turn dialogue and assign an intent level score from 0 to 5\.The score reflects how strong the user’s current willingness is to:continue communication, provide information, move toward a transaction, or complete conversion\.Important:Only annotate the current user utterance\. Do not output explanations, reasoning, action labels, or any other fields\. ##### Intent Level Definitions 0 — Churn / Conversation EndedAssign 0 when the user gives an empty reply, indicating the conversation has ended\.Example:User: ""1 — Very Low Intent / Negative or PerfunctoryAssign 1 when the user shows clear rejection, impatience, distrust, or extremely low willingness to communicate\.Typical cases include:Direct rejection of the offer; Expressing that the product/service is not needed; Suspecting a scam; Complaining about price; Showing impatience toward the agent; Providing meaningless or perfunctory responses\.Examples:“Not needed\.” “Too expensive\.” “Is this a scam?” “Whatever\.” “Later\.” “Oh\.” “Hmm\.” “Just looking\.”2 — Low Intent / Passive Basic ResponseAssign 2 when the user provides a minimal, passive answer without actively moving the conversation forward\.Characteristics:Cooperation is minimal; Typically very short factual answers; User is still responsive but not proactive\.Important rule:If the user asks a question or includes a question mark, do not assign 2; consider level 3 or above instead\.Examples:“Beijing\.” “SUV\.” “No\.” “Gasoline car\.” “Around 100,000\.”3 — Medium Intent / Active Response or Simple QuestionAssign 3 when the user actively communicates or asks a simple question related to the product/service\.Situations include:User gives more cooperative or detailed responses than a passive reply\.Examples:“I don’t have any brand requirements\.” “Any model is fine as long as the price is suitable\.” “I mainly want something for commuting\.” “I’m just comparing options for now\.” User asks a short/basic question\.Examples:“Do you have used cars?” “Is this a gasoline car?” “Can I pay in installments?” “Is the car still available?” “What models do you have?”4 — Strong Intent / Concrete Need or Transaction\-Oriented QuestionAssign 4 when the user shows strong interest or enters a concrete transaction scenario\.Typical behaviors include:Asking for agent’s contact information; Requesting to add WeChat or get a phone number; Providing detailed personal needs or constraints; Asking about vehicle appraisal, trade\-in, store address, delivery process, loan calculation, or installment details\.Examples:“What’s your WeChat?” “Add me on WeChat\.” “Send me your phone number\.” “I have a 2014 car and want to trade it in for this model\.” “Where is your store?” “How much can my old car be valued at?” “How much is the monthly payment?” “What is the delivery process?”Note:If the user only asks for contact info without providing their own, assign 4 \(not 5\)\.5 — Conversion Achieved / Contact Information ProvidedAssign 5 when the user provides concrete contact information in the current utterance\.Valid contact information includes:Phone number; WeChat ID; QQ number; Any clear string intended as contact information\.Examples:“My phone number is 138xxxx8888\.” “Add my WeChat: abc123\.” “Contact me at 186xxxx6666\.” “My WeChat ID is carbuyer2024\.”Important rule:Only assign 5 if actual contact information is provided\. If the user only says “I can give you my phone number” or “Let’s add WeChat” without giving the number/ID, assign 4 instead\.

##### Rationale Plausibility Rating

Objective:Assess whether the LLM\-generated rationale explaining the user intent is reasonable given the dialogue history and user profile\.

Instructions:

Guide to Rationale Plausibility Rating##### Task Goal Your task is to evaluate the plausibility of a rationale for a user’s intent in a multi\-turn dialogue\.Input includes the full dialogue history up to the current turn and the user profile description\. The user profile contains: Objective attributes: age, occupation, demographics Subjective behavioral priors ##### Procedure For each turn, assign a plausibility score that reflects the probability or confidence that the rationale aligns with the user’s intent\. Scores between 0 and 5, if specified Take context from previous turns into account, as user intent may evolve or persist\. Use the user profile to adjust expectations\. If uncertain, annotate conservatively and optionally provide notes\. ##### Considerations Take context from previous turns into account, since a user’s intent may evolve or persist across multiple turns\. Use the user profile to adjust expectations\. If uncertain, annotate conservatively and provide notes if needed\. ##### Rating Description 5 Completely reasonable: fully consistent with dialogue history and user profile 4 Mostly reasonable: minor inconsistencies or omissions, but overall plausible 3 Neutral: some support from dialogue and profile, but substantial ambiguity or partial mismatch 2 Mostly unreasonable: contains incorrect assumptions or contradicts dialogue/profile context 1 Completely unreasonable: no alignment with dialogue history or user profile; misleading rationale

## Appendix HImplementation Details of Dynamic Time Warping for Intent Trajectory Alignment

### H\.1Motivation

In multi\-turn user simulation, the intent trajectory generated by the simulator is not necessarily synchronized with the real user trajectory at the turn level\. Even when the simulator follows an overall intent evolution pattern similar to that of a real user, the two trajectories may still differ in local pacing\. For example, a real user may express a need and advance their intent within a single turn, whereas the simulator may take two turns to complete the same transition\. Conversely, the simulator may compress into one turn an intent transition that unfolds over two adjacent turns in the real dialogue\. Therefore, a strict turn\-by\-turn alignment, which forces the simulator’stt\-th turn to correspond to the real user’stt\-th turn, may incorrectly penalize behaviorally plausible trajectories as misaligned\.

To mitigate this issue, we adopt Dynamic Time Warping \(DTW\) to perform nonlinear alignment between simulated and real intent trajectories\. DTW has been widely used for speech recognition and time\-series matching, where it enables elastic alignment between two sequences that share similar global shapes but may differ in phase, progression speed, or sequence length\([Sakoe & Chiba, 1978](https://arxiv.org/html/2609.28690#bib.bib17);[Myers et al\., 1980](https://arxiv.org/html/2609.28690#bib.bib10)\)\. Its main advantage is that it reduces the influence of temporal shifts and local distortions on sequence similarity measurement through elastic transformation, while allowing the globally optimal alignment to be computed via dynamic programming inO⁡\(N​M\)O\(NM\)time\([Senin, 2008](https://arxiv.org/html/2609.28690#bib.bib18)\)\.

In our task, the goal is not to require the user simulator to reproduce the real dialogue word by word or turn by turn\. Instead, we aim to encourage the simulator to generate an intent evolution path that is consistent with real user behavioral logic\. DTW is therefore well suited for measuring the global consistency between simulated and real intent trajectories while allowing reasonable local stretching or compression across dialogue turns\.

### H\.2General Formulation of Dynamic Time Warping

Dynamic Time Warping is a nonlinear alignment method for measuring the similarity between two sequences\. Unlike pointwise comparison, DTW allows local stretching and compression along the temporal axis, making it suitable for sequences with different lengths, different progression speeds, or local phase shifts\. Classical DTW has been widely applied to time series, speech recognition, handwriting recognition, and motion sequence matching\. Its core idea is to find a warping path with the minimum cumulative cost, such that two sequences are optimally aligned while preserving their original temporal order\. The optimal path can be solved efficiently using dynamic programming, with a time complexity ofO⁡\(N​M\)O\(NM\)\.

![Refer to caption](https://arxiv.org/html/2609.28690v1/figure/dtw.png)Figure 8:Illustration of pointwise alignment and Dynamic Time Warping \(DTW\)\. In the left panel, strict pointwise alignment forces pointaato be matched with pointbbaccording to their temporal indices, although the more plausible counterpart ofaais the locally shifted pointb′b^\{\\prime\}\. The right panel illustrates DTW\-based nonlinear alignment, where one\-to\-one, one\-to\-many, and many\-to\-one correspondences are allowed under temporal\-order constraints\.Figure[8](https://arxiv.org/html/2609.28690#A8.F8)provides an intuitive comparison between strict pointwise alignment and DTW\-based nonlinear alignment\. In the left panel, strict pointwise comparison forces pointaain the upper sequence to be aligned with pointbbin the lower sequence because they share the same temporal index\. However, due to local phase shifts or different progression speeds, the more plausible counterpart ofaamay beb′b^\{\\prime\}rather thanbb\. In this case, pointwise alignment may overestimate the local discrepancy between the two sequences\. In contrast, DTW allows such locally shifted patterns to be matched through an order\-preserving nonlinear alignment path, as illustrated in the right panel\. This motivates the formal definition of the warping path and its constraints below\.

Given two general sequences

X=\(x1,x2,…,xN\),X=\(x\_\{1\},x\_\{2\},\\dots,x\_\{N\}\),\(8\)and

Y=\(y1,y2,…,yM\),Y=\(y\_\{1\},y\_\{2\},\\dots,y\_\{M\}\),\(9\)wherexix\_\{i\}andyjy\_\{j\}may be scalars, vectors, or other comparable sequence elements, we first define a local cost function

d⁡\(xi,yj\):𝒳×𝒴→ℝ≥0,d\(x\_\{i\},y\_\{j\}\):\\mathcal\{X\}\\times\\mathcal\{Y\}\\rightarrow\\mathbb\{R\}\_\{\\geq 0\},\(10\)which measures the local discrepancy betweenxix\_\{i\}andyjy\_\{j\}\. Based on this local cost function, we construct a local cost matrixC∈ℝN×MC\\in\\mathbb\{R\}^\{N\\times M\}, where

C⁡\(i,j\)=d⁡\(xi,yj\)\.C\(i,j\)=d\(x\_\{i\},y\_\{j\}\)\.\(11\)The objective of DTW is to find a warping path

P=\(p1,p2,…,pK\),P=\(p\_\{1\},p\_\{2\},\\dots,p\_\{K\}\),\(12\)where each path elementpk=\(ik,jk\)p\_\{k\}=\(i\_\{k\},j\_\{k\}\)indicates that elementxikx\_\{i\_\{k\}\}in sequenceXXis aligned with elementyjky\_\{j\_\{k\}\}in sequenceYY\. A valid warping path typically satisfies the following three constraints\.

First, the boundary constraint requires the path to start from the beginning of both sequences and end at the end of both sequences:

p1=\(1,1\),pK=\(N,M\)\.p\_\{1\}=\(1,1\),\\quad p\_\{K\}=\(N,M\)\.\(13\)
Second, the monotonicity constraint ensures that the alignment does not violate the original temporal order of either sequence:

i1≤i2≤⋯≤iK,j1≤j2≤⋯≤jK\.i\_\{1\}\\leq i\_\{2\}\\leq\\dots\\leq i\_\{K\},\\quad j\_\{1\}\\leq j\_\{2\}\\leq\\dots\\leq j\_\{K\}\.\(14\)
Third, the step\-size constraint restricts each transition to a local move\. A common setting is

pk\+1−pk∈\{\(1,0\),\(0,1\),\(1,1\)\}\.p\_\{k\+1\}\-p\_\{k\}\\in\\\{\(1,0\),\(0,1\),\(1,1\)\\\}\.\(15\)
These three moves correspond to one\-to\-many, many\-to\-one, and one\-to\-one local alignments, respectively\. Together, these constraints allow DTW to provide flexible nonlinear alignment while preserving the internal temporal order of both sequences\.

To compute the optimal path efficiently, DTW constructs a cumulative cost matrixDD, whereD⁡\(i,j\)D\(i,j\)denotes the minimum cumulative cost required to align the prefix\(x1,…,xi\)\(x\_\{1\},\\dots,x\_\{i\}\)with the prefix\(y1,…,yj\)\(y\_\{1\},\\dots,y\_\{j\}\)\. At the boundary, we setD⁡\(1,1\)=C⁡\(1,1\)D\(1,1\)=C\(1,1\)and initialize the first row and first column cumulatively\. For internal positions, the recurrence is

D⁡\(i,j\)=C⁡\(i,j\)\+min⁡\{D⁡\(i−1,j\),D⁡\(i,j−1\),D⁡\(i−1,j−1\)\}\.D\(i,j\)=C\(i,j\)\+\\min\\\{D\(i\-1,j\),D\(i,j\-1\),D\(i\-1,j\-1\)\\\}\.\(16\)
The DTW alignment cost between two sequences is then defined as the minimum cumulative cost over all valid warping paths:

DTW⁡\(X,Y\)=min⁡∑\(i,j\)∈PP⁡d⁡\(xi,yj\)\.\\mathrm\{DTW\}\(X,Y\)=\\min\_\{P\}\\sum\_\{\(i,j\)\\in P\}d\(x\_\{i\},y\_\{j\}\)\.\(17\)
Thus, DTW does not requireXXandYYto have the same length, nor does it require theii\-th element of one sequence to align with theii\-th element of the other\. Instead, it searches for the most plausible nonlinear correspondence between the two sequences under a globally optimal alignment path\.

### H\.3DTW\-based Intent Trajectory Alignment in Our Method

In this work, we formulate multi\-turn user simulation as a conditional sequential decision\-making problem\. Given static user variablesuuand a dynamic interaction contexthth\_\{t\}, the simulator policy generates a latent intent state and a natural\-language response at each turn\. We denote the real user trajectory as

τ=\{\(zt,yt\)\}t=1T,\\tau=\\\{\(z\_\{t\},y\_\{t\}\)\\\}\_\{t=1\}^\{T\},\(18\)and the simulator\-generated trajectory as

τ^=\{\(z^t,y^t\)\}t=1T∗\.\\hat\{\\tau\}=\\\{\(\\hat\{z\}\_\{t\},\\hat\{y\}\_\{t\}\)\\\}\_\{t=1\}^\{T^\{\*\}\}\.\(19\)
Here,ztz\_\{t\}andz^t\\hat\{z\}\_\{t\}denote the latent intent states of the real user and the simulator at turntt, respectively, whileyty\_\{t\}andy^t\\hat\{y\}\_\{t\}denote the corresponding natural\-language responses\.

Our DTW alignment is not directly applied to the full natural\-language response sequences\. Instead, it is applied to the intent trajectories extracted from multi\-turn interactions\. Specifically, we define the real user’s intent trajectory as

Z=\(z1,z2,…,zT\),Z=\(z\_\{1\},z\_\{2\},\\dots,z\_\{T\}\),\(20\)and the simulator\-generated intent trajectory as

Z^=\(z^1,z^2,…,z^T∗\)\.\\hat\{Z\}=\(\\hat\{z\}\_\{1\},\\hat\{z\}\_\{2\},\\dots,\\hat\{z\}\_\{T^\{\*\}\}\)\.\(21\)
In our data, each intent state is discretized into an ordered level from 0 to 5:

z^i,zj∈\{0,1,2,3,4,5\}\.\\hat\{z\}\_\{i\},z\_\{j\}\\in\\\{0,1,2,3,4,5\\\}\.\(22\)
These intent levels have a clear ordinal structure: smaller numerical differences indicate more similar user states, while larger differences indicate stronger intent deviation\. Therefore, we instantiate the local cost function in DTW as the absolute difference between the simulated intent level and the real intent level:

d⁡\(z^i,zj\)=\|z^i−zj\|\.d\(\\hat\{z\}\_\{i\},z\_\{j\}\)=\|\\hat\{z\}\_\{i\}\-z\_\{j\}\|\.\(23\)
This design preserves the ordinal structure of intent states\. For example, the discrepancy between intent levels 3 and 4 is smaller than that between intent levels 1 and 5\. Compared with a binary mismatch indicator, the absolute\-difference cost better captures gradual shifts in user intent evolution\.

Based on this local cost function, we construct a local cost matrixC∈ℝT∗×TC\\in\\mathbb\{R\}^\{T^\{\*\}\\times T\}:

C⁡\(i,j\)=d⁡\(z^i,zj\)\.C\(i,j\)=d\(\\hat\{z\}\_\{i\},z\_\{j\}\)\.\(24\)We then construct a cumulative cost matrixD∈ℝT∗×TD\\in\\mathbb\{R\}^\{T^\{\*\}\\times T\}, whereD⁡\(i,j\)D\(i,j\)denotes the minimum cumulative cost required to align the simulated intent prefix\(z^1,…,z^i\)\(\\hat\{z\}\_\{1\},\\dots,\\hat\{z\}\_\{i\}\)with the real intent prefix\(z1,…,zj\)\(z\_\{1\},\\dots,z\_\{j\}\)\.

For dynamic programming, we initialize the boundary as follows:

D⁡\(1,1\)=C⁡\(1,1\),D\(1,1\)=C\(1,1\),\(25\)D\(i,1\)=C\(i,1\)\+D\(i−1,1\),i=2,…,T∗,D\(i,1\)=C\(i,1\)\+D\(i\-1,1\),\\quad i=2,\\dots,T^\{\*\},\(26\)D\(1,j\)=C\(1,j\)\+D\(1,j−1\),j=2,…,T\.D\(1,j\)=C\(1,j\)\+D\(1,j\-1\),\\quad j=2,\\dots,T\.\(27\)The first row and first column correspond to the case where one trajectory prefix contains only a single element, so the other trajectory can only be aligned to it through consecutive local stretching\.

After boundary initialization, fori\>1i\>1andj\>1j\>1, the cumulative cost matrix is computed as

D⁡\(i,j\)=C⁡\(i,j\)\+min⁡\{D⁡\(i−1,j\),D⁡\(i,j−1\),D⁡\(i−1,j−1\)\}\.D\(i,j\)=C\(i,j\)\+\\min\\\{D\(i\-1,j\),D\(i,j\-1\),D\(i\-1,j\-1\)\\\}\.\(28\)The three transitions correspond to one\-to\-many, many\-to\-one, and one\-to\-one local alignment between the simulated and real trajectories\. Finally, the DTW cumulative cost between the two intent trajectories is defined as

DDTW​\(Z^,Z\)=min⁡∑\(i,j\)∈PP⁡d⁡\(z^i,zj\),D\_\{\\mathrm\{DTW\}\}\(\\hat\{Z\},Z\)=\\min\_\{P\}\\sum\_\{\(i,j\)\\in P\}d\(\\hat\{z\}\_\{i\},z\_\{j\}\),\(29\)wherePPdenotes a valid warping path\. This path preserves the temporal order within both trajectories while allowing local turn\-level stretching or compression\.

##### DTW\-based Trajectory Plausibility Reward\.

The above definition corresponds directly to the DTW alignment costDDTW\(z^1:T∗,z1:T\)D\_\{\\mathrm\{DTW\}\}\(\\hat\{z\}\_\{1:T^\{\*\}\},z\_\{1:T\}\)used in the trajectory plausibility rewardRtrajR\_\{\\mathrm\{traj\}\}in Section[4\.2](https://arxiv.org/html/2609.28690#S4.SS2.SSS0.Px2)\. Specifically, the trajectory plausibility reward is defined as

Rtraj=exp\(−α⋅DDTW\(z^1:T∗,z1:T\)\+β\|T∗−T\|T\)⋅η\.R\_\{\\mathrm\{traj\}\}=\\exp\\left\(\-\\alpha\\cdot\\frac\{D\_\{\\mathrm\{DTW\}\}\(\\hat\{z\}\_\{1:T^\{\*\}\},z\_\{1:T\}\)\+\\beta\|T^\{\*\}\-T\|\}\{T\}\\right\)\\cdot\\eta\.\(30\)
First,DDTW\(z^1:T∗,z1:T\)D\_\{\\mathrm\{DTW\}\}\(\\hat\{z\}\_\{1:T^\{\*\}\},z\_\{1:T\}\)denotes the optimal alignment cost between the simulated and real intent trajectories\. It measures the overall deviation between the two trajectories in terms of intent evolution, rather than the discrepancy between fixed turn positions\.

Second,β​\|T∗−T\|\\beta\|T^\{\*\}\-T\|is a trajectory\-length penalty\. Although DTW allows nonlinear alignment between sequences of different lengths, relying only on the DTW path cost may allow the simulator to avoid certain local deviations by generating trajectories that are overly short or overly long\. We therefore explicitly penalize the difference in the number of turns to prevent the simulated trajectory from deviating excessively from the real trajectory length\. Here,TTdenotes the number of turns in the real trajectory,T∗T^\{\*\}denotes the number of turns in the simulated trajectory, andβ\\betacontrols the strength of the length penalty\.

Third, the denominatorTTnormalizes the cost by the real trajectory length, improving comparability across sessions of different lengths\. Since longer real trajectories naturally tend to accumulate larger DTW costs, omitting this normalization would introduce a systematic length bias into the reward\.

Finally,η∈\(0,1\]\\eta\\in\(0,1\]is a flattening coefficient used to suppress degenerate intent trajectories\. We observe that, if only the DTW alignment cost and the length penalty are used, the simulator may tend to generate nearly constant intent sequences in order to reduce path deviation\. To discourage this behavior, when the real trajectory contains substantial intent variation but the simulated trajectory remains at a single or nearly single intent level for a long period, we setη\\etato a discount value smaller than 1\. This design encourages the simulator not only to match the terminal state and reduce the overall path cost, but also to preserve the dynamic variation pattern observed in real user trajectories\.

Therefore,RtrajR\_\{\\mathrm\{traj\}\}becomes high when the simulated and real trajectories have similar intent evolution paths, comparable numbers of turns, and no overly flat degenerate behavior\. Conversely, the reward decreases when the simulated trajectory reaches the same terminal state but follows a substantially different intermediate intent path, or when its number of turns deviates significantly from the real trajectory\.

##### Turn\-level Local Deviation from the DTW Optimal Path\.

Our fine\-grained advantage modulation mechanism further uses the DTW alignment result to construct a local deviation measuredtd\_\{t\}for each generated turn\. This corresponds to the future deviation score defined in Section[4\.3](https://arxiv.org/html/2609.28690#S4.SS3):

ct=∑k=0T∗−tρk⋅dt\+k\.c\_\{t\}=\\sum\_\{k=0\}^\{T^\{\*\}\-t\}\\rho^\{k\}\\cdot d\_\{t\+k\}\.\(31\)Here,dtd\_\{t\}is not a strict pointwise difference\|z^t−zt\|\|\\hat\{z\}\_\{t\}\-z\_\{t\}\|\. Instead, it is an aligned local deviation derived from the DTW optimal path\.

Let the optimal DTW alignment path be

P∗=\(\(i1,j1\),\(i2,j2\),…,\(iK,jK\)\),P^\{\*\}=\\big\(\(i\_\{1\},j\_\{1\}\),\(i\_\{2\},j\_\{2\}\),\\dots,\(i\_\{K\},j\_\{K\}\)\\big\),\(32\)where each path element\(ik,jk\)\(i\_\{k\},j\_\{k\}\)indicates that theiki\_\{k\}\-th simulated turnz^ik\\hat\{z\}\_\{i\_\{k\}\}is aligned with thejkj\_\{k\}\-th real turnzjkz\_\{j\_\{k\}\}\. The path satisfies the boundary, monotonicity, and step\-size constraints, thereby preserving the temporal order of both trajectories while allowing local one\-to\-many, many\-to\-one, and one\-to\-one alignments\.

Because DTW permits one\-to\-many and many\-to\-one alignments, the same simulated turniimay correspond to multiple real turns\. To obtain a local deviation measure for each simulated turn, we first collect all real\-turn indices aligned with simulated turniiunder the optimal path:

𝒜i=\{jk∣\(ik,jk\)∈P∗,ik=i\}\.\\mathcal\{A\}\_\{i\}=\\\{j\_\{k\}\\mid\(i\_\{k\},j\_\{k\}\)\\in P^\{\*\},\\ i\_\{k\}=i\\\}\.\(33\)We then define the local deviation of theii\-th simulated turn as the average absolute deviation over all aligned real turns:

di=1\|𝒜i\|​∑j∈𝒜i\|z^i−zj\|\.d\_\{i\}=\\frac\{1\}\{\|\\mathcal\{A\}\_\{i\}\|\}\\sum\_\{j\\in\\mathcal\{A\}\_\{i\}\}\|\\hat\{z\}\_\{i\}\-z\_\{j\}\|\.\(34\)
This definition converts the globally optimal DTW alignment path into a turn\-level deviation signal over the simulated trajectory\. Intuitively, if a simulated intent state at a certain turn differs substantially from the real intent stages aligned to it by DTW, then its local deviationdid\_\{i\}will be large\. If the simulated turn matches the corresponding real trajectory stage well, thendid\_\{i\}will be small\.

Based on these local deviations, the future deviation score accumulates downstream deviations starting from turntt:

ct=∑k=0T∗−tρk​dt\+k,c\_\{t\}=\\sum\_\{k=0\}^\{T^\{\*\}\-t\}\\rho^\{k\}d\_\{t\+k\},\(35\)
whereρ∈\[0,1\]\\rho\\in\[0,1\]is a discount factor\. This score measures the extent of intent deviation that may accumulate in the future trajectory after the current turn\. A largerctc\_\{t\}indicates that the simulated trajectory exhibits more severe or more persistent deviation after turntt\.

In turn\-level advantage modulation,ctc\_\{t\}controls the strength of advantage reweighting\. For positive trajectory\-level advantages, turns with largerctc\_\{t\}receive relatively weaker positive reinforcement, because they are followed by larger downstream deviations despite the overall trajectory being favorable\. For negative trajectory\-level advantages, turns with largerctc\_\{t\}receive stronger negative updates, because they are more likely to be associated with persistent downstream drift\.

It is important to distinguish the roles ofDDTWD\_\{\\mathrm\{DTW\}\}anddid\_\{i\}\. The termDDTWD\_\{\\mathrm\{DTW\}\}is a trajectory\-level cumulative alignment cost used to construct the trajectory plausibility rewardRtrajR\_\{\\mathrm\{traj\}\}\. In contrast,did\_\{i\}is a turn\-level local deviation derived from the same DTW optimal path and is used to construct the future deviation scorectc\_\{t\}, which further modulates the advantage at each turn\. Thus, the two quantities share the same DTW alignment process but serve different levels of training signal: reward design at the trajectory level and credit assignment at the turn level\.

##### Illustrative Example\.

Consider a real intent trajectory

Z=\[3,4,5\],Z=\[3,4,5\],\(36\)
and a simulator\-generated trajectory

Z^=\[3,3,4,5\]\.\\hat\{Z\}=\[3,3,4,5\]\.\(37\)
Under strict turn\-by\-turn alignment, the simulated trajectory would become misaligned from the second turn onward due to the length difference\. In contrast, DTW can naturally align the first two simulated intent states with the first real intent stage, and then align the subsequent states 4 and 5 with the corresponding real intent states 4 and 5\. Therefore, the extra simulated turn is not treated as a complete trajectory\-level error\. Instead, it is interpreted as a local pacing delay\.

Similarly, if the simulator completes need confirmation and intent advancement within a single turn, whereas the real user completes the same transition across two adjacent turns, DTW allows this simulated turn to be flexibly aligned with the neighboring stages in the real trajectory\. In this way, DTW better reflects natural pacing variation in multi\-turn dialogue, rather than penalizing every local turn\-level shift as an error\.

Through DTW alignment, our method assigns each simulated turn a more reasonable reference position in the real trajectory\. This yields both a trajectory\-level cumulative alignment costDDTWD\_\{\\mathrm\{DTW\}\}and a turn\-level local deviation signaldid\_\{i\}\. The former is used to construct the trajectory plausibility reward, while the latter is used for the future deviation score and turn\-level advantage modulation\. As a result, the simulator is no longer optimized under the overly rigid assumption that “thett\-th simulated turn must correspond to thett\-th real turn\.” Instead, it is encouraged to match the overall intent evolution path of real users\.

## Appendix ICase Study

CASE STUDY: Correct Conversion Does Not Guarantee Behavioral FidelityUser Profile & Task BackgroundThis is an 18–23\-year\-old platform user with a high\-school education or below\. The user has a partner and has no children\. The user works as a general worker and commonly uses an Android phone priced around RMB 1,700\. The user currently lives in Chaozhou, Guangdong Province, a fourth\-tier city in China\. The user’s personality trait is decisive and concise\. In this dialogue, the user intends to buy a car, asks about the down\-payment threshold and available vehicle models, and hopes to schedule a test drive\.Ground TruthAssistant:Hello, click the free consultation below, "Youjia" will serve you immediately\.User:What’s the minimum down payment?Assistant:Hello, the down payment ratio varies for different models, and prices change dynamically~ You can leave your contact info, and our specialist will provide one\-on\-one service to help you choose the most suitable car\-buying plan\.User:What cars do you have?Assistant:Hello, we have a variety of popular models on sale\. Do you have any preferences or intended models?User:Can I test drive?Assistant:Hello, yes you can\. Do you have a specific model in mind for the test drive? I can arrange it for you in advance\. Please leave your contact info, I’ll send you the store’s location\. You’re welcome to come and experience it~User:132xx9875TRACER\-7B \(Ours\)Assistant:Hello, click the free consultation below, "Youjia" will serve you immediately\.User:What’s the minimum down payment?Assistant:The down payment plan depends on the model and policies, and needs to be customized based on your budget\. Do you have a target budget range?User:What cars are there?Assistant:We have multiple popular models on sale with various configurations\. We need to recommend based on your needs, which brand do you prefer?User:Can I schedule a test driveAssistant:Test drives are supported, and you can schedule the time yourself\. Leave your contact info, and we will contact you as soon as possible to arrange the test drive~User:132xx9875Gemini\-3\.1\-ProAssistant:Hello, click the free consultation below, "Youjia" will serve you immediately\.User:I want to buy a car, what’s the minimum down payment? Any model recommendations? I want to test drive first\.Assistant:Down payments vary depending on the model and plan~ Is it convenient to leave your contact info? We can arrange a consultant to introduce in detail\. What’s your budget?User:Budget is under 100k, the lower the down payment, the better\. My phone number is 132xx9875, same on WeChat, just have the consultant add me and send the models, and arrange a test drive by the way\.Doubao\-Seed\-2\.0\-ProAssistant:Hello, click the free consultation below, "Youjia" will serve you immediately\.User:I’d like to ask, how much is the down payment generally for buying a car now? Are there any cost\-effective models suitable for ordinary workers like us? Can a test drive be arranged?Assistant:Down payments differ depending on the model, and test drives require a reservation\. We can arrange a one\-on\-one consultation to match plans for you~User:My budget is around 100k\. I’m in Guxiang Town, Chao’an, Chaozhou\. I just want a fuel\-efficient and durable commuter car\. How can I contact your consultants?Assistant:Understood your needs~ We will help you find suitable models\. Leave a contact number, we’ll arrange a one\-on\-one recommendation and contact you~User:Okay, my phone is 132xx9875, just arrange for someone to contact me\.Doubao\-Seed\-CharacterAssistant:Hello, click the free consultation below, "Youjia" will serve you immediately\.User:I want to buy a car, may I ask what the down payment threshold is, and are there any recommended models suitable for ordinary wage earners? I also want to book a test drive\.Assistant:Down payments vary depending on the model~ Leave your contact info, and we’ll arrange a consultant to introduce in detail~ What price range are you considering?User:I want a commuter car under 100k, my phone number is 132xx9875, WeChat is lalala, please arrange a consultant to connect with me, and I also want to book a test drive\.❖Baseline Failure Analysis•Lack of flaws\.•Overly cooperative, responses are too mechanical\.•Out of Character \(OOC\): Speech and behavior seriously violate the provided profile\.

## Appendix JPrompt Details

Persona and Task Background Extraction Prompt

Intent\-Level Annotation Prompt

Data Filtering and Cleaning Prompt

Merchant Agent System Prompt for Training and Test Data

User Simulator System Prompt for Training and Test Data

相似文章

SimTrace:面向在线用户建模的有据多模态用户轨迹生成

arXiv cs.AI

SimTrace 是一个开源框架,利用基于匿名真实轨迹和模拟网络环境的 computer-use agent,生成忠实、细粒度的合成多模态用户点击流。在 8 项保真度指标中有 7 项超过基线,并且在增强真实数据时将下一动作预测效果提升了 11.0%。

基于概率信念追踪的多轮人类可说服性模型

arXiv cs.CL

本文介绍了PersuasionTrace,一个用于研究人机交互中多轮说服的框架,采用贝叶斯网络模拟目标来建模信念更新。该框架揭示了大语言模型在多种主题和模态下具有说服力,并且贝叶斯目标比普通大语言模型模拟器更符合人类信念动态。

RealUserSim:通过真实用户模拟弥合智能体基准测试中的现实差距

arXiv cs.AI

本文介绍了RealUserSim,一个将基于LLM的用户模拟扎根于来自14,000+真实对话的人类行为数据中的框架,旨在弥合智能体基准测试中的现实差距。研究表明,基于真实数据的模拟将行为匹配率从24.2%提升至45.3%,并揭示了协作型模拟器无法发现的失效机制。