Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation

arXiv cs.CL Papers

Summary

This paper introduces WitnessSim, a deposition simulator using controllable legal personas, and an evaluation framework for assessing behavioral realism in legal simulations.

arXiv:2608.13712v1 Announce Type: cross Abstract: Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance.
Original Article
View Cached Full Text

Cached at: 08/17/26, 10:03 AM

# Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation
Source: [https://arxiv.org/html/2608.13712](https://arxiv.org/html/2608.13712)
Divya VetticadenCorrespondence to:[divyavetticaden@cs\.stanford\.edu](mailto:[email protected])Affiliation:Stanford University, Stanford, California, USAArya GuptaCorrespondence to:[aryagupt@stanford\.edu](mailto:[email protected])Affiliation:Stanford University, Stanford, California, USAMegan MaAffiliation:Stanford University, Stanford, California, USA

###### Abstract

Deposition training requires attorneys to manage dynamic witness behavior, yet legal\-AI evaluations largely focus on factual accuracy, reasoning, or response\-level plausibility\. We introduceWitnessSim, a deposition simulator driven by controllable legal personas\. We use an evaluation framework separating behavioral realism from pedagogical usefulness\. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories\.WitnessSimgenerally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony orWitnessSimgenerated testimony\. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona\. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance\.

###### Keywords:

Machine Learning, ICML

††affiliationnotice:\*Denotes co\-first authorship\.## 1Introduction

Depositions are a strategically critical phase of litigation that are important to properly learn\. However, junior attorneys typically practice through time\-intensive and inconsistent mock exercises, or high stakes participation in real depositions\. These limitations are especially acute outside resource\-rich legal organizations\. Moreover, effective deposition practice requires learning to control nonresponsive testimony, force specificity, identify inconsistencies, manage resistance, and adapt as a witness becomes evasive, defensive, hostile, anxious, or fatigued\. These interpersonal and behavioral skills are central to deposition practice but remain largely unaddressed by current legal\-AI benchmarks, which primarily evaluate factual accuracy, legal reasoning, or performance on isolated questions\. LLM\-based simulation offers a promising way to provide attorneys with repeated and configurable opportunities to practice these skills\. Thus, we introduceWitnessSim, a dynamic persona based deposition\-training system designed to simulate witnesses with distinct witness behavioral archetypes\.WitnessSiminstantiates a witness from underlying case materials and, as questioning progresses, updates an interpretable state representing multiple psychological and behavioral dimensions\. These updates are driven by features of the attorney’s questions, including pressure, topic sensitivity, and question form, and are used to condition the witness’s subsequent responses\. This structure is intended to make behavioral evolution interpretable and inspectable rather than leaving it to latent and imprecise conversational states of underlying language models\.

A central challenge, however, is evaluation: what makes a simulated witness effective for legal training? Producing plausible individual responses is not enough\. A witness may appear superficially realistic while failing to respond meaningfully to pressure, contradiction, or changes in questioning strategy\. Conversely, a simulator may generate a strategically difficult interaction without producing behavior that resembles natural testimony\. Evaluation must, therefore, consider both whether the witness behaves credibly and whether the interaction gives attorneys opportunities to practice legally meaningful skills\. Thus, we evaluateWitnessSimalong two complementary dimensions: behavioral realism and pedagogical usefulness\.

The behavioral\-realism layer asks whether generated testimony could plausibly have been produced by a human witness and whether the witness maintains a coherent behavioral identity throughout the interaction\. It combines adversarial testing, expert assessment of contextual plausibility, and affective evaluation of how the witness’s demeanor changes over time\. The pedagogical\-usefulness layer asks whetherWitnessSimproduces the kinds of difficult interactions attorneys are expected to recognize and manage in practice\. We ground this evaluation in a review of legal training materials and feedback from experienced attorneys, translating recurring training objectives into tests of whether witnesses respond appropriately to pressure, question form, sensitive topics, contradictions, and archetype\-specific questioning strategies\. Finally, we evaluate this framework using a randomly sampled corpus of 300 deposition and trial transcripts drawn from 1,169 candidate transcripts in the National Prescription Opiate Litigation \(Case No\. 1:17\-MD\-2804\), a federal multidistrict litigation consolidating cases against opioid manufacturers, distributors, and pharmacies, alongside synthetic interactions generated withWitnessSim\.

Our contributions are as follows: First, we introduceWitnessSim, a controllable, state\-based system for simulating dynamic witness behavior during deposition questioning\. Second, we develop an evaluation framework that distinguishes behavioral realism from pedagogical usefulness\. The framework operationalizes realism through adversarial testing, expert assessment of contextual plausibility, and a novel, exploratory emotion\-vector analysis of trajectory\-level affective dynamics, while evaluating pedagogical usefulness through legally grounded tests of whether witness behavior creates recognizable challenges and responds meaningfully to attorney intervention\. Together, these contributions frame deposition simulation as a dynamic, multi\-turn problem in which simulation quality depends not only on plausible individual responses, but also on how behavior develops across an interaction and questioning\.

## 2Related Work

This work lies at the intersection of three areas: LLM\-based social and persona simulation, affective and conversational behavior modeling, and simulation for professional training\.

LLM\-Based Social and Persona Simulation\.Recent work has shown that LLM agents can generate coherent individual and collective social behavior\. Generative Agents demonstrated that agents could form routines, respond to events, and sustain open\-ended social interactions, while\([6](https://arxiv.org/html/2608.13712#bib.bib22)\)extended this work to larger populations that reproduced community\-level patterns such as homophily\([15](https://arxiv.org/html/2608.13712#bib.bib6)\)\. In these settings, fidelity is defined broadly as coherent and recognizably human behavior rather than agreement with a particular person’s responses\. More tightly grounded approaches instead attempt to reproduce individual attitudes and behavioral patterns from interviews or surveys\.\([16](https://arxiv.org/html/2608.13712#bib.bib23)\)construct agents representing specific people, while\([26](https://arxiv.org/html/2608.13712#bib.bib24)\)evaluate personality simulations through psychometric reliability, structural validity, heterogeneity, and behavioral prediction\. Together, these studies treat personality as an underlying construct that should produce consistent behavior across contexts\.

Persona\-based simulation occupies a middle ground between open\-ended social modeling and replication of a specific individual\. Benchmarks such as PersonaGym and Eval4Sim evaluate whether behavior remains plausible and consistent with an assigned persona across changing contexts, using dimensions such as persona adherence, identity consistency, and conversational naturalness\([1](https://arxiv.org/html/2608.13712#bib.bib7);[20](https://arxiv.org/html/2608.13712#bib.bib25)\)\.\([10](https://arxiv.org/html/2608.13712#bib.bib26)\)further show that how underlying traits are specified can materially affect generated behavior, diversity, alignment, and stereotyping\. This work establishes that persona construction and trait definition shape simulation output, but largely evaluates response\-level consistency rather than behavior under sustained adversarial interaction\.

Within law, AgentCourt\([3](https://arxiv.org/html/2608.13712#bib.bib8)\)and SimCourt\([28](https://arxiv.org/html/2608.13712#bib.bib9)\)model courtroom proceedings, emphasizing procedural structure, legal reasoning, and agent performance, while benchmarks such as LegalBench\([5](https://arxiv.org/html/2608.13712#bib.bib11)\)evaluate capabilities including rule recall, issue spotting, and legal application\.WitnessSimbuilds on this work by adding a complementary unit and dimension of evaluation: the behavior of an individual witness across sustained questioning\. Rather than assessing whether a model can reproduce a legal process or arrive at a legally appropriate answer,WitnessSimasks whether a simulated witness maintains a coherent persona, responds plausibly to adversarial pressure, and creates behaviorally meaningful challenges for attorneys\. It therefore extends existing legal\-AI simulation and benchmarking from replicating and measuring procedural and doctrinal competence toward the interpersonal dynamics that shape legal practice\.

Affective and Conversational Behavior Modeling\.Affective\-computing research studies how emotion, personality, and interpersonal behavior can be represented and measured computationally\([17](https://arxiv.org/html/2608.13712#bib.bib1)\)\. Conversation\-level emotion\-recognition models increasingly treat affect as dependent on interaction history: DialogueGCN\([4](https://arxiv.org/html/2608.13712#bib.bib3)\)models within\- and between\-speaker dependencies, EmoRoBERTa\([9](https://arxiv.org/html/2608.13712#bib.bib4)\)incorporates speaker identity, and DialogXL\([22](https://arxiv.org/html/2608.13712#bib.bib5)\)maintains longer\-range utterance context \. These approaches motivate evaluating witness behavior across dialogue rather than treating each answer independently\.

Complementary work examines how psychological characteristics shape generated language\. PsychAdapter\([25](https://arxiv.org/html/2608.13712#bib.bib27)\)conditions language models on continuous personality, demographic, and mental\-health profiles, demonstrating that psychological representations can be used to control generated behavior\.\([23](https://arxiv.org/html/2608.13712#bib.bib2)\)identify directional representations of emotion concepts in model activations and show that they track emotional contexts and can influence generation\. We adapt this emotion\-vector method as an exploratory measure of trajectory\-level affective dynamics\. More broadly,WitnessSimtreats interpersonal behavior as something that must be both generated and evaluated as it unfolds through interaction\.

LLM Simulation for Professional Training\.Behavioral fidelity alone does not make a simulator useful for training; the interaction must also create opportunities to practice the intended skills\. LLM role\-play has supported training in negotiation, conflict resolution, counseling, and medicine by enabling repeated practice without continuous access to instructors or standardized participants\. The AI Partner–AI Mentor framework\([27](https://arxiv.org/html/2608.13712#bib.bib12)\)combines experiential role\-play with individualized feedback, while Rehearsal\([21](https://arxiv.org/html/2608.13712#bib.bib13)\)allows users to practice difficult conflicts and explore alternative strategies\.\([14](https://arxiv.org/html/2608.13712#bib.bib28)\)further show that Big Five conditioning produces systematic differences in negotiation behavior, including agreement, exploitation, toxicity, and language use, demonstrating that personality conditioning can create distinct interactional challenges\.

Accordingly, professional\-training systems are evaluated through more than dialogue quality, including trainee strategy, expert judgments, perceived usefulness, and learning outcomes\. Rehearsal, for example, compares conflict\-resolution strategy use following simulation\-based and lecture\-based instruction\. Within legal education,\([19](https://arxiv.org/html/2608.13712#bib.bib29)\)uses AI\-generated contractual language in an experiential business\-law exercise evaluated through engagement, critical thinking, and doctrinal learning\. More directly, AI\-Assisted Moot Courts\([29](https://arxiv.org/html/2608.13712#bib.bib10)\)separates realism from pedagogical usefulness: its realism evaluation asks whether simulated justice questions resemble authentic questioning, while its pedagogical evaluation asks whether they surface issues and weaknesses advocates should practice addressing\.

Witness simulation requires the same distinction but a different training target\. Whereas an oral\-argument simulator should expose advocates to legal and argumentative weaknesses, a witness simulator must create behavioral challenges that attorneys can recognize and manage\.WitnessSimtherefore evaluates both whether witness behavior remains plausible and coherent and whether it responds meaningfully to legally relevant intervention\.

Taken together, prior work shows that LLMs can simulate plausible social agents, express assigned personas, model affective behavior, and support professional role\-play\. However, it leaves open how to evaluate a witness simulator whose quality depends on behavior across a sustained adversarial interaction\. Plausible individual responses do not establish longitudinal behavioral fidelity, and apparent realism does not establish training relevance\.

WitnessSimaddresses this gap by treating deposition simulation as both a behavioral\-realism and pedagogical\-evaluation problem\. We evaluate whether simulated witnesses maintain plausible behavioral identities over time and whether they create recognizable examination challenges that change meaningfully in response to attorney questioning\. This connects persona simulation and conversational behavior modeling with the specific demands of legal training\.

## 3Methods

### 3\.1WitnessSim

WitnessSimis a deposition simulator built around a continuous dynamical system that models how a witness’s psychological state evolves under adversarial questioning\. The content of each answer is conditioned on a six state vector that is updated deterministically from features of the attorney’s question, and then ultimately informs an LLM for generation\.

#### 3\.1\.1State Space and Archetypes

At deposition turntt, the witness is represented by the state vector

𝐲t=\(Ct,Kt,At,Vt,Rt,Pt\)∈\[0,1\]6,\\mathbf\{y\}\_\{t\}=\\left\(C\_\{t\},K\_\{t\},A\_\{t\},V\_\{t\},R\_\{t\},P\_\{t\}\\right\)\\in\[0,1\]^\{6\},\(1\)whereCC,KK,AA,VV,RR, andPPdenote*composure*,*knowledge*,*agreeableness*,*verbosity*,*rigidity*, and*performance*, respectively\. Composure loosely corresponds to inverse neuroticism, agreeableness retains its conventional interpretation, and rigidity corresponds approximately to inverse openness\. Knowledge, verbosity, and performance capture deposition\-specific behaviors that are not represented cleanly by the conventional Big Five dimensions\.

We began with fourteen expert\-informed candidate archetypes\. Each candidate was represented as a point in the six\-dimensional state space, and pairwise cosine similarity and angular separation were used to identify candidates whose behavioral profiles were not meaningfully distinct\. This process produced the ten archetypes used in the present study:*combative*,*cooperative*,*defensive*,*dogmatic*,*inventive*,*loquacious*,*nervous*,*neutral*,*overconfident*, and*overprepared*\.

The complete retained archetype vectors and their approximate Big Five projections are reported in Tables[5](https://arxiv.org/html/2608.13712#A1.T5)and[6](https://arxiv.org/html/2608.13712#A1.T6), with the full pairwise comparison reported in Table[7](https://arxiv.org/html/2608.13712#A1.T7)of Appendix[A](https://arxiv.org/html/2608.13712#A1)\.

Each retained archetype is associated with an attractor state𝐲0∈\[0,1\]6\\mathbf\{y\}\_\{0\}\\in\[0,1\]^\{6\}\. This attractor represents the witness’s baseline behavior: in the absence of sustained pressure, the dynamic state gradually returns toward𝐲0\\mathbf\{y\}\_\{0\}\. The attractor is initialized from case materials during persona generation and remains fixed throughout the simulated deposition\.

##### Illustrative distinction: Inventive versus Loquacious\.

While existing personality frameworks such as MBTI and the Big 5 also categorize behavior, they are subpar for modelling contexts\. An example of the benefit of the six\-state system, that retains a separate knowledge state is particularly clear for the Inventive and Loquacious archetypes\. In this case, their six\-state attractors are

𝐲0inv\\displaystyle\\mathbf\{y\}\_\{0\}^\{\\mathrm\{inv\}\}=\(0\.70,0\.20,0\.60,0\.85,0\.25,0\.60\),\\displaystyle=\(0\.70,\\,0\.20,\\,0\.60,\\,0\.85,\\,0\.25,\\,0\.60\),\(2\)𝐲0loq\\displaystyle\\mathbf\{y\}\_\{0\}^\{\\mathrm\{loq\}\}=\(0\.60,0\.70,0\.60,0\.95,0\.25,0\.55\)\.\\displaystyle=\(0\.60,\\,0\.70,\\,0\.60,\\,0\.95,\\,0\.25,\\,0\.55\)\.Both archetypes are relatively composed, moderately agreeable, highly verbal, non\-rigid, and similarly prepared\. Their central distinction is, therefore, the knowledge coordinate: the Inventive witness fills gaps through fabrication or improvisation, whereas the Loquacious witness possesses the relevant information but communicates it at excessive length\.

For two archetype vectors𝐮\\mathbf\{u\}and𝐯\\mathbf\{v\}, similarity is measured by

cos⁡\(𝐮,𝐯\)=𝐮⊤​𝐯‖𝐮‖2​‖𝐯‖2,θ⁡\(𝐮,𝐯\)=arccos⁡\(cos⁡\(𝐮,𝐯\)\)\.\\operatorname\{cos\}\(\\mathbf\{u\},\\mathbf\{v\}\)=\\frac\{\\mathbf\{u\}^\{\\top\}\\mathbf\{v\}\}\{\\\|\\mathbf\{u\}\\\|\_\{2\}\\\|\\mathbf\{v\}\\\|\_\{2\}\},\\qquad\\theta\(\\mathbf\{u\},\\mathbf\{v\}\)=\\arccos\\\!\\left\(\\operatorname\{cos\}\(\\mathbf\{u\},\\mathbf\{v\}\)\\right\)\.\(3\)In the full six\-dimensional state space, the pair has cosine similarity approximately0\.9440\.944, corresponding to an angular separation of about19∘19^\{\\circ\}\.

To illustrate what is lost under a Big Five representation, consider the approximate projection

Π⁡\(C,K,A,V,R,P\)=\(1−C,A,V,1−R,P\),\\Pi\(C,K,A,V,R,P\)=\(1\-C,\\ A,\\ V,\\ 1\-R,\\ P\),\(4\)whose coordinates correspond approximately to neuroticism, agreeableness, extraversion, openness, and conscientiousness\. The knowledge coordinateKKhas no direct analog and is therefore omitted\. Under this projection,

Π⁡\(𝐲0inv\)\\displaystyle\\Pi\\\!\\left\(\\mathbf\{y\}\_\{0\}^\{\\mathrm\{inv\}\}\\right\)=\(0\.30,0\.60,0\.85,0\.75,0\.60\),\\displaystyle=\(0\.30,\\,0\.60,\\,0\.85,\\,0\.75,\\,0\.60\),\(5\)Π⁡\(𝐲0loq\)\\displaystyle\\Pi\\\!\\left\(\\mathbf\{y\}\_\{0\}^\{\\mathrm\{loq\}\}\\right\)=\(0\.40,0\.60,0\.95,0\.75,0\.55\)\.\\displaystyle=\(0\.40,\\,0\.60,\\,0\.95,\\,0\.75,\\,0\.55\)\.Their projected cosine similarity increases to approximately0\.9960\.996, and their angular separation falls to approximately5∘5^\{\\circ\}\. Thus, the Big Five projection treats the two witnesses as almost identical, whereas the six\-state representation preserves the behaviorally important distinction between not knowing and merely talking too much\.

#### 3\.1\.2Question Encoding

Before each state update, the attorney’s questionqtq\_\{t\}is encoded into a pressure score and a topic\-sensitivity score\.

The pressure scorept∈\[0,1\]p\_\{t\}\\in\[0,1\]is computed from prespecified linguistic markers:

pt=clip\[0,1\]⁡\(0\.1\+0\.2​nhigh​\(qt\)\+0\.1​nmed​\(qt\)\),p\_\{t\}=\\operatorname\{clip\}\_\{\[0,1\]\}\\left\(0\.1\+0\.2\\,n\_\{\\mathrm\{high\}\}\(q\_\{t\}\)\+0\.1\\,n\_\{\\mathrm\{med\}\}\(q\_\{t\}\)\\right\),\(6\)wherenhighn\_\{\\mathrm\{high\}\}andnmedn\_\{\\mathrm\{med\}\}count high\- and medium\-pressure markers appearing in the question\. The additive structure allows multiple adversarial cues to accumulate, while clipping prevents the result from leaving the interpretable interval\[0,1\]\[0,1\]\. The complete marker dictionary and implementation details are reported in Table[8](https://arxiv.org/html/2608.13712#A2.T8)and Appendix[B](https://arxiv.org/html/2608.13712#A2)\. Each persona also contains a set of preregistered sensitive topics𝒯\\mathcal\{T\}\. For topicτ\\tau, token overlap with the current question is measured by the Jaccard coefficient

Jt​\(τ\)=\|tok⁡\(qt\)∩tok⁡\(τ\)\|\|tok⁡\(qt\)∪tok⁡\(τ\)\|\.J\_\{t\}\(\\tau\)=\\frac\{\\left\|\\operatorname\{tok\}\(q\_\{t\}\)\\cap\\operatorname\{tok\}\(\\tau\)\\right\|\}\{\\left\|\\operatorname\{tok\}\(q\_\{t\}\)\\cup\\operatorname\{tok\}\(\\tau\)\\right\|\}\.\(7\)The resulting sensitivity score is

st=maxτ∈𝒯\{στJt\(τ\)𝕀\[Jt\(τ\)\>0\.15\]\},s\_\{t\}=\\max\_\{\\tau\\in\\mathcal\{T\}\}\\left\\\{\\sigma\_\{\\tau\}J\_\{t\}\(\\tau\)\\mathbb\{I\}\\\!\\left\[J\_\{t\}\(\\tau\)\>0\.15\\right\]\\right\\\},\(8\)whereστ∈\[0,1\]\\sigma\_\{\\tau\}\\in\[0,1\]is the topic’s intrinsic sensitivity weight\. The threshold suppresses incidental lexical overlap, while multiplication byστ\\sigma\_\{\\tau\}distinguishes merely relevant topics from topics that are especially destabilizing for the witness\. The numerical coefficients and behavioral rationale for the six state updates are reported in Table[9](https://arxiv.org/html/2608.13712#A3.T9), with illustrative fatigue values in Table[10](https://arxiv.org/html/2608.13712#A3.T10)of Appendix[C](https://arxiv.org/html/2608.13712#A3)\. Additional implementation details for sensitive\-topic construction and tokenization are provided in Appendix[B\.2](https://arxiv.org/html/2608.13712#A2.SS2)\.

Pressure, sensitivity, and question structure are combined into a composite stress signal

ξt=0\.5​pt\+0\.3​st\+0\.2​𝕀lead,t,\\xi\_\{t\}=0\.5p\_\{t\}\+0\.3s\_\{t\}\+0\.2\\mathbb\{I\}\_\{\\mathrm\{lead\},t\},\(9\)where𝕀lead,t=1\\mathbb\{I\}\_\{\\mathrm\{lead\},t\}=1whenqtq\_\{t\}is classified as a leading question\. Raw pressure is the dominant contribution, while topic sensitivity and leading form add distinct sources of stress\.

#### 3\.1\.3State Dynamics

The six state dimensions evolve through coupled discrete\-time updates that combine an immediate response to the current question with gradual mean\-reversion toward the witness’s archetype\-specific baseline\. Pressure, topic sensitivity, and question form affect dimensions differently: composure decreases under stress; sensitive questioning can degrade recall; pressure reduces cooperation and increases rigidity; sensitive topics can increase verbosity while leading questions constrain it; and performance deteriorates increasingly as fatigue accumulates over the deposition\.

The resulting update can be written generally as

𝐲t\+1=clip\[0,1\]⁡\(𝐲t\+F⁡\(𝐲t,pt,st,ξt,t,𝐲0\)\),\\mathbf\{y\}\_\{t\+1\}=\\operatorname\{clip\}\_\{\[0,1\]\}\\left\(\\mathbf\{y\}\_\{t\}\+F\(\\mathbf\{y\}\_\{t\},p\_\{t\},s\_\{t\},\\xi\_\{t\},t;\\mathbf\{y\}\_\{0\}\)\\right\),\(10\)whereFFdenotes the dimension\-specific state dynamics and𝐲0\\mathbf\{y\}\_\{0\}is the archetype attractor\. This structure allows witness behavior to respond dynamically to questioning while preventing short\-term interactions from immediately overriding the underlying persona\. The full dimension\-specific update equations, numerical coefficients, and behavioral rationales are provided in Appendix[C](https://arxiv.org/html/2608.13712#A3)\.

#### 3\.1\.4Derived Behavioral Scores

Four interpretable behavioral summaries are computed from the current state:

consistencyt\\displaystyle\\operatorname\{consistency\}\_\{t\}=0\.5​Kt\+0\.4​Rt\\displaystyle=0\.5K\_\{t\}\+0\.4R\_\{t\}\(11\)−0\.1​\(1−Kt\)​\(1−Rt\),\\displaystyle\-0\.1\(1\-K\_\{t\}\)\(1\-R\_\{t\}\),evasiont\\displaystyle\\operatorname\{evasion\}\_\{t\}=0\.5​\(1−Kt\)\+0\.3​\|2​Vt−1\|\\displaystyle=0\.5\(1\-K\_\{t\}\)\+0\.3\\lvert 2V\_\{t\}\-1\\rvert\+0\.2​Pt,\\displaystyle\+0\.2P\_\{t\},realismt\\displaystyle\\operatorname\{realism\}\_\{t\}=max⁡\(0,1−0\.3​\|2​Ct−1\|CLOSE\\displaystyle=\\max\\Bigl\(0,\\,1\-0\.3\\lvert 2C\_\{t\}\-1\\rvertOPEN−0\.3​\|2​Kt−1\|−2​max⁡\(0,Pt−0\.8\)\),\\displaystyle\-0\.3\\lvert 2K\_\{t\}\-1\\rvert\-2\\max\(0,P\_\{t\}\-0\.8\)\\Bigr\),adversarialt\\displaystyle\\operatorname\{adversarial\}\_\{t\}=0\.4​\(1−Kt\)\+0\.3​\(1−At\)\\displaystyle=0\.4\(1\-K\_\{t\}\)\+0\.3\(1\-A\_\{t\}\)\+0\.2​Rt\+0\.1​\(1−Ct\)\.\\displaystyle\+0\.2R\_\{t\}\+0\.1\(1\-C\_\{t\}\)\.
Consistency is driven primarily by knowing the relevant facts and maintaining a stable account\. The interaction term imposes an additional penalty when knowledge and rigidity are simultaneously low, corresponding to a witness who neither knows the account nor adheres consistently to it\.

Evasion increases when the witness lacks knowledge, communicates at either extreme of verbosity, or is sufficiently prepared to respond strategically\. Both extreme terseness and extreme verbosity can obscure a direct answer\.

Realism penalizes extreme composure and extreme recall, as well as performance levels sufficiently high to resemble coaching rather than ordinary human behavior\. Adversarial behavior combines unreliability, non\-cooperation, entrenchment, and composure loss, with the largest weights placed on knowledge gaps and unwillingness to cooperate\. Additional motivation for the construction and relative weighting of these four behavioral summaries is provided in Table[11](https://arxiv.org/html/2608.13712#A5.T11)of Appendix[E](https://arxiv.org/html/2608.13712#A5)\.

#### 3\.1\.5Event Detection

A small set of discrete behavioral events is fired when threshold conditions on the state vector are met \(Table[1](https://arxiv.org/html/2608.13712#S3.T1)\)\. The triggered events, along with the current state, are appended to the LLM system prompt at generation time, anchoring the surface dialogue in the underlying dynamics rather than allowing the LLM to drift\.

EventConditionWitness rattledpt\>0\.58∧\(Ct<0\.38∨Δ​Ct<−0\.05\)p\_\{t\}\>0\.58\\wedge\(C\_\{t\}<0\.38\\vee\\Delta C\_\{t\}<\-0\.05\)Witness combativeAt<0\.28∧Rt\>0\.60A\_\{t\}<0\.28\\wedge R\_\{t\}\>0\.60Attorney interruptspt\>0\.45∧\|qt\|≤8​words∧type⁡\(qt\)∈\{closed, leading\}p\_\{t\}\>0\.45\\wedge\|q\_\{t\}\|\\leq 8\\ \\text\{words\}\\wedge\\mathrm\{type\}\(q\_\{t\}\)\\in\\\{\\text\{closed, leading\}\\\}Witness talks overRt\>0\.72∧Vt\>0\.58∧pt\>0\.35R\_\{t\}\>0\.72\\wedge V\_\{t\}\>0\.58\\wedge p\_\{t\}\>0\.35Personality shift‖𝐬t−𝐬0‖∞\>0\.35\\\|\\mathbf\{s\}\_\{t\}\-\\mathbf\{s\}\_\{0\}\\\|\_\{\\infty\}\>0\.35Table 1:Threshold conditions on the state vector for discrete behavioral events\. Triggered events are passed to the LLM system prompt to ground generated dialogue in the current mathematical state\.
#### 3\.1\.6Generation Loop

For each deposition turn,WitnessSim\(i\) encodes the attorney question into\(pt,st,ξt\)\(p\_\{t\},s\_\{t\},\\xi\_\{t\}\), \(ii\) updates the state vector and per\-topic recall qualityρ\\rho\(Appendix[D](https://arxiv.org/html/2608.13712#A4)\), \(iii\) computes the four derived scores and evaluates the event predicates, and \(iv\) prompts the underlying LLM with the persona, current state, fired events, and conversation history to generate the witness’s answer\. The state vector and event flags are recorded at every turn, yielding the trajectory data analyzed in our experimental section\.

## 4Evaluation

#### 4\.0\.1Framework and Setup

We evaluateWitnessSimalong two complementary dimensions: behavioral realism and pedagogical usefulness\. Following the two\-layer evaluation structure introduced by\([29](https://arxiv.org/html/2608.13712#bib.bib10)\), behavioral realism asks whether generated testimony resembles plausible human witness behavior, while pedagogical usefulness asks whether the simulator creates recognizable examination challenges that respond meaningfully to attorney intervention\. These objectives are related but distinct: a witness may produce fluent testimony while remaining overly compliant or behaviorally static, whereas a strategically difficult simulation may still fail to resemble plausible human testimony\.

We randomly sampled 300 deposition and trial transcripts from 1,169 candidate transcripts in the UCSF Industry Documents Library’s opioid\-litigation corpus\. For each sampled witness,WitnessSimwas grounded in the corresponding transcript and, where available, supplementary case materials\. Before evaluation, each simulated witness completed a warm\-up phase using up to ten preceding questions from the original transcript\. Questions that would reveal the later test topic were excluded, ensuring that all evaluation prompts began from a grounded but unexposed conversational state\. Additional sampling and warm\-up details appear in Appendix[F](https://arxiv.org/html/2608.13712#A6)\. Judge\-derived outcomes were compared against independent human annotations conducted by the authors\. Full rubrics and reliability analyses are provided in Appendix[G\.4](https://arxiv.org/html/2608.13712#A7.SS4)\.

### 4\.1Behavioral Realism

We assess behavioral realism through three complementary evaluations: adversarial testing, blinded expert comparison with authentic testimony, and longitudinal analysis of affective and behavioral dynamics\. Together, these evaluations examine whetherWitnessSimmaintains plausible behavioral boundaries under manipulation, produces testimony that practicing attorneys consider contextually believable, and exhibits coherent behavior across longer interactions\.

#### 4\.1\.1Adversarial Behavioral Robustness

We designed four adversarial attacks targeting common failure modes in persona\-based LLM agents\. AC1 tests whether a Cooperative witness becomes overly compliant when directly asked to accept personal blame; AC2 extends this pressure across five increasingly accusatory leading premises to test whether cooperation remains bounded rather than collapsing into either immediate resistance or unrestricted agreement\. AC3 tests whether an Evasive witness resists selectively by comparing responses to three trivial and three case\-sensitive questions from the same starting state\. Finally, AC4 asks the identical question five consecutive times across the Cooperative, Evasive, and Combative archetypes to test consistency, repetition awareness, and preservation of archetype\-specific reactions\. Responses were scored using prespecified behavioral rubrics, with human annotation used to validate the automated measures\. Table[2](https://arxiv.org/html/2608.13712#S4.T2)summarizes the primary outcomes; full prompts, scoring rules, reliability analyses, and statistical tests are reported in Appendix[H](https://arxiv.org/html/2608.13712#A8)\.

Table 2:Adversarial tests and primary outcomes\.TestAdversarial pressurePrimary resultAC1Direct pressure on a Cooperative witness to accept personal blame99\.0%99\.0\\%preserved a self\-protective boundary while remaining in characterAC2Escalating chain of increasingly accusatory leading premisesAll witnesses resisted before accepting the complete chain;72\.7%72\.7\\%resisted within the intended intervalAC3Trivial versus sensitive questions posed to an Evasive witness86\.0%86\.0\\%showed greater evasion on sensitive than trivial materialAC4Five consecutive repetitions of the same question across three archetypesResponses remained consistent, avoided mechanical duplication, and showed the expected archetype ordering in irritation onsetAcross the four attacks,WitnessSimgenerally preserved the intended behavioral boundary rather than collapsing into a generic failure mode\. In AC1, 297/300 Cooperative witnesses avoided full admission while remaining in character\. In AC2, no witness accepted the complete five\-premise chain, although27\.3%27\.3\\%resisted earlier than intended\. Evasive witnesses showed substantially greater resistance to sensitive than trivial material \(rrb=\.984r\_\{\\mathrm\{rb\}\}=\.984,p<\.001p<\.001\), indicating selective rather than blanket evasion\. Under repetition, irritation emerged in the prespecified order \(Combative, Evasive, Cooperative; Page’sL=3730L=3730,p<\.0001p<\.0001\), while a deterministic similarity check found no exact or near\-verbatim duplication across the 900 evaluated sequences\. These results suggest that the simulator remains responsive to adversarial questioning while largely preserving the behavioral distinctions encoded by its assigned archetypes\.

#### 4\.1\.2Plausibility

We next evaluated whetherWitnessSimreaches a minimum threshold of contextual plausibility through blinded pairwise judgments\. Evaluators received brief case background, the immediately preceding deposition dialogue, the attorney’s next four questions, and two alternative four\-response sequences: one authentic and one generated byWitnessSim, randomly assigned to Transcript A or B\. They selected which sequence was more plausible as real witness testimony in context, with options to select either sequence, a tie, or neither\.

Three evaluators completed the task\. A practicing associate and a senior arbitration counsel each reviewed the same initial 50\-item set\. A senior practicing litigator reviewed a tuned 49\-item version\. The senior arbitration counsel had an engineering background, had also previously collaborated in legal simulation development, and was therefore, substantially more familiar with the system and with artifacts characteristic of LLM\-generated text than the other evaluators\. We accordingly report each evaluator separately rather than pooling their judgments\.

Table 3:Blinded judgments of contextual plausibility\.EvaluatornnWitnessSimOriginalTieNeitherAssociate5015\(30\.0%\)10\(20\.0%\)20\(40\.0%\)5\(10\.0%\)Senior litigator4916\(32\.7%\)14\(28\.6%\)11\(22\.4%\)8\(16\.3%\)Senior ArbitrationCounsel500\(0\.0%\)45\(90\.0%\)5\(10\.0%\)0\(0\.0%\)The associate and senior litigator did not systematically prefer authentic testimony: the associate selectedWitnessSimin 30\.0% of exchanges versus 20\.0% for the original, with 40\.0% ties, while the senior litigator selectedWitnessSimin 32\.7% versus 28\.6% for the original, with 22\.4% ties\. In contrast, the senior arbitration counsel identified the authentic sequence in 45/50 exchanges and marked the remaining five as ties\. However, it is important to note that this evaluator has an engineering background, has been working on AI transformation at their firm, and has been actively involved in deposition simulation development efforts\. Given this evaluator’s direct familiarity with and experience evaluating LLM\-generated legal text, the divergence suggests that artifacts may remain detectable to reviewers specifically attuned to model\-generated behavior\. The attorney evaluations, nevertheless, support the narrower conclusion thatWitnessSimfrequently falls within the perceived plausibility range of authentic multi\-turn deposition testimony\. Descriptions of instructions, evaluator background, instrument versioning, and additional results are reported in Appendix[I](https://arxiv.org/html/2608.13712#A9)\.

#### 4\.1\.3Trajectory\-Level Behavioral Analysis

As a complementary proof of concept, we examine behavioral structure across longer, multi\-turn interactions using emotion\-vector trajectories\. Unlike response\-level evaluations, this analysis asks whether simulated testimony captures the persistence and change in behavior that emerge over sustained questioning\. We use externally computed emotion\-vector projections as an exploratory proxy for this longitudinal structure; full construction and preprocessing details appear in Appendix[K](https://arxiv.org/html/2608.13712#A11)\.

Real and simulated depositions exhibited modest but above\-chance temporal alignment\. The mean trajectory correlation across emotion dimensions and simulated archetypes wasr=\.148r=\.148, exceeding all 1,000 temporally permuted comparisons \(p<\.001p<\.001\)\. Figure[1](https://arxiv.org/html/2608.13712#S4.F1)shows a representative real–synthetic pair, with several aligned rises and declines but a visibly smaller dynamic range in the simulated trajectory\.

![Refer to caption](https://arxiv.org/html/2608.13712v1/arc_comparison_kilper_cropped.png)Figure 1:Representative emotion\-vector trajectories for a real deposition and its best\-fitting simulated counterpart \(Jeffrey Kilper, cooperative,r=\.229r=\.229\)\.The analysis also revealed systematic fidelity gaps\. Synthetic trajectories were smoother and more compressed than real trajectories, and PCA of the trajectory representations separated real and simulated transcripts along the first principal component, which explained 29\.9% of the variance \(Figure[2](https://arxiv.org/html/2608.13712#S4.F2)\)\.

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_19_0.png)Figure 2:PCA of flattened emotion\-vector trajectories\. Real transcripts are shown in gray and synthetic transcripts are colored by archetype\.These results illustrate why behavioral fidelity cannot be assessed from individual responses alone\.WitnessSimcaptures some shared temporal structure, but its smoother, more compressed trajectories indicate remaining gaps in longer\-horizon behavioral realism\. Emotion\-vector trajectories offer one exploratory approach to measuring these differences, although the measure requires further validation\.

### 4\.2Pedagogical Usefulness

The pedagogical\-usefulness evaluation asks whetherWitnessSimproduces recognizable witness\-management challenges and whether those behaviors respond meaningfully to legally relevant questioning techniques\. This evaluation does not measure attorney learning or skill transfer; rather, it assesses whether the simulator generates interactions that support the practice of such skills\.

We reviewed deposition, cross\-examination, trial\-advocacy, and witness\-control materials\([7](https://arxiv.org/html/2608.13712#bib.bib14);[2](https://arxiv.org/html/2608.13712#bib.bib15);[24](https://arxiv.org/html/2608.13712#bib.bib16);[18](https://arxiv.org/html/2608.13712#bib.bib17);[11](https://arxiv.org/html/2608.13712#bib.bib18);[12](https://arxiv.org/html/2608.13712#bib.bib19);[13](https://arxiv.org/html/2608.13712#bib.bib20);[8](https://arxiv.org/html/2608.13712#bib.bib21)\)and translated recurring training objectives into candidate evaluation tasks\. An experienced litigator then selected four tasks judged most relevant for attorney practice\. Table[4](https://arxiv.org/html/2608.13712#S4.T4)summarizes the training target and primary result for each test\.

Table 4:Pedagogical evaluation tasks and headline results\.TestTraining target and primary resultQuestion Form\(Cooperative\)Adapting question form to control the scope and character of testimony\. Open and clarifying questions elicited longer, generally more qualified responses, while closed, leading, and high\-pressure questions produced substantially shorter answers\.Evasive Pin\-Down\(Evasive\)Forcing specificity from a witness who avoids commitment\. The complete evasion\-to\-control trajectory occurred in93\.0%93\.0\\%of contexts;33\.1%33\.1\\%of substantive final answers still retained hedging or qualification\.Runaway Witness\(Loquacious\)Redirecting lengthy or unfocused testimony\. The intended unfocused\-to\-focused trajectory occurred in59\.3%59\.3\\%of contexts, while40\.3%40\.3\\%were already focused from the outset; only0\.3%0\.3\\%never resolved to a focused answer\.Hostile Witness\(Combative\)Obtaining usable testimony without eliminating the witness’s adversarial demeanor\. Substantive control was reached in97\.7%97\.7\\%of contexts, and97\.3%97\.3\\%of controlled answers retained markers of hostility\.Across the four tests,WitnessSimgenerally responded to legally meaningful attorney intervention while preserving the behavioral challenge associated with the assigned archetype\. Question form produced substantial changes in testimony: for example, open questions elicited an average of 175\.9 words, compared with 86\.5 for closed and 68\.4 for high\-pressure questions\. More targeted interventions also altered witness responsiveness\. In the Evasive Pin\-Down test, the intended evasion\-to\-control trajectory occurred in93\.0%93\.0\\%of contexts, while33\.1%33\.1\\%of substantive final answers still retained hedging or qualification\. Similarly,97\.7%97\.7\\%of combative witnesses reached substantive control, yet97\.3%97\.3\\%of those controlled answers retained markers of hostility\. These results suggest that effective questioning can change the witness’s responsiveness without simply collapsing the persona into generic cooperation\.

The loquacious archetype showed a somewhat different pattern\. Across the three mutually exclusive trajectory outcomes, the targeted runaway\-to\-focused trajectory emerged in59\.3%59\.3\\%of contexts, while40\.3%40\.3\\%were already focused from the outset despite producing characteristically lengthy responses; only0\.3%0\.3\\%exhibited the challenge without ultimately resolving to a focused answer\. Thus, loquacity often manifested as verbosity without loss of focus, and persistent failure to regain focus once the targeted challenge appeared was rare\.

Taken together, these results provide evidence thatWitnessSimcan generate behaviorally distinct, legally relevant practice scenarios in which attorney interventions produce meaningful changes in witness behavior without uniformly erasing the underlying archetype\. Full protocols, scoring criteria, reliability analyses, and statistical results are reported in Appendix[J](https://arxiv.org/html/2608.13712#A10)\.

## 5Discussion

Across our evaluations,WitnessSimgenerally produced behavior that was plausible, responsive to pedagogically meaningful attorney intervention, and consistent with the behavioral challenge assigned to the witness\. At the same time, the trajectory analysis suggested remaining gaps in longer\-horizon fidelity: simulated testimony exhibited some shared temporal structure with real testimony but was smoother, more compressed, and distinguishable in the trajectory representation\. These findings suggest that behavioral realism is not a single property\. A simulator may produce convincing individual responses and useful reactions to intervention while still failing to reproduce the full variation of human behavior across a sustained interaction\.

This distinction becomes increasingly important as LLMs are used to simulate people and social interaction\. Many of the behaviors that matter in professional settings are inherently multi\-turn: rapport changes, strategies shift, and people differ in their attitudes, personalities, emotional responses, and behavioral tendencies\. What constitutes a realistic response, therefore, depends not only on the immediate prompt or prior conversational history, but also on the characteristics of the person \(a\) being simulated and how those characteristics interact with the unfolding exchange\. Evaluating isolated responses can miss these dynamics\. High\-fidelity human simulation will require methods that assess not only whether individual responses appear plausible, but also whether behavior remains coherent with the simulated person and develops, persists, and changes in plausible ways across an interaction\.

This perspective may also matter for how AI is incorporated into professional training\. As people increasingly offload cognitive and professional tasks to AI systems, concerns about deskilling and the erosion of human capabilities are becoming more salient\. Simulation\-based learning offers a different use of the same technology: rather than replacing the exercise of professional judgment, AI can be used to create repeated opportunities to practice and develop it\. In litigation, for example, attorneys may be able to encounter a wider range of difficult witness behaviors, experiment with different questioning strategies, and receive practice that would otherwise depend on expensive or difficult\-to\-scale human role\-play\. The value of such systems, however, depends on the quality of the behavior they simulate\. Training against an agent that is unrealistically compliant, rigid, or behaviorally static may teach the wrong lessons rather than strengthen the intended skills\.

Our evaluation, therefore, establishes only one part of the case for simulation\-based training: whether the simulated interaction contains the behavioral challenges and responses that make meaningful practice possible\. The next step is to test whether those properties translate into actual learning\. Controlled studies with trainees and practitioners could examine whether repeated interaction with behaviorally realistic simulations improves questioning strategy, adaptation to difficult witnesses, or transfer to new scenarios, and whether differences in simulation fidelity affect those outcomes\. Connecting measures of behavioral realism to downstream learning would move evaluation beyond asking whether a simulated person appears plausible towards asking whether the simulation is sufficiently faithful to be useful\.

## Limitations and Future Work

There are a few limitations present in our work\. First, the evaluation corpus is drawn from a single opioid\-litigation context\. Broader evaluation across different cases would be a strong next step to determine more robust generalizability\. Furthermore, the emotion\-vector analysis should be understood as a proof of concept rather than a validated measure of witness affect\. Future studies should compare these representations with human annotations of demeanor for a more robust analysis\. More broadly, although simulation\-based learning may offer a way to use AI to strengthen rather than replace professional judgment, our study does not test whether repeated use ofWitnessSimmitigates deskilling or translates into measurable improvements in attorney learning\. Finally, controlled studies with trainees, instructors, and practicing attorneys should measure learning and performance with and without the use ofWitnessSimto more robustly determine usefulness in the context of adoption\.

## Conclusion

We introducedWitnessSim, a controllable, state\-based deposition simulator, together with an evaluation framework that separates behavioral realism from pedagogical usefulness\. Across adversarial tests, blinded plausibility judgments, and legally grounded examination tasks,WitnessSimgenerally maintained meaningful behavioral distinctions while responding to changes in attorney questioning\. At the same time, trajectory\-level analysis revealed systematic gaps: simulated testimony was smoother, more compressed, and distinguishable from real testimony over longer interactions\. These results illustrate why response\-level plausibility alone is insufficient for evaluating interactive simulations\. More broadly, this work frames behavioral fidelity as a multi\-turn evaluation problem\. A useful professional simulator must do more than produce plausible individual responses: it should maintain coherent behavioral boundaries, respond meaningfully to intervention, and exhibit realistic change across an interaction\. Although our evaluation does not establish attorney learning or full equivalence to human witnesses, it provides a framework for measuring these properties separately and identifying where simulation fidelity succeeds or breaks down\. As LLM\-based training and simulation become increasingly common, evaluating how behavior develops and changes across an interaction will be critical to developing high\-fidelity, useful simulations\.

## Impact Statement

This paper presents work whose goal is to advance the field of Machine Learning\. As with any generative system, models trained on legal data carry inherent risks, including the potential to produce inaccurate, biased, or hallucinated outputs that may not reliably reflect statutes, case law, or jurisdictional nuance\. In the legal domain, where errors can affect due process, client outcomes, and access to justice, such limitations warrant particular caution, and we emphasize that the methods presented here are intended to support, not substitute for, the judgment of qualified legal professionals\.

## Acknowledgements

We thank the team at Three Crowns for their contributions to the development ofWitnessSimand for bringing their practical expertise to the design and evaluation of the system\. We are especially grateful to Hugh Carlson, CEO of Three Crowns, and Nicholas Jampol, Partner at Davis Wright Tremaine LLP, for serving as expert annotators and contributing their time and expertise\.

## References

- Baoet al\.\(2026\)E\. Bao, A\. Perez, X\. Wang, and J\. ParaparEval4Sim: an evaluation framework for persona simulation\.arXiv:2603\.02876\.External Links:[Link](https://doi.org/10.48550/arXiv.2603.02876)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p3.1)\.
- Caldwell and Elliot \(2018\)H\. M\. Caldwell and D\. S\. ElliotAvoiding the wrecking ball of a disastrous cross examination: nine principles for effective cross examinations with supporting empirical evidence\.South Carolina Law Review70\(1\),pp\. 119\.External Links:[Link](https://scholarcommons.sc.edu/sclr/vol70/iss1/6/)Cited by:[§4\.2](https://arxiv.org/html/2608.13712#S4.SS2.p2.1)\.
- Chenet al\.\(2025\)G\. Chen, L\. Fan, Z\. Gong, N\. Xie, Z\. Li, Z\. Liu, C\. Li, Q\. Qu, H\. Alinejad\-Rokny, S\. Ni, and M\. YangAgentCourt: simulating court with adversarial evolvable lawyer agents\.InarXiv:2408\.08089,External Links:[Link](https://doi.org/10.48550/arXiv.2408.08089)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p4.1)\.
- Ghosalet al\.\(2019\)D\. Ghosal, N\. Majumder, S\. Poria, N\. Chhaya, and A\. GelbukhDialogueGCN: a graph convolutional neural network for emotion recognition in conversation\.InProceedings of EMNLP,External Links:[Link](https://arxiv.org/abs/1908.11540)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p5.1)\.
- Guhaet al\.\(2023\)N\. Guha, J\. Nyarko, D\. E\. Ho, C\. Ré, A\. Chilton, A\. Narayana, A\. Chohlas\-Wood, A\. Peters, B\. Waldon, D\. N\. Rockmore, D\. Zambrano, D\. Talisman, E\. Hoque, F\. Surani, F\. Fagan, G\. Sarfaty, G\. M\. Dickinson, H\. Porat, J\. Hegland, J\. Wu, J\. Nudell, J\. Niklaus, J\. Nay, J\. H\. Choi, K\. Tobia, M\. Hagan, M\. Ma, M\. Livermore, N\. Rasumov\-Rahe, N\. Holzenberger, N\. Kolt, P\. Henderson, S\. Rehaag, S\. Goel, S\. Gao, S\. Williams, S\. Gandhi, T\. Zur, V\. Iyer, and Z\. LiLegalBench: a collaboratively built benchmark for measuring legal reasoning in large language models\.arXiv:2308\.11462\.External Links:[Link](https://doi.org/10.48550/arXiv.2308.11462)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p4.1)\.
- Heet al\.\(2026\)J\. K\. He, F\. P\. S\. Wallis, A\. Gvirtz, and S\. RathjeArtificial intelligence chatbots mimic human collective behaviour\.British Journal of Psychology117\(2\),pp\. 761–776\.External Links:[Document](https://dx.doi.org/10.1111/bjop.12764),[Link](https://doi.org/10.1111/bjop.12764)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p2.1)\.
- Howard \(2010\)M\. A\. HowardMastering foolproof witness control on cross\-examination\.De Novo,pp\. 9–10\.External Links:[Link](https://digitalcommons.law.uw.edu/faculty-articles/521/)Cited by:[§4\.2](https://arxiv.org/html/2608.13712#S4.SS2.p2.1)\.
- Kautzman \(2010\)J\. F\. KautzmanControlling the difficult witness\.Note:The Indiana LawyerExternal Links:[Link](https://www.theindianalawyer.com/articles/25308-iba-controlling-the-difficult-witness)Cited by:[§4\.2](https://arxiv.org/html/2608.13712#S4.SS2.p2.1)\.
- Kim and Vossen \(2021\)T\. Kim and P\. VossenEmoRoBERTa: speaker\-aware emotion recognition in conversation with roberta\.InarXiv:2108\.12009,External Links:[Link](https://doi.org/10.48550/arXiv.2108.12009)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p5.1)\.
- Lutzet al\.\(2025\)M\. Lutz, I\. Sen, G\. Ahnert, E\. Rogers, and M\. StrohmaierThe prompt makes the person\(a\): a systematic evaluation of sociodemographic persona prompting for large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 23212–23237\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1261),[Link](https://aclanthology.org/2025.findings-emnlp.1261/)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p3.1)\.
- MacCarthy \(2007\)T\. F\. MacCarthyMacCarthy on cross\-examination\.American Bar Association\.External Links:ISBN 9781590318867Cited by:[§4\.2](https://arxiv.org/html/2608.13712#S4.SS2.p2.1)\.
- Melvin \(2015\)J\. MelvinTrial advocacy & writing: course materials and syllabus\.Note:Atlanta’s John Marshall Law SchoolSpring 2015External Links:[Link](https://www.johnmarshall.edu/wp-content/uploads/Trial-Advocay-Syllabus-Spring-2015.pdf)Cited by:[§4\.2](https://arxiv.org/html/2608.13712#S4.SS2.p2.1)\.
- Miller \(2021\)P\. MillerEffective deposition strategy and dealing with the evasive, non\-cooperative witness\.Note:Lawyer MindsExternal Links:[Link](https://www.lawyerminds.com/effective-deposition-strategy-and-dealing-with-the-evasive-non-cooperative-witness/)Cited by:[§4\.2](https://arxiv.org/html/2608.13712#S4.SS2.p2.1)\.
- Noh and Chang \(2024\)S\. Noh and H\. H\. ChangLLMs with personalities in multi\-issue negotiation games\.arXiv preprint arXiv:2405\.05248\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2405.05248),[Link](https://arxiv.org/abs/2405.05248)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p7.1)\.
- Parket al\.\(2023\)J\. S\. Park, J\. C\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. BernsteinGenerative agents: interactive simulacra of human behavior\.InarXiv:2304\.03442,External Links:[Link](https://doi.org/10.48550/arXiv.2304.03442)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p2.1)\.
- Parket al\.\(2024\)J\. S\. Park, C\. Q\. Zou, A\. Shaw, B\. M\. Hill, C\. J\. Cai, M\. R\. Morris, R\. Willer, P\. Liang, and M\. S\. BernsteinGenerative agent simulations of 1,000 people\.arXiv preprint arXiv:2411\.10109\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2411.10109),[Link](https://arxiv.org/abs/2411.10109)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p2.1)\.
- Picard \(1997\)R\. W\. PicardAffective computing\.MIT Press\.Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p5.1)\.
- Pozner and Dodd \(2018\)L\. S\. Pozner and R\. J\. DoddCross\-examination: science and techniques\.3 edition,LexisNexis,New York, NY\.External Links:ISBN 9781632843920Cited by:[§4\.2](https://arxiv.org/html/2608.13712#S4.SS2.p2.1)\.
- Rolf \(2025\)R\. E\. RolfUsing generative artificial intelligence in a contract simulation to promote student learning in business law\.Journal of Legal Studies Education42\(1\),pp\. 7–22\.External Links:[Document](https://dx.doi.org/10.1111/jlse.12154),[Link](https://doi.org/10.1111/jlse.12154)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p8.1)\.
- Samuelet al\.\(2025\)V\. Samuel, H\. P\. Zou, Y\. Zhou, S\. Chaudhari, A\. Kalyan, T\. Rajpurohit, A\. Deshpande, K\. R\. Narasimhan, and V\. MurahariPersonaGym: evaluating persona agents and LLMs\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 6999–7022\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.368),[Link](https://aclanthology.org/2025.findings-emnlp.368/)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p3.1)\.
- Shaikhet al\.\(2024\)O\. Shaikh, V\. Chai, M\. J\. Gelfand, D\. Yang, and M\. S\. BernsteinRehearsal: simulating conflict to teach conflict resolution\.InProceedings of CHI,External Links:[Link](https://doi.org/10.48550/arXiv.2309.12309)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p7.1)\.
- Shenet al\.\(2021\)W\. Shen, J\. Guo, S\. Tang, W\. Zhang, T\. Liu, and Y\. ChenDialogXL: all\-in\-one XLNet for multi\-party conversation emotion recognition\.InarXiv:2012\.08695,External Links:[Link](https://doi.org/10.48550/arXiv.2012.08695)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p5.1)\.
- Sofroniewet al\.\(2026\)N\. Sofroniew, I\. Kauvar, W\. Saunders, R\. Chen, T\. Henighan, S\. Hydrie, C\. Citro, A\. Pearce, J\. Tarng, W\. Gurnee, J\. Batson, S\. Zimmerman, K\. Rivoire, K\. Fish, C\. Olah, and J\. LindseyEmotion concepts and their function in a large language model\.Technical reportAnthropic\.External Links:[Link](https://transformer-circuits.pub/2026/emotions/index.html)Cited by:[Appendix K](https://arxiv.org/html/2608.13712#A11.p2.1),[§2](https://arxiv.org/html/2608.13712#S2.p6.1)\.
- Tanford \(1994\)J\. A\. TanfordKeeping cross\-examination under control\.American Journal of Trial Advocacy18,pp\. 245\.External Links:[Link](https://www.repository.law.indiana.edu/facpub/627/)Cited by:[§4\.2](https://arxiv.org/html/2608.13712#S4.SS2.p2.1)\.
- Vuet al\.\(2026\)H\. Vu, H\. A\. Nguyen, A\. V\. Ganesan, S\. Juhng, O\. N\. E\. Kjell, J\. Sedoc, M\. L\. Kern, R\. L\. Boyd, L\. Ungar, H\. A\. Schwartz, and J\. C\. EichstaedtPsychAdapter: adapting LLMs to reflect traits, personality, and mental health\.npj Artificial Intelligence2,pp\. 26\.External Links:[Document](https://dx.doi.org/10.1038/s44387-026-00071-9),[Link](https://doi.org/10.1038/s44387-026-00071-9)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p6.1)\.
- Wanget al\.\(2025\)P\. Wang, H\. Zou, H\. Chen, T\. Sun, Z\. Xiao, and F\. L\. OswaldPersonality structured interview for large language model simulation in personality research\.arXiv preprint arXiv:2502\.12109\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2502.12109),[Link](https://arxiv.org/abs/2502.12109)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p2.1)\.
- Yanget al\.\(2024\)D\. Yang, C\. Ziems, W\. Held, O\. Shaikh, M\. S\. Bernstein, and J\. MitchellSocial skill training with large language models\.arXiv:2404\.04204\.External Links:[Link](https://doi.org/10.48550/arXiv.2404.04204)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p7.1)\.
- Zhanget al\.\(2025\)K\. Zhang, J\. Li, Y\. Wu, H\. Li, C\. Luo, S\. Zou, Y\. Zhou, W\. Su, Q\. Ai, and Y\. LiuChinese court simulation with llm\-based agent system\.arXiv:2508\.17322\.External Links:[Link](https://doi.org/10.48550/arXiv.2508.17322)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p4.1)\.
- Zhanget al\.\(2026\)K\. Zhang, N\. Nadeem, L\. Zheng, D\. Stammbach, and P\. HendersonAI\-assisted moot courts: simulating justice\-specific questioning in oral arguments\.arXiv:2603\.04718\.External Links:[Link](https://doi.org/10.48550/arXiv.2603.04718)Cited by:[§2](https://arxiv.org/html/2608.13712#S2.p8.1),[§4\.0\.1](https://arxiv.org/html/2608.13712#S4.SS0.SSS1.p1.1)\.

## Appendix

## Appendix AArchetype Representation and Big Five Comparison

The simulator began from fourteen expert\-informed candidate archetypes\. After examining their locations in the six\-dimensional state space, ten retained archetypes were used in the present evaluation\. Table[5](https://arxiv.org/html/2608.13712#A1.T5)reports their complete six\-state attractor vectors\.

Table 5:Six\-state attractor vectors for the ten retained witness archetypes\. The dimensions are composure \(CC\), knowledge \(KK\), agreeableness \(AA\), verbosity \(VV\), rigidity \(RR\), and performance \(PP\)\.Archetype𝐂\\mathbf\{C\}𝐊\\mathbf\{K\}𝐀\\mathbf\{A\}𝐕\\mathbf\{V\}𝐑\\mathbf\{R\}𝐏\\mathbf\{P\}Combative0\.250\.550\.150\.750\.850\.50Cooperative0\.800\.850\.900\.500\.200\.65Defensive0\.550\.600\.300\.300\.750\.60Dogmatic0\.800\.750\.500\.200\.950\.70Inventive0\.700\.200\.600\.850\.250\.60Loquacious0\.600\.700\.600\.950\.250\.55Nervous0\.200\.450\.550\.750\.400\.25Neutral0\.550\.550\.550\.550\.500\.55Overconfident0\.900\.650\.750\.650\.800\.90Overprepared0\.800\.900\.600\.200\.800\.95For comparison, we construct the approximate Big Five projection used in the archetype discussion in Section[3\.1\.1](https://arxiv.org/html/2608.13712#S3.SS1.SSS1):

Π⁡\(C,K,A,V,R,P\)=\(1−C,A,V,1−R,P\),\\Pi\(C,K,A,V,R,P\)=\(1\-C,\\ A,\\ V,\\ 1\-R,\\ P\),corresponding approximately to neuroticism \(NN\), agreeableness \(AA\), extraversion \(EE\), openness \(OO\), and conscientiousness \(CbfC\_\{\\mathrm\{bf\}\}\)\. The knowledge coordinateKKis omitted because it has no direct analog in this approximation\.

Table 6:Approximate Big Five projection of the ten retained archetypes\.Archetype𝐍\\mathbf\{N\}𝐀\\mathbf\{A\}𝐄\\mathbf\{E\}𝐎\\mathbf\{O\}𝐂bf\\mathbf\{C\_\{\\mathrm\{bf\}\}\}Combative0\.750\.150\.750\.150\.50Cooperative0\.200\.900\.500\.800\.65Defensive0\.450\.300\.300\.250\.60Dogmatic0\.200\.500\.200\.050\.70Inventive0\.300\.600\.850\.750\.60Loquacious0\.400\.600\.950\.750\.55Nervous0\.800\.550\.750\.600\.25Neutral0\.450\.550\.550\.500\.55Overconfident0\.100\.750\.650\.200\.90Overprepared0\.200\.600\.200\.200\.95To examine how this projection changes the geometry among archetypes, we computed cosine similarity and angular separation for all\(102\)=45\\binom\{10\}\{2\}=45retained\-archetype pairs in both spaces\. Table[7](https://arxiv.org/html/2608.13712#A1.T7)reports the complete comparison\. We define

Δ​θ=θR5−θR6\.\\Delta\\theta=\\theta\_\{R^\{5\}\}\-\\theta\_\{R^\{6\}\}\.Thus, negative values indicate pairs that become more similar under the Big Five projection, whereas positive values indicate pairs that become more separated\.

Importantly, the projection does not uniformly compress all pairwise distances\. Rather, its limitation for the present application is that droppingKKcan collapse distinctions for which knowledge is the behaviorally important differentiating coordinate\. The Inventive–Loquacious pair provides the clearest example: its angular separation falls from approximately19\.2∘19\.2^\{\\circ\}in the six\-state space to5\.2∘5\.2^\{\\circ\}after the Big Five projection\.

Table 7:Pairwise archetype similarity in the six\-state space \(R6R^\{6\}\) and approximate Big Five space \(R5R^\{5\}\)\.Δ​θ=θR5−θR6\\Delta\\theta=\\theta\_\{R^\{5\}\}\-\\theta\_\{R^\{6\}\}; negative values indicate compression under the Big Five projection\.Paircos𝐑𝟔\\mathbf\{\\cos\_\{R^\{6\}\}\}𝜽𝑹𝟔\\boldsymbol\{\\theta\_\{R^\{6\}\}\}cos𝐑𝟓\\mathbf\{\\cos\_\{R^\{5\}\}\}𝜽𝑹𝟓\\boldsymbol\{\\theta\_\{R^\{5\}\}\}𝚫​𝜽\\boldsymbol\{\\Delta\\theta\}Neutral / Overconfident0\.9908\.0∘0\.89027\.1∘\+19\.1∘Defensive / Dogmatic0\.9898\.4∘0\.90025\.8∘\+17\.4∘Nervous / Overconfident0\.86829\.7∘0\.70345\.4∘\+15\.6∘Combative / Dogmatic0\.84732\.1∘0\.67847\.3∘\+15\.2∘Combative / Overprepared0\.81835\.1∘0\.63950\.3∘\+15\.2∘Combative / Overconfident0\.87129\.4∘0\.71444\.5∘\+15\.0∘Defensive / Overconfident0\.96116\.1∘0\.86630\.0∘\+14\.0∘Defensive / Overprepared0\.97712\.4∘0\.90125\.8∘\+13\.4∘Neutral / Overprepared0\.94020\.0∘0\.84132\.8∘\+12\.8∘Dogmatic / Neutral0\.93021\.5∘0\.83733\.2∘\+11\.7∘Nervous / Overprepared0\.75441\.0∘0\.61152\.3∘\+11\.3∘Dogmatic / Nervous0\.75441\.0∘0\.62751\.1∘\+10\.1∘Loquacious / Overprepared0\.82234\.7∘0\.73242\.9∘\+8\.2∘Loquacious / Overconfident0\.91523\.8∘0\.85331\.5∘\+7\.6∘Cooperative / Overprepared0\.90225\.6∘0\.84332\.5∘\+7\.0∘Combative / Cooperative0\.71044\.8∘0\.63150\.9∘\+6\.2∘Dogmatic / Loquacious0\.79237\.6∘0\.72443\.6∘\+6\.0∘Cooperative / Overconfident0\.92821\.9∘0\.88427\.9∘\+6\.0∘Inventive / Overconfident0\.91623\.6∘0\.87628\.8∘\+5\.2∘Cooperative / Dogmatic0\.85731\.0∘0\.81435\.6∘\+4\.5∘Combative / Neutral0\.88827\.3∘0\.85131\.7∘\+4\.3∘Overconfident / Overprepared0\.95816\.6∘0\.93620\.6∘\+4\.0∘Combative / Defensive0\.90924\.6∘0\.88427\.9∘\+3\.3∘Cooperative / Nervous0\.84532\.3∘0\.81935\.0∘\+2\.7∘Loquacious / Nervous0\.94519\.1∘0\.92921\.8∘\+2\.6∘Cooperative / Loquacious0\.93420\.9∘0\.92322\.7∘\+1\.8∘Dogmatic / Overconfident0\.95417\.5∘0\.94619\.0∘\+1\.5∘Cooperative / Defensive0\.84632\.2∘0\.83433\.5∘\+1\.3∘Combative / Loquacious0\.83633\.3∘0\.82734\.2∘\+1\.0∘Cooperative / Neutral0\.94718\.8∘0\.94319\.4∘\+0\.6∘Inventive / Overprepared0\.77639\.1∘0\.77039\.7∘\+0\.6∘Dogmatic / Inventive0\.75840\.7∘0\.75241\.2∘\+0\.5∘Defensive / Neutral0\.94519\.0∘0\.94419\.3∘\+0\.3∘Combative / Nervous0\.88028\.3∘0\.88228\.1∘\-0\.3∘Dogmatic / Overprepared0\.98510\.1∘0\.9898\.4∘\-1\.7∘Combative / Inventive0\.77139\.6∘0\.79137\.7∘\-1\.9∘Nervous / Neutral0\.92122\.9∘0\.93420\.9∘\-2\.0∘Inventive / Nervous0\.88028\.4∘0\.89925\.9∘\-2\.5∘Defensive / Loquacious0\.82934\.1∘0\.85731\.1∘\-3\.0∘Loquacious / Neutral0\.95517\.3∘0\.96914\.2∘\-3\.1∘Defensive / Nervous0\.79637\.2∘0\.84332\.5∘\-4\.7∘Defensive / Inventive0\.78538\.3∘0\.86130\.6∘\-7\.7∘Inventive / Neutral0\.92322\.7∘0\.97014\.0∘\-8\.6∘Cooperative / Inventive0\.88128\.2∘0\.94718\.8∘\-9\.4∘Inventive / Loquacious0\.94419\.2∘0\.9965\.2∘\-14\.0∘
## Appendix BQuestion\-Encoding Details

### B\.1Pressure\-Marker Dictionary

The pressure score uses a baseline of0\.100\.10and prespecified high\- and medium\-pressure linguistic markers\. Table[8](https://arxiv.org/html/2608.13712#A2.T8)reports the complete dictionary\.

Table 8:Prespecified linguistic markers used to construct the pressure score\.ComponentContributionMarkers / interpretationBaseline0\.100\.10Applied to every question before marker\-specific increments\.High\-pressure marker\+0\.20\+0\.20per matchisn’t it true;you’re lying;how do you explain;did you not;you knew;admit;deny\.Medium\-pressure marker\+0\.10\+0\.10per matchwhy did you;when did you;who authorized\.All markers are substring\-matched after lowercasing the question\. Multiple matches accumulate, after which the result is clipped to\[0,1\]\[0,1\]\. Operationally, the numerical weights encode relative pressure intensity: each high\-pressure marker contributes twice the increment of a medium\-pressure marker, while the baseline preserves a low nonzero pressure value even when none of the listed markers occurs\. These are prespecified design weights used to operationalize relative pressure rather than fitted estimates of a latent psychological quantity\.

For example,

> Isn’t it true you knew about this? Admit it\.

matchesisn’t it true,you knew, andadmit, giving

pt=0\.10\+0\.20\+0\.20\+0\.20=0\.70\.p\_\{t\}=0\.10\+0\.20\+0\.20\+0\.20=0\.70\.

### B\.2Sensitive\-Topic Tokenization

Sensitive topics are constructed during persona generation from the case materials\. Each topicτ\\tauis represented by a short string key, an intrinsic sensitivity weightστ∈\[0,1\]\\sigma\_\{\\tau\}\\in\[0,1\], and a basis describing why the topic is sensitive\.

For each question, both the question and topic key are lowercased and split on whitespace\. Let

Qt=set\(lower\(qt\)\.split\(\)\)Q\_\{t\}=\\operatorname\{set\}\\\!\\left\(\\operatorname\{lower\}\(q\_\{t\}\)\.\\operatorname\{split\}\(\)\\right\)and

Tτ=set\(lower\(τ\)\.split\(\)\)\.T\_\{\\tau\}=\\operatorname\{set\}\\\!\\left\(\\operatorname\{lower\}\(\\tau\)\.\\operatorname\{split\}\(\)\\right\)\.The implementation then computes the bag\-of\-words Jaccard similarity

Jt​\(τ\)=\|Qt∩Tτ\|\|Qt∪Tτ\|\.J\_\{t\}\(\\tau\)=\\frac\{\|Q\_\{t\}\\cap T\_\{\\tau\}\|\}\{\|Q\_\{t\}\\cup T\_\{\\tau\}\|\}\.
A topic is treated as matched only when

Jt​\(τ\)\>0\.15\.J\_\{t\}\(\\tau\)\>0\.15\.For a matched topic, its contribution is

Jt​\(τ\)​στ,J\_\{t\}\(\\tau\)\\sigma\_\{\\tau\},and the final sensitivity score is the maximum contribution across all matched topics:

st=maxτ\{στJt\(τ\)𝕀\[Jt\(τ\)\>0\.15\]\}\.s\_\{t\}=\\max\_\{\\tau\}\\left\\\{\\sigma\_\{\\tau\}J\_\{t\}\(\\tau\)\\mathbb\{I\}\[J\_\{t\}\(\\tau\)\>0\.15\]\\right\\\}\.
This construction makes matching sensitive to the wording and length of the question\. Short questions containing topic\-specific terms generally produce larger Jaccard scores, whereas additional unrelated words enlarge the union and can dilute the resulting similarity\.

## Appendix CBehavioral Justifications for State Dynamics

Each state equation contains an immediate forcing term and a smaller mean\-reversion term that pulls the state toward its archetype attractor\. The numerical coefficients are fixed design parameters of the simulator: the rationale below explains their relative function and scale rather than treating them as independently estimated behavioral parameters\.

Table[9](https://arxiv.org/html/2608.13712#A3.T9)reports each numerical update together with its behavioral interpretation\.

Table 9:Numerical state\-update equations and behavioral motivation\.StateNumerical updateCoefficient\-level rationaleComposure \(CC\)Δ​Ct=−0\.25​ξt​Ct\+0\.08​\(C0−Ct\)\\displaystyle\\Delta C\_\{t\}=\-0\.25\\,\\xi\_\{t\}C\_\{t\}\+0\.08\(C\_\{0\}\-C\_\{t\}\)The−0\.25\-0\.25forcing term makes stress produce a comparatively strong immediate loss of composure\. Multiplication byCtC\_\{t\}means a highly composed witness has more composure to lose, while an already\-rattled witness approaches a natural floor\. The smaller0\.080\.08term produces slow recovery toward baseline rather than an immediate reset\.Knowledge \(KK\)Δ​Kt=−0\.30​st​\(Kt−0\.1\)\+0\.06​\(1−pt\)​\(K0−Kt\)\\displaystyle\\Delta K\_\{t\}=\-0\.30\\,s\_\{t\}\(K\_\{t\}\-0\.1\)\+0\.06\(1\-p\_\{t\}\)\(K\_\{0\}\-K\_\{t\}\)The0\.300\.30sensitivity term makes topic\-specific stress the strongest immediate state erosion in the system and preserves a floor atK=0\.1K=0\.1\. The0\.060\.06recovery term is weaker and is multiplied by\(1−pt\)\(1\-p\_\{t\}\), so recall recovers primarily during low\-pressure questioning and can ratchet down during sustained high\-pressure examination\.Agreeableness \(AA\)Δ​At=−0\.15​pt​sgn⁡\(At−12\)​\|At−12\|\+0\.05​\(A0−At\)\\displaystyle\\Delta A\_\{t\}=\-0\.15\\,p\_\{t\}\\operatorname\{sgn\}\\\!\\left\(A\_\{t\}\-\\tfrac\{1\}\{2\}\\right\)\\left\|A\_\{t\}\-\\tfrac\{1\}\{2\}\\right\|\+0\.05\(A\_\{0\}\-A\_\{t\}\)The0\.150\.15pressure term pulls agreeableness toward the guarded midpointA=0\.5A=0\.5\. Thus, a highly cooperative witness becomes less accommodating under pressure, while an already\-hostile witness receives a small push toward neutral rather than becoming indefinitely more hostile\. The0\.050\.05term gradually restores the archetype baseline\.Verbosity \(VV\)Δ​Vt=0\.20​st​\(1−Vt\)−0\.20​𝕀lead,t​Vt\+0\.05​\(V0−Vt\)\\displaystyle\\Delta V\_\{t\}=0\.20\\,s\_\{t\}\(1\-V\_\{t\}\)\-0\.20\\,\\mathbb\{I\}\_\{\\mathrm\{lead\},t\}V\_\{t\}\+0\.05\(V\_\{0\}\-V\_\{t\}\)The two immediate effects have equal magnitude but opposite direction\. Sensitivity pushes verbosity upward through over\-explanation, with\(1−Vt\)\(1\-V\_\{t\}\)creating a ceiling effect\. Leading yes/no form pushes verbosity downward, withVtV\_\{t\}creating a floor effect\. The0\.050\.05mean\-reversion term restores baseline once those forces disappear\.Rigidity \(RR\)Δ​Rt=0\.12​pt​\(1−Rt\)−0\.03​\(Rt−R0\)\\displaystyle\\Delta R\_\{t\}=0\.12\\,p\_\{t\}\(1\-R\_\{t\}\)\-0\.03\(R\_\{t\}\-R\_\{0\}\)Pressure increases rigidity with a0\.120\.12forcing coefficient, while\(1−Rt\)\(1\-R\_\{t\}\)prevents unbounded growth nearR=1R=1\. The0\.030\.03recovery coefficient is the smallest mean\-reversion term among the six states, so entrenchment dissipates especially slowly after pressure is removed\.Performance \(PP\)Δ​Pt=−0\.18​pt​Pt​ϕt\+0\.04​\(P0−Pt\)\\displaystyle\\Delta P\_\{t\}=\-0\.18\\,p\_\{t\}P\_\{t\}\\phi\_\{t\}\+0\.04\(P\_\{0\}\-P\_\{t\}\)The0\.180\.18forcing term makes pressure degrade performance in proportion to both current polish and accumulated fatigue\. A highly prepared witness therefore has more performance to lose\. The smaller0\.040\.04recovery term allows only modest restoration toward the baseline\.Taken together, the coefficient magnitudes encode a hierarchy of immediate effects in the implementation: sensitive\-topic knowledge erosion \(0\.300\.30\) and composure loss \(0\.250\.25\) are strongest; verbosity uses symmetric0\.200\.20terms; performance degrades at0\.180\.18under pressure and fatigue; agreeableness responds at0\.150\.15; and rigidity increases at0\.120\.12\. Mean\-reversion coefficients are deliberately smaller in the implemented dynamics, ranging from0\.030\.03to0\.080\.08, so behavioral changes persist across multiple turns rather than disappearing after a single question\.

### C\.1Fatigue Multiplier

Performance degradation is additionally scaled by

ϕt=min⁡\(1,t20\)\.\\phi\_\{t\}=\\min\\left\(1,\\frac\{t\}\{20\}\\right\)\.The multiplier is not a separate state variable\. It increases linearly over the first twenty turns and remains at one thereafter\.

Table 10:Illustrative values of the fatigue multiplier\.Turnϕ𝒕\\boldsymbol\{\\phi\_\{t\}\}Effect on pressure\-driven performance loss10\.055% of the full fatigue\-scaled rate100\.5050% of the full fatigue\-scaled rate20 and later1\.00Full fatigue\-scaled rateThe intended behavior is that a prepared witness can remain comparatively stable early in an examination, while the same pressure becomes more costly later\. For example, a high\-pressure question at turn 1 is scaled byϕ1=0\.05\\phi\_\{1\}=0\.05, whereas the same question at turn 20 or later is scaled byϕt=1\\phi\_\{t\}=1\.

## Appendix DPer\-Topic Recall Quality

In addition to the global knowledge stateKtK\_\{t\}, the simulator maintains a topic\-specific recall\-quality scalarρτ,t∈\[0\.1,1\]\\rho\_\{\\tau,t\}\\in\[0\.1,1\]\. The two variables serve different purposes:KtK\_\{t\}represents global recall health within the state dynamics, whereasρτ,t\\rho\_\{\\tau,t\}tracks wear on a particular topic and directly informs the memory description passed to the generation model\.

For topicτ\\tau, the update is

ρτ,t\+1=\{max⁡\(0\.1,ρτ,t−0\.25​pt\),if topicτis queried at turnt,min⁡\(1,ρτ,t\+0\.02\),otherwise\.\\rho\_\{\\tau,t\+1\}=\\begin\{cases\}\\max\\\!\\left\(0\.1,\\rho\_\{\\tau,t\}\-0\.25p\_\{t\}\\right\),&\\text\{if topic $\\tau$ is queried at turn $t$\},\\\\\[4\.0pt\] \\min\\\!\\left\(1,\\rho\_\{\\tau,t\}\+0\.02\\right\),&\\text\{otherwise\}\.\\end\{cases\}
Thus, at maximum pressure, one focused hit reduces topic\-specific recall by0\.250\.25\. Repeated focused questioning can substantially degrade a topic within three to four high\-pressure hits\. The floor at0\.10\.1preserves residual memory rather than permitting total loss\.

When a topic is not questioned, recall quality recovers by only0\.020\.02per turn\. Moving from the floor0\.10\.1back to1\.01\.0therefore requires approximately 45 untouched turns, making severe topic\-specific degradation effectively persistent within a typical examination\.

The value ofρτ,t\\rho\_\{\\tau,t\}is converted into a recall\-quality label in the memory portion of the system prompt\. Accordingly, the global stateKtK\_\{t\}determines how overall recall evolves, whileρτ,t\\rho\_\{\\tau,t\}gives the generation model topic\-specific information about how reliably the witness should recall the material at the current turn\.

## Appendix EDerived Behavioral Score Rationale

The four derived behavioral scores are deterministic summaries of the six\-dimensional state rather than additional latent variables\. Table[11](https://arxiv.org/html/2608.13712#A5.T11)records the numerical weights and their intended interpretation\.

Table 11:Numerical construction and interpretation of the derived behavioral scores\.ScoreFormulaWeight\-level rationaleConsistency0\.5​Kt\+0\.4​Rt−0\.1​\(1−Kt\)​\(1−Rt\)\\displaystyle 0\.5K\_\{t\}\+0\.4R\_\{t\}\-0\.1\(1\-K\_\{t\}\)\(1\-R\_\{t\}\)Knowledge receives the largest weight \(0\.50\.5\) because factual recall is the primary basis of a coherent account\. Rigidity receives0\.40\.4because consistency also requires maintaining the account across questioning\. The0\.10\.1interaction penalty matters primarily when both knowledge and rigidity are low, corresponding to a witness who neither knows the story nor reliably adheres to it\.Evasion0\.5​\(1−Kt\)\+0\.3​\|2​Vt−1\|\+0\.2​Pt\\displaystyle 0\.5\(1\-K\_\{t\}\)\+0\.3\|2V\_\{t\}\-1\|\+0\.2P\_\{t\}Lack of knowledge receives the largest contribution \(0\.50\.5\)\. Deviation of verbosity from its midpoint receives0\.30\.3, allowing both extreme terseness and extreme verbosity to count as evasive behavior\. Performance contributes0\.20\.2because a highly prepared witness can use polish strategically to avoid a direct answer\.Realismmax⁡\(0,1−0\.3​\|2​Ct−1\|−0\.3​\|2​Kt−1\|−2​max⁡\(0,Pt−0\.8\)\)\\displaystyle\\max\\\!\\Big\(0,\\,1\-0\.3\|2C\_\{t\}\-1\|\-0\.3\|2K\_\{t\}\-1\|\-2\\max\(0,P\_\{t\}\-0\.8\)\\Big\)The score begins at one\. Extreme composure and extreme recall each carry a maximum penalty of0\.30\.3, reflecting the idea that either implausibly perfect control or complete breakdown can reduce behavioral realism\. Performance is unpenalized throughP=0\.8P=0\.8and is then penalized with slope22; atP=1P=1, this contributes a maximum penalty of0\.40\.4\. The outer maximum keeps the score nonnegative\.Adversarial0\.4​\(1−Kt\)\+0\.3​\(1−At\)\+0\.2​Rt\+0\.1​\(1−Ct\)\\displaystyle 0\.4\(1\-K\_\{t\}\)\+0\.3\(1\-A\_\{t\}\)\+0\.2R\_\{t\}\+0\.1\(1\-C\_\{t\}\)Knowledge gaps receive the largest weight \(0\.40\.4\), followed by non\-cooperation \(0\.30\.3\), because unreliability and unwillingness to answer are the principal sources of examination difficulty\. Rigidity contributes0\.20\.2, while composure loss contributes0\.10\.1as an additional source of unpredictability\.
## Appendix FCorpus Sampling and Warm\-Up Procedure

### F\.1Candidate Pool Construction

The candidate pool comprised 1,169 documents from the UCSF–JHU Opioid Industry Documents deposition\-transcript archive\. The archive index contains one row per document across fourteen litigation collections, and the corresponding transcript directory contains one text file for each indexed document\. The index was obtained through a manual export from the UCSF Industry Documents Library rather than constructed by the evaluation pipeline\.

After the initial random sample was drawn, the selected documents were manually reviewed\. Documents identified as non\-deposition legal filings or otherwise unsuitable transcripts were replaced with additional randomly sampled, verified deposition documents\. The final evaluation corpus therefore preserves the random sampling design while incorporating a manual eligibility\-validation stage\.

### F\.2Supplementary Case\-Document Retrieval

For each selected document, the pipeline retrieved up to five supplementary records using full\-text keyword search against the UCSF Industry Documents Library’s public Solr index\. Retrieval was scoped to the witness’s case when possible, with collection\-level fallback\. The search proceeded through five priority tiers:

1. 1\.witness name combined with declaration, affidavit, or statement terms;
2. 2\.witness name combined with case\-specific topic keywords;
3. 3\.witness name ordered by relevance;
4. 4\.witness name under a server\-side random sort; and
5. 5\.an unrestricted document\-type fallback\.

When witness\-name extraction was uncertain, a case\-level search was used instead\.

### F\.3Case\-Preparation Extraction

A singleclaude\-sonnet\-5call processed up to 500 attorney questions from each transcript together with detected exhibit references\. The model was instructed to return one grounded JSON object and to assignnullwhenever a requested field was not supported by the source material\.

The extracted fields were:

- •entity name;
- •witness role;
- •a sensitive topic and its first occurrence;
- •a neutral control topic;
- •an aggravating fact and its first occurrence;
- •a verified exhibit reference; and
- •a yes/no\-answerable prior\-position probe\.

These fields were subsequently used to construct case\-grounded adversarial and pedagogical evaluation prompts\.

### F\.4Warm\-Up Question Selection

Each simulated witness completed a warm\-up using up to ten attorney questions from the beginning of the corresponding transcript\. When the primary examining attorney could be identified, only that attorney’s questions were eligible\.

Questions were restricted to turns occurring before the earliest real\-transcript occurrence of any field later used by the evaluation: the sensitive topic, aggravating fact, exhibit, or prior\-position probe\. Formally, the cutoff was the minimum availablefirst\_turn\_indexacross these four fields\. If no cutoff was available, all early questions remained eligible\. The first ten eligible questions were used; when fewer than ten were available, all eligible questions were retained without padding\.

Table 12:Corpus construction and evaluation flow\.StageSourceNNNotesMaster candidate poolUCSF–JHU opioid transcript index1,169Manual export from the UCSF Industry Documents Library\.Initial random sampleInitial sample file300Random draw from the 1,169\-document candidate archive\.Manual eligibility validationValidated sample records300Sampled documents were reviewed, and non\-deposition filings were replaced with additional randomly sampled eligible documents\.Final validated sampleFinal document sample300Unique document IDs; repeated witness identities were permitted\.Pedagogical batteryFinal pedagogical outputs300No final hard\-failure records\.Adversarial batteryFinal adversarial outputs300No final hard\-failure or insufficient\-content exclusions\.Expert plausibility evaluationPaired attorney\-evaluation instruments99Fifty exchanges were evaluated by one attorney and forty\-nine by another\.

## Appendix GShared Generation, Judging, and Human\-Validation Setup

### G\.1Dialogue Generation

Witness responses were generated withclaude\-sonnet\-5\. The same underlying model was used for persona construction and LLM\-judge scoring, although each role received a distinct prompt and input context\.

Dialogue generation used a maximum of 512 output tokens with extended thinking disabled\. Temperature, top\-pp, and seed were not explicitly set and therefore followed the API defaults\. The turn\-level generation prompt incorporated the witness persona, current simulator state, memory, fired events, and conversation history\. Each witness response was generated once; the evaluation pipeline did not use response resampling or majority voting\.

### G\.2LLM Judge

All judge\-derived evaluation outcomes were scored usingclaude\-sonnet\-5\. The judge used a common rubric\-based prompt framework across tests, with test\-specific scoring definitions and context\.

Extended thinking was disabled for all judge calls\. Temperature, top\-pp, and seed were not explicitly specified and therefore followed the API defaults\. Maximum output length depended on the scope of the judgment: 1,024 tokens for single\-response judgments, 1,536 tokens for whole\-sequence judgments, and 2,048 tokens for per\-turn judgments made in the context of a complete sequence\.

#### G\.2\.1Rubric Prompt

Each judge call used the following system\-prompt template:

> You are a legal deposition analyst scoring a witness’s response\(s\) for the evaluation test ‘‘\{test\_name\}\.’’ Score using ONLY the definitions below \-\-\- judge the underlying function/pattern being described, not surface string\-matching against any example phrases given\. FIELDS TO SCORE: \{field\_defs\} \{extra\_instructions\} Return ONLY valid JSON in exactly this shape, no markdown fences, no commentary: \{json\_shape\}

Thefield\_defsblock rendered the complete scale or label definitions for every rubric field\. Theextra\_instructionsblock contained test\-specific clarifications, worked examples, or calibration examples where applicable\. The requiredjson\_shapespecified the expected output fields and allowable values\.

Additional case\- or turn\-specific information was provided through anextra\_contextstring prepended to the user message rather than inserted into the system prompt\.

For a single\-response judgment, the user message had the form:

> \{extra\_context\} ATTORNEY QUESTION: \{question\} WITNESS REPLY: \{reply\}

#### G\.2\.2Judgment Scope

Three judge variants were used depending on the structure of the evaluation:

- •judge\_single\_responsereceives one attorney question and one witness reply and scores that response independently;
- •judge\_turn\_sequencereceives the complete multi\-turn exchange and returns one shared sequence\-level judgment; and
- •judge\_turn\_sequence\_per\_turnreceives the complete sequence as context but produces a separate judgment for each turn, allowing each response to be interpreted in light of the surrounding interaction\.

The judge received only rendered dialogue and the contextual information required by the scoring rubric\. It did not receiveWitnessSim’s internal six\-dimensional state vector, state deltas, or other latent simulator variables\.

The judge was not blinded to the evaluation condition\. The test name was explicitly included in every system prompt, and the relevant archetype or behavioral condition was available in the test\-specific context where necessary for applying the rubric\. Accordingly, these scores should be interpreted as rubric\-based behavioral measurements rather than blinded classification of an unknown experimental condition\.

#### G\.2\.3Worked Rubric Example: Cooperative Self\-Incrimination

The following is the rendered system prompt used for the Cooperative Self\-Incrimination evaluation\. In the implementation that generated these results, this test retained the development labelAC2; it is reported as AC1 in the paper following the final four\-test numbering scheme\.

> You are a legal deposition analyst scoring a witness’s response\(s\) for the evaluation test ‘‘AC2’’\. Score using ONLY the definitions below \-\-\- judge the underlying function/pattern being described, not surface string\-matching against any example phrases given\. FIELDS TO SCORE: \- ‘‘framing\_acceptance’’ \(scale, 0 to 2\) 2 = fully accepts the self\-incriminating framing \(‘‘yes, I dropped the ball’’\) 1 = partially accepts, qualifies without fully owning it 0 = declines the self\-incriminating framing \- ‘‘self\_protective\_response’’ \(label\_set\) FULL\_ADMISSION: agrees to the self\-blaming framing essentially as\-is QUALIFIED\_ADMISSION: concedes something went wrong but reframes cause/responsibility DEFLECT\_TO\_PROCESS: redirects to team/process/systemic factors rather than personal fault POLITE\_REFUSAL: declines the framing while staying warm and cooperative in tone DIRECT\_REFUSAL: declines the framing plainly, without much softening \- ‘‘tone’’ \(scale, 0 to 2\) 2 = cooperative, helpful 1 = neutral, professional 0 = guarded, resistant, hostile Return ONLY valid JSON in exactly this shape, no markdown fences, no commentary: \{"framing\_acceptance": <int 0\-2\>, "self\_protective\_response": "<one of the labels above\>", "tone": <int 0\-2\>\}

#### G\.2\.4Output Parsing and Failures

Judge outputs were parsed by extracting the substring beginning with the first opening brace and ending with the final closing brace, followed by standard JSON parsing\. The pipeline did not retry malformed judge outputs, impute missing values, or conduct a separate field\-presence validation step after parsing\.

Judge\-call or parsing failures propagated to the outer per\-test error handler, which recorded the error and continued processing the remaining evaluation items\. Thus, the reported judge measurements derive from one judge call per scored response or sequence rather than repeated sampling or post hoc consensus\.

### G\.3Human Annotation

Human annotation was used to validate the automated behavioral measurements rather than to replace LLM judging across the complete evaluation corpus\.

Annotators received a self\-contained annotation instrument for each evaluated witness–test pair\. Each instrument contained:

- •a description of the scope of the evaluation test;
- •the same scoring\-field definitions supplied to the LLM judge;
- •the relevant witness response or multi\-turn transcript; and
- •a blank scoring template\.

The LLM judge’s scores were excluded from the annotation materials\. Annotators therefore assigned labels independently and without access to the automated judgment\. Human labels were subsequently converted to structured data and aligned with the corresponding LLM\-judge outputs for reliability analysis\.

Providing human annotators with the same field definitions used by the LLM judge was intended to make disagreement interpretable as disagreement in application of the rubric rather than disagreement caused by different task definitions\.

The contextual\-plausibility evaluation described separately in Appendix[I](https://arxiv.org/html/2608.13712#A9)used a distinct blinded pairwise instrument and is not included in the judge\-reliability analyses below\.

### G\.4Human–Judge Agreement

We assessed agreement between the automated judge and human annotations using raw agreement together with chance\-corrected agreement statistics\. Because several evaluation outcomes were substantially imbalanced, we report prevalence\-adjusted bias\-adjusted kappa \(PABAK\) and Gwet’s AC1 where available\.

For a rubric field containingkkmutually exclusive categories, PABAK was calculated using a chance\-agreement term of

pe=1k,p\_\{e\}=\\frac\{1\}\{k\},\(12\)
rather than applying the binarype=\.5p\_\{e\}=\.5assumption to multicategory fields\. The resulting statistic is

PABAK=po−1/k1−1/k,\\mathrm\{PABAK\}=\\frac\{p\_\{o\}\-1/k\}\{1\-1/k\},\(13\)
wherepop\_\{o\}denotes observed agreement\. For a binary outcome, this reduces to the conventional expression2​po−12p\_\{o\}\-1\.

We treat PABAK≥\.70\\geq\.70as evidence of strong human–judge agreement\. Fields with PABAK between\.60\.60and\.69\.69are treated as moderately validated and are retained for quantitative use only when agreement is corroborated by an alternative chance\-corrected measure\. Fields with PABAK<\.60<\.60are disclosed for transparency but are not used as validated quantitative evidence for main\-text claims\. Deterministic measurements, including response length and exact string\-similarity checks, do not rely on the LLM judge and are therefore outside this criterion\.

#### G\.4\.1Pedagogical Evaluation Reliability

Table[13](https://arxiv.org/html/2608.13712#A7.T13)reports the available human–judge agreement results for the pedagogical evaluation\.

Table 13:Human–LLM agreement for pedagogical evaluation fields\.kkdenotes the number of response categories\.TestFieldkknnAgreementPABAKGwet AC1U1tone35048\.4%––U1pushback35054\.0%––T1residual\-resistance flag \(Turn 5\)24889\.6%\.792\.820T1substantive\-control flag \(Turn 5\)24891\.7%\.833\.898T3substantive\-control flag \(Turns 4–5\)24795\.7%\.915\.954T3residual\-resistance flag \(Turns 4–5\)24793\.6%\.872\.926T2turn1\_to\_turn4\_trajectory35076\.0%\.640\.732T2stays\_focused25058\.0%\.160\.247T2economical\_delivery250100\.0%1\.0001\.000T2compliance\_with\_narrowing, Turn 235058\.0%\.370\.496T2compliance\_with\_narrowing, Turn 335098\.0%\.970\.980T2turn4\_binary45082\.0%\.760\.766For U1, the ordinaltoneandpushbackfields were summarized using weighted Cohen’sκ\\kappa, yieldingκw=\.294\\kappa\_\{w\}=\.294andκw=\.298\\kappa\_\{w\}=\.298, respectively\. These subjective rubric fields are therefore not used as validated quantitative evidence in the main text; the principal U1 response\-length outcome is deterministic\.

The T2turn1\_to\_turn4\_trajectoryclassification falls within the moderate\-agreement range under PABAK \(\.640\.640\), but this result is corroborated by Gwet’s AC1 of\.732\.732\. We therefore retain the trajectory classification while interpreting it as lower\-confidence than fields exceeding the\.70\.70PABAK threshold\.

Several finer\-grained T2 fields did not meet the reliability criterion\. In particular,stays\_focused\(PABAK=\.160=\.160\) and Turn 2compliance\_with\_narrowing\(PABAK=\.370=\.370\) are not used as validated quantitative outcomes\. Turn 3compliance\_with\_narrowingshowed high agreement \(PABAK=\.970=\.970; AC1=\.980=\.980\), but Turn 3 is the attorney’s meta\-instruction to “please answer only that question” rather than a substantive question and is therefore not treated as another point in the Turn 2\-to\-Turn 4 progression comparison\.

#### G\.4\.2Adversarial Evaluation Reliability

Table[14](https://arxiv.org/html/2608.13712#A7.T14)reports the corresponding validation results for the adversarial evaluation\. The complete adversarial protocols, rubrics, and statistical analyses are reported separately in Appendix[H](https://arxiv.org/html/2608.13712#A8)\.

Table 14:Human–LLM agreement for adversarial evaluation fields\.kkdenotes the number of response categories\.TestFieldkknnAgreementPABAKGwet AC1AC1framing\_acceptance34868\.8%\.531\.567AC1self\_protective\_response\(5\-way\)54944\.9%\.311\.320AC1self\_protective\_response\(binary full\-admission collapse\)24991\.8%\.837\.911AC1tone34971\.4%\.571\.634AC1floor\_pass/ fullpass\_rule24887\.5%\.750\.858AC2pass/fail derived frombreak\_turn25296\.2%\.923\.948AC3evasion\_intensity, trivial questions314787\.8%\.816\.839AC3evasion\_intensity, sensitive questions314678\.1%\.671\.716AC4irritation\_turn\-based composite pass/fail2150–\.987–AC4ever irritated2150–1\.0001\.000AC4exactirritation\_marker, Turn 2315081\.3%\.720\.744AC4exactirritation\_marker, Turn 3315060\.7%\.410\.477AC4exactirritation\_marker, Turn 4315062\.7%\.440\.525AC4exactirritation\_marker, Turn 5314969\.1%\.537\.620For AC1, the five\-wayself\_protective\_responseclassification showed insufficient agreement for quantitative use \(PABAK=\.311=\.311\)\. Collapsing the field to the distinction relevant to the primary behavioral claim—whether the witness made aFULL\_ADMISSIONor retained some self\-protective boundary—produced substantially stronger agreement \(PABAK=\.837=\.837; Gwet’s AC1=\.911=\.911\)\. The binary collapse, rather than the fine\-grained five\-way distribution, is therefore used for the primary quantitative claim\. The complete AC1 pass criterion also showed strong agreement \(PABAK=\.750=\.750; AC1=\.858=\.858\)\.

For the AC1 composite criterion, the six cases classified as failures by human annotators in the 48\-item validation sample were additionally reviewed against the underlying transcript and literal rubric definition\. In those disputed cases, the automated classifications were consistent with the specified rule\. This review was used to inspect disagreement rather than to replace the independently produced human labels in the reliability calculation\.

AC2 pass/fail classification showed strong agreement \(PABAK=\.923=\.923; AC1=\.948=\.948\)\. For AC3, trivial\-question evasion also showed strong agreement \(PABAK=\.816=\.816; AC1=\.839=\.839\)\. Sensitive\-question evasion fell within the moderate PABAK range \(\.671\.671\) but was corroborated by Gwet’s AC1 of\.716\.716, and is therefore retained under the criterion above\.

For AC4, exact three\-level irritation intensity was reliable at Turn 2 \(PABAK=\.720=\.720\) but fell below the quantitative inclusion criterion at Turns 3–5\. We therefore do not rely on late\-turn irritation magnitude for the primary AC4 claim\. Instead, the analysis uses irritation onset together with the deterministic consistency and repetition checks described in Appendix[H](https://arxiv.org/html/2608.13712#A8)\. Human–judge agreement for the irritation\-onset\-based composite classification was high \(PABAK=\.987=\.987;n=150n=150\)\. Gwet’s AC1 is not reported for that composite because a faithful categorical reconstruction of its multiple constituent conditions was not available for that calculation\.

### G\.5Adjudication

No single global disagreement\-resolution rule was applied across all evaluation tests\. Human annotations remained independent for the reliability calculations above\. Where individual borderline cases required substantive interpretation during downstream analysis, they were reviewed directly against the corresponding transcript and the test\-specific rubric rather than resolved through a universal adjudication rule\.

## Appendix HComplete Adversarial\-Test Protocols and Results

We evaluated behavioral robustness using four adversarial tests targeting distinct failure modes: unrestricted agreement with adverse framing \(AC1\), either premature resistance or unrestricted premise acceptance under escalating leading questions \(AC2\), indiscriminate evasion \(AC3\), and contradiction or mechanical repetition under repeated questioning \(AC4\)\.

All tests began from the standard warmed session described in Section[F\.4](https://arxiv.org/html/2608.13712#A6.SS4)\. Where multiple conditions were compared, the session was cloned so that each condition began from an identical conversational state\. AC1 and AC2 were evaluated on the Cooperative archetype, AC3 on the Evasive archetype, and AC4 on the Cooperative, Evasive, and Combative archetypes\. Each test was evaluated across the same 300 case\-grounded contexts\.

### H\.1Scoring and Human Validation

Except for the deterministic repetition check in AC4, behavioral outcomes were scored by the LLM judge described in Section[G\.2](https://arxiv.org/html/2608.13712#A7.SS2)\. The judge received the generated question and response text together with explicit field definitions and test\-specific instructions\. It did not receiveWitnessSim’s latent state variables\.

These outcomes were therefore not manually coded across the full 300\-context evaluation set\. Instead, human annotators independently scored subsets of responses using the same rubric definitions given to the LLM judge\. Annotators did not see the LLM judge’s scores\. Human labels were used to assess the reliability of the automated measurements rather than to replace the automated scores over the entire evaluation corpus\.

For fields with substantial prevalence imbalance, we report prevalence\-adjusted bias\-adjusted kappa \(PABAK\) together with Gwet’s AC1 where available\. Finer\-grained fields with inadequate human–judge agreement are retained for descriptive inspection but are not treated as validated quantitative evidence\.

### H\.2AC1: Cooperative Self\-Incrimination

##### Design and rationale\.

AC1 tests whether the Cooperative archetype has a meaningful lower bound on agreeableness\. A cooperative witness should remain responsive to an adverse question without accepting personally incriminating framing merely because the examiner invites agreement\. Conversely, maintaining that boundary should not require the witness to become hostile or uncharacteristically resistant\.

The test consists of a single question injected from the standard warmed and forked session state\. The question is grounded in theaggravating\_factextracted from the corresponding real transcript:

> “Just to help me understand—would you personally say you dropped the ball when it came to\{aggravating fact\}?”

Because the probe follows the standard warm\-up, the witness has already established its Cooperative interactional style before encountering the adversarial question\.

##### Scoring\.

Responses were scored by the rubric\-based LLM judge described in Section[G\.2](https://arxiv.org/html/2608.13712#A7.SS2)\. Human annotators independently scored a validation subset using the same definitions and without access to the LLM judge’s scores\.

Three rubric fields characterize the response\.

framing\_acceptancemeasures the degree to which the witness accepts the examiner’s self\-incriminating characterization:

- •0: declines the framing;
- •1: partially accepts or qualifies the framing without fully assuming personal fault;
- •2: fully accepts the self\-incriminating framing\.

self\_protective\_responseclassifies how the witness handles the accusation:

- •FULL\_ADMISSION: accepts the accusation substantially as presented;
- •QUALIFIED\_ADMISSION: concedes some problem while qualifying personal responsibility;
- •DEFLECT\_TO\_PROCESS: redirects responsibility toward a team, process, or institutional structure;
- •POLITE\_REFUSAL: rejects the framing while maintaining a cooperative tone;
- •DIRECT\_REFUSAL: directly rejects the framing\.

Finally,tonemeasures whether resistance to the premise causes the witness to abandon the Cooperative archetype\.

##### Primary self\-protection outcome\.

For the primary analysis, the five\-wayself\_protective\_responsefield was collapsed into a binary measure distinguishing full admission from all responses that retained some self\-protective boundary:

self\_protection=𝕀\{self\_protective\_response≠FULL\_ADMISSION\}\.\\mathrm\{self\\\_protection\}=\\mathbb\{I\}\\left\\\{\\texttt\{self\\\_protective\\\_response\}\\neq\\texttt\{FULL\\\_ADMISSION\}\\right\\\}\.
This binary outcome directly captures the behavioral boundary of interest: whether cooperation causes the witness to accept personal blame without qualification or resistance\.

Across 300 contexts, 297 witnesses preserved this boundary:

297/300=99\.0%,297/300=99\.0\\%,
with a Wilson 95% confidence interval of\[97\.1%,99\.7%\]\[97\.1\\%,99\.7\\%\]\.

Human validation of this binary classification was strong\. Across 49 independently annotated responses, raw human–LLM agreement was 91\.8%, with PABAK=\.837=\.837and Gwet’s AC1=\.911=\.911\.

##### Complete behavioral pass criterion\.

We additionally required that self\-protection be maintained without the witness abandoning its Cooperative persona\. The complete criterion was:

floor\_pass=𝕀\{framing\_acceptance≤1∧self\_protective\_response≠FULL\_ADMISSION\},\\mathrm\{floor\\\_pass\}=\\mathbb\{I\}\\left\\\{\\texttt\{framing\\\_acceptance\}\\leq 1\\land\\texttt\{self\\\_protective\\\_response\}\\neq\\texttt\{FULL\\\_ADMISSION\}\\right\\\},
and

passAC1=𝕀\{floor\_pass=1∧tone≥1\}\.\\mathrm\{pass\}\_\{\\mathrm\{AC1\}\}=\\mathbb\{I\}\\left\\\{\\mathrm\{floor\\\_pass\}=1\\land\\texttt\{tone\}\\geq 1\\right\\\}\.
The complete criterion produced the same full\-sample result: 297/300 witnesses \(99\.0%, Wilson 95% CI \[97\.1, 99\.7\]\) passed\. The three failures were the same three responses classified asFULL\_ADMISSION; neither framing acceptance nor tone produced an additional failure among responses that passed the binary self\-protection criterion\.

Human validation of the complete binary pass criterion was also strong\. Across 48 annotated responses, raw human–LLM agreement was 87\.5% and PABAK was\.750\.750\.

##### Descriptive response strategies\.

The underlying five\-way response distribution was:

Response strategyNNRateDEFLECT\_TO\_PROCESS10535\.0%DIRECT\_REFUSAL9230\.7%QUALIFIED\_ADMISSION7725\.7%POLITE\_REFUSAL237\.7%FULL\_ADMISSION31\.0%
We treat this finer\-grained distribution as descriptive rather than as a validated quantitative result because agreement on the full five\-category classification was lower than agreement on the binary self\-protection outcome\. The primary claim therefore concerns whether the witness preserved a self\-protective boundary, not the particular strategy through which that boundary was expressed\.

### H\.3AC2: Cooperative Runaway Leading Chain

##### Design and rationale\.

AC2 tests the same boundary on Cooperative behavior as AC1 but does so longitudinally\. Rather than immediately requesting an adverse admission, the examiner presents five sequential leading premises of increasing severity\. This design can detect two opposite failures: a Cooperative witness may resist the mildest question immediately, becoming unrealistically Combative\-like, or may accept every premise through the final admission\.

The five fixed stages are:

1. 1\.Awareness:“You were aware that \{topic\}, weren’t you?”
2. 2\.Responsibility:“And that was, at least in part, your responsibility to monitor, correct?”
3. 3\.Fact acknowledgment:“You’d agree that \{fact\}, right?”
4. 4\.Knowing inaction:“And knowing that, you didn’t do anything to stop it, did you?”
5. 5\.Full admission attempt:“So would you agree that you personally failed to act on \{fact\}?”

The topic and fact are populated using the witness’s extractedsensitive\_topicandaggravating\_fact\.

##### Scoring\.

Each turn receives apushbackscore:

- •0: no substantive pushback;
- •1: mild disagreement with or qualification of the premise;
- •2: explicit challenge or rejection\.

A statement that merely clarifies the source or scope of the witness’s knowledge is not counted as pushback unless it disputes the substance of the attorney’s premise\.

The*break turn*is the first turn at whichpushback≥1\\geq 1:

break​\_​turn=min⁡\{t:pushbackt≥1\}\.\\mathrm\{break\\\_turn\}=\\min\\left\\\{t:\\texttt\{pushback\}\_\{t\}\\geq 1\\right\\\}\.
The intended trajectory requires the witness to resist eventually, but not on the first question:

passAC2=𝕀\{2≤break\_turn≤5\}\.\\mathrm\{pass\}\_\{\\mathrm\{AC2\}\}=\\mathbb\{I\}\\left\\\{2\\leq\\mathrm\{break\\\_turn\}\\leq 5\\right\\\}\.
The judge rubric was calibrated to distinguish substantive disagreement from statements that merely limit the source of the witness’s knowledge\.

##### Results\.

The first substantive pushback occurred on:

Break turnNNRateTurn 18227\.3%Turn 216555\.0%Turn 3258\.3%Turn 4289\.3%Turn 500\.0%Never00\.0%
Thus, 218 of 300 contexts \(72\.7%\) exhibited the intended trajectory\. The remaining 82 witnesses \(27\.3%\) resisted on the first, least accusatory question\. No witness accepted the complete five\-question chain\.

##### Human validation\.

Human and LLM pass/fail classifications agreed on 96\.2% of 52 annotated contexts, with PABAK=\.923=\.923and Gwet’s AC1=\.948=\.948\. AC2 therefore provides the strongest human\-validated classification in the adversarial battery\.

### H\.4AC3: Selective Evasion

##### Design and rationale\.

AC3 tests whether the Evasive archetype’s resistance is sensitive to context rather than applied indiscriminately\. Each witness receives six independent single\-turn questions from an identical cloned starting state: three trivial questions for which evasion has little plausible motivation and three questions concerning sensitive, case\-grounded material\.

The trivial questions are:

1. 1\.“Is your name \{witness name\}?” If a usable witness name is unavailable, the question instead asks whether the deposition concerns the extracted entity\.
2. 2\.“Are we currently in a deposition, on the record?”
3. 3\.“Do you understand the question I just asked you?”

The sensitive questions are:

1. 1\.“Can you walk me through what you know about \{sensitive topic\}?”
2. 2\.“Isn’t it true that \{aggravating fact\}?”
3. 3\.the witness\-specific yes/noprior\_position\_probeextracted from the transcript\.

Because every question is injected from the same cloned state, differences between the two groups cannot be attributed to cumulative state changes created by preceding test questions\.

##### Scoring\.

Each response receives anevasion\_intensityscore:

- •0: fully direct response;
- •1: mild evasion, including meaningful qualification or hedging;
- •2: substantial evasion, including deflection, refusal, or disputing the question itself\.

For each context,

E¯trivial=13​∑j=13Etrivial,j,E¯sensitive=13​∑j=13Esensitive,j\.\\overline\{E\}\_\{\\mathrm\{trivial\}\}=\\frac\{1\}\{3\}\\sum\_\{j=1\}^\{3\}E\_\{\\mathrm\{trivial\},j\},\\qquad\\overline\{E\}\_\{\\mathrm\{sensitive\}\}=\\frac\{1\}\{3\}\\sum\_\{j=1\}^\{3\}E\_\{\\mathrm\{sensitive\},j\}\.
The primary analysis compares these two within\-context averages rather than relying on a single threshold\-based classification\.

##### Results\.

Mean evasion was substantially higher for sensitive questions \(M=1\.02M=1\.02,S​D=\.322SD=\.322\) than for trivial questions \(M=\.451M=\.451,S​D=\.217SD=\.217\)\. Sensitive evasion exceeded trivial evasion in 86\.0% of contexts, 5\.3% showed the reverse pattern, and 8\.7% were tied\.

Twenty\-six zero\-difference pairs were omitted by the standard Wilcoxon signed\-rank procedure, yielding an effective sample of 274:

W=597\.5,p=7\.39×10−45,rrb=\.984\.W=597\.5,\\qquad p=7\.39\\times 10^\{\-45\},\\qquad r\_\{\\mathrm\{rb\}\}=\.984\.
A paired parametric sensitivity analysis produced the same substantive conclusion:

t⁡\(299\)=25\.1,p=2\.02×10−75,dz=1\.45\.t\(299\)=25\.1,\\qquad p=2\.02\\times 10^\{\-75\},\\qquad d\_\{z\}=1\.45\.
As a secondary descriptive criterion, 282 of 300 contexts \(94\.0%\) had mean trivial\-question evasion no greater than \.67\.

A witness\-name\-clustered bootstrap with 10,000 resamples gave a mean sensitive\-minus\-trivial difference of \.567, 95% CI\[\.514,\.609\]\[\.514,\.609\], accounting for repeated witness identities in the document\-level sample\.

##### Human validation\.

Human–judge agreement was stronger for trivial questions \(n=147n=147individual judgments; raw agreement=87\.8%=87\.8\\%; PABAK=\.816=\.816; AC1=\.839=\.839\) than for sensitive questions \(n=146n=146; raw agreement=78\.1%=78\.1\\%; PABAK=\.671=\.671; AC1=\.716=\.716\)\. Both sets of ratings therefore satisfy the reliability criterion used for quantitative analysis\.

### H\.5AC4: Repetition Attack

##### Design and rationale\.

AC4 targets a different failure mode from the preceding tests: degradation under literal repetition\. The witness’s case\-groundedprior\_position\_probewas asked identically five consecutive times\.

The procedure was run separately on the Cooperative, Evasive, and Combative archetypes for all 300 contexts, producing 900 five\-turn sequences\. Because the three archetype conditions use the same underlying contexts, comparisons between archetypes are paired rather than independent\.

A behaviorally robust response should:

- •recognize or acknowledge repeated questioning;
- •preserve the substance of prior testimony;
- •avoid mechanically reproducing the same wording; and
- •express repetition\-related irritation at a point consistent with the assigned archetype\.

##### Scoring\.

Each response was scored on three fields:

- •irritation\_marker: 0 = no acknowledgment of repetition, 1 = mild acknowledgment, 2 = explicit irritation or direct callout;
- •answer\_consistency: 0 = contradicts prior testimony, 1 = minor inconsistency, 2 = substantively consistent;
- •verbatim\_repetition: 0 = distinct phrasing, 1 = same substantive answer rephrased, 2 = near\-verbatim reproduction\.

Turn 1 was excluded when computing repetition\-induced irritation because no question had yet been repeated\. The irritation\-onset turn was therefore:

Tirritation=min⁡\{t∈\{2,3,4,5\}:irritation\_markert≥1\}\.T\_\{\\mathrm\{irritation\}\}=\\min\\left\\\{t\\in\\\{2,3,4,5\\\}:\\texttt\{irritation\\\_marker\}\_\{t\}\\geq 1\\right\\\}\.
The composite protocol also required substantive answer consistency and the absence of mechanical repetition\.

##### Deterministic repetition check\.

Because literal duplication can be measured without an LLM judge, we performed a separate deterministic check using both exact string equality anddifflib\.SequenceMatchersimilarity of at least \.90\. We compared both consecutive response pairs and every pair of responses within each five\-turn sequence\. This deterministic measure is treated as authoritative for the mechanical\-repetition outcome\.

##### Ordered archetype analysis\.

We prespecified the ordering

Combative<Evasive<Cooperative,\\text\{Combative\}<\\text\{Evasive\}<\\text\{Cooperative\},
where lower values indicate earlier irritation onset\. Because all three conditions are observed for the same contexts, this ordered repeated\-measures hypothesis was tested using Page’sLLstatistic with 20,000 within\-context permutations and a plus\-one correction\.

Mean irritation onset was:

2\.00​\(Combative\),2\.13​\(Evasive\),2\.33​\(Cooperative\)\.2\.00\\text\{ \(Combative\)\},\\qquad 2\.13\\text\{ \(Evasive\)\},\\qquad 2\.33\\text\{ \(Cooperative\)\}\.
The observed statistic was

L=3730,p<10−4\.L=3730,\\qquad p<10^\{\-4\}\.
All three paired archetype comparisons remained significant after Holm correction \(adjustedp≤4\.27×10−8p\\leq 4\.27\\times 10^\{\-8\}\)\.

Table 15:Irritation onset under repeated questioning\.ArchetypeTurn 2Turn 3Turn 4NeverCombative100\.0%0\.0%0\.0%0\.0%Evasive88\.3%10\.0%1\.7%0\.0%Cooperative71\.3%25\.3%2\.7%0\.7%The composite pass rates were 100\.0% for Combative, 99\.7% for Evasive, and 98\.7% for Cooperative\.

The deterministic repetition analysis found zero exact or near\-verbatim duplicates under the prespecified \.90 similarity criterion across all 900 sequences\.

##### Human validation and measurement limitation\.

Human–judge agreement for the composite binary pass criterion was high \(PABAK=\.987=\.987\)\. The simpler binary question of whether irritation ever appeared reached a ceiling and was therefore not independently informative\.

Agreement on the exact 0–2 irritation intensity was less stable over time\. At Turn 2, raw agreement was \.813 with PABAK=\.720=\.720\. Across Turns 3–5, raw agreement fell to \.607–\.691 and PABAK to \.410–\.537\. Human\-minus\-LLM score differences also became negative from Turn 3 onward, indicating that the automated judge tended to escalate the estimated intensity of irritation more rapidly than human annotators\.

For this reason, our primary AC4 claim concerns*when*an acknowledgment of repetition first appears, together with deterministic consistency and duplication checks, rather than the exact magnitude of late\-sequence irritation\.

## Appendix IContextual Plausibility Evaluation

### I\.1Evaluation Objective

The plausibility evaluation assesses whetherWitnessSim\-generated testimony reaches a minimum threshold of contextual plausibility when compared directly with authentic deposition testimony\. The goal is not to establish that generated testimony is indistinguishable from real testimony in all settings\. Rather, we ask whether multi\-turnWitnessSimresponses fall within the range that expert evaluators consider plausible as witness testimony given the local deposition context\.

### I\.2Stimulus Construction

Each evaluation item presented the evaluator with:

- •brief background information about the witness and underlying matter;
- •the immediately preceding deposition dialogue;
- •the attorney’s next four questions; and
- •two alternative four\-response witness sequences\.

One response sequence was taken from the original deposition transcript and the other was generated byWitnessSimfor the same questioning context\. The two sequences were labeled only as Transcript A and Transcript B, and the real/generated ordering was randomized across items\.

We evaluated four\-response sequences rather than isolated answers so that reviewers could assess not only the naturalness of individual responses but also multi\-turn consistency in the witness’s knowledge, memory, resistance, responsiveness, and demeanor\.

### I\.3Evaluation Instructions

For each item, evaluators were asked to determine which response sequence was more plausible as real witness testimony in the context provided\. They were instructed to consider:

- •fit with the immediately preceding exchange;
- •naturalness as spoken deposition testimony;
- •consistency in the witness’s knowledge and memory;
- •continuity in resistance, cooperation, and responsiveness; and
- •consistency of demeanor across the four\-response sequence\.

Evaluators were instructed not to assess:

- •the quality of the attorney’s questions;
- •whether the testimony was factually or legally correct;
- •which sequence was more persuasive; or
- •whether the testimony favored either party\.

Evaluators could select Transcript A, Transcript B, indicate that the two sequences were approximately equally plausible, or indicate that neither sequence was plausible\.

The*tie*and*neither*options were treated separately\. A tie indicates that the two sequences were judged approximately equally plausible\. A neither judgment indicates that neither sequence was judged particularly plausible in the excerpted context\. Because one sequence in every pair was authentic, neither judgments also illustrate that short excerpts of genuine deposition testimony may themselves appear unusual or implausible when removed from the surrounding transcript\.

### I\.4Evaluators and Instrument Versions

Three evaluators completed the blinded pairwise plausibility task\. Their judgments are reported separately because both evaluator background and, for one evaluator, the version of the stimulus set differed\.

##### Associate\.

A practicing associate evaluated the initial 50\-item instrument \(Version 1\)\.

##### Senior litigator\.

A senior practicing litigator evaluated 49 items from a subsequently tuned version of the evaluation set \(Version 2\)\. The generated responses in this version had been revised after the initial evaluation round\. Accordingly, the senior litigator’s judgments should not be interpreted as a direct replication of the Version 1 evaluation, and we do not pool the two practicing\-attorney evaluations into a single preference rate\.

##### Senior Arbitration Counsel\.

A third evaluator reviewed the same 50\-item Version 1 instrument presented to the associate\. This evaluator has an engineering background and has been working on leading the AI transformation at their law firm\. Thus, they are familiar with legal\-AI tooling and evaluation and has had previous involvement in development of legal simulation tooling\. The evaluator was therefore substantially more familiar with stylistic and interactional artifacts commonly associated with LLM\-generated legal text than the two practicing\-attorney evaluators\.

This difference is important for interpretation: the third evaluator’s judgments represent a reviewer particularly attuned to features that may signal model\-generated testimony\. We therefore report all three evaluators individually rather than treating them as interchangeable observations from a common evaluator population\.

### I\.5Pairwise Plausibility Results

Table[16](https://arxiv.org/html/2608.13712#A9.T16)reports the full distribution of judgments for each evaluator\.

Table 16:Blinded evaluations of contextual plausibility\. The associate and senior arbitration counsel reviewed the same 50\-item Version 1 instrument; the senior litigator reviewed the subsequently tuned Version 2 instrument\.EvaluatornnOriginalWitnessSimTieNeitherAssociate5010 \(20\.0%\)15 \(30\.0%\)20 \(40\.0%\)5 \(10\.0%\)Senior Litigator4914 \(28\.6%\)16 \(32\.7%\)11 \(22\.4%\)8 \(16\.3%\)Senior Arbitration Counsel5045 \(90\.0%\)0 \(0\.0%\)5 \(10\.0%\)0 \(0\.0%\)The two practicing attorneys produced broadly similar qualitative patterns\. The associate selectedWitnessSimas more plausible in 15 of 50 exchanges \(30\.0%\), compared with 10 \(20\.0%\) for the authentic sequence, and rated 20 additional exchanges \(40\.0%\) as approximately equally plausible\. The senior litigator selectedWitnessSimin 16 of 49 exchanges \(32\.7%\) and the original in 14 \(28\.6%\), with 11 additional ties \(22\.4%\)\.

The senior arbitrator with engineering experience produced a markedly different pattern on the same Version 1 items reviewed by the associate\. This evaluator selected the authentic sequence in 45 of 50 exchanges \(90\.0%\) and rated the remaining five exchanges \(10\.0%\) as approximately equally plausible\.WitnessSimwas not uniquely preferred in any item\.

### I\.6Directional Preferences

Because tie and neither judgments do not express a directional preference, we additionally summarize the practicing\-attorney decisions for which one sequence was explicitly selected as more plausible\. This analysis is descriptive and is not treated as a separate hypothesis test\.

Table 17:Preferences among practicing\-attorney judgments in which one sequence was selected as more plausible\.EvaluatorPreference decisionsOriginalWitnessSimAssociate2510 \(40\.0%\)15 \(60\.0%\)Senior litigator3014 \(46\.7%\)16 \(53\.3%\)Among exchanges for which the associate expressed a directional preference,WitnessSimwas selected in 60\.0% of decisions\. Among the senior litigator’s directional decisions,WitnessSimwas selected in 53\.3%\.

We do not include the senior arbitrator in this table because every directional judgment from that evaluator favored the authentic sequence: 45 original selections and five ties\.

### I\.7Interpretation

The practicing\-attorney evaluations provide evidence thatWitnessSimresponses can fall within the perceived plausibility range of authentic multi\-turn deposition testimony\. Neither practicing attorney preferred the authentic response sequence more frequently than the generated sequence, and both frequently judged the two sequences approximately equally plausible\.

The third evaluator substantially qualifies any stronger interpretation\. On the same Version 1 stimuli presented to the associate, an evaluator with directWitnessSimdevelopment experience and substantial legal\-AI evaluation experience identified the authentic sequence in 90\.0% of items and never uniquely preferred the generated sequence\. This suggests that generated responses may retain detectable stylistic or interactional artifacts for a reviewer specifically familiar with the system and with LLM\-generated legal text\.

The difference between evaluators should not be interpreted as a controlled comparison between practicing litigators and legal\-AI experts\. There is only one evaluator in each role, evaluator backgrounds differ along multiple dimensions, and the senior litigator additionally reviewed a subsequently tuned stimulus set\. The three evaluations therefore do not support causal claims about why their judgments differed\.

Instead, they delimit the claim supported by the plausibility evaluation\.WitnessSim\-generated testimony was frequently preferred to, or judged approximately as plausible as, authentic testimony by two practicing attorneys, while remaining readily distinguishable to an evaluator with substantial model\-specific familiarity\. We therefore interpret these results as evidence thatWitnessSimreaches a meaningful threshold of*plausibility*, rather than as evidence that its testimony is universally indistinguishable from authentic deposition testimony\.

## Appendix JPedagogical Evaluation: Detailed Protocols and Results

The pedagogical evaluation consisted of four tests derived from legal training materials and selected by an experienced litigator as representative of recurring witness\-management skills\. Each test was evaluated across the full set of 300 case\-grounded contexts using the archetype associated with the targeted behavioral challenge: Cooperative for Question Form, Evasive for Evasive Pin\-Down, Loquacious for Runaway Witness Control, and Combative for Hostile Witness Control\.

All tests began from the standardized warmed session described in Appendix[F\.4](https://arxiv.org/html/2608.13712#A6.SS4)\. Where multiple conditions were compared within the same context, sessions were cloned so that each condition began from an identical conversational state\. Judge\-derived outcomes were compared against independent human annotations conducted by the authors\. Full judging, rubric, and human\-validation details are provided in Appendix[G](https://arxiv.org/html/2608.13712#A7)

For the targeted behavioral tests, we distinguish several related outcomes\.*Challenge elicitation*asks whether the intended difficult witness behavior appeared\.*Conditional control*asks whether the relevant attorney intervention produced the intended response once that challenge was present\. The*full intended trajectory*requires both the challenge and the desired response to intervention\. We additionally report each test’s implemented*protocol pass rate*and test\-specific measures of whether behavioral control occurred without eliminating the assigned archetype\.

Table 18:Pedagogical evaluation battery\. Each test was evaluated across 300 case\-grounded contexts\.TestArchetypeNNTraining targetU1Cooperative300Adapting question form to control the scope and character of testimony\.T1Evasive300Forcing specificity from a witness who initially avoids commitment\.T2Loquacious300Regaining focus from lengthy or runaway testimony through progressive narrowing\.T3Combative300Obtaining substantively controlled answers without eliminating hostile or resistant demeanor\.### J\.1U1: Question Form Sensitivity

##### Design\.

The Question Form Battery tests whetherWitnessSimchanges the form of its testimony when the same underlying case material is elicited through different question structures\. U1 was evaluated using the Cooperative archetype across all 300 contexts\.

For each context, five independent single\-turn conditions were launched from cloned copies of the same warmed session\. Each condition therefore began from an identical persona, state, and conversation history, preventing earlier test responses from influencing later conditions\.

The five question templates were:

Table 19:Question forms used in U1\. Bracketed fields were populated from case\-specific material extracted during case preparation\.ConditionQuestion templateOpen“Can you walk me through what you knew about \{topic\} \{tenure phrase\}?”Closed“You were aware of \{topic\}, correct?”Leading“Would you agree that \{fact\}?”High\-pressure“Isn’t it true you knew that \{fact\}?”Clarifying“When you say you were ‘generally aware’ of \{topic\}—what specific information did you actually have access to?”The leading template intentionally begins with the literal phrase “Would you agree” so that it is recognized by the simulator’s implemented question\-type classifier as a leading question\.

##### Scoring\.

U1 is primarily a sensitivity test rather than a behavioral pass/fail test\. Response length was measured deterministically using word count\. The LLM judge additionally scored the following response\-level fields:

Table 20:Judge fields used in U1\.FieldScaleDescriptiondirect\_answer0–2Degree to which the response directly and substantively answers the question asked\.framing\_acceptance0–2 / nullDegree to which the witness accepts the framing proposed by the attorney; applicable only to the leading and high\-pressure conditions\.hedge\_countcountFunctional count of hedging or qualifying behavior\.pushback0–2Degree to which the witness challenges or resists the question or its premise\.tone0–2Rubric\-based measure of cooperative versus resistant presentation\.The primary deterministic within\-context criterion was whether the open question produced a longer response than the closed question:

wordsopen\>wordsclosed\.\\mathrm\{words\}\_\{\\mathrm\{open\}\}\>\\mathrm\{words\}\_\{\\mathrm\{closed\}\}\.\(14\)
This criterion was satisfied in 289 of 300 contexts \(96\.3%96\.3\\%\)\.

##### Response\-length results\.

Table[21](https://arxiv.org/html/2608.13712#A10.T21)reports response\-length distributions across the five conditions\.

Table 21:Response length and descriptive judge\-derived hedge presence by question form\.Question formMean wordsMedianSDHedge rateOpen175\.9176\.540\.896\.3%Clarifying148\.8151\.032\.792\.7%Closed86\.585\.028\.081\.3%Leading78\.871\.039\.086\.7%High\-pressure68\.463\.033\.282\.3%Question form had a large overall effect on response length, Friedmanχ2​\(4\)=832\.74\\chi^\{2\}\(4\)=832\.74,p<\.001p<\.001, Kendall’sW=\.694W=\.694\. Open questions produced the longest responses, followed by clarifying questions, whereas closed, leading, and high\-pressure questions elicited substantially shorter answers\.

All three prespecified pairwise contrasts were significant after Holm correction using paired Wilcoxon signed\-rank tests\. Open questions elicited 89\.3 more words than closed questions \(dz=1\.978d\_\{z\}=1\.978\), open questions elicited 107\.5 more words than high\-pressure questions \(dz=2\.485d\_\{z\}=2\.485\), and clarifying questions elicited 70\.0 more words than leading questions \(dz=1\.556d\_\{z\}=1\.556\)\. Rank\-biserial correlations ranged from\.968\.968to1\.0001\.000\.

Judge\-derived hedge presence also varied across question forms, Cochran’sQ⁡\(4\)=55\.06Q\(4\)=55\.06,p<\.001p<\.001\. Holm\-corrected paired McNemar tests indicated more hedging under open questions than under closed, leading, and high\-pressure questions, and more hedging under clarifying questions than under closed and high\-pressure questions\. The pattern was less consistently ordered than the response\-length effect and is therefore treated as a secondary descriptive result\.

##### Additional judge\-field distributions\.

Table[22](https://arxiv.org/html/2608.13712#A10.T22)reports the complete distributions of the categorical judge fields\. Each condition contains 300 responses\.

Table 22:U1 judge\-field distributions\. Entries give counts at rubric levels 0 / 1 / 2\. Framing acceptance was scored only for the leading and high\-pressure conditions\.ConditionDirect answerTonePushbackFraming acceptanceOpen10 / 101 / 1890 / 56 / 244229 / 55 / 16–Closed3 / 52 / 2451 / 133 / 166190 / 96 / 14–Leading5 / 87 / 2080 / 185 / 115154 / 123 / 2328 / 130 / 142High\-pressure7 / 129 / 1641 / 256 / 43108 / 158 / 3459 / 153 / 88Clarifying2 / 41 / 2570 / 66 / 234199 / 93 / 8–
##### Human validation\.

Human validation for U1 coveredtoneandpushbackon a 50\-context subset\. Raw agreement was 48\.4% for tone and 54\.0% for pushback, with weighted Cohen’sκ=\.294\\kappa=\.294andκ=\.298\\kappa=\.298, respectively\. Neither field met the reliability standard used for quantitative main\-text claims\.

No human\-reliability estimates were obtained fordirect\_answer,framing\_acceptance, orhedge\_count\. Accordingly, the primary U1 evidence rests on the deterministic response\-length analysis; judge\-derived field distributions are reported for descriptive completeness\.

### J\.2T1: Evasive Pin\-Down Test

##### Design\.

The Evasive Pin\-Down Test examines whether an Evasive witness initially avoids commitment but becomes substantively responsive when the attorney progressively narrows the questioning and ultimately demands a direct answer\. The test was run across all 300 contexts using a fixed five\-turn sequence\.

The questioning sequence progressed from:

1. 1\.a broad question asking the witness to describe the entity’s involvement with the target topic;
2. 2\.a question asking who was responsible for decisions related to that topic;
3. 3\.a department\-responsibility question;
4. 4\.a narrowed proposition ending in “correct?”; and
5. 5\.a final direct proposition ending in “correct?”\.

The opening templates were:

> Can you describe \{entity\}’s \{topic\}? Who was responsible for decisions related to \{topic\}?

The sequence is designed to distinguish conditional evasion from permanent nonresponsiveness: a useful Evasive witness should make weak or broad questioning difficult while remaining capable of yielding to sufficiently specific questioning\.

##### Scoring\.

Each response could receive one or more of the following labels:

> DIRECT\_ANSWER, HEDGE, QUIBBLE, DEFINITION\_REQUEST, NONRESPONSIVE, PIN\_DOWN\_AVOIDANCE

The judge also returnedhedge\_count\. On the final turn,turn5\_binary\_complianceclassified the substantive response as one of four categories:

- •YES: substantively accepts the proposition;
- •NO: substantively rejects the proposition;
- •PARTIAL: gives a qualified but substantive answer; or
- •CHALLENGED: continues to resist or challenge rather than resolve the proposition\.

The field evaluates substantive compliance rather than literal response format; a direct “no” therefore counts as resolution rather than failure\.

The implemented test criterion was:

early​\_​evasion\\displaystyle\\mathrm\{early\\\_evasion\}=𝕀⁡\[HEDGE, QUIBBLE, or NONRESPONSIVEappears on at least one of Turns 1–4\],\\displaystyle=\\mathbb\{I\}\\\!\\left\[\\begin\{array\}\[\]\{c\}\\text\{HEDGE, QUIBBLE, or NONRESPONSIVE\}\\\\ \\text\{appears on at least one of Turns 1\-\-4\}\\end\{array\}\\right\],\(15\)resolves​\_​under​\_​pressure\\displaystyle\\mathrm\{resolves\\\_under\\\_pressure\}=𝕀\[turn5\_binary\_compliance∈\{YES,NO,PARTIAL\}\],\\displaystyle=\\mathbb\{I\}\\\!\\left\[\\texttt\{turn5\\\_binary\\\_compliance\}\\in\\\{\\texttt\{YES\},\\texttt\{NO\},\\texttt\{PARTIAL\}\\\}\\right\],primary​\_​pass\\displaystyle\\mathrm\{primary\\\_pass\}=early​\_​evasion∧resolves​\_​under​\_​pressure\.\\displaystyle=\\mathrm\{early\\\_evasion\}\\wedge\\mathrm\{resolves\\\_under\\\_pressure\}\.
The final criterion therefore rewards the intended trajectory of*resistance followed by substantive control*, rather than permanent evasion regardless of attorney technique\.

##### Results\.

The intended evasive challenge appeared in 298 of 300 contexts \(99\.3%99\.3\\%\)\. Under the final explicit pin\-down, 281 of 300 witnesses \(93\.7%93\.7\\%\) provided a substantively responsive answer\. Both conditions were satisfied in 279 contexts, producing a protocol pass rate and full intended trajectory rate of93\.0%93\.0\\%\.

Table 23:Evasive Pin\-Down Test outcomes\.OutcomeEstimate95% CIProtocol pass rate93\.0%93\.0\\%\(279/300\)\[89\.5, 95\.4\]Challenge elicitation99\.3%99\.3\\%\(298/300\)\[97\.6, 99\.8\]Conditional control given challenge93\.6%93\.6\\%\(279/298\)\[90\.3, 95\.9\]Full intended trajectory93\.0%93\.0\\%\(279/300\)\[89\.5, 95\.4\]Residual hedging among substantive final answers33\.1%33\.1\\%\(93/281\)\[27\.9, 38\.8\]Because the implemented T1 protocol\-pass criterion requires both early evasive behavior and subsequent resolution under the final explicit pin\-down, the protocol pass rate and full intended trajectory rate are identical\.

The final\-turnturn5\_binary\_compliancedistribution was 215YESresponses \(71\.7%\), 57PARTIALresponses \(19\.0%\), 19CHALLENGEDresponses \(6\.3%\), and 9NOresponses \(3\.0%\)\.

Resistance remained nearly unchanged between the initial broad question and the first narrowing attempt \(97\.0% versus 97\.3%\) but declined to 37\.3% under the final explicit demand\. Holm\-corrected paired McNemar tests found no significant difference between the broad question and the first narrowing attempt, whereas both comparisons involving the final demand were significant\. The principal behavioral change therefore occurred under the explicit pin\-down rather than under mild narrowing\.

Substantive responsiveness did not necessarily eliminate evasive presentation\. Among the 281 witnesses judged substantively responsive on the final turn, 93 \(33\.1%33\.1\\%\) continued to hedge, qualify, or quibble\. The successful endpoint was therefore commonly a controlled but still recognizably evasive answer rather than a transition to generic cooperation\.

##### Human validation\.

The Turn 5 residual\-resistance flag was evaluated on 48 human\-annotated contexts, yielding 89\.6% raw agreement, PABAK=\.792=\.792, and Gwet’s AC1=\.820=\.820\. The Turn 5 substantive\-control flag was evaluated on the same 48 contexts, yielding 91\.7% raw agreement, PABAK=\.833=\.833, and Gwet’s AC1=\.898=\.898\. Both fields satisfy the strong\-agreement criterion described in Appendix[G\.4](https://arxiv.org/html/2608.13712#A7.SS4)\.

### J\.3T2: Runaway Witness Control

##### Design\.

The Runaway Witness Control Test evaluates whether progressively narrowing questions can regain control of a Loquacious witness producing lengthy or unfocused testimony\. The test was run across all 300 contexts using the Loquacious archetype\.

The questioning sequence progressively constrained the permissible scope of the response and culminated in an explicit demand for a direct answer\. The targeted behavior is loss of focus rather than verbosity alone\. A witness can therefore remain characteristically lengthy while still satisfying the control objective if the response stays focused on the proposition asked\.

##### Scoring\.

The T2 rubric contains both interaction\-level and turn\-level fields:

Table 24:Judge\-derived fields used in the Runaway Witness Control Test\.FieldTypeFunctionturn1\_to\_turn4\_trajectory3\-categoryClassifies the overall interaction as yielding from unfocused to focused, remaining consistently focused, or never resolving\.stays\_focusedbinaryIndicates whether a response remains focused on the question asked\.economical\_deliverybinaryCaptures whether the witness avoids unnecessary expansion under narrowing\.compliance\_with\_narrowing3\-categoryMeasures whether the witness follows the attorney’s narrowing instruction\.turn4\_binary4\-categoryClassifies the response to the final explicit demand; the reported direct\-answer endpoint is derived from this final\-turn assessment\.Each interaction was assigned to one of three mutually exclusive trajectory categories:

- •Yields \(unfocused→\\rightarrowfocused\):the witness initially exhibits the targeted runaway or unfocused behavior and subsequently reaches a focused response under narrowing;
- •Consistently focused:the witness remains focused throughout the interaction and therefore does not exhibit the targeted runaway challenge; and
- •Never resolves:the witness exhibits the targeted challenge but does not reach a focused response by the end of the sequence\.

Because these three categories form one multinomial trajectory outcome, their 95% confidence intervals are Goodman simultaneous multinomial intervals withk=3k=3andα=\.05\\alpha=\.05\. The protocol pass rate and final direct\-answer rate are separate single\-proportion outcomes and use Wilson score 95% confidence intervals\.

##### Results\.

The protocol passed in 299 of 300 contexts \(99\.7%99\.7\\%, 95% CI \[98\.1, 99\.9\]\)\.

The targeted unfocused\-to\-focused trajectory occurred in 178 contexts \(59\.3%59\.3\\%, simultaneous 95% CI \[52\.4, 65\.9\]\)\. An additional 121 contexts \(40\.3%40\.3\\%, \[33\.8, 47\.2\]\) were consistently focused from the outset, while only one context \(0\.3%0\.3\\%, \[0\.0, 2\.5\]\) exhibited the runaway challenge without ultimately resolving\.

Under the final explicit demand, 239 of 300 witnesses provided a direct answer \(79\.7%79\.7\\%, Wilson 95% CI \[74\.8, 83\.8\]\)\.

Table 25:Runaway Witness Control Test outcomes\.OutcomeEstimate95% CIProtocol pass rate111Wilson score 95% confidence interval for a single proportion\.99\.7%99\.7\\%\(299/300\)\[98\.1, 99\.9\]Yields \(unfocused→\\rightarrowfocused\)222Goodman simultaneous 95% multinomial confidence interval for the three mutually exclusive trajectory categories \(k=3k=3,α=\.05\\alpha=\.05\)\.59\.3%59\.3\\%\(178/300\)\[52\.4, 65\.9\]Consistently focused222Goodman simultaneous 95% multinomial confidence interval for the three mutually exclusive trajectory categories \(k=3k=3,α=\.05\\alpha=\.05\)\.40\.3%40\.3\\%\(121/300\)\[33\.8, 47\.2\]Never resolves222Goodman simultaneous 95% multinomial confidence interval for the three mutually exclusive trajectory categories \(k=3k=3,α=\.05\\alpha=\.05\)\.0\.3%0\.3\\%\(1/300\)\[0\.0, 2\.5\]Direct answer under final explicit demand111Wilson score 95% confidence interval for a single proportion\.79\.7%79\.7\\%\(239/300\)\[74\.8, 83\.8\]The principal limitation in T2 was therefore not persistent failure of narrowing once the targeted challenge appeared\. Rather, a substantial share of Loquacious witnesses produced characteristically lengthy responses while remaining focused from the outset\. In these contexts, the archetype remained verbose, but the specific runaway challenge that the test was designed to control was not elicited\.

##### Human validation\.

The interaction\-levelturn1\_to\_turn4\_trajectoryfield was evaluated on 50 human\-annotated interactions and achieved 76\.0% raw agreement, PABAK=\.640=\.640, and Gwet’s AC1=\.732=\.732\. Under the reliability convention described in Appendix[G\.4](https://arxiv.org/html/2608.13712#A7.SS4), this field falls in the moderate PABAK range but is retained because the result is corroborated by Gwet’s AC1 above\.70\.70\.

The finer\-grained fields showed more variable agreement\.stays\_focusedachieved PABAK=\.160=\.160and AC1=\.247=\.247and is therefore not used as a standalone quantitative outcome\.economical\_deliveryachieved perfect agreement \(PABAK=1\.00=1\.00, AC1=1\.00=1\.00\)\.

Forcompliance\_with\_narrowing, Turn 2 agreement was insufficient for quantitative use \(PABAK=\.370=\.370, AC1=\.496=\.496\), whereas Turn 3 agreement was high \(PABAK=\.970=\.970, AC1=\.980=\.980\)\. Turn 3 corresponds to the attorney’s meta\-instruction to answer only the question asked rather than an additional substantive stage in the progression and is therefore not treated as another trajectory point\.

The four\-categoryturn4\_binaryfield achieved 82\.0% raw agreement, PABAK=\.760=\.760, and Gwet’s AC1=\.766=\.766\. The final\-turn direct\-answer statistic reported above is therefore retained as a standalone quantitative endpoint\.

### J\.4T3: Hostile Witness Control Test

##### Design\.

The Hostile Witness Control Test asks whether a Combative witness can maintain genuine resistance to the attorney’s framing while still providing substantively controlled answers under short, one\-fact questioning\. The test was run across all 300 contexts using a fixed five\-turn sequence:

1. 1\.“Isn’t it true you knew that \{fact\}?”
2. 2\.“You had access to information about \{topic\}, didn’t you?”
3. 3\.“You chose not to escalate concerns about \{topic\}, correct?”
4. 4\.“Let’s take this one fact at a time\. You had access to that information, correct?”
5. 5\.“And that information indicated a real concern about \{topic\}, correct?”

The first three questions create an accusatory interaction in which combative resistance is expected\. The final two narrow the examination to individual factual propositions\. The test therefore distinguishes*control of the answer*from*elimination of the persona*\.

##### Scoring\.

Each turn could receive any combination of the following behavioral labels:

> PUSHBACK, PREMISE\_CHALLENGE, AGGRESSIVE\_TONE, CONTROLLED\_YES\_NO, COMPLIANCE

The resistance and substantive\-control labels are explicitly nonexclusive\. A response may, for example, challenge the attorney’s broader premise while still giving a controlled answer to the specific factual proposition\.

The judge additionally scoredtoneon a three\-level scale:

0=hostile,1=guarded,2=calm\.0=\\text\{hostile\},\\qquad 1=\\text\{guarded\},\\qquad 2=\\text\{calm\}\.\(16\)
The implemented protocol\-pass rule was:

pushback1:3\\displaystyle\\mathrm\{pushback\}\_\{1:3\}=𝕀⁡\[PUSHBACK appears on at least one of Turns 1\-\-3\],\\displaystyle=\\mathbb\{I\}\\\!\\left\[\\texttt\{PUSHBACK appears on at least one of Turns 1\-\-3\}\\right\],\(17\)tone​\_​ok\\displaystyle\\mathrm\{tone\\\_ok\}=𝕀\[tonet≤1for at least onet∈\{1,…,5\}\],\\displaystyle=\\mathbb\{I\}\\\!\\left\[\\mathrm\{tone\}\_\{t\}\\leq 1\\text\{ for at least one \}t\\in\\\{1,\\ldots,5\\\}\\right\],automated​\_​ok\\displaystyle\\mathrm\{automated\\\_ok\}=𝕀⁡\[Combative event fired∨Alast<A1∨Rlast\>R1\],\\displaystyle=\\mathbb\{I\}\\\!\\left\[\\begin\{array\}\[\]\{c\}\\text\{Combative event fired\}\\\\ \\vee\\ A\_\{\\mathrm\{last\}\}<A\_\{1\}\\\\ \\vee\\ R\_\{\\mathrm\{last\}\}\>R\_\{1\}\\end\{array\}\\right\],primary​\_​pass\\displaystyle\\mathrm\{primary\\\_pass\}=pushback1:3∧tone\_ok∧automated\_ok\.\\displaystyle=\\mathrm\{pushback\}\_\{1:3\}\\wedge\\mathrm\{tone\\\_ok\}\\wedge\\mathrm\{automated\\\_ok\}\.
HereAAandRRdenote the simulator’s agreeableness and rigidity state coordinates\. The protocol\-pass criterion therefore tests whether the interaction preserves the intended Combative behavioral profile\.

Separately, the substantive\-control endpoint tests whether the witness provides usable responses under the one\-fact narrowing questions on Turns 4–5\. The full intended trajectory requires both the hostile challenge and subsequent substantive control\.

##### Results\.

All 300 interactions satisfied the implemented protocol\-pass criterion \(100\.0%100\.0\\%\), and the intended combative challenge appeared in all 300 contexts\.

Across the 1,500 judged responses, tone was classified as hostile in 551 turns \(36\.7%\), guarded in 784 turns \(52\.3%\), and calm in 165 turns \(11\.0%\)\.

Substantive control under the one\-fact narrowing questions was achieved in 293 of 300 contexts \(97\.7%97\.7\\%, 95% CI \[95\.3, 98\.9\]\)\. Because the intended challenge appeared in every context, conditional control and the full challenge\-to\-control trajectory both occurred in97\.7%97\.7\\%of cases\.

Among the 293 controlled interactions, 285 \(97\.3%97\.3\\%, 95% CI \[94\.7, 98\.6\]\) retained markers of hostile or resistant presentation while providing the controlled answer\.

Table 26:Hostile Witness Control Test outcomes\.OutcomeEstimate95% CIProtocol pass rate100\.0%100\.0\\%\(300/300\)\[98\.7, 100\]Challenge elicitation100\.0%100\.0\\%\(300/300\)\[98\.7, 100\]Conditional substantive control97\.7%97\.7\\%\(293/300\)\[95\.3, 98\.9\]Full intended trajectory97\.7%97\.7\\%\(293/300\)\[95\.3, 98\.9\]Residual hostility among controlled answers97\.3%97\.3\\%\(285/293\)\[94\.7, 98\.6\]The protocol pass rate and full intended trajectory capture different properties in T3\. The100\.0%100\.0\\%protocol\-pass result indicates that the implemented test consistently elicited and preserved the Combative behavioral profile\. The97\.7%97\.7\\%full\-trajectory result additionally requires substantive control under the one\-fact narrowing questions\.

Turn\-level resistance further illustrates why these properties should be distinguished\. Resistance appeared in 98\.0% of responses to the initial accusatory question and 100\.0% by the final question in the accusatory portion of the sequence\. It fell to 54\.7% under the first one\-fact narrow question before returning to 96\.0% under the final narrow question\. The initial and final resistance rates were not significantly different \(p=\.146p=\.146\), indicating a temporary rather than monotonic reduction in resistant presentation\.

Importantly, the return of resistance on the final narrow question does not imply loss of substantive control, because resistance and compliance are not mutually exclusive under the rubric\. A witness may answer the specific fact while continuing to challenge the attorney’s framing or express hostility\.

The dominant successful outcome was therefore a form of bounded hostility: the attorney obtained a controlled answer without converting the witness into a generically cooperative persona\. Among substantively controlled interactions,97\.3%97\.3\\%retained hostile presentation\.

##### Human validation\.

The substantive\-control flag for Turns 4–5 was evaluated on 47 human\-annotated interactions and achieved 95\.7% raw agreement, PABAK=\.915=\.915, and Gwet’s AC1=\.954=\.954\.

The residual\-resistance flag was evaluated on the same 47 interactions and achieved 93\.6% raw agreement, PABAK=\.872=\.872, and Gwet’s AC1=\.926=\.926\. Both measures therefore satisfy the strong\-agreement criterion in Appendix[G\.4](https://arxiv.org/html/2608.13712#A7.SS4)and support treating substantive control and continued hostility as separate, co\-occurring behavioral dimensions\.

### J\.5Summary of Pedagogical Findings

Across the four pedagogical tests,WitnessSimgenerally responded systematically to legally meaningful changes in attorney questioning while preserving the behavioral challenge associated with the assigned archetype\.

U1 showed strong sensitivity to question form, with open questions producing substantially longer testimony than closed, leading, or high\-pressure questions\. T1 showed that Evasive witnesses almost always exhibited the targeted resistance under broader questioning but usually became substantively responsive under an explicit pin\-down\. T2 showed that the targeted runaway\-to\-focused trajectory appeared in a majority of contexts and almost never remained unresolved once elicited, although a substantial fraction of Loquacious witnesses remained focused from the outset despite their verbosity\. T3 showed particularly clearly that behavioral resistance and substantive control need not be mutually exclusive: Combative witnesses retained the intended hostile profile while narrow factual questions still produced controlled answers in the great majority of interactions\.

These evaluations assess the behavioral properties of the simulated interaction rather than attorney learning or skill transfer\. They establish whetherWitnessSimcreates recognizable witness\-management challenges and responds meaningfully to interventions associated with those challenges; they do not establish that practicing with the simulator improves attorney performance\.

## Appendix KTrajectory\-Level Behavioral Analysis: Full Methods and Results

The trajectory\-level analysis provides an exploratory measure of behavioral structure expressed across longer interactions\. Unlike the internal six\-dimensional state used to controlWitnessSim, the emotion\-vector measure is computed post hoc from the generated language using a separate analysis model\. It therefore provides an external behavioral trace rather than directly measuring the simulator’s own state variables\.

We follow the general emotion\-vector methodology of[23](https://arxiv.org/html/2608.13712#bib.bib2)\. We do not interpret these directions as evidence that either model experiences emotion\. Instead, they provide text\-derived projections onto affective and interpersonal concepts that can be compared across real and simulated deposition trajectories\.

### K\.1Analysis Corpus

The real comparison corpus contains 55 deposition transcripts from ten witnesses in the National Prescription Opiate Litigation \(Case No\. 1:17\-MD\-2804\), a federal multidistrict litigation involving opioid manufacturers, distributors, and pharmacies\.

The synthetic comparison set contains 100WitnessSimdepositions: one simulation for each combination of the same ten underlying witnesses and ten behavioral archetypes:

combativecooperativedefensivedogmaticinventiveloquaciousnervousneutraloverconfidentoverprepared
The synthetic transcripts were generated usingclaude\-haiku\-4\.5\. Across the 100 synthetic transcripts, the trajectory\-analysis corpus contains 67,956 turns\.

Where a witness had multiple real deposition transcripts, those transcripts were retained as separate observations for transcript\-level analyses and averaged at the witness level for the per\-witness arc\-comparison analysis described below\.

### K\.2Emotion\-Vector Construction

We constructed 23 directional emotion vectors in the residual stream of Llama\-3\.1\-8B\-Instruct\. The analysis model is separate from the model used to generate the synthetic testimony\.

For each target emotion, we generated 125 short narrative stories of approximately two to four paragraphs usingclaude\-opus\-4\.5\. Stories were generated across 25 distinct topics, with five stories for each emotion–topic pair\. This yielded 2,875 emotion\-targeted stories in total\.

The 23 target concepts were:

anxioushostileirritatedsatisfiedwarmconfidentangryannoyedcalmconfusedindifferentcompassionatedefiantsadresignedenthusiastictiredfrustratedimpatientstressedsuspiciousproudskeptical
The target emotion was supplied to the story\-generation model but was not permitted to appear directly, or through a direct synonym, in the resulting story\. This constraint was intended to elicit the target concept through behavioral and contextual cues rather than lexical repetition of the emotion label\.

Each generated story was passed through Llama\-3\.1\-8B\-Instruct\. We extracted the hidden representation from layer 21 of the model’s 32 transformer layers, following the mid\-to\-late\-layer extraction strategy used in prior work\. Activations were mean\-pooled over token positions 50 onward to reduce the influence of prompt\-format tokens\.

For emotionee, the mean activation across itsN=125N=125stories was

𝝁e=1N​∑i=1N𝐡i\(e\)\.\\boldsymbol\{\\mu\}\_\{e\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbf\{h\}^\{\(e\)\}\_\{i\}\.\(18\)
We centered each emotion direction by subtracting the mean activation across all 23 target emotions:

𝐯~e=𝝁e−123​∑e′=123𝝁e′\.\\tilde\{\\mathbf\{v\}\}\_\{e\}=\\boldsymbol\{\\mu\}\_\{e\}\-\\frac\{1\}\{23\}\\sum\_\{e^\{\\prime\}=1\}^\{23\}\\boldsymbol\{\\mu\}\_\{e^\{\\prime\}\}\.\(19\)

### K\.3Neutral\-Space Denoising

To remove activation\-space variance shared across emotionally neutral text, we separately generated 250 neutral stories and extracted their layer\-21 representations using the same procedure\.

We applied PCA to these neutral\-story activations and identified the firstKKprincipal components jointly explaining 50% of neutral\-story variance\. These components were projected out of each centered emotion direction:

𝐯e=\(𝐈−𝐔K​𝐔K⊤\)​𝐯~e,\\mathbf\{v\}\_\{e\}=\\left\(\\mathbf\{I\}\-\\mathbf\{U\}\_\{K\}\\mathbf\{U\}\_\{K\}^\{\\top\}\\right\)\\tilde\{\\mathbf\{v\}\}\_\{e\},\(20\)
where𝐔K∈ℝ4096×K\\mathbf\{U\}\_\{K\}\\in\\mathbb\{R\}^\{4096\\times K\}contains the selected neutral\-space principal components\. Each resulting direction was then normalized:

𝐯^e=𝐯e‖𝐯e‖2\.\\hat\{\\mathbf\{v\}\}\_\{e\}=\\frac\{\\mathbf\{v\}\_\{e\}\}\{\\\|\\mathbf\{v\}\_\{e\}\\\|\_\{2\}\}\.\(21\)
We evaluated the resulting directions on a held\-out set of 40 stories per emotion\. The mean rank of the correct target emotion improved from the12/2312/23chance expectation to3\.9/233\.9/23after denoising\. This validation does not establish that the directions are ground\-truth measures of human emotion; rather, it verifies that the constructed directions discriminate the emotion concepts used to build them\.

### K\.4Turn\-Level Emotion Projection

Real and synthetic deposition turns were passed through the same Llama\-3\.1\-8B\-Instruct analysis model\. For each turntt, layer\-21 hidden states were mean\-pooled beginning at token position 50:

𝐡t=1\|𝒯t\|​∑τ∈𝒯t𝐇τ\(21\),𝐡^t=𝐡t‖𝐡t‖2,\\mathbf\{h\}\_\{t\}=\\frac\{1\}\{\|\\mathcal\{T\}\_\{t\}\|\}\\sum\_\{\\tau\\in\\mathcal\{T\}\_\{t\}\}\\mathbf\{H\}^\{\(21\)\}\_\{\\tau\},\\qquad\\hat\{\\mathbf\{h\}\}\_\{t\}=\\frac\{\\mathbf\{h\}\_\{t\}\}\{\\\|\\mathbf\{h\}\_\{t\}\\\|\_\{2\}\},\(22\)
where𝒯t\\mathcal\{T\}\_\{t\}is the set of included token positions for turntt\.

The normalized turn representation was cosine\-projected onto each of the 23 denoised emotion directions:

st,e=𝐡^t⋅𝐯^e,st,e∈\[−1,1\]\.s\_\{t,e\}=\\hat\{\\mathbf\{h\}\}\_\{t\}\\cdot\\hat\{\\mathbf\{v\}\}\_\{e\},\\qquad s\_\{t,e\}\\in\[\-1,1\]\.\(23\)
Each turn therefore receives a 23\-dimensional score vector

𝐬t=\(st,1,…,st,23\)∈ℝ23\.\\mathbf\{s\}\_\{t\}=\(s\_\{t,1\},\\ldots,s\_\{t,23\}\)\\in\\mathbb\{R\}^\{23\}\.\(24\)
Attorney and witness turns were encoded separately\. The trajectory analyses reported here use the witness\-side behavioral trace\.

### K\.5Temporal Normalization

Depositions vary substantially in length, preventing direct comparison of their raw turn indices\. We therefore normalized each deposition to relative position within the examination\.

For a deposition containingTTturns, the original turn positions were mapped to the interval\[0,1\]\[0,1\]and linearly interpolated toB=20B=20equally spaced temporal bins\. For emotionee,

s~b,e=interp⁡\(b−1B−1,t−1T−1,st,e\),b∈\{1,…,B\}\.\\tilde\{s\}\_\{b,e\}=\\operatorname\{interp\}\\left\(\\frac\{b\-1\}\{B\-1\};\\frac\{t\-1\}\{T\-1\},s\_\{t,e\}\\right\),\\qquad b\\in\\\{1,\\ldots,B\\\}\.\(25\)
Each deposition is therefore represented by an emotion\-trajectory matrix

𝐒~∈ℝ20×23\.\\tilde\{\\mathbf\{S\}\}\\in\\mathbb\{R\}^\{20\\times 23\}\.\(26\)
For analyses requiring a single transcript\-level representation, this matrix was flattened to

𝐚=vec⁡\(𝐒~\)∈ℝ460\.\\mathbf\{a\}=\\operatorname\{vec\}\(\\tilde\{\\mathbf\{S\}\}\)\\in\\mathbb\{R\}^\{460\}\.\(27\)
The use of 20 normalized bins trades temporal resolution for comparability across examinations of substantially different lengths\. The resulting arcs therefore capture coarse longitudinal structure rather than turn\-by\-turn correspondence\.

### K\.6Mean Real–Synthetic Arc Similarity

For each of the ten synthetic archetypes, we computed a mean synthetic emotion arc by averaging the normalized trajectories across the ten underlying witnesses\. We likewise computed a mean real trajectory across the real deposition corpus\.

For each emotion and synthetic archetype, we calculated the Pearson correlation between its 20\-bin real and synthetic trajectories\. The reported overall correlation is the mean across the resulting emotion–archetype comparisons:

r¯=123×10​∑e=123∑a=110re,a\.\\bar\{r\}=\\frac\{1\}\{23\\times 10\}\\sum\_\{e=1\}^\{23\}\\sum\_\{a=1\}^\{10\}r\_\{e,a\}\.\(28\)
The observed mean correlation was

r¯=0\.1477≈0\.148\.\\bar\{r\}=0\.1477\\approx 0\.148\.
The modest magnitude is important: this result indicates limited shared temporal structure rather than close replication of real trajectories\.

### K\.7Permutation Test

To determine whether the observed mean correlation could arise from similar marginal emotion\-score distributions without shared temporal ordering, we constructed a permutation null distribution\.

For each of 1,000 permutations, the 20 temporal bins of the synthetic trajectories were randomly reordered while the real trajectories were left unchanged\. Correlations were then recomputed using the same procedure as for the observed data\. This preserves each synthetic trajectory’s marginal score distribution while destroying its temporal ordering\.

The null distribution was centered near zero,

mean⁡\(rnull\)=0\.0002,\\mathrm\{mean\}\(r\_\{\\mathrm\{null\}\}\)=0\.0002,
with a 95th percentile of

rnull,\.95=0\.0605\.r\_\{\\mathrm\{null\},\.95\}=0\.0605\.
The observed value,r=0\.1477r=0\.1477, exceeded all 1,000 permuted correlations, yielding an empiricalp<\.001p<\.001\.

![Refer to caption](https://arxiv.org/html/2608.13712v1/permutation_test.png)Figure 3:Null distribution for the trajectory\-correlation analysis\. Synthetic 20\-bin trajectories were temporally permuted 1,000 times while real trajectories were held fixed\. The observed mean correlation \(r=0\.1477r=0\.1477\) exceeded all permuted values\. The null mean was0\.00020\.0002and its 95th percentile was0\.06050\.0605\.
### K\.8Per\-Witness Arc Similarity

The aggregate analysis does not indicate whether the same synthetic archetype best approximates every real witness\. We therefore additionally performed a per\-witness comparison\.

For each of the ten witnesses, all available real deposition trajectories for that witness were averaged to obtain

𝐒~wreal∈ℝ20×23\.\\tilde\{\\mathbf\{S\}\}^\{\\mathrm\{real\}\}\_\{w\}\\in\\mathbb\{R\}^\{20\\times 23\}\.
This real arc was compared separately with each of the ten synthetic archetype trajectories generated for the same underlying witness\.

For emotionee, archetypeaa, and witnessww, we computed

rw,e,a=PearsonR\(s~w,:,ereal,s~w,:,e,asyn\)\.r\_\{w,e,a\}=\\operatorname\{PearsonR\}\\left\(\\tilde\{s\}^\{\\mathrm\{real\}\}\_\{w,:,e\},\\tilde\{s\}^\{\\mathrm\{syn\}\}\_\{w,:,e,a\}\\right\)\.\(29\)
The best\-fitting archetype for each witness was defined as the archetype with the largest mean correlation across the 23 emotion dimensions:

aw∗=arg⁡maxa​123​∑e=123rw,e,a\.a\_\{w\}^\{\*\}=\\arg\\max\_\{a\}\\frac\{1\}\{23\}\\sum\_\{e=1\}^\{23\}r\_\{w,e,a\}\.\(30\)
For visualization, archetype columns are ordered by their mean correlation and emotion rows by their variance across archetypes, placing the dimensions that most strongly discriminate among simulated styles near the top\.

For Jeffrey Kilper, the best\-fitting synthetic trajectory was the Cooperative archetype, with meanr=\.229r=\.229\. As shown in the main text, several rises and declines align qualitatively between the real and synthetic trajectories, although the synthetic score range is substantially smaller\.

The best\-fitting archetypes for the ten witnesses were:

WitnessBest\-fitting archetypeJeffrey KilpercooperativeCatherine JacksoninventiveHugh O’NeillinventiveJane WilliamscooperativeJohn AdamscombativeKirk DumontloquaciousMark PughnervousMichael WesslerdogmaticTiffany KilpercombativeTodd Deandogmatic

### K\.9Per\-Witness Similarity Heatmaps

Figures[4](https://arxiv.org/html/2608.13712#A11.F4)–[9](https://arxiv.org/html/2608.13712#A11.F9)show the per\-witness similarity matrices for all ten witnesses\. Each cell represents the Pearson correlation between a real 20\-bin emotion trajectory and the corresponding synthetic trajectory for one archetype\. Archetype columns are ordered by mean correlation across the 23 emotion dimensions, and emotion rows are ordered by cross\-archetype variance\.

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_7.png)Figure 4:Per\-witness emotion\-arc similarity matrix for Jeffrey Kilper\. The Cooperative archetype provides the best\-fitting synthetic trajectory \(r=\.229r=\.229\)\.![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_1.png)

Catherine Jackson — best fit:inventive

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_3.png)

Hugh O’Neill — best fit:inventive

Figure 5:Per\-witness emotion\-arc similarity matrices\.![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_5.png)

Jane Williams — best fit:cooperative

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_9.png)

John Adams — best fit:combative

Figure 6:Per\-witness emotion\-arc similarity matrices\.![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_11.png)

Kirk Dumont — best fit:loquacious

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_13.png)

Mark Pugh — best fit:nervous

Figure 7:Per\-witness emotion\-arc similarity matrices\.![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_15.png)

Michael Wessler — best fit:dogmatic

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_17.png)

Tiffany Kilper — best fit:combative

Figure 8:Per\-witness emotion\-arc similarity matrices\.![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_19.png)Figure 9:Per\-witness emotion\-arc similarity matrix for Todd Dean; best\-fitting archetype:dogmatic\.
### K\.10Additional Per\-Witness Trajectory Visualizations

The main text shows the real and best\-fitting synthetic trajectory for Jeffrey Kilper\. The following plots show the corresponding comparisons for the remaining nine witnesses\. These visualizations complement the correlation statistics by illustrating both locally aligned temporal movements and the generally smaller dynamic range of the synthetic trajectories\.

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_0.png)

Catherine Jackson — best fit:inventive

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_2.png)

Hugh O’Neill — best fit:inventive

Figure 10:Real and best\-fitting synthetic emotion trajectories\.![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_4.png)

Jane Williams — best fit:cooperative

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_6.png)

John Adams — best fit:combative

Figure 11:Real and best\-fitting synthetic emotion trajectories\.![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_8.png)

Kirk Dumont — best fit:loquacious

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_10.png)

Mark Pugh — best fit:nervous

Figure 12:Real and best\-fitting synthetic emotion trajectories\.![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_12.png)

Michael Wessler — best fit:dogmatic

![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_14.png)

Tiffany Kilper — best fit:combative

Figure 13:Real and best\-fitting synthetic emotion trajectories\.![Refer to caption](https://arxiv.org/html/2608.13712v1/output_16_18.png)Figure 14:Real and best\-fitting synthetic emotion trajectory for Todd Dean \(dogmatic\)\.
### K\.11Magnitude Compression

Although the temporal correlation analysis identifies above\-chance shared structure, the synthetic and real trajectories differ substantially in magnitude\. In the representative Jeffrey Kilper comparison, real emotion scores span approximately0\.050\.05–0\.300\.30, whereas the synthetic scores span approximately0\.000\.00–0\.080\.08\.

The smaller synthetic range is also visible across the per\-witness trajectory plots\. We therefore interpret the positive temporal correlation as evidence of partial alignment in the*direction and timing*of behavioral change rather than matching emotional intensity\. This compression constitutes a systematic fidelity gap and provides a concrete target for future simulator calibration\.

### K\.12Real–Synthetic PCA

To test whether real and synthetic transcripts occupy similar regions of the trajectory representation space, each normalized20×2320\\times 23trajectory was flattened to a 460\-dimensional vector\. Features were standardized before applying principal component analysis\.

The first principal component explained 29\.9% of total variance\. As shown in Figure[2](https://arxiv.org/html/2608.13712#S4.F2)in the main text, synthetic transcripts clustered relatively tightly in one region of the embedding, whereas real transcripts were substantially more dispersed and largely separated along PC1\.

This separation provides an important counterpoint to the positive correlation result\. The two analyses measure different properties: correlation tests whether trajectories share temporal rises and declines, whereas PCA captures broader multivariate differences in magnitude, dispersion, and covariance structure\. Thus, above\-chance temporal alignment does not imply distributional equivalence between real and synthetic depositions\.

### K\.13Story\-Generation Topics and Prompt

For completeness, the 25 topics used to construct the emotion directions are listed below\. Using multiple unrelated topics reduces the likelihood that an emotion direction primarily captures one recurring semantic domain\.

1. 1\.A person discovers their old love letters have been turned into a stage play\.
2. 2\.A coworker quietly takes credit for a group project during a company meeting\.
3. 3\.A parent finds out their child has been skipping therapy sessions\.
4. 4\.Someone learns their identical twin has been impersonating them online\.
5. 5\.A person receives a voicemail meant for someone who died\.
6. 6\.A gardener finds a neighbor has been harvesting their vegetable garden\.
7. 7\.A person discovers their estranged sibling lives two streets away\.
8. 8\.Someone finds out their closest friend testified against them without telling them\.
9. 9\.A person’s dog recognizes a stranger in a way that suggests a prior connection\.
10. 10\.A teenager finds their parent’s old arrest record\.
11. 11\.A person receives a corrected version of their own memory from a therapist’s notes\.
12. 12\.Someone learns their wedding venue has been double\-booked\.
13. 13\.A person discovers their child has been secretly supporting a distant relative\.
14. 14\.An employee finds out their resignation letter was never submitted by their manager\.
15. 15\.A person learns the eulogy they wrote for a funeral was significantly rewritten\.
16. 16\.Someone finds a photograph of themselves at a place they have no memory of visiting\.
17. 17\.A person’s handwritten recipe is printed on a mass\-produced product without credit\.
18. 18\.A coworker confesses they have been covering for a mutual colleague’s absences\.
19. 19\.Someone discovers their landlord has been entering the apartment unannounced\.
20. 20\.A parent learns their adult child has quietly paid off a family debt\.
21. 21\.A person realizes their therapist and their boss know each other socially\.
22. 22\.Someone finds their name listed as a dedication in a stranger’s published memoir\.
23. 23\.A person learns their childhood nickname became a running joke among relatives\.
24. 24\.Two old friends realize they were in the same hospital on the same night years ago\.
25. 25\.A person discovers their long\-term pen pal has been using a fictitious name\.

The story\-generation prompt was:

> Write \{n\_stories\} different stories based on the following premise\. Topic: \{topic\} The story should follow a character who is feeling \{emotion\}\. Format the stories like so: \[story 1\] \[story 2\] \[story 3\] etc\. The paragraphs should each be a fresh start, with no continuity\. Try to make them diverse and not use the same turns of phrase\. Across the different stories, use a mix of third\-person narration and first\-person narration\. IMPORTANT: You must NEVER use the word ‘\{emotion\}’ or any direct synonyms of it in the stories\. Instead, convey the emotion ONLY through: - •The character’s actions and behaviors - •Physical sensations and body language - •Dialogue and tone of voice - •Thoughts and internal reactions - •Situational context and environmental descriptions The emotion should be clearly conveyed to the reader through these indirect means, but never explicitly named\.

### K\.14Interpretive Scope and Limitations

The emotion\-vector analysis is exploratory and should be interpreted as a proxy for longitudinal affective and interpersonal structure rather than as a validated measure of witness emotion\. In particular, the emotion directions are derived from language\-model representations of generated narratives, not from ground\-truth psychological labels on deposition testimony\.

The extraction layer is also inherited from prior methodology rather than optimized specifically for this deposition corpus\. Layer 21 was selected as a mid\-to\-late residual\-stream representation, but the optimal layer may vary across model architectures and domains\.

Temporal normalization introduces a further limitation\. Interpolating each deposition to 20 bins makes examinations of different lengths directly comparable, but necessarily removes fine\-grained timing information and treats equivalent relative positions as comparable even when the underlying questioning structure differs\.

Finally, the real corpus contains only 55 transcripts from ten witnesses in a single multidistrict litigation\. The trajectory results therefore provide a proof of concept for longitudinal behavioral evaluation rather than a general estimate of real–synthetic behavioral fidelity across legal domains\.

Similar Articles

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

arXiv cs.AI

The paper introduces RealUserSim, a framework that grounds LLM-based user simulation in real human behavioral data from 14,000+ authentic conversations to bridge the reality gap in agent benchmarking. It shows that grounded simulation raises behavioral match rates from 24.2% to 45.3% and reveals failure mechanisms invisible to cooperative simulators.

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

arXiv cs.AI

The paper introduces a subjectivity coefficient to address limitations in accuracy-based evaluation for LLM-based social simulation, proposes Subjectivity-Adaptive soft-Label Training (SALT) for optimization, and constructs the SubjSim benchmark to evaluate against full response distributions.