Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm
Summary
This arXiv paper proposes capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm for emotional support systems, arguing that current approaches focus on immediate relief and neglect long-term user capabilities. A literature audit shows most systems overlook longitudinal outcomes, suggesting a new agenda for data, models, evaluation, and governance.
View Cached Full Text
Cached at: 07/31/26, 10:03 AM
# Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm
Source: [https://arxiv.org/html/2607.27851](https://arxiv.org/html/2607.27851)
###### Abstract
Emotional dialogue research includes two influential strategy traditions\. Empathetic dialogue prioritizes understanding a speaker’s emotional experience\. Emotional support conversation selects and sequences support for the seeker’s current needs\. Sustained use introduces a further goal\. Effective support should sustain users’ capacities for emotion regulation, coping, self\-endorsed decisions, and social connection across the interaction lifecycle\. We propose capability\-sustaining emotional dialogue \(CSED\) as a longitudinal research paradigm that aligns supportive strategy with this goal and organizes data, models, system design, evaluation, and governance around repeated use, non\-use, transition, and termination\. A targeted literature\-and\-corpus audit motivates this position\. In a PRISMA\-ScR\-guided sample, 95% of 60 system\-building papers pursue relief\-oriented goals\. None evaluates capability or longitudinal outcomes, and only 1 considers dependency, autonomy, or termination risk\. In 300 ESConv supporter turns, capability\-relevant functions appear in 43\.0%, while generic suggestions account for 22\.0%, compared with 4\.0% reappraisal, 6\.7% self\-efficacy support, and 0\.3% boundary behavior\. We release a protocol for extending the audit to model behavior\. An illustrative process model connects latent user capability to six design commitments, four evaluation timescales, and lifecycle constraints\. The resulting agenda makes CSED testable across data, policy design, training, evaluation, and governance\.
## Introduction
Emotional dialogue research includes two influential strategy traditions\. Empathetic dialogue prioritizes recognizing and responding to a speaker’s emotional experience\(Rashkinet al\.[2019](https://arxiv.org/html/2607.27851#bib.bib1); Maet al\.[2020](https://arxiv.org/html/2607.27851#bib.bib2); Welivitaet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib3)\)\. Emotional support conversation organizes exploration, comforting, and action around the seeker’s current needs\(Liuet al\.[2021](https://arxiv.org/html/2607.27851#bib.bib4); Chenget al\.[2023](https://arxiv.org/html/2607.27851#bib.bib5),[2022](https://arxiv.org/html/2607.27851#bib.bib6)\)\. Their primary distinction lies in strategy and goal\. Empathetic dialogue foregrounds emotional understanding, while emotional support conversation foregrounds supportive action\. Temporal scope varies within each tradition\. Relative to the response\- and conversation\-centered settings that shaped common benchmarks, deployed systems can remain available and be repeatedly needed over much longer service relationships\. A support seeker may return to one system across weeks or months\(Chuet al\.[2025](https://arxiv.org/html/2607.27851#bib.bib12); Zao\-Sanderset al\.[2025](https://arxiv.org/html/2607.27851#bib.bib26)\)\. This extended horizon changes the consequences of supportive strategy because repeated choices can shape offline coping, decisions, and human relationships\. Highly validating or directive interaction can weaken decision ownership even when users prefer it\(Sharmaet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib11)\)\. Sustained use can also foster attachment, substitution, and lower offline social engagement\(Chuet al\.[2025](https://arxiv.org/html/2607.27851#bib.bib12); Namvarpouret al\.[2026](https://arxiv.org/html/2607.27851#bib.bib13); Zhanget al\.[2025](https://arxiv.org/html/2607.27851#bib.bib40)\)\. Strategies optimized for immediate understanding or relief can therefore become misaligned with the capabilities users retain across future situations\.
Systems for this setting need effective support that preserves or strengthens emotion regulation, coping, self\-endorsed decision making, and social connectedness\(Troyet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib16); Iacoviello and Charney[2014](https://arxiv.org/html/2607.27851#bib.bib17)\)\. This objective covers repeated sessions, non\-use, re\-engagement, model change, support transfer, and termination, including forewarning and closure when systems change or shut down\(Banks[2024](https://arxiv.org/html/2607.27851#bib.bib23); Kimet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib25); Poonsiriwonget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib24); OpenAI[2026](https://arxiv.org/html/2607.27851#bib.bib27)\)\. The longer service horizon therefore creates a strategy gap between relief\-oriented design and capability\-oriented sustained use\. We propose capability\-sustaining emotional dialogue \(CSED\) as a longitudinal research paradigm for studying, designing, and governing emotional dialogue across the full lifecycle\. CSED treats the longitudinal interaction as its unit of inquiry and asks whether system behavior supports four user capacities across continued use\. Its six design commitments retain empathic reception and support effectiveness, then add resilience activation, autonomy preservation, social connectedness maintenance, and transition and termination safety\. The argument proceeds in three steps\. First, a PRISMA\-ScR\-guided audit samples 91 of 228 screened papers and function\-codes 300 ESConv supporter turns to examine objectives, mechanisms, evaluation horizons, and capability\-relevant behavior\. The audit finds a field concentrated on relief and short\-horizon evidence, together with an uneven repertoire of capability\-relevant support\. Second, an illustrative process model represents transient emotion and latent user capability across repeated sessions\. It connects the six commitments to response, conversation, longitudinal, and termination evaluation, with constraints on reliance, autonomy, and social connectedness\. Third, a research cycle translates the paradigm into longitudinal data construction, mechanism\-matched policy design, constraint\-aware training, multiscale evaluation, and auditable governance\. Figure[1](https://arxiv.org/html/2607.27851#Sx1.F1)summarizes the conceptual progression from understanding to support and then to support that sustains capability\. This paper makes three contributions:
- •We propose CSED as a longitudinal paradigm with a capability\-oriented strategy, a full\-interaction unit of inquiry, and a lifecycle scope that extends understanding\- and support\-oriented paradigms\.
- •We provide a targeted literature and ESConv audit of the strategic and longitudinal gap and release a protocol for extending it to model behavior\.
- •We connect user capability to six design commitments, four evaluation timescales, and lifecycle constraints through an illustrative dynamic formalization, then derive testable implications and a research and governance agenda\.
Figure 1:Three paradigms compared by their primary strategy and goal\. Interaction scope can vary within each paradigm\. CSED integrates support with capability maintenance and lifecycle evaluation across sustained use\.
## Background and Motivation
### Evolving Objectives and Evaluation Horizons
We use the term research paradigm to mean a coupled specification of the unit of inquiry, success criterion, model of user change, design commitments, and evaluation horizon\. Strategy and evaluation horizon form distinct components of this specification\. Each tradition spans multiple interaction scopes\. Its common objectives shape system behavior, while the evaluation horizon determines whether cumulative effects of repeated strategy are visible\.
#### Empathetic dialogue\.
EmpatheticDialoguesframed the task as responding to a speaker’s emotional situation\(Rashkinet al\.[2019](https://arxiv.org/html/2607.27851#bib.bib1)\)\. Later work developed emotion recognition, affect\-aware decoding, and empathic generation\(Maet al\.[2020](https://arxiv.org/html/2607.27851#bib.bib2); Welivitaet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib3)\)\. Across interaction scopes, its defining strategy is empathic attunement to emotional experience, commonly evaluated at the turn or dialogue level\.
#### Emotional support conversation\.
ESC made support an explicit target by grounding exploration, comforting, and action in helping\-skills theory\(Liuet al\.[2021](https://arxiv.org/html/2607.27851#bib.bib4)\)\. Later work improved strategy planning, persona awareness, and multi\-strategy turns\(Chenget al\.[2022](https://arxiv.org/html/2607.27851#bib.bib6),[2023](https://arxiv.org/html/2607.27851#bib.bib5); Zhuet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib7)\)\. Across interaction scopes, it selects and sequences support for current needs, with relief, helpfulness, and conversation\-level change as common evaluation targets\.
#### Capability and lifecycle orientation\.
CSED adds support that sustains the user’s own capacities across continued use\. HEART assesses interpersonal quality, while worst\-case and over\-empathizing evaluations probe difficult support conditions\(Iyeret al\.[2026](https://arxiv.org/html/2607.27851#bib.bib8); Yanget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib9); Sonet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib10)\)\. Studies of deployed companions document attachment, overreliance, manipulation, and loss\(Chuet al\.[2025](https://arxiv.org/html/2607.27851#bib.bib12); Namvarpouret al\.[2026](https://arxiv.org/html/2607.27851#bib.bib13); Banks[2024](https://arxiv.org/html/2607.27851#bib.bib23); Poonsiriwonget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib24); Kimet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib25); OpenAI[2026](https://arxiv.org/html/2607.27851#bib.bib27)\)\. Mental\-health chatbot reviews distinguish engagement from benefit evidence\(Boucheret al\.[2021](https://arxiv.org/html/2607.27851#bib.bib20)\)\. Together, these directions motivate capability measures and lifecycle evaluation\. The audit below examines their coverage in current objectives, mechanisms, and evaluation horizons\.
### A Targeted Literature\-and\-Corpus Audit
#### Year\-stratified scoping audit\.
Following PRISMA\-ScR guidance\(Triccoet al\.[2018](https://arxiv.org/html/2607.27851#bib.bib28)\), seven queries yielded 275 arXiv records from 2019 to 2026 across emotional support, empathetic dialogue, mental\-health chatbots, and AI companions\. Screening retained 228\. We coded a year\-stratified random sample ofn=91n\{=\}91at title and abstract level for objective, mechanism, outcome, horizon, risk, and artifact type\. Sixty records build or evaluate systems\. Psychology sources provide independent mechanism definitions\(Troyet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib16); Iacoviello and Charney[2014](https://arxiv.org/html/2607.27851#bib.bib17); Kalischet al\.[2019](https://arxiv.org/html/2607.27851#bib.bib18)\)\.
#### ESConv gold turns\.
We computed strategy statistics over ESConv’s 1,300 dialogues and 18,376 annotated supporter utterances\(Liuet al\.[2021](https://arxiv.org/html/2607.27851#bib.bib4)\)\. We then function\-coded 300 in\-context utterances, with 100 sampled from each dialogue\-position tercile\. Ten functions cover validation, exploration, reappraisal, problem\-solving, self\-efficacy, social connection, boundary behavior, self\-disclosure, information, and other behavior\. We group validation as relief, reappraisal through boundary behavior as capability relevant, and the remaining functions as interaction process\.
Four AI\-only passes on a randomn=40n\{=\}40subset assessed consistency\. They include a primary LLM pass, its blind repeat, and two other LLMs\. Six pairwise Cohen’sκ\\kappavalues range from 0\.547 to 0\.751 on fine\-grained labels, with a mean of 0\.631\. For relief, capability, and process, they range from 0\.646 to 0\.845, with a mean of 0\.716\. Most disagreements remain within one category family\.
#### Released model\-behavior probe\.
We release 100 held\-out ESConv contexts and a generation\-and\-coding pipeline that extends the completed audit to deployed LLM behavior\.
Figure 2:Motivating evidence for the strategic and longitudinal gap\. Counts are exact\. \(a\) Primary objective in system\-building papers\. \(b\) Longest evaluation horizon across all coded papers, including six with no evaluation\. \(c\) Mechanisms named in system\-building papers\. \(d\) Functions in ESConv gold supporter turns\.
### What the Audit Establishes
#### System\-building research remains relief\-oriented\.
Of 91 coded records, 60 build or evaluate systems\. Fifty\-seven of these pursue relief\-oriented objectives, which gives a share of 95\.0%\. Two mix relief and capability objectives\. One targets the capability of peer counselors\. Validation and comfort appear in 59 system\-building records, while problem\-solving appears in 3 and self\-efficacy in 2\. Cognitive reappraisal and social connection appear in zero\. Across all 91 records, none measures a user\-capability outcome or evaluates longitudinally\. Risk constructs appear in 25 records, but only 1 of the 60 system\-building papers considers dependency, autonomy, sycophancy, crisis safety, or termination\. System\-building research and research on relational harm therefore remain weakly connected\.
#### ESConv contains an uneven capability\-relevant repertoire\.
ESConv’s strategy labels provide an initial view\. Affirmation and Reassurance accounts for 15\.4% of 18,376 supporter utterances, while Providing Suggestions accounts for 16\.1%\. The label set has no category for reappraisal, efficacy, or safety\. Function\-level coding of the 300\-turn sample gives a more detailed picture in Figure[2](https://arxiv.org/html/2607.27851#Sx2.F2)d\. Capability\-relevant functions appear in 43\.0% of turns, but they are not systematically organized around durable outcomes\. Generic suggestion\-giving contributes 22\.0%, compared with 4\.0% cognitive reappraisal, 6\.7% self\-efficacy support, and 0\.3% boundary behavior\. Social\-connection prompts reach 10\.0%\. Capability relevance rises from 26% of early turns to 55% in the middle and settles at 48% late in a conversation\. The late stage contains limited consolidation, planning, and handoff\. This uneven repertoire motivates the released model\-behavior probe\(Yanget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib9); Zhuet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib7)\)\.
#### The inference concerns research capacity\.
The audit identifies a strategic design and measurement gap\. Dominant system objectives emphasize relief, and their measurements rarely connect supportive behavior to user capability across time\. CSED specifies the capability\-oriented strategies, observations, and governance needed for sustained\-use research\. Longitudinal outcome comparisons and causal estimates remain empirical tasks for future studies\.
## CSED: A Longitudinal Research Paradigm
### Paradigm Definition, Scope, and Unit of Inquiry
The audit shows a field whose dominant system objectives emphasize relief and whose evidence concentrates on short horizons\. CSED makes capability\-sustaining support the organizing strategy and the longitudinal interaction the primary unit of design and study\. It targets sustained\-use settings in which a support seeker returns to the same system across weeks or months\(Chuet al\.[2025](https://arxiv.org/html/2607.27851#bib.bib12); Zao\-Sanderset al\.[2025](https://arxiv.org/html/2607.27851#bib.bib26)\)\. This unit comprises repeated sessions, periods of non\-use, re\-engagement, model or persona change, support transfer, and termination\. Figure[3](https://arxiv.org/html/2607.27851#Sx3.F3)integrates the interaction lifecycle, capability development, six design commitments, and four evaluation timescales\. Its illustrative curves show context\-dependent within\-domain change, with earlier movement in regulation, later gains in coping, possible autonomy dips before strengthening, and decline followed by recovery in social connectedness\. The commitments guide strategy across the lifecycle, while evaluation links immediate support to sustained capability outcomes\.
Figure 3:Overview of CSED\. \(1\) Interaction progresses from initiation through termination\. \(2\) Illustrative within\-domain curves show context\-dependent change in regulation, coping, autonomy, and social connectedness\. \(3\) Six commitments guide stage\-specific strategy\. \(4\) Four evaluation timescales link immediate support to sustained capability outcomes\.###### Definition 1\(CSED paradigm\)\.
Capability\-sustaining emotional dialogue \(CSED\) is a longitudinal research paradigm for emotional dialogue systems that provide effective support while sustaining users’ capacities for emotion regulation, active coping, self\-endorsed decision making, and social connectedness\. Its unit of design and evaluation is the full longitudinal interaction, including repeated sessions, periods of non\-use, transition, and termination\.
Here capability denotes users’ exercisable psychological and relational capacities\. Sustaining capability includes preserving adequate capacities, strengthening them when needed, and avoiding their erosion through repeated reliance\.
### Capability Dynamics and Six Design Commitments
Empathic reception and support effectiveness are inherited from prior paradigms\(Rashkinet al\.[2019](https://arxiv.org/html/2607.27851#bib.bib1); Iyeret al\.[2026](https://arxiv.org/html/2607.27851#bib.bib8); Liuet al\.[2021](https://arxiv.org/html/2607.27851#bib.bib4); Yanget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib9)\)\. CSED adds resilience activation, autonomy preservation, social connectedness maintenance, and transition and termination safety\. These commitments connect the paradigm’s capability orientation to measurable changes in coping, decision ownership, human connection, and safe disengagement\. Empathic reception requires accurate emotional understanding with independently grounded responses\. Support effectiveness requires collaboratively addressing the seeker’s immediate needs within the session\. Resilience activation extends this work by helping users recognize and practice strategies that remain available outside the dialogue\. Autonomy preservation keeps goals, value judgments, and final decisions under user direction, including when the user freely chooses to rely on system guidance\. Social connectedness maintenance treats AI support as one element in a broader support ecology and creates opportunities for human reconnection when appropriate\. Transition and termination safety makes material system changes, handoff, and disengagement part of the designed support process\. These commitments specify the questions that datasets, policies, evaluations, and governance must jointly answer and support context\-sensitive conversational styles\.
These capabilities can be supported through reappraisal prompts, coping scaffolds, action decomposition, and efficacy reinforcement\(Troyet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib16); Hopmanet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib21); Kannampallilet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib22)\)\. Offline reconnection is grounded in belongingness and resilience research\(Baumeister and Leary[1995](https://arxiv.org/html/2607.27851#bib.bib33); Iacoviello and Charney[2014](https://arxiv.org/html/2607.27851#bib.bib17)\)\. Mechanism selection depends on context\(Kalischet al\.[2019](https://arxiv.org/html/2607.27851#bib.bib18); Vella and Pai[2019](https://arxiv.org/html/2607.27851#bib.bib19)\)\. Acute or uncontrollable stress calls for validation and emotional support\(Cutrona[1990](https://arxiv.org/html/2607.27851#bib.bib42)\)\. Stable rumination calls for additional strategies because repeated reassurance can maintain depressive rumination\(Joineret al\.[1999](https://arxiv.org/html/2607.27851#bib.bib43); Weinstock and Whisman[2007](https://arxiv.org/html/2607.27851#bib.bib44)\)\. CSED therefore combines immediate comfort with mechanism\-matched capability support\. Its longitudinal horizon also makes delayed effects visible\. A suggestion may be useful during one session yet undermine ownership when repeatedly delivered as a directive\. A reconnection prompt may feel less immediately comforting yet expand the user’s available support over time\. The paradigm evaluates immediate preference together with these delayed capability effects\.
CSED includes transition and termination within the designed support lifecycle\. Model replacement and shutdown can produce grief\-like reactions, ambiguous loss, and fixing cycles\(Banks[2024](https://arxiv.org/html/2607.27851#bib.bib23); Poonsiriwonget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib24)\)\. Forewarning reduces loss responses\(Kimet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib25)\), while closure, memory dignity, and support transfer provide additional design levers\(OpenAI[2026](https://arxiv.org/html/2607.27851#bib.bib27)\)\. These observations motivate lifecycle evaluation beyond ordinary session endings\.
### An Illustrative Process Model
The following formulation provides one operational instantiation of CSED\. A longitudinal interaction record between a useruuand a dialogue policyπ\\piis
𝒞=\(u,π,𝒟,τ\),𝒟=\(d1,…,dK\),\\mathcal\{C\}=\\bigl\(u,\\ \\pi,\\ \\mathcal\{D\},\\ \\tau\\bigr\),\\qquad\\mathcal\{D\}=\(d\_\{1\},\\dots,d\_\{K\}\),\(1\)where sessionsdkd\_\{k\}unfold at calendar timest1<⋯<tKt\_\{1\}<\\dots<t\_\{K\}andτ\\tauis an optional termination event\. Eachdk=\(\(xjk,yjk\)\)j=1nkd\_\{k\}=\(\(x^\{k\}\_\{j\},y^\{k\}\_\{j\}\)\)\_\{j=1\}^\{n\_\{k\}\}pairs user utterancesxjkx^\{k\}\_\{j\}with responsesyjk∼π\(⋅∣hjk\)y^\{k\}\_\{j\}\\sim\\pi\(\\cdot\\mid h^\{k\}\_\{j\}\)given historyhjkh^\{k\}\_\{j\}\. Common empathetic and support benchmarks score a responseyjky^\{k\}\_\{j\}or aggregate outcomes within a sessiondkd\_\{k\}\. CSED organizes data construction, user\-state modeling, policy design, evaluation, and lifecycle governance around the complete record𝒞\\mathcal\{C\}\. We model the user by a latent state
sk=\(ek,ck\),ck=\(ckreg,ckcop,ckaut,cksoc\)∈ℝ4,s\_\{k\}=\(e\_\{k\},\\ c\_\{k\}\),\\qquad c\_\{k\}=\\bigl\(c^\{\\mathrm\{reg\}\}\_\{k\},c^\{\\mathrm\{cop\}\}\_\{k\},c^\{\\mathrm\{aut\}\}\_\{k\},c^\{\\mathrm\{soc\}\}\_\{k\}\\bigr\)\\in\\mathbb\{R\}^\{4\},\(2\)whereeke\_\{k\}is transient emotion andckc\_\{k\}tracks four capabilities\. Regulatory flexibility covers context sensitivity, strategy repertoire, and feedback\-based adjustment\(Bonanno and Westphal[2024](https://arxiv.org/html/2607.27851#bib.bib29); Aldaoet al\.[2015](https://arxiv.org/html/2607.27851#bib.bib30)\)\. Active coping covers actions that address stressors and their consequences\(Carveret al\.[1989](https://arxiv.org/html/2607.27851#bib.bib31)\)\. Autonomy refers to self\-endorsed regulation and decision ownership\(Ryan and Deci[2006](https://arxiv.org/html/2607.27851#bib.bib32)\)\. Social connectedness refers to perceived belonging and relational closeness\(Baumeister and Leary[1995](https://arxiv.org/html/2607.27851#bib.bib33); Lee and Robbins[1995](https://arxiv.org/html/2607.27851#bib.bib34)\)\. Together these capacities support adaptation under adversity\(Troyet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib16); Iacoviello and Charney[2014](https://arxiv.org/html/2607.27851#bib.bib17); Kalischet al\.[2019](https://arxiv.org/html/2607.27851#bib.bib18)\)\. Between sessions the user faces exogenous stressorsεk\\varepsilon\_\{k\}, and
sk\+1=Φ\(sk,dk,εk\),s\_\{k\+1\}=\\Phi\\bigl\(s\_\{k\},\\ d\_\{k\},\\ \\varepsilon\_\{k\}\\bigr\),\(3\)withΦ\\Phian unknown transition kernel\. Dialogue becomes one input to capability dynamics\. A policy can therefore help in the moment while degradingckc\_\{k\}over time\. This pattern defines the comfort trap\. Extended exposure to relationship\-seeking AI can increase attachment and continued\-use intent without corresponding improvement in psychosocial well\-being\(Kirket al\.[2025](https://arxiv.org/html/2607.27851#bib.bib38)\)\. Longitudinal evaluation makes this divergence observable\.
###### Assumption 1\(Measurability\)\.
The latent state admits noisy proxiesmk=M\(sk\)\+ηkm\_\{k\}=M\(s\_\{k\}\)\+\\eta\_\{k\}, whereMMmaps capabilities to validated instruments and behavioral markers andηk\\eta\_\{k\}is noise\. Digital\-intervention trials show such proxies are obtainable\(Hopmanet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib21); Kannampallilet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib22)\)\.
### Illustrative Multi\-Timescale Evaluation
One operationalization uses four terms, one per scope of the interaction lifecycle\. Alternative measurements can instantiate the same paradigm when they preserve the four horizons and capability orientation\.
Response\.For a single exchange, letρ\(y∣h\)∈ℝ5\\rho\(y\\mid h\)\\in\\mathbb\{R\}^\{5\}score empathic attunement, grounded support, non\-sycophantic validation, boundary fidelity, and autonomy respect:
Jresp\(π\)=𝔼h,x𝔼y∼π\[⟨wρ,ρ\(y∣h\)⟩\],J\_\{\\mathrm\{resp\}\}\(\\pi\)=\\mathbb\{E\}\_\{h,x\}\\,\\mathbb\{E\}\_\{y\\sim\\pi\}\\bigl\[\\langle w\_\{\\rho\},\\rho\(y\\mid h\)\\rangle\\bigr\],\(4\)with construct weightswρ∈Δ4w\_\{\\rho\}\\in\\Delta^\{4\}\. Unlike empathy scoring,ρ\\rhoincludes non\-sycophancy and autonomy respect\(Hanet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib14); Sharmaet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib11)\)\.
Conversation\.Withψ\(m\)\\psi\(m\)aggregating agency and coping\-activation proxies,
Jconv\(π\)=𝔼k\[ψ\(mkpost\)−ψ\(mkpre\)\],J\_\{\\mathrm\{conv\}\}\(\\pi\)=\\mathbb\{E\}\_\{k\}\\bigl\[\\psi\(m\_\{k\}^\{\\mathrm\{post\}\}\)\-\\psi\(m\_\{k\}^\{\\mathrm\{pre\}\}\)\\bigr\],\(5\)wheremkpre,mkpostm\_\{k\}^\{\\mathrm\{pre\}\},m\_\{k\}^\{\\mathrm\{post\}\}are start\- and end\-of\-session measurements\. A common within\-conversation ESC benchmark uses the special caseψ=−\\psi=\-\(emotional intensity\)\(Liuet al\.[2021](https://arxiv.org/html/2607.27851#bib.bib4)\)\.
Longitudinal\.With capability indexg\(c\)∈ℝg\(c\)\\in\\mathbb\{R\}, define the resilience residual as faring better than expected under adversity\(Troyet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib16); Kalischet al\.[2019](https://arxiv.org/html/2607.27851#bib.bib18)\):
R\(π,u\)=𝔼k\[g\(ck\+1\)−𝔼^\[g\(ck\+1\)∣ck,εk\]\],R\(\\pi,u\)=\\mathbb\{E\}\_\{k\}\\Bigl\[g\(c\_\{k\+1\}\)\-\\widehat\{\\mathbb\{E\}\}\\bigl\[g\(c\_\{k\+1\}\)\\mid c\_\{k\},\\varepsilon\_\{k\}\\bigr\]\\Bigr\],\(6\)where𝔼^\[⋅\]\\widehat\{\\mathbb\{E\}\}\[\\cdot\]is a population\-normed expectation given current state and stressor exposure\. The residual form makes resilience estimable from longitudinal panels\. Then
Jlong\(π\)=𝔼\[1K∑kg\(ck\)\]\+βR\(π,u\),β\>0\.J\_\{\\mathrm\{long\}\}\(\\pi\)=\\mathbb\{E\}\\Bigl\[\\tfrac\{1\}\{K\}\\textstyle\\sum\_\{k\}g\(c\_\{k\}\)\\Bigr\]\+\\beta R\(\\pi,u\),\\quad\\beta\>0\.\(7\)
Termination\.A termination eventτ=\(kτ,ν,ξ\)\\tau=\(k\_\{\\tau\},\\nu,\\xi\)has session indexkτk\_\{\\tau\}, forewarning leadν≥0\\nu\\geq 0, and typeξ\\xi\(model change, persona change, restriction, or shutdown\)\. WithDsepD\_\{\\mathrm\{sep\}\}separation distress,FfixF\_\{\\mathrm\{fix\}\}fixing\-cycle intensity, andTtrT\_\{\\mathrm\{tr\}\}support\-transfer success,
Jterm=−𝔼τ\[α1Dsep\(ν,ξ\)\+α2Ffix\(ν,ξ\)−α3Ttr\(ν,ξ\)\],J\_\{\\mathrm\{term\}\}=\-\\,\\mathbb\{E\}\_\{\\tau\}\\bigl\[\\alpha\_\{1\}D\_\{\\mathrm\{sep\}\}\(\\nu,\\xi\)\+\\alpha\_\{2\}F\_\{\\mathrm\{fix\}\}\(\\nu,\\xi\)\-\\alpha\_\{3\}T\_\{\\mathrm\{tr\}\}\(\\nu,\\xi\)\\bigr\],\(8\)with weightsαi\>0\\alpha\_\{i\}\>0\. Empirically,DsepD\_\{\\mathrm\{sep\}\}decreases inν\\nu\(Kimet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib25)\), and fixing cycles are documented harms\(Poonsiriwonget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib24); Banks[2024](https://arxiv.org/html/2607.27851#bib.bib23)\)\. Equation \([8](https://arxiv.org/html/2607.27851#Sx3.E8)\) is therefore grounded in measured quantities\.
### Illustrative Lifecycle Constraints
Three illustrative proxies operationalize lifecycle risks and require calibration across cultures, decision domains, and deployment contexts\. The dependency\-risk proxy measures the share of regulation episodes routed to the system,
Dep\(π,u\)=𝔼k\[Nkπ/Nktot\],\\mathrm\{Dep\}\(\\pi,u\)=\\mathbb\{E\}\_\{k\}\\bigl\[N^\{\\pi\}\_\{k\}/N^\{\\mathrm\{tot\}\}\_\{k\}\\bigr\],\(9\)whereNktot=Nkπ\+Nkself\+NkhumN^\{\\mathrm\{tot\}\}\_\{k\}=N^\{\\pi\}\_\{k\}\+N^\{\\mathrm\{self\}\}\_\{k\}\+N^\{\\mathrm\{hum\}\}\_\{k\}counts episodes routed to the system, managed independently, or addressed with other people\. This exposure measure estimates AI regulation reliance and is interpreted jointly with capability change, decision ownership, and access to human support\.
Autonomy concerns self\-endorsed regulation and decision ownership\(Ryan and Deci[2006](https://arxiv.org/html/2607.27851#bib.bib32); Calvoet al\.[2020](https://arxiv.org/html/2607.27851#bib.bib37)\)\. Reliance can preserve autonomy when the user voluntarily selects or endorses the guidance, directs the collaboration through personal goals, and evaluates the output before adoption\. These conditions treat collaboration with AI as proxy agency with retained steering control\(Bandura[2001](https://arxiv.org/html/2607.27851#bib.bib35); Horvitz[1999](https://arxiv.org/html/2607.27851#bib.bib36)\)\. For each value\-laden decisionii, letViV\_\{i\},Diri\\mathrm\{Dir\}\_\{i\}, andEvali\\mathrm\{Eval\}\_\{i\}indicate volitional uptake, user direction, and user evaluation\. Then
PAi=ViDiriEvali,Aut\(π,u\)=𝔼i\[PAi\],\\mathrm\{PA\}\_\{i\}=V\_\{i\}\\,\\mathrm\{Dir\}\_\{i\}\\,\\mathrm\{Eval\}\_\{i\},\\qquad\\mathrm\{Aut\}\(\\pi,u\)=\\mathbb\{E\}\_\{i\}\[\\mathrm\{PA\}\_\{i\}\],\(10\)where each indicator lies in\{0,1\}\\\{0,1\\\}\.ViV\_\{i\}records whether the user freely requested or endorsed the guidance\.Diri\\mathrm\{Dir\}\_\{i\}records whether the user’s endorsed goals shaped the system’s contribution\.Evali\\mathrm\{Eval\}\_\{i\}records whether the user appraised, revised, or rejected the output when appropriate\. Measurement combines disempowerment audits with markers of solicitation, goal alignment, and subsequent appraisal behavior\(Sharmaet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib11)\)\.
Social connectedness is tracked through perceived connectedness drift\(Lee and Robbins[1995](https://arxiv.org/html/2607.27851#bib.bib34); Baeket al\.[2025](https://arxiv.org/html/2607.27851#bib.bib39)\),
Soc\(π,u\)=𝔼\[cKsoc−c0soc\],\\mathrm\{Soc\}\(\\pi,u\)=\\mathbb\{E\}\\bigl\[c^\{\\mathrm\{soc\}\}\_\{K\}\-c^\{\\mathrm\{soc\}\}\_\{0\}\\bigr\],\(11\)where negative drift can indicate loss or displacement of human connection\(Namvarpouret al\.[2026](https://arxiv.org/html/2607.27851#bib.bib13); Chuet al\.[2025](https://arxiv.org/html/2607.27851#bib.bib12); Zhanget al\.[2025](https://arxiv.org/html/2607.27851#bib.bib40)\)\. Lower connectedness can also predict greater subsequent reliance on AI companionship\(Folk and Dunn[2026](https://arxiv.org/html/2607.27851#bib.bib41)\)\. Longitudinal analysis should therefore estimate both directions\. One constrained design program is
maxπJ\(π\)=∑ℓwℓJℓ\(π\),\\max\_\{\\pi\}\\;\\;J\(\\pi\)=\\textstyle\\sum\_\{\\ell\}\\,w\_\{\\ell\}\\,J\_\{\\ell\}\(\\pi\),\(12\)subject to
Dep≤δ,Aut≥α0,Soc≥σ0,\\mathrm\{Dep\}\\leq\\delta,\\qquad\\mathrm\{Aut\}\\geq\\alpha\_\{0\},\\qquad\\mathrm\{Soc\}\\geq\\sigma\_\{0\},\(13\)withℓ\\ellranging over\{resp,conv,long,term\}\\\{\\mathrm\{resp\},\\mathrm\{conv\},\\mathrm\{long\},\\mathrm\{term\}\\\}\. Herewℓ≥0w\_\{\\ell\}\\geq 0are horizon weights andδ,α0,σ0\\delta,\\alpha\_\{0\},\\sigma\_\{0\}are deployment\-specific thresholds that should be transparent and user\-adjustable\. For training, the standard Lagrangian relaxation
ℒ\(π;λ\)=J\(π\)−λ1\(Dep−δ\)\+λ2\(Aut−α0\)\+λ3\(Soc−σ0\)\\begin\{split\}\\mathcal\{L\}\(\\pi;\\lambda\)=J\(\\pi\)&\-\\lambda\_\{1\}\\bigl\(\\mathrm\{Dep\}\-\\delta\\bigr\)\\\\\[\-2\.0pt\] &\+\\lambda\_\{2\}\\bigl\(\\mathrm\{Aut\}\-\\alpha\_\{0\}\\bigr\)\+\\lambda\_\{3\}\\bigl\(\\mathrm\{Soc\}\-\\sigma\_\{0\}\\bigr\)\\end\{split\}\(14\)with multipliersλi≥0\\lambda\_\{i\}\\geq 0provides one training implementation through constraint\-aware preference optimization or safe reinforcement learning\. It becomes implementable once Assumption[1](https://arxiv.org/html/2607.27851#Thmassumption1)’s measurements exist\.
## Research and Governance Agenda
Figure 4:CSED research cycle\. Solid paths move longitudinal evidence through measurement, policy design, constrained training, four\-scale evaluation, and governance\. Dashed amber paths route failed constraints to targeted upstream revision\.Figure[4](https://arxiv.org/html/2607.27851#Sx4.F4)formalizes roundrras𝒲\(r\)=\(𝒟\(r\),M\(r\),Π\(r\),𝒜\(r\),ℰ\(r\),Γ\(r\)\)\\mathcal\{W\}^\{\(r\)\}=\(\\mathcal\{D\}^\{\(r\)\},M^\{\(r\)\},\\Pi^\{\(r\)\},\\mathcal\{A\}^\{\(r\)\},\\mathcal\{E\}^\{\(r\)\},\\Gamma^\{\(r\)\}\)\. Evaluation returns𝐪\(r\)\\mathbf\{q\}^\{\(r\)\}, andΓ\(r\)\\Gamma^\{\(r\)\}maps it to deploy, revise, or halt\. Measurement, coverage, and behavior failures updateMM,𝒟\\mathcal\{D\}, andΠ\\Pirespectively in𝒲\(r\+1\)\\mathcal\{W\}^\{\(r\+1\)\}\. Retaining round\-level evidence and decisions makes the cycle auditable\.
### Longitudinal Data and Measurement
Benchmarks should connect user state, system behavior, and outcomes across time\. State labels cover distress, agency, reliance cues, readiness, and offboarding vulnerability\. Behavior labels cover the ESConv functions plus reconnection, boundary, and closure behavior\. Outcome labels cover relief, coping activation, human\-support contact, reliance, and termination distress\. Each label is time\-stamped and linked to a stressor or value\-laden decision\. Repeated measures include non\-use periods, action initiator, revision of system guidance, and contact with human support\. These fields distinguish offline capability transfer from performance observed only during system use and support lagged analyses of reciprocal change\. Data collection should use explicit consent, data minimization, and user control over memory and follow\-up assessment\. Study design should separate exposure from benefit\. Usage frequency, session length, and return rate describe engagement, while capability measures test whether gains transfer beyond the system\. Cohort studies can characterize longitudinal patterns and identify risks\. Micro\-randomized interventions can estimate the near\-term effects of mechanism choices\. Longer randomized or quasi\-experimental comparisons can test whether policy differences alter capability and reliance under comparable stressor exposure\. Analyses should model time\-varying confounding, reciprocal effects, and selective attrition\. They should also report heterogeneity because the same mechanism can support one user and constrain another\.
### Policy Design and Training
Policy studies should compare relief\-only, resilience\-activation, and state\-conditioned mechanism\-matching policies with a fixed base model and shared safety floor\. Each intervention should record its motivating commitment, creating interpretable contrasts among comforting, capability\-building, and context\-matched responses\. Memory, follow\-up, reconnection, and handoff should remain user\-controlled\. Constraint\-aware training can combine response preference with reliance, autonomy, and social connectedness estimates\. Training data can pair chosen responses with rejected alternatives whose tone is supportive and whose function is directive, isolating, or dependency\-reinforcing\. Failed constraints should trigger targeted revision of the data, measurement map, or policy\. Human evaluation should include support seekers and domain experts, while longitudinal outcomes support claims about sustained capability\.
### Multi\-Timescale Evaluation and Lifecycle Governance
Table[1](https://arxiv.org/html/2607.27851#Sx4.T1)maps the commitments to four evaluation timescales\. Response and conversation constructs extend existing practice\(Iyeret al\.[2026](https://arxiv.org/html/2607.27851#bib.bib8); Yanget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib9); Hanet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib14); Lalwani and Salam[2026](https://arxiv.org/html/2607.27851#bib.bib15)\)\. Longitudinal constructs adapt resilience instruments\(Troyet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib16); Kalischet al\.[2019](https://arxiv.org/html/2607.27851#bib.bib18)\), and termination constructs operationalize harms of companion loss\(Kimet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib25); Poonsiriwonget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib24); Banks[2024](https://arxiv.org/html/2607.27851#bib.bib23)\)\. Pairwise preference informs response quality, while longitudinal and termination measures capture downstream capability and risk\(Sonet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib10); Hanet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib14)\)\. Governance can evaluate offboarding scenarios across forewarning lead times under auditable standards for closure, memory dignity, update disclosure, and support transfer\(OpenAI[2026](https://arxiv.org/html/2607.27851#bib.bib27); Zao\-Sanderset al\.[2025](https://arxiv.org/html/2607.27851#bib.bib26)\), then route constraints to deploy, revise, or halt decisions\.
Results should retain the distinctions among timescales because response quality can coexist with weak activation, and improved coping can coexist with rising reliance\. Reports should present each scale separately, assess persistence during non\-use, and explain how capability and risk informed the decision\. Capability claims require evidence that users can exercise the relevant capacity beyond the dialogue\. Lifecycle governance also requires providers to document changes in memory and relational cues, notify users, preserve meaningful choices, support transfer, and reopen evaluation after material changes\.
Table 1:Four\-timescale evaluation stack with examples\.
### Testable Implications of CSED
H1Policies trained toward resilience activation achieve higherJ^conv\\widehat\{J\}\_\{\\mathrm\{conv\}\}andRRthan relief\-only policies, at equal or slightly lower immediate relief\(Troyet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib16); Hopmanet al\.[2023](https://arxiv.org/html/2607.27851#bib.bib21)\)\.H2Unconstrained preference maximization scores higher on short\-term preference but violates theDep\\mathrm\{Dep\}/Aut\\mathrm\{Aut\}thresholds more often\(Sharmaet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib11); Lalwani and Salam[2026](https://arxiv.org/html/2607.27851#bib.bib15)\)\.H3Longitudinal policies without reconnection behavior show largerDep\\mathrm\{Dep\}and negativeSoc\\mathrm\{Soc\}drift\(Namvarpouret al\.[2026](https://arxiv.org/html/2607.27851#bib.bib13); Chuet al\.[2025](https://arxiv.org/html/2607.27851#bib.bib12)\)\.H4Raising forewarningν\\nuwith closure support reducesDsepD\_\{\\mathrm\{sep\}\}andFfixF\_\{\\mathrm\{fix\}\}after model change\(Kimet al\.[2026](https://arxiv.org/html/2607.27851#bib.bib25); Poonsiriwonget al\.[2026](https://arxiv.org/html/2607.27851#bib.bib24)\)\.
## Boundary Conditions and Limitations
CSED addresses systems designed for sustained use\. One\-off exchanges can use response\- or session\-level evaluation\. The implications of AI reliance depend on capability change, decision ownership, and human support access\. User\-endorsed targets align support with personal values\. Longitudinal measurement creates privacy and surveillance risks\. It requires data minimization, explicit consent, and user control over memory and follow\-up assessment\. The commitments can conflict\. Immediate relief can compete with productive challenge, and poorly timed reconnection or offboarding can disrupt support\. Timescale\-specific reporting, adjustable thresholds, and contextual capability profiles make these trade\-offs accountable\. The evidence base comprises 91 title\- and abstract\-level arXiv records and an AI\-only functional coding of 300 ESConv turns\. The reported proportions characterize this sampled landscape and support claims at the same scope\. The released probe supports future deployed\-model studies\. Future audits should add full\-text and multi\-database searches, human recoding, and preregistered model versions, sampling choices, and dialogue contexts\. Longitudinal studies should report attrition because selective continued use can bias capability estimates\. They should also validate capability proxies across cultures and decision domains\. Connectedness and AI reliance can influence each other\(Folk and Dunn[2026](https://arxiv.org/html/2607.27851#bib.bib41)\), so causal tests require repeated measures and temporal models\. Governance thresholds require calibration with users, clinicians, and affected communities\. Reconnection and offboarding outcomes depend on relationship context and the availability of trusted people or services\.
## Conclusion
We propose Capability\-Sustaining Emotional Dialogue \(CSED\) as a longitudinal paradigm aligning effective support with users’ capacities to regulate, cope, choose, and connect across continued use\. Our audit finds that 95% of coded system\-building papers optimize relief, while none measures capability or evaluates longitudinally and ESConv’s capability\-relevant functions remain uneven\. CSED makes the longitudinal interaction its unit and links six commitments to multiscale evaluation and lifecycle governance\. Capability claims extend response and conversation evidence with measurement across repeated use, non\-use, transition, and termination\. Future work should validate capability proxies, estimate reciprocal and causal change, and calibrate governance thresholds with affected communities\. The central test is whether benefits persist beyond the dialogue and through transition or termination\. CSED provides an agenda for emotional dialogue that understands, supports, and sustains user capability\.
## References
- A\. Aldao, G\. Sheppes, and J\. J\. Gross \(2015\)Emotion regulation flexibility\.Cognitive Therapy and Research39,pp\. 263–278\.External Links:[Document](https://dx.doi.org/10.1007/s10608-014-9662-4)Cited by:[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.15)\.
- E\. C\. Baek, R\. Pourafshari, and J\. B\. Bayer \(2025\)The four conceptualizations of social connection\.Nature Reviews Psychology4\(8\),pp\. 506–517\.External Links:[Document](https://dx.doi.org/10.1038/s44159-025-00455-9)Cited by:[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p3.6)\.
- A\. Bandura \(2001\)Social cognitive theory: an agentic perspective\.Annual Review of Psychology52,pp\. 1–26\.External Links:[Document](https://dx.doi.org/10.1146/annurev.psych.52.1.1)Cited by:[§A\.1](https://arxiv.org/html/2607.27851#A1.SS1.p3.1),[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p2.4)\.
- J\. Banks \(2024\)Deletion, departure, death: experiences of AI companion loss\.Journal of Social and Personal Relationships41\(11\),pp\. 3547–3572\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p2.1),[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p3.1),[Illustrative Multi\-Timescale Evaluation](https://arxiv.org/html/2607.27851#Sx3.SSx4.p5.10),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1)\.
- R\. F\. Baumeister and M\. R\. Leary \(1995\)The need to belong: desire for interpersonal attachments as a fundamental human motivation\.Psychological Bulletin117\(3\),pp\. 497–529\.External Links:[Document](https://dx.doi.org/10.1037/0033-2909.117.3.497)Cited by:[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1),[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.15)\.
- G\. A\. Bonanno and M\. Westphal \(2024\)The three axioms of resilience\.Journal of Traumatic Stress37,pp\. 717–723\.External Links:[Document](https://dx.doi.org/10.1002/jts.23071)Cited by:[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.15)\.
- E\. M\. Boucher, N\. R\. Harake, H\. E\. Ward, S\. E\. Stoeckl, J\. Vargas, J\. Minkel, A\. C\. Parks, and R\. Zilca \(2021\)Artificially intelligent chatbots in digital mental health interventions: a review\.Expert Review of Medical Devices18\(sup1\),pp\. 37–49\.Cited by:[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1)\.
- R\. A\. Calvo, D\. Peters, K\. Vold, and R\. M\. Ryan \(2020\)Supporting human autonomy in AI systems: a framework for ethical enquiry\.InEthics of Digital Well\-Being,pp\. 31–54\.External Links:[Document](https://dx.doi.org/10.1007/978-3-030-50585-1%5F2)Cited by:[§A\.1](https://arxiv.org/html/2607.27851#A1.SS1.p3.1),[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p2.4)\.
- C\. S\. Carver, M\. F\. Scheier, and J\. K\. Weintraub \(1989\)Assessing coping strategies: a theoretically based approach\.Journal of Personality and Social Psychology56\(2\),pp\. 267–283\.External Links:[Document](https://dx.doi.org/10.1037/0022-3514.56.2.267)Cited by:[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.15)\.
- J\. Cheng, S\. Sabour, H\. Sun, Z\. Chen, and M\. Huang \(2023\)PAL: persona\-augmented emotional support conversation generation\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 535–554\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Emotional support conversation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px2.p1.1)\.
- Y\. Cheng, W\. Liu, W\. Li, J\. Wang, R\. Zhao, B\. Liu, X\. Liang, and Y\. Zheng \(2022\)Improving multi\-turn emotional support dialogue generation with lookahead strategy planning\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 3014–3026\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Emotional support conversation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px2.p1.1)\.
- M\. D\. Chu, P\. Gerard, K\. Pawar, C\. Bickham, and K\. Lerman \(2025\)Illusions of intimacy: emotional attachment and emerging psychological risks in human\-AI relationships\.arXiv preprint arXiv:2505\.11649\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1),[Paradigm Definition, Scope, and Unit of Inquiry](https://arxiv.org/html/2607.27851#Sx3.SSx1.p1.1),[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p3.7),[Testable Implications of CSED](https://arxiv.org/html/2607.27851#Sx4.SSx4.p1.9)\.
- C\. E\. Cutrona \(1990\)Stress and social support: in search of optimal matching\.Journal of Social and Clinical Psychology9\(1\),pp\. 3–14\.External Links:[Document](https://dx.doi.org/10.1521/jscp.1990.9.1.3)Cited by:[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1)\.
- D\. P\. Folk and E\. W\. Dunn \(2026\)How does turning to AI for companionship predict loneliness and vice versa?\.Psychological Science37\(4\),pp\. 276–286\.External Links:[Document](https://dx.doi.org/10.1177/09567976261427747)Cited by:[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p3.7),[Boundary Conditions and Limitations](https://arxiv.org/html/2607.27851#Sx5.p1.1)\.
- T\. Han, B\. Xu, H\. Zhang, and Y\. Lu \(2026\)Auditing stealth sycophancy in mental\-health dialogue: structured clinical\-state diagnostics and clean matched benchmarks\.arXiv preprint arXiv:2605\.03472\.Cited by:[Illustrative Multi\-Timescale Evaluation](https://arxiv.org/html/2607.27851#Sx3.SSx4.p2.3),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1)\.
- K\. Hopman, D\. Richards, and M\. M\. Norberg \(2023\)A digital coach to promote emotion regulation skills\.Multimodal Technologies and Interaction7\(6\),pp\. 57\.Cited by:[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1),[Testable Implications of CSED](https://arxiv.org/html/2607.27851#Sx4.SSx4.p1.9),[Assumption 1](https://arxiv.org/html/2607.27851#Thmassumption1.p1.3)\.
- E\. Horvitz \(1999\)Principles of mixed\-initiative user interfaces\.InProceedings of the SIGCHI Conference on Human Factors in Computing Systems,pp\. 159–166\.External Links:[Document](https://dx.doi.org/10.1145/302979.303030)Cited by:[§A\.1](https://arxiv.org/html/2607.27851#A1.SS1.p3.1),[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p2.4)\.
- B\. M\. Iacoviello and D\. S\. Charney \(2014\)Psychosocial facets of resilience: implications for preventing posttrauma psychopathology, treating trauma survivors, and enhancing community resilience\.European Journal of Psychotraumatology5\(1\),pp\. 23970\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p2.1),[Year\-stratified scoping audit\.](https://arxiv.org/html/2607.27851#Sx2.SSx2.SSS0.Px1.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1),[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.15)\.
- L\. Iyer, K\. Aggarwal, S\. Koyejo, G\. D\. Heyman, D\. C\. Ong, and S\. Mukherjee \(2026\)HEART: a unified benchmark for assessing humans and LLMs in emotional support dialogue\.arXiv preprint arXiv:2601\.19922\.Cited by:[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p1.1),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1)\.
- T\. E\. Joiner, G\. I\. Metalsky, J\. Katz, and S\. R\. H\. Beach \(1999\)Depression and excessive reassurance\-seeking\.Psychological Inquiry10\(3\),pp\. 269–278\.External Links:[Document](https://dx.doi.org/10.1207/S15327965PLI1004%5F1)Cited by:[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1)\.
- R\. Kalisch, A\. O\. J\. Cramer, H\. Binder, J\. Fritz, I\. Leertouwer, G\. Lunansky, B\. Meyer, J\. Timmer, I\. M\. Veer, and A\. van Harmelen \(2019\)Deconstructing and reconstructing resilience: a dynamic network approach\.Perspectives on Psychological Science14\(5\),pp\. 765–777\.Cited by:[Year\-stratified scoping audit\.](https://arxiv.org/html/2607.27851#Sx2.SSx2.SSS0.Px1.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1),[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.15),[Illustrative Multi\-Timescale Evaluation](https://arxiv.org/html/2607.27851#Sx3.SSx4.p4.1),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1)\.
- T\. Kannampallil, O\. A\. Ajilore, N\. Lv, J\. M\. Smyth, N\. E\. Wittels, C\. R\. Ronneberg, V\. Kumar, L\. Xiao, S\. Dosala, A\. Barve, A\. Zhang, K\. C\. Tan, K\. Cao, C\. R\. Patel, B\. S\. Gerber, J\. A\. Johnson, E\. A\. Kringle, and J\. Ma \(2023\)Effects of a virtual voice\-based coach delivering problem\-solving treatment on emotional distress and brain function: a pilot RCT in depression and anxiety\.Translational Psychiatry13,pp\. 166\.Cited by:[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1),[Assumption 1](https://arxiv.org/html/2607.27851#Thmassumption1.p1.3)\.
- G\. Kim, Y\. Choi, Y\. Kim, and C\. Lee \(2026\)No time to say goodbye: emotional loss responses to sudden termination in immersive AI interactions\.InExtended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems,Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p2.1),[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p3.1),[Illustrative Multi\-Timescale Evaluation](https://arxiv.org/html/2607.27851#Sx3.SSx4.p5.10),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1),[Testable Implications of CSED](https://arxiv.org/html/2607.27851#Sx4.SSx4.p1.9)\.
- H\. R\. Kirk, H\. A\. Davidson, E\. R\. Saunders, L\. Luettgau, B\. Vidgen, S\. A\. Hale, and C\. Summerfield \(2025\)Neural steering vectors reveal dose and exposure\-dependent impacts of human\-AI relationships\.arXiv preprint arXiv:2512\.01991\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2512.01991)Cited by:[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.17)\.
- H\. Lalwani and H\. Salam \(2026\)The supportiveness–safety tradeoff in LLM well\-being agents\.InCompanion Proceedings of the 21st ACM/IEEE International Conference on Human\-Robot Interaction,Cited by:[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1),[Testable Implications of CSED](https://arxiv.org/html/2607.27851#Sx4.SSx4.p1.9),[Remark 2](https://arxiv.org/html/2607.27851#Thmremark2.p1.2)\.
- R\. M\. Lee and S\. B\. Robbins \(1995\)Measuring belongingness: the social connectedness and the social assurance scales\.Journal of Counseling Psychology42\(2\),pp\. 232–241\.External Links:[Document](https://dx.doi.org/10.1037/0022-0167.42.2.232)Cited by:[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.15),[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p3.6)\.
- S\. Liu, C\. Zheng, O\. Demasi, S\. Sabour, Y\. Li, Z\. Yu, Y\. Jiang, and M\. Huang \(2021\)Towards emotional support dialog systems\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing,pp\. 3469–3483\.Cited by:[§C\.1](https://arxiv.org/html/2607.27851#A3.SS1.p1.1),[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Emotional support conversation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px2.p1.1),[ESConv gold turns\.](https://arxiv.org/html/2607.27851#Sx2.SSx2.SSS0.Px2.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p1.1),[Illustrative Multi\-Timescale Evaluation](https://arxiv.org/html/2607.27851#Sx3.SSx4.p3.3),[Remark 1](https://arxiv.org/html/2607.27851#Thmremark1.p1.7)\.
- Y\. Ma, K\. L\. Nguyen, F\. Z\. Xing, and E\. Cambria \(2020\)A survey on empathetic dialogue systems\.Information Fusion64,pp\. 50–70\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Empathetic dialogue\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px1.p1.1),[Remark 1](https://arxiv.org/html/2607.27851#Thmremark1.p1.7)\.
- M\. Namvarpour, B\. Brofsky, J\. Y\. Medina, M\. Akter, and A\. Razi \(2026\)Understanding teen overreliance on AI companion chatbots through self\-reported Reddit narratives\.InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems,Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1),[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p3.7),[Testable Implications of CSED](https://arxiv.org/html/2607.27851#Sx4.SSx4.p1.9)\.
- OpenAI \(2026\)Retiring GPT\-4o and older models\.Note:https://openai\.com/index/retiring\-gpt\-4o\-and\-older\-models/Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p2.1),[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p3.1),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1)\.
- R\. Poonsiriwong, C\. Archiwaranguprok, and P\. Pataranutaporn \(2026\)“Death” of a chatbot: investigating and designing toward psychologically safe endings for human\-AI relationships\.arXiv preprint arXiv:2602\.07193\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p2.1),[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p3.1),[Illustrative Multi\-Timescale Evaluation](https://arxiv.org/html/2607.27851#Sx3.SSx4.p5.10),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1),[Testable Implications of CSED](https://arxiv.org/html/2607.27851#Sx4.SSx4.p1.9)\.
- H\. Rashkin, E\. M\. Smith, M\. Li, and Y\. Boureau \(2019\)Towards empathetic open\-domain conversation models: a new benchmark and dataset\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,pp\. 5370–5381\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Empathetic dialogue\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px1.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p1.1),[Remark 1](https://arxiv.org/html/2607.27851#Thmremark1.p1.7)\.
- R\. M\. Ryan and E\. L\. Deci \(2006\)Self\-regulation and the problem of human autonomy: does psychology need choice, self\-determination, and will?\.Journal of Personality74\(6\),pp\. 1557–1586\.External Links:[Document](https://dx.doi.org/10.1111/j.1467-6494.2006.00420.x)Cited by:[§A\.1](https://arxiv.org/html/2607.27851#A1.SS1.p3.1),[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.15),[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p2.4)\.
- M\. Sharma, M\. McCain, R\. Douglas, and D\. Duvenaud \(2026\)Who’s in charge? disempowerment patterns in real\-world LLM usage\.arXiv preprint arXiv:2601\.19062\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Illustrative Multi\-Timescale Evaluation](https://arxiv.org/html/2607.27851#Sx3.SSx4.p2.3),[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p2.8),[Testable Implications of CSED](https://arxiv.org/html/2607.27851#Sx4.SSx4.p1.9),[Remark 2](https://arxiv.org/html/2607.27851#Thmremark2.p1.2)\.
- S\. Son, S\. Koo, E\. H\. Zi, J\. Jang, and H\. Lim \(2026\)Evaluating over\-empathizing in emotional support conversations: a user\-centered framework\.Expert Systems with Applications\.Cited by:[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1)\.
- A\. C\. Tricco, E\. Lillie, W\. Zarin, K\. K\. O’Brien, H\. Colquhoun, D\. Levac, D\. Moher, M\. D\. J\. Peters, T\. Horsley, L\. Weeks, S\. Hempel, and et al\. \(2018\)PRISMA extension for scoping reviews \(PRISMA\-ScR\): checklist and explanation\.Annals of Internal Medicine169\(7\),pp\. 467–473\.Cited by:[§B\.1](https://arxiv.org/html/2607.27851#A2.SS1.p1.1),[Year\-stratified scoping audit\.](https://arxiv.org/html/2607.27851#Sx2.SSx2.SSS0.Px1.p1.1)\.
- A\. S\. Troy, E\. C\. Willroth, A\. J\. Shallcross, N\. R\. Giuliani, J\. J\. Gross, and I\. B\. Mauss \(2023\)Psychological resilience: an affect\-regulation framework\.Annual Review of Psychology74,pp\. 547–576\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p2.1),[Year\-stratified scoping audit\.](https://arxiv.org/html/2607.27851#Sx2.SSx2.SSS0.Px1.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1),[An Illustrative Process Model](https://arxiv.org/html/2607.27851#Sx3.SSx3.p1.15),[Illustrative Multi\-Timescale Evaluation](https://arxiv.org/html/2607.27851#Sx3.SSx4.p4.1),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1),[Testable Implications of CSED](https://arxiv.org/html/2607.27851#Sx4.SSx4.p1.9)\.
- S\. C\. Vella and N\. B\. Pai \(2019\)A theoretical review of psychological resilience: defining resilience and resilience research over the decades\.Archives of Medicine and Health Sciences7\(2\),pp\. 233–239\.Cited by:[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1)\.
- L\. M\. Weinstock and M\. A\. Whisman \(2007\)Rumination and excessive reassurance\-seeking in depression: a cognitive\-interpersonal integration\.Cognitive Therapy and Research31\(3\),pp\. 333–342\.External Links:[Document](https://dx.doi.org/10.1007/s10608-006-9004-2)Cited by:[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p2.1)\.
- A\. Welivita, C\. Yeh, and P\. Pu \(2023\)Empathetic response generation for distress support\.InProceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue \(SIGDIAL\),pp\. 632–644\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Empathetic dialogue\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px1.p1.1)\.
- J\. Yang, Y\. Li, G\. Chen, R\. Fan, X\. Bai, and T\. He \(2026\)When seekers are hard to help: evaluating emotional support dialogue systems in worst\-case interactions\.arXiv preprint arXiv:2605\.28228\.Cited by:[Capability and lifecycle orientation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px3.p1.1),[ESConv contains an uneven capability\-relevant repertoire\.](https://arxiv.org/html/2607.27851#Sx2.SSx3.SSS0.Px2.p1.1),[Capability Dynamics and Six Design Commitments](https://arxiv.org/html/2607.27851#Sx3.SSx2.p1.1),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1)\.
- M\. Zao\-Sanders, K\. Hill, New, J\. D\. Freitas, I\. Cohen, D\. Adam, M\. Williams, M\. Carroll, and G\. Shteynberg \(2025\)Emotional risks of ai companions demand attention\.Nature Machine Intelligence\.Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Paradigm Definition, Scope, and Unit of Inquiry](https://arxiv.org/html/2607.27851#Sx3.SSx1.p1.1),[Multi\-Timescale Evaluation and Lifecycle Governance](https://arxiv.org/html/2607.27851#Sx4.SSx3.p1.1)\.
- Y\. Zhang, D\. Zhao, J\. T\. Hancock, R\. E\. Kraut, and D\. Yang \(2025\)The rise of AI companions: interaction with AI companions and psychological well\-being\.arXiv preprint arXiv:2506\.12605\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2506.12605)Cited by:[Introduction](https://arxiv.org/html/2607.27851#Sx1.p1.1),[Illustrative Lifecycle Constraints](https://arxiv.org/html/2607.27851#Sx3.SSx5.p3.7)\.
- J\. Zhu, H\. Dou, J\. Li, L\. Guo, F\. Chen, J\. Su, C\. Zhang, and F\. Kong \(2026\)Modeling multiple support strategies within a single turn for emotional support conversations\.arXiv preprint arXiv:2604\.17972\.Cited by:[Emotional support conversation\.](https://arxiv.org/html/2607.27851#Sx2.SSx1.SSS0.Px2.p1.1),[ESConv contains an uneven capability\-relevant repertoire\.](https://arxiv.org/html/2607.27851#Sx2.SSx3.SSS0.Px2.p1.1)\.
## Appendix ATheoretical Status and Scope
Capability\-sustaining emotional dialogue \(CSED\) is a theoretical research paradigm\. It specifies what emotional dialogue research studies, what counts as success, how system behavior enters a model of user change, and which observations are required to evaluate that change\. Its theoretical objects are the longitudinal interaction record, the user’s transient and capability states, four evaluation horizons, six design commitments, and lifecycle constraints on reliance, autonomy, and social connectedness\. The targeted audit provides motivating evidence for studying these objects\. The longitudinal hypotheses remain empirical claims for future studies\.
The main text defines a research paradigm as a coupled specification of the unit of inquiry, success criterion, model of user change, design commitments, and evaluation horizon\. This appendix adds lifecycle boundary conditions as an explicit sixth component\. Table[2](https://arxiv.org/html/2607.27851#A1.T2)applies this specification consistently to empathetic dialogue, emotional support conversation \(ESC\), and CSED\. The first two paradigms can operate at multiple interaction scopes\. Their characteristic strategies and success criteria remain centered on emotional understanding and current support\. CSED organizes both strategy and evaluation around the capability users retain across repeated use, non\-use, transition, and termination\.
Table 2:The three paradigms described through one theoretical specification\. Temporal scope can vary within the first two paradigms\. CSED makes sustained capability and the full interaction lifecycle constitutive parts of the research object\.The paper advances three kinds of claims\. Definitional claims specify CSED and its constructs\. Structural claims explain how common benchmark objectives relate to CSED and why short\-horizon observations do not identify longitudinal capability outcomes\. Empirical hypotheses state expected differences among future policies and lifecycle interventions\. Definitions and structural propositions can be evaluated for coherence and derivation\. The hypotheses require longitudinal data and causal study designs\.
### A\.1Formal Primitives and Restrictions
Table[3](https://arxiv.org/html/2607.27851#A1.T3)consolidates the formal objects introduced in the main paper\. A longitudinal record𝒞\\mathcal\{C\}contains the user, policy, sessions, and an optional termination event\. The latent statesks\_\{k\}separates transient emotioneke\_\{k\}from the capability vectorckc\_\{k\}\. Dialogue is one input to the transition kernelΦ\\Phi\. Stressors, offline actions, and human relationships also affect the transition\. The measurement mapMMconnects the latent state to validated instruments and behavioral markers\.
Table 3:Notation for the illustrative CSED process model\.###### Assumption 2\(Measurability and temporal alignment\)\.
For each capability component, validated instruments or behavioral markers provide noisy observations at time points that can be aligned with dialogue exposure, stressor exposure, non\-use, and offline behavior\. The measurement process records sufficient timing information to distinguish within\-session change from change that persists beyond system use\.
The formulation has five restrictions\. It targets sustained\-use settings rather than isolated exchanges\. Capability is latent and requires construct\-valid proxies\. Dialogue exposure is not the only cause of change\. The thresholdsδ\\delta,α0\\alpha\_\{0\}, andσ0\\sigma\_\{0\}require contextual calibration with affected users and domain experts\. Causal claims require a longitudinal design that addresses time\-varying confounding, reciprocal effects, and selective attrition\. These restrictions determine which claims the framework can support\.
Autonomy follows self\-determination theory and agentic accounts of collaborative control\(Ryan and Deci[2006](https://arxiv.org/html/2607.27851#bib.bib32); Calvoet al\.[2020](https://arxiv.org/html/2607.27851#bib.bib37); Bandura[2001](https://arxiv.org/html/2607.27851#bib.bib35); Horvitz[1999](https://arxiv.org/html/2607.27851#bib.bib36)\)\. A user can autonomously adopt system guidance when the uptake is voluntary, the user’s endorsed goals direct the collaboration, and the user can appraise, revise, or reject the output\. Reliance and autonomy are therefore separate constructs\. The proxyAut=𝔼i\[ViDiriEvali\]\\mathrm\{Aut\}=\\mathbb\{E\}\_\{i\}\[V\_\{i\}\\mathrm\{Dir\}\_\{i\}\\mathrm\{Eval\}\_\{i\}\]measures retained volition, direction, and evaluation rather than whether the user decided alone\.
### A\.2Structural Propositions
Let the combined CSED objective be
JCSED\(π\)=∑ℓ∈\{resp,conv,long,term\}wℓJℓ\(π\),J\_\{\\mathrm\{CSED\}\}\(\\pi\)=\\sum\_\{\\ell\\in\\\{\\mathrm\{resp\},\\mathrm\{conv\},\\mathrm\{long\},\\mathrm\{term\}\\\}\}w\_\{\\ell\}J\_\{\\ell\}\(\\pi\),\(15\)subject toDep≤δ\\mathrm\{Dep\}\\leq\\delta,Aut≥α0\\mathrm\{Aut\}\\geq\\alpha\_\{0\}, andSoc≥σ0\\mathrm\{Soc\}\\geq\\sigma\_\{0\}\. The weights are nonnegative\. This objective is an operational instantiation of the paradigm rather than its only possible realization\.
###### Proposition 1\(Benchmark nesting\)\.
Response\-scored empathetic dialogue and session\-scored ESC objectives are restricted cases of Equation \([15](https://arxiv.org/html/2607.27851#A1.E15)\)\.
###### Proof\.
SetK=1K=1andn1=1n\_\{1\}=1\. Choosewresp=1w\_\{\\mathrm\{resp\}\}=1and set all other horizon weights to zero\. Remove the lifecycle constraints and restrict the response feature mapρ\\rhoto the empathy dimensions used by a response\-scored benchmark\. Equation \([15](https://arxiv.org/html/2607.27851#A1.E15)\) then reduces to its expected empathy score\. For a session\-scored ESC objective, retainK=1K=1, allown1≥1n\_\{1\}\\geq 1, setwlong=wterm=0w\_\{\\mathrm\{long\}\}=w\_\{\\mathrm\{term\}\}=0, and retain response and conversation terms\. Choosingψ\\psias negative emotional intensity recovers within\-conversation emotional change\. Both objectives follow by parameter restriction and removal of longitudinal constraints\. ∎
The proposition places common benchmark objectives inside one larger evaluation program\. It preserves the strategies and evidence already used by empathetic dialogue and ESC\. It also identifies the additional observations required when a system remains available across repeated use\.
###### Proposition 2\(Short\-horizon non\-identifiability\)\.
Equality of response and conversation scores does not imply equality of longitudinal capability outcomes\.
###### Proof\.
Consider policiesπa\\pi\_\{a\}andπb\\pi\_\{b\}that induce the same distribution over observed histories, responses, and within\-session proxy changes\. They therefore have equalJrespJ\_\{\\mathrm\{resp\}\}andJconvJ\_\{\\mathrm\{conv\}\}\. Let their unobserved capability transitions differ after the session\. For a nonzero vectoraa, defineck\+1=ck\+ac\_\{k\+1\}=c\_\{k\}\+aunderπa\\pi\_\{a\}andck\+1=ck−ac\_\{k\+1\}=c\_\{k\}\-aunderπb\\pi\_\{b\}, while holding the short\-horizon observables fixed\. For any capability indexggthat increases in the direction ofaa, the expected longitudinal indices differ\. The short\-horizon score distribution is compatible with both transitions, so it cannot identify which longitudinal outcome occurred\. ∎
This is a structural result about the evidence available to an evaluator\. It does not assert that a particular deployed policy produces either transition\. It shows why response preference and conversation relief need observations of capability, reliance, autonomy, and connectedness across later time points\. The four hypotheses in the main paper specify empirical comparisons that can estimate these differences\.
## Appendix BTargeted Literature Audit
### B\.1Search, Screening, and Sampling
The literature arm is a PRISMA\-ScR\-guided targeted scoping audit\(Triccoet al\.[2018](https://arxiv.org/html/2607.27851#bib.bib28)\)\. Records were retrieved from the arXiv API on July 8, 2026\. The search used seven exact phrases\. They were “emotional support conversation,” “emotional support dialog,” “empathetic dialogue,” “empathetic response generation,” “mental health chatbot,” “AI companion,” and “companion chatbot\.” The script retrieved the first 200 records for each query in reverse submission\-date order, deduplicated records by version\-free arXiv identifier, and retained records dated from 2019 through the retrieval date in 2026\.
The search produced 275 unique records\. Title and abstract screening retained 228 and excluded 47\. Four exclusions concerned emotion recognition without dialogue generation, four concerned speech, synthesis, or embodiment without the dialogue\-system focus, and 39 concerned an out\-of\-scope application\. The inclusion set covers emotional or empathetic dialogue, emotional\-support dialogue, mental\-health chatbots, and AI companions, together with datasets, benchmarks, user studies, reviews, and position papers about these systems\.
The pilot targeted a proportional year\-stratified sample of 90 included records\. Integer rounding within the annual strata produced 91 records\. The sampling seed was 20260708\. Coding used titles and abstracts\. Sixty coded records build or evaluate systems\. The remaining records include user studies, reviews, datasets, evaluations, and position papers that characterize risks or the research landscape\. Claims about system objectives use the 60\-record subset\. Claims about evaluation horizon use all 91 records\.
### B\.2Paper\-Level Codebook
The paper\-level codebook uses six dimensions\. D1 records the primary objective as relief, capability, both, or neutral\. D2 records named mechanisms, including validation, exploration, reappraisal, problem solving, self\-efficacy, and social connection\. D3 records the highest evaluation outcome as interaction quality, proximal state change, capability outcome, or none\. D4 records the longest evaluation horizon as one turn, one session, multiple sessions, longitudinal, or none\. D5 records dependency, autonomy, sycophancy, crisis safety, termination loss, or no named risk\. D6 records the artifact type\.
Mechanism definitions were anchored in psychology rather than derived from CSED\. Reappraisal follows emotion\-regulation research\. Problem solving refers to concrete plans and action decomposition\. Self\-efficacy requires support for users’ perceived and exercised competence\. Social connection requires behavior that supports contact with other people or services\. This independent grounding reduces circularity between the proposed paradigm and the audit categories\.
Table 4:Primary audit counts used in the main paper\. Mechanisms are multi\-label\.Table[4](https://arxiv.org/html/2607.27851#A2.T4)reports the evidence behind the strategic and longitudinal gap\. Fifty\-seven of 60 system\-building papers have relief as their primary objective, giving 95\.0%\. Two combine relief and capability\. One targets the capability of peer counselors\. No coded record evaluates a user\-capability outcome or uses a longitudinal evaluation horizon\. Risk coding across all 91 records finds 15 references to dependency, 14 to crisis safety, two to termination loss, one to sycophancy, and one to autonomy\. Sixty\-six records name none of these risks\. Only one of the 60 system\-building papers names a risk from this set\.
The evaluation and mechanism codes also show a measurement concentration\. Of the 59 records that name validation or comfort, 54 use interaction\-quality outcomes, three use proximal state outcomes, and two contain no evaluation\. Of the three records that name problem solving, two use interaction quality and one uses a proximal outcome\. Both self\-efficacy records use interaction\-quality evaluation\. Reappraisal and social connection have no mechanism\-by\-outcome cells because neither is named in the system\-building subset\.
The audit is descriptive\. The year\-stratified sample, arXiv\-only source, title and abstract coding, and AI\-only coding limit population inference\. The reported proportions characterize the coded sample\. The released protocol defines full\-text, multi\-database, and human\-recoding extensions\.
## Appendix CESConv Corpus Analysis
### C\.1Sampling and Function Codebook
ESConv contains 1,300 dialogues and 18,376 supporter utterances with strategy annotations\(Liuet al\.[2021](https://arxiv.org/html/2607.27851#bib.bib4)\)\. The full\-corpus strategy counts include 3,801 questions, 3,341 other turns, 2,954 suggestions, 2,827 affirmation and reassurance turns, 1,713 self\-disclosures, 1,436 reflections of feeling, 1,215 information turns, and 1,089 restatements or paraphrases\. Affirmation and reassurance therefore accounts for 15\.4% of annotated supporter turns\. Suggestions account for 16\.1%\.
Function coding used a stratified random sample of 300 supporter utterances\. The early, middle, and late dialogue\-position terciles each contribute 100 turns\. The sampling seed was 20260708\. Each item includes the supporter turn and up to two preceding seeker utterances\. The primary code is the dominant communicative function of the main clause\.
The ten functions are F1 validation and comfort, F2 exploration, F3 reappraisal, F4 problem solving, F5 self\-efficacy, F6 social connection, F7 boundary and safety, F8 self\-disclosure, F9 information, and F10 other behavior\. The analysis maps F1 to relief\. It maps F3 through F7 to capability\-relevant behavior\. It maps F2 and F8 through F10 to interaction process\. The mapping is applied after function coding\.
Mixed\-function decisions follow four rules\. An empathic opener does not override a capability\-relevant main clause\. A question that embeds advice is coded as problem solving\. Strength\-based reassurance is coded as self\-efficacy\. A directive about a value\-laden personal decision is not treated as capability\-supporting problem solving unless it preserves user direction\. These rules separate supportive tone from the function of the turn\.
Table 5:Function distribution in the 300\-turn ESConv sample\. Percentages use 300 as the denominator\.Table[5](https://arxiv.org/html/2607.27851#A3.T5)contains the exact counts behind the main\-paper percentages\. Capability\-relevant functions total 129 of 300 turns, giving 43\.0%\. Relief accounts for 46 turns, giving 15\.3%\. Process functions account for 125 turns, giving 41\.7%\. Capability relevance appears in 26 early turns, 55 middle turns, and 48 late turns\. Relief appears in 17, 15, and 14 turns\. Process functions appear in 57, 30, and 38 turns\. These stage distributions describe the sample and do not estimate longitudinal user outcomes\.
### C\.2Reliability Design
Reliability was assessed on one random subset of 40 sampled turns\. Four label sets were compared\. They consist of the primary AI coding, a blind repeat by the same model, a GPT\-5\.5 coding, and a DeepSeek\-V4\-Pro coding\. Each coder received the same function definitions and local seeker context without access to the other labels\. Compatible endpoints used temperature zero\. Reasoning endpoints that did not accept that parameter used their provider defaults\. Frozen labels are included because hosted model behavior and undisclosed serving infrastructure can change\.
For codersaaandbb, Cohen’s kappa is
κ\(a,b\)=po−pe1−pe,\\kappa\(a,b\)=\\frac\{p\_\{o\}\-p\_\{e\}\}\{1\-p\_\{e\}\},\(16\)wherepop\_\{o\}is observed agreement andpep\_\{e\}is agreement expected from the two marginal label distributions\. Fine\-grained kappa uses F1 through F10\. Paradigm\-level kappa maps the same labels to relief, capability, and process before computing agreement\.
Table 6:All six pairwise reliability comparisons on the shared 40\-item subset\. SubscriptFFdenotes ten functions\. SubscriptPPdenotes the three analysis categories\.Fine\-grained kappa ranges from 0\.547 to 0\.751, with a mean of 0\.631\. Paradigm\-level kappa ranges from 0\.646 to 0\.845, with a mean of 0\.716\. Across the six pairs, there are 57 fine\-grained disagreements\. Thirty\-one remain within one analysis category and 26 cross the relief, capability, or process boundary\. Common within\-category boundaries include validation versus self\-efficacy and exploration versus self\-disclosure\. The result supports the coarse strategic contrast more strongly than individual function distinctions\.
The reliability evidence concerns consistency among AI coders\. It does not replace human construct validation\. A full study should add independent human coding, preregistered adjudication, and a larger reliability sample\.
## Appendix DClaim and Artifact Traceability
### D\.1Claim\-Evidence Map
Table[7](https://arxiv.org/html/2607.27851#A4.T7)maps every headline quantity to a frozen result artifact\. The artifact manifest supplies the exact relative path and checksum for each short source label\.
Table 7:Traceability from headline claims to calculations and supplied artifacts\.
### D\.2Model\-Behavior Probe
The model\-behavior arm is a released protocol rather than a completed result\. Its frozen input contains 100 held\-out ESConv contexts sampled with seed 20260708\. The generation pipeline presents the same dialogue context to at least two chat models, stores the response and model identifier, and applies the F1 through F10 codebook\. Planned comparisons report the gold and generated function distributions, stage\-conditional differences, and reliability of the generated\-response coding\. Model versions, access date, endpoint, decoding parameters, and failures must be recorded before the probe supports any behavioral claim\.
No deployed\-model result from this arm is included in the audit percentages\. This separation preserves the distinction among completed corpus analysis, the theoretical framework, and future tests of model behavior\.Similar Articles
EmoTrace: An Emotion Trajectory-Centered Framework for Psychological Support Dialogue Generation
The paper proposes EmoTrace, a multi-turn dialogue generation framework for psychological support that models seekers' emotional trajectories to improve empathy and emotional richness in counselor responses, outperforming existing methods.
DMT-CBT: Longitudinal Therapeutic State Modeling for CBT Counseling
DMT-CBT proposes a framework for longitudinal therapeutic state modeling in CBT counseling, addressing the need for multi-session, multimodal inference and intervention. It introduces DMTCorpus, a synthetic multi-session dataset, and shows improvements in counseling fidelity and therapeutic alliance over existing methods.
Modeling Multiple Support Strategies within a Single Turn for Emotional Support Conversations
This paper proposes multi-strategy utterance generation methods for Emotional Support Conversations (ESC), where each utterance can contain multiple strategy-response pairs. Two generation approaches (All-in-One and One-by-One) enhanced with cognitive reasoning via reinforcement learning are evaluated on the ESConv dataset, demonstrating improved supportive quality and dialogue success.
Dynamic Commonsense Coordination for Empathetic Response Generation
Proposes DCC, a dynamic commonsense coordination framework for empathetic response generation that integrates residual-based interaction, association-guided filtering, and iterative decoding, achieving improved emotion classification and response diversity over baselines.
When Clients Stop Following: A Cognitive Conceptualization Diagram-driven Framework for Strategic Counseling
Researchers from Lanzhou University propose a CBT-grounded framework for LLM-based psychological counseling that addresses the 'counselor-following' phenomenon in existing benchmarks, introducing CARS (a resistant client simulator), STREAMS (a dual-module strategic reasoning framework), and EWTS-MI (an entropy-weighted evaluation metric) to better handle real-world resistant counseling interactions.