Empath: 追踪危机咨询对话中的多层次情绪动态
摘要
介绍了Empath,一个用于分析危机咨询对话中情绪动态的框架,该框架应用于与黑人文本使用者的悲伤讨论,揭示了诸如持续性负面情感和渐进希望转变等模式。
arXiv:2609.29056v1 Announce Type: new
Abstract: Emotion dynamics are critical for understanding crisis-support conversations, yet most computational work treats emotion as static utterance-level labels. We introduce EMPATH, a framework for understanding affective dynamics in mental health dialogues across three granularities: turn-level labels, transition probabilities, and global conversation archetypes. Applying EMPATH to text-based crisis conversations with self-identified Black texters discussing grief, we find persistent negative affect, gradual hope-ward transitions, distinct texter-volunteer emotional roles, and heterogeneous recovery trajectories. These results highlight the informative patterns that emerge from computationally understanding crisis support and expressions of grief as dynamic processes within conversations, as well as the overall value of emotion-dynamic analysis for analyzing and comparing affect in dialogues.
查看缓存全文
缓存时间: 2026/09/25 09:16
# Empath: Tracing Multi-Level Emotion Dynamics in Crisis Counseling Dialogues
Source: [https://arxiv.org/html/2609.29056](https://arxiv.org/html/2609.29056)
Yuchen HuangWen LiangAffiliation:Columbia University, USANicholas DeasAffiliation:Columbia University, USAMelanie SubbiahAffiliation:Columbia University, USAKathleen McKeownAffiliation:Columbia University, USAJulia HirschbergEmail:[\{sara\.ziweigong, ndeas, kathy, julia\}@cs\.columbia\.edu∗Equal contributions\.](mailto:Equal%20contributions.)Affiliation:Columbia University, USAAffiliation:Barnard College, USA
###### Abstract
Emotion dynamics are critical for understanding crisis\-support conversations, yet most computational work treats emotion as static utterance\-level labels\. We introduceEmpath, a framework for understanding affective dynamics in mental health dialogues across three granularities: turn\-level labels, transition probabilities, and global conversation archetypes\. ApplyingEmpathto text\-based crisis conversations with self\-identified Black texters discussing grief, we find persistent negative affect, gradual hope\-ward transitions, distinct texter–volunteer emotional roles, and heterogeneous recovery trajectories\. These results highlight the informative patterns that emerge from computationally understanding crisis support and expressions of grief as dynamic processes within conversations, as well as the overall value of emotion\-dynamic analysis for analyzing and comparing affect in dialogues\.
## 1Introduction
Understanding how emotions change over the course of crisis\-support conversations is critical for studying online grief support\. In text\-based crisis settings, texters may move through persistent distress, disclosure, moments of gratitude, and gradual shifts toward hope over the course of a single interaction\. These changes are especially important in counseling conversations, where support is rarely a simple linear movement from negative to positive emotion\. However, much computational work on mental health dialogue treats emotion as a static utterance\-level label, following broader trends in emotion recognition and affective state identification[Jordan et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib2);[Wu et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib44);[Gong et al\. \(2023\)](https://arxiv.org/html/2609.29056#bib.bib43), leaving the temporal structure of emotional change underexplored\. Related work also examines the extraction of texters’ explicit emotion expressions in crisis conversations[Buda et al\. \(2026\)](https://arxiv.org/html/2609.29056#bib.bib42)\.
In this work, we study emotion dynamics in conversations from Crisis Text Line \(CTL\),111[https://www\.crisistextline\.org](https://www.crisistextline.org/)a mental health organization where volunteer counselors provide support over text messages to those in crisis\. As part of a larger collaboration between computational linguists and social work researchers, we focus specifically on CTL conversations with Black texters discussing or expressing grief\. Prior work has emphasized that grief in Black communities is shaped by social, historical, and structural contexts[Wilson and O’Connor \(2022\)](https://arxiv.org/html/2609.29056#bib.bib1), and emotional expressions of grief are known to be highly complex, sometimes involving simultaneous expressions of joy and deep distress[Patton et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib29)\. This setting is therefore a challenging domain for analyzing highly complex emotion expression, and provides an opportunity to examine how distress and support unfold in crisis conversations\. This setting also raises important methodological challenges: the data is highly sensitive, privacy\-restricted, and cannot be processed using API\-based systems or publicly released at scale\.
To analyze these conversations, we proposeEmpath, a framework for studying emotion dynamics in therapeutic and crisis\-support dialogue across three levels of granularity\. First,Empathidentifies utterance\-level emotions using a locally runnable emotion\-labeling pipeline designed for privacy\-restricted data\. Second, it analyzes micro\-dynamics, including turn\-level emotion transitions, polarity shifts, persistence, and recovery pivots\. Third, it characterizes macro\-dynamics, including volunteer strategies and conversation\-level trajectory archetypes\. Together, these levels allow us to move beyond aggregate emotion distributions and examine how emotional states persist, shift, and resolve over interactions\.
ApplyingEmpathto CTL grief conversations, we surface a range of different emotion dynamics patterns, including repeated negative emotion expressions, shifts from negative emotions to hope, and mixtures of these patterns\. Such emotion dynamics also highlight distinctions between the emotional roles of texters and volunteer counselors in dialogues\. Examining these patterns suggests that successful crisis support is better understood as a dynamic process rather than a terminal shift from distress to resolution\. In particular, the timing and structure of emotional transitions reveal aspects of support that are not visible from traditional utterance\-level labels or conversation endpoints alone\.
Finally, we use synthetic crisis\-support dialogues to show how the framework also applies beyond CTL\. Synthetic data has become increasingly attractive for developing analysis tools while preserving the confidentiality of real conversations[Kurakin et al\. \(2023\)](https://arxiv.org/html/2609.29056#bib.bib21);[Flemings and Annavaram \(2024\)](https://arxiv.org/html/2609.29056#bib.bib20);[Cabrera Lozoya et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib22), and recent work has explored LLM\-based patient simulations for clinical and counselor training[Louie et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib41);[Louie et al\. \(2026\)](https://arxiv.org/html/2609.29056#bib.bib3);[Wang et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib45)\. We show thatEmpathcan reveal meaningful differences in emotion dynamics, surfacing some similarities in aggregate emotional trends but gaps in fine\-grained patterns between synthetic and authentic dialogues\.
We summarize our contributions as follows:
1. 1\.We propose anovel, three\-level evaluation framework,Empath, to assess emotion dynamicsin mental health and crisis dialogues\.
2. 2\.UsingEmpath, we conduct alarge\-scale analysis of text\-based crisis conversations with Black texters about grieffrom Crisis Text Line\. We discuss how trends in emotion dynamics are reflective of successful dialogues in this context\. We identify key dynamic signatures of grief\-support conversations, including persistent distress, hope\-ward pivots, role\-differentiated emotional profiles, and heterogeneous recovery trajectories\.
## 2Related Work
Emotion Recognition in Conversation\(ERC\) is a foundational NLP task and key tool for tracking crisis intervention trajectories[Tripodi et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib40);[Xu et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib4)\. In text\-based therapeutic dialogues, model evaluations reveal clear tradeoffs: closed\-source models like GPT\-4 perform well in zero\-shot diagnostic settings but often struggle with nuanced, culturally sensitive emotional states[Wu et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib39);[Xie et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib5)\. Open\-weight LLaMA models, when lightly fine\-tuned, achieve competitive performance in assessing emotional safety[Badawi et al\. \(2026\)](https://arxiv.org/html/2609.29056#bib.bib38)\. Domain\-specific models, such as MentalBERT and MentalRoBERTa[Ji et al\. \(2022\)](https://arxiv.org/html/2609.29056#bib.bib14), consistently outperform general models on targeted clinical tasks, while crisis\-focused models like BERT\-EV leverage continuous valence scoring to track turn\-level de\-escalation in real\-world crisis conversations[Tripodi et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib40)\. Efforts on crisis de\-escalation and emotional support dialogue suggest that interactional trajectories, support strategies, and changes in distress over time are central to understanding support effectiveness[Tripodi et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib40);[Liu et al\. \(2021\)](https://arxiv.org/html/2609.29056#bib.bib37);[Wan et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib36);[Zhang et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib35);[Liu et al\. \(2026\)](https://arxiv.org/html/2609.29056#bib.bib34)\.
Grief and bereavementare highly complex experiences that are often misunderstood[Hall \(2014\)](https://arxiv.org/html/2609.29056#bib.bib24)and accompanied by constantly evolving theories including dual process models[Margaret Stroebe \(1999\)](https://arxiv.org/html/2609.29056#bib.bib25)and meaning making\-focused perspectives[Stroebe and Schut \(2001\)](https://arxiv.org/html/2609.29056#bib.bib26);[Neimeyer et al\. \(2002\)](https://arxiv.org/html/2609.29056#bib.bib27)\. Understanding these experiences has been further complicated by the movement of their expression to online, digital spaces[Moore et al\. \(2017\)](https://arxiv.org/html/2609.29056#bib.bib28);[Patton et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib29)\. As expressions of grief in digital counseling are highly dynamic and complex, we focus on this domain as a potential case where understanding may be aided through emotion dynamics\.
Figure 1:Summary ofEmpathframework for assessing conversation dynamics in mental health/crisis dialogues\. The example dialogue is simulated and contains no real CTL messages\.
## 3Crisis Dialogue Data
Crisis Text Line \(CTL\) is a non\-profit mental health organization that provides real\-time support for those in crisis over text messages\. Those in crisis \(Texters\) that reach out to CTL are connected with a volunteer crisis counselor \(Volunteers\) to receive support\. All crisis counselors are trained by CTL to provide effective support to texters and are supervised live by trained mental health professional staff\.
We use a corpus of 2,478 de\-identified conversations focused on expressions of grief in the Black community\. We collect only conversations with texters that self\-identified as Black or African American in an optional post\-conversation survey\. CTL volunteers may tag conversations with a topic, and we collect only conversations tagged as discussinggrief\. From this corpus, we collect a random sample of 100 de\-identified conversations for human annotation and validation of the emotion models\. Importantly, due to the sensitive nature of the data and in order to protect the confidentiality of CTL users, this data is de\-identified prior to access, is only accessed under a signed DUA, and all data is processed locally \(i\.e\., no CTL data are provided to API\-based LLMs or other online services\)\.
We collaborate with a team of social work researchers who conduct an inductive thematic analysis of the conversations–the researchers thoroughly read and label conversation turns according to themes that are qualitatively derived from the data\. All conversations within the 100\-conversation annotation sample are annotated by two researchers, and all disagreements between them are resolved through discussion\. While these themes cover a variety of categories, in this work, we focus on a set of 37 emotions \(e\.g\.,hopeful,numbness,longing\) identified by the social work researchers, as defined in Table[10](https://arxiv.org/html/2609.29056#A1.T10)\. Text\-level statistics summarizing the annotation sample are included in Table[1](https://arxiv.org/html/2609.29056#S3.T1)\. The 100 conversations span a period from December 2016 to October 2023, with an average of 56\.3 dialogue turns each\.
Table 1:Message summary statistics of the 100\-conversation CTL annotation sample\.
## 4EmpathFramework
Our primary methodology is theEmotion andMultidimensionalPatternAnalysis forTherapeuticHelp \(Empath\) framework\. We introduceEmpathas a tool for studying patterns in the emotion dynamics of crisis, mental health, and other therapeutic dialogues at multiple levels\. This three\-level evaluation framework \(summarized in Figure[1](https://arxiv.org/html/2609.29056#S2.F1)\) processes raw dialogue through an analysis ranging from individual utterances to holistic conversation:1\)Utterance\-Level Emotion Detection;2\)Micro\-Dynamics\(Transitions and Polarity\); andiii\)Macro\-Dynamics\(Strategies, Archetypes, and Diversity\)\.
### 4\.1Level 1: Emotion Detection
Table 2:Micro\-Dynamics Statistics derived from emotion\-label sequences at the category level of 37 emotion labels \(C\) and polarity level of positive/negative/neutral \(P\)\.We frame utterance\-level emotion detection as a constrained text generation task, similar to recent emotion analysis work[Deas et al\. \(2024b\)](https://arxiv.org/html/2609.29056#bib.bib23)\. Given an utterance and optional conversational context, a language model is prompted to select exactly one label from a list of the 37 emotion categories \(Table[10](https://arxiv.org/html/2609.29056#A1.T10)\) and provide reasoning\. These labels describe emotional content expressed or reflected in an utterance; volunteer labels may reflect texters’ emotions rather than the volunteers’ own emotional states\.
#### Prompt Design\.
Each prompt consists of three components:i\)the list of 37 emotion labels with brief descriptions,ii\)the conversational history \(a sliding window of up toNNpreceding utterances from both speakers\), andiii\)an instruction asking the model to output a structured\(label, reason\)response with post\-hoc explanations[Camburu et al\. \(2018\)](https://arxiv.org/html/2609.29056#bib.bib18);[Limpijankit et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib19)\. The instruction explicitly directs the model to weight the current message most heavily and to use preceding messages only as supporting context\. Prompts provided in §[A](https://arxiv.org/html/2609.29056#A1)\.
#### Model Selection and Validation\.
To safeguard the confidentiality of the CTL conversations, we exclusively evaluate locally runnable models within a secure, offline environment\. Specifically, we usemeta\-llama/Llama\-3\.2\-3B\-Instruct[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib11)as the backbone model, with experimental setups detailed in §[A](https://arxiv.org/html/2609.29056#A1.SS0.SSS0.Px2)\. We validate the pipeline by comparing generative models \(Llama\-3\.1/3\.2, Mistral\-7B\), encoder\-only models \(BERT Emotions, MentalBERT\), and a random baseline across the human\-annotated subset of the CTL conversations\. Performance is measured using exact\-match accuracy and semantic similarity\.
Our results demonstrate that generative models significantly outperform BERT\-based models across all taxonomy levels\. We selected Llama\-3\.2\-3B\-Instruct for all primary analyses as it offered the best trade\-off between performance and computational efficiency\. Notably, it achieved semantic similarity scores \(0\.97–0\.98\) comparable to larger models, indicating that its predictions are semantically aligned with human references even when exact labels differ\. Full candidate model descriptions, validation results, and the multi\-level taxonomy mapping are detailed in §[B](https://arxiv.org/html/2609.29056#A2)\.
We also test two configurations per corpus: \(i\)*without context*\(single\-utterance\), and \(ii\)*with context*\(includes preceding utterances\)\. We select \(i\)*with context*for our main analysis\. While context\-free models score slightly higher on single\-utterance accuracy, the with\-context mode produces the temporally coherent trajectories necessary for studying conversation dynamics \(see further discussion and full comparison in §[G](https://arxiv.org/html/2609.29056#A7)\)\.
### 4\.2Level 2: Micro\-Dynamics \(Transitions and Polarity\)
Once labels are assigned, this level characterizes the dynamics of emotional states across conversations at two complementary levels of granularity:*category\-level*, and*polarity\-level*\. Table[2](https://arxiv.org/html/2609.29056#S4.T2)summarizes the full set of derived statistics\.Category\-level analysistracks transitions among all 37 emotion labels\. Each conversation is modeled as a sequence of emotion labels from which we compute adjacent\-pair transitions\.Polarity\-level analysisprojects labels onto a polarity scale to capture coarse sentiment trajectories\. To complement the fine\-grained emotion category view, we project each of the 37 emotion labels onto a continuous valence axis using the NRC Valence–Arousal–Dominance \(VAD\) Lexicon[Mohammad \(2018\)](https://arxiv.org/html/2609.29056#bib.bib10)and discretize into three polarity classes:positive\(valence≥0\.55\\geq 0\.55\);negative\(valence≤0\.45\\leq 0\.45\); andneutral\(0\.45<valence<0\.550\.45<\\text\{valence\}<0\.55\)\. The full mapping is provided in §[C](https://arxiv.org/html/2609.29056#A3)\.
### 4\.3Level 3: Macro\-Dynamics \(Strategies and Archetypes\)
Moving beyond surface\-level labels, we analyze the interactional structure and global composition of conversations\. To support this, we derive a per\-turn distress scoreddfrom negated valence scorevv,d=1−v∈\[0,1\]d=1\-v\\in\[0,1\], so that negative\-valence labels \(e\.g\.,fear,sadness\) yield high distress values and positive\-valence labels \(e\.g\.,gratitude,joy\) yield low ones\.
#### Volunteer Strategy Analysis\.
Surface\-level emotional arcs may look similar across real and synthetic conversations while masking differences in*how*support is delivered\. To probe this, we examine whether the same volunteer strategy produces comparable downstream effects on texter distress in synthetic and real CTL data\. We analyze local three\-turn windows\(ut,ht,ut\+1\)\(u\_\{t\},h\_\{t\},u\_\{t\+1\}\), where a texter turnutu\_\{t\}is followed by a volunteer turnhth\_\{t\}and the next texter turnut\+1u\_\{t\+1\}\. Each volunteer turn is classified into one of eight ESConv support\-strategy categories\([Bai et al\., 2025](https://arxiv.org/html/2609.29056#bib.bib16)\):Affirmation and Reassurance,Information,Others,Providing Suggestions,Question,Reflection of Feelings,Restatement or Paraphrasing, andSelf\-disclosure\.
For each window, we compute the immediate downstream distress changeΔd=d\(ut\+1\)−d\(ut\)\\Delta d=d\(u\_\{t\+1\}\)\-d\(u\_\{t\}\), where more negative values indicate distress reduction, and define a binary de\-escalation indicator equal to 1 whenΔd<−0\.1\\Delta d<\-0\.1\. We estimate per\-strategy frequency, meanΔd\\Delta dwith 95% bootstrap confidence intervals, and de\-escalation rate, and aggregate to the conversation level for comparison across ablation conditions \(*model family*,*topic condition*,*length condition*,*label context*\)\.
#### Conversation Archetype Analysis\.
To test whether synthetic and real crisis conversations differ in the*composition*of trajectory shapes, we cluster individual conversations into interpretable archetypes\. We summarize each interaction as a length\-normalized distress trajectory by interpolating sequences onto a shared grid\. We then concatenate trajectories and apply pooledKK\-means clustering\. Centroids are assigned names based on shape features\. We name each centroid from its mean distress level, total fall, and when that fall occurs, yielding five archetypes:*Early Resolution*,*Late Recovery*,*Persistent Moderate Distress*,*Steady De\-escalation \(high distress\)*, and*Unresolved High Distress*\. Details on the trajectory normalization and feature\-based labeling of archetypes in §[E](https://arxiv.org/html/2609.29056#A5)\.
## 5Results: Emotion Dynamics in CTL Grief Conversations
The following analyses use the full CTL corpus of conversations with self\-identified Black or African American texters tagged as discussing grief\. The annotated subset is used to validate the emotion models\.
### 5\.1Level 1: Utterance\-Level Emotion Profiles
We focus primarily ontexterroles, as texter emotion dynamics are the main signal of interest in crisis\-support conversation analysis\. Texters and volunteers nevertheless exhibit complementary emotion profiles: texters show high negative\-polarity persistence \(84\.1%\), whereas volunteers more consistently maintain positive states, with positive\-polarity persistence of 72\.4%\. Volunteers’ most frequent cross\-emotion transition ishopeless→\\tohopeful, consistent with movement from acknowledging distress toward a more hopeful frame\. The full role\-differentiation analysis is provided in §[F\.1](https://arxiv.org/html/2609.29056#A6.SS1)\.
Table 3:Top\-10 emotion labels for CTL texter utterances \(left,N=60,788N\{=\}60\{,\}788labeled utterances from 2,478 conversations\) and volunteer utterances \(right,N=56,460N\{=\}56\{,\}460labeled utterances\)\. For texters, 60,788 of 62,723 raw utterances received parseable labels\.#### Emotion Distribution\.
Table[3](https://arxiv.org/html/2609.29056#S5.T3)presents the top\-10 emotion labels for texter and volunteer utterances in CTL dialogues\. The texter distribution is dominated byhopeless\(25\.0%\), with the top\-5 labels accounting for 66\.8% of all texter utterances\. Notably,hopefulranks third \(12\.6%\), reflecting moments of positive engagement interspersed with distress\. This coexistence is particularly relevant to digital expressions of grief by Black texters, where expressions of hope and joy more broadly do not necessarily replace grief but may emerge alongside it over the course of the interaction[Patton et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib29)\. The remaining labels form a long tail, indicating that although a small set of emotions dominates the corpus, the full 37\-label inventory inductively derived from the data captures a broader range of grief\-related and crisis experiences\. Volunteer utterances show a complementary profile, with greater concentration in positive and supportive emotional states, particularlyhopeful\(46\.4%\), whilehopeless\(21\.0%\) andoverwhelm\(11\.1%\) also remain prominent\. Together, these distributions illustrate the different interactional roles occupied by texters and volunteers, while motivating the transition\-based analyses below: aggregate prevalence alone does not reveal how these emotional states unfold over time\.
#### African American Language Analysis\.
Given that the CTL conversations involve self\-identified Black texters, we note that some conversations also involve the use of African American Language \(AAL\)–the variety of English used by many, but not all and not exclusively, African Americans in the US[Grieser \(2022\)](https://arxiv.org/html/2609.29056#bib.bib32)\. We qualitatively observe cases where the model appears to misinterpret the use of AAL; for example, one texter says, "I feel so alone n having to keep dis away for my kids, bout only person that bout understand I’m hurting at time is my son,"222Note, this example is paraphrased to protect texter privacy, but the use of AAL features is conserved\.which is classified ashopeful\. The model attributes this to the mention of the texter’s son understanding, but in the original comment the texter emphasizes their loneliness and the lack of others understanding\. We argue that these select cases do not significantly impact the overarching patterns identified given the validation of the model against expert annotations \(§[B](https://arxiv.org/html/2609.29056#A2)\) and that dense AAL features are not common in the corpus: the demographic alignment classifier introduced in[Blodgett et al\. \(2016\)](https://arxiv.org/html/2609.29056#bib.bib33)predicts AAL as the most likely label for∼\\sim6% of texts, and a probability exceeding \.8 for less than 1% of texts\. We do, however, note that models’ emotion labels on AAL texts are likely to be unreliable as also shown in prior work[Deas et al\. \(2023\)](https://arxiv.org/html/2609.29056#bib.bib30);[Deas et al\. \(2024a\)](https://arxiv.org/html/2609.29056#bib.bib31)\.
### 5\.2Level 2: Micro\-Dynamics
#### Emotion Transitions\.
Table[4](https://arxiv.org/html/2609.29056#S5.T4)presents the most frequent emotion\-label transitions for CTL texter utterances\. Six of the ten most frequent transitions are self\-transitions, reflecting substantial emotional inertia across adjacent texter turns\. Distress often persists rather than resolving immediately:hopeless→\\tohopeless,worthlessness→\\toworthlessness, andoverwhelm→\\tooverwhelmare among the most common patterns\.
Table 4:Top\-10 emotion transitions for CTL texters \(N=57,147N\{=\}57\{,\}147\)\. Self\-transitions dominate;hopeless→\\tohopeful\(rank 8\) is the most frequent recovery transition\. Percentages are calculated over all transitions\.At the same time, the transition structure contains recurring movement toward hope\.Hopeless→\\tohopefulranks eighth overall and is the most frequent recovery transition\. This pattern suggests that hope\-ward movement is not limited to conversation endpoints, but also appears locally within the interaction\. This also aligns with work identifying frequent discussions of self\-care and joy among Black social media users discussing grief\. Recovery in these CTL conversations therefore involves both persistent distress and repeated affective pivots rather than a simple replacement of negative emotion with positive emotion\.
Figure[2](https://arxiv.org/html/2609.29056#S5.F2)presents the conversation\-level emotion flow from start to end, where conversation start refers to the first substantive texter emotion after leading neutral\-labelled and sub\-three\-word opener turns are removed\. Conversations beginning inhopelessfrequently end inhopefulorgratitude, while conversations beginning in other distress\-related states disperse across both positive and negative end states\. This visualization complements the adjacent\-turn analysis by showing the net emotional movement across entire conversations\.
Figure 2:Conversation\-start to conversation\-end emotion flows for CTL texter utterances \(n=2,475n=2\{,\}475\)\. Left nodes show the first*substantive*emotion, after leading neutral\-labelled and sub\-three\-word opener turns are removed; right nodes show the last\. Prominent flows fromhopelesstowardhopefulandgratitudeillustrate hope\-ward movement at the conversation level\.
#### Emotional Volatility\.
Texters also exhibit greater emotional volatility than volunteers \(Table[15](https://arxiv.org/html/2609.29056#A6.T15)\), consistent with fluctuating affect during crisis\. Volunteers maintain more stable emotional stances across turns, further illustrating the complementary roles of the two speakers\.
#### Persistence and Transition Probabilities\.
Table[5](https://arxiv.org/html/2609.29056#S5.T5)reports polarity persistence for CTL texters and volunteers, while Table[6](https://arxiv.org/html/2609.29056#S5.T6)presents the transition matrix for texter utterances\. Three key patterns emerge for CTL\. First, negative states are highly persistent for texters \(84\.1%\), consistent with sustained distress during crisis, but substantially less so for volunteers \(64\.6%\)\. Second, positive states, once reached, are moderately stable for texters \(64\.5%\) and more so for volunteers \(72\.4%\), reflecting the counselor’s role in anchoring conversations in a supportive frame\. Third, neutral states are transient for both roles \(P\(stay\)≤25\.0%P\(\\text\{stay\}\)\\leq 25\.0\\%\)\. For texters, neutral states transition to positive emotion in 41\.4% of cases, compared with 14\.9% of transitions from negative states, suggesting that neutral states can function as intermediate points in broader affective movement\.
Table 5:Polarity persistence for CTL texter and volunteer utterances\.Table 6:Row\-normalized polarity transition matrix for CTL texter utterances\.Table 7:Conversation\-start and conversation\-end polarity distributions for CTL texter utterances\.
#### Conversation Arc\.
Table[7](https://arxiv.org/html/2609.29056#S5.T7)shows the starting and ending polarity distributions for CTL texter conversations, where conversation start refers to the first substantive texter utterance after leading neutral\-labelled and sub\-three\-word opener turns are removed\. CTL conversations overwhelmingly begin in negative states \(90\.9%\), while 70\.9% end in positive states\.
These results suggest a period of continued disclosure and emotional persistence before positive movement becomes visible in the texter’s language\. The CTL conversation arc is therefore not simply a difference between negative beginnings and positive endings; it is produced through repeated local transitions over the course of the interaction\.
### 5\.3Level 3: Macro\-Dynamics
The previous analyses characterize how texter emotions evolve over the course of CTL grief conversations\. We next examine the interactional role of volunteer responses: which support strategies are associated with downstream reductions in texter distress, and which are followed by continued or heightened distress?
#### Role Dynamics\.
Table[8](https://arxiv.org/html/2609.29056#S5.T8)shows that CTL volunteer responses are dominated byAffirmation and ReassuranceandQuestion, which together account for more than half of all predicted support strategies\. This distribution reflects two central functions of crisis support: providing immediate emotional validation and eliciting enough context to understand the texter’s situation\. More directive or resource\-oriented strategies, such asProviding SuggestionsandInformation, occur less frequently, suggesting that concrete advice and resource\-sharing are used more selectively\.
Table 8:Distribution of predicted volunteer support strategies in CTL conversations\.We also compute the downstream change in texter distress for each strategy\. Figure[3](https://arxiv.org/html/2609.29056#S5.F3)shows that volunteer strategies differ in their association with next\-turn texter distress change\. Some strategies are more often followed by de\-escalation, while others are associated with distress persistence or continued escalation\. In particular,Information,Providing Suggestions, andAffirmation and Reassuranceare associated with downstream distress reduction, withInformationshowing the largest mean decrease\.Reflection of Feelingsis also followed by a smaller reduction in distress\. However, the wider uncertainty intervals forInformationand the relatively infrequentSelf\-disclosurestrategy suggest that these estimates should be interpreted cautiously\.
By contrast,QuestionandRestatement or Paraphrasingare followed by small mean increases in texter distress\. We do not interpret this as evidence that these strategies are ineffective\. Rather, in crisis\-support conversations, questions and restatements often invite elaboration, clarify the texter’s situation, or keep the conversation open before de\-escalation occurs\. Their association with short\-term distress increases may therefore reflect their role in supporting continued disclosure rather than immediately reducing distress\.
These results highlight the importance of modeling crisis support as an interactional process\. Volunteer strategies are not interchangeable surface forms: the same texter emotion can be followed by different support moves, and those moves are associated with different short\-term affective trajectories\. This role\-level analysis complements the transition and archetype analyses by showing how local counselor behavior is associated with the emotional path of the conversation\.
Figure 3:Volunteer strategy effects in CTL conversations\. Bars show the downstream change in texter distress after each volunteer strategy, computed over local three\-turn windows\(ut,ht,ut\+1\)\(u\_\{t\},h\_\{t\},u\_\{t\+1\}\)\. Negative values indicate reduced texter distress after the volunteer response\.
#### Conversation\-Level Trajectory Archetypes\.
The trajectory analysis further shows that CTL conversations do not follow a single de\-escalation pattern\. Conversations are distributed relatively evenly across the identified archetypes \(Fig\.[5](https://arxiv.org/html/2609.29056#S6.F5)\), with each accounting for approximately 19–22% of the corpus\. These include early resolution, late recovery, late recovery under high distress, steady de\-escalation under high distress, and unresolved high distress\. This balance indicates substantial heterogeneity in how grief\-related crisis conversations unfold\. Some texters show early or gradual movement toward lower distress, while others recover late or remain distressed at the end of the interaction\. These archetypes reinforce the micro\-dynamic findings: authentic crisis support is not adequately represented by a single average recovery curve\.
#### Together,
the three\-level analysis shows that emotion dynamics in CTL grief conversations are neither static nor uniformly linear\. Distress is highly persistent, yet conversations also contain recurring hope\-ward pivots, gradual movement toward positive affect\. Texters and volunteers occupy complementary emotional roles with variations in how recovery unfolds\. These findings highlight why crisis support cannot be understood through utterance\-level labels or conversation endpoints alone\.Empathprovides a unified framework for examining how affect persists, shifts, and resolves across interaction\. More broadly, it offers an evaluation system for studying counseling conversations as dynamic processes rather than collections of isolated responses\.
## 6Secondary Analysis: Do Synthetic Dialogues Preserve Emotion Dynamics?
The preceding sections useEmpathto characterize emotion dynamics in authentic CTL grief conversations\. We next use the framework to ask whether synthetic crisis\-support dialogues preserve these dynamics\. We generate zero\-shot and dual\-agent conversations using three frontier LLMs and apply the same emotion, transition, strategy, and trajectory analyses used for CTL\. Full generation settings and prompts are provided in §[H](https://arxiv.org/html/2609.29056#A8)\.
#### Micro Dynamics\.
CTL conversations show de\-escalation composed of repeated local changes in emotional state\. As Figure[4](https://arxiv.org/html/2609.29056#S6.F4)illustrates, synthetic conversations also show substantial distress reduction, but often follow different paths: distress remains elevated for longer or changes more abruptly depending on the generation setup\.
This distinction is also visible in the transition structure\. CTL conversations contain recurring hope\-ward pivots such ashopeless→\\tohopeful, alongside substantial persistence within emotional states\. Synthetic texters show greater negative\-state persistence than CTL texters \(89\.4% vs\. 84\.1%\) and fewer direct negative→\\topositive transitions \(10\.3% vs\. 14\.9%\)\. Their most frequent fine\-grained transitions are also dominated by self\-loops and movement among negative states, with no cross\-polarity recovery transition appearing in the top ten\. The CTL data therefore suggest that de\-escalation is not simply an endpoint shift, but a sequence of smaller emotional transitions unfolding throughout the interaction\.
Figure 4:Mean texter distress trajectories for the full CTL analysis corpus and synthetic conditions\. All conditions show distress reduction, but differ in the pacing and structure of de\-escalation\.
#### Contextual Grounding\.
Table[9](https://arxiv.org/html/2609.29056#S6.T9)illustrates another property of the CTL interactions: support strategies are grounded in the texter’s specific circumstances\. Here, the CTL volunteer offers a grief\-specific resource and explains why it may be relevant\. The synthetic exchange uses the same broad*Information*strategy but provides more generic crisis resources\. This example illustrates why strategy labels alone do not capture how support is adapted to the preceding interaction\.
Table 9:Representative*Information*exchanges in CTL and synthetic conversations\. The CTL volunteer offers a specific, contextually matched resource \(GriefNet\); the synthetic volunteer offers generic escalation options \(988, ER\) without tailoring\. Texts drawn from CTL are paraphrased\.Figure 5:Distribution of conversation\-level distress trajectory archetypes across the full CTL analysis corpus, zero\-shot synthetic, and dual\-agent synthetic conversations\.
#### Macro Dynamics\.
The CTL corpus also contains substantial variation at the conversation level\. As Figure[5](https://arxiv.org/html/2609.29056#S6.F5)shows, CTL conversations are distributed relatively evenly across the five trajectory archetypes \(19–22%\), reflecting heterogeneous paths through distress and recovery\. Synthetic trajectories are more concentrated and depend strongly on generation setup: zero\-shot conversations overrepresent*Steady De\-escalation \(high distress\)*\(34% vs\. 19% in CTL\), whereas 64% of dual\-agent conversations fall into*Unresolved High Distress*, partly due to the dual\-agent termination procedure\. Neither synthetic condition reproduces the balanced trajectory distribution observed in CTL\.
Overall, this secondary analysis further highlights the structure observed in CTL: emotional change is gradual, locally patterned, contextually grounded, and heterogeneous across conversations\. Synthetic dialogues can reproduce some aggregate properties of these interactions, but do not consistently recover these finer\-grained dynamics\.Empathprovides a way to characterize these properties directly rather than relying on surface plausibility alone\.
## 7Conclusion
We introducedEmpath, a framework for analyzing emotion dynamics in crisis\-support dialogue through utterance\-level emotions, turn\-level transitions, and conversation\-level trajectory archetypes\. ApplyingEmpathto Crisis Text Line conversations with Black texters discussing grief, we find that these crisis support conversations involve persistent distress but also gradual movement toward hope, distinct texter–volunteer emotional roles, and heterogeneous recovery trajectories\. These results support that computationally, crisis support is better understood as a dynamic, interactional process than as a simple shift from negative to positive emotion\. Beyond CTL, we show through secondary analyses thatEmpathprovides a general perspective to understand counseling dialogues, measuring how affective states persist, shift, and resolve over time\.
While we focus primarily on CTL grief conversations in this study, in future work,Empathcan be extended to broader counseling contexts, including peer support, therapy dialogue, counselor training simulations, and mental health dialogue systems\. As privacy\-preserving synthetic and simulated data become increasingly common, emotion\-dynamic evaluation can help understand such dialogues and differences among them to ensure effective support\. Overall, our work positionsEmpathas a tool for studying digital grief support, evaluating realism in counseling technologies, and computationally understanding broader mental health dialogues\.
## Limitations
Our analyses characterize emotional expression within text\-based crisis\-support conversations and do not measure longitudinal changes in grief or establish clinical recovery\. Findings are specific to the samples and inclusion criteria described in Section[3](https://arxiv.org/html/2609.29056#S3); they should not be generalized to all CTL texters or to Black communities more broadly\. Emotion labels are model predictions, and the single\-label formulation may miss simultaneous emotions\. Performance on the 493 utterances with explicit human emotion annotations does not establish equivalent performance on all utterances or on African American Language\. Volunteer labels may reflect acknowledgment of texter emotions rather than volunteers’ own emotional states\. Associations between support strategies and subsequent label\-derived distress do not establish causal effects or counseling effectiveness\. Synthetic comparisons are also sensitive to topic composition, prompt design, context\-window length, and conversation termination criteria\.
## Ethics
Prior to data access, all Crisis Text Line data were de\-identified\. CTL data were accessed only under a signed Data Use Agreement and processed locally; no CTL data were provided to API\-based LLMs or other online services\. The example dialogue in Figure[1](https://arxiv.org/html/2609.29056#S2.F1)is simulated, and CTL excerpts presented in the paper are paraphrased to protect texter privacy\. No CTL data or CTL\-derived material were used in synthetic generation, including the construction of prompts or topic seeds\. The analyses are intended to study patterns of emotional expression and should not be treated as direct assessments of individual clinical states or of volunteer performance\.
## References
- Anthropic \(2025\)AnthropicClaude haiku 4\.5\.External Links:[Link](https://www.anthropic.com/claude/haiku)Cited by:[Appendix H](https://arxiv.org/html/2609.29056#A8.p2.1)\.
- Badawiet al\.\(2026\)A\. Badawi, E\. Rahimi, M\. T\. R\. Laskar, S\. Grach, L\. Bertrand, L\. Danok, P\. Dhanesh, J\. Huang, F\. Rudzicz, and E\. DolatabadiWhen can we trust LLMs in mental health? large\-scale benchmarks for reliable LLM evaluation\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 3873–3896\.External Links:[Link](https://aclanthology.org/2026.eacl-long.180/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.180),ISBN 979\-8\-89176\-380\-7Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
- Baiet al\.\(2025\)X\. Bai, G\. Chen, T\. He, C\. Zhou, and Y\. LiuEmotional supporters often use multiple strategies in a single turn\.External Links:2505\.15316,[Link](https://arxiv.org/abs/2505.15316)Cited by:[§4\.3](https://arxiv.org/html/2609.29056#S4.SS3.SSS0.Px1.p1.1)\.
- Blodgettet al\.\(2016\)S\. L\. Blodgett, L\. Green, and B\. O’ConnorDemographic dialectal variation in social media: a case study of African\-American English\.InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing,J\. Su, K\. Duh, and X\. Carreras \(Eds\.\),Austin, Texas,pp\. 1119–1130\.External Links:[Link](https://aclanthology.org/D16-1120/),[Document](https://dx.doi.org/10.18653/v1/D16-1120)Cited by:[§5\.1](https://arxiv.org/html/2609.29056#S5.SS1.SSS0.Px2.p1.1)\.
- Budaet al\.\(2026\)G\. Buda, I\. J\. Tripodi, K\. L\. Zuromski, M\. Meagher, and E\. A\. OlsonExtraction of texters’ explicit emotion expressions in crisis conversations\.InFindings of the Association for Computational Linguistics: ACL 2026,M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 27–44\.External Links:[Link](https://aclanthology.org/2026.findings-acl.2/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.2),ISBN 979\-8\-89176\-395\-1Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p1.1)\.
- Cabrera Lozoyaet al\.\(2025\)D\. Cabrera Lozoya, E\. Hernandez Lua, J\. A\. Barajas Perches, M\. Conway, and S\. D’AlfonsoSynthetic empathy: generating and evaluating artificial psychotherapy dialogues to detect empathy in counseling sessions\.InProceedings of the 10th Workshop on Computational Linguistics and Clinical Psychology \(CLPsych 2025\),A\. Zirikly, A\. Yates, B\. Desmet, M\. Ireland, S\. Bedrick, S\. MacAvaney, K\. Bar, and Y\. Ophir \(Eds\.\),Albuquerque, New Mexico,pp\. 157–171\.External Links:[Link](https://aclanthology.org/2025.clpsych-1.13/),[Document](https://dx.doi.org/10.18653/v1/2025.clpsych-1.13),ISBN 979\-8\-89176\-226\-8Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p5.1)\.
- Camburuet al\.\(2018\)O\. Camburu, T\. Rocktäschel, T\. Lukasiewicz, and P\. BlunsomE\-snli: natural language inference with natural language explanations\.InProceedings of the 32nd International Conference on Neural Information Processing Systems,NIPS’18,Red Hook, NY, USA,pp\. 9560–9572\.Cited by:[§4\.1](https://arxiv.org/html/2609.29056#S4.SS1.SSS0.Px1.p1.1)\.
- Deaset al\.\(2024a\)N\. Deas, J\. A\. Grieser, X\. Hou, S\. Kleiner, T\. Martin, S\. Nandanampati, D\. U\. Patton, and K\. McKeownPhonATe: impact of type\-written phonological features of african american language on generative language modeling tasks\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=rXEwxmnGQs)Cited by:[§5\.1](https://arxiv.org/html/2609.29056#S5.SS1.SSS0.Px2.p1.1)\.
- Deaset al\.\(2023\)N\. Deas, J\. Grieser, S\. Kleiner, D\. Patton, E\. Turcan, and K\. McKeownEvaluation of African American language bias in natural language generation\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 6805–6824\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.421/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.421)Cited by:[§5\.1](https://arxiv.org/html/2609.29056#S5.SS1.SSS0.Px2.p1.1)\.
- Deaset al\.\(2024b\)N\. Deas, E\. Turcan, I\. E\. P\. Mejia, and K\. McKeownMASIVE: open\-ended affective state identification in English and Spanish\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 20467–20485\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.1139/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.1139)Cited by:[§4\.1](https://arxiv.org/html/2609.29056#S4.SS1.p1.1)\.
- Devlinet al\.\(2019\)J\. Devlin, M\. Chang, K\. Lee, and K\. ToutanovaBERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 4171–4186\.Cited by:[item 1](https://arxiv.org/html/2609.29056#A2.I1.i1.p1.1)\.
- Flemings and Annavaram \(2024\)J\. Flemings and M\. AnnavaramDifferentially private knowledge distillation via synthetic text generation\.InAnnual Meeting of the Association for Computational Linguistics,External Links:[Link](https://api.semanticscholar.org/CorpusID:268230792)Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p5.1)\.
- Gonget al\.\(2024\)Z\. Gong, X\. Hu, M\. Yao, X\. Zhu, and J\. HirschbergA mapping on current classifying categories of emotions used in multimodal models for emotion recognition\.InProceedings of the 18th Linguistic Annotation Workshop \(LAW\-XVIII\),pp\. 19–28\.Cited by:[Appendix B](https://arxiv.org/html/2609.29056#A2.SS0.SSS0.Px3.p2.1),[Appendix B](https://arxiv.org/html/2609.29056#A2.SS0.SSS0.Px3.p3.1),[Table 11](https://arxiv.org/html/2609.29056#A2.T11),[Table 14](https://arxiv.org/html/2609.29056#A4.T14),[Appendix D](https://arxiv.org/html/2609.29056#A4.p1.1)\.
- Gonget al\.\(2023\)Z\. Gong, Q\. Min, and Y\. ZhangEliciting rich positive emotions in dialogue generation\.InProceedings of the First Workshop on Social Influence in Conversations \(SICon 2023\),K\. Chawla and W\. Shi \(Eds\.\),Toronto, Canada,pp\. 1–8\.External Links:[Link](https://aclanthology.org/2023.sicon-1.1/),[Document](https://dx.doi.org/10.18653/v1/2023.sicon-1.1)Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p1.1)\.
- Google \(2025\)GoogleGemini 2\.5 pro\.External Links:[Link](https://deepmind.google/technologies/gemini/)Cited by:[Appendix H](https://arxiv.org/html/2609.29056#A8.p2.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri,et al\.The Llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[Appendix A](https://arxiv.org/html/2609.29056#A1.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.29056#S4.SS1.SSS0.Px2.p1.1)\.
- Grieser \(2022\)J\. A\. GrieserThe black side of the river: race, language, and belonging in washington, dc\.Georgetown University Press\.Cited by:[§5\.1](https://arxiv.org/html/2609.29056#S5.SS1.SSS0.Px2.p1.1)\.
- Hall \(2014\)C\. HallBereavement theory: recent developments in our understanding of grief and bereavement\.Bereavement Care33\(1\),pp\. 7–12\.External Links:[Document](https://dx.doi.org/10.1080/02682621.2014.902610),[Link](https://doi.org/10.1080/02682621.2014.902610),https://doi\.org/10\.1080/02682621\.2014\.902610Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p2.1)\.
- Jiet al\.\(2022\)S\. Ji, T\. Zhang, L\. Ansari, J\. Fu, P\. Tiwari, and E\. CambriaMentalBERT: publicly available pretrained language models for mental health\.arXiv preprint arXiv:2110\.15621\.Cited by:[item 2](https://arxiv.org/html/2609.29056#A2.I1.i2.p1.1),[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
- Jianget al\.\(2023\)A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. d\. l\. Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. SayedMistral 7B\.arXiv preprint arXiv:2310\.06825\.Cited by:[item 5](https://arxiv.org/html/2609.29056#A2.I1.i5.p1.1)\.
- Jordanet al\.\(2025\)E\. Jordan, R\. Terrisse, V\. Lucarini, M\. Alrahabi, M\. Krebs, J\. Desclés, and C\. LemeySpeech emotion recognition in mental health: systematic review of voice\-based applications\.JMIR mental health12\(1\),pp\. e74260\.Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p1.1)\.
- Kurakinet al\.\(2023\)A\. Kurakin, N\. Ponomareva, U\. Syed, L\. MacDermed, and A\. TerzisHarnessing large\-language models to generate private synthetic text\.ArXivabs/2306\.01684\.External Links:[Link](https://api.semanticscholar.org/CorpusID:259063934)Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p5.1)\.
- Limpijankitet al\.\(2025\)M\. Limpijankit, Y\. Chen, M\. Subbiah, N\. Deas, and K\. McKeownCounterfactual simulatability of LLM explanations for generation tasks\.InProceedings of the 18th International Natural Language Generation Conference,L\. Flek, S\. Narayan, L\. H\. Phương, and J\. Pei \(Eds\.\),Hanoi, Vietnam,pp\. 659–683\.External Links:[Link](https://aclanthology.org/2025.inlg-main.38/)Cited by:[§4\.1](https://arxiv.org/html/2609.29056#S4.SS1.SSS0.Px1.p1.1)\.
- Liuet al\.\(2021\)S\. Liu, C\. Zheng, O\. Demasi, S\. Sabour, Y\. Li, Z\. Yu, Y\. Jiang, and M\. HuangTowards emotional support dialog systems\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 3469–3483\.External Links:[Link](https://aclanthology.org/2021.acl-long.269/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.269)Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
- Liuet al\.\(2026\)Z\. Liu, Z\. Gong, L\. Ai, Z\. Hui, R\. Chen, C\. W\. Leach, M\. R\. Greene, and J\. HirschbergA review of incorporating psychological theories in LLMs\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 7459–7495\.External Links:[Link](https://aclanthology.org/2026.eacl-long.350/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.350),ISBN 979\-8\-89176\-380\-7Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
- Louieet al\.\(2024\)R\. Louie, A\. Nandi, W\. Fang, C\. Chang, E\. Brunskill, and D\. YangRoleplay\-doh: enabling domain\-experts to create LLM\-simulated patients via eliciting and adhering to principles\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 10570–10603\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.591/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.591)Cited by:[Appendix H](https://arxiv.org/html/2609.29056#A8.p4.1),[§1](https://arxiv.org/html/2609.29056#S1.p5.1)\.
- Louieet al\.\(2025\)R\. Louie, I\. H\. Orney, J\. P\. Pacheco, R\. S\. Shah, E\. Brunskill, and D\. YangCan LLM\-simulated practice and feedback upskill human counselors? A randomized study with 90\+ novice counselors\.External Links:2505\.02428,[Link](https://arxiv.org/abs/2505.02428)Cited by:[Appendix H](https://arxiv.org/html/2609.29056#A8.p4.1)\.
- Louieet al\.\(2026\)R\. Louie, R\. S\. Shah, I\. H\. Orney, J\. P\. Pacheco, E\. Brunskill, and D\. YangCan llm\-simulated practice and feedback upskill human counselors? a randomized study with 90\+ novice counselors\.InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems,CHI ’26,pp\. 1–31\.External Links:[Link](http://dx.doi.org/10.1145/3772318.3791821),[Document](https://dx.doi.org/10.1145/3772318.3791821)Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p5.1)\.
- Margaret Stroebe \(1999\)H\. S\. Margaret StroebeTHE dual process model of coping with bereavement: rationale and description\.Death Studies23\(3\),pp\. 197–224\.Note:PMID: 10848151External Links:[Document](https://dx.doi.org/10.1080/074811899201046),[Link](https://doi.org/10.1080/074811899201046),https://doi\.org/10\.1080/074811899201046Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p2.1)\.
- Mohammad \(2018\)S\. M\. MohammadObtaining reliable human ratings of valence, arousal, and dominance for 20,000 English words\.InProceedings of ACL,Cited by:[§4\.2](https://arxiv.org/html/2609.29056#S4.SS2.p1.1)\.
- Mooreet al\.\(2017\)J\. Moore, S\. Magee, E\. Gamreklidze, and J\. KowalewskiSocial media mourning: using grounded theory to explore how people grieve on social networking sites\.OMEGA \- Journal of Death and Dying79\(3\),pp\. 231–259\.External Links:ISSN 1541\-3764,[Link](http://dx.doi.org/10.1177/0030222817709691),[Document](https://dx.doi.org/10.1177/0030222817709691)Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p2.1)\.
- Neimeyeret al\.\(2002\)R\. A\. Neimeyer, H\. G\. Prigerson, and B\. DaviesMourning and meaning\.American Behavioral Scientist46\(2\),pp\. 235–251\.External Links:ISSN 1552\-3381,[Link](http://dx.doi.org/10.1177/000276402236676),[Document](https://dx.doi.org/10.1177/000276402236676)Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p2.1)\.
- OpenAI \(2024\)OpenAIGPT\-4o system card\.External Links:[Link](https://openai.com/index/gpt-4o-system-card/)Cited by:[Appendix H](https://arxiv.org/html/2609.29056#A8.p2.1)\.
- Pattonet al\.\(2025\)D\. U\. Patton, S\. Kleiner, S\. Miller, N\. Deas, F\. J\. Edwards, J\. A\. Grieser, J\. Shepard, E\. Turcan, and K\. McKeownDigital narratives of grief and resilience: insights from the integrating emotional stories online \(ieso\) platform\.InNew Trends in Disruptive Technologies, Tech Ethics and Artificial Intelligence,D\. H\. de la Iglesia, J\. F\. de Paz Santana, and A\. J\. López Rivero \(Eds\.\),Cham,pp\. 368–378\.External Links:ISBN 978\-3\-031\-99474\-6Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p2.1),[§2](https://arxiv.org/html/2609.29056#S2.p2.1),[§5\.1](https://arxiv.org/html/2609.29056#S5.SS1.SSS0.Px1.p1.1)\.
- Stroebe and Schut \(2001\)M\. S\. Stroebe and H\. SchutMeaning making in the dual process model of coping with bereavement\.\.InMeaning reconstruction & the experience of loss\.,pp\. 55–73\.External Links:ISBN 1557987424,[Link](http://dx.doi.org/10.1037/10397-003),[Document](https://dx.doi.org/10.1037/10397-003)Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p2.1)\.
- Tripodiet al\.\(2025\)I\. J\. Tripodi, G\. Buda, M\. Meagher, and E\. A\. OlsonAssessing effective de\-escalation of crisis conversations using transformer\-based models and trend statistics\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 29763–29777\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1512/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1512),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
- Wanet al\.\(2025\)C\. Wan, M\. Labeau, and C\. ClavelEmoDynamiX: emotional support dialogue strategy prediction by modelling MiXed emotions and discourse dynamics\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 1678–1695\.External Links:[Link](https://aclanthology.org/2025.naacl-long.81/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.81),ISBN 979\-8\-89176\-189\-6Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
- Wanget al\.\(2024\)R\. Wang, S\. Milani, J\. C\. Chiu, J\. Zhi, S\. M\. Eack, T\. Labrum, S\. M\. Murphy, N\. Jones, K\. V\. Hardy, H\. Shen, F\. Fang, and Z\. ChenPATIENT\-ψ\\psi: using large language models to simulate patients for training mental health professionals\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 12772–12797\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.711/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.711)Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p5.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-consistency improves chain of thought reasoning in language models\.arXiv preprint arXiv:2203\.11171\.Cited by:[item 3](https://arxiv.org/html/2609.29056#A1.I1.i3.p1.1)\.
- Wilson and O’Connor \(2022\)D\. T\. Wilson and M\. O’ConnorFrom grief to grievance: combined axes of personal and collective grief among black americans\.Frontiers in PsychiatryVolume 13 \- 2022\.External Links:[Link](https://www.frontiersin.org/journals/psychiatry/articles/10.3389/fpsyt.2022.850994),[Document](https://dx.doi.org/10.3389/fpsyt.2022.850994),ISSN 1664\-0640Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p2.1)\.
- Wuet al\.\(2025\)C\. Wu, Y\. Cai, Y\. Liu, P\. Zhu, Y\. Xue, Z\. Gong, J\. Hirschberg, and B\. MaMultimodal emotion recognition in conversations: a survey of methods, trends, challenges and prospects\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 6257–6274\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.332/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.332),ISBN 979\-8\-89176\-335\-7Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
- Wuet al\.\(2024\)Z\. Wu, Z\. Gong, J\. Koo, and J\. HirschbergMultimodal multi\-loss fusion network for sentiment analysis\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),K\. Duh, H\. Gomez, and S\. Bethard \(Eds\.\),Mexico City, Mexico,pp\. 3588–3602\.External Links:[Link](https://aclanthology.org/2024.naacl-long.197/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.197)Cited by:[§1](https://arxiv.org/html/2609.29056#S1.p1.1)\.
- Xieet al\.\(2025\)S\. J\. Xie, S\. Zhai, Y\. Liang, J\. Li, X\. Fan, T\. Cohen, and W\. YuwenCultural prompting improves the empathy and cultural responsiveness of gpt\-generated therapy responses\.External Links:2512\.00014,[Link](https://arxiv.org/abs/2512.00014)Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
- Xuet al\.\(2024\)X\. Xu, B\. Yao, Y\. Dong, S\. Gabriel, H\. Yu, J\. Hendler, M\. Ghassemi, A\. K\. Dey, and D\. WangMental\-llm: leveraging large language models for mental health prediction via online text data\.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies8\(1\),pp\. 1–32\.External Links:ISSN 2474\-9567,[Link](http://dx.doi.org/10.1145/3643540),[Document](https://dx.doi.org/10.1145/3643540)Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
- Zhanget al\.\(2025\)X\. Zhang, W\. Wang, and Q\. JinIntentionESC: an intention\-centered framework for enhancing emotional support in dialogue systems\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 26494–26516\.External Links:[Link](https://aclanthology.org/2025.findings-acl.1358/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.1358),ISBN 979\-8\-89176\-256\-5Cited by:[§2](https://arxiv.org/html/2609.29056#S2.p1.1)\.
## Appendix AEmotion Detection Prompt, Label Inventory, and Experimental Set\-Ups
#### Data Pre\-processing\.
Before inference, we exclude three categories of utterances: \(i\) system messages, \(ii\) single\-character Y/N texter feedback, and \(iii\) trivial first texter messages, defined as the first texter utterance in a conversation when it is empty, whitespace\-only, or consists of a single token \(e\.g\., “Hi”\)\. These trivial openers carry little or no emotional content and would otherwise receive unreliable*neutral*labels, potentially distorting downstream transition statistics\.
For sequence\-level analyses, we apply an additional opener filter after inference\. Starting from the beginning of each conversation, leading utterances are removed iteratively until reaching the first utterance that is both non\-*neutral*and substantive\. An utterance is considered non\-substantive if it contains fewer than three whitespace\-delimited words or, for space\-free CJK text, at most three characters\. This criterion is necessary because short procedural openers such as “HOME”, “WARM”, or “Ok” may receive non\-neutral labels due primarily to surrounding conversational context rather than their own semantic content\. Importantly, this filtering is restricted to conversation openers: mid\-conversation utterances are never removed, and all mid\-conversation*neutral*labels are retained\.
This second filter is applied only to sequence\-level statistics\. Label\-frequency analyses are computed over all labeled utterances after the initial pre\-processing step, since removing conversation openers would alter the overall label distribution and therefore bias what is intended to be a corpus\-level register statistic\.
#### Experimental Setups\.
We use meta\-llama/Llama\-3\.2\-3B\-Instruct[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib11)as the backbone model, loaded in FP16 precision on a single CUDA\-capable GPU\. Inference uses greedy decoding \(Stage 1\) with a maximum of 128 new tokens\. For self\-consistency runs \(Stage 3\), we sample with temperature=0\.7=0\.7and top\-p=0\.9p\{=\}0\.9overK=5K\{=\}5independent generations\. The batch size is set to 64 and the random seed to 42 for reproducibility\. Total inference time for the CTL dataset was approximately 27\.5 hours \(117,248 items across 1,832 batches\)\.
We run two configurations per corpus: \(i\)*without context*\(single\-utterance input\), and \(ii\)*with context*using a sliding window of the 5 most recent preceding utterances\. In both settings, system messages, Y/N feedback rows, and trivial first texter messages are excluded from inference\. For the synthetic data, we use context length 10 to match the longer average turn lengths in generated conversations\.
#### Prompt Template \(with\-context mode\)\.
The following template is instantiated per utterance\.\{label\_definitions\}is expanded to the full 37\-label list with definitions;\{conversation\}contains the sliding window of preceding utterances plus the target;\{author\}is the speaker role of the target utterance;\{label\_list\}is the comma\-separated list of allowed labels\.
```
You are an expert in emotion recognition
and mental health support.
Below is a conversation between a texter
and a volunteer. Here are the possible
emotion labels:
{label_definitions}
Conversation so far:
{conversation}
Focus especially on the last message by the
{author} shown above.
When determining the most appropriate emotion
label, give the most weight to the content
and tone of the current (last) message,
and only use the earlier conversation as
supporting context if needed.
You MUST respond in exactly this format
and nothing else:
Label: <one label chosen from this list:
{label_list}>
Reason: <one short sentence, referencing
the definition and the current
message>
```
#### Label Inventory\.
Table[10](https://arxiv.org/html/2609.29056#A1.T10)lists all 37 emotion labels and their definitions as provided to the model\.
Table 10:The 37 emotion labels and abridged definitions used in the detection prompt\. Full definitions are provided verbatim to the model\. Theselflabel merges the original coding\-book categoriesself\-doubtandself\-aware\.
#### Emotion Detection Pipeline\.
We use a three\-stage cascading procedure, applied independently to each batch so that later stages process only unresolved or uncertain examples:
1. 1\.Deterministic generation\.The model generates greedily \(temperature=0=0\)\. A prediction is accepted only from an explicitly completedLabel:field\. The generated value is matched to one of the 37 valid labels after normalization; if no exact match is found, we accept a valid label appearing as the leading whole word \(e\.g\., “anxiety, because…”\)\. Unfilled placeholders and outputs listing three or more distinct labels are rejected\. We do not perform whole\-output substring matching\.
2. 2\.First\-token scoring\.For rows that fail parsing, we score all labels in a single forward pass using the logits at theLabel:position and select the label with the highest first\-token score\.
3. 3\.Self\-consistency refinement\.If the margin between the top two scores is small \(Δ<1\.0\\Delta<1\.0\), we drawK=5K\{=\}5stochastic generations \(temperature=0\.7=0\.7, top\-p=0\.9p\{=\}0\.9\) and use majority vote[Wang et al\. \(2023\)](https://arxiv.org/html/2609.29056#bib.bib12)\. Unparseable samples are discarded; if none parse, the stage\-2 prediction is retained\. Uncertain rows are batched together for each of theKKsampling passes\.
If a row reaches a fallback stage when no local model is available, it is assigned*neutral*and explicitly flagged as a fallback case\.
## Appendix BModel Validation
We validate the emotion detection model by comparing three candidate architectures on the 100\-conversation CTL annotation sample\. Each model’s predictions are evaluated against human reference annotations from six independently coded CTL conversation subsets using two complementary metrics:*semantic similarity*\(cosine similarity between label embeddings\) and*exact\-match accuracy*\. We favor semantic similarity as the primary metric because, with a 37\-category inventory, exact match is overly strict: clinically similar predictions \(e\.g\.,hopelessvs\.worthlessness\) are penalized equally to entirely wrong ones \(e\.g\.,hopelessvs\.joy\)\. Semantic similarity provides graded credit that better reflects the practical quality of predictions\. Standard per\-class precision, recall, and F1 are not reported because \(i\) BERT Emotions uses a different label taxonomy, making class\-level comparison infeasible, and \(ii\) with 37 fine\-grained categories and limited per\-subset sizes \(503–2,203 utterances\), many individual classes have too few samples for stable per\-class estimates\. All models are evaluated on utterances with explicit human emotion annotations\.
#### Candidate Models\.
We compare five approaches spanning the encoder\-only and generative paradigms:
1. 1\.BERT Emotions[Devlin et al\. \(2019\)](https://arxiv.org/html/2609.29056#bib.bib13): A BERT\-base model fine\-tuned on emotion classification, serving as the reference baseline\. Because BERT Emotions uses a different label taxonomy than our 37\-category inventory, exact\-match accuracy cannot be computed; only semantic similarity between its predicted labels and the reference is reported\.
2. 2\.MentalBERT[Ji et al\. \(2022\)](https://arxiv.org/html/2609.29056#bib.bib14): A BERT variant pre\-trained on mental\-health corpora \(Reddit counseling, psychological forums\), hypothesized to better capture therapeutic language\.
3. 3\.Llama\-3\.2\-3B\-Instruct: The pipeline described in §[A](https://arxiv.org/html/2609.29056#A1.SS0.SSS0.Px5), using Llama\-3\.2\-3B\-Instruct with retrieval\-augmented label definitions\. Evaluated under two context conditions: without conversational context \(single\-utterance input\) and with context \(sliding window of preceding turns\)\.
4. 4\.Llama\-3\.1\-8B\-Instruct: A larger Llama variant \(8B parameters\) to assess whether increased model capacity improves emotion classification within the same pipeline\.
5. 5\.Mistral\-7B\-Instruct\-v0\.3[Jiang et al\. \(2023\)](https://arxiv.org/html/2609.29056#bib.bib15): A 7B\-parameter instruction\-tuned model from a different model family, included to evaluate cross\-architecture generalization of the pipeline\.
#### Comparison Results\.
BERT Emotions and MentalBERT are both BERT\-based architectures with a maximum input length of 512 tokens \(64 tokens for BERT Emotions\), designed for single\-sentence classification\. They cannot naturally incorporate multi\-turn conversational context: BERT Emotions was fine\-tuned on short single\-sentence inputs, and MentalBERT was continually pre\-trained on individual Reddit posts rather than multi\-turn dialogues\. In contrast, the generative models \(Llama\-3\.2\-3B, Llama\-3\.1\-8B, Mistral\-7B\) support large context windows and were instruction\-tuned on conversational formats\. We compare all five models in the without\-context setting for a fair evaluation\. Table[11](https://arxiv.org/html/2609.29056#A2.T11)reports performance on 493 human\-annotated utterances with explicit emotion labels\.
Mistral\-7B\-Instruct\-v0\.3 achieves the highest accuracy at all taxonomy levels\. However, we use Llama\-3\.2\-3B\-Instruct for all primary analyses\. The accuracy gap narrows as the taxonomy becomes coarser: at the Ekman 7\-category level the difference is only 1\.5 percentage points \(0\.682 vs\. 0\.667\), and at the sentiment level the two models are nearly indistinguishable \(0\.844 vs\. 0\.842\)\. Moreover, Llama\-3\.2\-3B\-Instruct achieves semantic similarity comparable to or higher than the other generative models at every mapped level \(0\.975–0\.983\), indicating that its predictions are semantically closest to the human annotations even when the exact label differs\. The accuracy differences therefore reflect fine\-grained synonym disagreements \(e\.g\.,hopelessvs\.worthlessness\) rather than systematic misclassification\. Because our dynamics analysis aggregates over tens of thousands of transitions, these per\-utterance differences are unlikely to alter corpus\-level patterns such as transition rankings, polarity persistence rates, or conversation\-arc statistics\. Additionally, Llama\-3\.2\-3B \(3B parameters\) outperforms the larger Llama\-3\.1\-8B \(8B parameters\) at all taxonomy levels despite being less than half the size, offering the best accuracy\-to\-compute tradeoff for processing 117K\+ utterances across both CTL and synthetic corpora\. The multi\-model validation thus serves to demonstrate that the pipeline generalizes across model families and scales, rather than to select a single backbone\.
#### Multi\-Level Taxonomy Evaluation\.
The 37\-label exact\-match accuracy reported above is a conservative lower bound: clinically similar predictions \(e\.g\.,hopelessvs\.worthlessness\) are penalized as errors even though both reflect the same broad emotional state\. To quantify how much of the apparent “error” is attributable to fine\-grained label confusion rather than genuine misclassification, we evaluate accuracy at progressively coarser emotion taxonomies\.
We adopt the cascading mapping method of[Gong et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib17), which maps fine\-grained emotion labels to coarser categories based on shared names, valence/arousal similarity, and human evaluation\. Our 37 labels are first mapped to Plutchik’s 14 fine\-grained categories \(e\.g\.,hopeless,worthlessness,loneliness→\\to*sadness*;anxiety,stress,overwhelm→\\to*fear*;happiness,love,gratitude→\\to*joy*\), then cascaded to Ekman’s 7 basic emotions and 3 sentiment classes following the paper’s validated hierarchy\. The full mapping is provided in §[D](https://arxiv.org/html/2609.29056#A4)\. Both the model’s predicted labels and the human reference annotations are mapped to the same coarser taxonomy before computing accuracy, so a prediction ofhopelessagainst a human label ofworthlessnesscounts as correct at the 14\-category level \(both map to*sadness*\) even though it is an error at the 37\-label level\.
For BERT Emotions, which predicts from a 13\-label taxonomy, exact\-match accuracy cannot be computed at the 37\-label or 14\-category levels due to taxonomy mismatch\. At the Ekman 7\-category and sentiment 3\-category levels, we map its 13 labels to the same targets using shared label names and valence/arousal alignment, then apply the 7→\\to3 cascade from[Gong et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib17)\. Table[11](https://arxiv.org/html/2609.29056#A2.T11)reports evaluation on the 493 utterances with explicit human emotion annotations\.
Table 11:Model evaluation \(N=493N\{=\}493\) across taxonomy levels: 37 \(original\), 14 \(Plutchik fine\-grained\), 7 \(Ekman basic\), and 3 \(sentiment\)\. BERT Emotions accuracy at levels 37 and 14 is not reported due to taxonomy mismatch; at levels 7 and 3, its 13 labels are mapped to the same targets via shared names and the[Gong et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib17)cascade\. Bold indicates best per metric\.At the Ekman 7\-category level, all three generative models exceed 0\.64 accuracy, with Mistral\-7B reaching 0\.682\. At the sentiment level \(3 categories\), accuracy rises above 0\.81 for all generative models\. BERT Emotions, now comparable at the mapped levels, achieves 0\.335 accuracy at Ekman 7 and 0\.582 at sentiment 3, substantially below the generative models but above MentalBERT at the 7\-category level\. This confirms that the majority of prediction errors are confusions between semantically adjacent labels within the same broad category \(e\.g\.,hopelessvs\.worthlessness, both mapping to*sadness*\), rather than gross misclassifications across emotional poles\. The convergence of semantic similarity across all taxonomy levels \(0\.974–0\.983 for generative models\) further confirms that model choice does not materially affect the analysis\. MentalBERT, while substantially better than the random baseline at the sentiment level \(0\.688 vs\. 0\.582\), still lags behind all other models at every taxonomy level\.
We further use the*with\-context*configuration for all primary analyses\. As we demonstrate in §[F\.1](https://arxiv.org/html/2609.29056#A6.SS1)and §[G](https://arxiv.org/html/2609.29056#A7), the without\-context configuration fails to capture sustained emotional states, produces erratic label sequences \(change rate 84\.2% vs\. 63\.9%\), and collapses the texter–volunteer role distinction\. The with\-context configuration sacrifices some single\-utterance accuracy but produces temporally coherent emotion trajectories essential for dynamics analysis\.
#### Design Rationale\.
These results motivate our choice of the generative prompt\-engineering approach\. First, the structured prompt, which supplies explicit definitions for all 37 categories alongside conversational context, enables nuanced distinctions \(e\.g\.,*grateful*vs\.*hopeful*,*anxious*vs\.*afraid*\) that fixed\-vocabulary classifiers collapse\. Second, the three\-stage cascade \(§[A](https://arxiv.org/html/2609.29056#A1.SS0.SSS0.Px5)\) provides a principled fallback that maintains label validity even when free\-form generation fails\. Third, theReason:field offers interpretable justifications auditable for clinical plausibility, a property absent from softmax\-based classifiers\. Finally, the pipeline’s advantage is robust across all three generative models, all six evaluation subsets, and both context conditions, suggesting that the prompt design generalizes well across architectures and conversational structures\.
#### Llama\-3\.2\-3B\-Instruct Context Ablation\.
Table[12](https://arxiv.org/html/2609.29056#A2.T12)compares the Llama\-3\.2\-3B\-Instruct pipeline under both context conditions on 493 human\-annotated utterances with explicit emotion labels\. The without\-context configuration achieves higher single\-utterance accuracy, as expected since human annotations were produced without conversational context\. However, the with\-context configuration produces more temporally coherent label sequences needed for dynamics analysis \(§[F\.1](https://arxiv.org/html/2609.29056#A6.SS1)\)\.
Table 12:Llama\-3\.2\-3B\-Instruct context ablation \(N=493N\{=\}493utterances with explicit human emotion labels\)\. The without\-context configuration scores higher on single\-utterance validation but fails to capture temporal dynamics \(see §[G](https://arxiv.org/html/2609.29056#A7)\)\.
## Appendix CEmotion to Polarity Mapping
Each of the 37 emotion labels is mapped to a polarity class \(negative, neutral, or positive\) via the NRC VAD Lexicon\. Labels with valence≤0\.45\\leq 0\.45are classified as negative, valence≥0\.55\\geq 0\.55as positive, and0\.45<valence<0\.550\.45<\\text\{valence\}<0\.55as neutral\. Table[13](https://arxiv.org/html/2609.29056#A3.T13)lists the full mapping\.
Emotion LabelVAD Term\(s\)ValencePolarity*Negative \(valence≤\\leq0\.45\), 21 labels*worthlessnessworthless, worthlessness0\.042Negativefearfear, afraid0\.042Negativeshameshame, shameful0\.050Negativedisgustdisgust, disgusted0\.052Negativehopelesshopeless, hopelessness0\.060Negativedistressdistress, distressed0\.108Negativedisappointmentdisappointment0\.115Negativetiredtired0\.125Negativeangeranger, angry0\.145Negativeboredombored, boredom0\.160Negativeworryworry, worried0\.170Negativestressstress, stressed0\.170Negativeguiltguilt, guilty0\.172Negativelonelinesslonely, loneliness0\.198Negativeregretregret, regretful0\.199Negativenumbnessnumb, numbness0\.207Negativeanxietyanxiety, anxious0\.214Negativesadnesssad0\.225Negativepreoccupiedpreoccupied0\.265Negativeoverwhelmoverwhelm, overwhelmed0\.296Negativedistractiondistraction, distracted0\.298Negative*Neutral \(0\.45<0\.45<valence<0\.55<0\.55\), 3 labels*neutralneutral0\.469Neutralmoodmood, moody0\.483Neutralresilientresilient0\.542Neutral*Positive \(valence≥\\geq0\.55\), 12 labels*longinglonging0\.604Positiveanticipationanticipation0\.698Positiveselfself0\.704Positiveserenityserenity, serene0\.851Positiveempathyempathy, empathetic0\.865Positivesurprisesurprise0\.875Positivegratitudegratitude, grateful0\.922Positivetrusttrust, trustworthy0\.929Positivehopefulhopeful0\.947Positivehappinesshappiness, happy0\.976Positivejoyjoy0\.980Positivelovelove0\.998PositiveTable 13:Mapping of 37 emotion labels to polarity classes via NRC\-VAD valence scores\. Labels with valence≤0\.45\\leq 0\.45are negative,≥0\.55\\geq 0\.55are positive, and intermediate values are neutral\. The labelchaotichas no direct entry in the NRC\-VAD lexicon and is omitted from polarity analysis\.
## Appendix DMulti\-Level Emotion Taxonomy Mapping
Table[14](https://arxiv.org/html/2609.29056#A4.T14)shows the mapping from our 37 emotion labels to the 14 Plutchik fine\-grained categories, following the method of[Gong et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib17)\. Labels sharing the same name are mapped directly; remaining labels are assigned based on valence and arousal similarity\. The 14 categories are then cascaded to coarser levels following the paper’s validated hierarchy: 14→\\to7 \(Ekman\) merges*serenity*→\\to*joy*,*annoyance*→\\to*anger*,*boredom*→\\to*disgust*,*distraction*→\\to*sadness*,*interest*→\\to*anticipation*→\\to*neutral*, and*trust*→\\to*joy*; 7→\\to3 \(sentiment\) maps*joy*→\\to*positive*,*anger/disgust/sadness/fear/surprise*→\\to*negative*, and*neutral*→\\to*neutral*\.
Table 14:Mapping of 37 emotion labels to 14 Plutchik fine\-grained categories following[Gong et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib17)\. Direct name matches are mapped as\-is; remaining labels are assigned via valence/arousal similarity\.
## Appendix EArchetype Analysis Details
#### Trajectory Representation\.
For each conversation, we extract the texter’s turn\-level distress sequence using the emotion scoring scheme, discarding conversations with fewer than four labelled texter turns\. Because conversations vary in length, we linearly interpolate each sequence onto a shared grid ofT=20T=20equally spaced points on the normalised progress axis\[0,1\]\[0,1\]\. This representation preserves trajectory shape while abstracting away absolute duration, enabling direct comparison across corpora\.
#### Pooled clustering\.
We concatenate trajectories from all three sources—zero\-shot synthetic, dual\-agent synthetic, and CTL—and fitKK\-means withK=5K=5clusters \(n\_init=20=20, seed4242\)\. Pooling ensures that archetype labels are shared across sources, so differences in archetype prevalence reflect genuine compositional shifts rather than source\-specific partitioning artifacts\.
#### Archetype Labeling\.
We characterize each cluster centroid using four shape features: mean distress level, total decline in distress from start to end, the fraction of that decline occurring in the final quarter, and volatility, measured as the standard deviation of first\-order differences\. Archetype names are assigned directly from these features rather than selected from a fixed candidate list\.
Centroids with a total distress decline below0\.200\.20are labeled*Unresolved High Distress*when their mean distress is at least0\.650\.65, and*Persistent Moderate Distress*otherwise\. For centroids showing a larger decline, labels depend on its timing:*Late Recovery*when at least60%60\\%of the decline occurs in the final quarter,*Early Resolution*when at most30%30\\%occurs there, and*Steady De\-escalation*otherwise\. These labels receive the suffix*\(high distress\)*when mean distress is at least0\.650\.65\. We additionally reserve*Rupture\-Repair*and*Oscillatory Distress*for unusually volatile centroids, defined as having volatility above0\.050\.05and at least0\.050\.05greater than the median volatility of the remaining centroids\. No centroid meets this relative\-volatility criterion in our data, so these labels are not assigned\. When multiple centroids receive the same archetype label, we distinguish them as*higher distress*and*lower distress*according to their mean levels\.
#### Cross\-source comparison\.
For each source, we compute the proportion of conversations assigned to each archetype\. Treating the CTL distribution as a reference, we quantify how strongly each synthetic condition over\- or under\-represents particular recovery patterns relative to real crisis conversations \(Figure[5](https://arxiv.org/html/2609.29056#S6.F5)\)\.
#### Inter\-conversation Diversity\.
To test whether synthetic conversations are less varied than real CTL conversations, we measure inter\-conversation diversity at the trajectory level\. For each source \(synthetic zero\-shot, synthetic dual\-agent, real CTL\), we restrict to with\-context\-labeled conversations, retain texter turns only, and use the distress score\. Conversations with fewer than four valid distress values are discarded\. Each remaining conversation is summarized as a length\-normalized trajectory by linearly interpolating its distress sequence onto a shared grid ofT=20T=20equally spaced points on\[0,1\]\[0,1\], yielding one 20\-dimensional vector per conversation\. To keep the three sources on comparable footing, we subsample each source to at mostN=1000N=1000conversations \(seed4242\) and compute the full pairwise Euclidean distance matrix between trajectories within each source\. We summarize inter\-conversation diversity as the mean of the off\-diagonal entries of this matrix; a higher mean indicates that conversations within the corpus are more dissimilar from one another\. Uncertainty is quantified by a nonparametric bootstrap \(1,000 resamples\) over conversation indices, and we test whether CTL is significantly more diverse than each synthetic source by bootstrapping the difference of means and reporting the 95% confidence interval\.
## Appendix FAdditional Results
### F\.1Speaker Role Differentiation
With conversational context, texter and volunteer emotion profiles are sharply distinguished along complementary axes\. Texter utterances tend to remain in negative states, with negative\-polarity persistence of 84\.1%, whereas volunteer utterances more consistently remain in positive states, with positive\-polarity persistence of 72\.4%; the two roles therefore exhibit distinct emotion\-label profiles\. The volunteer’s most frequent cross\-emotion transition ishopeless→\\tohopeful, consistent with movement from acknowledging distress toward a more hopeful frame\. Texter labels are also more variable, with a per\-conversation emotion change rate of 62\.8% compared with 52\.8% for volunteers\. Texter conversations are additionally highly likely to begin in a negative state, with 90\.9% of first substantive texter emotions labeled negative\.
Without conversational context, this role differentiation largely collapses: persistence gaps narrow \(negative: 22\.0 pp→\\to12\.0 pp; positive: 8\.9 pp→\\to3\.1 pp\), dominant self\-transitions converge on low\-specificity labels \(neutral,sadness\), and both roles end positive at identical rates \(37\.6%\)\. These findings confirm that conversational context is essential for the detection model to distinguish speaker roles and recover the therapeutic arc; we therefore use the with\-context configuration for all primary analyses\. The full with\- vs\. without\-context comparison is provided in §[G](https://arxiv.org/html/2609.29056#A7)\.
### F\.2Additional Tables
Table[15](https://arxiv.org/html/2609.29056#A6.T15)reports per\-conversation emotion change rates\.
Table 15:Per\-conversation emotion change rate for the full CTL analysis corpus \(fraction of adjacent utterance pairs with different labels\)\.
## Appendix GWith\- vs\. Without\-Context Comparison
#### Contexts\.
In preliminary study to compare between with\- vs\. without\- context prediction, we test two inference modes on the full CTL analysis corpus:
- •Without context:The model receives only the target utterance\.
- •With context:The model receives a sliding window of theNNmost recent utterances \(from both speakers\) preceding the target, plus the target itself\.
Table[16](https://arxiv.org/html/2609.29056#A7.T16)presents the full comparison of texter and volunteer emotion dynamics under both context conditions\.
Table 16:Texter vs\. volunteer emotion dynamics across context conditions\.Δ\\Delta= Texter−\-Volunteer\. With context, the two roles are sharply differentiated; without context, the distinction largely collapses\.
#### With Context: Clear Differentiation\.
When conversational context is available, the gap in negative persistence between texters and volunteers is 22\.0 pp \(0\.848 vs\. 0\.628\), reflecting the fundamental asymmetry between the texter’s sustained distress and the counselor’s active redirection\. The top texter self\-transition ishopeless→\\tohopeless\(6,758 counts\), while for volunteers it ishopeful→\\tohopeful\(14,714 counts\)\. The starting polarity distributions also diverge sharply: 92\.2% of texter sequences begin negative versus only 51\.8% of volunteer sequences\.
#### Without Context: Role Collapse\.
Without conversational context, the model treats each utterance as an isolated text fragment, and the texter–volunteer distinction degrades substantially\. The negative persistence gap narrows from 22\.0 pp to 12\.0 pp \(0\.684 vs\. 0\.564\), and the positive persistence gap shrinks from 8\.9 pp to 3\.1 pp \(0\.318 vs\. 0\.349\)\. Many counselor messages \(e\.g\., “That sounds really hard,” “I can hear how much pain you’re in”\) read as negative when stripped of their empathetic, redirective function in the conversational flow\.
The change\-rate ordering inverts: with context, texter emotion labels change more often than volunteer labels \(63\.9% vs\. 56\.0%\); without context, volunteers show a slightly*higher*change rate \(86\.8% vs\. 84\.2%\), because topically diverse counselor utterances receive highly variable labels when classified in isolation\. The dominant self\-transitions converge on low\-specificity labels \(texter:neutral→\\toneutral, 2,558; volunteer:sadness→\\tosadness, 2,117\), and both roles end positive at identical rates \(37\.6%\), compared to a 2\.2 pp gap with context\.
Overall, without conversational context, the model treats each utterance in isolation, inflating the emotion change rate \(84\.2% vs\. 63\.9% with context\) and collapsing the texter–volunteer role distinction: both roles end positive at identical rates \(37\.6%\), and the model defaults toneutralfor contextually ambiguous utterances such as a volunteer’s empathetic reflection of distress\. With context, the model can sustain emotional labels across turns, capture the therapeutic arc \(92\.2% negative starts shifting to 63\.7% positive endings\), and distinguish the complementary roles of texter \(anchored in distress\) and volunteer \(anchored in hope\)\.
## Appendix HSynthetic Generation Conditions
We generate synthetic crisis\-support conversations using two complementary strategies\. In the*dual\-agent*setup, separate LLM\-backed agents take asymmetric Texter and Volunteer roles and generate the conversation turn by turn\. This paradigm is designed to encourage interactional emergence, allowing each agent to respond to the developing conversation\. In the*zero\-shot*setup, a single LLM generates a complete conversation in one pass, providing tighter control over the global arc and conversation structure\. No CTL messages, paraphrases, summaries, or other CTL\-derived materials were used in either generation setup, including in prompts or seeds\.
Across both setups, we use three frontier instruction\-tuned models from distinct model families:claude\-haiku\-4\-5\-20251001[Anthropic \(2025\)](https://arxiv.org/html/2609.29056#bib.bib8),gpt\-4o\-2024\-08\-06[OpenAI \(2024\)](https://arxiv.org/html/2609.29056#bib.bib7), andGemini\-2\.5\-pro[Google \(2025\)](https://arxiv.org/html/2609.29056#bib.bib9)\. These models span distinct commercial training pipelines and allow us to examine how model\-specific alignment, safety tuning, and stylistic priors affect generated crisis dialogues\.
To control topical coverage, we construct a seed inventory spanning two broad classes\. Grief\-related concerns include parental loss, pet loss, relationship dissolution, job loss, miscarriage, and identity loss\. Non\-grief concerns include anxiety, depression, academic stress, social anxiety, family conflict, workplace stress, self\-esteem, and LGBTQ\+ struggles\. Each generated texter is conditioned on a sampled seed and prompted to produce an opening message grounded in the assigned situation\. The volunteer role is prompted to provide empathetic listening, non\-directive questioning, and appropriately paced safety assessment\.
For the texter role, we adapt established prompts for LLM\-simulated patients[Louie et al\. \(2024\)](https://arxiv.org/html/2609.29056#bib.bib41);[Louie et al\. \(2025\)](https://arxiv.org/html/2609.29056#bib.bib6)\. Because no directly analogous prompt exists for the volunteer role, we design volunteer prompts that encode counselor\-oriented behaviors while maintaining asymmetric role grounding\.
We apply the sameEmpathpipeline to CTL and synthetic conversations, including utterance\-level emotion labeling, polarity mapping, transition analysis, distress trajectory construction, volunteer strategy classification, and archetype assignment\.相似文章
EmoTrace:一种以情感轨迹为中心的心理支持对话生成框架
该论文提出了EmoTrace,这是一个面向心理支持的多轮对话生成框架,通过对求助者情感轨迹进行建模,提升咨询师回应中的共情能力和情感丰富度,其性能优于现有方法。
超越感觉好转:能力维持型情感对话作为纵向研究范式
这篇arXiv论文提出将能力维持型情感对话(CSED)作为情感支持系统的纵向研究范式,认为当前方法侧重于即时缓解而忽视了用户的长期能力。文献审计显示,大多数系统忽略了纵向结果,从而为数据、模型、评估和治理提出了新的议程。
心理健康对话中的专家级危机检测
介绍了CRADLE-Dialogue,一个由临床医生标注的基准数据集,用于心理健康对话中的对话轮次级危机检测,同时包含Alert–Confirm评估协议、合成训练语料库以及一个32B参数模型,该模型在性能上优于现有的开放源代码和专有模型。
用于共情回应的动态常识协调
提出了DCC,一个用于共情回应生成的动态常识协调框架,集成了基于残差的交互、关联引导的过滤和迭代解码,相比于基线模型提高了情感分类准确率和回应多样性。
共情作为预测性错位容忍:一种协同调节框架与对话修复的机制结构
本文重新定义了AI对话系统中的共情为'预测性错位容忍',并提出了一种解释性错误容忍(IET)启发式方法。实验表明,对话修复具有依赖状态的机制结构,在不同噪声水平下在判别保真度与主旨保留之间进行权衡。