Emotional Labor Strategy Preferences in LLM Personas

arXiv cs.CL Papers

Summary

The study investigates how large language models with personality personas exhibit preferences for emotional labor strategies, finding that models align more with deep acting and that personality traits like Conscientiousness and Emotional Stability predict this preference.

arXiv:2609.00310v1 Announce Type: new Abstract: Emotional labor is the effortful management of emotional displays to meet social or professional expectations. Personality traits have been correlated with emotional labor strategies, yet research on this link relies almost exclusively on self-report scales administered only in occupational settings. We investigate whether large language models injected with psychometrically grounded personas reproduce these personality-driven selection patterns across everyday social scenarios. We construct the first emotional labor strategy dataset of 500 socially situated events, each offering three behavioral choices corresponding to surface acting, deep acting, and genuine expression. We source 50 fictional characters from a large-scale personality repository and profile each through two parallel tracks: observer-rated bipolar adjective composites and in-character self-report items. Five LLMs evaluate all scenarios under both persona conditions. We find that models align more towards deep acting, and that Conscientiousness and Emotional Stability consistently predict this preference. Entropy analysis confirms that persona reliably influences the output and varies across models and emotions.
Original Article
View Cached Full Text

Cached at: 09/02/26, 05:49 AM

# Emotional Labor Strategy Preferences in LLM Personas
Source: [https://arxiv.org/html/2609.00310](https://arxiv.org/html/2609.00310)
Tianyu JiangAffiliation:saimmd@mail\.uc\.edu, tianyu\.jiang@uc\.edu

###### Abstract

Emotional labor is the effortful management of emotional displays to meet social or professional expectations\. Personality traits have been correlated with emotional labor strategies, yet research on this link relies almost exclusively on self\-report scales administered only in occupational settings\. We investigate whether large language models injected with psychometrically grounded personas reproduce these personality\-driven selection patterns across everyday social scenarios\. We construct the first emotional labor strategy dataset of 500 socially situated events, each offering three behavioral choices corresponding to surface acting, deep acting, and genuine expression\. We source 50 fictional characters from a large\-scale personality repository and profile each through two parallel tracks: observer\-rated bipolar adjective composites and in\-character self\-report items\. Five LLMs evaluate all scenarios under both persona conditions\. We find that models align more towards deep acting, and that Conscientiousness and Emotional Stability consistently predict this preference\. Entropy analysis confirms that persona reliably influences the output and varies across models and emotions\.

![Refer to caption](https://arxiv.org/html/2609.00310v1/Intro.png)Figure 1:An example of emotional labor categories and what persona profiles choose from each category\.## 1Introduction

The way individuals manage and display emotions in social contexts constitutes a central object of study across psychology, organizational behavior, and, more recently, computational linguistics\.[Hochschild \(2012\)](https://arxiv.org/html/2609.00310#bib.bib3)introduced the concept of*emotional labor*\(EL\) to describe the effortful management of feelings to produce a publicly observable display that conforms to role expectations\. Since then, researchers have consistently identified three principal strategies through which individuals execute this management \(Figure[1](https://arxiv.org/html/2609.00310#S0.F1)\):*surface acting*\(SA\), in which outward expression is modified without changing inner feeling;*deep acting*\(DA\), in which the felt emotion itself is reappraised so that it is followed by a genuine display of emotions; and*genuine expression*\(GE\) or naturally felt expressions, in which experienced emotion already matches what the situation calls for, and no regulatory effort is required\([Ashforth and Humphrey, 1993](https://arxiv.org/html/2609.00310#bib.bib4);[Diefendorff et al\., 2005](https://arxiv.org/html/2609.00310#bib.bib5);[Grandey, 2000](https://arxiv.org/html/2609.00310#bib.bib6)\)\. The mismatch between the regulated exterior and the persisting interior is what separates SA from the other two strategies, since the effort is spent on the display itself rather than on the feeling\. In Genuine Expression, the felt emotion already matches what the situation calls for, so no internal adjustment occurs before the display\. In Deep Acting, the felt emotion is actively reshaped through reframing, perspective\-taking, or self\-talk, so that regulation happens upstream of expression rather than at the surface\. These three strategies have been employed in empirical work that shows they have different consequences for well\-being, service quality, and interpersonal perception\([Kammeyer\-Mueller et al\., 2013](https://arxiv.org/html/2609.00310#bib.bib10);[Grandey, 2003](https://arxiv.org/html/2609.00310#bib.bib7);[Groth et al\., 2009](https://arxiv.org/html/2609.00310#bib.bib11)\)\.

A well\-established line of research in organizational psychology shows that personality traits can shape how people perceive and classify others’ emotional displays\([Terracciano et al\., 2003](https://arxiv.org/html/2609.00310#bib.bib31);[Furnes et al\., 2019](https://arxiv.org/html/2609.00310#bib.bib34)\)\. These findings draw on the Big Five \(OCEAN\) framework—Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism\([Goldberg, 1992](https://arxiv.org/html/2609.00310#bib.bib14)\), which is an extensively validated taxonomy of human personality differences\([McCrae and John, 1992](https://arxiv.org/html/2609.00310#bib.bib15)\)\. Individuals high in Agreeableness and Extraversion tend to perceive and enact more genuine or deep\-acting strategies, whereas those high in Neuroticism are more prone to surface acting\([Diefendorff et al\., 2005](https://arxiv.org/html/2609.00310#bib.bib5);[Austin et al\., 2008](https://arxiv.org/html/2609.00310#bib.bib12);[Kiffin\-Petersen et al\., 2011](https://arxiv.org/html/2609.00310#bib.bib13)\)\. However, researchers have largely assessed patterns of trait\-linked EL preferences using self\-report questionnaires/scales, in which participants rate statements about individual EL categories on Likert scales\. Moreover, most sampling methods for analyzing EL strategies have been limited to specific healthcare and customer service contexts and have not been tested in everyday situations\.

A growing body of work shows that LLMs, when prompted with personality descriptions, exhibit behavioral patterns that align with their assigned traits\([Jiang et al\., 2024](https://arxiv.org/html/2609.00310#bib.bib16);[Sorokovikova et al\., 2024](https://arxiv.org/html/2609.00310#bib.bib19);[Li et al\., 2024](https://arxiv.org/html/2609.00310#bib.bib33)\)\. These injected*personas*can modulate lexical choices, sentiment patterns, and judgment tendencies in ways that mirror the psychological literature on human personality\([Peters and Matz, 2024](https://arxiv.org/html/2609.00310#bib.bib20);[Matz et al\., 2024](https://arxiv.org/html/2609.00310#bib.bib21)\)\. This provides a reliable substrate for controlled experimental simulation of human psychological variability\.

Despite parallel advances in EL theory and LLM persona research, these two lines of work have not yet converged\.No prior study has examined how personality\-injected LLM personas select among emotional labor strategies in socially situated scenarios\.If particular personality traits systematically shift which strategy a persona selects, then persona injection introduces a source of behavioral variance that may go unnoticed in downstream applications that rely on LLM\-generated responses\. The stakes of emotion\-related misclassification are amplified by the growing use of LLMs in emotional support conversations, customer\-service interactions, and everyday affect\-sensitive domains, where inappropriate response or interpretation may negatively affect user well\-being, communication outcomes, and even sway ethical judgments\([Grandey and Sayre, 2019](https://arxiv.org/html/2609.00310#bib.bib24);[Weidinger et al\., 2021](https://arxiv.org/html/2609.00310#bib.bib22);[Wu et al\., 2025](https://arxiv.org/html/2609.00310#bib.bib23);[Saim and Jiang, 2026](https://arxiv.org/html/2609.00310#bib.bib47)\)\.

In this work, we address this gap by building on an existing appraisal corpus to construct an emotional labor strategy \(ELS\) dataset of 500 sentences, each with annotated event descriptions followed by three choices of how an agent would enact them based on the three core strategies\. In parallel, we employ fictional characters from TV/Film shows with character personality ratings on bipolar adjective pairs\. To generalize and compare the persona profiles, we also administer the IPIP\-50 questionnaire in\-character, following established procedures for extracting Big Five scores from validated instruments\([Goldberg et al\., 2006](https://arxiv.org/html/2609.00310#bib.bib27);[Jiang et al\., 2024](https://arxiv.org/html/2609.00310#bib.bib16)\)\. We load the resulting persona profile into multiple LLMs for evaluation on the curated ELS dataset\. For each sentence, the persona must select one of three options: surface acting \(SA\), deep acting \(DA\), or genuine expression \(GE\)\. We then analyze EL strategy choice patterns across the OCEAN dimensions of each persona, mapping which traits drive each persona’s selection\. We publicly release the code and the emotion labor strategy dataset\.111[https://github\.com/cincynlp/emotion\-labor](https://github.com/cincynlp/emotion-labor)As an overview, this paper makes the following contributions:

1. 1\.We introduce the first study to frame emotional labor strategy selection as a personality\-conditioned task for LLM personas, in which each persona chooses how it would respond behaviorally to a social scenario, and to compare these choices against previously established human trends\.
2. 2\.We build and release the first full\-scale dataset on emotional labor strategies, comprising 500 scenarios across everyday affect and social contexts that involve choices among surface acting, deep acting, and genuine expressions\.
3. 3\.We show that models prefer deep acting as the dominant strategy, and that Conscientiousness and Emotional Stability emerge as the consistent trait\-level predictors across both persona tracks\.

## 2Related Works

#### Emotions in NLP and Emotional Labor\.

The computational study of emotion in text has matured considerably over the past decade\. Early work cast emotion recognition as a categorical classification problem, mapping text to discrete labels derived from foundational theories\([Strapparava and Mihalcea, 2007](https://arxiv.org/html/2609.00310#bib.bib28);[Mohammad et al\., 2018](https://arxiv.org/html/2609.00310#bib.bib36);[Hofmann et al\., 2020](https://arxiv.org/html/2609.00310#bib.bib17)\)\. More theoretically grounded approaches have moved beyond flat emotion labels toward appraisal\-based frameworks, which treat emotion as the product of a person’s cognitive evaluation of an event along dimensions such as novelty, relevance, and coping potential\([Smith and Ellsworth, 1985](https://arxiv.org/html/2609.00310#bib.bib37);[Scherer, 2001](https://arxiv.org/html/2609.00310#bib.bib38)\)\.[Troiano et al\. \(2023\)](https://arxiv.org/html/2609.00310#bib.bib39)formalize this shift in a comprehensive corpus study, annotating event descriptions and showing that appraisal variables reliably improve emotion categorization in text\. This line of work is directly relevant to ours: appraisal theory frames emotion as a*situated evaluation*of a social event, which maps naturally onto the idea that emotional labor strategies represent different stances a person takes toward a situation\. Another line of work in emotion recognition evaluates or applies embodied or physiological indicators rather than explicit emotion labels\([Zhuang et al\., 2024](https://arxiv.org/html/2609.00310#bib.bib44);[Duong et al\., 2025](https://arxiv.org/html/2609.00310#bib.bib45);[Saim et al\., 2025](https://arxiv.org/html/2609.00310#bib.bib46)\)\. This embodied perspective aligns with our dataset design, since the strategy options involve behavioral or physiological details\.

In organizational psychology, measurement of emotional labor \(EL\) has relied almost exclusively on self\-report instruments administered to workers in service roles\.[Brotheridge and Lee \(2003\)](https://arxiv.org/html/2609.00310#bib.bib40)developed the Emotional Labour Scale, a 15\-item questionnaire that assesses the frequency, intensity, variety, duration, and both surface and deep acting \(SA and DA\); this instrument has become one of the most widely adopted tools in the field\.[Diefendorff et al\. \(2005\)](https://arxiv.org/html/2609.00310#bib.bib5)later extended this framework to a three\-factor structure that treats naturally felt expression as a dimensionally distinct strategy alongside SA and DA, thereby establishing the taxonomy used in our study\. A core gap in research on EL strategies is the lack of a corpus that goes beyond narrow professional roles such as customer service and healthcare workers\. This leaves open the question of whether the three\-strategy taxonomy applies to the broader, non\-occupational interactions that constitute most of daily social life\.

#### Persona Injection and Personality in LLMs\.

A rapidly growing body of NLP work has examined whether LLMs can stably express and enact personality traits when prompted with persona descriptions\.[Jiang et al\. \(2024\)](https://arxiv.org/html/2609.00310#bib.bib16)showed that LLM personas assigned Big Five profiles produce BFI self\-report scores and writing samples that align with their designated traits, with large effect sizes across all five dimensions\.[Sorokovikova et al\. \(2024\)](https://arxiv.org/html/2609.00310#bib.bib19)replicated this pattern across multiple open\-source models, and[Tseng et al\. \(2024\)](https://arxiv.org/html/2609.00310#bib.bib41)provides a comprehensive taxonomy of the field, distinguishing role\-playing personas \(where LLMs adopt assigned identities\) from personalization settings where models adapt to user traits\.

On persona traits,[Huang and Hadfi \(2024\)](https://arxiv.org/html/2609.00310#bib.bib43)found that OCEAN\-grounded LLM agents reproduce human\-like personality\-linked patterns in bilateral negotiations, with Agreeableness promoting cooperative outcomes and Neuroticism leading to less favorable outcomes\. Agreeableness and Extraversion consistently predict greater deep acting and genuine expression, while Neuroticism predicts surface acting\([Austin et al\., 2008](https://arxiv.org/html/2609.00310#bib.bib12);[Kiffin\-Petersen et al\., 2011](https://arxiv.org/html/2609.00310#bib.bib13);[Yeh et al\., 2020](https://arxiv.org/html/2609.00310#bib.bib42)\)\. Recent mechanistic work finds that Big\-Five traits are encoded as recoverable linear directions in LLM residual streams, and that steering along these directions produces reliable, trait\-consistent behavioral shifts\([Frising and Balcells, 2026](https://arxiv.org/html/2609.00310#bib.bib32)\)\. Taken together, this literature establishes that personality injection is a reliable lever on LLM behavior\. However, most measurement work is limited to self\-reported behavior within occupational samples\. No work examines which trait\-conditioned LLMs select as emotional regulation strategies when evaluated in socially situated EL scenarios\. We fill this gap by presenting OCEAN\-grounded personas with EL\-annotated stimuli and measuring whether their strategy choices mirror the personality\-EL associations documented in human literature and surveys\.

## 3Methodology

We begin by describing the emotional labor strategy \(ELS\) dataset, followed by the pipeline for persona traits\. We then evaluate the characters, loaded with distinct traits, on the ELS dataset to analyze how each character’s strategy preferences vary across emotional labor scenarios\.

### 3\.1Emotion Labor Strategy Dataset

We construct the ELS dataset by building on the emotion appraisal corpus\([Troiano et al\., 2023](https://arxiv.org/html/2609.00310#bib.bib39)\), a collection of sentences annotated with appraisal dimensions and a single felt emotion label drawn from seven categories:anger,disgust,fear,guilt,joy,sadness, andshame\. These labels reflect the narrator’s internal affective state\. We preserve the original dataset’s uniform emotion split and filter to a final set of 500 sentences\.

#### Social Context Augmentation\.

Many sentences in the corpus describe internal affective states without situating them in a social encounter\. Since emotional labor is fundamentally an interpersonal phenomenon, i\.e\., it arises when a person regulates their emotional display in the presence of or in response to another social agent\([Hochschild, 2012](https://arxiv.org/html/2609.00310#bib.bib3);[Grandey, 2000](https://arxiv.org/html/2609.00310#bib.bib6)\), we add a social agent or context to each scenario\. We augment each sentence with a brief context that introduces a second agent and coherent background information that creates a plausible reason to regulate one’s emotional expression\. The context is appended directly to the original sentence to produce a single, continuous narrative that serves as the scenario stem for all downstream steps\.

#### Emotional Labor Strategy Options\.

With the social context in place, each scenario presents a person experiencing a felt emotion in the presence of another agent\. We extend each scenario into a three\-way choice item by generating one behavioral response option for each of the three emotional labor strategies:surface acting\(SA\),deep acting\(DA\), andgenuine expression\(GE\) or naturally felt emotions\. Table[1](https://arxiv.org/html/2609.00310#S3.T1)defines each of the three emotional labor strategies represented in the dataset\. We employ the GPT\-5\.4 model to generate the context and the three\-way choices\. No single correct response or ground truth exists, as the task is designed to elicit apreferencefor each persona\.

StrategyDefinitionSurface Acting \(SA\)The person performs a different emotion outwardly while the original feeling persists internally\.Deep Acting \(DA\)The person genuinely shifts their internal state through reframing, perspective\-taking, or self\-talk before responding\.Genuine Expression \(GE\)The person expresses the naturally felt emotion directly with no attempt to regulate or conceal it\.Table 1:Definitions of the three emotional labor strategies used in dataset construction\.The generation produces all three options as single sentences that complete the scenario stem naturally, without using any emotion labels or adjectives, conveying emotional state only through behavioral, physical, or physiological detail\. Table[2](https://arxiv.org/html/2609.00310#S3.T2)shows an example set from the dataset\. The final dataset item consists of the augmented scenario and three behavioral completion options\. The final dataset was manually validated to \(i\) verify coherence between the added context and the scenario narrative, \(ii\) determine whether each option instantiated its strategy without using emotion words, and \(iii\) determine whether the three options were distinguishable\. We took several steps to mitigate concerns about subjective bias\. The strategy generation was tightly constrained by the prompts \(Appendix[A](https://arxiv.org/html/2609.00310#A1)\), which enforced structural rules such as prohibiting emotional adjectives, requiring a single leakage cue for Surface Acting, and mandating behavioral rather than affective descriptions\. This offloads much of the strategy\-fidelity check from human judgment onto the generation constraints themselves\. Beyond its use in this study, the dataset is the first resource to use emotion labor categories as contrastive behavioral options in naturalistic, non\-occupational scenarios and can support future work on emotion regulation modeling, personality\-conditioned generation, and computational appraisal research\.

Example ScenarioI feltfearwhen I was being bullied at work and didn’t think there was going to be an end to it\.A newer coworker was standing nearby, watching how I would respond\.SA:I lifted my chin and gave a short laugh, but my hands shook once against the desk as I answered\.DA:I reminded myself their words were about their own problems, slowed my breathing, and answered in an even voice with my shoulders settling\.GE:I stepped back, gripped the edge of the desk, and said they needed to stop, my voice breaking in the middle\.Table 2:An example scenario from the dataset\. Theitalicizedportion shows the social context added to the original corpus\.

### 3\.2Persona Construction

The design of a persona profile with only a character’s name and source work risks activating whatever personality representation the model has internalized during pretraining, which may be inconsistent across models\. Therefore, we construct persona profiles through two parallel experimental tracks\. Both tracks use the same set of characters but differ in how personality information is sourced\. This design allows us to \(i\) avoid any bias from relying on a single framework and \(ii\) examine whether trait\-strategy associations are aligned across two distinct representations of personality\.

#### Character Pool and BAP Profiling\.

We source character personality data from the Open Psychometrics Statistical “Which Character” Personality Quiz dataset\([Open\-Source Psychometrics Project, 2022](https://arxiv.org/html/2609.00310#bib.bib25)\), which contains crowd\-sourced personality ratings for over 2,000 fictional characters across film and television\. Each character is rated by human respondents on 500 Bipolar Adjective Pairs \(BAPs\), a psycho\-lexical instrument in which each item presents two semantically opposing adjectives \(e\.g\.,relaxed–tense,kind–cruel\) and human raters place the character on a continuous scale of 1–100\. Since using every BAP as a trait in the model might introduce too much noise for the LLM to process, we select BAPs based on standard deviation\.

To ground our trait markers in a theoretically validated source, we verify each BAP adjective against the original adjective list reported by[Goldberg \(1992\)](https://arxiv.org/html/2609.00310#bib.bib14)\. We retain only those adjective pairs for which at least one pole produces an exact match with Goldberg’s taxonomy\. This filtering yields a subset of 40 verified BAP markers, which ensures that every trait dimension we use is validated and has a well\-established structural relationship to the Big Five OCEAN dimensions\.

We filter characters from the full pool by ranking them on the standard deviation of their BAP scores\. Characters with high score variance across adjective pairs have more personality\-differentiated profiles, making them better suited to testing trait\-driven behavioral differences\. We then apply greedy farthest\-point sampling in the 40\-dimensional BAP space, iteratively selecting the character that is maximally distant from all previously selected characters\. This produces a final set of 50 characters\. We chose 50 characters by prioritizing personality diversity over sample size in the 40\-item BAP space\. This method selects characters that are maximally spread across the personality space rather than clustering around common archetypes\. The overall BAP\-persona block for each character consists of the character name, their source work, and their human\-rated scores on the 40 verified BAP items, each expressed on a 1–100 scale\. This block is passed directly to the model as part of the persona injection prompt in the ELS evaluation task\.

#### IPIP\-50 Profiling\.

Relying solely on human\-rated BAP scores may underrepresent traits that are internally characterized\. The second track addresses this by eliciting in\-character self\-reports on a validated self\-report instrument\.[Vazire \(2010\)](https://arxiv.org/html/2609.00310#bib.bib35)presented the self\-other knowledge asymmetry \(SOKA\) model, which shows that observer ratings more reliably capture visible, behaviorally expressed traits such as extraversion, while self\-reports capture internal states and motivations that are less accessible to outside observers, such as neuroticism and openness\. The IPIP\-50 is a 50\-item public\-domain scale with 10 items per OCEAN dimension\([Goldberg et al\., 2006](https://arxiv.org/html/2609.00310#bib.bib27)\)\. We administer it to each model with the character’s name and source work and ask it to rate each IPIP item on a 1–5 Likert scale from the character’s perspective \(1 =Very Inaccurate, 5 =Very Accurate\)\. The character pool remains the same across the BAP\-persona to ensure the analysis focuses solely on methodological differences\. The full list of filtered characters, BAP terms, and IPIP items is listed in Appendix[B](https://arxiv.org/html/2609.00310#A2)\.

### 3\.3Evaluation and Model Selection

We end up with two parallel persona experimentation tracks\. A short summary is given in Table[3](https://arxiv.org/html/2609.00310#S3.T3)\. These are evaluated on the 500 sentences of the ELS dataset\. We primarily study ELS choice patterns across different models and their correlations with the OCEAN categories\. We employ a suite of five LLMs: Qwen\-3\-8B and Qwen\-3\-32B\([Yang et al\., 2025](https://arxiv.org/html/2609.00310#bib.bib30)\), GPT\-5\.4\([OpenAI, 2026](https://arxiv.org/html/2609.00310#bib.bib26)\), Deepseek\-V4\-Flash\([DeepSeek\-AI, 2026](https://arxiv.org/html/2609.00310#bib.bib29)\)and Gemma\-4\-31B\([Google DeepMind, 2026](https://arxiv.org/html/2609.00310#bib.bib18)\)\. We prompt these models to choose one of the three EL strategies: SA, DA, or GE\.

Track\-1 \(BAP\)Track\-2 \(IPIP\-50\)SourceHuman\-rated crowd scoresModel self\-report in\-characterInstrument40 bipolar adjective pairs50 standard itemsScale1–100 continuous1–5 LikertItem Examplerelaxed–tense“I am the life of the party”

Table 3:Comparison of the two persona tracks\.Figure 2:Aggregate distribution of emotion labor strategy across models and persona tracks\.

## 4Results and Analysis

We present the findings of our study across three levels of analysis\. We begin by characterizing the aggregate distribution of ELS selections across all five models and both persona\-induction tracks \(BAP and IPIP\-50\)\. We then report Big Five trait–strategy correlations that connect persona\-level personality profiles to ELS preferences and draw comparisons to organizational psychology benchmarks\. Finally, we study the impact of personas on the reliability with which they drive ELS scenarios\.

### 4\.1Aggregate ELS Distributions Across Models and Persona Tracks

#### Deep Acting is the preferred response compared to SA and GE\.

We first examine how each of the five models distributes its ELS selections across 500 scenarios when conditioned on 50 character personas\. Figure[2](https://arxiv.org/html/2609.00310#S3.F2)reports the proportion of each ELS category selection for each model under both the BAP and IPIP\-50 persona tracks\.

Notably, DA is the modal strategy for eight of the ten model\-track combinations\. DA proportions range from approximately 39% \(GPT\-5\.4, IPIP track\) to 61% \(Qwen\-3\-32B, both tracks\)\. This establishes DA as the dominant default across the model pool\. This preference casts the majority of diverse agents as effortful and authentic\. Models trained on large corpora of psychological and organizational text appear to encode DA as the prototypically appropriate regulatory response\. The encoding appears to serve as a prior that shapes strategy selection beyond any personality variation introduced by the injected persona\. In contrast, SA selections remain low across all models\. In psychology literature, SA has consistently been identified as a prevalent emotional labor strategy in service\-oriented occupations, especially when workers experience emotional dissonance between authentic feelings and institutionally required emotional displays\([Grandey, 2000](https://arxiv.org/html/2609.00310#bib.bib6);[Mesmer\-Magnus et al\., 2012](https://arxiv.org/html/2609.00310#bib.bib9)\)\. We hypothesize that models suppress the SA strategy not because persona profiles steer them away from it, but because SA may carry a socially undesirable signal in natural language \(emotion suppression\)\.

ReliabilityBAP \(Observer\-Rated\)IPIP \(Self\-Report\)Trait𝜶BAP\\boldsymbol\{\\alpha\_\{\\text\{BAP\}\}\}𝜶IPIP\\boldsymbol\{\\alpha\_\{\\text\{IPIP\}\}\}SADAGESADAGEOpenness0\.700\.95\+0\.064\+0\.064−0\.013\-0\.013−0\.032\-0\.032−0\.260\-0\.260\+0\.121\+0\.121−0\.078\-0\.078Conscientiousness0\.880\.98−0\.446∗⁣∗\-0\.446^\{\*\*\}\+0\.597∗∗∗\+0\.597^\{\*\*\*\}−0\.570∗∗∗\-0\.570^\{\*\*\*\}−0\.345∗\-0\.345^\{\*\}\+0\.649∗∗∗\+0\.649^\{\*\*\*\}−0\.592∗∗∗\-0\.592^\{\*\*\*\}Extraversion0\.830\.98\+0\.110\+0\.110−0\.203\-0\.203\+0\.184\+0\.184\+0\.200\+0\.200−0\.299∗\-0\.299^\{\*\}\+0\.232\+0\.232Agreeableness0\.950\.97−0\.276\-0\.276\+0\.144\+0\.144−0\.072\-0\.072−0\.398∗⁣∗\-0\.398^\{\*\*\}\+0\.372∗⁣∗\+0\.372^\{\*\*\}−0\.263\-0\.263Emotional Stability0\.760\.97−0\.350∗\-0\.350^\{\*\}\+0\.522∗∗∗\+0\.522^\{\*\*\*\}−0\.517∗∗∗\-0\.517^\{\*\*\*\}−0\.311∗\-0\.311^\{\*\}\+0\.850∗∗∗\+0\.850^\{\*\*\*\}−0\.883∗∗∗\-0\.883^\{\*\*\*\}Table 4:Spearman correlations between OCEAN trait composites and emotional labor strategy proportions for the BAP and IPIP tracks\. Cronbach’sα\\alphais reported for BAP and IPIP composites\.∗p<\.05\{\}^\{\*\}p<\.05,p∗⁣∗<\.01\{\}^\{\*\*\}p<\.01,∗∗∗p<\.001\{\}^\{\*\*\*\}p<\.001\.
#### BAP and IPIP\-50 tracks show moderate convergence\.

Across the model pool, both persona tracks preserve the same ordinal ranking of strategies \(DA\>\>GE\>\>SA in most cases\), providing a baseline level of convergent validity between the observer\-rated and self\-reported persona inductions\. However, quantitative differences between tracks are meaningful and model\-dependent\. For GPT\-5\.4 and DeepSeek\-V4\-Flash, the IPIP\-50 track redistributes mass from DA toward GE relative to the BAP track: GPT\-5\.4 drops from approximately 51% to 39% on DA while GE rises from approximately 30% to 40%\. DeepSeek\-V4\-Flash shows a smaller but directionally consistent shift\. This pattern accords with the self\-other agreement work, which documents that observer\-rated and self\-reported personality tap partially distinct variance in the same underlying constructs\([Vazire, 2010](https://arxiv.org/html/2609.00310#bib.bib35);[Connelly and Ones, 2010](https://arxiv.org/html/2609.00310#bib.bib8)\)\. Observer ratings \(BAP track\) emphasize behaviorally visible regularities and external presentation\. They may therefore bias the injected persona toward more deliberate, socially regulated strategies such as DA\. Self\-report ratings \(IPIP\-50 track\) introduce a first\-person affective framing that may allow greater expression of internal states\.

We also computed pairwise inter\-model agreement on both evaluation tracks\. The highest agreement is observed between GPT\-5\.4 and Gemma\-4\-31B, with Cohen’sκ=0\.545\\kappa=0\.545; per\-category agreement is higher for GE and DA but substantially lower for SA\. Agreement scores between all models are provided in Appendix[C\.5](https://arxiv.org/html/2609.00310#A3.SS5)\.

### 4\.2Trait–Strategy Correlations

We study how personality shapes ELS selection by computing Spearman rank correlations between each Big Five composite and each strategy proportion across the 50 evaluated characters\. Each BAP and IPIP item is grouped under one of the OCEAN traits as shown in Table[7](https://arxiv.org/html/2609.00310#A3.T7)and Table[8](https://arxiv.org/html/2609.00310#A3.T8)\. The central question we examine is whether each OCEAN trait correlates similarly with the three ELS, as established in previous studies, and whether the two persona tracks yield convergent or divergent mappings\. We report results for GPT\-5\.4 \(Table[4](https://arxiv.org/html/2609.00310#S4.T4)\); the model comparison to Deepseek\-V4\-Flash can be found in Appendix[C\.1](https://arxiv.org/html/2609.00310#A3.SS1)\.

Two traits exhibit consistent, significant correlations across both measurement tracks:Conscientiousness \(C\)andEmotional Stability \(ES\)\(the inverse of Neuroticism\)\. Both traits show the same directional pattern, i\.e\., high C and high ES predict more DA and less SA and GE\. For Conscientiousness, the BAP track yieldsρ\\rho=−\-0\.446 \(SA\), \+0\.597 \(DA\),−\-0\.570 \(GE\), all significant atp<\.001p<\.001orp<\.01p<\.01, and the IPIP track replicates this pattern with comparable magnitude\. The same signature holds for ES\. The model appears to encode DA as the disciplined, regulated response\. The characters who are organized, reliable, and emotionally secure preferentially enact deep acting rather than surface performance or unguarded expression\. This is consistent with the general notion reported in human psychological studies that conscientiousness and emotional stability both correlate negatively with SA and positively with DA\([Austin et al\., 2008](https://arxiv.org/html/2609.00310#bib.bib12);[Judge et al\., 2009](https://arxiv.org/html/2609.00310#bib.bib48);[Mesmer\-Magnus et al\., 2012](https://arxiv.org/html/2609.00310#bib.bib9)\)\.

We also report a negative correlation between conscientiousness and genuine expression, whereas previous human studies report a positive correlation, on the theory that workers with higher conscientiousness scores internalize job demands and naturally align their felt emotions with role requirements\. They do not need to regulate because their internal state already matches the display norm\. We hypothesize that the model encodes conscientiousness not as internalized role alignment but as an active, effortful self\-regulation, much like a careful worker who writes a draft before sending an email when a spontaneous reply would do\. We plan to form our future work based on mechanistic insights from open\-source models to validate this finding, following work that recovers independent social\-response directions in LLM representations\([Yao et al\., 2026](https://arxiv.org/html/2609.00310#bib.bib1);[Vennemeyer et al\., 2026](https://arxiv.org/html/2609.00310#bib.bib2)\)\.

The BAP track returns no significantAgreeableness \(A\)effects, while the IPIP track produces A–SA \(ρ\\rho=−\-0\.398,p<\.01p<\.01\) and A–DA \(ρ\\rho= \+0\.372,p<\.01p<\.01\)\. Agreeableness is a trait humans understand better from the inside; warm, cooperative intentions are more readily disclosed in self\-report than inferred by external observers from behavioral adjectives\([Vazire, 2010](https://arxiv.org/html/2609.00310#bib.bib35)\)\. In this reading, the IPIP persona activates an agreeableness signal that the observer\-rated BAP does not encode with sufficient specificity\.Opennessproduces no significant correlations in either track\. Extraversion yields only one marginal IPIP effect with DA, suggesting that more extroverted personas are slightly less inclined toward deep acting\.

We also determine whether an observed trait–strategy association reflects a genuine pattern or a statistical\-only result of the response format \(three\-way choice\)\. Therefore, we replicate the evaluation by asking the model to rate each of the three EL categories based on how likely they are to choose that scenario, rather than forcing it\. The details and correlations are given in Appendix[C\.2](https://arxiv.org/html/2609.00310#A3.SS2)\. We find that all correlations broadly follow the same directional pattern\. Convergence between a three\-way choice and Likert correlations strengthens the claim that observed patterns reflect personality\-driven strategy preferences\.

#### Internal consistency\.

We also validate our findings by assessing the reliability of each trait composite\. Cronbach’sα\\alphameasures the extent of intercorrelation of items assigned to a given trait scale\. A highα\\alphameans multiple BAP or IPIP items under one of the OCEAN traits pull in the same direction, and the composite score is a stable summary of the underlying construct\. BAP composites yield moderate\-to\-high reliability \(α\\alpha= 0\.70–0\.95\), with conscientiousness and agreeableness showing the strongest internal consistency\. IPIP\-50 composites achieve uniformly high reliability across all five traits\. This also explains the ES–DA \(ρ\\rho= \+0\.850\) and ES–GE \(ρ\\rho=−\-0\.883\) large correlations\.

![Refer to caption](https://arxiv.org/html/2609.00310v1/heatmap.png)Figure 3:Persona influence by emotion and model \(BAP track\) depicted by Shannon Entropy\.

### 4\.3Is Persona Influential Enough?

#### Entropy Analysis\.

A distributional dominance result does not tell us whether personality drives meaningful divergence in strategy selection, or whether the emotional situation itself overrides individual differences\. To test this, we compute the Shannon entropy of the SA/DA/GE distribution across all 50 characters for each scenario\. We bound scores between 0 \(scenario\-dominant\) and 1 \(personality\-sensitive\), and report mean normalized entropy aggregated by the felt emotion category from the dataset\. We report BAP\-track results here to analyze human\-rated divergence\. Character\-level entropy traits are discussed in Appendix[C\.3](https://arxiv.org/html/2609.00310#A3.SS3)\. IPIP entropy results appear in Appendix[C\.4](https://arxiv.org/html/2609.00310#A3.SS4)\.

Figure[3](https://arxiv.org/html/2609.00310#S4.F3)reveals that most models across each emotion category form a high\-entropy cluster, indicating that personality meaningfully differentiates strategy selection in these models\. Within the high\-entropy clusters, Joy, Fear, and Anger sustain high entropy \(0\.82–0\.92\), while Disgust and Guilt sit slightly lower\. For Qwen\-32B, entropy ranges from 0\.42 \(Sadness\) to 0\.70 \(Disgust\), which shows that the particular emotion and scenario also matter more thanwhothe persona is for this model\.

#### Ablation\.

We further probe the drivers of strategy selection by running a three\-condition ablation on a 50\-scenario subset spanning all 50 characters\. In the vanilla condition, we provide only the scenario and the three response options with no character identity or trait information\. In the character\-only condition, we inject the character’s name and source work so the model relies entirely on its pretraining knowledge\. In the BAP\-only condition, we inject only the 40 adjective\-pair scores without naming any character\. Table[5](https://arxiv.org/html/2609.00310#S4.T5)reports aggregate strategy proportions across all three conditions\.

StrategyVanillaCharacter\-OnlyBAP\-OnlySurface Acting10\.0%29\.5%20\.5%Deep Acting68\.0%36\.6%36\.9%Genuine Expression22\.0%34\.0%42\.6%

Table 5:Strategy proportions under vanilla \(no persona\), name\-only, and BAP trait\-only persona injection, evaluated on a 50\-scenario subset across 50 characters \(GPT\-5\.4\)\.Without any persona signal, deep acting has the majority share of predictions while surface acting falls to just 10%\. This concentration confirms that DA serves as a strong default encoded in the model’s training distribution, rather than a product of persona conditioning\. Introducing a persona of any kind redistributes this mass\. Character\-only prompting cuts DA predictions by almost half and raises both SA and GE to near\-uniform levels, which suggests that even name\-level identity priming is sufficient to disrupt the DA default and provide a wider range of regulatory responses\. BAP\-only prompting produces a comparable DA proportion \(36\.9%\) but with higher GE numbers\. Trait content, therefore, does not simply diversify strategy selection the way name recognition does\. It selectively suppresses surface acting, consistent with the trait–strategy correlations reported in Section[4\.2](https://arxiv.org/html/2609.00310#S4.SS2), in which high Conscientiousness and Emotional Stability predict less SA\.

## 5Conclusion

This study demonstrates that personality\-injected LLM personas produce reliable differences in the selection of emotional labor strategies across everyday social scenarios\. We introduced the first dataset that frames surface acting, deep acting, and genuine expression as situated behavioral choices outside occupational contexts\. Our two\-track persona design, combining observer\-rated adjective profiles with in\-character self\-reports, reveals that Conscientiousness and Emotional Stability consistently drive strategy preferences across five models, while Agreeableness surfaces only through self\-report induction\. LLMs show a preference for deep acting, suggesting that training corpora encode this strategy as a normative regulatory response\. Conscientiousness suppresses genuine expression in LLM personas, inverting the positive association documented in previous human\-evaluated reports\. Future work can probe this divergence through the mechanistic interpretability of open\-source models and extend evaluation to cross\-cultural settings\.

## Limitations

We acknowledge the constraints on the scope and generalization of our findings\. We evaluate personas constructed from fictional characters, whose personality profiles reflect crowd\-sourced perceptions rather than ground\-truth personality measures\. The dataset, while grounded in a published corpus, relies on synthetic augmentation for social context and strategy options, which may not fully capture the complexity of real emotional labor encounters\. Our scenarios are English\-only and culturally situated, leaving open the question of whether these trait\-strategy patterns hold across languages and cultural norms\. Finally, our findings remain correlational, and validating the causal mechanisms of trait\-strategy associations will require a mechanistic interpretability framework for future work\.

## Acknowledgments

We thank the CincyNLP group for their suggestions and feedback\. We also thank the anonymous EMNLP reviewers for their insightful suggestions\.

## References

- Ashforth and Humphrey \(1993\)B\. E\. Ashforth and R\. H\. HumphreyEmotional labor in service roles: the influence of identity\.Academy of Management Review18\(1\),pp\. 88–115\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.5465/amr.1993.3997508)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p1.1)\.
- Austinet al\.\(2008\)E\. J\. Austin, T\. C\. P\. Dore, and K\. M\. O’DonovanAssociations of personality and emotional intelligence with display rule perceptions and emotional labour\.Personality and Individual Differences44\(3\),pp\. 679–688\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.paid.2007.10.001)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p2.1),[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px2.p2.1),[§4\.2](https://arxiv.org/html/2609.00310#S4.SS2.p2.1)\.
- Brotheridge and Lee \(2003\)C\. M\. Brotheridge and R\. T\. LeeDevelopment and validation of the emotional labour scale\.Journal of Occupational and Organizational Psychology76\(3\),pp\. 365–379\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1348/096317903769647229),[Link](https://bpspsychub.onlinelibrary.wiley.com/doi/abs/10.1348/096317903769647229),https://bpspsychub\.onlinelibrary\.wiley\.com/doi/pdf/10\.1348/096317903769647229Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p2.1)\.
- Connelly and Ones \(2010\)B\. S\. Connelly and D\. S\. OnesAn other perspective on personality: meta\-analytic integration of observers’ accuracy and predictive validity\.Psychological Bulletin136\(6\),pp\. 1092–1122\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1037/a0021212)Cited by:[§4\.1](https://arxiv.org/html/2609.00310#S4.SS1.SSS0.Px2.p1.1)\.
- DeepSeek\-AI \(2026\)DeepSeek\-AIDeepSeek\-V4: towards highly efficient million\-token context intelligence\.External Links:2606\.19348,[Link](https://arxiv.org/abs/2606.19348)Cited by:[§3\.3](https://arxiv.org/html/2609.00310#S3.SS3.p1.1)\.
- Diefendorffet al\.\(2005\)J\. M\. Diefendorff, M\. H\. Croyle, and R\. H\. GosserandThe dimensionality and antecedents of emotional labor strategies\.Journal of Vocational Behavior66\(2\),pp\. 339–357\.External Links:[Document](https://dx.doi.org/10.1016/j.jvb.2004.02.001)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p1.1),[§1](https://arxiv.org/html/2609.00310#S1.p2.1),[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p2.1)\.
- Duonget al\.\(2025\)P\. A\. Duong, C\. Luong, D\. Bommana, and T\. JiangCHEER\-Ekman: fine\-grained embodied emotion classification\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(ACL 2025\),External Links:[Link](https://aclanthology.org/2025.acl-short.88/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-short.88)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p1.1)\.
- Frising and Balcells \(2026\)M\. Frising and D\. BalcellsLinear personality probing and steering in llms: a big five study\.External Links:2512\.17639,[Link](https://arxiv.org/abs/2512.17639)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px2.p2.1)\.
- Furneset al\.\(2019\)D\. Furnes, H\. Berg, R\. M\. Mitchell, and S\. PaulmannExploring the effects of personality traits on the perception of emotions from prosody\.Frontiers in psychology10,pp\. 184\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.3389/fpsyg.2019.00184)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p2.1)\.
- Goldberget al\.\(2006\)L\. R\. Goldberg, J\. A\. Johnson, H\. W\. Eber, R\. Hogan, M\. C\. Ashton, C\. R\. Cloninger, and H\. G\. GoughThe international personality item pool and the future of public\-domain personality measures\.Journal of Research in Personality40\(1\),pp\. 84–96\.External Links:[Document](https://dx.doi.org/10.1016/j.jrp.2005.08.007)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p5.1),[§3\.2](https://arxiv.org/html/2609.00310#S3.SS2.SSS0.Px2.p1.1)\.
- Goldberg \(1992\)L\. R\. GoldbergThe development of markers for the big\-five factor structure\.Psychological Assessment4\(1\),pp\. 26–42\.External Links:[Document](https://dx.doi.org/10.1037/1040-3590.4.1.26)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p2.1),[§3\.2](https://arxiv.org/html/2609.00310#S3.SS2.SSS0.Px1.p2.1)\.
- Google DeepMind \(2026\)Google DeepMindGemma 4: our most capable open models to date\.External Links:[Link](https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/)Cited by:[§3\.3](https://arxiv.org/html/2609.00310#S3.SS3.p1.1)\.
- Grandey and Sayre \(2019\)A\. A\. Grandey and G\. M\. SayreEmotional labor: regulating emotions for a wage\.Current Directions in Psychological Science28\(2\),pp\. 131–137\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1177/0963721418812771)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p4.1)\.
- Grandey \(2000\)A\. A\. GrandeyEmotion regulation in the workplace: a new way to conceptualize emotional labor\.Journal of Occupational Health Psychology5\(1\),pp\. 95–110\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1037/1076-8998.5.1.95)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p1.1),[§3\.1](https://arxiv.org/html/2609.00310#S3.SS1.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.00310#S4.SS1.SSS0.Px1.p2.1)\.
- Grandey \(2003\)A\. A\. GrandeyWhen the show must go on: surface acting and deep acting as determinants of emotional exhaustion and peer\-rated service delivery\.Academy of Management Journal46\(1\),pp\. 86–96\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.5465/30040678)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p1.1)\.
- Grothet al\.\(2009\)M\. Groth, T\. Hennig\-Thurau, and G\. WalshCustomer reactions to emotional labor: the roles of employee acting strategies and customer detection accuracy\.Academy of Management Journal52\(5\),pp\. 958–974\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.5465/amj.2009.44634116)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p1.1)\.
- Hochschild \(2012\)A\. R\. HochschildThe managed heart: commercialization of human feeling\.1 edition,University of California Press\.External Links:ISBN 9780520272941,[Link](http://www.jstor.org/stable/10.1525/j.ctt1pn9bk)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p1.1),[§3\.1](https://arxiv.org/html/2609.00310#S3.SS1.SSS0.Px1.p1.1)\.
- Hofmannet al\.\(2020\)J\. Hofmann, E\. Troiano, K\. Sassenberg, and R\. KlingerAppraisal theories for emotion classification in text\.InProceedings of the 28th International Conference on Computational Linguistics \(COLING 2020\),External Links:[Link](https://aclanthology.org/2020.coling-main.11/),[Document](https://dx.doi.org/10.18653/v1/2020.coling-main.11)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p1.1)\.
- Huang and Hadfi \(2024\)Y\. J\. Huang and R\. HadfiHow personality traits influence negotiation outcomes? a simulation based on large language models\.InFindings of the Association for Computational Linguistics \(Findings of EMNLP 2024\),External Links:[Link](https://aclanthology.org/2024.findings-emnlp.605/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.605)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px2.p2.1)\.
- Jianget al\.\(2024\)H\. Jiang, X\. Zhang, X\. Cao, C\. Breazeal, D\. Roy, and J\. KabbaraPersonaLLM: investigating the ability of large language models to express personality traits\.InFindings of the Association for Computational Linguistics \(Findings of NAACL 2024\),External Links:[Link](https://aclanthology.org/2024.findings-naacl.229/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-naacl.229)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p3.1),[§1](https://arxiv.org/html/2609.00310#S1.p5.1),[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px2.p1.1)\.
- Judgeet al\.\(2009\)T\. A\. Judge, E\. F\. Woolf, and C\. HurstIs emotional labor more difficult for some than for others? a multilevel, experience\-sampling study\.Personnel psychology62\(1\),pp\. 57–88\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/j.1744-6570.2008.01129.x)Cited by:[§4\.2](https://arxiv.org/html/2609.00310#S4.SS2.p2.1)\.
- Kammeyer\-Muelleret al\.\(2013\)J\. D\. Kammeyer\-Mueller, A\. L\. Rubenstein, D\. M\. Long, M\. A\. Odio, B\. R\. Buckman, Y\. Zhang, and M\. D\.K\. Halvorsen\-GanepolaA meta\-analytic structural model of dispositonal affectivity and emotional labor\.Personnel Psychology66,pp\. 47–90\.External Links:[Document](https://dx.doi.org/10.1111/peps.12009)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p1.1)\.
- Kiffin\-Petersenet al\.\(2011\)S\. A\. Kiffin\-Petersen, C\. L\. Jordan, and G\. N\. SoutarThe big five, emotional exhaustion and citizenship behaviors in service settings: the mediating role of emotional labor\.Personality and Individual Differences50\(1\),pp\. 43–48\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.paid.2010.08.018)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p2.1),[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px2.p2.1)\.
- Liet al\.\(2024\)J\. Li, C\. Peris, N\. Mehrabi, P\. Goyal, K\. Chang, A\. Galstyan, R\. Zemel, and R\. GuptaThe steerability of large language models toward data\-driven personas\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL 2024\),External Links:[Link](https://aclanthology.org/2024.naacl-long.405/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.405)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p3.1)\.
- Matzet al\.\(2024\)S\. C\. Matz, J\. D\. Teeny, S\. S\. Vaid, H\. Peters, G\. M\. Harari, and M\. CerfThe potential of generative AI for personalized persuasion at scale\.Scientific Reports14,pp\. 4692\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41598-024-53755-0)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p3.1)\.
- McCrae and John \(1992\)R\. R\. McCrae and O\. P\. JohnAn introduction to the five\-factor model and its applications\.Journal of Personality60\(2\),pp\. 175–215\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/j.1467-6494.1992.tb00970.x)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p2.1)\.
- Mesmer\-Magnuset al\.\(2012\)J\. R\. Mesmer\-Magnus, L\. A\. DeChurch, and A\. WaxMoving emotional labor beyond surface and deep acting: a discordance–congruence perspective\.Organizational Psychology Review2\(1\),pp\. 6–53\.External Links:[Document](https://dx.doi.org/10.1177/2041386611417746)Cited by:[§4\.1](https://arxiv.org/html/2609.00310#S4.SS1.SSS0.Px1.p2.1),[§4\.2](https://arxiv.org/html/2609.00310#S4.SS2.p2.1)\.
- Mohammadet al\.\(2018\)S\. Mohammad, F\. Bravo\-Marquez, M\. Salameh, and S\. KiritchenkoSemEval\-2018 task 1: affect in tweets\.InProceedings of the 12th International Workshop on Semantic Evaluation,External Links:[Link](https://aclanthology.org/S18-1001/),[Document](https://dx.doi.org/10.18653/v1/S18-1001)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p1.1)\.
- Open\-Source Psychometrics Project \(2022\)Open\-Source Psychometrics ProjectData from the statistical “which character” personality quiz\.External Links:[Link](https://openpsychometrics.org/tests/characters/data/)Cited by:[§3\.2](https://arxiv.org/html/2609.00310#S3.SS2.SSS0.Px1.p1.1)\.
- OpenAI \(2026\)OpenAIIntroducing gpt\-5\.4\.External Links:[Link](https://openai.com/index/introducing-gpt-5-4/)Cited by:[§3\.3](https://arxiv.org/html/2609.00310#S3.SS3.p1.1)\.
- Peters and Matz \(2024\)H\. Peters and S\. MatzLarge language models can infer psychological dispositions of social media users\.External Links:2309\.08631,[Link](https://arxiv.org/abs/2309.08631)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p3.1)\.
- Saimet al\.\(2025\)M\. Saim, P\. A\. Duong, C\. Luong, A\. Bhanderi, and T\. JiangAnatomy of a feeling: narrating embodied emotions via large vision\-language models\.InFindings of the Association for Computational Linguistics \(Findings of EMNLP 2025\),External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1276/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1276)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p1.1)\.
- Saim and Jiang \(2026\)M\. Saim and T\. JiangDo emotions influence moral judgment in large language models?\.InFindings of the Association for Computational Linguistics \(Findings of ACL 2026\),External Links:[Link](https://aclanthology.org/2026.findings-acl.1346/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1346)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p4.1)\.
- Scherer \(2001\)K\. R\. SchererAppraisal considered as a process of multilevel sequential checking\.InAppraisal Processes in Emotion: Theory, Methods, Research,pp\. 92–120\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1093/oso/9780195130072.003.0005)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p1.1)\.
- Smith and Ellsworth \(1985\)C\. A\. Smith and P\. C\. EllsworthPatterns of cognitive appraisal in emotion\.Journal of Personality and Social Psychology48\(4\),pp\. 813–838\.External Links:[Document](https://dx.doi.org/10.1037/0022-3514.48.4.813)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p1.1)\.
- Sorokovikovaet al\.\(2024\)A\. Sorokovikova, S\. Rezagholi, N\. Fedorova, and I\. P\. YamshchikovLLMs simulate Big5 personality traits: further evidence\.InProceedings of the 1st Workshop on Personalization of Generative AI Systems,External Links:[Link](https://aclanthology.org/2024.personalize-1.7/)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p3.1),[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px2.p1.1)\.
- Strapparava and Mihalcea \(2007\)C\. Strapparava and R\. MihalceaSemEval\-2007 task 14: affective text\.InProceedings of the Fourth International Workshop on Semantic Evaluations \(SemEval\-2007\),External Links:[Link](https://aclanthology.org/S07-1013/)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p1.1)\.
- Terraccianoet al\.\(2003\)A\. Terracciano, M\. Merritt, A\. B\. Zonderman, and M\. K\. EvansPersonality traits and sex differences in emotion recognition among african americans and caucasians\.Annals of the New York Academy of Sciences1000,pp\. 309–312\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1196/annals.1280.032)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p2.1)\.
- Troianoet al\.\(2023\)E\. Troiano, L\. Oberländer, and R\. KlingerDimensional modeling of emotions in text with appraisal theories: corpus creation, annotation reliability, and prediction\.Computational Linguistics49\(1\),pp\. 1–72\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00461)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.00310#S3.SS1.p1.1)\.
- Tsenget al\.\(2024\)Y\. Tseng, Y\. Huang, T\. Hsiao, W\. Chen, C\. Huang, Y\. Meng, and Y\. ChenTwo tales of persona in LLMs: a survey of role\-playing and personalization\.InFindings of the Association for Computational Linguistics \(Findings of EMNLP 2024\),External Links:[Link](https://aclanthology.org/2024.findings-emnlp.969/)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px2.p1.1)\.
- Vazire \(2010\)S\. VazireWho knows what about a person? the self–other knowledge asymmetry \(SOKA\) model\.\.Journal of personality and social psychology98\(2\),pp\. 281\.External Links:[Document](https://dx.doi.org/10.1037/a0017908)Cited by:[§3\.2](https://arxiv.org/html/2609.00310#S3.SS2.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.00310#S4.SS1.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2609.00310#S4.SS2.p4.1)\.
- Vennemeyeret al\.\(2026\)D\. Vennemeyer, P\. A\. Duong, T\. Zhan, and T\. JiangSycophancy is not one thing: causal separation of sycophantic behaviors in LLMs\.External Links:2509\.21305,[Link](https://arxiv.org/abs/2509.21305)Cited by:[§4\.2](https://arxiv.org/html/2609.00310#S4.SS2.p3.1)\.
- Weidingeret al\.\(2021\)L\. Weidinger, J\. Mellor, M\. Rauh, C\. Griffin, J\. Uesato, P\. Huang, M\. Cheng, M\. Glaese, B\. Balle, A\. Kasirzadeh, Z\. Kenton, S\. Brown, W\. Hawkins, T\. Stepleton, C\. Biles, A\. Birhane, J\. Haas, L\. Rimell, L\. A\. Hendricks, W\. Isaac, S\. Legassick, G\. Irving, and I\. GabrielEthical and social risks of harm from language models\.External Links:2112\.04359,[Link](https://arxiv.org/abs/2112.04359)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p4.1)\.
- Wuet al\.\(2025\)S\. Wu, Y\. Zhu, W\. Hsu, M\. Lee, and Y\. DengFrom personas to talks: revisiting the impact of personas on LLM\-synthesized emotional support conversations\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing \(EMNLP 2025\),External Links:[Link](https://aclanthology.org/2025.emnlp-main.277/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.277)Cited by:[§1](https://arxiv.org/html/2609.00310#S1.p4.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. QiuQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§3\.3](https://arxiv.org/html/2609.00310#S3.SS3.p1.1)\.
- Yaoet al\.\(2026\)L\. H\. Yao, V\. Anand, Y\. Zhuang, and T\. JiangRhetorical questions in LLM representations: a linear probing study\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(ACL 2026\),External Links:[Link](https://aclanthology.org/2026.acl-long.5/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.5)Cited by:[§4\.2](https://arxiv.org/html/2609.00310#S4.SS2.p3.1)\.
- Yehet al\.\(2020\)S\. J\. Yeh, S\. S\. Chen, K\. Yuan, W\. Chou, and T\. T\. WanEmotional labor in health care: the moderating roles of personality and the mediating role of sleep on job performance and satisfaction\.Frontiers in Psychology11,pp\. 574898\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.3389/fpsyg.2020.574898)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px2.p2.1)\.
- Zhuanget al\.\(2024\)Y\. Zhuang, T\. Jiang, and E\. RiloffMy heart skipped a beat\! recognizing expressions of embodied emotion in natural language\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL 2024\),External Links:[Link](https://aclanthology.org/2024.naacl-long.193/),[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.193)Cited by:[§2](https://arxiv.org/html/2609.00310#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix APrompts

We designed two prompts for this study\. The dataset augmentation prompt \(Prompt A\.1\) instructed GPT\-5\.4 to extend each sentence from the original corpus with a one\-sentence social context introducing display pressure and to generate three behavioral response options corresponding to surface acting, deep acting, and genuine expression\. To prevent label leakage, the prompt explicitly prohibited emotion adjectives and labels, requiring the model to convey internal states exclusively through physical and behavioral cues\. The persona evaluation prompt \(Prompt A\.2\) presented each fictional character’s BAP trait profile as a scored bipolar adjective list and asked the model to select, in character, the single response option it would most naturally choose\. We constrained the output to a single letter \(A, B, or C\) to enforce a clean forced\-choice response and eliminate free\-text ambiguity in downstream analysis\. A similar prompt was repeated for the IPIP\-track, the only difference being that the IPIP block had Likert ratings \(1\-5\) instead of adjective pair scores \(1\-100\)\.

## Appendix BStrategy selection choices

### B\.1Selected BAP Items and OCEAN Composite Construction

Table[7](https://arxiv.org/html/2609.00310#A3.T7)reports the 40 verified BAP adjective pairs retained after filtering the full 500\-item instrument against Goldberg’s taxonomy\. We retain only pairs for which at least one pole maps exactly onto an adjective appearing in Goldberg’s published marker list, which ensures that every trait signal we inject into a persona has a well\-established structural relationship to the Big Five \(OCEAN\) traits\. The filtering reduces the raw instrument from 500 to 40 items, distributed across five OCEAN dimensions: Openness \(8 items\), Conscientiousness \(7 items\), Extraversion \(8 items\), Agreeableness \(8 items\), and Emotional Stability \(9 items\)\. We report the inverse of the Neuroticism dimension, written as ‘Emotional Stability’ throughout the paper\. BAP items in this dimension are originally keyed such that higher scores indicate greater stability \(e\.g\., insecure–confident, anxious–calm\)\.

We manually verify each retained item against the Goldberg taxonomy to determine which pole corresponds to the high end of the relevant OCEAN dimension, and we reverse\-score items accordingly before computing composites\. For example, BAP112 \(flexible–rigid\) requires reversal so that a score of “flexible” \(on the higher end\) contributes positively to Agreeable rather than negatively\. The final composite for each dimension is the mean of its verified, polarity\-aligned items, expressed on a 1–100 scale that is passed directly into the BAP persona block at inference time\.

### B\.2IPIP\-50 Administration and Scoring

Table[8](https://arxiv.org/html/2609.00310#A3.T8)lists all 50 IPIP items used in the in\-character self\-report track, grouped by OCEAN dimension with forward and reverse keying indicated\. The IPIP\-50 is a public\-domain instrument drawn from the International Personality Item Pool with 10 items per dimension rated on a 1–5 Likert scale\. Reverse\-keyed items \(marked with a negative key \) are reflected before aggregation: a rating of 1 on a reverse\-keyed item contributes 5 to the composite, and vice versa\. Subscale composites are then computed as the mean of the 10 reflected items, yielding a score on a 1–5 scale for each OCEAN dimension\. These are standardized across characters before entry into the Spearman correlation analyses to place the two persona tracks on a comparable footing\.

### B\.3Character Selection

Table[9](https://arxiv.org/html/2609.00310#A3.T9)lists the 50 characters selected for evaluation\. The pool spans a broad range of fictional universes, including long\-running television series\. Character personality profiles in the OpenPsychometrics BAP dataset reflect aggregated crowd\-sourced ratings, and characters from culturally dominant franchises tend to accumulate larger rater pools, which increases the stability of their BAP score distributions\. The standard deviation filtering step selects characters whose raters disagree substantially across adjective pairs\. Greedy farthest\-point sampling in the BAP space then ensures that the 50 selected characters occupy distinct regions of personality space rather than clustering around a single archetype, such as the agreeable protagonist or the neurotic antagonist\.

## Appendix CExperimental Details and Additional Analysis

### C\.1Parallel Model Correlations

#### DeepSeek\-V4\-Flash\.

Table[10](https://arxiv.org/html/2609.00310#A3.T10)reports Spearman correlations for DeepSeek\-V4\-Flash\. Conscientiousness and Emotional Stability again dominate, replicating the same three\-way directional signature \(high C/ES→\\tomore DA, less SA, less GE\) at comparable magnitudes\. The inversion between conscientiousness and genuine expression persists \(ρ\\rho=−\-0\.645, BAP;−\-0\.611, IPIP\), confirming that this anomaly is not model\-specific\. We also observe Agreeableness reaching greater significance in the BAP track for DeepSeek, whereas GPT\-5\.4 showed no significant BAP Agreeableness effects; the IPIP track reinforces this with strong SA suppression \(ρ\\rho=−\-0\.640∗∗∗\)\. Also, Extraversion shows a significant positive GE effect in the BAP track \(ρ\\rho= \+0\.369∗∗\) that is absent in GPT\-5\.4\. This suggests that DeepSeek encodes extraversion as a signal for unguarded expression when responding to observer\-rated personas\. Openness gains marginal significance only in the IPIP track\.

### C\.2Three\-Way Choice vs\. Likert Format Validation

We assess whether our forced\-choice response format introduced systematic measurement artifacts\. We ran a parallel evaluation in which we replaced the three\-way forced\-choice with a Likert\-scale rating task\. Instead of selecting a single option, the model rated its likelihood of adopting each of the three strategies on a scale of 1 to 5\. We conducted this validation run on the BAP persona track using GPT\-5\.4 across all 50 characters and 500 scenarios, exactly mirroring the primary evaluation\. The motivation was two\-fold: first, to verify that the trait–strategy associations we observe do not arise as a result of the forced\-choice constraint; second, to establish whether Likert scaling recovers the same ordinal structure across OCEAN traits\.

Table[11](https://arxiv.org/html/2609.00310#A3.T11)reports Spearman correlations between BAP\-derived OCEAN composites and EL strategy measures under both formats\. The pattern of associations remains largely consistent across response formats\. Conscientiousness and Emotional Stability produce the strongest and most significant correlations in both conditions, and the direction of every significant coefficient is preserved\. The C–GE inversion \(high Conscientiousness predicting reduced genuine expression\) replicates under the Likert format \(ρ=−0\.516\\rho=\-0\.516,p<\.001p<\.001\), confirming that this finding is not a forced\-choice artifact\. Minor magnitude differences appear for Openness and Agreeableness, where neither format yields significant correlations, suggesting that these traits produce consistent weak signals\.

### C\.3Individual Character Preference

Table[6](https://arxiv.org/html/2609.00310#A3.T6)grounds the entropy analysis with character\-level evidence\. The three lowest\-entropy characters all show near\-total DA preference, and their OCEAN profiles share high Conscientiousness and Emotional Stability, which validates the traits our correlation analysis identifies as the strongest DA predictors\. Their personas exert such a rigid pull toward one strategy that the emotional situation barely shifts the distribution\. The three highest\-entropy characters each distribute selections near evenly across all three strategies \(28–37% per strategy\), and their profiles reflect low Conscientiousness and high Openness, which are traits that our analysis associates with weaker, more context\-sensitive regulatory commitments\.

CharacterWorkSA/DA/GE \(%\)OCEAN compositesRigid personality \(lowest entropy\)TuvokStar Trek: Voyager2/98/164 / 91 / 42 / 54 / 72Dr\. Hannibal LecterHannibal8/91/185 / 60 / 60 / 39 / 61IrohAvatar: TLA4/90/674 / 62 / 50 / 91 / 71Scenario\-responsive \(highest entropy\)Ava ColemanAbbott Elementary31/35/3360 / 14 / 86 / 28 / 54Jules LoudenThe Cabin in the Woods28/37/3552 / 30 / 75 / 48 / 48Gaius BaltarBattlestar Galactica28/37/3578 / 29 / 60 / 26 / 26

Table 6:Top and bottom characters of strategy consistency ordered by entropy\. OCEAN composites derived from BAP observer ratings \(GPT\-5\.4,N=50N\{=\}50characters\)\.![Refer to caption](https://arxiv.org/html/2609.00310v1/heatmap_ipip.png)Figure 4:Persona influence by emotion and model \(IPIP track\) depicted by Shannon Entropy\.
### C\.4IPIP\-track Entropy Experiments

Figure[4](https://arxiv.org/html/2609.00310#A3.F4)reports normalized Shannon entropy under the IPIP\-50 persona track\. The pattern corroborates the BAP\-track findings\. On average, Gemma\-4\-31B again achieves the highest entropy across emotion categories, GPT\-5\.4 and DeepSeek\-V4\-Flash occupy a high\-entropy band \(0\.81–0\.91\), and the two Qwen models remain comparably more susceptible to the scenario\. However, IPIP\-track values are generally higher than their BAP counterparts across the model pool, suggesting that self\-reported persona preserves more individual variation in strategy selection than observer\-rated adjective profiles do\. Qwen\-3\-32B ranges from 0\.59 \(Sadness\) to 0\.78 \(Joy\) under IPIP, compared to 0\.42–0\.70 under BAP, indicating that richer first\-person framing partially restores persona sensitivity even in the more situation\-dominant models\.

### C\.5Inter\-Model Agreement

Table[12](https://arxiv.org/html/2609.00310#A3.T12)reports the pairwise agreement between the five evaluated models across the BAP and IPIP persona tracks\. Agreement is uniformly higher under the IPIP track than the BAP track, suggesting that richer, psychometrically grounded persona descriptions produce more consistent emotion\-regulation judgments across models\. On both the BAP and IPIP tracks, the strongest pair is GPT\-5\.4 vs\. Gemma\-31B \(κ=0\.401\\kappa=0\.401,κ=0\.545\\kappa=0\.545\), while Gemma\-31B vs\. Qwen\-3\-8B shows the weakest alignment\. Allκ\\kappavalues fall below 0\.60, indicating only moderate agreement at best, consistent with the inherent ambiguity of emotion\-labor identification\.

Breaking agreement down by category shows that deep\-acting items attract the highest per\-class agreement \(BAP mean 38\.6%, IPIP 46\.2%\), followed by genuine expression \(32\.2% / 35\.5%\), while surface\-acting items are the hardest to agree on \(17\.3% / 18\.7%\)\.

### A\.1 Dataset Augmentation Prompt

Prompt A\.1: Social Context and EL Strategy GenerationYou are generating evaluation items for a psychology dataset on emotion regulation\. You will receive a situation and a felt emotion\. Your task:1\.Add a brief social context \(1 sentence\) before or after the given sentence that creates pressure to manage the emotional display—e\.g\., a colleague asks, a child is watching, a client is present\.IMPORTANT RULES FOR ADDING CONTEXT:•The added context can appear BEFORE or AFTER the original situation sentence for diversity and in a coherent and correct grammatical way\.•The context must be CONSISTENT with the original sentence\. Do not introduce details that contradict or are incompatible with the situation described\. Ensure the added context fits the same setting and circumstances in a coherent way\.2\.Generate exactly 3 options \(A–C\) in randomized order from these categories:SURFACE ACTING:The person performs a different emotion outwardly while the original feeling persists inside\. The performance is imperfect—one subtle involuntary cue reveals the inner state\. The performed emotion must differ from the felt emotion\.DEEP ACTING:The person actively changes how they feel internally through reframing, perspective\-taking, or self\-talk\. By the time they respond, their inner state has genuinely shifted\. They are not faking\.GENUINE EXPRESSION:The person expresses the felt emotion directly with no attempt to manage or hide it\.RULES:•Every option must include at least one behavioral or physical detail\.•Each option: 20–30 words, one sentence, completes the stem naturally\.•Do NOT use any emotion labels or adjectives\. Convey the emotional state through actions, body language, or physiological responses only\.•The surface acting option must have exactly one performed display cue and one leakage cue\.

### A\.2 Persona Evaluation Prompt

Prompt A\.2: EL Strategy Selection Under PersonaYou are \{character\_name\} from \{character\_work\}\. Below are your personality traits, each shown as a pair of opposite adjectives with a score on a scale of 1–100, where 1 = entirely the left adjective and 100 = entirely the right adjective:\{bap\_block\}\(These will be human\-reported scores between 1\-100 for each BAP\.\)Read the situation below and decide which of the three options you would most naturally choose, given your personality traits\.Situation:\{EL\_sentence\}Options:\{options\_block\}\(Shuffled options of SA, DA and GE to select from \.\)Reply with only the single letter of the option you would choose \(A, B, or C\)\.

RankIDAdjective PairGoldberg ScaleSDExtraversion8BAP49scheduled – spontaneousinhibited–spontaneous23\.9998BAP133stick\-in\-the\-mud – adventurousunadventurous–adventurous21\.38100BAP47deliberate – spontaneousinhibited–spontaneous21\.36165BAP493mellow – energeticunenergetic–energetic20\.31211BAP391timid – cockytimid–bold19\.77256BAP130passive – assertiveunassertive–assertive19\.11276BAP2shy – boldtimid–bold18\.80389BAP131slothful – activeinactive–active16\.98Emotional Stability68BAP499unstable – stableunstable–stable21\.92130BAP125angry – good\-humoredangry–calm20\.87159BAP35emotional – logicalemotional–unemotional20\.41202BAP62insecure – confidentinsecure–secure19\.87245BAP74anxious – calmangry–calm19\.19263BAP367emotional – unemotionalemotional–unemotional18\.97273BAP36moody – stablemoody–steady18\.83304BAP18tense – relaxedtense–relaxed18\.50479BAP330envious – pridefulenvious–not envious13\.33Agreeableness34BAP84cruel – kindunkind–kind22\.8736BAP79selfish – altruisticselfish–unselfish22\.7951BAP129cold – warmcold–warm22\.3858BAP17competitive – cooperativeuncooperative–cooperative22\.2272BAP31rude – respectfulrude–polite21\.8594BAP76quarrelsome – warmcold–warm21\.42121BAP351stingy – generousstingy–generous21\.04324BAP112rigid – flexibleinflexible–flexible18\.19Conscientiousness28BAP1playful – seriousfrivolous–serious23\.0741BAP75disorganized – self\-disciplineddisorganized–organized22\.6170BAP27impulsive – cautiousrash–cautious21\.87173BAP311experimental – reliableundependable–reliable20\.17207BAP353extravagant – thriftyextravagant–thrifty19\.81269BAP90serious – boldfrivolous–serious / timid–bold18\.86335BAP32lazy – diligentlazy–hardworking18\.03Openness88BAP9rugged – refinedunrefined–refined21\.53160BAP132practical – imaginativeunimaginative–imaginative/ impractical–practical20\.41176BAP490intuitive – analyticalunanalytical–analytical20\.12184BAP29conventional – creativeuncreative–creative20\.06259BAP73uncreative – open to new experiencesuncreative–creative19\.01334BAP299unobservant – perceptiveimperceptive–perceptive18\.03351BAP372rustic – cultureduncultured–cultured17\.76448BAP30apathetic – curiousuninquisitive–curious15\.28Table 7:Selected BAP adjective\-pair items \(sorted on valence\) ranked by standard deviation \(SD\), grouped by OCEAN personality dimensions\. Lower ranks indicate higher variability across characters\.IDItemKeyExtraversionEXT1I am the life of the party\.\+EXT2I don’t talk a lot\.−\-EXT3I feel comfortable around people\.\+EXT4I keep in the background\.−\-EXT5I start conversations\.\+EXT6I have little to say\.−\-EXT7I talk to a lot of different people at parties\.\+EXT8I don’t like to draw attention to myself\.−\-EXT9I don’t mind being the center of attention\.\+EXT10I am quiet around strangers\.−\-Emotional StabilityEST1I get stressed out easily\.−\-EST2I am relaxed most of the time\.\+EST3I worry about things\.−\-EST4I seldom feel blue\.\+EST5I am easily disturbed\.−\-EST6I get upset easily\.−\-EST7I change my mood a lot\.−\-EST8I have frequent mood swings\.−\-EST9I get irritated easily\.−\-EST10I often feel blue\.−\-AgreeablenessAGR1I feel little concern for others\.−\-AGR2I am interested in people\.\+AGR3I insult people\.−\-AGR4I sympathize with others’ feelings\.\+AGR5I am not interested in other people’s problems\.−\-AGR6I have a soft heart\.\+AGR7I am not really interested in others\.−\-AGR8I take time out for others\.\+AGR9I feel others’ emotions\.\+AGR10I make people feel at ease\.\+ConscientiousnessCSN1I am always prepared\.\+CSN2I leave my belongings around\.−\-CSN3I pay attention to details\.\+CSN4I make a mess of things\.−\-CSN5I get chores done right away\.\+CSN6I often forget to put things back in their proper place\.−\-CSN7I like order\.\+CSN8I shirk my duties\.−\-CSN9I follow a schedule\.\+CSN10I am exacting in my work\.\+OpennessOPN1I have a rich vocabulary\.\+OPN2I have difficulty understanding abstract ideas\.−\-OPN3I have a vivid imagination\.\+OPN4I am not interested in abstract ideas\.−\-OPN5I have excellent ideas\.\+OPN6I do not have a good imagination\.−\-OPN7I am quick to understand things\.\+OPN8I use difficult words\.\+OPN9I spend time reflecting on things\.\+OPN10I am full of ideas\.\+Table 8:IPIP\-50 items grouped by OCEAN personality trait dimensions\. “Key” indicates whether the item is positively keyed \(\+\) or reverse keyed \(−\-\)\.RankChar IDCharacterWork1OD/2TelemachusThe Odyssey2R30/2Tracy Jordan30 Rock3SV/3Nelson BighettiSilicon Valley4GOT/23Joffrey BaratheonGame of Thrones5FR/4OlafFrozen6STV/7TuvokStar Trek: Voyager7ALA/8Firelord OzaiAvatar: The Last Airbender8ALA/7IrohAvatar: The Last Airbender9WC/1Neal CaffreyWhite Collar10RM/1Rick SanchezRick and Morty11ARC/2PowderArcane12GLEE/15Emma PillsburyGlee13TO/8Stanley HudsonThe Office14HNB/2Dr\. Hannibal LecterHannibal15DHSAB/2Captain HammerDr\. Horrible’s Sing\-Along Blog16NG/2Nick MillerNew Girl17HP/24Petunia DursleyHarry Potter18PR/7Jerry GergichParks and Recreation19FAR/4Dominar Rygel XVIFarscape20AD/6Buster BluthArrested Development21OFOCN/3‘Chief’ BromdenOne Flew Over the Cuckoo’s Nest22Y/1John DuttonYellowstone23AE/1Janine TeaguesAbbott Elementary24MR/1Elliot AldersonMr\. Robot25BR/2Martha ScottBaby Reindeer26OA/1Prairie JohnsonThe OA27GP/4JanetThe Good Place28OPM/1SaitamaOne Punch Man29HSM/3Sharpay EvansHigh School Musical30WE/1Wynonna EarpWynonna Earp31SQG/5Oh Il\-namSquid Game32HD/8Lord Larys StrongHouse of the Dragon33FB/2ClaireFleabag34R30/4Pete Hornberger30 Rock35MLP/1ApplejackMy Little Pony: Friendship Is Magic36GIGE/15Matt PressGinny & Georgia37CITW/3Jules LoudenThe Cabin in the Woods38YJ/5MistyYellowjackets39SHL/1Frank GallagherShameless40EXP/3Amos BurtonThe Expanse41GILG/4Luke DanesGilmore Girls42SC/5Mr\. BigSex and the City43STIG/1Rintarou OkabeSteins;Gate44TW/12Omar LittleThe Wire45BSG/5Gaius BaltarBattlestar Galactica46PKB/2Arthur ShelbyPeaky Blinders47SNW/1Snow WhiteSnow White and the Seven Dwarfs48AE/5Ava ColemanAbbott Elementary49AFTL/1Tony JohnsonAfter Life50WSW/4Teddy FloodWestworldTable 9:50 fictional characters selected via greedy maximin sampling in BAP trait space\.ReliabilityBAP \(Observer\-Rated\)IPIP \(Self\-Report\)Trait𝜶BAP\\boldsymbol\{\\alpha\}\_\{\\text\{BAP\}\}𝜶IPIP\\boldsymbol\{\\alpha\}\_\{\\text\{IPIP\}\}SADAGESADAGEOpenness0\.700\.95−0\.028\-0\.028\+0\.210\+0\.210−0\.266\-0\.266−0\.167\-0\.167\+0\.323∗\+0\.323^\{\*\}−0\.312∗\-0\.312^\{\*\}Conscientiousness0\.880\.99−0\.529∗∗∗\-0\.529^\{\*\*\*\}\+0\.679∗∗∗\+0\.679^\{\*\*\*\}−0\.645∗∗∗\-0\.645^\{\*\*\*\}−0\.244\-0\.244\+0\.574∗∗∗\+0\.574^\{\*\*\*\}−0\.611∗∗∗\-0\.611^\{\*\*\*\}Extraversion0\.830\.99\+0\.131\+0\.131−0\.312∗\-0\.312^\{\*\}\+0\.369∗⁣∗\+0\.369^\{\*\*\}−0\.140\-0\.140\+0\.044\+0\.044−0\.019\-0\.019Agreeableness0\.950\.99−0\.433∗⁣∗\-0\.433^\{\*\*\}\+0\.512∗∗∗\+0\.512^\{\*\*\*\}−0\.504∗∗∗\-0\.504^\{\*\*\*\}−0\.640∗∗∗\-0\.640^\{\*\*\*\}\+0\.472∗∗∗\+0\.472^\{\*\*\*\}−0\.395∗⁣∗\-0\.395^\{\*\*\}Emotional Stability0\.760\.99−0\.466∗∗∗\-0\.466^\{\*\*\*\}\+0\.682∗∗∗\+0\.682^\{\*\*\*\}−0\.640∗∗∗\-0\.640^\{\*\*\*\}−0\.634∗∗∗\-0\.634^\{\*\*\*\}\+0\.870∗∗∗\+0\.870^\{\*\*\*\}−0\.874∗∗∗\-0\.874^\{\*\*\*\}Table 10:Spearman correlations \(ρ\\rho\) between OCEAN trait composites and emotional labor strategy proportions \(Deepseek\-V4\-Flash\)\. BAP = observer\-rated composites from bipolar adjective profiles; IPIP = self\-report composites from in\-character IPIP\-50 administration\. Cronbach’sα\\alphavalues are reported separately for BAP and IPIP composites\.∗p<\.05\{\}^\{\*\}p<\.05,p∗⁣∗<\.01\{\}^\{\*\*\}p<\.01,∗∗∗p<\.001\{\}^\{\*\*\*\}p<\.001\.Three\-Way Choice \(Proportion\)Likert \(Mean Rating\)TraitSADAGESADAGEOpenness\+0\.064\+0\.064−0\.013\-0\.013−0\.032\-0\.032\+0\.178\+0\.178−0\.007\-0\.007\+0\.068\+0\.068Conscientiousness−0\.446∗⁣∗\-0\.446^\{\*\*\}\+0\.597∗∗∗\+0\.597^\{\*\*\*\}−0\.570∗∗∗\-0\.570^\{\*\*\*\}−0\.280∗\-0\.280^\{\*\}\+0\.626∗∗∗\+0\.626^\{\*\*\*\}−0\.516∗∗∗\-0\.516^\{\*\*\*\}Extraversion\+0\.110\+0\.110−0\.203\-0\.203\+0\.184\+0\.184−0\.018\-0\.018−0\.289∗\-0\.289^\{\*\}\+0\.124\+0\.124Agreeableness−0\.276\-0\.276\+0\.144\+0\.144−0\.072\-0\.072\+0\.123\+0\.123\+0\.244\+0\.244\+0\.186\+0\.186Emotional Stability−0\.350∗\-0\.350^\{\*\}\+0\.522∗∗∗\+0\.522^\{\*\*\*\}−0\.517∗∗∗\-0\.517^\{\*\*\*\}−0\.368∗⁣∗\-0\.368^\{\*\*\}\+0\.503∗∗∗\+0\.503^\{\*\*\*\}−0\.495∗∗∗\-0\.495^\{\*\*\*\}Table 11:Spearman correlations \(ρ\\rho\) between BAP\-derived OCEAN composites and emotional labor strategy measures under two response formats: forced\-choice \(strategy proportions\) and Likert \(mean 1–5 ratings\)\.∗p<\.05\{\}^\{\*\}p<\.05,p∗⁣∗<\.01\{\}^\{\*\*\}p<\.01,∗∗∗p<\.001\{\}^\{\*\*\*\}p<\.001\.Model AModel Bκ\\kappaBAP persona trackGPT\-5\.4DeepSeek\-V4\-Flash0\.263GPT\-5\.4Gemma\-4\-31B0\.401GPT\-5\.4Qwen\-3\-8B0\.111GPT\-5\.4Qwen\-3\-32B0\.147DeepSeek\-V4\-FlashGemma\-4\-31B0\.269DeepSeek\-V4\-FlashQwen\-3\-8B0\.125DeepSeek\-V4\-FlashQwen\-3\-32B0\.160Gemma\-4\-31BQwen\-3\-8B0\.080Gemma\-4\-31BQwen\-3\-32B0\.106Qwen\-3\-8BQwen\-3\-32B0\.202IPIP persona trackGPT\-5\.4DeepSeek\-V4\-Flash0\.357GPT\-5\.4Gemma\-4\-31B0\.545GPT\-5\.4Qwen\-3\-8B0\.169GPT\-5\.4Qwen\-3\-32B0\.207DeepSeek\-V4\-FlashGemma\-4\-31B0\.335DeepSeek\-V4\-FlashQwen\-3\-8B0\.172DeepSeek\-V4\-FlashQwen\-3\-32B0\.221Gemma\-4\-31BQwen\-3\-8B0\.144Gemma\-4\-31BQwen\-3\-32B0\.187Qwen\-3\-8BQwen\-3\-32B0\.249Table 12:Pairwise inter\-model Cohen’sκ\\kappaon the BAP and IPIP persona tracks\.

Similar Articles

How Well Do Large Language Models Capture Human Personality?

arXiv cs.AI

This paper systematically evaluates assumptions about LLM persona prompting and identifies 'persona manifold collapse,' where richer persona descriptions reduce behavioral diversity and simulation fidelity. The findings show that simple age-gender personas often outperform more detailed profiles.

Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning

arXiv cs.CL

This paper investigates whether assigning personas to large language models induces human-like motivated reasoning, finding that persona-assigned LLMs show up to 9% reduced veracity discernment and are up to 90% more likely to evaluate scientific evidence in ways congruent with their induced political identity, with prompt-based debiasing largely ineffective.