PTEI: Integrating Personality Traits to Enhance Emotional Intelligence in Large Language Models
Summary
This paper presents PTEI, a framework that integrates personality traits (MBTI and OCEAN) into LLMs to enhance emotional intelligence, using contrastive learning and personality-aware prompts. Experiments show significant improvements in emotional understanding, especially when combined with Chain-of-Thought reasoning.
View Cached Full Text
Cached at: 07/14/26, 04:21 AM
# PTEI: Integrating Personality Traits to Enhance Emotional Intelligence in Large Language Models
Source: [https://arxiv.org/html/2607.10245](https://arxiv.org/html/2607.10245)
Amir Reza Jafari,Praboda RajapakshaDepartment of Computer Science, Aberystwyth UniversityAberystwythWales,Reza FarahbakhshSamovar, Telecom SudParis, Institut Polytechnique de ParisPalaiseauFranceandNoel CrespiSamovar, Telecom SudParis, Institut Polytechnique de ParisPalaiseauFrance
###### Abstract\.
Despite advances inEmotional Intelligence \(EI\), Large Language Models \(LLMs\) still significantly underperform humans in complex emotional reasoning\. This gap originates partly from the limited incorporation of individual differences, particularly personality traits, which are fundamental to human emotional inference\. To address this, we proposePTEI, a novel framework for integrating Personality Traits into Emotional Intelligence tasks using LLMs\. In PTEI, MBTI and OCEAN personality traits are first extracted directly from the given emotional scenarios and then utilized as contextual knowledge within personality\-aware prompts, guiding LLMs to accurately infer emotions and their underlying causes\. To ensure optimal contextual grounding, we employContrastive Learningto construct an optimized retrieval system that surfaces emotionally and personally aligned scenarios, enhancing reasoning quality\. Extensive experiments on established EI benchmarks show that PTEI enhances Emotional Understanding \(EU\) capabilities of various LLMs in EI, with the strongest improvement observed in GPT models, where combining PTEI withChain\-of\-Thought \(CoT\) reasoningyields an additional 4% increase in accuracy\. These findings underscore PTEI’s contribution toward advancing AI systems with more sophisticated social and psychological grounding\.
Emotional Intelligence, Emotion Detection and Analysis, Language Modeling, Personality Traits, Social Science, Large Language Models
## 1\.Introduction
Figure 1\.Illustration of PTEI’s impact on emotional inference in LLMs\. \(a\) An emotionally ambiguous scenario featuring Jacob\. \(b\) The LLM’s task: a multiple\-choice question asks for both Jacob’s emotion and its underlying cause\. \(c\) Personality trait information extracted for Jacob\. \(d\) The baseline LLM without personality knowledge misinterprets Jacob’s solitude assadness\. \(e\) Our PTEI framework correctly infers the emotion asjoy, recognizing that Jacob’s introverted and self\-reliant personality makes solitary time emotionally rewarding rather than lonely\.Emotional intelligence \(EI\), the ability to perceive, understand, regulate, and express emotions, is essential for effective communication, social interaction, and decision making\(Salovey and Mayer,[1990](https://arxiv.org/html/2607.10245#bib.bib1); Goleman,[1996](https://arxiv.org/html/2607.10245#bib.bib14); Hess and Bacigalupo,[2011](https://arxiv.org/html/2607.10245#bib.bib4)\)\. As Large Language Models \(LLMs\) are increasingly deployed in human\-facing applications, there is growing interest in evaluating and improving their emotional capabilities\(Wanget al\.,[2023](https://arxiv.org/html/2607.10245#bib.bib12)\)\. While recent studies show that models like GPT\-4 can perform well on tasks such as emotional awareness and understanding\(Elyosephet al\.,[2023](https://arxiv.org/html/2607.10245#bib.bib13)\), their abilities remain limited, particularly in scenarios involving implicit emotional situations or subjective interpretation\(Marufet al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib15)\)\. This presents an ongoing challenge in Natural Language processing \(NLP\): enabling LLMs to reliably interpret and reason about human emotions in context\. Enhancing EI in LLMs is therefore critical for more natural and effective human\-AI collaboration\.
Recent EI benchmarks such as EQ\-Bench\(Paech,[2023](https://arxiv.org/html/2607.10245#bib.bib11)\)and EmoBench\(Sabouret al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib27)\)offer structured evaluations of LLMs’ EI, but they still struggle with complex aspects such as emotional reasoning, regulation, and application in ambiguous social contexts\. While these benchmarks represent an important step forward, a key limitation is their lack of personal context; they overlook individual characteristics such as personality traits, which are known to significantly shape emotional inference and behavior\(Robinson and Clore,[2002](https://arxiv.org/html/2607.10245#bib.bib23); Sapet al\.,[2022](https://arxiv.org/html/2607.10245#bib.bib28)\)\. In psychology, personality traits refer to enduring individual differences in patterns of thinking, feeling, and behaving, as described by trait theory\(McCrae and Costa,[1997](https://arxiv.org/html/2607.10245#bib.bib24)\)\. This gap reflects an issue in how LLMs are typically prompted or fine\-tuned for EI tasks such as Emotional Understanding \(EU\): most current methodologies operate without incorporating sufficient personal context and tend to treat all emotional scenarios as one\-size\-fits\-all\. Hence, EI evaluations often remain surface\-level and fail to capture the individualized, psychologically grounded reasoning required for real\-world EU, as illustrated in Figure[1](https://arxiv.org/html/2607.10245#S1.F1)\.
Personality has long been recognized as a critical factor in shaping emotional perception and behavior\(McCrae and John,[1992](https://arxiv.org/html/2607.10245#bib.bib2); Myers,[1987](https://arxiv.org/html/2607.10245#bib.bib3)\)\. In humans, individual differences in traits such as openness, neuroticism, or extraversion are correlated with how emotions are interpreted, regulated, and expressed\(Izardet al\.,[1993](https://arxiv.org/html/2607.10245#bib.bib19)\)\. Recent work has shown that LLMs can exhibit consistent and measurable personality traits in their responses, and these traits can be shaped and aligned with desired profiles\(Safdariet al\.,[2023](https://arxiv.org/html/2607.10245#bib.bib16)\)\. Despite extensive research on personality prediction from text\(Stajner and Yenikent,[2020](https://arxiv.org/html/2607.10245#bib.bib6); Mehtaet al\.,[2020](https://arxiv.org/html/2607.10245#bib.bib7); Amirhosseini and Kazemian,[2020](https://arxiv.org/html/2607.10245#bib.bib9); Sorokovikovaet al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib8)\), most existing studies have either focused speaker characteristic for emotion recognition\(Fuet al\.,[2025](https://arxiv.org/html/2607.10245#bib.bib29)\)or explored the interaction between personality and emotion in narrow settings relying on a single framework \(e\.g\., OCEAN\(McCrae and John,[1992](https://arxiv.org/html/2607.10245#bib.bib2)\), MBTI\(Myers,[1987](https://arxiv.org/html/2607.10245#bib.bib3)\)\) and rarely addressing EI as a broader construct\(Wanget al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib30)\)\. Thus, the role of personality traits in enhancing the EI of LLMs remains largely underexplored\.
This paper proposesPTEI \(Personality Traits in Emotional Intelligence\), a novel framework to systematically integrates OCEAN and MBTI personality traits to enhance EI in LLMs\. PTEI extracts individual personality traits directly from textual scenarios and leverages this information through personality\-aware prompting to improve emotion and cause prediction\. Additionally, PTEI employs a Contrastive Learning\-based embedding method and a retrieval mechanism to identify emotionally and personally similar scenarios, which helps ground the model’s reasoning in psychologically aligned examples and improves its contextual sensitivity\. Our approach specifically targets implicit and ambiguous emotional scenarios, significantly improving LLMs’ capabilities in EI tasks and promoting more psychologically grounded inference strategies\.
In summary, our contributions are as follows:
- •We proposePTEI, first comprehensive framework utilizing MBTI and OCEAN personality traits into EU tasks for LLMs, addressing both type and trait theories of personality\.
- •We design an efficient, personality detection module that leverages structured few\-shot prompting to infer MBTI and OCEAN traits and incorporates this knowledge into customized prompts for emotion and cause prediction\.
- •We construct a synthetic memory bank of similar EU scenarios to the EI benchmark and enriched it with fine\-grained personality annotations\. We also introduce a personality\-aware Contrastive Learning \(CL\) objective to structure the scenario embedding space, enabling more effective retrieval of emotionally and personally aligned examples to support contextual reasoning in our few\-shot setup\.
- •We demonstrate that integrating personality traits via few\-shot and Chain\-of\-Thought \(CoT\) prompting enhances EI in LLMs, consistently outperforming personality\-agnostic baselines and substantially narrowing the performance gap to human\-level inference on challenging EI benchmarks\.
## 2\.Related Work
### 2\.1\.Personality\-based Methods
Analyzing personality traits from a psychological perspective plays a crucial role in understanding and predicting human behavior and emotions\. Among the various models, the Big Five personality framework \(known as OCEAN\), encompassing Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism\(McCrae and John,[1992](https://arxiv.org/html/2607.10245#bib.bib2)\), and the Myers\-Briggs Type Indicator \(MBTI\), based on four categories: Introversion versus Extraversion, Sensing vs Intuition, Thinking vs Feeling, and Judging vs Perceiving\(Myers,[1987](https://arxiv.org/html/2607.10245#bib.bib3)\), are two of the most widely used approaches for characterizing individual personality profiles\.
Personality prediction from text has emerged as a prominent task in NLP\(Stajner and Yenikent,[2020](https://arxiv.org/html/2607.10245#bib.bib6); Mehtaet al\.,[2020](https://arxiv.org/html/2607.10245#bib.bib7); Amirhosseini and Kazemian,[2020](https://arxiv.org/html/2607.10245#bib.bib9)\)\. Extensive research has focused on enhancing the detection of personality traits in human\-generated text using LLMs\(Sorokovikovaet al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib8)\)\. For example, PADO\(Yeoet al\.,[2025](https://arxiv.org/html/2607.10245#bib.bib31)\)introduces personality\-induced agents that estimate OCEAN trait levels by using GPT\-4o and LLaMA3\-8B\. Similarly, PsyCoT\(Yanget al\.,[2023](https://arxiv.org/html/2607.10245#bib.bib32)\)employs LLMs as AI assistants, utilizing a CoT approach based on specially designed questionnaires to facilitate personality inference\. Beyond identifying personality traits, it is crucial to understand their interplay with other cognitive functions, such as emotional processing, which is central to our work\.
Emotion features have been shown to enhance personality prediction performance in LLMs\(Liet al\.,[2025](https://arxiv.org/html/2607.10245#bib.bib33),[2022](https://arxiv.org/html/2607.10245#bib.bib10)\), while personality traits themselves serve as valuable features for emotion recognition, particularly in conversational scenarios; LaERC\-S\(Fuet al\.,[2025](https://arxiv.org/html/2607.10245#bib.bib29)\)exploits speaker characteristics to improve emotion prediction in dialogue and ERC\-DP\(Wanget al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib30)\)proposes a dynamic personality detection module that extracts OCEAN traits of a speaker from conversations rather than assuming static traits, thereby improving conversational emotion recognition\.
While prior work explores the interplay between emotions and personality and are often focus on either OCEAN or MBTI, we examine their combined impact on recognizing implicit emotional expressions, enabling more nuanced emotional inference through a fuller psychological profile\.

Figure 2\.PTEI system architecture\. First, the personality detection module extracts personality traits for both memory bank and test scenarios\. Then, the memory bank scenarios are encoded and used in a Contrastive Learning setup to generate the contrastive embedding library\[blue arrow\]\. For each test case, its encoded representation is used to retrieve similar scenarios from contrastive embedding\[green arrow\]\. Finally, the test scenario is combined with its personality traits and the retrieved examples to evaluate PTEI’s framework\[yellow arrow\]\.
### 2\.2\.Emotional Intelligence \(EI\)
EI, the ability to recognize, understand, and regulate emotions, is key in psychology and social computing\(Salovey and Mayer,[1990](https://arxiv.org/html/2607.10245#bib.bib1)\)\. As LLMs enter emotionally sensitive domains, EI has gained prominence in AI\. Early work\(Schuller and Schuller,[2018](https://arxiv.org/html/2607.10245#bib.bib18)\)identified emotion recognition, generation, and augmentation as pillars of Artificial Emotional Intelligence \(AEI\)\.
LLMs have achieved high performance on emotion\-related tasks in practical domains, such as emotion\-cause pair extraction using CoT prompting\(Wuet al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib17)\), and emotionally supportive dialogue generation via explicit strategy modeling\(Wanet al\.,[2025](https://arxiv.org/html/2607.10245#bib.bib34)\)\. Despite these promising results, rigorous evaluation of emotional reasoning in LLMs has remained limited\. Recent EI benchmarks like EmoBench\(Sabouret al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib27)\)and EQ\-Bench\(Paech,[2023](https://arxiv.org/html/2607.10245#bib.bib11)\)were developed to assess deeper EI capabilities such as EU, management, and social reasoning, thus these benchmarks show LLMs lag behind humans on EI tasks\.
However, existing EI benchmarks and systems often treat EI as a generic skill, overlooking individual\-level factors that shape emotional responses\. Personality traits, as defined by OCEAN and MBTI, are crucial in how emotions are perceived, interpreted, and expressed, yet remain underutilized in current LLM\-based EI evaluations\. We address this gap by introducing a personality\-aware framework that promotes more personalized and psychologically grounded emotional reasoning\. To enhance contextual grounding, we employ CL to build a retrieval system that surfaces emotionally and personally aligned scenarios\.
## 3\.Methodology
### 3\.1\.Problem Definition
Given a dataset of EU scenarios𝒮=\{s1,s2,…,sN\}\\mathcal\{S\}=\\\{s\_\{1\},s\_\{2\},\.\.\.,s\_\{N\}\\\}, each scenariosis\_\{i\}is a natural language description involving a subjectaia\_\{i\}who experiences one or more emotions in a specific context\. The task is to infer the pair\(e^i,r^i\)∈ℰ×ℛ\(\\hat\{e\}\_\{i\},\\hat\{r\}\_\{i\}\)\\in\\mathcal\{E\}\\times\\mathcal\{R\}, wheree^i\\hat\{e\}\_\{i\}denotes the predicted emotion label from a predefined set of emotion categoriesℰ\\mathcal\{E\}\(e\.g\.,grateful,anxious,frustrated\) andr^i\\hat\{r\}\_\{i\}denotes the corresponding predicted cause or trigger\. Each scenario is also assigned a category labelci∈𝒞c\_\{i\}\\in\\mathcal\{C\}, specifying the reasoning type required for interpretation\.
Our key innovation is conditioning inference on personality traits\. The objective is to learn a functionfEIf\_\{\\text\{EI\}\}whereP\(si,ai\)=\(Mi,Oi\)P\(s\_\{i\},a\_\{i\}\)=\(M\_\{i\},O\_\{i\}\)represent the personality profile for the subjectaia\_\{i\}in scenariosis\_\{i\},MiM\_\{i\}is the MBTI type andOiO\_\{i\}is the OCEAN profile\.
\(1\)fEI:\(si,ai,P\(ai\)\)↦\(e^i,r^i\)f\_\{\\text\{EI\}\}:\(s\_\{i\},a\_\{i\},P\(a\_\{i\}\)\)\\mapsto\(\\hat\{e\}\_\{i\},\\hat\{r\}\_\{i\}\)FunctionfEIf\_\{\\text\{EI\}\}maps a scenariosis\_\{i\}with subjectaia\_\{i\}and its associated subject’s personality profileP\(ai\)P\(a\_\{i\}\)to its corresponding emotion\-cause pair through context\-aware and personality\-informed reasoning\.
### 3\.2\.Architecture Overview
Figure[2](https://arxiv.org/html/2607.10245#S2.F2)illustrates the overall architecture of our proposed PTEI framework\. The system is designed to improve EI in LLMs by systematically incorporating psychological and contextual signals\. It consists of 3 primary components: \(1\) a personality detection module, \(2\) CL for embedding optimization and \(3\) scenario retrieval and final inference\. The pipeline begins by inferring the subject’s personality traitsP\(ai\)P\(a\_\{i\}\)\(both OCEAN and MBTI\) from input scenarios, which are used to construct personality\-aware prompts for the final inference LLM\. Moreover, a CL module optimizes the scenario embedding space from a memory bank by aligning similar emotion/personality pairs and separating dissimilar ones\. In parallel, top\-kksimilar scenarios are retrieved to support our base and CoT prompting setup\. Finally, these components feed into an LLM that performs joint emotion and cause prediction, enabling the model to reason across diverse and psychologically grounded contexts\.
### 3\.3\.Memory Bank Construction
To support few\-shot learning and enrich EU via contextual analogies, we construct a memory bank𝔹\\mathbb\{B\}ofN𝔹=500N\_\{\\mathbb\{B\}\}=500diverse EU scenarios, each annotated with an emotion label, its corresponding cause, and inferred personality traits\. Since manual creation of these complex scenarios is costly and requires expert annotation, we synthesized the scenarios using GPT\-4\(OpenAI,[2023](https://arxiv.org/html/2607.10245#bib.bib20)\)\. The generation process was guided by the high\-level task categories defined in EmoBench \(e\.g\., complex emotions, perspective\-taking\) to ensure coverage, but no scenarios, text segments, or lexical material were copied or paraphrased from EmoBench\. This design preserves structural similarity while maintaining full independence from the test set, thereby avoiding any risk of overlap or data leakage\. We define the memory bank𝔹\\mathbb\{B\}as:
\(2\)𝔹=\{\(si,ai,ei,ri,P\(ai\)\)\}i=1N𝔹\\mathbb\{B\}=\\\{\(s\_\{i\},a\_\{i\},e\_\{i\},r\_\{i\},P\(a\_\{i\}\)\)\\\}\_\{i=1\}^\{N\_\{\\mathbb\{B\}\}\}wheresis\_\{i\}is a generated scenario,aia\_\{i\}is the subject,eie\_\{i\}denotes emotion labels,rir\_\{i\}denotes cause labels, andP\(ai\)P\(a\_\{i\}\)denotes personality profile extracted for the subjectaia\_\{i\}in the scenario using the personality detection module\. This memory bank𝔹\\mathbb\{B\}serves as a reference library from which top\-kksimilar examples are retrieved based on proximity in the personality\-aware embedding space to a given test scenario\. The retrieved examplesSretrievedS\_\{\\text\{retrieved\}\}are then used to construct few\-shot prompts, enabling the LLM to perform personality\-aware emotional inference with improved contextual grounding\.
\(3\)Sretrieved\(sj\)=\{\(s^j,a^j,e^j,r^j,P\(a^j\)\)\}j=1kS\_\{\\text\{retrieved\}\}\(s\_\{j\}\)=\\\{\(\\hat\{s\}\_\{j\},\\hat\{a\}\_\{j\},\\hat\{e\}\_\{j\},\\hat\{r\}\_\{j\},P\(\\hat\{a\}\_\{j\}\)\)\\\}\_\{j=1\}^\{k\}
More results of the memory bank analysis are available in Appendix[A](https://arxiv.org/html/2607.10245#A1)\.
### 3\.4\.Evaluation of Generated Scenarios
To ensure the quality of the synthetic scenarios used in our memory bank for contrastive learning and few\-shot retrieval, we designed a hybrid evaluation pipeline combining LLM\-as\-a\-judge and human annotators\.
Scenario Generation Design\.To avoid overlap or potential leakage from EmoBench, scenario generation followed a structured prompt template developed independently of any benchmark items, as shown in Table[3](https://arxiv.org/html/2607.10245#A3.T3)in Appendix[C\.1](https://arxiv.org/html/2607.10245#A3.SS1)\. The template defines the scenario categories along with their descriptions and ensures consistency across generated examples\. Each generation request specifies a goal and category, with constraints on emotion and cause choices drawn from a predefined emotion list\. No scenarios or textual content from EmoBench were directly used in the prompts; only the overall structure and category taxonomy were adopted to maintain task alignment\. The generated output follows a JSON\-like schema with explicit emotion and cause fields, enabling automated parsing and validation in subsequent processing stages\.
LLM\-Based Quality Evaluation\.For each generated scenariosis\_\{i\}, we first performed an automated quality screening using two independent LLM evaluators, Gemma2 and Claude\. Each evaluatorm∈ℳm\\in\\mathcal\{M\}assigns a score vector over four criteria: coherence, category alignment, emotion\-label correctness, and cause\-label correctness\. Formally, the score vector is defined as:
\(4\)𝐪i\(m\)=\[qicoh,qicat,qiemo,qicause\]∈\[1,5\]4\.\\mathbf\{q\}\_\{i\}^\{\(m\)\}=\\left\[q\_\{i\}^\{\\text\{coh\}\},q\_\{i\}^\{\\text\{cat\}\},q\_\{i\}^\{\\text\{emo\}\},q\_\{i\}^\{\\text\{cause\}\}\\right\]\\in\[1,5\]^\{4\}\.
The aggregated LLM quality score for scenariosis\_\{i\}is computed as:
\(5\)QiLLM=1\|ℳ\|∑m∈ℳ14∑k=14qi,k\(m\),Q\_\{i\}^\{\\text\{LLM\}\}=\\frac\{1\}\{\|\\mathcal\{M\}\|\}\\sum\_\{m\\in\\mathcal\{M\}\}\\frac\{1\}\{4\}\\sum\_\{k=1\}^\{4\}q\_\{i,k\}^\{\(m\)\},whereℳ=\{Gemma2,Claude\}\\mathcal\{M\}=\\\{\\text\{Gemma2\},\\text\{Claude\}\\\}\. Only scenarios that passed the quality screening from both evaluators were retained for further consideration\.
Human Evaluation Setup\.From the filtered set of scenarios that passed automated screening, we randomly selected 10% for human evaluation\. Human annotators, with backgrounds in psychology and linguistics, assessed the scenarios in a blind setup, where synthetic scenarios were mixed with real examples from the EmoBench benchmark\. Each scenario was rated on the same four criteria using a 1–5 Likert scale, along with an additional binary overall acceptability judgment\.
For a human annotatorh∈ℋh\\in\\mathcal\{H\}, the average quality score and acceptability label are defined as:
\(6\)Qi\(h\)=14∑k=14qi,k\(h\),Ai\(h\)∈\{0,1\}\.Q\_\{i\}^\{\(h\)\}=\\frac\{1\}\{4\}\\sum\_\{k=1\}^\{4\}q\_\{i,k\}^\{\(h\)\},\\quad A\_\{i\}^\{\(h\)\}\\in\\\{0,1\\\}\.
Filtering and Acceptance Criteria\.A scenariosis\_\{i\}is accepted into the final memory bank𝔹\\mathbb\{B\}if it satisfies the following conditions:
\(7\)QiLLM≥τq∧1\|ℋ\|∑h∈ℋQi\(h\)≥τh∧1\|ℋ\|∑h∈ℋAi\(h\)=1,Q\_\{i\}^\{\\text\{LLM\}\}\\geq\\tau\_\{q\}\\;\\land\\;\\frac\{1\}\{\|\\mathcal\{H\}\|\}\\sum\_\{h\\in\\mathcal\{H\}\}Q\_\{i\}^\{\(h\)\}\\geq\\tau\_\{h\}\\;\\land\\;\\frac\{1\}\{\|\\mathcal\{H\}\|\}\\sum\_\{h\\in\\mathcal\{H\}\}A\_\{i\}^\{\(h\)\}=1,whereτq=τh=3\.5\\tau\_\{q\}=\\tau\_\{h\}=3\.5\. Scenarios failing to meet these criteria were discarded and regenerated\. This filtering process ensures that only scenarios validated by both automated and human evaluation are retained\.
To minimize stylistic or conceptual leakage from EmoBench, all scenario prompts were created independently, without direct reuse of benchmark content\. To further reduce bias and maintain representativeness, the scenario generation process was guided by the category distribution of the EmoBench test set used in downstream evaluation\. For each category, a proportional number of synthetic scenarios were generated and combined with diverse personality trait profiles drawn from both MBTI and OCEAN spaces\.
Agreement Analysis\.To assess the consistency between automated and human judgments, we compute inter\-annotator agreement using Cohen’s kappa coefficient:
\(8\)κ=po−pe1−pe,\\kappa=\\frac\{p\_\{o\}\-p\_\{e\}\}\{1\-p\_\{e\}\},wherepop\_\{o\}denotes the observed agreement andpep\_\{e\}denotes the expected agreement by chance\. Results indicate strong alignment between LLM and human judgments \(κ=0\.92\\kappa=0\.92\), supporting the reliability of automated evaluation at scale\.
Findings\.Overall, synthetic scenarios scored comparably to EmoBench items in terms of emotion\-label correctness and clarity, while exhibiting slightly lower plausibility in more abstract categories such asstrange story\. Human annotators confirmed that the retained scenarios are coherent, category\-consistent, and label\-appropriate, validating the use of the synthetic memory bank as a high\-quality resource\.
For the EmoBench test set used in downstream evaluation, we rely on the gold\-standard labels provided by the benchmark authors, which were validated by human experts in the original release\.
### 3\.5\.Personality Detection Module
To provide structured psychological context for each EU scenario, we implement a personality detection module that infers both MBTI and OCEAN profiles directly from scenario textsis\_\{i\}\. For subjectaia\_\{i\}insis\_\{i\}, the module outputs: \(1\) an MBTI typeMi∈ℳMBTIM\_\{i\}\\in\\mathcal\{M\}\_\{\\text\{MBTI\}\}, whereℳMBTI\\mathcal\{M\}\_\{\\text\{MBTI\}\}is a set of 16 Myers\-Briggs personality types\(Myers,[1987](https://arxiv.org/html/2607.10245#bib.bib3)\), and \(2\) a set of Big Five \(OCEAN\) trait levels\(McCrae and John,[1992](https://arxiv.org/html/2607.10245#bib.bib2)\)oi=\{oi\(1\),…,oi\(5\)\}o\_\{i\}=\\\{o\_\{i\}^\{\(1\)\},\.\.\.,o\_\{i\}^\{\(5\)\}\\\}, where eachoi\(j\)∈\{low,medium,high\}o\_\{i\}^\{\(j\)\}\\in\\\{\\textit\{low\},\\textit\{medium\},\\textit\{high\}\\\}\.
We adopt a prompt\-based annotation strategy using GPT\-4o\-mini, where each prompt is designed to elicit structured MBTI and OCEAN predictions based on the subject’s inferred behavior and contextual cues in the scenario\. This approach allows for efficient and scalable personality annotation without requiring manually labeled personality data for every scenario\. Full prompt templates and examples are provided in Appendix[C\.2](https://arxiv.org/html/2607.10245#A3.SS2)\.
### 3\.6\.Personality\-Aware Contrastive Learning
To enhance the ability of our framework to understand and reason about emotions within individual psychological contexts, we propose a personality\-aware contrastive learning approach\. It learns a robust scenario embedding spaceE:s↦z∈ℝdE:s\\mapsto z\\in\\mathbb\{R\}^\{d\}such that scenarios are embedded based on both emotional content and personality traits\.
#### Pair construction\.
We construct positive and negative pairs from our memory bank𝔹\\mathbb\{B\}\. Given a pair of scenarios\(\(si,ai\),\(sj,aj\)\)\(\(s\_\{i\},a\_\{i\}\),\(s\_\{j\},a\_\{j\}\)\):
- •Positive pair: if their emotion labels match\(ei=ej\)\(e\_\{i\}=e\_\{j\}\)and their personality profiles are similar\(Spersonality\(P\(ai\),P\(aj\)\)≥θs\)\(S\_\{\\text\{personality\}\}\(P\(a\_\{i\}\),P\(a\_\{j\}\)\)\\geq\\theta\_\{s\}\), where the similarity thresholdθs=0\.7\\theta\_\{s\}=0\.7\.
- •Negative pair: if their emotion labels differ\(ei≠ej\)\(e\_\{i\}\\neq e\_\{j\}\)or they share the same emotion label but have dissimilar personalities\(Spersonality\(P\(ai\),P\(aj\)\)<θd\)\(S\_\{\\text\{personality\}\}\(P\(a\_\{i\}\),P\(a\_\{j\}\)\)<\\theta\_\{d\}\), with dissimilarity thresholdθd=0\.3\\theta\_\{d\}=0\.3\.
#### Personality similarity metric\.
To quantify personality similaritySps=Spersonality\(P1,P2\)S\_\{\\text\{ps\}\}=S\_\{\\text\{personality\}\}\(P\_\{1\},P\_\{2\}\)between two profilesP1=\(M1,O1\)P\_\{1\}=\(M\_\{1\},O\_\{1\}\)andP2=\(M2,O2\)P\_\{2\}=\(M\_\{2\},O\_\{2\}\), we use the following composite metric:
\(9\)Sps=α⋅SMBTI\(M1,M2\)\+\(1−α\)⋅SOCEAN\(O1,O2\)S\_\{\\text\{ps\}\}=\\alpha\\cdot S\_\{\\text\{MBTI\}\}\(M\_\{1\},M\_\{2\}\)\+\(1\-\\alpha\)\\cdot S\_\{\\text\{OCEAN\}\}\(O\_\{1\},O\_\{2\}\)whereα\\alphabalances the influence of MBTI and OCEAN similarities \(defaultα=0\.5\\alpha=0\.5\)\.SMBTIS\_\{\\text\{MBTI\}\}is computed as the fraction of matching MBTI dimensions, andSOCEANS\_\{\\text\{OCEAN\}\}is the cosine similarity between trait vectors, with each trait mapped to a numeric scale:high= 1\.0,medium= 0\.5, andlow= 0\.0\.
#### Contrastive training\.
Scenarios are encoded into vector representations using a pretrained sentence encoder \(all\-mpnet\-base\-v2111[https://www\.sbert\.net/docs/pretrained\_models\.html](https://www.sbert.net/docs/pretrained_models.html)\) from the SentenceTransformers library to generate semantically and personally aligned embeddings from our training data in the memory bank\(Reimers and Gurevych,[2019](https://arxiv.org/html/2607.10245#bib.bib35)\)\. We selected this model due to its strong performance on semantic similarity tasks and its effectiveness in producing dense, high\-quality embeddings well\-suited for contrastive learning frameworks\. We fine\-tune the encoder on scenario pairs using a contrastive objective, normalized temperature\-scaled cross\-entropy \(NT\-Xent\) loss\(Chenet al\.,[2020](https://arxiv.org/html/2607.10245#bib.bib25)\):
\(10\)ℒcontrastive=−logexp\(sim\(i,j\)/τ\)∑k=1Nexp\(sim\(i,k\)/τ\)\\mathcal\{L\}\_\{\\text\{contrastive\}\}=\-\\log\\frac\{\\exp\(\\text\{sim\}\(i,j\)/\\tau\)\}\{\\sum\_\{k=1\}^\{N\}\\exp\(\\text\{sim\}\(i,k\)/\\tau\)\}Here,sim\(i,j\)\\text\{sim\}\(i,j\)is the cosine similarity between positive pair embeddings, andτ\\tau\(typically 0\.07\) is the temperature hyperparameter\.
#### Scenario retrieval\.
After training, the learned embedding space enables efficient retrieval of psychologically aligned examples through k\-nearest neighbor \(KNN\) search\(Johnsonet al\.,[2021](https://arxiv.org/html/2607.10245#bib.bib26)\)at inference\. These retrieved examples are then incorporated into few\-shot prompts to enhance the personality\-aware EU capability of the framework\.
Table 1\.Evaluation results on the Emotional Understanding \(EU\) task across LLMs of different sizes\. TheBase\(no personality or retrieval\) andCoT\(no personality or retrieval with CoT\) configurations are personality\-agnostic results reported from the EmoBench benchmark\. We compare these with our proposed methods:PTEI\-Base\(personality\-aware with retrieval\) andPTEI\-CoT\(personality\-aware with retrieval and CoT\)\. Results cover four reasoning categories \(CE, PBE, PT, EC\), along with Emotion, Cause, and Overall accuracy\.↑\\uparrowand↓\\downarrowindicate PTEI changes relative to Base or CoT, andboldvalues denote the best scores per model\.
## 4\.Experiments
### 4\.1\.Experimental Setup
All experiments were conducted using an NVIDIA A100 GPU with 40GB VRAM\. We used a combination of open and closed source LLMs, including GPT\-4\(OpenAI,[2023](https://arxiv.org/html/2607.10245#bib.bib20)\), LLaMA\-3\(Grattafioriet al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib22)\), and Qwen\(Baiet al\.,[2023](https://arxiv.org/html/2607.10245#bib.bib21)\), accessed through API interfaces\.
For the contrastive learning module \(Section[3\.6](https://arxiv.org/html/2607.10245#S3.SS6)\), we fine\-tuned a sentence encoder with a lightweight projection head using the NT\-Xent loss \(τ=0\.07\\tau=0\.07\)\. Scenario pairs were sampled from our personality\-enriched memory bank𝔹\\mathbb\{B\}\(Section[3\.3](https://arxiv.org/html/2607.10245#S3.SS3)\), withα=0\.5\\alpha=0\.5, a positive similarity threshold = 0\.7 and a negative threshold = 0\.3\. We trained the model with a batch size of 32, using the Adam optimizer \(learning rate = 2e\-5\), selecting the checkpoint with the best validation retrieval accuracy\.
### 4\.2\.Dataset
We primarily evaluated our PTEI framework on the EmoBench benchmark\(Sabouret al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib27)\)\. EmoBench features emotionally complex scenarios for EU and Emotional Application \(EA\) tasks\. We use 200 English multiple\-choice scenarios from the EU task, selected for their focus on inferring emotions and causes from rich textual descriptions\. Each scenario includes two questions: one on the subject’s primary emotion and another on its cause\. To support few\-shot prompting and retrieval, we also generated 500 synthetic scenarios using GPT\-4 \(Section[3\.3](https://arxiv.org/html/2607.10245#S3.SS3)\)\. These synthetic examples were annotated with emotion labels, causes, and inferred MBTI and OCEAN personality traits, forming a personality\-enriched memory bank𝔹\\mathbb\{B\}to improve emotional reasoning during inference\.
### 4\.3\.Baselines
To evaluate the effectiveness of the PTEI framework, we analyze performance across a selection of LLMs \(see Section[4\.1](https://arxiv.org/html/2607.10245#S4.SS1)\)\. For each model, we implement two configurations:PTEI\-Base, which applies personality\-aware few\-shot prompting using retrieved examples, andPTEI\-CoT, which extends this with Chain\-of\-Thought \(CoT\) reasoning\. We compare these with corresponding non\-personality\-aware variants:Base\(zero\-shot\) andCoT\(zero\-shot with CoT\), both excluding personality conditioning\. Our comparisons include the best\-performing LLMs from EmoBench as personality\-agnostic baselines, representing small\-scale \(¡14B\), mid\-scale \(14B\), and large\-scale \(¿14B\) models, providing a robust reference for quantifying the added value of incorporating personality traits\.
TheBaseandCoTresults are based on those reported in the EmoBench benchmark\(Sabouret al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib27)\), ensuring consistency with prior evaluation protocols\. The only exceptions areGPT\-4oandLLaMA 3\.1 8B\. ForGPT\-4o, we replace the originalGPT\-4used in EmoBench to evaluate PTEI’s performance under the latest state\-of\-the\-art model\. ForLLaMA 3\.1 8B, which was not part of EmoBench, we reproduce the benchmark evaluation process to generate comparableBaseandCoTresults before applying our PTEI framework\.
We evaluate across four EU task categories from EmoBench:Complex Emotions \(CE\),Personal Beliefs and Experiences \(PBE\),Perspective\-Taking \(PT\), andEmotional Cues \(EC\)\. Detailed definitions of these categories are provided in Appendix[B](https://arxiv.org/html/2607.10245#A2)\.
## 5\.Result and Analysis
In this section, we comprehensively evaluate the proposed PTEI framework through experiments covering main results \(Section[5\.1](https://arxiv.org/html/2607.10245#S5.SS1)\), ablation studies \(Section[5\.2](https://arxiv.org/html/2607.10245#S5.SS2)\), robustness analysis \(Section[5\.3](https://arxiv.org/html/2607.10245#S5.SS3)\), and qualitative case studies \(Section[5\.4](https://arxiv.org/html/2607.10245#S5.SS4)\)\.
### 5\.1\.Main Results
Table[1](https://arxiv.org/html/2607.10245#S3.T1)presents evaluation results \(accuracy\) for the EU task across various LLMs and prompting configurations\.GPT\-4o achieves the highest combined accuracy \(63\.62%\)under the personality\-informed CoT configuration \(PTEI\-CoT\), significantly outperforming all other models\. In contrast, smaller models such asQwen\-7BandLLaMA3\.1\-8Bgenerally struggle to surpass majority\-class heuristics, particularly under standard zero\-shot or CoT prompting\.
Notably,CoT prompting alone yields limited or negative effectson smaller\-scale models\. For instance, accuracy forLLaMA3\.1\-8Bdecreases notably from the Base \(16\.62%\) to CoT \(12\.25%\)\. This suggests constrained structured reasoning abilities without sufficient grounding\. However,integrating personality context via PTEI\-CoT considerably enhances CoT effectiveness, particularly for larger models such asGPT\-4o, which improves by \+3\.3 points over its base and \+4\.7 points over standard CoT\.
Performance across the four EU categories:Complex Emotions \(CE\),Personal Beliefs and Experiences \(PBE\),Perspective\-Taking \(PT\), andEmotional Cues \(EC\), demonstrates thatPT and PBE scenarios consistently pose the greatest challenges\. These tasks demand reasoning about mental states and personal beliefs\. Conversely, scenarios categorized as CE and EC, which contain more explicit emotional indicators, yield comparatively higher accuracy\.
Lastly, model scale notably impacts performance:larger models like GPT\-4onot only achieve higher baseline accuracy but alsoshow greater relative improvements from personality\-aware prompting\. This emphasizes the synergy between increased model capacity, structured reasoning \(CoT\), and psychologically grounded personality context in enhancing emotional reasoning capabilities\.
### 5\.2\.Ablation Study
Figure 3\.Average accuracy under different personality input configurations\. Injecting both MBTI and OCEAN traits leads to the highest performance across models\.
Figure 4\.Case study comparison for GPT\-4o, showing how personality\-aware prompting improves emotional prediction accuracy\.To evaluate the effect of personality conditioning, we conduct an ablation study on the Personality Detection Module \(Section[3\.5](https://arxiv.org/html/2607.10245#S3.SS5)\)\. Four LLMs are tested:GPT\-4o,Qwen\-7B,Qwen\-14B, andLLaMA 3\.1\-8Bunder four input settings: \(1\) no personality, \(2\) MBTI\-only, \(3\) OCEAN\-only, and \(4\) MBTI\+OCEAN \(full PTEI\)\. Models are assessed on emotion and cause prediction, reporting average accuracy across both\. The personality weighting parameterα\\alphacontrols the influence of MBTI and OCEAN in the similarity function:α=1\.0\\alpha=1\.0\(MBTI\-only\),α=0\.0\\alpha=0\.0\(OCEAN\-only\), andα=0\.5\\alpha=0\.5\(combined\)\. This weighting affects both contrastive training and retrieval during inference\.
As shown in Figure[3](https://arxiv.org/html/2607.10245#S5.F3), personality information consistently improves results over the no\-personality baseline\.GPT\-4oattains the highest accuracy with MBTI\+OCEAN, surpassing MBTI\-only and no\-personality by \+1\.29% and \+1\.87%, respectively\.Qwen\-14Bshows similar gains, with the combined setup outperforming OCEAN\-only and no\-personality by \+2\.92% and \+0\.62%\. While MBTI\-only and OCEAN\-only remain competitive, the combined approach yields the best overall performance across all model scales\.
### 5\.3\.Robustness Analysis
Table 2\.Ablation results with and without CoT\. The*TraitsOnly*rows correspond to personality traits injected without retrieval, while the*RAG\-only*rows include retrieval examples but omit personality traits\. Bold indicates the best performance within each block\.To evaluate the stability of model predictions, we prompt each LLM three times per question and apply majority voting over the responses\. To reduce sensitivity to answer ordering, we additionally permute the multiple\-choice options three times, yielding four total permutations \(original \+ 3\)\. Final accuracy is averaged across these runs\.
We further perform an ablation to isolate the effect of personality traits and retrieval\. As shown in Table[2](https://arxiv.org/html/2607.10245#S5.T2), we compare four variants under both without\-CoT and with\-CoT setups:*Base, CoT*\(no traits or retrieval\),*TraitsOnly*\(traits only, zero\-shot\),*RAG\-only*\(retrieval only, two\-shot without traits\), and*PTEI base/CoT*\(traits \+ retrieval\)\. Results show that traits and retrieval each contribute improvements over the base, while their combination yields the best performance\.
### 5\.4\.Case Study: Qualitative Analysis
To illustrate the reasoning process of the PTEI framework, we analyze a representative scenario \(Figure[4](https://arxiv.org/html/2607.10245#S5.F4)\) featuring ’Helena’, who recently experienced a breakup and finds that her mother has torn up a letter from her ex\. The emotional inference depends on sentimental value, perceived intrusion, and the mother’s emotional intent\.
Without personality awareness, baseline LLMs interpret the act as a boundary violation, predictingAngerand emphasizing intrusion\. In contrast, personality\-informed models \(PTEI\-Base and PTEI\-CoT\) leverage Helena’s ISFJ type, high Agreeableness, and low Neuroticism to contextualize her response as family\-oriented and emotionally stable\. PTEI\-CoT further infers the mother’s protective intent, shielding her daughter from pain, leading to the correct prediction ofGratitudeand a cause aligned with emotional intent rather than surface cues\.
This case demonstrates how personality conditioning and structured reasoning allow PTEI to interpret emotionally complex scenarios with greater human\-like precision in both emotion and cause prediction\.
## 6\.Conclusion
This paper tackled the limitations of current LLMs in complex emotional intelligence tasks, particularly their neglect of psychological factors such as personality\. We introduced PTEI, a personality\-aware emotional inference framework that integrates MBTI and OCEAN traits into emotion understanding \(EU\) scenarios\. The framework employs structured prompting for personality detection, builds a synthetic memory bank with personality annotations, and leverages contrastive learning to enhance scenario retrieval and inference\. Experiments show that incorporating personality traits through few\-shot and CoT prompting notably improves emotion and cause reasoning, surpassing personality\-agnostic baselines\. By shaping an embedding space sensitive to both emotional and psychological cues, PTEI enables more human\-aligned emotional inference\. Future work will extend the framework to dynamic personality modeling and multi\-turn EI reasoning\.
## 7\.Limitations
While our PTEI framework demonstrates significant improvements in EU through personality\-aware reasoning, several limitations remain\.
First, our experiments are limited to English\-language, text\-based scenarios\. Emotional reactions and personality interpretations may vary across cultures and languages, which constrains the generalizability of our framework to multilingual or multimodal settings\. Additionally, emotional inference is inherently subjective\. While we adopt multiple\-choice evaluation formats with predefined correct answers, some scenarios may reasonably allow for multiple plausible interpretations\.
Second, due to the limited number of test cases \(200 EU scenarios\) in the EmoBench benchmark, the reported results may not fully capture the generalizability and impact of our method across the full spectrum of emotionally intelligent behavior\. We plan to expand our evaluation to a broader range of benchmarks and real\-world EI tasks in future work\.
Third, despite careful prompt engineering for personality\-aware reasoning and CoT prompting, model performance remains sensitive to prompt phrasing and structure\. The prompt templates used may not be optimal for all LLM architectures or tasks\.
Finally, our use of MBTI and OCEAN frameworks offers only a coarse approximation of personality\. MBTI, in particular, has been criticized for its limited empirical robustness and test\-retest reliability\. We adopt it here primarily as a widely used and interpretable typology that provides categorical variation for computational experiments, while relying on OCEAN traits for a more established psychological foundation\. Moreover, the inference of MBTI and OCEAN traits from short text scenarios is inherently approximate: such profiles should be seen as pragmatic proxies that enable structured reasoning, rather than clinically validated ground\-truth traits\. Both frameworks also remain static abstractions that do not capture dynamic, context\-sensitive aspects of personality\. Future work could explore adaptive or conversational personality modeling to more accurately reflect real\-world psychological variability\.
## Ethical Statement
This work investigates the use of personality traits to improve large language models’ \(LLMs\) reasoning about emotionally complex situations\. We emphasize that our framework, PTEI, focuses on modelingperceivedemotional intelligence through structured prompts and retrieval\-based reasoning, rather than suggesting that LLMs possess genuine emotions or self\-awareness\. The goal is to study how personality knowledge can inform more human\-aligned predictions in emotion and cause inference tasks\.
While our method involves predicting MBTI and OCEAN personality traits based on textual descriptions, we do not use real user data, nor do we attempt to profile individuals in real\-world settings\. All personality information is synthetic and used strictly for academic experimentation in controlled benchmark scenarios\.
We acknowledge that the use of personality modeling in NLP may raise ethical concerns, particularly around privacy, profiling, and fairness in downstream applications\. Appropriate safeguards must be ensured in future deployments, including transparency, user consent, and bias auditing\. Our current system is intended solely for research purposes, and we advocate for careful consideration of the psychological and social implications when applying similar methods in real\-world contexts\.
## References
- M\. H\. Amirhosseini and H\. Kazemian \(2020\)Machine learning approach to personality type prediction based on the myers–briggs type indicator®\.Multimodal Technologies and Interaction4\(1\),pp\. 9\.External Links:[Link](https://doi.org/10.3390/mti4010009),[Document](https://dx.doi.org/10.3390/MTI4010009)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p2.1)\.
- J\. Bai, S\. Bai, Y\. Chu, Z\. Cui, K\. Dang, X\. Deng, Y\. Fan, W\. Ge, Y\. Han, F\. Huang, and et al\. \(2023\)Qwen technical report\.arXiv preprint arXiv:2309\.16609\.External Links:[Link](https://arxiv.org/abs/2309.16609),[Document](https://dx.doi.org/10.48550/arXiv.2309.16609)Cited by:[§4\.1](https://arxiv.org/html/2607.10245#S4.SS1.p1.1)\.
- T\. Chen, S\. Kornblith, M\. Norouzi, and G\. Hinton \(2020\)A simple framework for contrastive learning of visual representations\.InProceedings of the 37th International Conference on Machine Learning,pp\. 1597–1607\.External Links:[Link](https://dl.acm.org/doi/10.5555/3524938.3525087),[Document](https://dx.doi.org/10.5555/3524938.3525087)Cited by:[§3\.6](https://arxiv.org/html/2607.10245#S3.SS6.SSS0.Px3.p1.3)\.
- Z\. Elyoseph, D\. Hadar\-Shoval, K\. Asraf, and M\. Lvovsky \(2023\)ChatGPT outperforms humans in emotional awareness evaluations\.Frontiers in Psychology14,pp\. 1199058\.External Links:[Document](https://dx.doi.org/10.3389/fpsyg.2023.1199058),[Link](https://doi.org/10.3389/fpsyg.2023.1199058)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p1.1)\.
- Y\. Fu, J\. Wu, Z\. Wang, M\. Zhang, L\. Shan, Y\. Wu, and B\. Liu \(2025\)LaERC\-S: improving LLM\-based emotion recognition in conversation with speaker characteristics\.InProceedings of the 31st International Conference on Computational Linguistics,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),Abu Dhabi, UAE,pp\. 6748–6761\.External Links:[Link](https://aclanthology.org/2025.coling-main.451/)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p3.1)\.
- D\. Goleman \(1996\)Emotional intelligence\. why it can matter more than iq\.\.Learning24\(6\),pp\. 49–50\.Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, and et al\. \(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.External Links:[Link](https://arxiv.org/abs/2407.21783),[Document](https://dx.doi.org/10.48550/arXiv.2407.21783)Cited by:[§4\.1](https://arxiv.org/html/2607.10245#S4.SS1.p1.1)\.
- J\. D\. Hess and A\. C\. Bacigalupo \(2011\)Enhancing decisions and decision\-making processes through the application of emotional intelligence skills\.Management decision49\(5\),pp\. 710–721\.External Links:[Document](https://dx.doi.org/10.1108/00251741111130805)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p1.1)\.
- C\. E\. Izard, D\. Z\. Libero, P\. Putnam, and O\. M\. Haynes \(1993\)Stability of emotion experiences and their relations to traits of personality\.Journal of Personality and Social Psychology64\(5\),pp\. 847–860\.External Links:[Document](https://dx.doi.org/10.1037//0022-3514.64.5.847),[Link](https://doi.org/10.1037//0022-3514.64.5.847)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1)\.
- J\. Johnson, M\. Douze, and H\. Jégou \(2021\)Billion\-scale similarity search with gpus\.Vol\.7\.External Links:[Document](https://dx.doi.org/10.1109/TBDATA.2019.2921572)Cited by:[§3\.6](https://arxiv.org/html/2607.10245#S3.SS6.SSS0.Px4.p1.1)\.
- Y\. Li, A\. Kazemeini, Y\. Mehta, and E\. Cambria \(2022\)Multitask learning for emotion and personality traits detection\.Neurocomputing493,pp\. 340–350\.External Links:[Link](https://doi.org/10.1016/j.neucom.2022.04.049),[Document](https://dx.doi.org/10.1016/J.NEUCOM.2022.04.049)Cited by:[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p3.1)\.
- Z\. Li, S\. Li, D\. Zhu, Q\. Ma, and W\. Xiong \(2025\)EERPD: leveraging emotion and emotion regulation for improving personality detection\.InProceedings of the 31st International Conference on Computational Linguistics,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),Abu Dhabi, UAE,pp\. 7721–7734\.External Links:[Link](https://aclanthology.org/2025.coling-main.516/)Cited by:[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p3.1)\.
- A\. A\. Maruf, F\. Khanam, Md\. M\. Haque, Z\. M\. Jiyad, M\. F\. Mridha, and Z\. Aung \(2024\)Challenges and opportunities of text\-based emotion detection: a survey\.IEEE Access12\(\),pp\. 18416–18450\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2024.3356357)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p1.1)\.
- R\. R\. McCrae and O\. P\. John \(1992\)An introduction to the five\-factor model and its applications\.Journal of personality60\(2\),pp\. 175–215\.External Links:[Document](https://dx.doi.org/10.1111/j.1467-6494.1992.tb00970.x)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p1.1),[§3\.5](https://arxiv.org/html/2607.10245#S3.SS5.p1.7)\.
- R\. R\. McCrae and P\. T\. Jr\. Costa \(1997\)Personality trait structure as a human universal\.American Psychologist52\(5\),pp\. 509–516\.External Links:[Document](https://dx.doi.org/10.1037//0003-066x.52.5.509),[Link](https://doi.org/10.1037//0003-066x.52.5.509)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p2.1)\.
- Y\. Mehta, N\. Majumder, A\. Gelbukh, and E\. Cambria \(2020\)Recent trends in deep learning based personality detection\.Artificial Intelligence Review53\(4\),pp\. 2313–2339\.External Links:[Link](https://doi.org/10.1007/s10462-019-09770-z),[Document](https://dx.doi.org/10.1007/S10462-019-09770-Z)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p2.1)\.
- I\. B\. Myers \(1987\)Introduction to type: a description of the theory and applications of the myers\-briggs type indicator\.Consulting Psychologists Press\.Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p1.1),[§3\.5](https://arxiv.org/html/2607.10245#S3.SS5.p1.7)\.
- OpenAI \(2023\)GPT\-4 technical report\.arXiv preprint arXiv:2303\.08774\.External Links:[Link](https://arxiv.org/abs/2303.08774),[Document](https://dx.doi.org/10.48550/arXiv.2303.08774)Cited by:[§3\.3](https://arxiv.org/html/2607.10245#S3.SS3.p1.3),[§4\.1](https://arxiv.org/html/2607.10245#S4.SS1.p1.1)\.
- S\. J\. Paech \(2023\)Eq\-bench: an emotional intelligence benchmark for large language models\.arXiv preprint arXiv:2312\.06281\.External Links:[Link](https://doi.org/10.48550/arXiv.2312.06281),[Document](https://dx.doi.org/10.48550/ARXIV.2312.06281)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.10245#S2.SS2.p2.1)\.
- N\. Reimers and I\. Gurevych \(2019\)Sentence\-BERT: sentence embeddings using Siamese BERT\-networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 3982–3992\.External Links:[Link](https://aclanthology.org/D19-1410/),[Document](https://dx.doi.org/10.18653/v1/D19-1410)Cited by:[§3\.6](https://arxiv.org/html/2607.10245#S3.SS6.SSS0.Px3.p1.3)\.
- M\. D\. Robinson and G\. L\. Clore \(2002\)Belief and feeling: evidence for an accessibility model of emotional self\-report\.Psychological Bulletin128\(6\),pp\. 934–960\.External Links:[Document](https://dx.doi.org/10.1037/0033-2909.128.6.934)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p2.1)\.
- S\. Sabour, S\. Liu, Z\. Zhang, J\. Liu, J\. Zhou, A\. Sunaryo, T\. Lee, R\. Mihalcea, and M\. Huang \(2024\)EmoBench: evaluating the emotional intelligence of large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 5986–6004\.External Links:[Link](https://aclanthology.org/2024.acl-long.326/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.326)Cited by:[Appendix B](https://arxiv.org/html/2607.10245#A2.p1.1),[§1](https://arxiv.org/html/2607.10245#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.10245#S2.SS2.p2.1),[§4\.2](https://arxiv.org/html/2607.10245#S4.SS2.p1.1),[§4\.3](https://arxiv.org/html/2607.10245#S4.SS3.p2.1)\.
- M\. Safdari, G\. Serapio\-García, C\. Crepy, S\. Fitz, P\. Romero, L\. Sun, M\. Abdulhai, A\. Faust, and M\. J\. Mataric \(2023\)Personality traits in large language models\.CoRRabs/2307\.00184\.External Links:[Link](https://doi.org/10.48550/arXiv.2307.00184),[Document](https://dx.doi.org/10.48550/ARXIV.2307.00184),2307\.00184Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1)\.
- P\. Salovey and J\. D\. Mayer \(1990\)Emotional intelligence\.Imagination, Cognition and Personality9\(3\),pp\. 185–211\.External Links:[Document](https://dx.doi.org/10.2190/DUGG-P24E-52WK-6CDG),[Link](https://doi.org/10.2190/DUGG-P24E-52WK-6CDG)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p1.1),[§2\.2](https://arxiv.org/html/2607.10245#S2.SS2.p1.1)\.
- M\. Sap, R\. Le Bras, D\. Fried, and Y\. Choi \(2022\)Neural theory\-of\-mind? on the limits of social intelligence in large LMs\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 3762–3780\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.248/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.248)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p2.1)\.
- D\. Schuller and B\. W\. Schuller \(2018\)The age of artificial emotional intelligence\.Computer51\(9\),pp\. 38–46\.External Links:[Document](https://dx.doi.org/10.1109/MC.2018.3620963)Cited by:[§2\.2](https://arxiv.org/html/2607.10245#S2.SS2.p1.1)\.
- A\. Sorokovikova, N\. Fedorova, S\. Rezagholi, and I\. P\. Yamshchikov \(2024\)Llms simulate big five personality traits: further evidence\.arXiv preprint arXiv:2402\.01765\.External Links:[Link](https://doi.org/10.48550/arXiv.2402.01765),[Document](https://dx.doi.org/10.48550/ARXIV.2402.01765)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p2.1)\.
- S\. Stajner and S\. Yenikent \(2020\)A survey of automatic personality detection from texts\.InProceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain \(Online\), December 8 13,2020,pp\. 6284–6295\.External Links:[Link](https://doi.org/10.18653/v1/2020.coling-main.553),[Document](https://dx.doi.org/10.18653/V1/2020.COLING-MAIN.553)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p2.1)\.
- C\. Wan, M\. Labeau, and C\. Clavel \(2025\)EmoDynamiX: emotional support dialogue strategy prediction by modelling MiXed emotions and discourse dynamics\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 1678–1695\.External Links:[Link](https://aclanthology.org/2025.naacl-long.81/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.81),ISBN 979\-8\-89176\-189\-6Cited by:[§2\.2](https://arxiv.org/html/2607.10245#S2.SS2.p2.1)\.
- X\. Wang, X\. Li, Z\. Yin, Y\. Wu, and J\. Liu \(2023\)Emotional intelligence of large language models\.Journal of Pacific Rim Psychology17\.External Links:[Document](https://dx.doi.org/10.1177/18344909231213958),[Link](https://doi.org/10.1177/18344909231213958)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p1.1)\.
- Y\. Wang, B\. Wang, Y\. Zhao, D\. Zhao, X\. Jin, J\. Zhang, R\. He, and Y\. Hou \(2024\)Emotion recognition in conversation via dynamic personality\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 5711–5722\.External Links:[Link](https://aclanthology.org/2024.lrec-main.507/)Cited by:[§1](https://arxiv.org/html/2607.10245#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p3.1)\.
- J\. Wu, Y\. Shen, Z\. Zhang, and L\. Cai \(2024\)Enhancing large language model with decomposed reasoning for emotion cause pair extraction\.arXiv preprint arXiv:2401\.17716\.External Links:[Link](https://arxiv.org/abs/2401.17716),[Document](https://dx.doi.org/10.48550/arXiv.2401.17716)Cited by:[§2\.2](https://arxiv.org/html/2607.10245#S2.SS2.p2.1)\.
- T\. Yang, T\. Shi, F\. Wan, X\. Quan, Q\. Wang, B\. Wu, and J\. Wu \(2023\)PsyCoT: psychological questionnaire as powerful chain\-of\-thought for personality detection\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 3305–3320\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.216/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.216)Cited by:[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p2.1)\.
- H\. Yeo, T\. Noh, S\. Jin, and K\. Han \(2025\)PADO: personality\-induced multi\-agents for detecting OCEAN in human\-generated texts\.InProceedings of the 31st International Conference on Computational Linguistics,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),Abu Dhabi, UAE,pp\. 5719–5736\.External Links:[Link](https://aclanthology.org/2025.coling-main.382/)Cited by:[§2\.1](https://arxiv.org/html/2607.10245#S2.SS1.p2.1)\.
## Appendix AMemory Bank Analysis
### A\.1\.Emotion Label Distribution
We analyzed the normalized distribution of emotions of the memory bank\. Labels were grouped into high\-level categories and split if mixed\.
Figure 5\.Flattened Emotion Category Distribution \(Normalized\)
### A\.2\.Emotion\-Personality Correlation
We computed Pearson correlations between normalized emotion categories and personality traits using both MBTI and OCEAN frameworks\. Emotion labels were one\-hot encoded, and personality traits were encoded either as binary \(MBTI\) or ordinal \(OCEAN\)\.
MBTI Correlation:Each of the 16 standard MBTI types \(e\.g\., INFP, ESTJ\) was represented as a binary feature\. We then calculated the Pearson correlation between these types and each individual emotion category\. This analysis highlights patterns such as higher emotional resonance of Joy with ENFP and stronger ties between Apprehension and ISFP types\.
Figure 6\.MBTI vs Emotion CorrelationOCEAN Correlation:OCEAN traits were originally qualitative and mapped numerically: Low = 1, Medium = 2, High = 3\. We computed correlations between each trait and each emotion category\. For instance, Neuroticism shows a positive correlation with emotions like Fear and Nervousness, while Agreeableness aligns more strongly with Trust and Caring emotions\.
Figure 7\.OCEAN Traits vs Emotion Correlation
## Appendix BEU Task Categories
We evaluate our framework across four emotional understanding \(EU\) task categories defined in EmoBench\(Sabouret al\.,[2024](https://arxiv.org/html/2607.10245#bib.bib27)\):
- •Complex Emotions \(CB\): This task category evaluates a model’s ability to reason about nuanced and layered emotional experiences\. It includes scenarios involving \(1\) emotional transitions in response to evolving events, \(2\) mixtures of emotions with potentially conflicting valence \(e\.g\., joy and disappointment\), and \(3\) unexpected emotional outcomes that challenge commonsense assumptions\. These situations require deeper inferential reasoning to correctly interpret how emotions manifest across dynamic or atypical contexts\.
- •Personal Beliefs and Experiences \(PBE\): This category evaluates how well a model understands that emotional responses are shaped by a person’s cultural values, sentimental attachments, and prior experiences\. Scenarios may involve culturally influenced norms \(e\.g\., attitudes toward punctuality\), differing levels of sentimental value assigned to objects or events, or reactions driven by personal traits and experiences, such as phobias or personality styles\. Accurate prediction in this category requires sensitivity to how internal beliefs and past experiences modulate emotional appraisal\.
- •Perspective\-Taking \(PT\): This category focuses on the model’s ability to simulate others’ emotional states by reasoning about their beliefs, knowledge, and social context\. It adapts classical theory\-of\-mind tasks: False Belief, Faux Pas, and Strange Stories to scenarios where emotional inference depends on understanding what different individuals know or assume\. For instance, detecting excitement in someone misled by false information, or recognizing that embarrassment does not occur when a faux pas is unknown to the speaker, requires the model to attribute distinct perspectives to each character in the scenario\.
- •Emotional Cues \(EC\): This task assesses a model’s ability to infer emotions from implicit textual descriptions of vocal and visual cues, such as tone, speech patterns, facial expressions, or bodily gestures\. Unlike explicit emotional statements, these scenarios require the model to recognize subtle indicators\. For instance, interpreting a sigh as annoyance or relief, or a flushed face as either anger or embarrassment, based on context\. This category tests whether LLMs can perceive and reason about nonverbal affective signals conveyed through text\.
## Appendix CPrompt Templates
This section outlines the prompt configurations used in our PTEI framework\. Prompts are grouped by their corresponding modules\.
### C\.1\.Prompts for Scenario Generation
The prompt template for scenario generation was developed independently of any benchmark items, as shown in Table[3](https://arxiv.org/html/2607.10245#A3.T3)\.
Table 3\.Scenario generation prompt template with structured goals, categories, and output fields\.
### C\.2\.Prompts for Personality Detection Module
These prompts are used to infer personality traits for each subject within the scenario\.
- •MBTI Personality Prompt: Table[4](https://arxiv.org/html/2607.10245#A3.T4)
- •OCEAN Personality Prompt: Table[5](https://arxiv.org/html/2607.10245#A3.T5)
### C\.3\.Prompts for EU Task
These prompts are designed to guide LLMs in predicting emotions and their causes for each scenario, with or without step\-by\-step reasoning\. They incorporate contextual retrieval and personality traits\.
- •PTEI\-Base Prompt: Table[6](https://arxiv.org/html/2607.10245#A3.T6)
- •PTEI\-CoT Prompt: Table[7](https://arxiv.org/html/2607.10245#A3.T7)
Table 4\.MBTI prompt template with structured instructions and two\-shot demonstration format\.Table 5\.OCEAN prompt template used for personality trait inference, including structured trait outputs, reasoning explanation, and a two\-shot demonstration format\.Table 6\.PTEI\-Base prompt template for emotion and cause prediction, incorporating memory\-based retrieval and structured personality conditioning \(MBTI \+ OCEAN\)\.Table 7\.PTEI\-CoT prompt template for emotion and cause prediction, combining retrieval, personality traits, and step\-by\-step reasoning\. The model is expected to output both a selected answer and an explanation\.Similar Articles
Beyond Static Personas: Situational Personality Steering for Large Language Models
This paper introduces IRiS, a training-free framework for situational personality steering in LLMs that moves beyond static persona modeling by identifying and leveraging situation-dependent persona neurons. The approach demonstrates that LLM behavior varies contextually and proposes neuron-based identification, retrieval, and weighted steering methods validated on PersonalityBench and a new SPBench benchmark.
Persona Cartography: Charting Language Model Personality Traits in Weight Space
This paper introduces a method for analyzing and controlling language model personalities using the OCEAN framework, training low-rank adapters to manipulate traits, and demonstrates additive composition and safety implications.
Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions
This paper introduces a mechanistic interpretability approach to steer LLM personality traits by identifying and intervening on latent features using sparse autoencoders, achieving controllable personality modulation while maintaining language performance.
How Well Do Large Language Models Capture Human Personality?
This paper systematically evaluates assumptions about LLM persona prompting and identifies 'persona manifold collapse,' where richer persona descriptions reduce behavioral diversity and simulation fidelity. The findings show that simple age-gender personas often outperform more detailed profiles.
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
This paper examines when and why self-reported psychometric measures predict the actual behavior of large language models, finding that fine-grained, behavior-specific instruments (Theory of Planned Behavior) achieve human-level coherence within a shared conversation, while broad traits like Big 5 do not.