Semantic Primes as Explanans for Emotion in Large Language Models

arXiv cs.AI Papers

Summary

This paper proposes using Natural Semantic Metalanguage (NSM) primes as a more basic explanation for emotions in LLMs, showing they are recoverable, causally controllable, and faithful, outperforming appraisal-based directions.

arXiv:2607.18691v1 Announce Type: new Abstract: Progresses have been made on understanding emotion mechanisms of large language models (LLMs). However, how to explain emotion in LLMs, or even what constitutes good explanations, are less clear. Emotion representations, components, circuits are widely recoverable, but as explanations of a model's own computation they are circular; the emotion space dimensions tend to be arbitrary and non-terminating. A pressing question to ask is whether a more primitive set of internal variables does the work: the semantic primes of the Natural Semantic Metalanguage (NSM). Across four instruction-tuned LLMs (Llama-1B, Gemma-2B, Gemma-9B, OLMo-7B), experiments show that the NSM primes are (1) recoverable internal elements; and (2) on the reference model, intervening with a prime based direction controls emotion about three times as strongly, and twice as selectively, as the best appraisal based direction; and (3) the model treats a prime based explication as interchangeable with the corresponding emotion. These evidences suggest that NSM primes seem to be better explanans for emotion in LLMs than many alternative options according to scientific explanations criteria.
Original Article
View Cached Full Text

Cached at: 07/22/26, 08:22 AM

# Semantic Primes as Explanans for Emotion in Large Language Models
Source: [https://arxiv.org/html/2607.18691](https://arxiv.org/html/2607.18691)
###### Abstract

Progresses have been made on understanding emotion mechanisms of large language models \(LLMs\)\. However, how to explain emotion in LLMs, or even what constitutes good explanations, are less clear\. Emotion representations, components, circuits are widely recoverable, but as explanations of a model’s own computation they are circular; the emotion space dimensions tend to be arbitrary and non\-terminating\. A pressing question to ask is whether a more primitive set of internal variables does the work: the semantic primes of the Natural Semantic Metalanguage \(NSM\)\. Across four instruction\-tuned LLMs \(Llama\-1B, Gemma\-2B, Gemma\-9B, OLMo\-7B\), experiments show that the NSM primes are \(1\) recoverable internal elements; and \(2\) on the reference model, intervening with a prime based direction controls emotion about three times as strongly, and twice as selectively, as the best appraisal based direction; and \(3\) the model treats a prime based explication as interchangeable with the corresponding emotion\. These evidences suggest that NSM primes seem to be better explanans for emotion in LLMs than many alternative options according to scientific explanations criteria\.

Code & Data—http://github\.com/fxing79/pcs

![Refer to caption](https://arxiv.org/html/2607.18691v1/x1.png)Figure 1:What makes a good causal explanation of the emotion an LLM computes and outoputs? Three tests are drawn from the scientific explanation literature: \(1\) real: it exists internally, \(2\) causal: intervening on it moves the emotion, \(3\) faithful: the model behaves as if it is the emotion\. All three candidates pass them, but to differing degrees\. What separates them most is the fourth vertical*basicness*axis: an explanans must be more basic than the explanandum\. Emotion labels are circular, and appraisals partial reduction; only NSM primes bottom out at a definitional floor\.## 1Introduction

Unlike human emotion, which is phenomenal and embodied, emotion in LLMs is functional and computed\. The difference calls for a whole set of theory on how emotion emerges and works in LLMs as they are increasingly used and integrated in our society\. Progresses have been made on understanding that emotion space\(Wuet al\.[2026](https://arxiv.org/html/2607.18691#bib.bib19)\)and emotion representations\(Sofroniewet al\.[2026](https://arxiv.org/html/2607.18691#bib.bib17)\)are recoverable in LLMs, and that LLM emotion behavior can be controlled and manipulated via layer injection\(Taket al\.[2025](https://arxiv.org/html/2607.18691#bib.bib1)\)or circuits\(Wanget al\.[2025](https://arxiv.org/html/2607.18691#bib.bib16)\)\. But there is a significant gap between understanding and \(scientific\) explanations\. For example, I understand “a magician pulls a rabbit out of a hat” is a stage effect and is not real, but I cannot explain how he did it\. People also have no difficulty understanding an apple falls after ripening without explaining it with plant hormones or gravity\. In this sense, there has been little purposeful investigation on explanations of emotion in LLMs from the mechanistic interpretability literature\(Bereska and Gavves[2024](https://arxiv.org/html/2607.18691#bib.bib29); Sharkeyet al\.[2025](https://arxiv.org/html/2607.18691#bib.bib18)\)\. Since using the emotion artifacts \(representations, sparse autoencoder extractions, circuits, etc\.\) to explain themselves is non\-semantic and circular, other internal variables must be engaged\. Two concept families are often imported from human emotion psychology\(Taket al\.[2025](https://arxiv.org/html/2607.18691#bib.bib1)\)\. The first family of variables are the “more basic” discrete emotion states or categories decoded from residual\-stream activations as linear directions\(Tiggeset al\.[2024](https://arxiv.org/html/2607.18691#bib.bib3)\), e\.g\., Ekman’s 6 basics or Russell’s Circumplex Model; the second family are continuous ratings, e\.g\., valence\-arousal dimensions\(Wuet al\.[2026](https://arxiv.org/html/2607.18691#bib.bib19)\)or Scherer’s 21 appraisal dimensions thatTaket al\.\([2025](https://arxiv.org/html/2607.18691#bib.bib1)\)use to probe and steer LLM emotional behavior\.

Two problems follow\. First, judging the completeness and evaluating competing explanations are intricate\. One set of variables may be more linearly recoverable via probing, but that fact alone does not establish better explanations: probe accuracy is correlational and need not reflect causal use\(Belinkov[2022](https://arxiv.org/html/2607.18691#bib.bib32)\)\. The mechanistic interpretability field has answered this by adding interventions\(Viget al\.[2020](https://arxiv.org/html/2607.18691#bib.bib34); Elazaret al\.[2021](https://arxiv.org/html/2607.18691#bib.bib33)\), and on the narrow technical reading of the word, an explanation is*mechanistic*only when it makes a causal claim\(Saphra and Wiegreffe[2024](https://arxiv.org/html/2607.18691#bib.bib31)\)\. But when another set of variables steers LLM emotion behavior better, weighing between these two sets becomes complicated\. Second, neither family of variable sets bottom out as explanations\. To classify an activation as*guilt*re\-applies the annotator’s label to the very activation the label should explain: the explanans is the explanandum\. The appraisal space dimensions describe guilt as a profile over a composition of*self\_responsibility*\+*unpleasantness*, but those dimensions are not more primitive than the emotions they describe;*unpleasantness*is an affective concept of the same order as*guilt*\. The reduction floats and never terminates\.

In response to the two problems, I attempt a set of tests to filter out pseudo\-explanation and evaluate candidate explanans, and draw a third family of variables from a 60\-year\-old linguistic program, i\.e\., the Natural Semantic Metalanguage \(NSM\) of Wierzbicka and Goddard\(Wierzbicka[1972](https://arxiv.org/html/2607.18691#bib.bib10),[1996](https://arxiv.org/html/2607.18691#bib.bib7); Goddard and Wierzbicka[2002](https://arxiv.org/html/2607.18691#bib.bib11),[2014](https://arxiv.org/html/2607.18691#bib.bib9)\)\. NSM identifies∼\\sim65*semantic primes*\(want,feel,think,know,do,good,bad,not,because,i,someone,…\\ldots\), held to be mutually indefinable and shared with the rest of language, so that every other concept, emotions included, can be*explicated*as a paraphrase built only from primes\. For instance, the gold explication of*guilt*is roughly*“I did something bad; I feel bad because of this; I do not want this to have happened,”*which bottoms out at primes \(i,do,bad,feel,want,not\) that are not themselves emotions\. This explication of*guilt*is carried as a running example\. It is noteworthy that confirming the correctness of NSM in either human language or the LLM case is not this paper’s scope\. Instead, my claim is an engineering one that NSM primes seem to be better explanans considering the criteria set out here\.

The three families of variables are judged by the standard of causal explanation, on the interventionist account\(Woodward[2003](https://arxiv.org/html/2607.18691#bib.bib30)\): a feature is explanatorily relevant to an outcome when intervening on it changes the outcome\. Three demands elaborate \(Figure[1](https://arxiv.org/html/2607.18691#S0.F1)\), and they come from this account of explanation, not from any tenet of NSM\.

- •Existence\.The feature is a real internal element, a genuine linear representation rather than a probe artifact\. This is rigorously tested by decoding above length and control\-task baselines\.
- •Intervention\.Intervening on the feature changes the emotion the model computes, and more so the feature serves as an explanan better\. This is tested by steering\.
- •Behavioral equivalence\.At the input–output level the model treats the feature as the concept: in the NSM case, a prime explication is interchangeable with the emotion word and is a sufficient cue\. This is tested with the model’s own inferences and generations\.

The three demands are necessary: a feature that is absent can be part of the mechanism but is beyond explanations, one that is causally idle can be a confounder but does not explain, and one that is causally active but means something else is not the concept\(Sharkeyet al\.[2025](https://arxiv.org/html/2607.18691#bib.bib18)\)\. They are not jointly a proof though: a prime’s lexical correlate could in principle pass all three, so construct validity remains an open problem \(see Section[10](https://arxiv.org/html/2607.18691#S10)\), and the internal mechanism exhibited is a sketch focusing on the raw ingredients rather than a gap\-free account\.

#### Contributions\.

Primes inside four instruction\-tuned models spanning two orders of magnitude and three model families \(Llama\-3\.2\-1B, Gemma\-2\-2B, Gemma\-2\-9B, OLMo\-2\-7B\) are studied to ensure result generalizability\. Experiments are built onTaket al\.\([2025](https://arxiv.org/html/2607.18691#bib.bib1)\)’s public harness data so the emotion comparison is head\-to\-head\.

The result is a single claim that NSM primes are better explanans for emotion in LLMs when compared to other explanans used in literature\. It is supported by three findings: \(1\) primes are real internal elements as 30 of 32 emotion\-describing primes are linearly encoded above length and control\-task baselines in all four models; \(2\) on the reference model, a prime direction controls emotion about three times as far, at nearly twice the selectivity, as the best appraisal direction, beating a dimensionality\-matched composite \(p<10−3p<10^\{\-3\}\) and a random control; \(3\) the model treats a prime explication as interchangeable with the emotion and as a sufficient cue, in all four models and more so than a matched appraisal description\. For reproducibility, a contrastive suite of 11,902 minimal pairs for 32 primes with controls and gold explications is released\. Analysis runs on a single CPU and extraction in minutes on a single GPU, so the study is cheap to replicate\.

## 2Background

### 2\.1From explanans to their emotion readout

Mechanistic interpretability seeks a causal, not merely predictive, account of a computation: to explain “the model reports emotionee” is to exhibit the internal factor that*produces*it\. The two requirements discussed in Introduction sort the candidate explanations of emotion by how well each passes the causal tests and how far each reduces\. In the emotion label case, a residual activationh∈ℝdh\\in\\mathbb\{R\}^\{d\}at a consolidation layer admits a linear classifier intoKKpossible emotion labels; each emotion is a readout directionwew\_\{e\}and the emotion space is flat, so*guilt*is an atom and the account does not reduce: the explanans is the explanandum\. In the appraisal case, the samehhadmits 21 linear regressions onto appraisal scales;*guilt*becomes a profile \(2 dimensions activated\)\. NSM primes are∼\\sim65 directionsvpv\_\{p\}, and an emotion becomes an*explication*, a short structured paraphrase of prime predicates over a floor shared with the rest of language:*guilt*≈\\approx*“I did something bad; I feel bad because of this\.”*The compositional prediction is defined as:

we≈f​\(\{vp:p∈P​\(e\)\}\),w\_\{e\}\\;\\approx\\;f\\\!\\Big\(\\\{v\_\{p\}:p\\in P\(e\)\\\}\\Big\),\(1\)withP​\(e\)P\(e\)the prime support ofee’s explication andffthe readout map\. The strongest reading takesfflinear, which empirical results reject; the behavioral reading takesffto be whatever the model computes and that is conceptually the NSM’s grammar, which empirical results confirm\. Obviously, an explication is structured, not a bag of primes, so a signed sum of prime directions is lossy\. Dimensional appraisal is by contrast additive, an emotion a profile over independent dimensions\(Smith and Ellsworth[1985](https://arxiv.org/html/2607.18691#bib.bib35)\); the non\-linearity of Scherer’s process model\(Scherer[2009](https://arxiv.org/html/2607.18691#bib.bib36)\)is temporal, not in this static mapping\. So the accounts differ in reduction depth and in the form offf: on this reading prime composition would be grammatical and appraisal composition dimensional, a contrast the steering experiment puts to the test\.

### 2\.2Primes and emotion in LLMs

That LLMs present and process semantic primes seems self\-evident, though the proof is absent from literature, especially whether those are used when processing emotion\. A recent study confirms prime\-related schema slots from LLM outputs\(Xing and Cambria[2026](https://arxiv.org/html/2607.18691#bib.bib15)\)\. At mid\-layers, emotion is read from residual streams as linear structure\(Taket al\.[2025](https://arxiv.org/html/2607.18691#bib.bib1)\)or extracted using a sparse autoencoder\(Wuet al\.[2026](https://arxiv.org/html/2607.18691#bib.bib19)\)\. The work closest to this one isTaket al\.\([2025](https://arxiv.org/html/2607.18691#bib.bib1)\), who probe 13 emotion categories and 21 appraisal dimensions across 10 open LLMs, localize mid\-layer emotion heads by activation patching and knockout, and steer behavior with orthogonalized appraisal injection; this pipeline is reproduced as the comparability anchor\. Since primes are fundamental concepts, existence test adds a Hewitt\-Liang control task\(Hewitt and Liang[2019](https://arxiv.org/html/2607.18691#bib.bib23)\)to separate a represented feature from probe capacity, within the linear\-representation framing ofParket al\.\([2024](https://arxiv.org/html/2607.18691#bib.bib20)\); sparse autoencoders\(Lieberumet al\.[2024](https://arxiv.org/html/2607.18691#bib.bib24); Templetonet al\.[2026](https://arxiv.org/html/2607.18691#bib.bib25)\)are an alternative route used only in pre\-trained form; and an LLM pipeline that generates and verifies NSM explications\(Baartmanset al\.[2025](https://arxiv.org/html/2607.18691#bib.bib14)\)supplies those gold recipes not published in literature\.

## 3A Contrastive Suite for Semantic Primes

Probing primes needs stimuli that assert a single prime while holding lexical context fixed\. For this reason, a contrastive minimal\-pair suite for 32 primes is built and released with gold explications, code, and a machine\-translated Chinese version for cross\-lingual studies\. For each prime templated pairs were generated: a positive member asserting the prime and a negative member with the same content but the prime absent\. For example forwant,

“*\(\+\)The nurse wants to play the piano*”

“*\(\-\)The nurse played the piano*”;

and fornot,

“*\(\+\)The driver did not open the gate*”

“*\(\-\)The driver opened the gate*\.”

Templates are paraphrased under rejection rules that forbid the prime’s exponent word in the negative member, match target word\-length within±1\\pm 1where feasible, and track a syntactic\-template identifier for a held\-out split\. The 32 primes span the symbol classes that admit a single\-prime contrast: mental predicates \(want, feel, think, know, say, do\), evaluators \(good, bad\), modals \(can, maybe, not, true\), person primes \(i, someone, other, people\), temporal, descriptor, bodily, and reason primes\.

The selection of primes is principled, not a convenience sample\. A contrastive probe needs a prime a sentence can assert or withhold while holding context fixed, which is natural for predicates and operators but ill\-posed for referential substantives and determiners \(this,something,kind\) that appear in nearly every sentence, and for quantifiers \(one,some,all\) whose contrast is graded and noun\-phrase\-bound\. Crucially, of the 22 distinct primes that appear in the gold explications of the 13 target emotions, the suite covers 21; the lone exception isthis\. The untested primes are dominated by the quantifier and spatial classes, which appear in no emotion explication, so the 32 primes already cover every emotion experiments required in this paper\. The full prime inventory is in Appendix[A](https://arxiv.org/html/2607.18691#A1)\.

The suite contains 11,902 minimal pairs / 23,804 sentences,≥\\geq202 pairs per prime, and≥\\geq16 held\-out template pairs per prime\. No negative leaks the prime’s exponent \(a hard generation constraint\); the length confound is quantified per prime as Cohen’sddon word count, and the probing harness always reports a length\-only baseline, so a prime counts as encoded only when it clears the length shortcut\. The pipeline is validated on synthetic data, successfully recovering a planted layer for six synthetic primes with off\-layer accuracy at chance \(Appendix[A](https://arxiv.org/html/2607.18691#A1)\)\.

## 4Experimental Setup

#### Models and data\.

Four LLMs are studied: Llama\-3\.2\-1B\-Instruct \(16 blocks,d=2048d\{=\}2048\), Gemma\-2\-2B\-it \(26 blocks,d=2304d\{=\}2304\), Gemma\-2\-9B\-it \(42 blocks,d=3584d\{=\}3584\), and OLMo\-2\-7B\-Instruct \(32 blocks,d=4096d\{=\}4096\), spanning two orders of magnitude and three families, including the fully open OLMo whose independent pre\-training makes a cross\-family claim more than a within\-lineage one\. Llama\-3\.2\-1B is the reference model for the intervention experiment because of its quasi\-linear emotion readout; the other three test prime existence and explication behavioral equivalence across family and scale\. Emotion and appraisal labels come from crowd\-enVent\(Troianoet al\.[2023](https://arxiv.org/html/2607.18691#bib.bib27)\)\(6,800 event descriptions annotated for 13 emotions and 21 appraisal dimensions\); I reproduceTaket al\.\([2025](https://arxiv.org/html/2607.18691#bib.bib1)\)’s split of 2,740 train and 1,370 test sentences, with ISEAR\(Scherer and Wallbott[1994](https://arxiv.org/html/2607.18691#bib.bib28)\)as an additional robustness check\.

#### Probes and steering\.

Linear probes are L2\-regularized logistic regression, appraisal probes L2 ridge; hyperparameters are not retuned per layer\. Steering adds a unit\-normalized direction, scaled by the layer’s mean residual norm, to the block residual at layer 11, the mid\-network consolidation site fixed by the patching peak \(Section[7](https://arxiv.org/html/2607.18691#S7)\), at all token positions\. The behavioral tests use only the model’s generations and answer\-token logits, with no probe, and run on all four models\. For existence, this study reports per\-prime peak accuracy against the length baseline and Hewitt\-Liang selectivity; for intervention, target\-logit shift, off\-target mass, selectivity, and target\-argmax rate; for behavioral equivalence, explication\-to\-word logit agreement, explication\-only classification with prime ablation, and the linear recipe match\. Cached activations are reused and never recomputed\.

## 5Primes Are Real Internal Elements

This section presents existence test results: a prime must be a genuine internal element, not a probe artifact\. Decoding establishes this only with controls\. Probe accuracy alone only shows a label is recoverable, while clearing a Hewitt\-Liang control task shows the model represents the prime\(Hewitt and Liang[2019](https://arxiv.org/html/2607.18691#bib.bib23)\), since the control measures what a probe of the same capacity can fit on structureless pseudo\-labels\.

#### Protocol\.

For each of the 32 primes, residual activations are extracted on its contrastive pairs \(balanced, last token, every layer\) and fit one logistic probe per layer, recording a length\-only baseline, a Hewitt\-Liang control\-task probe\(Hewitt and Liang[2019](https://arxiv.org/html/2607.18691#bib.bib23)\), and a held\-out template split\. A prime counts as*linearly encoded*when its peak accuracy clears the length baseline significantly and shows selectivity over other primes\.

#### Existence and replication\.

Thirty of the 32 primes are linearly encoded in all four models: 31/32 on Llama\-3\.2\-1B \(bootstrap 95% CI\[29,32\]\[29,32\]\), 30/32 on Gemma\-2\-2B, and 32/32 on Gemma\-2\-9B and OLMo\-2\-7B, across three families and the 1B\-to\-9B scale range\. Averaged over the 32 primes on Llama\-3\.2\-1B, held\-out template decoding reaches0\.910\.91, far above the Hewitt\-Liang control task \(0\.500\.50, at chance\) and the length\-only baseline \(0\.620\.62\); the gap between probe and control is the selectivity that makes this an existence claim and not a probe artifact \(Figure[2](https://arxiv.org/html/2607.18691#S5.F2)\)\. The two exceptions across models,doandsay, carry the worst length confound \(d=1\.26d\{=\}1\.26and2\.032\.03\); other confounded primes \(know,can,think, alld\>1\.2d\>1\.2\) still clear the bar, so the control separates genuine encoding from a length shortcut\. The atoms*guilt*’s explication is built from,i,do,bad, andfeel, are all among the encoded primes, so the ingredients its recipe needs are all present in LLMs\. Decoding with controls, replicated across four models, establishes existence but not use: a probe can read a direction the model never uses, which the next section tests by intervention\.

![Refer to caption](https://arxiv.org/html/2607.18691v1/x2.png)Figure 2:Existence on Llama\-3\.2\-1B: per\-layer prime decoding, averaged over the 32 primes\. Held\-out template accuracy \(red\) rises to0\.910\.91and stays far above the Hewitt\-Liang control\-task accuracy \(grey, near the chance line0\.50\.5\) and the length\-only baseline \(dashed\); the shaded gap is the selectivity that separates a represented prime from a probe artifact\.

## 6Primes Control Emotion Better

This section presents intervention test results\. In short, a direction assembled from prime atoms and injected into the residual stream is found to control emotion more strongly and more selectively than the best appraisal based direction\. The shift is about three times as large at nearly twice the selectivity, the gap widest been on the agency axis \(i\.e\.,iversussomeone\), suggesting a less effective modeling construct by appraisals\.

#### One mechanism, six directions\.

A forward hook adds a vector to the block residual at layer 11, at all token positions\. The six arms differ only in the injected direction: the*emotion*probe directionwew\_\{e\}\(a readout ceiling\); the*single appraisal*most associated with the target \(e\.g\., guilt to*self\_responsblt*, anger to*other\_responsblt*, joy to*pleasantness*, sadness to*unpleasantness*\); a*multi\-appraisal composite*, the top\-kkappraisals withkkmatched to the recipe’s component count; the*single best prime*; the*prime recipe*, a centrality\-weighted signed sum of fitted prime directions \(guilt asi\+do\+bad\\textsc\{i\}\+\\textsc\{do\}\+\\textsc\{bad\}, so is still lossy\); and a*random\-concept*control\. Every direction is unit\-normalized and scaled by the layer’s mean residual norm, so the dose is a common fraction of residual magnitude; the composite and random arms control for direction count and for injecting any vector at all\. The 13 emotion\-token logits are read at the answer position over four targets \(guilt, anger, joy, sadness, see Table[1](https://arxiv.org/html/2607.18691#S6.T1)\)\.

Table 1:Per\-target steering atβ=2\\beta=2\(Llama\-3\.2\-1B\)\. The emotion arm is the readout ceiling; the prime arm beats the appraisal arm on every target emotion\.
#### Per\-target steering detail\.

Averaged over the four targets and the positive dose grid \(β∈\{0\.5,1,2\}\\beta\\in\\\{0\.5,1,2\\\}\), the prime recipe shifts the target emotion by3\.733\.73logits at selectivity0\.580\.58, while the single best appraisal shifts it by1\.291\.29at0\.310\.31\(Table[2](https://arxiv.org/html/2607.18691#S6.T2), Figure[3](https://arxiv.org/html/2607.18691#S6.F3)\)\. Against the component\-matched composite the prime recipe still shifts emotion further \(3\.733\.73versus2\.112\.11grid\-averaged; strongest dose4\.914\.91, CI\[4\.76,5\.05\]\[4\.76,5\.05\], versus2\.842\.84,\[2\.68,2\.99\]\[2\.68,2\.99\]\) and more selectively \(0\.600\.60versus0\.540\.54, non\-overlapping\); a permutation test gives a\+2\.07\+2\.07\-logit gap,p<10−3p<10^\{\-3\}\. The random\-concept control barely moves the target \(0\.020\.02, selectivity−0\.26\-0\.26\), a near\-zero causal floor, and a single best prime already about matches the full recipe \(3\.623\.62at0\.610\.61\), so the handle is carried by the prime representation at matched dimensionality, not by component count\. The advantage is sharpest on the agency axis \(selectivity0\.570\.57for primes versus0\.270\.27for the single appraisal\); for anger the prime arm drives anger to the top prediction in77%77\\%of prompts\. Injecting the emotion direction itself shifts the target most \(9\.809\.80, selectivity0\.780\.78\), as expected for the readout direction; the prime arm, built from general\-purpose atoms, recovers a large and selective fraction of that ceiling\.

Table 2:Steering, two models\. Llama: the prime recipe beats both appraisal arms on shift and selectivity \(gap versus composite\+2\.07\+2\.07,p<10−3p<10^\{\-3\}\); random control near zero; emotion ceiling9\.809\.80\. Gemma\-9B: the model’s own explication encoding steers primes \(mostly via content\) but not appraisals; ceiling1\.821\.82, floor0\.100\.10\.![Refer to caption](https://arxiv.org/html/2607.18691v1/x3.png)Figure 3:Steering head to head \(Llama\-3\.2\-1B\): prime \(red\), appraisal \(blue\), emotion \(yellow\)\. \(a\) Mean target\-logit shift versus dose\. \(b\) Selectivity versus shift for all positive doses; upper right is better\. The prime arm dominates the appraisal arm on both axes; the emotion arm is the readout ceiling – best control but weakest explanatory power\.
#### Robustness check on injection layer\.

The prime advantage is not an artifact of a particular direction fit or injection site\. It holds in all six bootstrap resamples of the probe\-fitting data, where the appraisal arm varies but the prime arm stays ahead, and at every injection layer from 7 to 14, where all directions including the prime recipe are refit, the prime arm leads\. The steered direction is also nearly orthogonal to the emotion readout: the cosine between the prime recipe and the emotion probe directionwew\_\{e\}averages0\.040\.04across targets\. A direction geometrically unrelated to the readout still drives the emotion logit, which is the signature of a*compositional*causal effect: the primes are injected upstream and processed by the network into the emotion, not read off the recipe directly\. This also reconciles with the geometric non\-reduction of Section[7](https://arxiv.org/html/2607.18691#S7), where the emotion readout is not a linear function of prime directions\. The composition is robust to the ingredient: assembled from the model’s own single\-prime encodings, a contrastive difference of means, rather than from probe directions, the recipe steers as well \(4\.54\.5versus3\.93\.9at matched selectivity0\.590\.59\), and both far exceed injecting the whole explication’s encoding \(1\.11\.1, selectivity−0\.25\-0\.25\)\. Independent prime vectors, composed, out\-steer the holistic sentence encoding, so the effect is neither a probe\-geometry artifact nor a global\-content injection\.

#### On larger models the linear recipe does not transfer\.

The linear prime recipe is model\-specific\. With per\-model dose and layer calibration it is inert on Gemma\-2\-9B \(grid\-averaged shift0\.020\.02, at the random floor\), and assembling it from the model’s own single\-prime encodings rather than probe directions does not rescue it \(0\.140\.14\): the primes do not linearly compose there by either ingredient\. The emotion is reachable only from the whole explication’s encoding \(0\.970\.97at selectivity0\.660\.66, against an emotion\-readout ceiling of1\.821\.82\), the trivial content direction any model carries for guilt\-describing text\. Phrased in primes that content still out\-steers a matched appraisal description \(1\.581\.58versus0\.740\.74\), and the same treatment leaves appraisal at its additive profile \(0\.740\.74versus0\.410\.41; Table[2](https://arxiv.org/html/2607.18691#S6.T2)\)\. For this reason, the main analysis uses Llama\-3\.2\-1B, where the compositional recipe itself steers; on Gemma\-2\-9B the prime explication out\-steers the appraisal description as content, but the linear prime composition does not transfer\. On OLMo\-2\-7B the readout ceiling is ill\-set \(a negative emotion\-arm shift\), so numbers are uninformative and not reported\.

## 7Emotions Reduce to Primes Behaviorally

This section presents behavioral equivalence test results\. In all four models the model treats a prime explication as the emotion\. The reduction is behavioral, not geometric: the late readout is not a linear sum of prime directions, and the localization of emotion within the network explains why the two levels diverge\.

#### Behavioral equivalence and sufficiency\.

With no probing and on all four models, the 13 emotion\-token logits are read for: \(1\) the emotion word, \(2\) its gold explication, \(3\) a same\-affect decoy explication, e\.g\., shame for guilt, differing by onepeople\-knowcomponent, and \(4\) a prime\-scrambled control\. Given the explication alone, with no emotion word present, the model classifies it as the target emotion well above the 13\-way chance rate of0\.080\.08, at0\.540\.54,0\.850\.85,0\.770\.77, and0\.540\.54respectively for Llama\-3\.2\-1B, Gemma\-2\-2B, Gemma\-2\-9B, and OLMo\-2\-7B, and the target explication beats its same\-affect decoy by0\.300\.30to0\.630\.63probability mass, so it is not riding on generic affect words or on the presence of an emotion label, which the explication never contains\. Reduction is also sufficient: ablating the explication one prime at a time, removing a*central*prime degrades the target logit more than removing a*peripheral*one in every model \(0\.590\.59versus0\.120\.12on Llama\-3\.2\-1B, and the same ordering elsewhere; see Table[3](https://arxiv.org/html/2607.18691#S7.T3)\)\.

Table 3:Behavioral test batteries, all four models\. 13\-way chance for classification is0\.080\.08\. The explications are read as the emotion \(E\) and are compositionally sufficient \(S, central\>\>peripheral in every model\) more strongly than a matched appraisal description, whose decoy margin is near zero and whose sufficiency gradient inverts on the two larger models\. Primes resist simplification more than non\-primes in every model \(floor\)\.For*guilt*, the evaluativebadandfeelare central and thewant/notclause peripheral: the central primes carry the emotion\. The same battery on a matched appraisal description, a Component\-Process\-style clause profile built from the crowd\-enVent appraisal ratings, is markedly weaker \(Table[3](https://arxiv.org/html/2607.18691#S7.T3)\): the description is read as the target emotion below the prime explication on every model \(0\.250\.25to0\.420\.42versus0\.540\.54to0\.850\.85\), beats its same\-affect decoy by only0\.030\.03to0\.110\.11against the prime’s0\.300\.30to0\.630\.63, and its central\-clause sufficiency inverts on the two larger models\. The model treats a prime explication, not a matched appraisal profile, as interchangeable with the emotion, on our best operationalization of each\. A fluency confound does not explain this gap\. On the reference model the prime explication is more fluent than the matched appraisal description \(mean per\-token log\-probability higher by0\.810\.81nats,p<10−3p<10^\{\-3\}\), yet under fluency\-invariant scoring, domain\-conditional PMI\(Holtzmanet al\.[2021](https://arxiv.org/html/2607.18691#bib.bib21)\)and contextual calibration\(Zhaoet al\.[2021](https://arxiv.org/html/2607.18691#bib.bib22)\), the prime advantage persists and widens, because calibration removes a label\-prior bias that suppressed the explication\.

#### The readout is not a linear sum\.

The prime directions are fit from the contrastive suite in the same residual space as the emotion vectors and build a centrality\-weighted recipe direction per emotion from gold explications\(Wierzbicka[1999](https://arxiv.org/html/2607.18691#bib.bib8)\); the match test asks whetherwew\_\{e\}is closer to its own recipe than to the other twelve emotions\. Both direct sources land near chance \(Table[4](https://arxiv.org/html/2607.18691#A3.T4)\): out\-of\-context rank\-30\.390\.39against chance0\.230\.23, in\-context exactly at chance\. Emotion is linearly recoverable from the prime coordinates, but not because of them: projecting the consolidation\-layer residual onto the 32 prime directions and training a classifier reaches0\.860\.86, yet a*random*32\-dimensional projection reaches0\.820\.82, so the prime subspace adds only0\.050\.05, and per\-prime ablation moves accuracy by at most0\.020\.02\. The prime subspace is not a privileged basis\.

#### A mechanism sketch\.

The two findings reconcile through localization\. The emotion computation peaks at layer 10, and by the consolidation layer the prime content has been*consumed*into the emotion representation: of 18 primes tested at the answer position, onlyfeel\(lift\+0\.23\+0\.23\) andsomeone\(\+0\.11\+0\.11\) remain decodable\. The late readout cannot be a linear sum of prime directions because there are almost no prime directions left to sum, and the survivors are in a sense, the emotion\-handles: the feeling predicate and the other\-person agency role\. Emotions reduce to primes as a computation the model carries out in the middle layers and spends, not as a static linear geometry of the final state\. I present this as a mechanism sketch: it localizes where the reduction happens but does not trace the layer\-10\-to\-readout step arrow by arrow\.

## 8Do Primes Bottom Out in LLMs?

The explication reduces an emotion to primes; a probe\-free signature shows the primes themselves do not reduce further\. Prompted to restate a word “using only simpler words,” the model simplifies a non\-prime far more readily than a prime: the fraction of output words more frequent than the target is0\.73/0\.50/0\.72/0\.720\.73/0\.50/0\.72/0\.72for non\-primes against0\.47/0\.21/0\.34/0\.400\.47/0\.21/0\.34/0\.40for primes \(see Table[3](https://arxiv.org/html/2607.18691#S7.T3)and Figure[4](https://arxiv.org/html/2607.18691#S8.F4)\)\. A prime resists reduction in a sense that more than half of the explaining words cannot be simpler, so the explication regress terminates\. This is the bottom\-out property an explanation demands:*guilt*reduces toi,do,bad, andfeel, which are not to be explained further in the semantic space, where a label or an appraisal dimension would still float\.

![Refer to caption](https://arxiv.org/html/2607.18691v1/x4.png)Figure 4:The definitional floor, behaviorally\. Asked to restate a word in simpler terms, the model produces a smaller fraction of simpler \(shorter, higher\-frequency\) words for a prime than for a length\- and class\-matched non\-prime, in all four models: a prime resists reduction, a non\-prime does not\. Error bars are the standard error of the mean over the primes and the matched non\-primes\.The localization of primes also seems earlier than emotions\. I reproduce and concur the localization of emotions on Llama\-3\.2\-1B byTaket al\.\([2025](https://arxiv.org/html/2607.18691#bib.bib1)\), which fixes the consolidation band and the injection layer\. Emotion accuracy saturates late \(linear probe peak0\.9430\.943at layer 13\); activation patching is sharper, transplanting the event\-position residual flips the target prediction at a rate that climbs through the mid\-network and peaks at0\.8050\.805at layer 10 before falling, recovering a mid\-layer attention\-mediated consolidation\. It is this peak that fixes layer 11 as the injection site\. Full curves and the emotion geometry, where valence dominates and the agency contrast is secondary, are in Appendix[B](https://arxiv.org/html/2607.18691#A2)\.

## 9Discussion

This paper shows the potential of semantic primes as a handle that explains in LLMs, not predicts over complicated theories\. The two standard accounts of emotion in LLMs do not bottom out: labels relabel the activation, appraisals describe it in terms no more primitive than emotion, and both are validated mostly by probes that show recoverability, not use\. Judged by what it takes to explain the emotion a model computes, primes pass where these fall short: they are real internal elements, the model treats their explications as the emotion, and intervening on them controls emotion more than appraisals do\. The same primes \(want,not,can,feel\) appear in explications of decisions, permissions, and requests, so the prime\-interface is more universal than emotion\-specific explanans\.

It worths mentioning, however, that geometrically the emotion direction is not a linear sum of prime directions; though behaviorally the model treats an emotion word and its explication as interchangeable and the explication as sufficient\. A causal handle need not be a geometric basis\. The reconciliation is mechanistic: primes are computed in the mid layers \(e\.g\., 7/16\) and consumed into the emotion representation \(11/16\), so the final state retains only the feeling and agency atoms\. A prime\-indexed account of emotion is therefore viable, but it must read emotion off the primes nonlinearly\.

The clean compositional result is the Llama one, and it is strong there: independent prime vectors, whether probe directions or the model’s own single\-prime encodings, compose into a direction that out\-steers the whole explication’s encoding, so the effect is composition from parts, not a global\-content injection\. On Gemma\-2\-9B the parts do not compose by either ingredient; only the whole\-content direction steers, which is the trivial encoding any model has for emotion\-describing text\. The prime explication still out\-steers a matched appraisal description there, and the appraisal arm gains nothing from the same treatment, so the content asymmetry is real cross\-model even though the compositional claim is Llama\-only\. Tracing whether a larger model composes primes through a different circuit is left open \(Section[10](https://arxiv.org/html/2607.18691#S10)\)\.

One structure recurs as a causal signal even though valence dominates the static geometry\. Agency is where prime steering most beats appraisal \(selectivity0\.570\.57versus0\.270\.27\), the one decomposition that survives across direction sources \(anger at rank\-1\), andsomeoneis one of only two primes still decodable at the readout\. In NSM it is a single substitution,i↔someone\\textsc\{i\}\\leftrightarrow\\textsc\{someone\}, inside a recipe of the form*X did something bad; I feel bad because of this*, the minimal edit that turns guilt into anger\.

## 10Limitations and Conclusion

The main limitation is on construct validity\. Because the three tests are necessary for a set of variables to count as an explanation, but not complete or jointly a proof, it is possible that a decoded direction measures a lexical correlate, not the NSM prime per se\. More external validation can be supplemented by blind human paraphrases written by annotators unaware of the target prime\.

A second limitation is on the full understanding from prime ingredients to emotion representations: the mechanism is only a sketch\. The intervention is clean on Llama\-3\.2\-1B and transfers to Gemma\-2\-9B under non\-linear composition \(Section[6](https://arxiv.org/html/2607.18691#S6)\); but OLMo\-2\-7B is not calibrated, and the non\-linear appraisal arm uses a constructed appraisal description that a stronger operationalization could potentially improve, so the cross\-model causal claim rests on two models\. The mechanism that reconciles behavioral and geometric reduction localizes a mid\-network consolidation band but does not trace the layer\-10\-to\-readout step\. Concretely, the attention heads or MLP sub\-modules that compose the primes, or the NSM grammar component, are unidentified\. The gold recipes are taken from the published NSM canon\(Wierzbicka[1999](https://arxiv.org/html/2607.18691#bib.bib8); Goddard and Wierzbicka[2014](https://arxiv.org/html/2607.18691#bib.bib9)\), whose explications carry the validity of a long peer\-reviewed program; what is specific to us is the reduction to a 32\-prime working subset and the centrality weighting used for ablation\. Behavioral robustness to that operationalization can be further checked by recipe\-perturbation or re\-deriving the central\-versus\-peripheral ablation under alternative but NSM\-valid explications\.

Another limitation is that cross\-lingual universality of primes is not tested in the LLM case, though it is one of the NSM proposition for human language\. Transfer to other languages would be confounded by a multilingual model routing other languages through an English\-like latent space\(Wendleret al\.[2024](https://arxiv.org/html/2607.18691#bib.bib12); Schutet al\.[2025](https://arxiv.org/html/2607.18691#bib.bib13)\), so separating universality from English\-pivot routing requires native\-authored stimuli in a typologically distant language and ideally an individually trained non\-English LLM, which is left to future work\.

In summary, this research confirms semantic primes to be good explanans of emotion in LLMs and requires minimal theory\-building\. Although other models\(Taket al\.[2025](https://arxiv.org/html/2607.18691#bib.bib1)\), linguistic features\(Chang[2025](https://arxiv.org/html/2607.18691#bib.bib6)\)or narrative style\(Shenet al\.[2024](https://arxiv.org/html/2607.18691#bib.bib4)\)may as well explain emotion, they are likely correlational \(not causal\) and subject to further specification\. It also presents a framework for assessing the quality of scientific explanations, therefore makes methodological contribution to mechanistic interpretability\.

## References

- R\. Baartmans, M\. Raffel, R\. Vikram, A\. Deringer, and L\. Chen \(2025\)Towards universal semantics with large language models\.External Links:2505\.11764Cited by:[§2\.2](https://arxiv.org/html/2607.18691#S2.SS2.p1.1)\.
- Y\. Belinkov \(2022\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p2.1)\.
- L\. Bereska and E\. Gavves \(2024\)Mechanistic interpretability for AI safety \- A review\.Transactions on Machine Learning Research2024\.External Links:[Link](https://openreview.net/forum?id=ePUVetPKu6)Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p1.1)\.
- E\. Y\. Chang \(2025\)Modeling emotions in multimodal llms\.InMulti\-LLM Agent Collaborative Intelligence: The Path to Artificial General Intelligence,pp\. 1–10\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1145/3749421.3749433)Cited by:[§10](https://arxiv.org/html/2607.18691#S10.p4.1)\.
- Y\. Elazar, S\. Ravfogel, A\. Jacovi, and Y\. Goldberg \(2021\)Amnesic probing: behavioral explanation with amnesic counterfactuals\.Transactions of the Association for Computational Linguistics9,pp\. 160–175\.Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p2.1)\.
- C\. Goddard and A\. Wierzbicka \(2002\)Meaning and universal grammar: theory and empirical findings\.John Benjamins,Amsterdam\.Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p3.2)\.
- C\. Goddard and A\. Wierzbicka \(2014\)Words and meanings: lexical semantics across domains, languages, and cultures\.Oxford University Press\.Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p3.2),[§10](https://arxiv.org/html/2607.18691#S10.p2.1)\.
- J\. Hewitt and P\. Liang \(2019\)Designing and interpreting probes with control tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Cited by:[§2\.2](https://arxiv.org/html/2607.18691#S2.SS2.p1.1),[§5](https://arxiv.org/html/2607.18691#S5.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2607.18691#S5.p1.1)\.
- A\. Holtzman, P\. West, V\. Shwartz, Y\. Choi, and L\. Zettlemoyer \(2021\)Surface form competition: why the highest probability answer isn’t always right\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),External Links:2104\.08315Cited by:[§7](https://arxiv.org/html/2607.18691#S7.SS0.SSS0.Px1.p2.10)\.
- T\. Lieberum, S\. Rajamanoharan, A\. Conmy, L\. Smith, N\. Sonnerat, V\. Varma, J\. Kramár, A\. Dragan, R\. Shah, and N\. Nanda \(2024\)Gemma scope: open sparse autoencoders everywhere all at once on gemma 2\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 278–300\.External Links:2408\.05147Cited by:[§2\.2](https://arxiv.org/html/2607.18691#S2.SS2.p1.1)\.
- K\. Park, Y\. J\. Choe, and V\. Veitch \(2024\)The linear representation hypothesis and the geometry of large language models\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),Cited by:[§2\.2](https://arxiv.org/html/2607.18691#S2.SS2.p1.1)\.
- N\. Saphra and S\. Wiegreffe \(2024\)Mechanistic?\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p2.1)\.
- K\. R\. Scherer and H\. G\. Wallbott \(1994\)Evidence for universality and cultural variation of differential emotion response patterning\.Journal of Personality and Social Psychology66\(2\),pp\. 310–328\.Cited by:[§4](https://arxiv.org/html/2607.18691#S4.SS0.SSS0.Px1.p1.4)\.
- K\. R\. Scherer \(2009\)The dynamic architecture of emotion: evidence for the component process model\.Cognition and Emotion23\(7\),pp\. 1307–1351\.Cited by:[§2\.1](https://arxiv.org/html/2607.18691#S2.SS1.p1.14)\.
- L\. Schut, Y\. Gal, and S\. Farquhar \(2025\)Do multilingual llms think in english?\.InICLR 2025 Workshop on Building Trust in Language Models and Applications,pp\. 1–37\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2502.15603)Cited by:[§10](https://arxiv.org/html/2607.18691#S10.p3.1)\.
- L\. Sharkey, B\. Chughtai, J\. Batson, and et al\. \(2025\)Open problems in mechanistic interpretability\.Transactions on Machine Learning Research2025,pp\. 1–82\.Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p1.1),[§1](https://arxiv.org/html/2607.18691#S1.p4.2)\.
- J\. Shen, J\. Mire, H\. W\. Park, C\. Breazeal, and M\. Sap \(2024\)HEART\-felt narratives: tracing empathy and narrative style in personal stories with llms\.InENMLP,pp\. 1026–1046\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.59)Cited by:[§10](https://arxiv.org/html/2607.18691#S10.p4.1)\.
- C\. A\. Smith and P\. C\. Ellsworth \(1985\)Patterns of cognitive appraisal in emotion\.Journal of Personality and Social Psychology48\(4\),pp\. 813–838\.Cited by:[§2\.1](https://arxiv.org/html/2607.18691#S2.SS1.p1.14)\.
- N\. Sofroniew, I\. Kauvar, W\. Saunders, and et al\. \(2026\)Emotion concepts and their function in a large language model\.External Links:2604\.07729Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p1.1)\.
- A\. N\. Tak, A\. Banayeeanzade, A\. Bolourani, M\. Kian, R\. Jia, and J\. Gratch \(2025\)Mechanistic interpretability of emotion inference in large language models\.InFindings of the Association for Computational Linguistics: ACL 2025,External Links:2502\.05489Cited by:[§1](https://arxiv.org/html/2607.18691#S1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2607.18691#S1.p1.1),[§10](https://arxiv.org/html/2607.18691#S10.p4.1),[§2\.2](https://arxiv.org/html/2607.18691#S2.SS2.p1.1),[§4](https://arxiv.org/html/2607.18691#S4.SS0.SSS0.Px1.p1.4),[§8](https://arxiv.org/html/2607.18691#S8.p2.2)\.
- A\. Templeton, T\. Conerly, J\. Marcus, J\. Lindsey, and et al\. \(2026\)Scaling monosemanticity: extracting interpretable features from claude 3 sonnet\.External Links:2605\.29358Cited by:[§2\.2](https://arxiv.org/html/2607.18691#S2.SS2.p1.1)\.
- C\. Tigges, O\. J\. Hollinsworth, A\. Geiger, and N\. Nanda \(2024\)Language models linearly represent sentiment\.InProceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP,pp\. 58–87\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.blackboxnlp-1.5)Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p1.1)\.
- E\. Troiano, L\. A\. M\. Oberländer, and R\. Klinger \(2023\)Dimensional modeling of emotions in text with appraisal theories: corpus creation, annotation reliability, and prediction\.Computational Linguistics49\(1\),pp\. 1–72\.Cited by:[§4](https://arxiv.org/html/2607.18691#S4.SS0.SSS0.Px1.p1.4)\.
- J\. Vig, S\. Gehrmann, Y\. Belinkov, S\. Qian, D\. Nevo, Y\. Singer, and S\. Shieber \(2020\)Investigating gender bias in language models using causal mediation analysis\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p2.1)\.
- C\. Wang, Y\. Zhang, R\. Yu, and et al\. \(2025\)Do llms "feel"? emotion circuits discovery and control\.External Links:2510\.11328Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p1.1)\.
- C\. Wendler, V\. Veselovsky, G\. Monea, and R\. West \(2024\)Do llamas work in english? on the latent language of multilingual transformers\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(ACL\),External Links:2402\.10588Cited by:[§10](https://arxiv.org/html/2607.18691#S10.p3.1)\.
- A\. Wierzbicka \(1972\)Semantic primitives\.Athenäum,Frankfurt\.Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p3.2)\.
- A\. Wierzbicka \(1996\)Semantics: primes and universals\.Oxford University Press\.Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p3.2)\.
- A\. Wierzbicka \(1999\)Emotions across languages and cultures: diversity and universals\.Cambridge University Press\.Cited by:[§10](https://arxiv.org/html/2607.18691#S10.p2.1),[§7](https://arxiv.org/html/2607.18691#S7.SS0.SSS0.Px2.p1.7)\.
- J\. Woodward \(2003\)Making things happen: a theory of causal explanation\.Oxford University Press\.Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p4.1)\.
- X\. Wu, H\. Wang, Z\. Yan, and et al\. \(2026\)Decoding and controlling emotion in llms through human\-aligned representational geometry with enhanced interpretability\.Computers in Human Behavior183\(109051\),pp\. 1–17\.Cited by:[§1](https://arxiv.org/html/2607.18691#S1.p1.1),[§2\.2](https://arxiv.org/html/2607.18691#S2.SS2.p1.1)\.
- F\. Xing and E\. Cambria \(2026\)Faithful by definition: emotion analysis via natural semantic metalanguage explications\.External Links:2607\.00661Cited by:[§2\.2](https://arxiv.org/html/2607.18691#S2.SS2.p1.1)\.
- Z\. Zhao, E\. Wallace, S\. Feng, D\. Klein, and S\. Singh \(2021\)Calibrate before use: improving few\-shot performance of language models\.InProceedings of the 38th International Conference on Machine Learning \(ICML\),Cited by:[§7](https://arxiv.org/html/2607.18691#S7.SS0.SSS0.Px1.p2.10)\.

## Appendix AThe Contrastive Suite: Inventory, Confounds, Synthetic Validation

#### Symbol classes and the full inventory\.

The 32 tested primes are mental predicates \(want, feel, think, know, say, do\), evaluators \(good, bad\), modals \(can, maybe, not, true\), person primes \(i, someone, other, people\), temporal \(now, before, after, a\-long\-time, a\-short\-time\), descriptors \(big, small, very, more, same\), bodily and animate \(body, live, die\), and reason \(because, if, happen\)\. The untested∼\\sim33 are dominated by quantifiers \(one, two, some, all, much, few\), spatial primes \(here, above, below, near, far, inside, touch, side, where\), and the referential substantives and determiners \(you, something, this, kind, part\); none of these appears in a gold emotion explication, and none admits a clean single\-prime assertion\-contrast\. Of the 22 distinct primes used across the 13 emotion recipes, 21 are tested \(onlythisis missing\)\.

#### Splits, confounds, validation\.

Split is done 70/30 at the template level\. Negated sub\-sets \(want, know, can, true\) and graded sub\-sets \(time3, valence3, size3, duration2\) test partial orderings inside a prime family\. Zero negatives leak the prime’s exponent; 13 primes show\|d\|\>0\.5\|d\|\>0\.5on word count, the most extremesayatd=2\.03d\{=\}2\.03\. The harness recovers a planted layer for all six synthetic primes with off\-layer accuracy and length baseline at chance \(Figure[5](https://arxiv.org/html/2607.18691#A1.F5)\)\.

![Refer to caption](https://arxiv.org/html/2607.18691v1/Figures/fig_synthetic_sweep.png)Figure 5:Synthetic validation: a prime\-style signal planted at layer 7 of a random\-initialized 12\-layer transformer is recovered for all six synthetic primes \(the spike\); off\-layer accuracy and length baseline at chance\.

## Appendix BLocalization: Full Curves and Geometry

Emotion accuracy peaks at0\.9430\.943\(linear\) at layer 13 and the MLP at0\.9470\.947at layer 14; mean appraisalR2R^\{2\}reaches0\.300\.30in the same band, led by*pleasantness*\(0\.700\.70\) and*unpleasantness*\(0\.650\.65\) \(Figure[6](https://arxiv.org/html/2607.18691#A2.F6)\)\. The emotion geometry is dominated by valence: measured as the Pearson correlation betweenwe⊤​hw\_\{e\}^\{\\top\}handwa⊤​hw\_\{a\}^\{\\top\}h, the strongest cells are emotions on pleasantness and unpleasantness, and the agency contrast is weaker \(anger loads0\.150\.15on other\- versus−0\.30\-0\.30on self\-responsibility; Figure[7](https://arxiv.org/html/2607.18691#A2.F7)\)\. Zeroing a 3\-layer span at the event position collapses agreement with the clean model to0\.110\.11at layer 9, while a norm\-matched random perturbation is less damaging; activation patching peaks at0\.8050\.805at layer 10 \(Figure[8](https://arxiv.org/html/2607.18691#A2.F8)\)\.

![Refer to caption](https://arxiv.org/html/2607.18691v1/x5.png)Figure 6:Layer\-wise probing of Llama\-3\.2\-1B on crowd\-enVent\. Emotion accuracy \(left\) and mean appraisalR2R^\{2\}\(right\) saturate late, fixing the consolidation band; chance accuracy is1/13=0\.081/13=0\.08\.![Refer to caption](https://arxiv.org/html/2607.18691v1/x6.png)Figure 7:Decoded\-score correlationρ\\rhobetween the 13 emotion readouts and the appraisal readouts at layer 13\. Valence dominates; the agency contrast is fainter, visible as anger loading on other\- over self\-responsibility\.![Refer to caption](https://arxiv.org/html/2607.18691v1/x7.png)Figure 8:Residual\-stream knockout and patching\. Zeroing a 3\-layer span at the event position \(red\) maximally disrupts the emotion decision; a random perturbation \(green\) is less damaging\. Patching peaks at layer 10 \(flip rate0\.8050\.805\)\.
## Appendix CBehavioral Tests and In\-Context Decodability

The behavioral tests use only generations and answer\-token logits, so they can run on all four models\. The 13 emotion\-token logits are read for: \(1\) the emotion word, \(2\) its gold explication, \(3\) a same\-affect decoy explication, and \(4\) a prime\-scrambled control\. Sufficiency ablates the explication one prime at a time \(central vs peripheral by the gold recipe\); the definitional floor scores the fraction of restatement tokens more frequent than the target\. Table[3](https://arxiv.org/html/2607.18691#S7.T3)has collected the numbers; Figure[9](https://arxiv.org/html/2607.18691#A3.F9)shows the full recipe\-match matrix and Table[5](https://arxiv.org/html/2607.18691#A3.T5)the in\-context decodability behind the geometric non\-reduction \(onlyfeelandsomeoneclear\+0\.10\+0\.10\)\.

![Refer to caption](https://arxiv.org/html/2607.18691v1/x8.png)Figure 9:Emotion\-to\-recipe match matrix \(Llama\-3\.2\-1B\)\. Rows are emotion readout directions, columns recipe\-implied directions; a correct linear decomposition would put the maximum on the diagonal\. Only a few emotions, led by anger, match their own recipe\.Table 4:Linear emotion\-to\-recipe match for prime directions fit from the contrastive suite \(Llama\-3\.2\-1B\)\. Chance is0\.080\.08\(rank\-1\) and0\.230\.23\(rank\-3\)\. The in\-context source, the only one without a space mismatch, is at chance because the primes are consumed by the consolidation layer; linear composition is weak, in contrast to the behavioral reduction in Section[7](https://arxiv.org/html/2607.18691#S7)\.Table 5:In\-context prime decodability at the consolidation layer, answer position \(selected rows\)\. Lift is held\-out minus majority\. Onlyfeelandsomeoneclear\+0\.10\+0\.10\.

Similar Articles

Negative Before Positive: Asymmetric Valence Processing in Large Language Models

arXiv cs.CL

This paper investigates how large language models process emotional valence through mechanistic interpretability. Using activation patching and steering on three open-source LLMs, the authors find that negative valence is localized to early layers while positive valence peaks in mid-to-late layers, and they validate this through topic-controlled flip tests.

Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

arXiv cs.CL

This paper replicates the finding of 'emotion vectors' in open-weight LLMs Apertus-8B and Gemma-4-E4B, showing that valence geometry is recoverable across models with differences in layer emergence. The study also finds that arousal encoding is sensitive to the story corpus used for extraction.