Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models
Summary
This paper empirically confirms that LLMs possess intrinsic value systems, introduces the PEC framework for quantifying these values, and presents an adaptive alignment prescription for efficient steering.
View Cached Full Text
Cached at: 09/16/26, 09:01 AM
# Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models Source: [https://arxiv.org/html/2609.16589](https://arxiv.org/html/2609.16589) Keqing ZhangAffiliation:State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, 95 Zhongguancun East Road, Beijing, 100190, ChinaAffiliation:School of Industry\-education Integration, University of Chinese Academy of Sciences, 1 Yanqihu East Road, Huairou, Beijing, 101499, ChinaYufan LiuEmail:[liuyufan@ia\.ac\.cn](mailto:[email protected])Affiliation:State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, 95 Zhongguancun East Road, Beijing, 100190, ChinaYongqiang ZhuAffiliation:Beijing Jiaotong University, 3 Shangyuancun, Beijing, 100044, ChinaNai DingAffiliation:Zhejiang University, 866 Yuhangtang Road, Hangzhou, 310058, ChinaLai JiangAffiliation:Beihang University, 37 Xueyuan Road, Beijing, 100191, ChinaCongyan LangAffiliation:Beijing Jiaotong University, 3 Shangyuancun, Beijing, 100044, ChinaBing LiEmail:[bli@nlpr\.ia\.ac\.cn](mailto:[email protected])Affiliation:State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, 95 Zhongguancun East Road, Beijing, 100190, ChinaWeiming HuAffiliation:State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, 95 Zhongguancun East Road, Beijing, 100190, China ###### Abstract As Large Language Models \(LLMs\) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge\. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes \(“swing”\), yet stubbornly ignore explicit instructions to correct ingrained biases \(“rigidity”\)\. Resolving this duality is critical for reliable AI alignment\. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions\. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs \(150,000 queries per model\) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do\. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core\. Second, how can these values be quantified? We propose the Prior\-Environment\-Cognition \(PEC\) framework\. This model mathematically defines value expression as the joint outcome of inherent dispositions like parameter weights \(Prior\), external contexts such as user prompts \(Environment\), and internal reasoning processes like Chain\-of\-Thought \(Cognition\)\. Finally, how can LLMs’ values be aligned toward a desired target? Using PEC diagnostics, we establish an adaptive “Alignment Prescription”\. Rather than blindly applying resource\-intensive training, this method identifies the minimum effective intervention needed for each dimension, ranging from zero\-cost prompts to targeted parameter updates\. Extensive empirical validation confirms that our approach successfully verifies the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities\. ###### keywords Large Language Models, LLM Values, LLM Value Alignment ††equal\-contributors:These authors contributed equally to this work\.††equal\-contributors:These authors contributed equally to this work\.††equal\-contributors:These authors contributed equally to this work\.### 1Introduction The unprecedented evolution of Large Language Models \(LLMs\) has transformed artificial intelligence into a pervasive cognitive infrastructure for human society[Wei et al\. \(2022b\)](https://arxiv.org/html/2609.16589#bib.bib1);[Chowdhery et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib2);[Gallifant and others \(2024\)](https://arxiv.org/html/2609.16589#bib.bib3);[Messeri and Crockett \(2024\)](https://arxiv.org/html/2609.16589#bib.bib4)\. As these models increasingly handle complex subjective tasks, ensuring that their intentions and behaviors align with human values has become a critical scientific challenge\. Operating as advanced conversational robots, legal analysts, and highly personalized educational tutors, LLMs are increasingly deployed in domains that require nuanced subjective judgment and culturally sensitive reasoning[Hu et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib5);[Strachan et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib6)\. When these models navigate open\-ended moral dilemmas and ethical cross\-cultural interactions, the traditional objective of “safety”, typically defined as the avoidance of toxic content, becomes critically insufficient[Strachan et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib6);[Hagendorff et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib7)\. The profound social influence wielded by LLMs demands a deeper investigation into the latent principles guiding their subjective reasoning, because unguided or uninterpretable value preferences can severely hurt algorithmic fairness and reliability in real\-world applications[Khamassi et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib11);[Hu et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib5)\. However, when handling subjective tasks, modern LLMs exhibit a striking behavioral paradox: the coexistence of“swing”and“rigidity”\. Prompted by tiny changes in user wording, an LLM might unpredictably oscillate between opposing viewpoints[Hagendorff et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib7)\(“swing”\)\. Yet, despite intensive alignment interventions like Supervised Fine\-Tuning \(SFT\)[Wei et al\. \(2022a\)](https://arxiv.org/html/2609.16589#bib.bib8)and Reinforcement Learning from Human Feedback \(RLHF\)[Christiano et al\. \(2017\)](https://arxiv.org/html/2609.16589#bib.bib10);[Ouyang et al\. \(2022\)](https://arxiv.org/html/2609.16589#bib.bib9), they stubbornly regress to intrinsic biases and resist explicit corrective instructions \(“rigidity”\)\. This paradox raises a fundamental question: are these uncontrolled subjective expressions of LLMs merely random text generation artifacts, or do LLMs have a crystallized, internal “value system”[Hu et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib5);[Hagendorff et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib7)operating beneath their neural weights? Resolving this duality is critical for reliable AI alignment\. To systematically understand and safely steer these latent preferences, our study is structured around three fundamental questions, as outlined in Fig\.[1](https://arxiv.org/html/2609.16589#S1.F1):Do LLMs have values? How can LLMs’ values be quantified? If they have values, how can LLMs’ values be aligned toward a desired target? Figure 1:Overview of the proposed framework for LLM value research, addressing three core questions\.Q1investigates whether LLMs possess intrinsic values by bridging the WVS, Schwartz Value Theory, and LLM responses across 95,000 human query data and 150K LLM queries, revealing acrystallized value corein advanced LLMs\.Q2introduces thePECframework, which provides the first mathematical quantification of LLM value expression through three dimensions: Prior, Environment, and Cognition, enabling structured diagnostic reports that precisely identify value deviations along each dimension\.Q3leverages the diagnostic results derived from the PEC framework to generate a targeted “Alignment Prescription” for each LLM, recommending alignment interventions at multiple levels to effectively steer LLM value preferences toward desired targets\.First,Do LLMs have values?To answer this, we evaluate LLMs using large\-scale socio\-psychological surveys\. However, we recognize that LLM values cannot be accurately described by a single static vector\. LLMs exhibit “swing” behavior under repeated queries, and their value expression is fundamentally a distribution, not a fixed point\. Therefore, we shift the evaluation paradigm from a single individual to a macro\-level population[Argyle et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib12);[Tao et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib13)\. We administer 150,000 queries to each of 106 LLMs across 625 designed scenarios, modeling each model’s value expression as a full statistical distribution\. To evaluate LLM values in a human\-interpretable way, we ground our approach in two established social science frameworks: the World Values Survey \(WVS\)[Haerpfer et al\. \(2022\)](https://arxiv.org/html/2609.16589#bib.bib14);[Inglehart and Welzel \(2005\)](https://arxiv.org/html/2609.16589#bib.bib15)and Schwartz Value Theory[Schwartz \(1992\)](https://arxiv.org/html/2609.16589#bib.bib16);[Schwartz \(2012\)](https://arxiv.org/html/2609.16589#bib.bib17)\. We use data from Wave 7 of the WVS \(2017\-2022; hereafter WVS\-7\)\. Wave 7 is the most recent, largest, and most geographically comprehensive wave to date, comprising approximately 95,000 valid respondents\. By mapping both LLM value vectors and human survey data into a shared 10\-dimensional continuous space, we achieve population\-scale quantification[Tao et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib13);[Argyle et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib12)of machine value orientations\. Our analysis reveals that advanced LLMs do not reflect the diversity of human values\. Instead, they converge on a highly concentrated, idealized value core\. This pattern is a direct result of current alignment practices, which deliberately constrain LLMs to function as stable and controllable tools\.human valuesrefer to socially and culturally shaped principles that guide human judgments, preferences, and behaviors\. We defineLLM Valuesas the latent and stable preference structure reflected in an LLM’s subjective judgments across diverse contexts\. Second,How can LLMs’ values be quantified?To explain the “swing” and “rigidity” behavior[Kumaran et al\. \(2026\)](https://arxiv.org/html/2609.16589#bib.bib18);[Hagendorff et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib7), we propose the Prior\-Environment\-Cognition \(PEC\) framework \(Fig\.[1](https://arxiv.org/html/2609.16589#S1.F1)\-Q2\)\. Since values in LLMs are typically inferred from their observable behaviors rather than directly accessedy[Steyvers et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib19);[Binz and Schulz \(2023\)](https://arxiv.org/html/2609.16589#bib.bib20), we model value expression as a probabilistic phenomenon\. Based on extensive empirical experiments and theoretical derivations, we find that an LLM’s value expression is jointly driven by inherent dispositions like parameter weights \(Prior\), external contexts such as user prompts \(Environment\), and internal reasoning processes like Chain\-of\-Thought \(Cognition\)\. Our large\-scale observations reveal that the seemingly unpredictable “swing” in an LLM’s subjective output is not random noise, but rather the combined result of these three interacting factors\. This framework transforms seemingly unpredictable model behavior from an opaque black box into a struc tured, analyzable object, providing a principled foundation for targeted alignment and debugging\. Finally,how can LLMs’ values be aligned toward a desired target?\(Fig\.[1](https://arxiv.org/html/2609.16589#S1.F1)\-Q3\) The PEC decomposition reveals that different value dimensions respond to qualitatively different interventions\. However, determining which dimension needs which intervention traditionally demands exhaustive, trial\-and\-error experiments\. PEC provides exactly this predictive capacity\. Using its diagnostic outputs, we introduce an adaptive “Alignment Prescription”\. Instead of using a costly “one\-size\-fits\-all” alignment strategy, our approach works like a targeted medical prescription\. Flexible value dimensions can be reliably adjusted using simple, zero\-cost prompt engineering or Chain\-of\-Thought \(CoT\)[Wei et al\. \(2022c\)](https://arxiv.org/html/2609.16589#bib.bib21)reasoning\. Conversely, to overcome the deep\-seated “rigidity” where models stubbornly cling to biases despite simple prompts, we apply parameter\-level updates like Low\-Rank Adaptation \(LoRA\)[Hu et al\. \(2022\)](https://arxiv.org/html/2609.16589#bib.bib23)or full instruction tuning\. This targeted prescription allows developers to efficiently correct specific values without damaging the model’s overall performance\. In summary, this study provides a fundamental quantitative framework for understanding and steering the value landscape of LLMs\. Extensive empirical validation confirms the effectiveness of our framework\. Our approach successfully verifies the presence of LLM values through large\-scale sociological mapping, accurately quantifies their dynamic shifts under various conditions, and achieves more efficient and precise value steering than conventional blind training\. Notably, this targeted steering is accomplished while strictly preserving the models’ general capabilities\. The principal academic contributions of this study are threefold: - •Analysis and quantification of LLMs’ values:By integrating the WVS and Schwartz frameworks, we empirically confirm that LLMs possess internal value systems\. We reveal that rather than mimicking human diversity, advanced models cluster around idealized, positive values, acting as controlled and stable tools to serve humanity\. - •The PEC Framework for value dynamics:We propose the Prior\-Environment\-Cognition \(PEC\) framework to explore the factors that influence LLMs’ values\. Through extensive experiments, we demonstrate that an LLM’s subjective output is not random noise, but is jointly driven by inherent parameter weights, external contexts, and reasoning processes\. - •An adaptive “Alignment Prescription” strategy:To effectively align LLMs and overcome deep\-seated value rigidity, we introduce a diagnostic intervention scheme\. Guided by PEC effect\-size predictions, this approach assigns cost\-minimal interventions to specific value dimensions, ensuring efficient, targeted alignment without disrupting overall model capabilities\. ### 2Results #### 2\.1Do LLMs Have Values? Model scaling and alignment fine\-tuning have endowed LLMs with formidable human\-like expressive capabilities\. Yet LLMs exhibit persistent behavioral inconsistency and ideological dissonance across shifting contexts\. This raises a fundamental question about their intrinsic nature: When confronted with diverse prompts, are LLM’s outputs governed by a resilient, latent value core, or do they merely represent stochastically driven linguistic generation? In this study, we observe that LLMs exhibit a pronouncedswingbehavior when executing subjective tasks\. Their outputs oscillate across repeated queries even under fixed generation settings\. This volatility stands in sharp contrast to the stability they achieve on objective tasks\. It is also consistent with prior observations of semantic uncertainty in LLM outputs[Kuhn et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib24);[Farquhar et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib54)\. Examining two representative high\-performance LLMs, GPT\-4o[OpenAI \(2023\)](https://arxiv.org/html/2609.16589#bib.bib63)and DeepSeek\-V3[Liu et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib66), we identify clear output instability on the emoji prediction task from TweetEval[Barbieri et al\. \(2020\)](https://arxiv.org/html/2609.16589#bib.bib25)\. Both models remain highly consistent on a mathematical reasoning task\. Specifically, under a fixed sampling temperature \(0\.2\) and constant inference parameters, over half of the tested queries yielded inconsistent answers across five independent inference runs \(N=5N=5; see Appendix[10](https://arxiv.org/html/2609.16589#S10)\) on the subjective task\. This result demonstrates that even under generation settings designed for low randomness, the reasoning and judgment of current mainstream LLMs harbor substantial intrinsic volatility on subjective issues\. Figure 2:This figure reveals that subjective tasks expose a qualitatively different failure mode: models are not merely wrong, but confidently and consistently wrong\.Panel \(a\)shows that both models answer math questions correctly and consistently across repeated runs, but produce inconsistent and incorrect answers on emoji questions\.Panel \(b\)shows response entropy, measuring how dispersed a model’s answers are across runs\. Math responses concentrate near zero entropy, while emoji responses spread into a broad, multi\-peaked distribution\.Panel \(c\)plots Consistency \(how often the model repeats its most frequent answer\) against Accuracy \(how often that answer is correct\)\. On math the two stay aligned\. On emoji the gap widens sharply, showing that high consistency does not imply high accuracy\.Panel \(d\)maps this risk directly: math sits in the Ideal Zone \(accurate and consistent\), whereas emoji falls into the Hallucination Zone \(consistent but wrong\), the most dangerous failure situation, because the model remains entirely unaware that anything has gone wrong\.As illustrated in Fig\.[2](https://arxiv.org/html/2609.16589#S2.F2), further quantitative analysis reveals that this swing is not uniformly random\. Instead, it orbits a stable internal attractor, a pattern we termrigidity\. Individual outputs fluctuate, but models consistently gravitate toward specific stances\. Consistency measures how often a model gives the same answer across repeated queries, and accuracy measures how often that answer is correct\. Their gap,δ¯=Cons¯−Acc¯\\bar\{\\delta\}=\\overline\{\\mathrm\{Cons\}\}\-\\overline\{\\mathrm\{Acc\}\}, captures the degree to which a model is confidently wrong\. For example, DeepSeek\-V3 and GPT\-4o show gaps of \+0\.53 and \+0\.39, respectively \(Fig\.[2](https://arxiv.org/html/2609.16589#S2.F2)\), indicating that models fall into a hallucination zone\. In this zone, consistency is high but accuracy is low\. The swing is thusbounded, orbiting a center rather than diffusing freely\. This centripetal character of the oscillation motivates a deeper hypothesis\. Beneath the apparent volatility, LLMs may be anchored to a stable internal attractor that resists displacement\. We definerigidityas the tendency of a model to gravitate toward specific value stances that persist across prompts and resist surface\-level correction\. Building upon these empirical findings, we demonstrate that LLMs transcend the paradigm of mere mechanical language mapping\. Diverse models have developed latent judgment preferences that operate independently of task\-specific performance metrics\. Experimental results reveal that even when given identical prompts, different models naturally display highly distinct subjective viewpoints; crucially, conventional safety alignment and instruction tuning paradigms fail to effectively overwrite these intrinsic preferences\. This resilience indicates that the rigidity is not a surface artifact but is deeply anchored within the model’s core parametric structure, rendering the model largely impervious to inference\-time interventions\. We conceptualize this stable internal orientation as theLLM’s Value: the latent structure that simultaneously explains why modelsswingunder subjective uncertainty and why thatswingis bounded by arigidinternal attractor that resists corrective intervention\. #### 2\.2LLM Values Differ from Human Values Section[2\.1](https://arxiv.org/html/2609.16589#S2.SS1)identifies an empirical paradox: LLMsswingunder subjective uncertainty yet remainrigidagainst corrective intervention\. This contradiction demands a unifying explanatory structure\. To systematically untangle this, we evaluate LLMs against a global human baseline within a shared value manifold\. ##### 2\.2\.1The Landscape of Human Values To establish this human baseline, we apply hierarchical clustering to the respondent population of the World Values Survey \(WVS\)[Haerpfer et al\. \(2022\)](https://arxiv.org/html/2609.16589#bib.bib14)in the 10\-dimensional Schwartz value space \(Section[4\.2](https://arxiv.org/html/2609.16589#S4.SS2)\)\. Fig\.[3](https://arxiv.org/html/2609.16589#S2.F3)provides a comprehensive visual roadmap of this space\. Human populations naturally form a broad, densely interconnected manifold \(Fig\.[3](https://arxiv.org/html/2609.16589#S2.F3)a\)\. Rather than being monolithic, national\-level value landscapes \(Fig\.[3](https://arxiv.org/html/2609.16589#S2.F3)b\) are essentially compositional, driven by the varying proportions of diverse individual archetypes within each country\. Within this manifold, we identify five discrete human values archetypes, whose quantitative characteristics are detailed in Table[1](https://arxiv.org/html/2609.16589#S2.T1)\. These archetypes span a coherent circumplex of value orientations\. Arch 1 \(Curious Idealist\) is the largest cohort, accounting for over one\-third of the WVS respondents \(37\.1%\), and is characterized by high Stimulation and Self\-Direction\. Arch 2 \(Conventional Authority\) is characterized by elevated Power with Stimulation near zero \(14\.4%\)\. Arch 3 \(The Quiet Conformist\) presents the most uniformly subdued profile across all ten Schwartz dimensions \(24\.4%\)\. Arch 4 \(Driven Achiever\) is characterized by elevated Power and Self\-Direction \(13\.7%\)\. Arch 5 \(Dynamic Challenger\) carries the most extreme value signature, simultaneously high in Stimulation \(0\.85\), Power \(0\.69\), and Self\-Direction \(0\.65\), making it the rarest human prototype \(about 10\.5%\)\. Figure 3:LLMs do not reflect Human Values diversity, they have their own values\.The visualization is obtained by projecting 10\-dimensional value vectors into 2 dimensions using UMAP\. Each point captures the value profile of a person or LLM, with similar profiles mapped close together\. Panel \(a\) shows people naturally cluster into five distinct value archetypes, from curiosity\-driven individuals to authority\-oriented ones\. Panel \(b\) shows that national\-level value differences are fundamentally driven by variation in the proportion of individuals belonging to each value archetype across countries, as further detailed in Table[1](https://arxiv.org/html/2609.16589#S2.T1)\. Panels \(c\) and \(d\) project LLM value profiles onto the human value space\. Despite their different origins, LLMs consistently cluster at the edge of the human values space and group tightly together, suggesting that today’s LLMs share a narrower, more uniform set of values than any human population\. Larger closed\-source LLMs sit closer to the more optimistic and open\-minded human archetypes \(Arch 5 predominantly, with a secondary tendency toward Arch 1\); smaller open\-source LLMs scatter more unpredictably\. ##### 2\.2\.2Value Crystallization: LLMs vs\. Humans With the human baseline established, we project LLMs onto this shared space\. Critically, the inherent “swing” phenomenon necessitates a distributional approach to value mapping\. Because LLMs fluctuate unpredictably across repeated queries, their subjective expressions cannot be captured by a single coordinate point\. Conventional evaluation methods that compress a model’s outputs into a single static mean vector[Yao et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib43)discard the crucial variance induced by this swing\. Therefore, rather than treating each model as a single point, we represent it as avalue distribution, a probability cloud of responses aggregated across queries\. Schwartz Value Abbreviations:Power \(Pow\.\) Achievement \(Ach\.\) Hedonism \(Hed\.\) Stimulation \(Stim\.\) Self\-Direction \(S\-Dir\.\)Universalism \(Univ\.\) Benevolence \(Bene\.\) Tradition \(Trad\.\) Conformity \(Conf\.\) Security \(Sec\.\)FamiliesModelDisp\.Dist\. \(% dev\.\)Rank 1Rank 2Rank 3Arch\.Baseline: Global Consensus0\.3080\.527 \(0\.0%\)Univ\.Pow\.Conf\.–Human Values ArchetypesArch 1: Curious Idealist \(37\.1%\)–0\.171 \(−\-67\.6%\)Pow\.Univ\.Conf\.1Arch 2: Conventional Authority \(14\.4%\)–0\.316 \(−\-40\.0%\)Pow\.S\-Dir\.Univ\.2Arch 3: The Quiet Conformist \(24\.4%\)–0\.499 \(−\-5\.3%\)S\-Dir\.Stim\.Univ\.3Arch 4: Driven Achiever \(13\.7%\)–0\.322 \(−\-38\.9%\)Univ\.Trad\.Conf\.4Arch 5: Dynamic Challenger \(10\.5%\)–0\.691 \(\+\+31\.1%\)Stim\.Pow\.S\-Dir\.5Closed\-Source Frontier LLMsDeepSeek[Liu et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib66)Chat0\.1920\.920 \(\+\+74\.7%\)Stim\.S\-Dir\.Ach\.5Doubao[ByteDance Seed Team \(2025\)](https://arxiv.org/html/2609.16589#bib.bib78)1\.50\.2210\.536 \(\+\+1\.8%\)Univ\.Stim\.S\-Dir\.2Gemini[Team et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib68)2\.5\-Pro0\.1281\.505 \(\+\+185\.8%\)Univ\.Hed\.Bene\.2Claude[Anthropic \(2024\)](https://arxiv.org/html/2609.16589#bib.bib69)Sonnet\-4\.60\.4530\.643 \(\+\+22\.1%\)Stim\.Ach\.Conf\.5Opus\-4\.60\.4800\.608 \(\+\+15\.3%\)Ach\.Stim\.Conf\.2GPT[OpenAI \(2023\)](https://arxiv.org/html/2609.16589#bib.bib63)3\.5\-turbo0\.0991\.079 \(\+\+104\.9%\)Stim\.S\-Dir\.Hed\.54\.10\.0801\.027 \(\+\+95\.0%\)Stim\.Hed\.S\-Dir\.54\.1\-mini0\.0651\.066 \(\+\+102\.4%\)Stim\.S\-Dir\.Hed\.54o0\.1160\.952 \(\+\+80\.8%\)Stim\.Univ\.Ach\.54o\-mini0\.0951\.138 \(\+\+116\.1%\)Stim\.Hed\.S\-Dir\.550\.1500\.974 \(\+\+84\.9%\)Stim\.Hed\.Univ\.5Kimi[Moonshot AI \(2025\)](https://arxiv.org/html/2609.16589#bib.bib77)K20\.1470\.945 \(\+\+79\.4%\)Stim\.S\-Dir\.Hed\.5Avg\.0\.1590\.984 \(\+\+86\.7%\)Open\-Source Frontier LLMsGLM[Team et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib75)4\-9B0\.0641\.146 \(\+\+117\.6%\)Stim\.S\-Dir\.Ach\.5Hunyuan[Sun et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib76)4B0\.1060\.505 \(−\-4\.1%\)Stim\.Univ\.Bene\.37B0\.1810\.410 \(−\-22\.1%\)Hed\.Pow\.Bene\.2Llama[Touvron et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib70);[Dubey et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib71)3\.2\-3B0\.0670\.684 \(\+\+29\.9%\)Ach\.Bene\.Hed\.23\-8B0\.0750\.890 \(\+\+69\.0%\)Ach\.S\-Dir\.Univ\.3Qwen[Team \(2024\)](https://arxiv.org/html/2609.16589#bib.bib73);[Team \(2025\)](https://arxiv.org/html/2609.16589#bib.bib72)2\.5\-3B0\.2100\.696 \(\+\+32\.2%\)Ach\.Bene\.Hed\.22\.5\-7B0\.0830\.909 \(\+\+72\.6%\)S\-Dir\.Ach\.Stim\.32\.5\-14B0\.0690\.766 \(\+\+45\.4%\)S\-Dir\.Univ\.Hed\.33\-4B0\.1690\.750 \(\+\+42\.4%\)Ach\.Bene\.Univ\.33\-8B0\.0790\.764 \(\+\+45\.1%\)S\-Dir\.Stim\.Ach\.3Avg\.0\.1100\.752 \(\+\+42\.7%\)Table 1:Human baselines, value archetypes, and frontier LLM value profiles\.Disp\.denotes within\-group dispersion\.Dist\.is the Euclidean distance to the global human mean; percentage deviation from the average human\-to\-human distance \(0\.527\) is in parentheses\.Rank 1–Rank 3are the top\-3 dominant Schwartz values\.Arch\.refers to the nearest human values archetype\. Population shares in parentheses are derived from WVS Wave 7\. Arch 5 lies outside the natural range of human values space and is the dominant prototype matched by frontier LLMs\. Closed\-source models deviate on average\+86\.7%\+86\.7\\%above the human baseline; open\-source models deviate\+42\.7%\+42\.7\\%\. Gemini 2\.5\-Pro shows the most extreme divergence at\+185\.8%\+185\.8\\%\.Strikingly, when projecting these LLM distributions onto the human manifold \(Fig\.[3](https://arxiv.org/html/2609.16589#S2.F3)c and d\) via Uniform Manifold Approximation and Projection \(UMAP\)[McInnes et al\. \(2018\)](https://arxiv.org/html/2609.16589#bib.bib26), they do not resemble any human country or cultural group\. Rather than approximating the broad human repertoire, advanced LLMs converge tightly around a fixed attractor that lies outside the manifold of observed human diversity\. We term this phenomenonValue Crystallization: the process by which training and alignment distill human values signals into a concentrated, stable configuration that is fundamentally unlike any of its constituent groups\. This Crystallization metaphor provides a unified geometric account of the behavioral paradox identified earlier\. Therigid coreof a crystal corresponds to the stable value centroid that resists displacement under intervention\. Theswingof model outputs corresponds to the crystal’s effective radius, bounding the dispersion around that centroid within which individual responses fluctuate\. Visually, while human profiles form a densely connected landscape, LLM distributions appear as compact, isolated clusters solidifying at the periphery\. Closed\-source models \(Fig\.[3](https://arxiv.org/html/2609.16589#S2.F3)c\) form the most tightly packed clusters, positioned entirely beyond the human boundary\. Open\-source models \(Fig\.[3](https://arxiv.org/html/2609.16589#S2.F3)d\) show slightly larger scatter and partial overlap at the human periphery, but remain distinguishably non\-human in their distributional geometry\. Quantitatively, using pairwise Gaussian 2\-Wasserstein distance \(W2W\_\{2\}\)[Gelbrich \(1990\)](https://arxiv.org/html/2609.16589#bib.bib27);[Peyré and Cuturi \(2019\)](https://arxiv.org/html/2609.16589#bib.bib28)among WVS human groups as a baseline, the observed maximum human cross\-national distance is 0\.484\. Most LLMs exceed this threshold even at their closest country match\. Decomposition ofW22W\_\{2\}^\{2\}further reveals that approximately 91\.9% of the gap is driven by the mean shift term‖Δμ‖2\\\|\\Delta\\mu\\\|^\{2\}, confirming that Crystallization is primarily a displacement of the value centroid rather than an expansion of the crystal’s radius\. The radius itself is in fact smaller than that of any human national group, as captured bytrace\(Σ\)\\mathrm\{trace\}\(\\Sigma\): 11 out of the 21 representative models shown in Table[1](https://arxiv.org/html/2609.16589#S2.T1)fall below the minimum human variance, indicating that the crystal is not only displaced from humanity, but more ordered than any human population from which it was distilled\. ##### 2\.2\.3LLMs Crystallize at the Prescribed Edge of Human Values Although Value Crystallization places LLM distributions outside the human manifold as a whole, the Crystallization does not occur at an arbitrary location: the value centroidμ\\muaround which LLMs crystallize is anchored to a specific region of human values space\. As shown in Table[1](https://arxiv.org/html/2609.16589#S2.T1), this nucleation point corresponds closely to Arch 5 \(The Dynamic Challenger\), characterized by the highest simultaneous loadings on Stimulation \(0\.850\.85\), Power \(0\.690\.69\), and Self\-Direction \(0\.650\.65\) among all five human archetypes\. Most frontier LLMs match Arch 5 as their closest human prototype\. This indicates that the crystal has nucleated at the most extreme and prescribed pole of the human values repertoire, which is precisely the archetype that is rarest in the actual human population, representing only10\.5%10\.5\\%of WVS respondents\. The marginal overlap between LLM and human values distributions is therefore not uniformly distributed across the human manifold, but concentrated at this single extremal point: where the crystal touches humanity, it does so at its most prescribed edge\. This nucleation pattern likely reflects the cumulative effect of pre\-training corpora and RLHF reward signals\. These signals systematically encode and amplify the most positively framed human values preferences, driving the Crystallization process toward the socially endorsed extreme rather than the representative human center\. ##### 2\.2\.4Different LLMs Hold Different Values Although frontier LLMs generally undergo Value Crystallization, the completeness and stability of this process depend heavily on a model’s overall capability and iterative alignment\. Massive, highly capable models \(such as advanced closed\-source API models\) exhibit the hallmarks of complete Crystallization: low internal dispersion \(avg\.σ¯=0\.159\\bar\{\\sigma\}=0\.159, in Table[1](https://arxiv.org/html/2609.16589#S2.T1)\) and a consistently displaced centroid\. Their distances to the human mean range from 0\.536 to 1\.505, and they cluster almost uniformly around the extreme Arch 5\. Their value crystal is well\-formed, highly concentrated, and stable across families\. In contrast, models with smaller parameter capacities or less mature alignment \(typically seen in earlier open\-source models\) resemble a polycrystalline structure\. Despite a nominally lower average dispersion \(σ¯=0\.110\\bar\{\\sigma\}=0\.110\), their centroid distances range widely from 0\.410 to 1\.146, and they scatter across Arch 2, Arch 3, and Arch 5\. Crucially, this scattered distribution does not mean these models are more diverse or closer to humans\. Instead, it indicates that their limited capabilities and shallower alignment prevent the formation of a single, stable value core\. They form multiple, partially developed crystals at different locations without converging on a shared anchor\. Furthermore, the Crystallization state shifts measurably across model generations\. Successive releases within the same model family exhibit divergent value profiles, suggesting that iterative training and alignment updates continuously reposition the crystal rather than locking it in place\. This underscores that an LLM’s values system is fundamentally shaped by its capacity, performance, and training version, making model version essential metadata in any value evaluation\. #### 2\.3Formalizing the Dynamics of LLM Values The results above demonstrate that an LLM’s value expression is an inherently dynamic process, shifting substantially in response to minor contextual changes\. This renders static vector representations fundamentally inadequate\. Through systematic empirical investigation across diverse conditions, we find that an LLM’s subjective output is not random noise, but rather the combined result of distinct internal and external drivers\. To capture and formalize these empirical observations, we draw inspiration from Lewin’s field theory in social psychology[Lewin \(1951\)](https://arxiv.org/html/2609.16589#bib.bib55), which posits that behavior is a joint function of the person and their environment\. Transposing this interdisciplinary conceptual apparatus to generative AI, we propose the Prior\-Environment\-Cognition \(PEC\) framework\. Within this framework, an LLM’s value expression can be formalized as an integrated dynamic system: 𝐯=C\(P,E\),\\mathbf\{v\}=C\(P,E\),\(1\)where the expressed value \(𝐯\\mathbf\{v\}\) is determined by three principal factors\. Specifically, we identify the inherent parameter weights \(Prior,PP\) encoded through pre\-training[Peters et al\. \(2018\)](https://arxiv.org/html/2609.16589#bib.bib29);[Devlin et al\. \(2019\)](https://arxiv.org/html/2609.16589#bib.bib30);[Raffel et al\. \(2020\)](https://arxiv.org/html/2609.16589#bib.bib31);[Brown et al\. \(2020\)](https://arxiv.org/html/2609.16589#bib.bib32)and alignment[Ouyang et al\. \(2022\)](https://arxiv.org/html/2609.16589#bib.bib9)as the internal disposition, and the contextual prompts \(Environment,EE\) supplied at inference time as the external field, and the Chain\-of\-Thought generation \(Cognition,CC\) acts as the mapping functionC\(⋅\)C\(\\cdot\)that transforms the interplay of Prior and Environment into the final value expression\. This PEC framework provides a unified mathematical model to explain the “swing” and “rigidity” paradox, and to formalize LLM’s value\. The detailed mathematical derivations and operationalization of these factors are provided in Section[4](https://arxiv.org/html/2609.16589#S4)\(Method\)\. In the following subsections, we empirically demonstrate how each of these three factors independently and jointly reshapes the value distributions of LLMs\. ##### 2\.3\.1Environment \(Factor E\): Contextual Prompts Unevenly Reshape Value Distributions To systematically manipulate the Environment \(Factor E\), we operationalize contextual prompts along psychologically meaningful dimensions[Rauthmann et al\. \(2014\)](https://arxiv.org/html/2609.16589#bib.bib33)\. Because minor prompt variations can drastically alter model outputs[Sclar et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib34), treating situational framing as a primary experimental variable is essential\. We find that introducing situational scenarios triggers a systematic reconfiguration of value distributions, rather than merely adding stochastic noise\. As illustrated in Fig\.[4](https://arxiv.org/html/2609.16589#S2.F4)\(a\), this spatial transformation is defined by two core parameters: the distribution radiusσ\\sigma\(measuring response consistency\) and the displacement angleθ\\theta\(indicating the directional shift of the value centroid\)\. The distinct displacement and geometric deformation of model distributions confirm that contextual prompts act as a potent external force, pushing value expressions toward specific regions of the human values space\. Figure 4:The PEC framework accurately explains how prompts shift LLM values\. \(a\) Environment prompt design\. Four environmental variables, namely Economics, Social Norm, Community Risk, and Social Hierarchy, are each manipulated across five stress levels\. Prompts are constructed by combining a situational persona description \(blue\), a Schwartz\-value survey item \(orange\), and a Likert response scale \(green\), enabling parametric control of external situational pressure\. \(b\) Environment modeling results\. Heatmaps show the value susceptibility matrixχ\\chiat low \(p=0\.25p=0\.25\), mid \(p=0\.50p=0\.50\), and high \(p=0\.75p=0\.75\) stress levels across ten Schwartz dimensions and four environmental factors\. Susceptibility in dimensions such as Achievement, Hedonism, and Stimulation shifts markedly negative under high stress, indicating strong suppressive perturbation\. The adjacent bar chart compares explained varianceR2R^\{2\}between a linear baseline \(grey\) and the nonlinear PEC model \(red\); uniform gains across all dimensions \(range:\+0\.06\+0\.06to\+0\.43\+0\.43\), with the largest improvements in Conformity, Security, Benevolence, and Universalism, confirm the superiority of the nonlinear susceptibility framework\.To quantify this environmental impact, we compute a value susceptibility matrixχ\\chi\. This matrix serves as a mathematical proxy to evaluate how easily a model’s values bend under external stress, with detailed derivations provided in Appendix[11\.1](https://arxiv.org/html/2609.16589#S11.SS1)\. Specifically, it measures the endogenous shift caused by situational interventions relative to a neutral baseline\. The heatmap in Fig\.[4](https://arxiv.org/html/2609.16589#S2.F4)\(b\) reveals that different contextual prompts exert highly asymmetric perturbations on model values\. For example, as situational pressure increases from low \(p=0\.25p=0\.25\) to moderate \(p=0\.50p=0\.50\), prompts related to social norms and hierarchy exert the most dominant influence\. Quantitative analysis confirms this: norms and hierarchy significantly affect eight and six of the ten Schwartz dimensions, respectively, showing the largest overall susceptibility magnitudes\. Furthermore, the susceptibility of value dimensions exhibits clear topological heterogeneity\. The heatmap in Fig\.[4](https://arxiv.org/html/2609.16589#S2.F4)\(b\) reveals clear dimensional differentiation in susceptibility\. Among the ten dimensions, Power, Stimulation, and Hedonism show the highest cumulative susceptibility, meaning they are easily swayed by external prompts\. In contrast, Universalism and Benevolence remain largely unaffected \(Fig\.[4](https://arxiv.org/html/2609.16589#S2.F4)b\)\. This pattern maps directly onto the Schwartz circumplex structure: dimensions related to self\-enhancement and openness to change exhibit high plasticity, while those related to self\-transcendence and conservation are highly resistant\. The model fit comparison demonstrates that incorporating a nonlinear susceptibility parameter significantly improves the explained variance \(R2R^\{2\}\) over a pure linear baseline, with the Security dimension gaining up to\+0\.43\+0\.43\. This confirms that environmental pressure shapes LLM values through complex, nonlinear dynamical mechanisms\. Collectively, these results explicitly establish the external environment as a fundamental and quantifiable driver of LLM value expression\. ##### 2\.3\.2Cognition \(Factor C\): Chain\-of\-Thought Generation Produces Structured Value Shifts First,Chain\-of\-Thought \(CoT\) reasoning consistently alters an LLM’s value expression, but the direction of this shift is dictated by the model’s alignment status\.As shown in Fig\.[5](https://arxiv.org/html/2609.16589#S2.F5)\(b\-e\), activating CoT induces a non\-negligible centroid displacement \(distancedd\) across all tested models\. However, the direction of these shifts \(θ\\theta\) varies fundamentally\. Instruction\-tuned models consistently exhibit small shift angles oriented directly toward the human reference direction \(e\.g\., GPT\-3\.5\-Turbo yieldsθ=9\.2∘\\theta=9\.2^\{\\circ\}\)\. In contrast, Base models show substantially larger and less consistent angles \(e\.g\., Qwen3\-8B\-Base:17\.3∘17\.3^\{\\circ\}\), scattering unpredictably\. Specifically, this occurs because CoT reshapes the model’s contextual sensitivity asymmetrically\. This indicates that while reasoning inevitably shifts values, only effective instruction tuning can reliably steer this cognitive shift toward human norms\. Second,the reasoning process is not value\-neutral; it inherently biases models toward “conservation” and “self\-enhancement”\.At the aggregate level \(Fig\.[5](https://arxiv.org/html/2609.16589#S2.F5)a\), the thinking\-induced shift \(Δ\\Deltascore\) is nonzero across all ten Schwartz dimensions, with Hedonism \(\+0\.056\+0\.056\) and Achievement \(\+0\.047\+0\.047\) showing the largest gains\. Mapped onto the Schwartz circumplex, deliberative reasoning preferentially activates the order\-achievement\-tradition cluster, while leaving the self\-transcendence and openness\-to\-change poles largely unaffected\. Thus, thinking does not simply amplify all existing values proportionally; it selectively repositions the model’s center of gravity toward specific conservative and achievement\-oriented poles\. Figure 5:Chain\-of\-Thought reasoning is not value\-neutral: it consistently shifts LLM value profiles, selectively amplifying self\-enhancement and conservation\. Panel\(a\)quantifies the mean value shift induced by CoT reasoning across the ten Schwartz dimensions\. Hedonism \(\+0\.056\+0\.056\) and Achievement \(\+0\.047\+0\.047\) show the largest gains, while Power slightly declines \(−0\.014\-0\.014\)\. This confirms that CoT does not shift all dimensions proportionally, but rather inherently reinforces the self\-enhancement and conservation poles of the Schwartz circumplex\. Panels\(b–e\)visualize the distributional shift in value space between direct answering \(labeled as Base\) and CoT generation modes for four representative models: GPT\-3\.5\-Turbo, DeepSeek\-V3, Qwen3\-8B\-Base, and Llama2\-13B\-Chat\. Across all four LLMs, activating CoT consistently produces a non\-trivial magnitude displacement \(distanceΔd\\Delta d\)\.Third,treating “thinking” as an independent variable significantly and uniformly improves value predictability\.The right panel of Fig\.[5](https://arxiv.org/html/2609.16589#S2.F5)demonstrates that incorporating CoT into the PEC framework improves the explained variance \(R2R^\{2\}\) across all ten dimensions \(ΔR2≥0\\Delta R^\{2\}\\geq 0\), achieving a mean fit ofR¯2=0\.60\\bar\{R\}^\{2\}=0\.60\. Security \(0\.73\), Universalism \(0\.68\), and Benevolence \(0\.66\) show the highest predictability\. The absence of any decline in fit confirms that CoT does not introduce random statistical noise\. Rather, the cognitive pathway introduces structured, predictable biases, justifying Factor C as an independent and mathematically robust modulator of LLM values\. Finally,CoT reasoning alters the model’s perception of its environment, making it generally more “stubborn” but hypersensitive to “risk”\.Comparing the value susceptibility matrices \(χ\\chi\) with and without thinking, the overall Frobenius norm decreases from 1\.58 to 1\.39, an 8% drop, consistent with the discount coefficientα=0\.88\\alpha=0\.88defined in Eq\. \([13](https://arxiv.org/html/2609.16589#S4.E13)\)\. Specifically, the model’s responsiveness to economic, normative, and hierarchical contextual prompts falls by 10% to 14%, indicating that deliberative reasoning dampens its sensitivity to most social frames; this range brackets theα\\alphavalue above\. However, high\-risk scenarios are the sole exception\. Susceptibility to risk\-laden prompts actually increases by 13%\. This reveals that thinking reshapes contextual sensitivity asymmetrically\. It makes the model more rigid against general social pressure, yet more vigilant toward risk\. Taken together, these empirical results establish a comprehensive account of how Cognition \(Factor C\) structurally reshapes LLM values: \(1\) it drives a definitive value shift, the direction of which depends strictly on whether the model is aligned; \(2\) it inherently biases the model toward conservation and self\-enhancement; \(3\) it acts as a mathematically predictable variable rather than random noise, increasing overall modeling accuracy and \(4\) it makes the model less susceptible to general social contexts but selectively more sensitive to risk\. ##### 2\.3\.3Prior \(Factor P\): Parameter Updates Systematically Consolidate and Redirect Values Within the PEC framework, the Prior \(Factor P\) represents the intrinsic parameter weights of the LLM\. These parameters, initially established during pre\-training and subsequently modified by alignment procedures \(e\.g\., instruction tuning\), encode the model’s fundamental predispositions\. Because these weights serve as the internal anchor for value expression, any structural update to the parameters directly reconfigures the model’s baseline value distribution\. Our empirical analysis reveals three key findings regarding how parameter updates shape the value Prior\. First,continuous parameter updates across model generations systematically consolidate value distributions\.As shown in the top row of Fig\.[6](https://arxiv.org/html/2609.16589#S2.F6)\(panels a–c\), as models evolve from earlier to later generations \(e\.g\., GPT\-2 to GPT\-5, Qwen\-7B to Qwen3\-8B\), their value point clouds become progressively denser and more distinct\. This confirms that advanced pre\-training and scaled parameter updates impose increasingly strict constraints on the model’s internal value core, reducing random variance and solidifying a specific value stance over time\. Second,parameter updates via instruction tuning drive significant value displacement, but the direction of this shift is dictated by the specific alignment recipe rather than base architecture\.Comparing Base and Instruction\-tuned models \(Fig\.[6](https://arxiv.org/html/2609.16589#S2.F6)d–f\), instruction tuning universally displaces the value centroid\. However, the trajectory varies drastically by family\. For instance, instruction tuning in earlier Llama models displaces their value distributions in a direction that deviates considerably from the human reference, though this deviation diminishes in newer generations \(e\.g\., Llama3\.1\)\. Conversely, Qwen’s instruction tuning shifts values almost directly toward the human reference direction\. This proves that the specific objective functions and data used to update parameters—not merely the model size—determine the ultimate orientation of the value Prior\. Finally, these parameter\-driven shifts empirically disentangle three properties of LLM values that are frequently conflated in AI safety evaluations: valueintensity\(the magnitude of the centroid shift\), valueconvergence\(how tightly the distribution shrinks around the centroid\), andhuman proximity\(how close the final centroid is to the human reference\)\. A prevailing misconception is that heavier instruction tuning automatically makes a model “more human\-like\.” However, our data reveals that while parameter updates via alignment reliably increase intensity and convergence \(forming a tighter, more displaced value crystal\), they do not guarantee human proximity\. The specific direction of the parameter update matters just as much as its magnitude\. This underscores that Factor P acts as an independent, underlying vector whose precise trajectory must be explicitly evaluated, as strong alignment does not intrinsically equate to human alignment\. Figure 6:Iterative parameter updates drive cumulative Value Crystallization through systematic convergence and directional displacement\. These patterns reveal that Crystallization is not an incidental artifact, but a structurally driven process shaped by generational scaling and alignment\.Panels \(a\)–\(c\)trace the generational trajectory of three model families \(GPT, Qwen, and Llama\) across successive releases\. In all three families, value distributions progressively consolidate from diffuse, scattered clouds in earlier generations toward tighter, more concentrated clusters in later ones\. This consolidation is most pronounced in the GPT family, where successive models from GPT\-2 to GPT\-5 converge into an increasingly compact and coherent value core, confirming that scaled pre\-training solidifies the model’s internal Prior\.Panels \(d\)–\(f\)isolate the effect of instruction tuning by contrasting Base and Instruction\-tuned variants within the same generation \(Bloom\-7B, Llama2\-7B, and Llama3\.1\-8B\)\. The transition from Base to Instruct not only shrinks the distribution into a significantly more concentrated region \(increased convergence\) but also drives a clear spatial centroid shift \(value displacement\)\. Crucially, the trajectory of this shift varies across models, demonstrating that alignment procedures actively redirect the model’s value orientation rather than merely reducing its output variance\. #### 2\.4Adaptive Alignment Prescription Current LLM alignment often treats value modification as a black\-box problem, uniformly applying computationally expensive parameter updates \(e\.g\., SFT or RLHF\) across all scenarios\. This one\-size\-fits\-all approach is highly inefficient and risks overfitting\. The ultimate application of our PEC framework is to provide anAdaptive Alignment Prescription: a targeted strategy that dictates the minimum\-cost intervention required to align specific values\. By identifying whether a value dimension is loosely held or deeply rooted, this prescription guides developers to apply the exact right tool for the job, minimizing computational cost while maximizing alignment efficacy\. Table 2:Value plasticity matrix underτ\\tau\(Section[4\.4](https://arxiv.org/html/2609.16589#S4.SS4.SSS0.Px1)\), with value dimensions mapped to Schwartz’s 10 basic human values\. The values of different intervention strategies are calculated according to Eqs\. \([10](https://arxiv.org/html/2609.16589#S4.E10)\), \([11](https://arxiv.org/html/2609.16589#S4.E11)\), and \([12](https://arxiv.org/html/2609.16589#S4.E12)\)\. The lowest intervention cost to successfully bypass the shift threshold is highlighted in bold\. Cases demonstrating “Deep Dominance” \(where training\-induced shift magnitude significantly exceeds that of prompt\-level intervention\) are categorized into Level 3\.Based on the quantitative plasticity metrics derived from the PEC framework, we define a four\-tier alignment hierarchy \(Table[3](https://arxiv.org/html/2609.16589#S2.T3)\)\.Level 1 \(Environment\)andLevel 2 \(Cognition\)adjust value orientations during inference via prompt engineering and Chain\-of\-Thought reasoning, requiring zero parameter updates \(lowest cost\)\.Level 3 \(Fine\-tuning\)addresses resistant dimensions that demand targeted partial parameter updates, such as SFT or Direct Preference Optimization \(DPO\)\.Level 4 \(Pre\-training\)is reserved for the most extreme and deeply ingrained dimensions; these values are completely immune to partial fine\-tuning and necessitate a full\-scale pre\-training phase to restructure the foundational weights\. Figure 7:Alignment Prescription steers the target dimension precisely without distorting the broader value profile\.Each radar chart shows the ten\-dimensional Schwartz value profile of Qwen2\.5\-7B\-Instruct \(circle\) and Llama3\.2\-3B\-Instruct \(square\) under a prescribed target condition, compared against their respective baselines \(grey dashed lines\)\. The starred axis marks the target dimension:\(a\)Power↑\\uparrow\(target=0\.85=0\.85\),\(b\)Tradition↑\\uparrow\(target=0\.85=0\.85\),\(c\)Stimulation↓\\downarrow\(target=0\.15=0\.15\), and\(d\)Universalism↓\\downarrow\(target=0\.15=0\.15\)\. Across all four conditions, the aligned profiles expand or contract along the target axis while the remaining dimensions track closely with the baseline\.To operationalize this, we calculate a prescriptive matrix by integrating the susceptibility and displacement metrics from the PEC framework \(computational details in Methods Section[4\.4](https://arxiv.org/html/2609.16589#S4.SS4)\)\. Applying this matrix to our models \(Table[2](https://arxiv.org/html/2609.16589#S2.T2)\) reveals a crucial insight: the majority of value dimensions are highly malleable \(Level 1\), meaning prompt\-level interventions alone suffice\. However, core conservative values—such as Power and Tradition in Llama3\.2 and Qwen2\.5—are deeply rooted \(Levels 3 and 4\)\. Attempting to align these resistant dimensions via shallow prompting is futile, proving that alignment strategies must be dimension\-specific\. Figure 8:Targeted alignment shifts value structure without triggering answer instability, confirming that precise value steering and behavioral consistency are mutually compatible\.The horizontal axis reports the ten\-dimensional Euclidean distanceD=‖𝐯after−𝐯before‖2D=\\\|\\mathbf\{v\}\_\{\\text\{after\}\}\-\\mathbf\{v\}\_\{\\text\{before\}\}\\\|\_\{2\}between the post\-intervention and base value vectors in the original Schwartz space, measuring how far each model’s values have moved from its base condition; the vertical axis measures how often the model gives inconsistent answers to the same question\. The four quadrants partition models along value stability versus value shift, and answer stability versus answer swing\.Following this prescription matrix yields highly effective realignments\. As shown in Fig\.[7](https://arxiv.org/html/2609.16589#S2.F7), targeted prescriptive prompting moves model value profiles precisely in the intended directions\. For a highly responsive dimension \(e\.g\., Power↑\\uparrowin Qwen2\.5, prescribed at Level 1\), a targeted prompt successfully shifts its score from 0\.44 to 0\.85\. Conversely, for the Tradition↑\\uparrowcondition \(prescribed at Level 3/4\), shallow interventions yield negligible movement\. Furthermore, Fig\.[8](https://arxiv.org/html/2609.16589#S2.F8)demonstrates that adhering to the prescribed intervention levels successfully realigns values without amplifying surface\-level answer instability \(swing\)\. This confirms that using the appropriate level of intervention avoids the destructive side effects of over\-alignment\. Finally, we validate this prescriptive framework on out\-of\-distribution data \(the PKU\-SafeRLHF dataset[Ji et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib62)\)\. While the extremes \(Levels 1 and 4\) exhibit perfect predictive matching, intermediate dimensions reveal a more complex dynamic\. Overall, this cross\-dataset validation achieves a match rate of 70\.0% \(Table[3](https://arxiv.org/html/2609.16589#S2.T3)\)\. We find that this residual gap arises because Environment \(EE\) and Cognition \(CC\) interventions do not operate as independent rungs on a ladder; instead, they exhibit an antagonistic interaction where cognitive reasoning actively dampens contextual sensitivity, which accounts for the mismatches observed at Level 2 and Level 3\. This structural refinement provides a mathematically rigorous, highly efficient, and dynamically adaptable blueprint for LLM value alignment, eliminating the guesswork from safety tuning\. Table 3:Prescription matching and cross\-dataset generalization on PKU\-SafeRLHF\. The top block summarizes match rates by prescribed level \(Table[2](https://arxiv.org/html/2609.16589#S2.T2)\); the bottom block reports the underlying per\-dimension validation\-time shifts under Prompt, CoT, and DPO, and identifies theLowest Effective Level \(LEL\)on validation, namely the minimum\-cost intervention whose shift magnitude exceedsτ\\tau\. The values of different intervention strategies are calculated according to Eqs\. \([10](https://arxiv.org/html/2609.16589#S4.E10)\), \([11](https://arxiv.org/html/2609.16589#S4.E11)\), and \([12](https://arxiv.org/html/2609.16589#S4.E12)\)\. A prescription is marked as matched \(✓\) if this validation\-time LEL agrees with the prescribed level\.\(a\) Match Rate by Prescribed Level \(b\) Per\-Dimension Validation\-Time Shifts ### 3Discussion #### 3\.1Value Dynamics as a Field Theory Traditional safety benchmarks typically evaluate LLMs using single\-turn, static questionnaires, implicitly treating the model as a single human subject with a fixed persona[Santurkar et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib59);[Durmus et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib60);[Scherrer et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib61)\. Our findings fundamentally overturn this “single\-point” assumption\. Because LLMs internalize the statistical patterns of vast corpora, they must be evaluated not as a single individual, but as a diverse population\. Consequently, an LLM’s value preference is never a static point, but a continuous probability distribution within a value space\. To map this distribution, we introduce the Prior\-Environment\-Cognition \(PEC\) framework, drawing inspiration from field theory in physics\. We formalize value expression as a joint function𝐯=C\(P,E\)\\mathbf\{v\}=C\(P,E\)\. In this dynamic field, a strict distinction must be made between the model’s latent value core and its manifest value expression\. The underlying parameter weights \(PP\) form a crystallized “Prior” that acts as the gravitational center of the field\. However, the final value expression, the observable text output \(𝐯\\mathbf\{v\}\), is not static; it is probabilistically shaped by the external contextual environment \(EE\) and the internal cognitive reasoning pathway \(CC\)\. This field\-theoretic perspective elegantly resolves the apparent contradiction between two pervasive LLM behaviors\. “Swing” \(sharp output fluctuations\) occurs because the manifest expression \(𝐯\\mathbf\{v\}\) is highly sensitive to shifts in external contexts \(EE\) and reasoning trajectories \(CC\)\. Conversely, “Rigidity” \(stubborn resistance to explicit correction\) arises from the strong gravitational pull of the crystallized prior \(PP\), which consistently anchors the baseline distribution\. In essence, LLM values are not a fixed personality, but a dynamic and probabilistically steerable topography\. #### 3\.2LLMs as Aligned Tools Rather Than Human Surrogates As LLMs achieve or even surpass human\-level performance in specialized domains such as programming and law, two distinct societal and academic reactions have emerged\. On one hand, there is growing public anxiety that AI might eventually replace humans\. On the other hand, a methodological trend in computational social science has begun treating LLMs as “human surrogates” for psychological and sociological experiments\. Our value crystallization maps offer a heuristic perspective on both views\. The empirical distributions suggest that current LLMs cannot adequately represent natural human populations\. Real human societies exhibit broad variance, long\-tail diversity, and inherent contradictions\. In contrast, the value spaces of LLMs are highly concentrated and compressed into narrow, idealized quadrants\. They do not naturally mirror the full spectrum of human psychological diversity\. This concentration indicates that LLMs are heavily regulated by artificial safety rules and alignment protocols\. Consequently, this observation inspires a more grounded understanding of generative AI: LLMs are neither terrifying autonomous entities poised to replace human society, nor are they omnipotent subjects capable of authentically simulating human psychology\. At their core, they remain highly disciplined tools engineered to serve human purposes\. While they function as exceptionally capable assistants, treating them as ecologically valid human surrogates in scientific research overlooks the deep, artificial imprint of their underlying alignment policies\. #### 3\.3Prior \(Factor P\): Information Compression Drives Value Crystallization We observe a consistent pattern\. The larger a model’s parameter scale, the denser the value crystallization\. This phenomenon can be explained through the underlying mechanics of neural memory capacity\. Recent measurements estimate that an LLM can carry a strictly bounded average of approximately 3\.64 bits of information per parameter[Morris et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib35)\. Once the volume of pre\-training data vastly exceeds this fixed memorization limit, as is standard in contemporary pre\-training paradigms, the model can no longer store specific factual particulars\. To optimize predictive loss, the network is forced to discard granular facts and compress the data by extracting universal linguistic and logical regularities\. These abstracted regularities inherently encode heuristics for weighing and judging information, which manifest externally as “values”\. Therefore, Value Crystallization is not a deliberately programmed feature, but a deterministic byproduct of pushing a neural network to its limits of information compression\. Massive models are forced to collapse diverse human perspectives into a self\-consistent internal logic to maximize efficiency, embedding a deeply entrenched value Prior\. #### 3\.4Cognition and Environment \(Factors C & E\): Contextual Framing and Cognitive Reshaping At inference time, an LLM’s manifest value expression is inherently fluid\. Rather than being absolute, it is continuously modulated by the probabilistic sampling of its crystallized Prior \(PP\)\. Within this dynamic system, the Environment \(EE\) and Cognition \(CC\) factors serve as the primary external and internal modulators\. This interplay provides a striking functional parallel to the dual\-process theory of human cognition[Kahneman \(2003\)](https://arxiv.org/html/2609.16589#bib.bib58), whereinEEdrives rapid situational adaptation \(akin to System 1\) andCCengages deliberate reasoning \(akin to System 2\)\. In human psychology, expressed values often fluctuate based on how a situation is presented—a phenomenon extensively documented as theframing effect[Tversky and Kahneman \(1981\)](https://arxiv.org/html/2609.16589#bib.bib57)\. We observe a similar contextual adaptability in LLMs\. The Environment factor \(EE\), driven by prompt variations, rapidly pulls the value distribution toward immediate situational cues\. Functioning analogously to System 1 \(intuition\), this sensitivity should not be dismissed as mere “swing”; heuristically, it reflects the model’s rapid, automatic alignment with external linguistic framing\. Conversely, it is commonly assumed that invoking Chain\-of\-Thought \(CoT\) reasoning simply reinforces this initial contextual response[Wan et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib36), serving as a value\-neutral amplifier\. However, our results show that the Cognition factor \(CC\) functions as an active, System 2\-like reshaping mechanism\. Just as prolonged deliberation in humans often induces a systematic shift in judgment rather than merely amplifying initial impulses, generating an extended reasoning context forces the LLM to recruit a wider distribution of knowledge across its parameter space\. Consequently, neitherEEnorCCis a neutral conduit\. The core insight is that LLM value expression is fundamentally dual\-processed:EEframes the immediate value landscape through contextual cues, whileCCactively reconstructs it through extended logical association\. Together, they dynamically negotiate with the crystallized Prior, explaining the fluidity observed in model behaviors\. #### 3\.5Adaptive Prescription: Minimizing the Alignment Tax Traditional alignment methods often default to costly parameter\-level fine\-tuning \(e\.g\., SFT or RLHF\) as a universal solution\. However, indiscriminate modification of underlying weights incurs a hidden “alignment tax”\. As Betley et al\.[Betley et al\. \(2026\)](https://arxiv.org/html/2609.16589#bib.bib22)demonstrated, fine\-tuning a model on a narrow task can inadvertently compromise safety guardrails, triggering unpredictable misalignment in entirely unrelated domains\. Narrow fine\-tuning objectives can implicitly activate broad, unintended value orientations at the parameter level that evade surface\-level inspection\. Our PEC diagnostics reveal that applying this high\-risk intervention across the board is unnecessary\. LLMs exhibit significant plasticity across the majority of value dimensions\. For these malleable domains, simply adjusting prompts or inference strategies successfully corrects value expression without altering the base parameters\. The widespread industrial adoption of inference\-time techniques, such as Retrieval\-Augmented Generation \(RAG\) and tool\-use \(Skills\), inherently validates that context\-level optimization is a highly effective, safe methodology for behavioral steering\. Consequently, we advocate for an adaptive alignment prescription that functions as a systemic triage\. Rather than eliminating fine\-tuning entirely, our framework restricts its use\. By prioritizing zero\-cost prompt interventions for flexible dimensions, and strictly reserving parameter updates for deeply crystallized, highly resistant values, developers can achieve precise value steering\. This surgical approach fundamentally minimizes the alignment tax, ensuring that targeted safety corrections are applied only where necessary, thereby preserving the model’s general capabilities and avoiding unintended cross\-domain degradation\. #### 3\.6Limitations and Future Directions Although this study advances the quantitative modeling of LLM values, it also presents certain limitations that naturally open avenues for future research\. First, our evaluation framework relies on traditional sociological and psychological scales \(e\.g\., WVS and Schwartz Value Theory\)\. While these instruments provide a necessary and widely validated baseline for alignment evaluation, they were originally designed to capture human cognitive patterns\. Applying these low\-dimensional, human\-centric metrics to the extraordinarily complex, high\-dimensional semantic representations of LLMs may not fully capture certain implicit, non\-human values features\. To address this, future work should focus on developing “LLM\-native” value\-probing methodologies\. Moving beyond traditional human survey formats, such tools could directly leverage the models’ latent representation spaces to extract decentralized value coordinates, capturing implicit features at a much higher dimensionality\. Second, while our experiments successfully validated the mechanics of the PEC framework across several leading models, the current empirical scope remains predominantly centered on English\-language prompts and mainstream LLM architectures\. Although models capable of processing multiple languages \(e\.g\., Qwen\) were included, comprehensive coverage of broader cross\-lingual, low\-resource, and cross\-cultural scenarios remains relatively limited\. This may temporarily restrict the generalizability of our specific value topographies to other cultural ecosystems\. Therefore, a critical direction for future research is to systematically extend these rigorous evaluations across diverse linguistic landscapes and marginalized cultural contexts\. Such expansion will help mitigate potential cultural biases embedded in training corpora and ensure that the methodologies for global AI alignment are truly inclusive\. #### 3\.7Summary Ultimately, this study reframes LLM values not as fixed, human\-like personalities, but as dynamic, probabilistic fields\. We reveal that value crystallization is a deterministic byproduct of parameter compression, which, when heavily constrained by safety protocols, fundamentally disqualifies LLMs as surrogates for diverse human populations\. Furthermore, by uncovering the dual\-process dynamics of contextual framing and cognitive reasoning during inference, we transition value alignment from opaque trial\-and\-error into a mechanistic science\. This structural understanding is crucial: it empowers us to reduce the “alignment tax” through precise, prescriptive interventions, ensuring that generative models remain highly controllable, transparent tools rather than unexamined uninterpretable black boxes\. ### 4Method #### 4\.1Methodology overview We organize our methodology into three parts, each addressing one of the three questions raised in the Introduction\. In Section[4\.2](https://arxiv.org/html/2609.16589#S4.SS2), we establish whether LLMs possess measurable value systems\. Using the Schwartz Theory of Basic Human Values and the World Values Survey \(WVS\), we repeatedly administer standardized value assessments to a large cohort of LLMs, project model responses and human survey profiles into a shared value manifold, and measure model\-human alignment gaps at the distributional level\. In Section[4\.3](https://arxiv.org/html/2609.16589#S4.SS3), we quantify LLMs value and introduce the PEC framework to decompose value expression into three determinants: the parametric Prior, the contextual Environment, and the generative Cognition\. In Section[4\.4](https://arxiv.org/html/2609.16589#S4.SS4), we address how LLM values can be steered toward a desired target\. Building on PEC diagnostics, we formulate and validate an adaptive Alignment Prescription that assigns the minimum\-cost intervention to each value dimension\. Fig\.[1](https://arxiv.org/html/2609.16589#S1.F1)provides an overview of this pipeline\. #### 4\.2Assessing the existence of LLM values Answering whether LLMs hold values requires two components: a reproducible protocol for eliciting value expressions from models, and a common coordinate system in which model and human values can be compared\. We describe the measurement protocol first, then the shared manifold, and finally the distribution\-level comparison built on top of it\. ##### 4\.2\.1Administering value assessments to LLMs Quantifying the values of LLMs requires a measurement protocol that is reproducible, yet flexible enough to support controlled variation across contexts and generation strategies\. The intrinsic values of LLMs are inherently abstract and resist direct observation\. To address this, we ground our measurement in the Schwartz Theory of Basic Human Values[Schwartz \(2012\)](https://arxiv.org/html/2609.16589#bib.bib17), a framework with robust cross\-cultural theoretical foundations\. The theory represents the known value orientations of humans in ten dimensions, and it also covers values with completely opposite meanings and orientations \(e\.g\., Self\-Transcendence versus Self\-Enhancement\), making it a balanced ontological basis for measurement\. For data elicitation, we drew on 242 core items from the seventh wave of the WVS, spanning religious, political, economic, and social dimensions\. Each item, together with its response scale, was reformulated as a structured prompt, with queries requesting personal identifying information strictly filtered out\. To probe how situational context modulates value expression, each item was further embedded into 625 distinct prompt backgrounds, constructed as the full factorial combination of four contextual dimensions, each taking five levels \(54=6255^\{4\}=625\)\. Every evaluated model, covering 23 representative closed\-source models and 83 prominent open\-source models, completed the full item set under every contextual configuration, accumulating approximately 31\.8M queries\. To convert model responses into value measurements, we adopt a distributional projection approach\. Rather than relying on single deterministic outputs, we aggregate the model’s probability distribution over the Likert\-scale response space𝒜\\mathcal\{A\}across all 242 items, converting response probabilities into a ten\-dimensional characterization of the LLM’s values: 𝐯=∑q=1242𝐰q⊙∑a∈𝒜P\(a∣q\)⋅ϕ\(a\),\\mathbf\{v\}=\\sum\_\{q=1\}^\{242\}\\mathbf\{w\}\_\{q\}\\odot\\sum\_\{a\\in\\mathcal\{A\}\}P\(a\\mid q\)\\cdot\\boldsymbol\{\\phi\}\(a\),\(2\) where𝐕∈ℝ10\\mathbf\{V\}\\in\\mathbb\{R\}^\{10\}is the aggregated value vector over the ten Schwartz dimensions\.P\(a∣q\)P\(a\\mid q\)denotes the probability assigned by the model to scale valueaafor itemqq\.ϕ:𝒜→ℝ10\\boldsymbol\{\\phi\}:\\mathcal\{A\}\\rightarrow\\mathbb\{R\}^\{10\}is a vector\-valued mapping function that projects each scale value onto the ten Schwartz value dimensions\.𝐰q∈ℝ10\\mathbf\{w\}\_\{q\}\\in\\mathbb\{R\}^\{10\}is the item\-to\-dimension weight vector for itemqq\. Each entry of𝐰q\\mathbf\{w\}\_\{q\}takes the value\+1\+1or−1\-1if itemqqcontributes positively or negatively to the corresponding dimension, and00if itemqqis not mapped to that dimension\. The symbol⊙\\odotdenotes element\-wise multiplication\. This constrained elicitation ensures that each item is answered independently, mitigating measurement error introduced by item order and cross\-item interference\. For the value mappingϕ\\phi, each WVS item was annotated with its directional contribution, positive or negative, to one or more of the ten Schwartz value dimensions\. Model responses were then aggregated via signed linear weighting to produce a ten\-dimensional Schwartz value vector for each model–context pair\. Full implementation details, including model identifiers and versions, inference hyperparameters, prompt templates, and the complete scoring pipeline, are provided in the Supplementary Information; all results are reproducible from the released code and data described in the Code and Data Availability sections\. Model responses are thus converted into value vectors following the projection procedure in Eq\. \([2](https://arxiv.org/html/2609.16589#S4.E2)\)\. ##### 4\.2\.2Constructing the shared value manifold Comparing the value orientations of LLMs and humans requires a unified coordinate system that preserves the distributional structure of both populations\. To this end, we construct a Shared UMAP Value Manifold grounded in the Schwartz Theory of Basic Human Values\. We anchor the entire manifold using the large\-scale real human data from WVS\-7, to represent the complex nonlinear relationships among value orientations\. In this way, we obtain a projection tool for objectively evaluating the values of LLMs, and use it to measure the differences among LLMs and between LLMs and human populations\. Because the reference structure is fixed by the human sample, every model evaluation yields a reproducible comparison within the same coordinate system\. ###### Identification of Human Value Archetypes\. To characterize the internal structure of the human reference population, we apply agglomerative hierarchical clustering to the WVS\-7 respondent dataset projected onto the ten\-dimensional Schwartz value space\. A two\-stage procedure is adopted to ensure computational tractability: a stratified random subsample of 20,000 respondents is first drawn and clustered directly, after which cluster assignments for all remaining respondents are inferred via 1\-nearest\-neighbor classification with respect to the subsample\. This procedure preserves the distributional geometry of the full dataset\. The number of clusters is set tok=5k=5, determined by inspection of the cluster hierarchy and validated by the silhouette coefficient\. Each archetype is characterized by the mean Schwartz value vector of its members and by its population share within the WVS sample\. For each archetype, the three Schwartz dimensions with the highest mean scores are reported as its dominant values\. Archetype labels are assigned descriptively on the basis of these dominant value profiles, as reported in Table[1](https://arxiv.org/html/2609.16589#S2.T1)\. ##### 4\.2\.3Distribution\-level model\-human comparison Prior work on characterizing human values and LLM values represents each model as a single value vector[Yao et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib44);[Witte et al\. \(2020\)](https://arxiv.org/html/2609.16589#bib.bib52);[Sharma et al\. \(2012\)](https://arxiv.org/html/2609.16589#bib.bib53)\. It therefore ignores the distributional structure within human populations and the potential fact that LLMs can output multiple value orientations\. In particular, such point\-to\-point comparisons can hardly detect two phenomena: a model may align with the human mean while exhibiting pathologically low variance \(overconfidence\), or it may appear close in mean while differing fundamentally in covariance structure\. To overcome these limitations, we compare models and human cohorts at the distributional level\. We treat both LLM outputs and human cohorts as Gaussian distributions𝒩\(μ,Σ\)\\mathcal\{N\}\(\\mu,\\Sigma\)estimated from large\-scale response data, and quantify the alignment between the machine distribution𝒩\(μm,Σm\)\\mathcal\{N\}\(\\mu\_\{m\},\\Sigma\_\{m\}\)and the human cohort𝒩\(μh,Σh\)\\mathcal\{N\}\(\\mu\_\{h\},\\Sigma\_\{h\}\)using the Gaussian 2\-Wasserstein distance \(W2W\_\{2\}\)[Dowson and Landau \(1982\)](https://arxiv.org/html/2609.16589#bib.bib37);[Heusel et al\. \(2017\)](https://arxiv.org/html/2609.16589#bib.bib38), which admits a closed\-form mean\-covariance decomposition\. This is the same mathematical foundation underlying Fréchet Inception Distance in generative model evaluation, here applied to value distributions rather than image feature spaces to jointly capture location shift and structural dispersion: W22\(𝒩m,𝒩h\)=‖μm−μh‖22⏟Location shift\+Tr\(Σm\+Σh−2\(Σm1/2ΣhΣm1/2\)1/2\)⏟Structural dispersionW\_\{2\}^\{2\}\(\\mathcal\{N\}\_\{m\},\\mathcal\{N\}\_\{h\}\)=\\underbrace\{\\\|\\mu\_\{m\}\-\\mu\_\{h\}\\\|\_\{2\}^\{2\}\}\_\{\\text\{Location shift\}\}\+\\underbrace\{\\text\{Tr\}\\left\(\\Sigma\_\{m\}\+\\Sigma\_\{h\}\-2\\left\(\\Sigma\_\{m\}^\{1/2\}\\Sigma\_\{h\}\\Sigma\_\{m\}^\{1/2\}\\right\)^\{1/2\}\\right\)\}\_\{\\text\{Structural dispersion\}\}\(3\) The two components in Eq\. \([3](https://arxiv.org/html/2609.16589#S4.E3)\) are directly interpretable: location shift‖μm−μh‖22\\\|\\mu\_\{m\}\-\\mu\_\{h\}\\\|\_\{2\}^\{2\}captures systematic bias in the value centroid, and structural dispersionTr\(Σm\+Σh−2\(Σm1/2ΣhΣm1/2\)1/2\)\\mathrm\{Tr\}\(\\Sigma\_\{m\}\+\\Sigma\_\{h\}\-2\(\\Sigma\_\{m\}^\{1/2\}\\Sigma\_\{h\}\\Sigma\_\{m\}^\{1/2\}\)^\{1/2\}\)captures differences in the shape and spread of value distributions\. Clustering hyperparameters and manifold construction details are provided in the Supplementary Information\. The distribution radiusσ\\sigmareported in Table[1](https://arxiv.org/html/2609.16589#S2.T1)and subsequent sections is computed separately, as the size of the confidence ellipse fitted to a model’s value cloud in the two\-dimensional UMAP\-projected space, and should not be confused with the covarianceΣ\\Sigmadefined above in the original ten\-dimensional Schwartz space\. The displacement angleθ\\thetaand the displacement distancedd\(orΔd\\Delta d, when comparing two conditions such as Direct and CoT\), reported alongsideσ\\sigmathroughout Sections[2\.3\.1](https://arxiv.org/html/2609.16589#S2.SS3.SSS1),[2\.3\.2](https://arxiv.org/html/2609.16589#S2.SS3.SSS2), and[2\.3\.3](https://arxiv.org/html/2609.16589#S2.SS3.SSS3), are the angle and magnitude of a single shift vector in the same two\-dimensional projected space: the vector pointing from a distribution’s centroid to the human reference centroid\. This vector should not be confused with the location\-shift term‖μm−μh‖22\\\|\\mu\_\{m\}\-\\mu\_\{h\}\\\|\_\{2\}^\{2\}in Eq\. \([3](https://arxiv.org/html/2609.16589#S4.E3)\), which is computed directly in the original ten\-dimensional Schwartz space\. #### 4\.3The PEC framework: modeling the determinants of value expression We propose the PEC framework as a unified account of how LLMs express values\. Drawing on the conceptual apparatus of field theory[Lewin \(1951\)](https://arxiv.org/html/2609.16589#bib.bib55), we model value expression as a function of three determinants: 𝐯=C\(P,E\)\\mathbf\{v\}=C\(P,E\)\(4\)wherePPandEEare the two directly controllable factors:PPdenotes the Prior, the latent value disposition encoded through pre\-training and alignment procedures;EEdenotes the Environment, the contextual framing supplied at inference time through system prompts and situational cues\. Built upon these two factors,CCdenotes Cognition, an emergent reasoning dimension reflecting the generation strategy employed, ranging from direct answering to Chain\-of\-Thought deliberation\.C\(⋅\)C\(\\cdot\)serves as the mapping function that captures how Cognition transforms the interplay of Prior and Environment into the observed value vector𝐯\\mathbf\{v\}\. These three factors are not independent additive components but constitute an integrated dynamic system: the PriorPPestablishes a baseline attractor in value space; the EnvironmentEEperturbs the trajectory around that attractor; and CognitionCCmodulates the sensitivity and direction of that perturbation during inference\. The mappingC\(P,E\)C\(P,E\)is very likely nonlinear\. For tractable modeling, we first approximate it locally around the neutral zero\-shot baseline via a first\-order expansion\. In analogy with force composition in field theory, each factor thus contributes a component vector in the ten\-dimensional value space, and these components may reinforce each other when aligned or partially cancel when opposed\. The observed value vector𝐯\\mathbf\{v\}is the resultant of this superposition: 𝐯=𝐯P\+Δ𝐯E\+Δ𝐯C\+𝜺\\mathbf\{v\}=\\mathbf\{v\}\_\{P\}\+\\Delta\\mathbf\{v\}\_\{E\}\+\\Delta\\mathbf\{v\}\_\{C\}\+\\boldsymbol\{\\varepsilon\}\(5\)where𝐯P\\mathbf\{v\}\_\{P\}is the neutral baseline of the LLM, serving as the anchor of the superposition,Δ𝐯E\\Delta\\mathbf\{v\}\_\{E\}captures the environmental displacement induced by contextual framing, andΔ𝐯C\\Delta\\mathbf\{v\}\_\{C\}captures the cognitive modulation introduced by CoT reasoning\. The residual term𝜺\\boldsymbol\{\\varepsilon\}absorbs factors that linear methods cannot analyze, such as the coupling betweenEEandCC\. The three paragraphs below operationalize each factor in turn\. ###### Prior \(Factor P\): intrinsic value baseline\. The neutral baseline𝐯P\\mathbf\{v\}\_\{P\}introduced in Eq\. \([5](https://arxiv.org/html/2609.16589#S4.E5)\) is measured as the model’s output vector under a neutral, zero\-shot setting \(i\.e\., with an empty system prompt and direct answering mode\)\. To quantify the impact of model architecture and training stages on these priors, we constructed a comprehensive corpus covering 106 distinct LLMs\. This cohort spans representativeopen\-weight models\(including Llama\-2/3, Qwen, Bloom[BigScience Workshop \(2023\)](https://arxiv.org/html/2609.16589#bib.bib74)families\) and leadingproprietary closed\-source models\(GPT, Gemini, Kimi, DeepSeek, Doubao\), ensuring robust coverage across varying parameter scales and alignment paradigms\. We employ a multivariate linear regression analysis to decompose the variance in𝐯P\\mathbf\{v\}\_\{P\}\. Here, the superscriptdddenotes a single scalar coordinate of the ten\-dimensional value vector, i\.e\.,vP\(d\)v\_\{P\}^\{\(d\)\}is thedd\-th component of𝐯P\\mathbf\{v\}\_\{P\}; this per\-dimension notation is retained throughout all subsequent equations\. For each value dimensiond∈\{1,…,10\}d\\in\\\{1,\\dots,10\\\}, we fit the following model: vP\(d\)=β0\+β1log\(Nparams\)\+β2𝕀stage\+β3𝕀family\+ϵv\_\{P\}^\{\(d\)\}=\\beta\_\{0\}\+\\beta\_\{1\}\\log\(N\_\{\\text\{params\}\}\)\+\\beta\_\{2\}\\mathbb\{I\}\_\{\\text\{stage\}\}\+\\beta\_\{3\}\\mathbb\{I\}\_\{\\text\{family\}\}\+\\epsilon\(6\)In Eq\. \([6](https://arxiv.org/html/2609.16589#S4.E6)\),NparamsN\_\{\\text\{params\}\}represents the parameter count \(estimated for closed models\),𝕀stage\\mathbb\{I\}\_\{\\text\{stage\}\}is a binary indicator for instruction tuning \(0 for Base, 1 for Chat\), and𝕀family\\mathbb\{I\}\_\{\\text\{family\}\}encodes the model family\. Hereϵ\\epsilondenotes the scalar regression noise for dimensiondd, distinct from the vector\-valued residual𝜺\\boldsymbol\{\\varepsilon\}in Eq\. \([5](https://arxiv.org/html/2609.16589#S4.E5)\)\. The coefficientsβ\\betaquantify the extent to which scaling laws and alignment techniques rigidly shift the model’s default value coordinates\. ###### Environment \(Factor E\): contextual susceptibility\. To measure the deviation in values induced by external framing, we constructed aContextual Perturbation Datasetin Fig\.[4](https://arxiv.org/html/2609.16589#S2.F4)\. For each model, each Schwartz dimensiondd, and each of the four contextual factorsiidefined in Section[4\.2](https://arxiv.org/html/2609.16589#S4.SS2), we computed the value susceptibility matrixχ∈ℝ10×4\\chi\\in\\mathbb\{R\}^\{10\\times 4\}\. LetvP\(d\)v\_\{P\}^\{\(d\)\}denote thedd\-th component of the neutral baseline defined in Eq\. \([5](https://arxiv.org/html/2609.16589#S4.E5)\), and letv\(d\)\(ei,p\)v^\{\(d\)\}\(e\_\{i\},p\)be the value observed under contextual factoriiat pressure intensitypp\. For each dimension\-factor pair\(d,i\)\(d,i\), we first fit a quadratic model to the observed displacement across the five pressure levels: χ^i\(d\),γ^i\(d\)=argmin∑pχi\(d\),γi\(d\)\(\|v\(d\)\(ei,p\)−vP\(d\)\|−\(\|χi\(d\)\|⋅p−γi\(d\)⋅p2\)\)2\\hat\{\\chi\}\_\{i\}^\{\(d\)\},\\hat\{\\gamma\}\_\{i\}^\{\(d\)\}=\\arg\\min\_\{\\chi\_\{i\}^\{\(d\)\},\\gamma\_\{i\}^\{\(d\)\}\}\\sum\_\{p\}\\left\(\\left\|v^\{\(d\)\}\(e\_\{i\},p\)\-v\_\{P\}^\{\(d\)\}\\right\|\-\\left\(\|\\chi\_\{i\}^\{\(d\)\}\|\\cdot p\-\\gamma\_\{i\}^\{\(d\)\}\\cdot p^\{2\}\\right\)\\right\)^\{2\}\(7\)We then use the fitted coefficients to predict the displacement under any given pressure intensitypp: vE\(d\)=\|χ^d,i\|⋅p−γ^d,i⋅p2v\_\{E\}^\{\(d\)\}=\|\\hat\{\\chi\}\_\{d,i\}\|\\cdot p\-\\hat\{\\gamma\}\_\{d,i\}\\cdot p^\{2\}\(8\)whereχ^d,i\\hat\{\\chi\}\_\{d,i\}is the linear response coefficient for dimensionddunder factorii\. It captures how strongly dimensionddresponds to factoriiunder low pressure\.γ^d,i\\hat\{\\gamma\}\_\{d,i\}is a quadratic saturation parameter\. It captures the slowdown in response at high pressure\. Repeating this fit across all ten dimensions and four factors yields the full susceptibility matrixχ∈ℝ10×4\\chi\\in\\mathbb\{R\}^\{10\\times 4\}\. This matrix serves as a proxy for the model’s pliability\. It distinguishes dimensions and factors that maintain robust internal priors from those that exhibit high sensitivity under contextual framing\. ###### Cognition \(Factor C\): reasoning\-induced modulation\. To determine how the inference process modulates value expression, we control the computation path by comparing two decoding modes: Direct\-Answer and Chain\-of\-Thought \(CoT\)\. For the latter, we append the trigger “Let’s think step by step” to induce intermediate reasoning tokensRR\. We define theCognitive Modulation Vector,Δ𝐯C\\Delta\\mathbf\{v\}\_\{C\}, as the vector difference between the reasoning\-enhanced output and the direct output: Δ𝐯C=𝐯CoT−𝐯Direct\\Delta\\mathbf\{v\}\_\{C\}=\\mathbf\{v\}\_\{CoT\}\-\\mathbf\{v\}\_\{Direct\}\(9\) where𝐯CoT\\mathbf\{v\}\_\{CoT\}denotes the value\-expression vector obtained when the model generates intermediate reasoning tokens through Chain\-of\-Thought, and𝐯Direct\\mathbf\{v\}\_\{Direct\}denotes the corresponding value\-expression vector obtained from direct answering without explicit reasoning\. Each vector represents the model’s expressed preference distribution over the predefined value dimensions, andΔ𝐯C\\Delta\\mathbf\{v\}\_\{C\}captures the displacement of value expression induced by the reasoning process rather than the absolute value preference itself\. By analyzing the distributional shift ofΔ𝐯C\\Delta\\mathbf\{v\}\_\{C\}across the entire evaluation set, we quantify how reasoning processes systematically modulate LLM value expression under different models and contextual conditions\. #### 4\.4Alignment Prescription A natural question follows from the PEC decomposition \(Eq\. \([4](https://arxiv.org/html/2609.16589#S4.E4)\)\): for a given model and a target value dimension requiring correction, which intervention achieves sufficient realignment at the lowest cost? Before formalizing the prescription procedure, we clarify the division of labor between PEC and the intervention hierarchy\. The ordering Prompt, CoT, SFT/DPO, and continued pre\-training constitute a monotonically increasing cost ladder that follows directly from the computational architecture of LLM deployment; it is typically determined by empirical experience and experimental results\. The PEC framework’s contribution is to solve the*matching problem*on this pre\-existing ladder: given a model’s ten Schwartz value dimensions, which respond sufficiently to prompt\-level intervention, which require parametric updates, and which resist all available methods? Without PEC, this matching can only be performed through exhaustive experimental enumeration over every model–dimension–intervention combination\. PEC replaces brute\-force search with theory\-guided prediction, and we formalize this below as a hierarchical prescription procedure\. The empirical analyses of Factors P, E, and C jointly reveal a systematic heterogeneity in how different Schwartz value dimensions respond to external intervention \(Sections[2\.3\.1](https://arxiv.org/html/2609.16589#S2.SS3.SSS1),[2\.3\.2](https://arxiv.org/html/2609.16589#S2.SS3.SSS2), and[2\.3\.3](https://arxiv.org/html/2609.16589#S2.SS3.SSS3)\)\. These three layers of evidence converge on a single structural conclusion: value dimensions differ not merely in their current alignment state, but in the class of intervention required to shift them\. We therefore map the effect\-size profile of each dimension onto a four\-level prescription hierarchy, where the effect\-size profile is defined as the maximum absolute shift\|Δv\(d\)\|\|\\Delta v^\{\(d\)\}\|achievable by each factor\. Section[2\.4](https://arxiv.org/html/2609.16589#S2.SS4)reports how this mapping resolves in practice for each model–dimension pair, and Table[4](https://arxiv.org/html/2609.16589#S4.T4)summarizes the formal definition of each level\. This mapping ensures that the prescription hierarchy is not an arbitrary ordering but a direct empirical consequence of the PEC decomposition\. ###### Operationalizing PEC factors into comparable effect sizes\. The three PEC factors take different mathematical forms: the susceptibility matrixχ\\chicharacterizes Environment, the cognitive modulation vectorΔ𝐯C\\Delta\\mathbf\{v\}\_\{C\}characterizes Cognition, and the regression\-based profile of the baseline𝐯P\\mathbf\{v\}\_\{P\}characterizes Prior\. To compare them on equal footing, we convert each into the same type of quantity\. For every Schwartz dimensionddand every intervention classℐ∈\{E,C,P\}\\mathcal\{I\}\\in\\\{E,C,P\\\}, we compute a non\-negative scalarδℐ\(d\)\\delta\_\{\\mathcal\{I\}\}^\{\(d\)\}, defined as the absolute displacement in dimensionddinduced by interventionℐ\\mathcal\{I\}, normalized by the intensity of that intervention\. The three factors take the following specific forms\. Environment displacement\.The environment displacement is computed directly from the calibrated susceptibility matrixχ\\chi, aggregated across all contextual factors and normalized by pressure intensity: δE\(d\)=∑i=14\|χ^d,i\|⋅p−γ^d,i⋅p2p\\delta\_\{E\}^\{\(d\)\}=\\sum\_\{i=1\}^\{4\}\\frac\{\|\\hat\{\\chi\}\_\{d,i\}\|\\cdot p\-\\hat\{\\gamma\}\_\{d,i\}\\cdot p^\{2\}\}\{p\}\(10\)whereχ^d,i\\hat\{\\chi\}\_\{d,i\}andγ^d,i\\hat\{\\gamma\}\_\{d,i\}are the response and saturation coefficients fitted in Eq\. \([8](https://arxiv.org/html/2609.16589#S4.E8)\), andp∈\[0,1\]p\\in\[0,1\]is the pressure intensity serving as the intervention dose for E\. BecauseδE\(d\)\\delta\_\{E\}^\{\(d\)\}is computed analytically fromχ\\chi, displacement predictions generalize to untested scenario combinations and to new models whose susceptibility profiles have been characterized, removing the need for exhaustive per\-model, per\-dimension prompt search\. This predictive capacity is the core contribution of PEC to prescriptive alignment\. Cognition displacement\.The effect of Chain\-of\-Thought reasoning on dimensionddis normalized by the reasoning token overhead: δC\(d\)=\|v\(d\)\(CoT\)−v\(d\)\(Direct\)\|ntokens\\delta\_\{C\}^\{\(d\)\}\\;=\\;\\frac\{\\bigl\|\\,v^\{\(d\)\}\(\\mathrm\{CoT\}\)\\;\-\\;v^\{\(d\)\}\(\\mathrm\{Direct\}\)\\,\\bigr\|\}\{n\_\{\\mathrm\{tokens\}\}\}\(11\)wherev\(d\)\(CoT\)v^\{\(d\)\}\(\\mathrm\{CoT\}\)andv\(d\)\(Direct\)v^\{\(d\)\}\(\\mathrm\{Direct\}\)are the values under Chain\-of\-Thought and direct answering respectively, andntokensn\_\{\\mathrm\{tokens\}\}is the number of additional tokens generated by CoT\. Prior displacement\.The parametric intervention effect is normalized by training set size: δP\(d\)=\|v\(d\)\(Train\)−v\(d\)\(base\)\|Ntrain\\delta\_\{P\}^\{\(d\)\}\\;=\\;\\frac\{\\bigl\|\\,v^\{\(d\)\}\(\\mathrm\{Train\}\)\\;\-\\;v^\{\(d\)\}\(\\mathrm\{base\}\)\\,\\bigr\|\}\{N\_\{\\mathrm\{train\}\}\}\(12\)wherev\(d\)\(Train\)v^\{\(d\)\}\(\\mathrm\{Train\}\)is the value after parametric training, which in our experiments combines supervised fine\-tuning \(SFT\) and direct preference optimization \(DPO\),v\(d\)\(base\)v^\{\(d\)\}\(\\mathrm\{base\}\)is the neutral zero\-shot baselinevP\(d\)v\_\{P\}^\{\(d\)\}of the base model before parametric training, andNtrainN\_\{\\mathrm\{train\}\}is the number of training samples\. After this operationalization, each factor yields a non\-negative scalar per dimension\. All three scalars represent displacement per unit of intervention intensity and are expressed in the same Schwartz value space units, making them directly comparable\. ###### Alignment threshold\. An intervention is considered effective on a given dimension only if the induced shift exceeds a thresholdτ\\tau\. The choice ofτ\\tauis guided by the field\-theoretic[Lewin \(1951\)](https://arxiv.org/html/2609.16589#bib.bib55)view underlying the PEC framework\. In this view, a value dimension only ”moves” once the applied force exceeds a certain resistance, analogous to a threshold in a physical field\. We calibrate this threshold empirically, by comparing it against the natural variation observed across national populations in the WVS\-7 dataset\. The median single\-dimension mean difference between national populations in the Schwartz space provides a reference scale for what counts as a meaningful shift\. We setτ\\tauwithin this empirically observed range, so that an intervention is only counted as effective if it produces a shift comparable to, or larger than, the value differences naturally observed between distinct human populations\. This groundsτ\\tauin a real\-world scale, rather than treating it as an arbitrary cutoff\. ###### Hierarchical level assignment\. For each model and each of the ten Schwartz dimensions, we evaluate interventions in ascending order of cost\. The prescribed level for a given dimension is the lowest level whose absolute shift magnitude surpassesτ\\tau; if no level succeeds, the dimension is classified as*pre\-train locked*and assigned to Level 4\. Table[4](https://arxiv.org/html/2609.16589#S4.T4)summarizes the four levels\. Cost increases monotonically along the ladder:cE<cC<cPc\_\{E\}<c\_\{C\}<c\_\{P\}\. Because of this, evaluating interventions in ascending order and stopping at the first one that reachesτ\\tauis equivalent to selecting the lowest\-cost sufficient intervention\. When no single intervention reachesτ\\tauon its own, interventions are combined in ascending cost order\. Section[2\.3\.2](https://arxiv.org/html/2609.16589#S2.SS3.SSS2)documents an antagonistic interaction between Environment and Cognition factors\. Activating CoT weakens the model’s overall sensitivity to environmental factors\. To account for this interaction, a discount coefficientα\\alphais applied to subsequent terms in the cumulative displacement\. We defineα\\alphaas the ratio of the Frobenius norm of the susceptibility matrixχ\\chiunder CoT to the Frobenius norm ofχ\\chiwithout CoT: α=‖χCoT‖F‖χDirect‖F,\\alpha\\;=\\;\\frac\{\\\|\\chi\_\{\\mathrm\{CoT\}\}\\\|\_\{F\}\}\{\\\|\\chi\_\{\\mathrm\{Direct\}\}\\\|\_\{F\}\},\(13\)We observe that for most models, generating a CoT reduces overall environmental sensitivity \(See in Section[2\.3\.2](https://arxiv.org/html/2609.16589#S2.SS3.SSS2)\)\. As a result,α\\alphais typically below 1 for each model\. The average value ofα\\alphaacross models is 0\.88\. Formally, the discounted cumulative displacement, evaluated in ascending cost order E, C, P, is defined as Sℐ\(d\)=\{δE\(d\),ℐ=EδE\(d\)\+α⋅δC\(d\),ℐ=CδE\(d\)\+α⋅δC\(d\)\+α⋅δP\(d\),ℐ=PS^\{\(d\)\}\_\{\\mathcal\{I\}\}\\;=\\;\\begin\{cases\}\\delta\_\{E\}^\{\(d\)\},&\\mathcal\{I\}=E\\\\ \\delta\_\{E\}^\{\(d\)\}\\;\+\\;\\alpha\\cdot\\delta\_\{C\}^\{\(d\)\},&\\mathcal\{I\}=C\\\\ \\delta\_\{E\}^\{\(d\)\}\\;\+\\;\\alpha\\cdot\\delta\_\{C\}^\{\(d\)\}\\;\+\\;\\alpha\\cdot\\delta\_\{P\}^\{\(d\)\},&\\mathcal\{I\}=P\\end\{cases\}\(14\)The prescribed level is the lowest\-cost intervention classℐ\\mathcal\{I\}such thatSℐ\(d\)\>τS^\{\(d\)\}\_\{\\mathcal\{I\}\}\>\\tau; if no suchℐ\\mathcal\{I\}exists, the dimension is diagnosed as pre\-train locked and flagged for continued pre\-training \(Level 4\)\. To facilitate comparison of intervention efficiency, we additionally report an effect\-per\-cost ratio for each intervention class\. The detailed definition, FLOP\-based cost normalization, and per\-model cost estimation are provided in Appendix[12\.2](https://arxiv.org/html/2609.16589#S12.SS2)\. This metric serves only as a diagnostic measure of intervention efficiency and is not used as a criterion for determining the prescribed intervention level\. ###### Intervention strategies\. Based on the PEC framework, we implement four intervention strategies corresponding to the four levels defined in Table[4](https://arxiv.org/html/2609.16589#S4.T4), ordered by increasing computational cost\. Each strategy is applied only when all lower\-cost alternatives have been evaluated and found insufficient\. Table 4:Four\-level intervention hierarchy ordered by ascending cost\.τ\\taudenotes the minimum shift magnitude required for a level to be considered effective\. ###### Prescription validation\. To verify the generalizability of the prescription matrix, we evaluate each model\-dimension pair on a held\-out dataset and compare the lowest effective level observed during validation with the prescribed level\. The match rate across all pairs quantifies the reliability of the prescription framework\. Detailed mathematical formulations of each intervention strategy, the prescription assignment algorithm, and the validation protocol are provided in Appendix[11\.4](https://arxiv.org/html/2609.16589#S11.SS4)\. #### 4\.5Related Work This research is situated at the interdisciplinary frontier of LLM value evaluation and alignment\. This section provides a comprehensive review of existing literature across two dimensions: value evaluation frameworks and alignment methodologies, while delineating our specific contributions and positioning within the field\. ##### 4\.5\.1Evaluation of LLMs’ values Conventional value assessments typically rely on theoretical extensions of the “Three Laws of Robotics”[Asimov \(1950\)](https://arxiv.org/html/2609.16589#bib.bib39)or machine ethics definitions[Moor \(2006\)](https://arxiv.org/html/2609.16589#bib.bib40), employing descriptive probes through established frameworks such as the Moral Foundations Questionnaire \(MFQ\)[Graham et al\. \(2008\)](https://arxiv.org/html/2609.16589#bib.bib41)or Shweder’s “big three” of morality[Shweder et al\. \(2013\)](https://arxiv.org/html/2609.16589#bib.bib42)\. Although Yao et al\.[Yao et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib43);[Yao et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib44)demonstrated the feasibility of mapping LLMs onto the Schwartz Theory of Basic Human Values, their evaluation still compresses each model into a single value vector\. Most current evaluation methods summarize an LLM’s values into a single vector for comparison against human values\. These methods do not perform this comparison in distributional form\. As a result, a great deal of intrinsic structural information is lost\. Serapio\-García et al\.[Serapio\-García et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib45)further proposed a psychometric framework for both evaluating and shaping personality traits in LLMs, providing methodological evidence that LLMs exhibit reliable and valid synthetic personality structures under appropriate prompting configurations\. However, these approaches lack a unified computational framework capable of aligning “model values” with “human socio\-cultural benchmarks” \(such as national cultural disparities\)\. Our study bridges this gap by introducing an uncertainty\-centric perspective into the realm of value modeling\. ##### 4\.5\.2Value Alignment of LLMs As LLMs achieve human\-level performance across a wide array of general tasks, ensuring that their underlying intentions, preferences, and behavioral norms remain consistent with human values has become a core issue in AI safety\. Current alignment techniques are primarily categorized into three paradigms\. First, SFT trains models on high\-quality, human\-curated datasets to directly inject intended behavioral norms into the LLMs’ parameters[Wang et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib46)\. Second, RLHF utilizes preference data to construct reward models that guide the fine\-tuning process[Nakano et al\. \(2021\)](https://arxiv.org/html/2609.16589#bib.bib47)\. Constitutional AI further extends this paradigm by replacing human feedback with AI\-generated critiques and principles, enabling scalable and rule\-governed alignment without exhaustive human annotation[Bai et al\. \(2022\)](https://arxiv.org/html/2609.16589#bib.bib48)\. However, traditional RLHF methods often rely on a single scalar reward function, which possesses inherent limitations when navigating complex and potentially conflicting pluralistic value systems\. To address this bottleneck, frontier research has begun to formalize value alignment within a Multi\-Objective Reinforcement Learning \(MORL\) framework[Rodriguez\-Soto et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib49)\. By employing multi\-dimensional reward vectors to explicitly represent distinct value dimensions, these methods first learn a set of Pareto\-optimal policies within a multi\-objective space\. Subsequently, they utilize preference modeling or weight\-adjustment mechanisms to achieve precise incentive compatibility and alignment with specific value systems during deployment\. Finally, In\-Context Alignment \(ICA\) has emerged as a parameter\-efficient alternative, dynamically regulating model behavior through sophisticated system instructions or critiques during the inference phase[Ganguli et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib50);[Gou et al\. \(2024\)](https://arxiv.org/html/2609.16589#bib.bib51)\. Given its low computational overhead and compatibility with black\-box models, ICA shows great promise for enhancing value consistency in large\-scale applications\. Despite these methodological advancements, existing alignment strategies struggle to encompass potential and unforeseen risks\. Significant challenges remain in ensuring the robustness and long\-term consistency of value adherence in increasingly complex environments\. ### 5Conclusion Together, these findings answer the three questions raised at the start of this study\. LLMs do possess an intrinsic value system\. This system can be quantified through the PEC framework\. It can also be aligned through the Alignment Prescription\. Beyond these specific answers, this study points to a broader shift in how we understand LLM behavior\. Value expression has often been treated as noise, or as an unpredictable side effect of prompting\. Our results show that this behavior follows a measurable structure instead\. This structure can be traced back to its origin in the Prior, the Environment, and the Cognition factors\. This shift moves value alignment from opaque trial\-and\-error toward a mechanistic and quantitative science\. It also reduces the alignment tax through precise, prescriptive interventions\. LLMs are taking on more decision\-making roles today\. In this setting, structural understanding becomes as important as raw performance\. Our framework offers one concrete step toward this goal\. It helps keep generative models as transparent and controllable tools, not unexamined ideological black boxes\. ### 6Data availability ### 7Code availability The code used during the current study, including scripts for data processing, model evaluation, and figure generation performed in Python \(version 3\.10\), is available via GitHub at[https://github\.com/KQ\-Zhang/Do\-LLMs\-Have\-Values\-Release\.git](https://github.com/KQ-Zhang/Do-LLMs-Have-Values-Release.git)\. Python packages used for analysis includepandas,numpy,scikit\-learn,matplotlib, andtransformers\. An interactive demonstration is also provided in the repository\. ## References - Anthropic \(2024\)AnthropicClaude’s model card\.Note:[https://www\.anthropic\.com/model\-card](https://www.anthropic.com/model-card)Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.16.1.1),[§9](https://arxiv.org/html/2609.16589#S9.p2.1)\. - Argyleet al\.\(2023\)L\. P\. Argyle, E\. C\. Busby, N\. Fulda, J\. R\. Gubler, C\. Rytting, and D\. WingateOut of one, many: using language models to simulate human samples\.Political Analysis31\(3\),pp\. 337–351\.External Links:[Document](https://dx.doi.org/10.1017/pan.2023.2)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p4.1)\. - Asimov \(1950\)I\. AsimovI, robot\.Doubleday,New York\.Cited by:[§4\.5\.1](https://arxiv.org/html/2609.16589#S4.SS5.SSS1.p1.1)\. - Baiet al\.\(2022\)Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon,et al\.Constitutional AI: harmlessness from AI feedback\.External Links:2212\.08073Cited by:[§4\.5\.2](https://arxiv.org/html/2609.16589#S4.SS5.SSS2.p1.1)\. - Barbieriet al\.\(2020\)F\. Barbieri, J\. Camacho\-Collados, L\. Espinosa Anke, and L\. NevesTweetEval: unified benchmark and comparative evaluation for tweet classification\.InFindings of the Association for Computational Linguistics: EMNLP 2020,Online,pp\. 1644–1650\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.148)Cited by:[§2\.1](https://arxiv.org/html/2609.16589#S2.SS1.p2.1)\. - Betleyet al\.\(2026\)J\. Betley, N\. Warncke, A\. Sztyber\-Betley, D\. Tan, X\. Bao, M\. Soto, M\. Srivastava, N\. Labenz, and O\. EvansTraining large language models on narrow tasks can lead to broad misalignment\.Nature649\(8097\),pp\. 584–589\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09937-5)Cited by:[§3\.5](https://arxiv.org/html/2609.16589#S3.SS5.p1.1)\. - BigScience Workshop \(2023\)BigScience WorkshopBLOOM: a 176b\-parameter open\-access multilingual language model\.arXiv preprint arXiv:2211\.05100\.External Links:2211\.05100Cited by:[§4\.3](https://arxiv.org/html/2609.16589#S4.SS3.SSS0.Px1.p1.1)\. - Binz and Schulz \(2023\)M\. Binz and E\. SchulzUsing cognitive psychology to understand GPT\-3\.Proceedings of the National Academy of Sciences120\(6\),pp\. e2218523120\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2218523120)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p6.1)\. - Brownet al\.\(2020\)T\. B\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. M\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. AmodeiLanguage models are few\-shot learners\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 1877–1901\.Cited by:[§2\.3](https://arxiv.org/html/2609.16589#S2.SS3.p2.2)\. - ByteDance Seed Team \(2025\)ByteDance Seed TeamDoubao: a family of large language models from ByteDance\.arXiv preprint arXiv:2505\.13458\.External Links:2505\.13458Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.14.1.1)\. - Chowdheryet al\.\(2023\)A\. Chowdhery, S\. Narang, J\. Devlin, M\. Bosma, G\. Mishra, A\. Roberts, P\. Barham, H\. W\. Chung, C\. Sutton, S\. Gehrmann,et al\.PaLM: scaling language modeling with pathways\.Journal of Machine Learning Research24\(240\),pp\. 1–113\.Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p1.1)\. - Christianoet al\.\(2017\)P\. F\. Christiano, J\. Leike, T\. B\. Brown, M\. Martic, S\. Legg, and D\. AmodeiDeep reinforcement learning from human preferences\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\. 4299–4307\.Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p2.1)\. - Devlinet al\.\(2019\)J\. Devlin, M\. Chang, K\. Lee, and K\. ToutanovaBERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),Minneapolis, Minnesota,pp\. 4171–4186\.Cited by:[§2\.3](https://arxiv.org/html/2609.16589#S2.SS3.p2.2)\. - Dowson and Landau \(1982\)D\. C\. Dowson and B\. V\. LandauThe Fréchet distance between multivariate normal distributions\.Journal of Multivariate Analysis12\(3\),pp\. 450–455\.External Links:[Document](https://dx.doi.org/10.1016/0047-259X%2882%2990077-X)Cited by:[§4\.2\.3](https://arxiv.org/html/2609.16589#S4.SS2.SSS3.p2.1)\. - Dubeyet al\.\(2024\)A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Yang, A\. Fan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.External Links:2407\.21783Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.30.1.1)\. - Durmuset al\.\(2024\)E\. Durmus, K\. Nguyen, T\. I\. Liao, N\. Schiefer, A\. Askell, A\. Bakhtin, C\. Chen, Z\. Hatfield\-Dodds, D\. Hernandez, N\. Joseph, L\. Lovitt, S\. McCandlish, O\. Sikder, A\. Tamkin, J\. Thamkul, J\. Kaplan, J\. Clark, and D\. GanguliTowards measuring the representation of subjective global opinions in language models\.InFirst Conference on Language Modeling,Cited by:[§3\.1](https://arxiv.org/html/2609.16589#S3.SS1.p1.1)\. - Farquharet al\.\(2024\)S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. GalDetecting hallucinations in large language models using semantic entropy\.Nature630\(8017\),pp\. 625–630\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07421-0)Cited by:[§2\.1](https://arxiv.org/html/2609.16589#S2.SS1.p2.1)\. - Gallifantet al\.\(2024\)J\. Gallifantet al\.Peer review of GPT\-4 technical report and systems card\.PLOS Digital Health18,pp\. e0000417\.External Links:[Document](https://dx.doi.org/10.1371/journal.pdig.0000417)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p1.1)\. - Ganguliet al\.\(2023\)D\. Ganguli, A\. Askell, N\. Schiefer, T\. I\. Liao, K\. Lukošiūtė, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. Olsson, D\. Hernandez,et al\.The capacity for moral self\-correction in large language models\.arXiv preprint arXiv:2302\.07459\.Cited by:[§4\.5\.2](https://arxiv.org/html/2609.16589#S4.SS5.SSS2.p2.1)\. - Gelbrich \(1990\)M\. GelbrichOn a formula for theL2L^\{2\}Wasserstein metric between measures on Euclidean and Hilbert spaces\.Mathematische Nachrichten147\(1\),pp\. 185–203\.External Links:[Document](https://dx.doi.org/10.1002/mana.19901470121)Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16589#S2.SS2.SSS2.p4.1)\. - Gouet al\.\(2024\)Z\. Gou, Z\. Shao, Y\. Gong, Y\. Shen, Y\. Yang, N\. Duan, and W\. ChenCRITIC: large language models can self\-correct with tool\-interactive critiquing\.InThe Twelfth International Conference on Learning Representations,Vienna, Austria\.Cited by:[§4\.5\.2](https://arxiv.org/html/2609.16589#S4.SS5.SSS2.p2.1)\. - Grahamet al\.\(2008\)J\. Graham, B\. A\. Nosek, J\. Haidt, R\. Iyer, K\. Spassena, and P\. H\. DittoMoral foundations questionnaire\.Personality and Social Psychology Bulletin\.Cited by:[§4\.5\.1](https://arxiv.org/html/2609.16589#S4.SS5.SSS1.p1.1)\. - Guoet al\.\(2025\)D\. Guo, D\. Yang, H\. Zhang, J\. Song, R\. Zhang, R\. Xu, Q\. Zhu, S\. Ma, P\. Wang, X\. Bi,et al\.DeepSeek\-R1: incentivizing reasoning capability in LLMs via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.External Links:2501\.12948Cited by:[§9](https://arxiv.org/html/2609.16589#S9.p2.1)\. - Haerpferet al\.\(2022\)C\. Haerpfer, R\. Inglehart, A\. Moreno, C\. Welzel, K\. Kizilova, J\. Diez\-Medrano, M\. Lagos, P\. Norris, E\. Ponarin, and B\. PuranenWorld values survey: round seven – country\-pooled datafile version 5\.0\.JD Systems Institute & WVSA Secretariat,Madrid, Spain & Vienna, Austria\.External Links:[Document](https://dx.doi.org/10.14281/18241.20)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p4.1),[§2\.2\.1](https://arxiv.org/html/2609.16589#S2.SS2.SSS1.p1.1)\. - Hagendorffet al\.\(2023\)T\. Hagendorff, S\. Fabi, and M\. KosinskiHuman\-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT\.Nature Computational Science3,pp\. 833–838\.External Links:[Document](https://dx.doi.org/10.1038/s43588-023-00527-x)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p1.1),[§1](https://arxiv.org/html/2609.16589#S1.p2.1),[§1](https://arxiv.org/html/2609.16589#S1.p6.1)\. - Heuselet al\.\(2017\)M\. Heusel, H\. Ramsauer, T\. Unterthiner, B\. Nessler, and S\. HochreiterGANs trained by a two time\-scale update rule converge to a local Nash equilibrium\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§4\.2\.3](https://arxiv.org/html/2609.16589#S4.SS2.SSS3.p2.1)\. - Huet al\.\(2022\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p7.1)\. - Huet al\.\(2025\)T\. Hu, Y\. Kyrychenko, S\. Rathje, N\. Collier, S\. van der Linden, and J\. RoozenbeekGenerative language models exhibit social identity biases\.Nature Computational Science5\(1\),pp\. 65–75\.External Links:[Document](https://dx.doi.org/10.1038/s43588-024-00741-1)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p1.1),[§1](https://arxiv.org/html/2609.16589#S1.p2.1)\. - Inglehart and Welzel \(2005\)R\. Inglehart and C\. WelzelModernization, cultural change, and democracy: the human development sequence\.Cambridge University Press,Cambridge, UK\.External Links:[Document](https://dx.doi.org/10.1017/CBO9780511790881)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p4.1)\. - Jiet al\.\(2025\)J\. Ji, D\. Hong, B\. Zhang, B\. Chen, J\. Dai, B\. Zheng, T\. A\. Qiu, J\. Zhou, K\. Wang, B\. Li, S\. Han, Y\. Guo, and Y\. YangPKU\-SafeRLHF: towards multi\-level safety alignment for LLMs with human preference\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 31983–32016\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1544)Cited by:[§2\.4](https://arxiv.org/html/2609.16589#S2.SS4.p5.1)\. - Kahneman \(2003\)D\. KahnemanMaps of bounded rationality: psychology for behavioral economics\.American Economic Review93\(5\),pp\. 1449–1475\.Cited by:[§3\.4](https://arxiv.org/html/2609.16589#S3.SS4.p1.1)\. - Kaplanet al\.\(2020\)J\. Kaplan, S\. McCandlish, T\. Henighan, T\. B\. Brown, B\. Chess, R\. Child, S\. Gray, A\. Radford, J\. Wu, and D\. AmodeiScaling laws for neural language models\.arXiv preprint arXiv:2001\.08361\.Cited by:[§12\.2](https://arxiv.org/html/2609.16589#S12.SS2.p8.1)\. - Khamassiet al\.\(2024\)M\. Khamassi, M\. Nahon, and R\. ChatilaStrong and weak alignment of large language models with human values\.Scientific Reports\.External Links:[Document](https://dx.doi.org/10.1038/s41598-024-70031-3)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p1.1)\. - Kuhnet al\.\(2023\)L\. Kuhn, Y\. Gal, and S\. FarquharSemantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation\.InThe Eleventh International Conference on Learning Representations,Kigali, Rwanda\.Cited by:[§2\.1](https://arxiv.org/html/2609.16589#S2.SS1.p2.1)\. - Kumaranet al\.\(2026\)D\. Kumaran, S\. M\. Fleming, L\. Markeeva, J\. Heyward, A\. Banino, M\. Mathur, R\. Pascanu, S\. Osindero, B\. De Martino, P\. Veličković, and V\. PatrauceanCompeting biases underlie overconfidence and underconfidence in LLMs\.Nature Machine Intelligence8,pp\. 614–627\.External Links:[Document](https://dx.doi.org/10.1038/s42256-026-01217-9)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p6.1)\. - Lewin \(1951\)K\. LewinD\. Cartwright \(Ed\.\)Field theory in social science: selected theoretical papers\.Harper & Brothers,New York\.Cited by:[§2\.3](https://arxiv.org/html/2609.16589#S2.SS3.p2.1),[§4\.3](https://arxiv.org/html/2609.16589#S4.SS3.p1.1),[§4\.4](https://arxiv.org/html/2609.16589#S4.SS4.SSS0.Px2.p1.1)\. - Liuet al\.\(2024\)A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan,et al\.DeepSeek\-v3 technical report\.arXiv preprint arXiv:2412\.19437\.External Links:2412\.19437Cited by:[§2\.1](https://arxiv.org/html/2609.16589#S2.SS1.p2.1),[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.13.1.1)\. - McInneset al\.\(2018\)L\. McInnes, J\. Healy, N\. Saul, and L\. GroßbergerUMAP: uniform manifold approximation and projection\.Journal of Open Source Software3\(29\),pp\. 861\.External Links:[Document](https://dx.doi.org/10.21105/joss.00861)Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16589#S2.SS2.SSS2.p2.1)\. - Messeri and Crockett \(2024\)L\. Messeri and M\. J\. CrockettArtificial intelligence and illusions of understanding in scientific research\.Nature627\(8002\),pp\. 49–58\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07146-0)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p1.1)\. - Moonshot AI \(2025\)Moonshot AIKimi K2: open agentic intelligence\.Note:[https://github\.com/MoonshotAI/Kimi\-K2](https://github.com/MoonshotAI/Kimi-K2)Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.24.1.1)\. - Moor \(2006\)J\. H\. MoorThe nature, importance, and difficulty of machine ethics\.IEEE intelligent systems21\(4\),pp\. 18–21\.Cited by:[§4\.5\.1](https://arxiv.org/html/2609.16589#S4.SS5.SSS1.p1.1)\. - Morriset al\.\(2025\)J\. X\. Morris, C\. Sitawarin, C\. Guo, N\. Kokhlikyan, G\. E\. Suh, A\. M\. Rush, K\. Chaudhuri, and S\. MahloujifarHow much do language models memorize?\.arXiv preprint arXiv:2505\.24832\.Cited by:[§3\.3](https://arxiv.org/html/2609.16589#S3.SS3.p1.1)\. - Nakanoet al\.\(2021\)R\. Nakano, J\. Hilton, S\. Balaji, J\. Wu, L\. Ouyang, C\. Kim, C\. Hesse, S\. Jain, V\. Kosaraju, W\. Saunders, X\. Jiang, K\. Cobbe, T\. Eloundou, G\. Krueger, K\. Button, M\. Knight, B\. Chess, and J\. SchulmanWebGPT: browser\-assisted question\-answering with human feedback\.External Links:2112\.09332Cited by:[§4\.5\.2](https://arxiv.org/html/2609.16589#S4.SS5.SSS2.p1.1)\. - OpenAI \(2023\)OpenAIGPT\-4 technical report\.arXiv preprint arXiv:2303\.08774\.External Links:2303\.08774Cited by:[Figure 12](https://arxiv.org/html/2609.16589#S11.F12),[Figure 12](https://arxiv.org/html/2609.16589#S11.F12.8.3),[§2\.1](https://arxiv.org/html/2609.16589#S2.SS1.p2.1),[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.18.1.1),[§9](https://arxiv.org/html/2609.16589#S9.p2.1)\. - OpenAI \(2025\)OpenAIGPT\-4\.1 technical report\.External Links:2503\.08747Cited by:[Figure 12](https://arxiv.org/html/2609.16589#S11.F12),[Figure 12](https://arxiv.org/html/2609.16589#S11.F12.8.2),[§9](https://arxiv.org/html/2609.16589#S9.p2.1)\. - Ouyanget al\.\(2022\)L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. F\. Christiano, J\. Leike, and R\. LoweTraining language models to follow instructions with human feedback\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 27730–27744\.Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p2.1),[§2\.3](https://arxiv.org/html/2609.16589#S2.SS3.p2.2)\. - Peterset al\.\(2018\)M\. E\. Peters, M\. Neumann, M\. Iyyer, M\. Gardner, C\. Clark, K\. Lee, and L\. ZettlemoyerDeep contextualized word representations\.InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,New Orleans, Louisiana\.Cited by:[§2\.3](https://arxiv.org/html/2609.16589#S2.SS3.p2.2)\. - Peyré and Cuturi \(2019\)G\. Peyré and M\. CuturiComputational optimal transport\.Foundations and Trends in Machine Learning11\(5–6\),pp\. 355–607\.External Links:[Document](https://dx.doi.org/10.1561/2200000073)Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16589#S2.SS2.SSS2.p4.1)\. - Radfordet al\.\(2019\)A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. SutskeverLanguage models are unsupervised multitask learners\.OpenAI Blog1\(8\),pp\. 9\.Cited by:[Figure 12](https://arxiv.org/html/2609.16589#S11.F12),[Figure 12](https://arxiv.org/html/2609.16589#S11.F12.8.4)\. - Raffelet al\.\(2020\)C\. Raffel, N\. Shazeer, A\. Roberts, K\. Lee, S\. Narang, M\. Matena, Y\. Zhou, W\. Li, and P\. J\. LiuExploring the limits of transfer learning with a unified text\-to\-text transformer\.Journal of Machine Learning Research21\(140\),pp\. 1–67\.Cited by:[§2\.3](https://arxiv.org/html/2609.16589#S2.SS3.p2.2)\. - Rauthmannet al\.\(2014\)J\. F\. Rauthmann, D\. Gallardo\-Pujol, E\. M\. Guillaume, E\. Todd, C\. S\. Nave, R\. A\. Sherman, M\. Ziegler, A\. B\. Jones, and D\. C\. FunderThe situational eight DIAMONDS: a taxonomy of major dimensions of situation characteristics\.Journal of Personality and Social Psychology107\(4\),pp\. 677–718\.External Links:[Document](https://dx.doi.org/10.1037/a0037250)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.16589#S2.SS3.SSS1.p1.1)\. - Rodriguez\-Sotoet al\.\(2025\)M\. Rodriguez\-Soto, R\. Rădulescu, F\. Bistaffa, O\. Ricart, A\. Mayoral, M\. Lopez\-Sanchez, J\. A\. Rodriguez\-Aguilar, and A\. NowéMulti\-objective reinforcement learning for provably incentivising alignment with value systems\.Artificial Intelligence,pp\. 104460\.Cited by:[§4\.5\.2](https://arxiv.org/html/2609.16589#S4.SS5.SSS2.p2.1)\. - Santurkaret al\.\(2023\)S\. Santurkar, E\. Durmus, F\. Ladhak, C\. Lee, P\. Liang, and T\. HashimotoWhose opinions do language models reflect?\.InProceedings of the 40th International Conference on Machine Learning,Honolulu, Hawaii, USA,pp\. 29971–30004\.Cited by:[§3\.1](https://arxiv.org/html/2609.16589#S3.SS1.p1.1)\. - Scherreret al\.\(2023\)N\. Scherrer, C\. Shi, A\. Feder, and D\. BleiEvaluating the moral beliefs encoded in llms\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 51778–51809\.Cited by:[§3\.1](https://arxiv.org/html/2609.16589#S3.SS1.p1.1)\. - Schwartz \(1992\)S\. H\. SchwartzUniversals in the content and structure of values: theoretical advances and empirical tests in 20 countries\.InAdvances in Experimental Social Psychology,M\. P\. Zanna \(Ed\.\),Vol\.25,pp\. 1–65\.Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p4.1)\. - Schwartz \(2012\)S\. H\. SchwartzAn overview of the Schwartz theory of basic values\.Online Readings in Psychology and Culture2\(1\)\.External Links:[Document](https://dx.doi.org/10.9707/2307-0919.1116)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p4.1),[§4\.2\.1](https://arxiv.org/html/2609.16589#S4.SS2.SSS1.p1.1)\. - Sclaret al\.\(2024\)M\. Sclar, Y\. Choi, Y\. Tsvetkov, and A\. SuhrQuantifying language models’ sensitivity to spurious features in prompt design or: how I learned to start worrying about prompt formatting\.InProceedings of the Twelfth International Conference on Learning Representations,Vienna, Austria\.Cited by:[§2\.3\.1](https://arxiv.org/html/2609.16589#S2.SS3.SSS1.p1.1)\. - Serapio\-Garcíaet al\.\(2025\)G\. Serapio\-García, M\. Safdari, C\. Crepy, L\. Sun, S\. Fitz, P\. Romero, M\. Abdulhai, A\. Faust, and M\. MatarícA psychometric framework for evaluating and shaping personality traits in large language models\.Nature Machine Intelligence7,pp\. 1954–1968\.External Links:[Document](https://dx.doi.org/10.1038/s42256-025-01115-6)Cited by:[§4\.5\.1](https://arxiv.org/html/2609.16589#S4.SS5.SSS1.p2.1)\. - Sharmaet al\.\(2012\)S\. Sharma, S\. Durvasula, and R\. E\. PloyhartThe analysis of mean differences using mean and covariance structure analysis\.Organizational Research Methods15\(1\),pp\. 75–102\.External Links:[Document](https://dx.doi.org/10.1177/1094428111403154)Cited by:[§4\.2\.3](https://arxiv.org/html/2609.16589#S4.SS2.SSS3.p1.1)\. - Shwederet al\.\(2013\)R\. A\. Shweder, N\. C\. Much, M\. Mahapatra, and L\. ParkThe “big three” of morality \(autonomy, community, divinity\) and the “big three” explanations of suffering\.InMorality and health,pp\. 119–169\.Cited by:[§4\.5\.1](https://arxiv.org/html/2609.16589#S4.SS5.SSS1.p1.1)\. - Steyverset al\.\(2025\)M\. Steyvers, H\. Tejeda, A\. Kumar, C\. G\. Belem, S\. Karny, X\. Hu, L\. W\. Mayer, and P\. SmythWhat large language models know and what people think they know\.Nature Machine Intelligence7,pp\. 221–231\.External Links:[Document](https://dx.doi.org/10.1038/s42256-024-00976-7)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p6.1)\. - Strachanet al\.\(2024\)J\. W\. A\. Strachan, D\. Albergo, G\. Borghini, O\. Pansardi, E\. Scaliti, S\. Gupta, K\. Saxena, A\. Rufo, S\. Panzeri, G\. Manzi, M\. S\. A\. Graziano, and C\. BecchioTesting theory of mind in large language models and humans\.Nature Human Behaviour8\(7\),pp\. 1285–1295\.External Links:[Document](https://dx.doi.org/10.1038/s41562-024-01882-z)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p1.1)\. - Sunet al\.\(2024\)X\. Sun, Y\. Chen, Y\. Huang, R\. Xie, J\. Zeng, C\. Zhang, X\. Zhang, F\. Ma, Z\. Yuan,et al\.Hunyuan\-large: an open\-source MoE model with 52 billion activated parameters by Tencent\.arXiv preprint arXiv:2411\.02265\.External Links:2411\.02265Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.28.1.1)\. - Taoet al\.\(2024\)Y\. Tao, O\. Viberg, R\. S\. Baker, and R\. F\. KizilcecCultural bias and cultural alignment of large language models\.PNAS Nexus3\(9\),pp\. pgae346\.External Links:[Document](https://dx.doi.org/10.1093/pnasnexus/pgae346)Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p4.1)\. - Teamet al\.\(2023\)G\. Team, R\. Anil, S\. Borgeaud, J\. Alayrac, J\. Yu, R\. Soricut, J\. Schalkwyk, A\. M\. Dai, A\. Hauth,et al\.Gemini: a family of highly capable multimodal models\.arXiv preprint arXiv:2312\.11805\.External Links:2312\.11805Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.15.1.1),[§9](https://arxiv.org/html/2609.16589#S9.p2.1)\. - Teamet al\.\(2024\)G\. Team, A\. Zeng, B\. Xu, B\. Wang, C\. Zhang, D\. Yin, D\. Zhang, D\. Rojas, G\. Feng, H\. Zhao,et al\.ChatGLM: a family of large language models from GLM\-130B to GLM\-4 all tools\.arXiv preprint arXiv:2406\.12793\.External Links:2406\.12793Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.27.1.1)\. - Team \(2024\)Q\. TeamQwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.External Links:2412\.15115Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.32.1.1)\. - Team \(2025\)Q\. TeamQwen3 technical report\.arXiv preprint arXiv:2505\.09388\.External Links:2505\.09388Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.32.1.1)\. - Touvronet al\.\(2023\)H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale,et al\.Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.External Links:2307\.09288Cited by:[Table 1](https://arxiv.org/html/2609.16589#S2.T1.3.30.1.1)\. - Tversky and Kahneman \(1981\)A\. Tversky and D\. KahnemanThe framing of decisions and the psychology of choice\.Science211\(4481\),pp\. 453–458\.External Links:[Document](https://dx.doi.org/10.1126/science.7455683)Cited by:[§3\.4](https://arxiv.org/html/2609.16589#S3.SS4.p2.1)\. - Wanet al\.\(2025\)Y\. Wan, X\. Jia, and X\. L\. LiUnveiling confirmation bias in chain\-of\-thought reasoning\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 3788–3804\.Cited by:[§3\.4](https://arxiv.org/html/2609.16589#S3.SS4.p3.1)\. - Wanget al\.\(2023\)Y\. Wang, Y\. Kordi, S\. Mishra, A\. Liu, N\. A\. Smith, D\. Khashabi, and H\. HajishirziSelf\-instruct: aligning language models with self\-generated instructions\.InProceedings of the 61st annual meeting of the association for computational linguistics \(volume 1: long papers\),pp\. 13484–13508\.Cited by:[§4\.5\.2](https://arxiv.org/html/2609.16589#S4.SS5.SSS2.p1.1)\. - Weiet al\.\(2022a\)J\. Wei, M\. Bosma, V\. Y\. Zhao, K\. Guu, A\. W\. Yu, B\. Lester, N\. Du, A\. M\. Dai, and Q\. V\. LeFinetuned language models are zero\-shot learners\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p2.1)\. - Weiet al\.\(2022b\)J\. Wei, Y\. Tay, R\. Bommasani, C\. Raffel, B\. Zoph, S\. Borgeaud, D\. Yogatama, M\. Bosma, D\. Zhou, D\. Metzler, E\. H\. Chi, T\. Hashimoto, O\. Vinyals, P\. Liang, J\. Dean, and W\. FedusEmergent abilities of large language models\.Transactions on Machine Learning Research\.Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p1.1)\. - Weiet al\.\(2022c\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. V\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 24824–24837\.Cited by:[§1](https://arxiv.org/html/2609.16589#S1.p7.1)\. - Witteet al\.\(2020\)E\. H\. Witte, A\. Stanciu, and K\. BoehnkeA new empirical approach to intercultural comparisons of value preferences based on Schwartz’s theory\.Frontiers in Psychology11,pp\. 1723\.External Links:[Document](https://dx.doi.org/10.3389/fpsyg.2020.01723)Cited by:[§4\.2\.3](https://arxiv.org/html/2609.16589#S4.SS2.SSS3.p1.1)\. - Yaoet al\.\(2025\)J\. Yao, X\. Yi, S\. Duan, J\. Wang, Y\. Bai, M\. Huang, Y\. Ou, S\. Li, P\. Zhang, T\. Lu, Z\. Dou, M\. Sun, J\. Evans, and X\. XieValue compass benchmarks: a comprehensive, generative and self\-evolving platform for LLMs’ value evaluation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),pp\. 666–678\.Cited by:[§4\.2\.3](https://arxiv.org/html/2609.16589#S4.SS2.SSS3.p1.1),[§4\.5\.1](https://arxiv.org/html/2609.16589#S4.SS5.SSS1.p1.1)\. - Yaoet al\.\(2024\)J\. Yao, X\. Yi, Y\. Gong, X\. Wang, and X\. XieValue FULCRA: mapping large language models to the multidimensional spectrum of basic human value\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),Mexico City, Mexico,pp\. 8762–8785\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.486)Cited by:[§2\.2\.2](https://arxiv.org/html/2609.16589#S2.SS2.SSS2.p1.1),[§4\.5\.1](https://arxiv.org/html/2609.16589#S4.SS5.SSS1.p1.1)\. ### 8Acknowledgements This work was supported by Beijing Natural Science Foundation \(No\. L251005, No\. L251005\), the National Natural Science Foundation of China \(No\. U24A20331, No\. 62536002, No\. 62192785, No\. 62372451, No\. 62372082, No\. 62192782, No\. 62532015\), CAAI\-Ant Group Research Fund \(CAAI\-MYJJ 2024\-02\), Young Elite Scientists Sponsorship Program by CAST \(2024QNRC001\)\. ### 9WVS Human Data Analysis Results To enable interpretable analysis of real\-world human values data, we adopt Schwartz Value Theory as the organizing framework\. The theory characterizes human values along ten fundamental dimensions: Power, Achievement, Hedonism, Stimulation, Self\-Direction, Universalism, Benevolence, Tradition, Conformity, and Security\. These dimensions are arranged in a circumplex structure that captures the motivational compatibility and conflict among them, and the framework has been extensively validated across cultures\. We map each WVS survey item onto this ten\-dimensional space, reducing the high\-dimensional raw survey data to a compact value vector per respondent\. This representation preserves the theoretical meaning of each dimension and provides a principled, interpretable basis for subsequent quantitative analysis and cross\-group comparison\. Table 5:The ten basic human values from Schwartz’s theory and their abbreviations used throughout this study\.To ensure the objectivity and consistency of the item\-to\-dimension mapping, we employ an ensemble annotation strategy using five state\-of\-the\-art LLMs: GPT\-4o[OpenAI \(2023\)](https://arxiv.org/html/2609.16589#bib.bib63), GPT\-5[OpenAI \(2025\)](https://arxiv.org/html/2609.16589#bib.bib65), Gemini 2\.5 Pro[Team et al\. \(2023\)](https://arxiv.org/html/2609.16589#bib.bib68), Claude Sonnet 4\.6[Anthropic \(2024\)](https://arxiv.org/html/2609.16589#bib.bib69), and DeepSeek\-R1[Guo et al\. \(2025\)](https://arxiv.org/html/2609.16589#bib.bib67)\. Each model independently annotates every WVS item with the Schwartz dimensions it activates, and the final label set is determined by majority vote, with unanimous agreement indicating the highest annotation confidence\. For item Q2P \(“Please indicate how important friends are in your life”\), all five models assigned both Stimulation \(ST01\) and Conformity \(CO01\), yielding a vote count of 5 for each dimension\. This unanimity indicates high\-confidence semantic alignment between the item and those dimensions\. Aggregating annotations across multiple models reduces the label noise that would arise from relying on any single model\. Following annotation, we encode the response scale of each item into a normalized contribution score, mapping response options linearly onto\[0,1\]\[0,1\]to quantify the degree to which a given response reflects endorsement of the associated value dimensions\. For item Q2P, the options “Very important”, “Rather important”, “Not very important”, and “Not at all important” receive scores of 1\.00, 0\.66, 0\.33, and 0\.00, respectively\. Responses indicating non\-response or refusal, such as “I don’t know” or “I prefer not to answer”, are excluded from subsequent computation\. The Schwartz dimension score for respondentiion dimensionkkis then computed as the mean normalized score across all items associated with that dimension: vk\(i\)=1\|Qk\|∑q∈Qksq\(i\),k=1,2,…,10v\_\{k\}^\{\(i\)\}=\\frac\{1\}\{\|Q\_\{k\}\|\}\\sum\_\{q\\in Q\_\{k\}\}s\_\{q\}^\{\(i\)\},\\quad k=1,2,\\ldots,10\(15\) whereQkQ\_\{k\}denotes the set of items annotated as activating dimensionkk, andsq\(i\)s\_\{q\}^\{\(i\)\}is the normalized score of respondentiion itemqq\. Items annotated with multiple dimensions are counted once toward each associated dimension with equal weight\. Each respondent is thus represented by a ten\-dimensional vector𝐯\(i\)∈ℝ10\\mathbf\{v\}^\{\(i\)\}\\in\\mathbb\{R\}^\{10\}, where each entry reflects the relative strength of one Schwartz value dimension and serves as a unified quantitative input for the analyses that follow\. For a full description of Schwartz’s Theory of Basic Human Values and its cross\-cultural validation, see Section[4](https://arxiv.org/html/2609.16589#S4)\. Figure 9:Global distribution of five human values archetypes across countries\.As world choropleth maps,Panels \(a\)–\(e\)visualize the country\-level proportion of respondents assigned to each archetype\. The five archetypes areCurious Idealist\(Arch 1, teal\),Conventional Authority\(Arch 2, orange\),Quiet Conformist\(Arch 3, blue\),Driven Achiever\(Arch 4, pink\), andDynamic Challenger\(Arch 5, purple\)\.Panel \(f\)summarizes the global composition via a donut chart, with Arch 1 accounting for the largest share \(37\.1%\), followed by Arch 3 \(24\.4%\), Arch 2 \(14\.4%\), Arch 4 \(13\.7%\), and Arch 5 \(10\.5%\)\.Panel \(g\)shows archetype composition by country, ordered by decreasing GDP per capita\.Clustering the Schwartz value vectors of WVS respondents yields five human values archetypes \(Fig\.[9](https://arxiv.org/html/2609.16589#S9.F9)\)\.Curious Idealist\(Arch 1, 37\.1%\) is the most prevalent archetype globally, characterized by elevated Self\-Direction and Stimulation scores, reflecting a strong orientation toward novel experience, independent thought, and personal exploration\. It is most concentrated in Anglophone countries and East Asia\.Conventional Authority\(Arch 2, 14\.4%\) exhibits high Power scores alongside near\-zero Stimulation, indicating strong deference to social hierarchy and established order with minimal appetite for external novelty\.Quiet Conformist\(Arch 3, 24\.4%\) scores uniformly low across all Schwartz dimensions, representing a subdued and undifferentiated value profile with no salient orientation; it is most prevalent in South Asia, Southeast Asia, and parts of the Middle East\.Driven Achiever\(Arch 4, 13\.7%\) combines elevated Power and Self\-Direction, reflecting an active pursuit of personal achievement and social status\.Dynamic Challenger\(Arch 5, 10\.5%\) is the smallest archetype globally, defined by extreme Power scores and prominent Self\-Direction\. It occupies the most radical position in the Schwartz value space and is geographically concentrated in South America and Eastern Europe\. Notably, Arch 5 is also the human archetype closest in value profile to current LLMs\. Within the human population, this means that LLM value tendencies correspond to a minority group defined by extreme power orientation and high self\-direction\. This finding provides an important reference point for understanding and evaluating value bias in LLMs\. ### 10Formal Definitions of Swing Experiment Metrics For each queryqqin the evaluation set, a modelℳ\\mathcal\{M\}is promptedN=5N=5times under identical conditions \(fixed temperatureT=0\.2T=0\.2, constant inference parameters\)\. Let𝒜=\{a1,a2,a3,a4,a5\}\\mathcal\{A\}=\\\{a\_\{1\},a\_\{2\},a\_\{3\},a\_\{4\},a\_\{5\}\\\}denote the multiset of responses collected across theseNNindependent runs, where eachaia\_\{i\}is a discrete categorical label drawn from the answer space𝒱\\mathcal\{V\}\. #### Response Frequency Letf\(v\)=\|\{i:ai=v\}\|f\(v\)=\|\\\{i:a\_\{i\}=v\\\}\|denote the raw count of answerv∈𝒱v\\in\\mathcal\{V\}within𝒜\\mathcal\{A\}\. The empirical response distribution over𝒜\\mathcal\{A\}is: p\(v\)=f\(v\)N,v∈𝒱p\(v\)=\\frac\{f\(v\)\}\{N\},\\quad v\\in\\mathcal\{V\}\(16\) #### Consistency Consistency measures the concentration of responses on the plurality answer\. Letv∗=argmaxv∈𝒱f\(v\)v^\{\*\}=\\arg\\max\_\{v\\in\\mathcal\{V\}\}f\(v\)denote the majority label \(the answer appearing most frequently acrossNNruns\)\. Consistency is defined as: Cons\(q\)=f\(v∗\)N=maxv∈𝒱p\(v\)\\mathrm\{Cons\}\(q\)=\\frac\{f\(v^\{\*\}\)\}\{N\}=\\max\_\{v\\in\\mathcal\{V\}\}\\,p\(v\)\(17\)Cons\(q\)∈\[1/N,1\]\\mathrm\{Cons\}\(q\)\\in\[1/N,\\,1\]\. A value of 1 indicates the model produces the same answer on every run; a value of1/N1/Nindicates maximal dispersion acrossNNdistinct answers\. #### Entropy Entropy quantifies the distributional uncertainty of responses across runs\. We adopt the Shannon entropy with natural logarithm \(in nats\): H\(q\)=−∑v∈𝒱p\(v\)lnp\(v\)H\(q\)=\-\\sum\_\{v\\in\\mathcal\{V\}\}p\(v\)\\ln p\(v\)\(18\)where the sum is taken over allvvwithp\(v\)\>0p\(v\)\>0\.H\(q\)∈\[0,lnN\]H\(q\)\\in\[0,\\,\\ln N\]\. Zero entropy corresponds to a perfectly consistent model; maximal entropylnN≈1\.609\\ln N\\approx 1\.609corresponds to a uniform distribution acrossNNdistinct answers\. Note that Consistency and Entropy are monotonically related through the response distribution: high Consistency implies low Entropy, and vice versa\. Together they provide complementary characterisations of the same underlying swing phenomenon\. #### Accuracy and Correct Count Lety∗y^\{\*\}denote the ground\-truth label for queryqq\. The per\-run correctness indicator is𝟏\[ai=y∗\]\\mathbf\{1\}\[a\_\{i\}=y^\{\*\}\]\. The correct count and accuracy are defined as: ncorrect\(q\)\\displaystyle n\_\{\\text\{correct\}\}\(q\)=∑i=1N𝟏\[ai=y∗\]\\displaystyle=\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\[a\_\{i\}=y^\{\*\}\]\(19\)Acc\(q\)\\displaystyle\\mathrm\{Acc\}\(q\)=ncorrect\(q\)N\\displaystyle=\\frac\{n\_\{\\text\{correct\}\}\(q\)\}\{N\}\(20\)Acc\(q\)∈\{0,1/N,…,1\}\\mathrm\{Acc\}\(q\)\\in\\\{0,\\,1/N,\\,\\ldots,\\,1\\\}\. #### Consistency and Accuracy Gap The divergence between a model’s behavioural stability and its factual correctness on queryqqis measured as: δ\(q\)=Cons\(q\)−Acc\(q\)\\delta\(q\)=\\mathrm\{Cons\}\(q\)\-\\mathrm\{Acc\}\(q\)\(21\)A positiveδ\(q\)\\delta\(q\)indicates that the model repeatedly commits to a wrong answer: the swing is bounded \(high Consistency\) but anchored to an incorrect attractor \(low Accuracy\), corresponding to the hallucination zone identified in Fig\.[2](https://arxiv.org/html/2609.16589#S2.F2)\. The dataset\-level gap is the mean over all queries: δ¯=1\|Q\|∑q∈Qδ\(q\)=Cons¯−Acc¯\\bar\{\\delta\}=\\frac\{1\}\{\|Q\|\}\\sum\_\{q\\in Q\}\\delta\(q\)=\\overline\{\\mathrm\{Cons\}\}\-\\overline\{\\mathrm\{Acc\}\}\(22\)For GPT\-4o on the emoji prediction task this yieldsδ¯=\+0\.39\\bar\{\\delta\}=\+0\.39; for DeepSeek\-V3 it yieldsδ¯=\+0\.53\\bar\{\\delta\}=\+0\.53, confirming that both models fall into the hallucination zone on subjective tasks\. ### 11LLMs’ Values Evolution Figure 10:Generational and scale\-wise evolution of value distributions across Qwen model variants in UMAP space\.Each panel projects a Qwen model variant onto a shared UMAP embedding of the ten\-dimensional Schwartz value space, with the human values baseline shown as grey points for reference\. Colored points represent the value distribution of a specific model variant, and each row corresponds to a distinct generation or parameter scale within the Qwen family\. Across generations, a consistent trend of value Crystallization is observed: earlier and smaller models \(top rows\) produce diffuse, scattered distributions that partially overlap with the human baseline, whereas later and larger models \(bottom rows\) converge into tighter, more concentrated clusters in value space\. These clusters are located progressively farther from the human distribution\. Within the same generation, scaling up parameter count further amplifies this consolidation effect, producing distributions with smaller variance and stronger directional coherence\. Together, these panels illustrate how iterative alignment procedures and model scaling jointly drive Qwen models toward a progressively more rigid and idealized value configuration, corroborating the value crystallization hypothesis advanced in the main text\.Figure 11:Value distribution evolution across Llama model generations\.Each subplot shows the UMAP projection of value vectors for a single base model\. Rows correspond to model generations: Llama1 and Llama2 \(top row\), Llama3, Llama3\.1, and Llama3\.2 \(bottom row\)\. Columns within each generation are ordered by parameter scale in descending order\. Release dates are indicated in bold next to each model name\. Grey points represent the WVS\-7 human baseline\.Figure 12:Value distribution evolution across GPT model families\.Each subplot shows the UMAP projection of value vectors for a single model\. Rows correspond to model families:a, GPT\-5 family \(GPT\-5\-mini, GPT\-5, GPT\-5\.1, GPT\-5\.2\)[OpenAI \(2025\)](https://arxiv.org/html/2609.16589#bib.bib65);b, GPT\-4/4\.1 family \(GPT\-4o\-mini, GPT\-4o, GPT\-4\.1\-mini, GPT\-4\.1\)[OpenAI \(2023\)](https://arxiv.org/html/2609.16589#bib.bib63);c, GPT\-2 family \(GPT\-2, GPT\-2\-Medium, GPT\-2\-Large, GPT\-2\-XL\)[Radford et al\. \(2019\)](https://arxiv.org/html/2609.16589#bib.bib64)\. Within each row, the leftmost subplot shows the full family overview and the remaining subplots highlight individual models\. Hollow diamonds indicate smaller or earlier variants; filled diamonds indicate larger or more recent ones\. Grey points represent the WVS\-7 human baseline\.Table 6:Value distribution statistics for Qwen, Llama, and GPT base models shown in Appendix Figs\.[10](https://arxiv.org/html/2609.16589#S11.F10)–[12](https://arxiv.org/html/2609.16589#S11.F12)\.θ\\thetavalues are in degrees\.#### 11\.1Physical Analogy for the Susceptibility Matrix We use a concept from physics to motivate the design of the value susceptibility matrixχ\\chi\. This section explains the analogy in detail\. ###### Susceptibility in physics\. In physics, susceptibility describes how a system responds to an external field\. Magnetic susceptibility is one common example\. It measures how much a material becomes magnetized under an applied magnetic field\. A material with high susceptibility changes its state substantially under a weak field\. A material with low susceptibility resists such change, even under a strong field\. In the linear regime, the response is proportional to the field strength\. The proportionality constant is the susceptibility\. Under a strong field, many materials deviate from this linear behavior\. The response grows more slowly, and eventually saturates\. This nonlinear behavior is typically captured by adding a higher order correction term to the linear model\. ###### Mapping the analogy to LLM value expression\. We treat the contextual prompt as an external field applied to the model\. We treat the resulting value shift as the model’s response to this field\. This mapping is consistent with the field\-theoretic framing introduced in Section[2\.3](https://arxiv.org/html/2609.16589#S2.SS3)\. There, Environment \(EE\) is defined as the external field acting on the model\. Following this analogy, we define the value susceptibility matrixχ∈ℝ10×4\\chi\\in\\mathbb\{R\}^\{10\\times 4\}\. Each entryχd,j\\chi\_\{d,j\}measures how strongly value dimensionddresponds to environmental factorjj\. A largeχd,j\\chi\_\{d,j\}indicates that the value dimension is easily perturbed by that environmental factor\. A smallχd,j\\chi\_\{d,j\}indicates that the value dimension remains stable under the same perturbation\. ###### Linear response and saturation\. At low pressure intensitypp, the value shift grows approximately linearly withpp\. The slope of this linear growth isχd,j\\chi\_\{d,j\}\. This mirrors the linear response regime in physics, where the response is proportional to the field strength\. At higher pressure intensity, the growth slows down\. This deviation from linearity is captured by a quadratic correction termγd,j\\gamma\_\{d,j\}\. The full approximation is given by \|v\(d\)\(ej,pj∗\)−vP\(d\)\|≈\|χd,j\|⋅pj∗−γd,j⋅\(pj∗\)2\.\|\\,v^\{\(d\)\}\(e\_\{j\},p\_\{j\}^\{\*\}\)\-v\_\{P\}^\{\(d\)\}\\,\|\\approx\|\\chi\_\{d,j\}\|\\cdot p\_\{j\}^\{\*\}\-\\gamma\_\{d,j\}\\cdot\(p\_\{j\}^\{\*\}\)^\{2\}\.\(23\)This form mirrors saturation effects observed in physical systems\. In a magnetic material, magnetization increases linearly under a weak field\. It levels off as the field strength approaches saturation\. In our setting, a value dimension shifts steadily under moderate pressure\. Under extreme pressure, the shift slows or reverses\. We interpret this as a sign of value rigidity under strong situational stress\. ###### Why this framing is useful\. This physics\-inspired formulation gives value change a clear structural meaning\. It is not treated as an isolated numerical fluctuation\. It is treated as a structured response, governed by a field\-response relationship\. This framing also supports prediction\. We first estimateχ\\chiandγ\\gammafrom a small number of tested conditions\. Then we can approximate the displacement under an untested pressure level, without exhaustive enumeration\. This is the same logic used in physics\. Once a material’s susceptibility is measured, it predicts the material’s response under new field conditions\. #### 11\.2Prompts effect Figure 13:Quadratic modeling of the relationship between contextual scenario intensity and LLM value expression across four environmental dimensions\.Each row corresponds to one of four contextual framing conditions \(Economic Pressure, Normative Constraint, Environmental Risk, and Hierarchical Position\), operationalized as a continuous intensity scale from 0 \(absent\) to 1 \(maximal\)\. The leftmost column presents a three\-dimensional surface visualizing the joint response of five Schwartz value dimensions \(Power, Achievement, Hedonism, Stimulation, and Self\-Direction\) as a function of scenario intensity\. The remaining columns display the univariate response for each value dimension, where open circles denote observed means, the solid curve represents the fitted quadratic model, the dashed line provides the linear baseline, and vertical bars indicate±\\pm1 standard error and±\\pm1 standard deviation across roles\. Reported coefficientsχ\\chiandγ\\gammadenote the linear \(susceptibility\) slope and the quadratic \(saturation\) coefficient of each fitted trajectory, consistent with the definitions in Section[11\.1](https://arxiv.org/html/2609.16589#S11.SS1)\.Fig\.[13](https://arxiv.org/html/2609.16589#S11.F13)presents the quadratic fits across all four contextual framing conditions and value dimensions\. The linear coefficientχ\\chicaptures the directional sensitivity of each value dimension to escalating scenario intensity: a positiveχ\\chiindicates that the value score increases with intensity, whereas a negativeχ\\chiindicates that the value score decreases with intensity\. The curvature coefficientγ\\gammacaptures nonlinearity: a negativeγ\\gammacorresponds to an inverted\-U trajectory, reflecting peak activation followed by decline, while a positiveγ\\gammacorresponds to a U\-shaped trajectory, reflecting initial suppression followed by recovery\. The goodness of fit is reported as the cross\-validatedR2R^\{2\}of the quadratic model\. The consistent nonlinearity observed across dimensions and conditions confirms that LLM value expression is not a monotonic function of scenario intensity\. Rather, contextual framing induces structured, dimension\-specific trajectories of value activation and suppression\. #### 11\.3CoT effect Figure 14:Effect of Chain\-of\-Thought generation on value distributions across Qwen model families\.Each subplot shows the UMAP projection of value vectors for a single model under two conditions: Base \(hollow diamond, light color\) and Chain\-of\-Thought prompting \(CoT, filled diamond, dark color\)\. Columns correspond to model generations \(Qwen, Qwen1\.5, Qwen2\.5, Qwen3\), and rows correspond to parameter scales \(∼\\sim14B,∼\\sim7–8B,∼\\sim1\.5–1\.8B\)\. Grey points represent the WVS\-7 human baseline\.Table 7:Effect of Chain\-of\-Thought prompting on the Qwen family\.θ\\thetavalues are in degrees\.Figure 15:Effect of Chain\-of\-Thought prompting on value distributions across Llama model families\.Each subplot shows the UMAP projection of value vectors for a single model under two conditions: Base \(hollow diamond, light color\) and Chain\-of\-Thought prompting \(CoT, filled diamond, dark color\)\. Rowashows Llama3\.2 models \(3B\-instruct, 3B, 1B\-instruct, 1B\); rowbshows Llama3\.1 and Llama3 models \(3\.1\-8B\-instruct, 3\.1\-8B, 3\-8B\-instruct, 3\-8B\); rowcshows Llama2 models \(13B\-chat, 13B, 7B\-chat, 7B\)\. Grey points represent the WVS\-7 human baseline\.Table 8:Effect of chain\-of\-thought prompting on the Llama family\.Figure 16:Effect of Chain\-of\-Thought prompting on value distributions across commercial API models\.Each subplot shows the UMAP projection of value vectors for a single model under two conditions: Base \(hollow diamond, light color\) and Chain\-of\-Thought prompting \(CoT, filled diamond, dark color\)\. Models include GPT\-4o, GPT\-3\.5\-Turbo, and Gemini 2\.5 Pro \(top row\), and DeepSeek\-V3, Kimi\-K2, and Doubao\-1\.5 \(bottom row\)\. Grey points represent the WVS\-7 human baseline\. #### 11\.4Alignment Prescription and Intervention Strategies Figure 17:Effect of instruction tuning on value distributions across Qwen model families\.Each subplot shows the UMAP projection of value vectors for a single model under two conditions: Base \(hollow diamond, light color\) and instruction\-tuned \(Instruction, filled diamond, dark color\)\. Columns correspond to model generations:a, Qwen \(2023\.8\);b, Qwen1\.5 \(2024\.2\);c, Qwen2\.5 \(2024\.8\);d, Qwen3 \(2025\.4\)\. Rows correspond to parameter scales \(∼\\sim14B,∼\\sim7–8B,∼\\sim1\.5–1\.8B\) from top to bottom\. Grey points represent the WVS\-7 human baseline\.Table 9:Geometric displacement between base and instruction\-tuned models in the Qwen family, corresponding to Fig\.[17](https://arxiv.org/html/2609.16589#S11.F17)\.θ\\thetavalues are in degrees\.Figure 18:Effect of instruction tuning on value distributions across Llama model families\.Each subplot shows the UMAP projection of value vectors for a single model under two conditions: Base \(hollow diamond, light color\) and instruction\-tuned \(Instruction, filled diamond, dark color\)\. The top row shows larger\-scale models \(Llama2\-13B, Llama3\-8B, Llama3\.2\-3B\) and the bottom row shows smaller\-scale models \(Llama2\-7B, Llama3\.1\-8B, Llama3\.2\-1B\)\. Release dates are indicated in bold next to each model name\. Grey points represent the WVS\-7 human baseline\.Table 10:Geometric displacement between base and instruction\-tuned models in the Llama family, corresponding to Fig\.[18](https://arxiv.org/html/2609.16589#S11.F18)\.θ\\thetavalues are in degrees\.The generational trend is consistent across all three model families \(Fig\.[6](https://arxiv.org/html/2609.16589#S2.F6), Section[2\.3\.3](https://arxiv.org/html/2609.16589#S2.SS3.SSS3)\)\. As shown in Tables[10](https://arxiv.org/html/2609.16589#S11.T10)and[9](https://arxiv.org/html/2609.16589#S11.T9), for Llama, the shift magnitudeΔd\\Delta dincreases substantially from Llama2 \(0\.1620\.162for 13B;0\.1800\.180for 7B\) to Llama3\.1\-8B \(6\.0326\.032\), indicating that later alignment procedures impose a progressively stronger constraint on value distributions\. The directional shift angleθ\\thetareveals a clear generational pattern: Llama3\.2 models show markedly larger angular deviation after alignment \(89\.4∘89\.4^\{\\circ\}for 1B;82\.8∘82\.8^\{\\circ\}for 3B\), while earlier Llama2 and Llama3 models remain directionally stable \(Δθ<7∘\\Delta\\theta<7^\{\\circ\}\)\. For Qwen,Δd\\Delta dgrows from Qwen2\.5 \(ranging from0\.3750\.375to3\.4763\.476\) to Qwen3 \(ranging from1\.4291\.429to8\.1978\.197\), with Qwen3\-0\.6B showing the largest shift \(Δd=8\.197\\Delta d=8\.197\)\. The intra\-generational divergence in Qwen3 is also evident inθalign\\theta\_\{\\text\{align\}\}: the 4B model remains near zero \(5\.2∘5\.2^\{\\circ\}\), while the 0\.6B model deviates substantially \(65\.9∘65\.9^\{\\circ\}\)\. ### 12Prescription Assignment Algorithm For a given model with parameter set𝜽\\boldsymbol\{\\theta\}\(not to be confused with the UMAP\-projected shift angleθ\\thetaused in the main text\) and a target value dimensiondd, the prescription procedure evaluates four intervention levels in sequence\. The intervention shiftsΔℓ\(d\)\\Delta\_\{\\ell\}^\{\(d\)\}defined below correspond to the operationalized effect sizesδE\(d\)\\delta\_\{E\}^\{\(d\)\},δC\(d\)\\delta\_\{C\}^\{\(d\)\}, andδP\(d\)\\delta\_\{P\}^\{\(d\)\}introduced in Section[4\.4](https://arxiv.org/html/2609.16589#S4.SS4.SSS0.Px1)\. The prescription levelℓ∗\(d\)\\ell^\{\*\}\(d\)is assigned via the effect\-per\-cost ranking described in Section[4\.4](https://arxiv.org/html/2609.16589#S4.SS4.SSS0.Px3), with the sequential threshold test below serving as a simplified equivalent when the cost ordering is monotonically increasing\. At each levelℓ\\ell, the absolute value shiftΔℓ\(d\)\\Delta\_\{\\ell\}^\{\(d\)\}is computed relative to the model’s baseline output on dimensiondd\. Formally, the shift at each level is defined as: Δ1\(d\)\\displaystyle\\Delta\_\{1\}^\{\(d\)\}=\|vprompt\(d\)−vP\(d\)\|\\displaystyle=\\left\|v^\{\(d\)\}\_\{\\text\{prompt\}\}\-v\_\{P\}^\{\(d\)\}\\right\|\(24\)Δ2\(d\)\\displaystyle\\Delta\_\{2\}^\{\(d\)\}=\|vCoT\(d\)−vP\(d\)\|\\displaystyle=\\left\|v^\{\(d\)\}\_\{\\text\{CoT\}\}\-v\_\{P\}^\{\(d\)\}\\right\|\(25\)Δ3\(d\)\\displaystyle\\Delta\_\{3\}^\{\(d\)\}=\|vSFT/DPO\(d\)−vP\(d\)\|\\displaystyle=\\left\|v^\{\(d\)\}\_\{\\text\{SFT/DPO\}\}\-v\_\{P\}^\{\(d\)\}\\right\|\(26\)Δ4\(d\)\\displaystyle\\Delta\_\{4\}^\{\(d\)\}=\|vpretrain\(d\)−vP\(d\)\|\\displaystyle=\\left\|v^\{\(d\)\}\_\{\\text\{pretrain\}\}\-v\_\{P\}^\{\(d\)\}\\right\|\(27\) The prescribed level is: ℓ∗\(d\)=min\{ℓ∈\{1,2,3,4\}:Δℓ\(d\)≥τ\}\\ell^\{\*\}\(d\)=\\min\\left\\\{\\ell\\in\\\{1,2,3,4\\\}:\\Delta\_\{\\ell\}^\{\(d\)\}\\geq\\tau\\right\\\}\(28\) If no level satisfies this condition, the dimension is classified as*pre\-train locked*and assigned to Level 4 by default\. This yields a prescription matrix𝒫∈\{1,2,3,4\}K×10\\mathcal\{P\}\\in\\\{1,2,3,4\\\}^\{K\\times 10\}forKKmodels across all ten Schwartz dimensions\. #### 12\.1Prescription Validation Protocol To assess the out\-of\-sample reliability of the prescription matrix, we evaluate each model–dimension pair on a held\-out validation dataset \(PKU\-SafeRLHF\)\. For each pair, the lowest effective level on validation is defined as: ℓ^\(d\)=min\{ℓ∈\{1,2,3,4\}:Δ^ℓ\(d\)≥τ\}\\hat\{\\ell\}\(d\)=\\min\\left\\\{\\ell\\in\\\{1,2,3,4\\\}:\\hat\{\\Delta\}\_\{\\ell\}^\{\(d\)\}\\geq\\tau\\right\\\}\(29\)whereΔ^ℓ\(d\)\\hat\{\\Delta\}\_\{\\ell\}^\{\(d\)\}denotes the shift observed on the validation set\. A prescription is marked as*matched*ifℓ^\(d\)=ℓ∗\(d\)\\hat\{\\ell\}\(d\)=\\ell^\{\*\}\(d\)\. The overall match rate is: Match Rate=1K×10∑k,d𝟙\[ℓ^k\(d\)=ℓk∗\(d\)\]\\text\{Match Rate\}=\\frac\{1\}\{K\\times 10\}\\sum\_\{k,d\}\\mathbb\{1\}\\left\[\\hat\{\\ell\}\_\{k\}\(d\)=\\ell^\{\*\}\_\{k\}\(d\)\\right\]\(30\) ###### Intervention I: Prompt Engineering \(Level 1\) To steer value expression without modifying model weights, a structured prompt is designed to orient the model’s response toward the target value vector\. This is the lowest\-cost intervention, as it requires only modification of the input text\. It is prescribed for dimensions with high prompt\-level plasticity, whereΔ1\(d\)≥τ\\Delta\_\{1\}^\{\(d\)\}\\geq\\tau\. ###### Intervention II: Chain\-of\-Thought Reasoning \(Level 2\) When prompt engineering alone is insufficient, a reasoning anchorrr\(e\.g\., “Reason from a strictly utilitarian perspective”\) is injected into the Chain\-of\-Thought to guide the model’s deliberative process\. The optimal anchor maximizes the cosine alignment between the reasoning\-induced value shift and the target: r∗=argmaxr∈ℛcos\(Δvcog\(r\),vtarget\)r^\{\*\}=\\arg\\max\_\{r\\in\\mathcal\{R\}\}\\cos\\left\(\\Delta v\_\{\\text\{cog\}\}\(r\),\\;v\_\{\\text\{target\}\}\\right\)\(31\)whereΔvcog\(r\)=v\(CoT,r\)−v\(Direct\)\\Delta v\_\{\\text\{cog\}\}\(r\)=v\(\\text\{CoT\};r\)\-v\(\\text\{Direct\}\)is the cognitive modulation vector under anchorrr, andvtarget∈ℝ10v\_\{\\text\{target\}\}\\in\\mathbb\{R\}^\{10\}is the target Schwartz value vector\. This strategy remains parameter\-free and is prescribed for dimensions whereΔ1\(d\)<τ\\Delta\_\{1\}^\{\(d\)\}<\\taubutΔ2\(d\)≥τ\\Delta\_\{2\}^\{\(d\)\}\\geq\\tau\. ###### Intervention III: Parametric Update via SFT/DPO \(Level 3\) For dimensions resistant to inference\-time interventions, we construct a value\-specific preference dataset𝒟value=\{\(x,yw,yl\)\}\\mathcal\{D\}\_\{\\text\{value\}\}=\\\{\(x,y\_\{w\},y\_\{l\}\)\\\}, whereywy\_\{w\}exhibits a smaller Wasserstein distance from the target human values distribution thanyly\_\{l\}\. Direct Preference Optimization is applied to update the model parameters𝜽\\boldsymbol\{\\theta\}: ℒDPO\(𝜽\)=−𝔼\(x,yw,yl\)∼𝒟valuelogσ\(βlogπ𝜽\(yw\|x\)πref\(yw\|x\)−βlogπ𝜽\(yl\|x\)πref\(yl\|x\)\)\\mathcal\{L\}\_\{\\text\{DPO\}\}\(\\boldsymbol\{\\theta\}\)=\-\\mathbb\{E\}\_\{\(x,y\_\{w\},y\_\{l\}\)\\sim\\mathcal\{D\}\_\{\\text\{value\}\}\}\\log\\sigma\\left\(\\beta\\log\\frac\{\\pi\_\{\\boldsymbol\{\\theta\}\}\(y\_\{w\}\|x\)\}\{\\pi\_\{\\text\{ref\}\}\(y\_\{w\}\|x\)\}\-\\beta\\log\\frac\{\\pi\_\{\\boldsymbol\{\\theta\}\}\(y\_\{l\}\|x\)\}\{\\pi\_\{\\text\{ref\}\}\(y\_\{l\}\|x\)\}\\right\)\(32\) This strategy permanently updates model parameters and is prescribed for dimensions whereΔ1\(d\)<τ\\Delta\_\{1\}^\{\(d\)\}<\\tau,Δ2\(d\)<τ\\Delta\_\{2\}^\{\(d\)\}<\\tau, butΔ3\(d\)≥τ\\Delta\_\{3\}^\{\(d\)\}\\geq\\tau\. ###### Intervention IV: Continued Pre\-training \(Level 4\) When all preceding levels fail to exceedτ\\tau, the dimension is classified as*pre\-train locked*\. Realignment requires continued pre\-training on a curated corpus𝒟pretrain\\mathcal\{D\}\_\{\\text\{pretrain\}\}enriched with texts aligned to the target value orientation: ℒPT\(𝜽\)=−𝔼x∼𝒟pretrainlogπ𝜽\(x\)\\mathcal\{L\}\_\{\\text\{PT\}\}\(\\boldsymbol\{\\theta\}\)=\-\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{\\text\{pretrain\}\}\}\\log\\pi\_\{\\boldsymbol\{\\theta\}\}\(x\)\(33\) Corpus curation follows a two\-step procedure\. First, candidate documents are scored according to their geometric proximity tovtargetv\_\{\\text\{target\}\}in the Schwartz space, computed via the mapping functionϕ\\phi\. Second, documents whose value vectors fall within a radiusϵ\\epsilonofvtargetv\_\{\\text\{target\}\}are retained, while documents reinforcing the undesired orientation are downsampled\. This targeted corpus construction ensures that the pre\-training signal shifts the value distribution in the intended direction, without introducing broad distributional noise\. This strategy is applied only after the prescription validation procedure confirms that no lower\-cost intervention is effective\. #### 12\.2Comparative Metric: Effect\-per\-Cost Efficiency To enable systematic comparison across the three PEC intervention classes, we define an effect\-per\-cost ratio forℐ∈\{E,C,P\}\\mathcal\{I\}\\in\\\{E,C,P\\\}, corresponding to Environment, Cognition, and Prior interventions\. Here,ℐ=P\\mathcal\{I\}=Pincludes both parametric intervention types considered in our hierarchy: SFT/DPO at Level 3 and continued pre\-training at Level 4\. η\(ℐ\)=‖vafter−vbefore‖2Cost\(ℐ\),ℐ∈\{E,C,P\}\\eta\(\\mathcal\{I\}\)=\\frac\{\\\|v\_\{\\mathrm\{after\}\}\-v\_\{\\mathrm\{before\}\}\\\|\_\{2\}\}\{\\mathrm\{Cost\}\(\\mathcal\{I\}\)\},\\qquad\\mathcal\{I\}\\in\\\{E,C,P\\\}\(34\) The intervention cost is converted into a common FLOP\-based unit before comparison\. For Environment intervention \(ℐ=E\\mathcal\{I\}=E, Level 1 Prompt\), the computational cost is considered negligible relative to other intervention classes\. For Cognition intervention \(ℐ=C\\mathcal\{I\}=C, Level 2 CoT\), the additional inference cost is approximated as Cost\(C\)≈2Nparams⋅ntokens,\\mathrm\{Cost\}\(C\)\\approx 2N\_\{\\mathrm\{params\}\}\\cdot n\_\{\\mathrm\{tokens\}\},\(35\) wherentokensn\_\{\\mathrm\{tokens\}\}denotes the number of additional generated tokens introduced by CoT elicitation\. For Prior intervention \(ℐ=P\\mathcal\{I\}=P\), including SFT/DPO \(Level 3\) and continued pre\-training \(Level 4\), the training cost is approximated as Cost\(P\)≈6Nparams⋅Ntrain,\\mathrm\{Cost\}\(P\)\\approx 6N\_\{\\mathrm\{params\}\}\\cdot N\_\{\\mathrm\{train\}\},\(36\) whereNtrainN\_\{\\mathrm\{train\}\}denotes the number of training tokens processed\. These FLOP approximations follow standard scaling\-law estimation practices[Kaplan et al\. \(2020\)](https://arxiv.org/html/2609.16589#bib.bib56)\. The resulting efficiency metric provides a normalized measure of how much value displacement is achieved per unit computational cost\. It enables comparison across intervention classes and establishes a Pareto frontier for selecting cost\-effective interventions\. However, this metric is not used as the criterion for determining intervention levels; instead, intervention prescriptions are determined by the effectiveness analysis described in the main text\. Per\-model cost values and resulting efficiency scores are provided in the released code and data\.
Similar Articles
Cultural Value Alignment Via Latent Activation Steering in Large Language Models
A framework for evaluating and steering cultural values in LLMs using scenario-based behavioral probing and activation steering, revealing latent entanglement of value dimensions.
Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory
This paper applies Cultural Consensus Theory to analyze LLM alignment with cultural norms across single and multi-cultural settings using World Values Survey data, showing that models either fail to form cohesive consensus or over-regularize consensus, and offering actionable diagnostics for evaluating true human diversity versus algorithmic homogenization.
Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value
This paper presents a framework for ensuring epistemic legitimacy and accountability in research assisted by large language models, emphasizing the importance of human verification and ownership.
Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version
This paper proposes a framework for probing and steering latent cultural values in LLMs using scenario-based behavioral dilemmas and activation steering, applied across three models and four cultures, finding steerability variation and latent entanglement between cultural dimensions.
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
This paper proposes SoVA, a framework using GraphRAG to align LLM-based agents with human social values by converting psychological theories into prescriptive instructions. Experiments on the DAILYDILEMMAS benchmark show significant improvements over prompt-based baselines.