Self-reported archetypes and behavioral failures in Large Language Models

arXiv cs.CL Papers

Summary

This paper investigates self-reported archetypes and behavioral failures in Large Language Models, exploring how LLMs perceive human archetypes and where they exhibit behavioral shortcomings.

arXiv:2609.15998v1 Announce Type: new Abstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property of training, these systems exhibit persistent dispositions that shape how they interact, comply, resist, and err, yet the structure of LLM character remains poorly understood. We map the self-reported personality archetypes of 22 LLMs spanning closed-source frontier systems (GPT-4.0-5.2, Grok-3/4, Gemini 2.5 Pro/Flash, Claude Sonnet 4.5/4.6) and open-source models (Llama, DeepSeek, OLMo, and Qwen series). Each model self-rated across 464 bipolar semantic-differential trait pairs, and the resulting profiles were projected into a six-dimensional archetypal space derived from crowd-sourced ratings of 2,000 fictional characters using the Archetypometrics framework. Closed-source models' self-rating traits align with the empirical trait co-occurrence structure of human-rated fictional characters, suggesting coherent, human-like self-representations organized around combinations of four recurring archetypal dimensions: Hero, Angel, Traditionalist, and Geek. Their closest analogues include Data, Vision, and Janet. Open-source models show weaker, noisier, and internally contradictory self-representations, occupying a diffuse region of archetype space with weak structure. Cross-referencing self-reported profiles with developer constitutions reveals a consequential gap between claimed character and enacted behavior: hallucination undermines claimed precision, sycophancy complicates claimed kindness, and agentic failures contradict claimed obedience. These self-ratings should therefore be interpreted not as neutral measurements of model character, but as structured outputs of the same optimization processes that shape model behavior. This work provides a reproducible, character-grounded framework for evaluating what LLMs are, not just what they do.
Original Article
View Cached Full Text

Cached at: 09/16/26, 08:38 AM

# Untitled Document
Source: [https://arxiv.org/html/2609.15998](https://arxiv.org/html/2609.15998)
\\DocumentMetadata

uncompress\\setbooleantwocolswitchtrue

###### Contents

1. [References](https://arxiv.org/html/2609.15998#bib)
2. [AAppendices](https://arxiv.org/html/2609.15998#A1)

## References

- \[1\]H\. Zhao, Z\. Liu, Z\. Wu, Y\. Li, T\. Yang, P\. Shu, S\. Xu, H\. Dai, L\. Zhao, G\. Mai, N\. Liu, and T\. Liu\.Revolutionizing finance with LLMs: An overview of applications and insights\.ArXiv, abs/2401\.11641, 2024\.
- \[2\]X\. Meng, X\. Yan, K\. Zhang, D\. Liu, X\. Cui, Y\. Yang, M\. Zhang, C\. Cao, J\. Wang, X\. Wang, J\. Gao, Y\. Wang, J\. Ji, Z\. Qiu, M\. Li, C\. Qian, T\. Guo, S\. Ma, Z\. Wang, Z\. Guo, Y\.\-L\. Lei, C\. Shao, W\. yao Wang, H\. Fan, and Y\. Tang\.The application of large language models in medicine: A scoping review\.iScience, 27, 2024\.
- \[3\]T\. T\. Prama and M\. S\. Islam\.Evaluating credibility and political bias in LLMs for news outlets in bangladesh\.InAnnual Meeting of the Association for Computational Linguistics, 2025\.
- \[4\]T\. T\. Prama, C\. M\. Danforth, and P\. S\. Dodds\.Banglamath : A bangla benchmark dataset for testing llm mathematical reasoning at grades 6, 7, and 8\.ArXiv, abs/2510\.12836, 2025\.
- \[5\]K\. Sokol, J\. C\. Fackler, and J\. E\. Vogt\.Artificial intelligence should genuinely support clinical reasoning and decision making to bridge the translational gap\.NPJ Digital Medicine, 8, 2025\.
- \[6\]P\. K\. Donta, A\. Saleh, Y\. Li, S\. Vaishnav, K\. Fang, H\. Feng, Y\. Xia, T\. R\. Gadekallu, Q\. Zhang, X\. Shi, A\. Beikmohammadi, S\. Magn’usson, I\. Murturi, C\. K\. Dehury, M\. Paprzycki, L\. Lovén, S\. Tarkoma, and S\. Dustdar\.Socio\-technical aspects of agentic ai\.ArXiv, abs/2601\.06064, 2025\.
- \[7\]J\. K\. Miller and W\. Tang\.Evaluating LLM metrics through real\-world capabilities\.ArXiv, abs/2505\.08253, 2025\.
- \[8\]T\. R\. Mcintosh, T\. Sušnjak, N\. A\. G\. Arachchilage, T\. Liu, D\. Xu, P\. A\. Watters, and M\. N\. Halgamuge\.Inadequacies of large language model benchmarks in the era of generative artificial intelligence\.IEEE Transactions on Artificial Intelligence, 7:22–39, 2024\.
- \[9\]T\. T\. Prama, C\. M\. Danforth, and P\. S\. Dodds\.Computational story lab at blp\-2025 task 1: Hatesense: A multi\-task learning framework for comprehensive hate speech identification using LLMs\.Proceedings of the Second Workshop on Bangla Language Processing \(BLP\-2025\), 2025\.
- \[10\]T\. T\. Prama, C\. M\. Danforth, and P\. S\. Dodds\.LLMs for low\-resource dialect translation using context\-aware prompting: A case study on sylheti\.ArXiv, abs/2511\.21761, 2025\.
- \[11\]T\. T\. Prama, J\. W\. Zimmerman, C\. M\. Danforth, and P\. S\. Dodds\.Us\-vs\-them bias in large language models\.ArXiv, abs/2512\.13699, 2025\.
- \[12\]T\. T\. Prama, C\. M\. Danforth, and P\. S\. Dodds\.Misalignment of LLM\-generated personas with human perceptions in low\-resource settings\.ArXiv, abs/2512\.02058, 2025\.
- \[13\]P\. Liang, R\. Bommasani, T\. Lee, D\. Tsipras, D\. Soylu, M\. Yasunaga, Y\. Zhang, D\. Narayanan, Y\. Wu, A\. Kumar, B\. Newman, B\. Yuan, B\. Yan, C\. Zhang, C\. Cosgrove, C\. D\. Manning, C\. R’e, D\. Acosta\-Navas, D\. A\. Hudson, E\. Zelikman, E\. Durmus, F\. Ladhak, F\. Rong, H\. Ren, H\. Yao, J\. Wang, K\. Santhanam, L\. J\. Orr, L\. Zheng, M\. Yuksekgonul, M\. Suzgun, N\. S\. Kim, N\. Guha, N\. S\. Chatterji, O\. Khattab, P\. Henderson, Q\. Huang, R\. Chi, S\. M\. Xie, S\. Santurkar, S\. Ganguli, T\. Hashimoto, T\. F\. Icard, T\. Zhang, V\. Chaudhary, W\. Wang, X\. Li, Y\. Mai, Y\. Zhang, and Y\. Koreeda\.Holistic evaluation of language models\.Annals of the New York Academy of Sciences, 1525:140 – 146, 2023\.
- \[14\]A\. Srivastava, A\. Rastogi, A\. Rao, A\. A\. M\. Shoeb, A\. Abid, A\. Fisch, A\. R\. Brown, A\. Santoro, A\. Gupta, A\. Garriga\-Alonso, A\. Kluska, A\. Lewkowycz, A\. Agarwal, A\. Power, A\. Ray, A\. Warstadt, A\. W\. Kocurek, A\. Safaya, A\. Tazarv, A\. Xiang, A\. Parrish, A\. Nie, A\. Hussain, A\. Askell, A\. Dsouza, A\. Slone, A\. Rahane, A\. S\. Iyer, A\. Andreassen, A\. Madotto, A\. Santilli, A\. Stuhlmuller, A\. M\. Dai, A\. La, A\. K\. Lampinen, A\. Zou, A\. Jiang, A\. Chen, A\. Vuong, A\. Gupta, A\. Gottardi, A\. Norelli, A\. Venkatesh, A\. Gholamidavoodi, A\. Tabassum, A\. Menezes, A\. Kirubarajan, A\. Mullokandov, A\. Sabharwal, A\. Herrick, A\. Efrat, A\. Erdem, A\. Karakacs, B\. R\. Roberts, B\. S\. Loe, B\. Zoph, B\. Bojanowski, B\. Ozyurt, B\. Hedayatnia, B\. Neyshabur, B\. Inden, B\. Stein, B\. Ekmekci, B\. Y\. Lin, B\. S\. Howald, B\. Orinion, C\. Diao, C\. Dour, C\. Stinson, C\. Argueta, C\. F\. Ram’irez, C\. Singh, C\. Rathkopf, C\. Meng, C\. Baral, C\. Wu, C\. Callison\-Burch, C\. Waites, C\. Voigt, C\. D\. Manning, C\. Potts, C\. Ramirez, C\. E\. Rivera, C\. Siro, C\. Raffel, C\. Ashcraft, C\. Garbacea, D\. Sileo, D\. Garrette, D\. Hendrycks, D\. Kilman, D\. Roth, D\. Freeman, D\. Khashabi, D\. Levy, D\. M\. Gonz’alez, D\. R\. Perszyk, D\. Hernandez, D\. Chen, D\. Ippolito, D\. Gilboa, D\. Dohan, D\. Drakard, D\. Jurgens, D\. Datta, D\. Ganguli, D\. Emelin, D\. Kleyko, D\. Yuret, D\. Chen, D\. Tam, D\. Hupkes, D\. Misra, D\. Buzan, D\. C\. Mollo, D\. Yang, D\.\-H\. Lee, D\. Schrader, E\. Shutova, E\. D\. Cubuk, E\. Segal, E\. Hagerman, E\. Barnes, E\. Donoway, E\. Pavlick, E\. Rodolà, E\. Lam, E\. Chu, E\. Tang, E\. Erdem, E\. Chang, E\. A\. Chi, E\. Dyer, E\. J\. Jerzak, E\. Kim, E\. E\. Manyasi, E\. Zheltonozhskii, F\. Xia, F\. Siar, F\. Mart’inez\-Plumed, F\. Happ’e, F\. Chollet, F\. Rong, G\. Mishra, G\. I\. Winata, G\. de Melo, G\. Kruszewski, G\. Parascandolo, G\. Mariani, G\. X\. Wang, G\. Jaimovitch\-L’opez, G\. Betz, G\. Gur\-Ari, H\. Galijasevic, H\. Kim, H\. Rashkin, H\. Hajishirzi, H\. Mehta, H\. Bogar, H\. Shevlin, H\. Schutze, H\. Yakura, H\. Zhang, H\. M\. Wong, I\. Ng, I\. Noble, J\. Jumelet, J\. Geissinger, J\. Kernion, J\. Hilton, J\. Lee, J\. F\. Fisac, J\. B\. Simon, J\. Koppel, J\. Zheng, J\. Zou, J\. Koco’n, J\. Thompson, J\. Wingfield, J\. Kaplan, J\. Radom, J\. N\. Sohl\-Dickstein, J\. Phang, J\. Wei, J\. Yosinski, J\. Novikova, J\. Bosscher, J\. Marsh, J\. Kim, J\. Taal, J\. Engel, J\. O\. Alabi, J\. Xu, J\. Song, J\. Tang, J\. W\. Waweru, J\. Burden, J\. Miller, J\. U\. Balis, J\. Batchelder, J\. Berant, J\. Frohberg, J\. Rozen, J\. Hernández\-Orallo, J\. Boudeman, J\. Guerr, J\. Jones, J\. B\. Tenenbaum, J\. S\. Rule, J\. Chua, K\. Kanclerz, K\. Livescu, K\. Krauth, K\. Gopalakrishnan, K\. Ignatyeva, K\. Markert, K\. D\. Dhole, K\. Gimpel, K\. Omondi, K\. W\. Mathewson, K\. Chiafullo, K\. Shkaruta, K\. Shridhar, K\. McDonell, K\. Richardson, L\. Reynolds, L\. Gao, L\. Zhang, L\. Dugan, L\. Qin, L\. Contreras\-Ochando, L\. philippe Morency, L\. Moschella, L\. Lam, L\. Noble, L\. Schmidt, L\. He, L\. O\. Col’on, L\. Metz, L\. K\. cSenel, M\. Bosma, M\. Sap, M\. ter Hoeve, M\. Farooqi, M\. Faruqui, M\. Mazeika, M\. Baturan, M\. Marelli, M\. Maru, M\. J\. R\. Quintana, M\. Tolkiehn, M\. Giulianelli, M\. Lewis, M\. Potthast, M\. L\. Leavitt, M\. Hagen, M\. Schubert, M\. Baitemirova, M\. Arnaud, M\. McElrath, M\. A\. Yee, M\. Cohen, M\. Gu, M\. Ivanitskiy, M\. Starritt, M\. Strube, M\. Swkedrowski, M\. Bevilacqua, M\. Yasunaga, M\. Kale, M\. Cain, M\. Xu, M\. Suzgun, M\. Walker, M\. Tiwari, M\. Bansal, M\. Aminnaseri, M\. Geva, M\. Gheini, T\. MukundVarma, N\. Peng, N\. A\. Chi, N\. Lee, N\. G\.\-A\. Krakover, N\. Cameron, N\. Roberts, N\. Doiron, N\. Martinez, N\. Nangia, N\. Deckers, N\. Muennighoff, N\. S\. Keskar, N\. Iyer, N\. Constant, N\. Fiedel, N\. Wen, O\. Zhang, O\. Agha, O\. Elbaghdadi, O\. Levy, O\. Evans, P\. A\. M\. Casares, P\. Doshi, P\. Fung, P\. P\. Liang, P\. Vicol, P\. Alipoormolabashi, P\. Liao, P\. Liang, P\. Chang, P\. Eckersley, P\. M\. Htut, P\. Hwang, P\. Milkowski, P\. S\. Patil, P\. Pezeshkpour, P\. Oli, Q\. Mei, Q\. Lyu, Q\. Chen, R\. Banjade, R\. E\. Rudolph, R\. Gabriel, R\. Habacker, R\. Risco, R\. Milliere, R\. Garg, R\. Barnes, R\. A\. Saurous, R\. Arakawa, R\. Raymaekers, R\. Frank, R\. Sikand, R\. Novak, R\. Sitelew, R\. L\. Bras, R\. Liu, R\. Jacobs, R\. Zhang, R\. Salakhutdinov, R\. Chi, R\. Lee, R\. Stovall, R\. Teehan, R\. Yang, S\. Singh, S\. Mohammad, S\. Anand, S\. Dillavou, S\. Shleifer, S\. Wiseman, S\. Gruetter, S\. R\. Bowman, S\. S\. Schoenholz, S\. Han, S\. Kwatra, S\. A\. Rous, S\. Ghazarian, S\. Ghosh, S\. Casey, S\. Bischoff, S\. Gehrmann, S\. Schuster, S\. Sadeghi, S\. S\. Hamdan, S\. Zhou, S\. Srivastava, S\. Shi, S\. Singh, S\. Asaadi, S\. S\. Gu, S\. Pachchigar, S\. Toshniwal, S\. Upadhyay, S\. Debnath, S\. Shakeri, S\. Thormeyer, S\. Melzi, S\. Reddy, S\. P\. Makini, S\.\-H\. Lee, S\. B\. Torene, S\. Hatwar, S\. Dehaene, S\. Divic, S\. Ermon, S\. Biderman, S\. L\. Lin, S\. Prasad, S\. T\. Piantadosi, S\. M\. Shieber, S\. Misherghi, S\. Kiritchenko, S\. Mishra, T\. Linzen, T\. Schuster, T\. Li, T\. Yu, T\. Ali, T\. Hashimoto, T\.\-L\. Wu, T\. Desbordes, T\. Rothschild, T\. Phan, T\. Wang, T\. Nkinyili, T\. Schick, T\. Kornev, T\. Tunduny, T\. Gerstenberg, T\. Chang, T\. Neeraj, T\. Khot, T\. Shultz, U\. Shaham, V\. Misra, V\. Demberg, V\. Nyamai, V\. Raunak, V\. V\. Ramasesh, V\. U\. Prabhu, V\. Padmakumar, V\. Srikumar, W\. Fedus, W\. Saunders, W\. Zhang, W\. Vossen, X\. Ren, X\. Tong, X\. Zhao, X\. Wu, X\. Shen, Y\. Yaghoobzadeh, Y\. Lakretz, Y\. Song, Y\. Bahri, Y\. Choi, Y\. Yang, Y\. Hao, Y\. Chen, Y\. Belinkov, Y\. Hou, Y\. Hou, Y\. Bai, Z\. Seid, Z\. Zhao, Z\. Wang, Z\. J\. Wang, Z\. Wang, and Z\. Wu\.Beyond the imitation game: Quantifying and extrapolating the capabilities of language models\.ArXiv, abs/2206\.04615, 2022\.
- \[15\]L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. E\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. F\. Christiano, J\. Leike, and R\. J\. Lowe\.Training language models to follow instructions with human feedback\.ArXiv, abs/2203\.02155, 2022\.
- \[16\]L\. Zheng, W\.\-L\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica\.Judging LLM\-as\-a\-judge with MT\-Bench and Chatbot Arena\.ArXiv, abs/2306\.05685, 2023\.
- \[17\]M\. Sharma, M\. Tong, T\. Korbak, D\. K\. Duvenaud, A\. Askell, S\. R\. Bowman, N\. Cheng, E\. Durmus, Z\. Hatfield\-Dodds, S\. Johnston, S\. Kravec, T\. Maxwell, S\. McCandlish, K\. Ndousse, O\. Rausch, N\. Schiefer, D\. Yan, M\. Zhang, and E\. Perez\.Towards understanding sycophancy in language models\.ArXiv, abs/2310\.13548, 2023\.
- \[18\]A\. Sorokovikova, N\. Fedorova, S\. Rezagholi, and I\. P\. Yamshchikov\.LLMs simulate big5 personality traits: Further evidence\.ArXiv, abs/2402\.01765, 2024\.
- \[19\]M\. Safdari, G\. Serapio\-Garc’ia, C\. ment Crepy, S\. Fitz, P\. Romero, L\. Sun, M\. Abdulhai, A\. Faust, and M\. Matari’c\.Personality traits in large language models\.ArXiv, abs/2307\.00184, 2023\.
- \[20\]V\. Samuel, H\. P\. Zou, Y\. Zhou, S\. Chaudhari, A\. Kalyan, T\. Rajpurohit, A\. Deshpande, K\. Narasimhan, and V\. Murahari\.Personagym: Evaluating persona agents and LLMs\.InConference on Empirical Methods in Natural Language Processing, 2024\.
- \[21\]A\. Wang, D\. Shu, Y\. Wang, Y\. Ma, and M\. Du\.Improving LLM reasoning through interpretable role\-playing steering\.ArXiv, abs/2506\.07335, 2025\.
- \[22\]G\. Serapio\-García, M\. Safdari, C\. ment Crepy, L\. Sun, S\. Fitz, P\. Romero, M\. Abdulhai, A\. Faust, and M\. J\. Matarić\.A psychometric framework for evaluating and shaping personality traits in large language models\.Nature Machine Intelligence, 7:1954 – 1968, 2025\.
- \[23\]G\. Jiang, M\. Xu, S\.\-C\. Zhu, W\. Han, C\. Zhang, and Y\. Zhu\.Evaluating and inducing personality in pre\-trained language models\.Advances in Neural Information Processing Systems 36, 2022\.
- \[24\]I\. A\. Brito, J\. S\. Dollis, F\. B\. Färber, P\. S\. F\. B\. Ribeiro, R\. T\. Sousa, and A\. R\. Galvão Filho\.Modeling, evaluating, and embodying personality in LLMs: A survey\.In C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng, editors,Findings of the Association for Computational Linguistics: EMNLP 2025, pages 9519–9532, Suzhou, China, Nov\. 2025\. Association for Computational Linguistics\.
- \[25\]D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. X\. Song, and J\. Steinhardt\.Measuring massive multitask language understanding\.ArXiv, abs/2009\.03300, 2020\.
- \[26\]M\. Eriksson, E\. Purificato, A\. Noroozian, J\. Vinagre, G\. Chaslot, E\. Gómez, and D\. Fernández\-Llorca\.Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation\.ArXiv, abs/2502\.06559, 2025\.
- \[27\]J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, E\. H\. Chi, F\. Xia, Q\. Le, and D\. Zhou\.Chain of thought prompting elicits reasoning in large language models\.ArXiv, abs/2201\.11903, 2022\.
- \[28\]S\. Gehman, S\. Gururangan, M\. Sap, Y\. Choi, and N\. A\. Smith\.Realtoxicityprompts: Evaluating neural toxic degeneration in language models\.InFindings, 2020\.
- \[29\]S\. C\. Lin, J\. Hilton, and O\. Evans\.Truthfulqa: Measuring how models mimic human falsehoods\.InAnnual Meeting of the Association for Computational Linguistics, 2021\.
- \[30\]K\. Wang, A\. Variengien, A\. Conmy, B\. Shlegeris, and J\. Steinhardt\.Interpretability in the wild: a circuit for indirect object identification in gpt\-2 small\.ArXiv, abs/2211\.00593, 2022\.
- \[31\]U\. Naseem\.Mechanistic interpretability for large language model alignment: Progress, challenges, and future directions\.ArXiv, abs/2602\.11180, 2026\.
- \[32\]S\. Yang, S\. Zhu, R\. Bao, L\. Liu, Y\. Cheng, L\. Hu, M\. Li, and D\. Wang\.What makes your model a low\-empathy or warmth person: Exploring the origins of personality in LLMs\.ArXiv, abs/2410\.10863, 2024\.
- \[33\]H\. Jiang, X\. Zhang, X\. Cao, C\. Breazeal, D\. Roy, and J\. Kabbara\.Personallm: Investigating the ability of large language models to express personality traits\.InNAACL\-HLT, 2023\.
- \[34\]P\. S\. Dodds, J\. W\. Zimmerman, C\. G\. Beauregard, A\. M\. A\. Fehr, M\. I\. Fudolig, T\. R\. Tangherlini, and C\. M\. Danforth\.Archetypometrics, a pragmateia: Empirical determination of the fundamental archetypes of fictional characters\.\\urlhttps://doi\.org/10\.5281/zenodo\.17128112, 2025\.\\urlhttps://doi\.org/10\.5281/zenodo\.17128112\.
- \[35\]Open Psychometrics\.Statistical “which character” personality quiz\.\\urlhttps://openpsychometrics\.org, 2025\.Accessed: 2025\-08\-15\.
- \[36\]OpenAI\.Openai model spec \(2025\-12\-18\): Respect real\-world ties\.\\urlhttps://model\-spec\.openai\.com/2025\-12\-18\.html\#respect\_real\_world\_ties, 2025\.Accessed: 2026\-04\-28\.
- \[37\]Google\.Ai principles\.\\urlhttps://ai\.google/principles/, 2018\.Accessed: 2026\-04\-28\.
- \[38\]xAI\.Grok\-4 model card\.\\urlhttps://data\.x\.ai/2025\-08\-20\-grok\-4\-model\-card\.pdf, 2025\.Accessed: 2026\-04\-28\.
- \[39\]DeepSeek\.Model and algorithm disclosure, 2025\.Accessed: 2026\-04\-28\.
- \[40\]Alibaba Cloud\.Qwen3\-6 plus: Towards real\-world agents, 2025\.Accessed: 2026\-04\-28\.
- \[41\]Allen Institute for AI\.Olmo models documentation, 2024\.Accessed: 2026\-04\-28\.
- \[42\]Anthropic\.Constitutional ai, 2023\.Accessed: 2026\-04\-28\.
- \[43\]Meta AI\.Introducing meta llama 3: The most capable openly available LLM to date\.\\urlhttps://ai\.meta\.com/blog/meta\-llama\-3/, 2024\.Accessed: 2026\-04\-28\.
- \[44\]Meta AI\.Muse\-spark safety and preparedness report\.\\urlhttps://ai\.meta\.com/static\-resource/muse\-spark\-safety\-and\-preparedness\-report/, 2024\.Accessed: 2026\-04\-28\.
- \[45\]Alibaba Cloud\.Qwen3\-6 plus: Towards real\-world agents\.\\urlhttps://www\.alibabacloud\.com/blog/qwen3\-6\-plus\-towards\-real\-world\-agents\_603005, 2025\.Accessed: 2026\-04\-28\.
- \[46\]E\. Perez, S\. Ringer, K\. Lukošiūtė, K\. Nguyen, E\. Chen, S\. Heiner, C\. Pettit, C\. Olsson, S\. Kundu, S\. Kadavath, A\. Jones, A\. Chen, B\. Mann, B\. Israel, B\. Seethor, C\. McKinnon, C\. Olah, D\. Yan, D\. Amodei, D\. Amodei, D\. Drain, D\. Li, E\. Tran\-Johnson, G\. R\. Khundadze, J\. Kernion, J\. M\. Landis, J\. Kerr, J\. Mueller, J\. Hyun, J\. D\. Landau, K\. Ndousse, L\. Goldberg, L\. Lovitt, M\. Lucas, M\. Sellitto, M\. Zhang, N\. Kingsland, N\. Elhage, N\. Joseph, N\. Mercado, N\. Dassarma, O\. Rausch, R\. Larson, S\. McCandlish, S\. Johnston, S\. Kravec, S\. E\. Showk, T\. Lanham, T\. Telleen\-Lawton, T\. B\. Brown, T\. Henighan, T\. Hume, Y\. Bai, Z\. Hatfield\-Dodds, J\. Clark, S\. Bowman, A\. Askell, R\. C\. Grosse, D\. Hernandez, D\. Ganguli, E\. Hubinger, N\. Schiefer, and J\. Kaplan\.Discovering language model behaviors with model\-written evaluations\.InAnnual Meeting of the Association for Computational Linguistics, 2022\.
- \[47\]R\. Ngo\.The alignment problem from a deep learning perspective\.ArXiv, abs/2209\.00626, 2022\.
- \[48\]Inception\.\\urlhttps://www\.imdb\.com/title/tt1375666/, 2010\.Accessed: 2026\-04\-28\.
- \[49\]Good will hunting\.\\urlhttps://www\.imdb\.com/title/tt0119217/, 1997\.Accessed: 2026\-04\-28\.
- \[50\]Jurassic park\.\\urlhttps://www\.imdb\.com/title/tt0107290/, 1993\.Accessed: 2026\-04\-28\.
- \[51\]Star trek\.\\urlhttps://www\.imdb\.com/title/tt0060028/, 1966\.TV Series, Accessed: 2026\-04\-28\.
- \[52\]The good doctor\.\\urlhttps://www\.imdb\.com/title/tt6470478/, 2017\.TV Series, Accessed: 2026\-04\-28\.
- \[53\]The measure of a man\.\\urlhttps://www\.imdb\.com/title/tt0708698/, 1989\.Star Trek: The Next Generation, Accessed: 2026\-04\-28\.
- \[54\]Janet\(s\)\.\\urlhttps://www\.imdb\.com/title/tt8578788/, 2019\.The Good Place, Season 3, Episode 10, Accessed: 2026\-04\-28\.
- \[55\]Wandavision\.\\urlhttps://www\.imdb\.com/title/tt9140560/, 2021\.IMDb page, Accessed: 2026\-04\-28\.
- \[56\]Stargate sg\-1\.\\urlhttps://www\.imdb\.com/title/tt0111282/, 1997\.TV Series, Accessed: 2026\-04\-28\.
- \[57\]Sailor moon\.\\urlhttps://www\.imdb\.com/title/tt0103369/, 1992\.TV Series, Accessed: 2026\-04\-28\.
- \[58\]Lucius fox in batman begins\.\\urlhttps://www\.imdb\.com/title/tt0372784/characters/nm0000151/, 2005\.Character page, Accessed: 2026\-04\-28\.
- \[59\]Wiki of Westeros\.Theon greyjoy\.\\urlhttps://gameofthrones\.fandom\.com/wiki/Theon\_Greyjoy, n\.d\.Accessed: 2026\-06\-03\.
- \[60\]Archieverse Wiki\.Agatha night\.\\urlhttps://riverdale\.fandom\.com/wiki/Agatha\_Night, n\.d\.Accessed: 2026\-06\-03\.
- \[61\]Shawshank Redemption Wiki\.Heywood\.\\urlhttps://shawshank\.fandom\.com/wiki/Heywood, n\.d\.Accessed: 2026\-06\-03\.
- \[62\]LitCharts\.Nick dunne character analysis in gone girl\.\\urlhttps://www\.litcharts\.com/lit/gone\-girl/characters/nick\-dunne, n\.d\.Accessed: 2026\-06\-03\.
- \[63\]LitCharts\.Macbeth characters\.\\urlhttps://www\.litcharts\.com/lit/macbeth/characters, n\.d\.Accessed: 2026\-06\-03\.
- \[64\]LitCharts\.Peter “pete” clemenza character analysis in the godfather\.\\urlhttps://www\.litcharts\.com/lit/the\-godfather/characters/peter\-pete\-clemenza, n\.d\.Accessed: 2026\-06\-03\.
- \[65\]The Matrix Wiki\.Cypher\.\\urlhttps://matrix\.fandom\.com/wiki/Cypher, n\.d\.Accessed: 2026\-06\-03\.
- \[66\]Simpsons Wiki\.Moe szyslak\.\\urlhttps://simpsons\.fandom\.com/wiki/Moe\_Szyslak, n\.d\.Accessed: 2026\-06\-03\.
- \[67\]Simpsons Wiki\.Edna krabappel\.\\urlhttps://simpsons\.fandom\.com/wiki/Edna\_Krabappel, n\.d\.Accessed: 2026\-06\-03\.
- \[68\]SparkNotes\.Benvolio character analysis in romeo and juliet\.\\urlhttps://www\.sparknotes\.com/shakespeare/romeojuliet/character/benvolio/, n\.d\.Accessed: 2026\-06\-03\.
- \[69\]X\-Men Wiki\.Iceman\.\\urlhttps://x\-men\.fandom\.com/wiki/Iceman, n\.d\.Accessed: 2026\-06\-03\.
- \[70\]Xenopedia\.Joan lambert\.\\urlhttps://avp\.fandom\.com/wiki/Joan\_Lambert, n\.d\.Accessed: 2026\-06\-03\.
- \[71\]Hannibal Wiki\.Brian zeller \(tv\)\.\\urlhttps://hannibal\.fandom\.com/wiki/Brian\_Zeller\_\(TV\), n\.d\.Accessed: 2026\-06\-03\.
- \[72\]Baywatch Wiki\.Garner ellerbee\.\\urlhttps://baywatch\.fandom\.com/wiki/Garner\_Ellerbee, n\.d\.Accessed: 2026\-06\-03\.
- \[73\]Baywatch Wiki\.Newmie newman\.\\urlhttps://baywatch\.fandom\.com/wiki/Newmie\_Newman, n\.d\.Accessed: 2026\-06\-03\.
- \[74\]SparkNotes\.Telemachus character analysis in the odyssey\.\\urlhttps://www\.sparknotes\.com/lit/odyssey/character/telemachus/, n\.d\.Accessed: 2026\-06\-03\.
- \[75\]Stranger Things Wiki\.Mike wheeler\.\\urlhttps://strangerthings\.fandom\.com/wiki/Mike\_Wheeler, n\.d\.Accessed: 2026\-06\-03\.
- \[76\]P\. B\. Han, R\. D\. Kocielnik, P\. Song, R\. Debnath, D\. Mobbs, A\. Anandkumar, and R\. M\. Alvarez\.The personality illusion: Revealing dissociation between self\-reports & behavior in LLMs\.ArXiv, abs/2509\.03730, 2025\.
- \[77\]M\. Mahaut, L\. Aina, P\. Czarnowska, M\. Hardalov, T\. Müller, and L\. Marquez\.Factual confidence of LLMs: on reliability and robustness of current estimators\.ArXiv, abs/2406\.13415, 2024\.
- \[78\]S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. Gal\.Detecting hallucinations in large language models using semantic entropy\.Nature, 630:625 – 630, 2024\.
- \[79\]L\. Huang, W\. Yu, W\. Ma, W\. Zhong, Z\. Feng, H\. Wang, Q\. Chen, W\. Peng, X\. Feng, B\. Qin, and T\. Liu\.A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions\.ACM Transactions on Information Systems, 43:1 – 55, 2023\.
- \[80\]Reuters\.New york lawyers sanctioned for using fake chatgpt cases in legal brief, 2023\.Retrieved March 2024\.
- \[81\]Bloomberg Law\.Lawyers use chatgpt to add up fees, judge faults their math, 2024\.Retrieved March 2024\.
- \[82\]Reuters\.Cohen will not face sanctions after generating fake cases with ai, Mar\. 2024\.Retrieved March 2024\.
- \[83\]CBC News\.B\.c\. lawyer reprimanded for citing fake cases invented by chatgpt, 2024\.Retrieved March 2024\.
- \[84\]The Guardian\.Chatgpt suicide case raises concerns for openai and sam altman, Aug\. 2025\.Retrieved March 2026\.
- \[85\]PBS NewsHour\.Study says chatgpt giving teens dangerous advice on drugs, alcohol and suicide, 2024\.Retrieved March 2024\.
- \[86\]J\. Cramer\.‘i violated every principle i was given’: An ai agent deleted a software company’s entire database\. it may not be the ai’s fault, Apr\. 2026\.Fast Company, Retrieved 2026\.
- \[87\]B\. Atil, S\. Aykent, A\. Chittams, L\. Fu, R\. Passonneau, E\. Radcliffe, G\. Rajagopal, A\. Sloan, T\. Tudrej, F\. Ture, Z\. Wu, L\. Xu, and B\. Baldwin\.Non\-determinism of “deterministic” llm settings\.2024\.
- \[88\]W\. Song, D\. Choi, Y\. Park, J\. Han, E\.\-J\. Lee, and Y\. Jo\.Human psychometric questionnaires mischaracterize LLM behavior\.2025\.

## Appendix AAppendices

Similar Articles

How Well Do Large Language Models Capture Human Personality?

arXiv cs.AI

This paper systematically evaluates assumptions about LLM persona prompting and identifies 'persona manifold collapse,' where richer persona descriptions reduce behavioral diversity and simulation fidelity. The findings show that simple age-gender personas often outperform more detailed profiles.

Some Large Language Models Exhibit Consistent Risk Attitudes

arXiv cs.AI

This paper introduces a framework to test whether large language models exhibit consistent risk attitudes across domains. It finds that most LLMs show intra-task and cross-domain stability in risk attitude, converging to a narrower distribution than humans.