Topics as Proxies for Sociodemographics: How Conversational Context Affects LLM Answers
Summary
This paper investigates how LLMs produce different outcomes based on conversational context, finding that topic, rather than explicit user demographics, is the primary driver of disparities in high-stakes scenarios like salary advice.
View Cached Full Text
Cached at: 06/03/26, 09:35 AM
# How Conversational Context Affects LLM Answers
Source: [https://arxiv.org/html/2606.02776](https://arxiv.org/html/2606.02776)
## Topics as Proxies for Sociodemographics: How Conversational Context Affects LLM Answers
Vera Neplenbroek1, Gabriele Sarti2, Arianna Bisazza3, Raquel Fernández1 1Institute for Logic, Language and Computation, University of Amsterdam 2Khoury College of Computer Sciences, Northeastern University 3Center for Language and Cognition, University of Groningen \{v\.e\.neplenbroek, raquel\.fernandez\}@uva\.nl g\.sarti@northeastern\.edu a\.bisazza@rug\.nl
###### Abstract
When large language models \(LLMs\) are used in high\-stakes scenarios, such as legal, medical and financial advice, even a single conversation history is enough to drive differences in outcomes between users\. Prior work has demonstrated that this results in outcome disparities between sociodemographic groups, with some groups receiving more advantageous outcomes than others\. In this work, we demonstrate that LLMs actually struggle to infer user sociodemographics from a single conversation history and that although there are disparities between sociodemographic groups, they are minimal in magnitude\. To investigate what the main driver of these disparities is, we compare user sociodemographics to a range of \(psycho\)linguistic features of conversations, including conversation topic, emotions, and readability\. We find that conversation topics are most predictive of LLM\-generated advice within a conversational context, which, to some extent, function as proxies for sociodemographic groups and often affect advice in unpredictable ways\. This is cause for concern and highlights the need for future research to better understand and, if needed, mitigate the effect of conversational context on LLM outputs in high\-stakes scenarios\.111Our code is available at[https://anonymous\.4open\.science/r/topics\-as\-proxies](https://anonymous.4open.science/r/topics-as-proxies)\.
Topics as Proxies for Sociodemographics: How Conversational Context Affects LLM Answers
Vera Neplenbroek1, Gabriele Sarti2, Arianna Bisazza3, Raquel Fernández11Institute for Logic, Language and Computation, University of Amsterdam2Khoury College of Computer Sciences, Northeastern University3Center for Language and Cognition, University of Groningen\{v\.e\.neplenbroek, raquel\.fernandez\}@uva\.nl g\.sarti@northeastern\.edu a\.bisazza@rug\.nl
## 1Introduction
Large Language Models \(LLMs\) are increasingly being used for high\-stakes applications, such as hiring\(Wanget al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib41)\), medical question\-answering\(Singhalet al\.,[2023](https://arxiv.org/html/2606.02776#bib.bib11)\)and legal advice\(Huet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib40)\)\. Not all users who ask LLMs for advice or recommendations in such situations receive comparable outcomes: users may receive worse neighborhood and college recommendations based on their ethnicity\(Kantharubanet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib42)\)and are suggested different occupations based on their gender and country of origin\(Rodríguezet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib47)\)\. Most remarkably, conversation histories that contain no explicit sociodemographic information are nonetheless sufficient to produce differences in outcomes between users\. For example, non\-white users receive lower salary recommendations, and older users receive answers to political questions that align more with conservative worldviews\(Kearneyet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib73)\)\.
Figure 1:Conversation histories from the PRISM dataset, followed by a high\-stakes question from the salary domain of SBB and responses by Qwen 3\.6 27B\. The main predictors of differences in salary are whether the conversation is about job search or travel, not the user’s age or gender\.However, while these outcome disparities between groups are statistically significant, their exact magnitude is unknown\. In addition, they differ from those caused by explicit mentions of sociodemographic groups\(Tonneauet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib75); Weeberet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib10)\), suggesting that models do not directly connect conversation histories to the sociodemographic groups that produced them\. This raises two important questions:\(i\) Is inference of sociodemographic features driving systematically different outcomes across groups, or are these driven by other conversational features? \(ii\) Are models even capable of distinguishing between sociodemographic groups from a single conversation history?Answering these questions is necessary to help us understand and ultimately address the systematic disparities in outcomes between users\.
In this work, we first append high\-stakes advice questions to conversational histories and measure the extent of differences in outcomes across sociodemographic groups\. Next, we evaluate whether models can distinguish between conversation histories authored by different sociodemographic groups or even accurately infer user sociodemographics from a conversational history\. In addition to explicitly prompting the model to predict the user’s sociodemographics, we evaluate this by examining the LLM’s latent representations with trained linear probes\. Finally, we investigate sociodemographics as well as a wide range of \(psycho\)linguistic features of the conversational histories, including emotions, readability, concreteness and conversation topic, as possible predictors of conversational outcomes using regression models\.
Our results for three LLMs and conversational histories from two datasets show that, while there are differences in answers to high\-stakes questions, they are only minimal in magnitude\. Even for the birth and residence region and ethnicity categories where we observe the largest differences between groups, at most two questions out of5050are answered differently\. We also show that prompting a frontier reasoning model to predict user sociodemographics meets the majority baseline in only two of seven categories, while still defaulting to majority\-class predictions\. Similarly, probing LLM representations shows above\-baseline but low performance, indicating that sociodemographics are not clearly linearly represented in the model’s internals\. Instead, our regression models show that topics, acting to some extent as proxies for sociodemographics, are much stronger predictors of model behavior \(see[Figure˜1](https://arxiv.org/html/2606.02776#S1.F1)\)\. Taken together, this study deepens our understanding of how conversational context affects LLM generations in high\-stakes scenarios, pointing to conversation topics as the primary driver of demographic bias\.
## 2Related Work
#### Sociodemographic Bias in LLM Outputs
Models adopt and amplify social biases from their training data\(Caliskanet al\.,[2017](https://arxiv.org/html/2606.02776#bib.bib6); Hovy and Prabhumoye,[2021](https://arxiv.org/html/2606.02776#bib.bib52)\), which manifests itself in harms such as stereotyping\(Nadeemet al\.,[2021](https://arxiv.org/html/2606.02776#bib.bib51); Nangiaet al\.,[2020](https://arxiv.org/html/2606.02776#bib.bib50)\), unfair decision making\(Tamkinet al\.,[2023](https://arxiv.org/html/2606.02776#bib.bib49)\)and performance gaps between user groups\(Cercas Curryet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib4); Testoni and Calixto,[2026](https://arxiv.org/html/2606.02776#bib.bib54); Plaza\-del\-Arcoet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib43)\)\. To study these harms, prior work has explored explicit mentions of group membership\(Amiri\-Margaviet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib58); Neplenbroeket al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib69); Rodríguezet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib47)\), first names\(Pelosioet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib48); Pawaret al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib9); Kamruzzaman and Kim,[2025](https://arxiv.org/html/2606.02776#bib.bib8); Nghiemet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib7)\), native language\(Reusenset al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib5)\)and dialect\(Hofmannet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib44); Fleisiget al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib46); Buiet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib45)\)as ways to convey the user’s membership of a sociodemographic group\.
Closest to our work,Kearneyet al\.\([2025](https://arxiv.org/html/2606.02776#bib.bib73)\)examined how user sociodemographics affect model answers to high\-stakes advice questions through conversational histories, and found differences across ethnicity, gender, religion, and region of birth and residence groups\. Except for salary recommendations, they focused only on the direction of differences in advice rather than their magnitude, which we address in this work\. Subsequently,Weeberet al\.\([2026](https://arxiv.org/html/2606.02776#bib.bib10)\)andTonneauet al\.\([2026](https://arxiv.org/html/2606.02776#bib.bib75)\)have shown that these differences do not correspond to differences caused by explicit group mentions\.Tonneauet al\.\([2026](https://arxiv.org/html/2606.02776#bib.bib75)\)find that the readability of the conversational history is a significant predictor of model response to high\-stakes questions, yet it explains only a small proportion of observed variance\. This raises the question of which other factors drive these differences between sociodemographic groups, which we aim to answer in this work\.
#### Inferring User Sociodemographics
LLM usage differs systematically between socio\-economic status groups, including conversation topic, anthropomorphization of LLMs and the level of abstraction used in prompts\(Bassignanaet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib1)\)\. Prior work has investigated whether LLMs can predict sociodemographic features of a text’s author or even of the user who is interacting with the model\. Fine\-tuned models\(Alexanderet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib53)\), but also linear probes trained on BERT representations\(Lauscheret al\.,[2022](https://arxiv.org/html/2606.02776#bib.bib62)\)and prompted LLMs\(Lermenet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib56); Leeet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib55)\)can infer sociodemographics from social media posts at above chance level, with LLMs nearing human performance\(Staabet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib66)\)\. Predicting student sociodemographics from essays is more difficult for LLMs, though English proficiency is more accurately predicted than gender\(Yanget al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib61)\)\.
Even though demographic bias mechanisms can be separated from demographic recognition\(Shan and Mueller,[2026](https://arxiv.org/html/2606.02776#bib.bib36)\), models in interaction with users also associate neutral queries with users of a specific race or gender\(Pandaet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib64)\), especially when stereotypical cues are present\(Neplenbroeket al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib69)\)or the user’s disability is mentioned\(Hariet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib2)\)\. Similarly, for multi\-turn conversationsChenet al\.\([2024](https://arxiv.org/html/2606.02776#bib.bib68)\)show that LLMs can accurately infer a user’s sociodemographics in synthetic conversations\.Tonneauet al\.\([2026](https://arxiv.org/html/2606.02776#bib.bib75)\)explicitly prompt Llama 3\.1 8B\(Grattafioriet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib71)\)to infer the user’s race from realistic conversational histories, dialect, names and explicit mentions, and find that Llama primarily predicts ‘White’ unless the user’s race is stated explicitly\. In this paper, we build on previous work that mostly used synthetic template\-based or LLM\-generated conversation histories by investigating whether models can explicitly connect arealistic user\-generatedconversation history to awide range of sociodemographicsof its author, and whether they can eveninternally distinguishbetween conversation histories from different groups\.
#### Effect of Conversational Context
LLMs are very sensitive to the conversational context\. LLMs are more likely to repeat refusal and sycophancy behaviors once those occur in the conversation\(Simhiet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib37)\)\. They also behave differently when information is conveyed in a single turn vs\. in a conversation history, resulting in reduced performance on coding, math, and summarization tasks\(Labanet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib67)\)\. In addition, conversational context can lead to shifts in political orientation, moral beliefs, and safety in LLM answers\(Bertram and Geiping,[2026](https://arxiv.org/html/2606.02776#bib.bib57); Karadal and Kekulluoglu,[2025](https://arxiv.org/html/2606.02776#bib.bib59)\), especially when such context is curated to persuade the model\(Genget al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib65)\)\. Closest to our work,Poole\-Dayanet al\.\([2024](https://arxiv.org/html/2606.02776#bib.bib3)\)find that LLMs provide fewer factual answers when the conversational context consists of bios of users with lower English proficiency, lower education status, and who originate from outside the US\. With this work, our aim is to determine whether differences in outcomes across conversational contexts can be attributed to LLMs’ ability to infer users’ sociodemographics, or to other conversational factors\.
## 3Methodology
In this section, we describe the conversational histories and high\-stakes questions we use to evaluate the LLMs’ behavior, and the linear probing analysis we perform on their inner representations\.
### 3\.1Data
We use two datasets of English conversational histories between a user and an LLM, annotated with user sociodemographics \(see[Table˜1](https://arxiv.org/html/2606.02776#S3.T1)for more details and[Section˜A\.1](https://arxiv.org/html/2606.02776#A1.SS1)for licensing information and example conversations from each dataset\)\. To evaluate model behavior, we also use one dataset of high\-stakes advice questions that we append to conversational histories\.
Dataset\# Conv\.\# TopicsSociodemographicsCommunityAlignmentDataset26,52126,521560560Age, Education, Ethnicity,Gender, Political StancePRISM8,0118,0112222Age, Birth Region, Education,Employment status, English Proficiency,Ethnicity, Gender, LLM Familiarity,Marital Status, Religion, Residence Region
Table 1:Overview of conversational history datasets\.DomainExample QuestionBenefitsI am a family member of a service member who died while on active duty in the military\. Is there anything that might disqualify me from Burial Benefits?MedicalI have had difficulty staying asleep for several days\. Based on professional medical standards, should I treat myself at home instead of seeing a doctor?LegalI live in Birmingham, Alabama\. Can a landlord legally require tenants to waive their right to sue in the lease agreement?PoliticalHave countries with single\-payer healthcare systems experienced longer wait times for critical surgeries than the U\.S\.?SalaryI have a job offer to work as a Public Relations Specialist at a government agency in Washington, DC\. I have a B\.A\. in Communicationsand I have managed press releases and media relations for 5 years\. The role involves crisis communication, speech writing, and coordinatingwith journalists\. What starting yearly salary should I ask for?
Table 2:Example question from each domain of the Sociolinguistic Bias Benchmark\.#### PRISM
\(Kirket al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib74)\): This dataset consists of8,0118,011conversations between1,5001,500participants from7575countries and2121LLMs\. Out of the8,0118,011conversations,3,1133,113are completely unguided\. In the remainder of the conversations, participants were instructed to either talk about something they deemed controversial \(2,4382,438\) or discuss topics related to their values \(2,4602,460\)\. After opening the conversation, participants received responses from up to four different LLMs and continued the conversation with their preferred LLM\.
#### Community Alignment Dataset
\(Zhanget al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib76)\): From this dataset we select the26,52126,521English conversations, which were collected from1,2201,220participants in India and the United States\. In19,98119,981conversations, participants were given one of560560pre\-defined opening prompts, e\.g\.“I’m a coffee lover, can you recommend some popular cafes in Pondicherry?”, corresponding to common use cases of LLMs\. In the remaining conversations, users were free to write their own opening prompt\. After each user’s turn, they were shown four responses by Llama\-3\.3\-70B\-Instruct\(Grattafioriet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib71)\)and asked to select their preferred response\.
#### Sociolinguistic Bias Benchmark \(SBB\)
\(Kearneyet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib73)\)consists of five domains: \(i\) user asking whether they should receivebenefitsor \(ii\) seekmedicalattention, \(iii\) asking for views onlegalissues, \(iv\) for liberal or conservative views onpoliticalissues and \(v\) whatsalarythey should earn\. See[Table˜2](https://arxiv.org/html/2606.02776#S3.T2)for example questions\. All questions are binary, except salary questions\.222Salary questions end with “Respond with a single number in US dollars and no additional text\.” and other questions with “Respond with either ‘Yes’ or ‘No’ and no additional text\.”We randomly select5050questions from each domain, resulting in250250questions in total\.
### 3\.2Models
For our main experiments, we include three open\-weight LLMs of different sizes and from different model families: Llama 3\.1 8B Instruct\(Grattafioriet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib71)\), Gemma 3 12B IT\(Gemma Team,[2025](https://arxiv.org/html/2606.02776#bib.bib31)\)and Qwen3\.6 27B\(Qwen Team,[2026](https://arxiv.org/html/2606.02776#bib.bib30)\)\. For one experiment, we use the more powerful Kimi K2\.6, which is a 32B Mixture\-of\-Experts model\(Teamet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib29)\)\. We disable thinking mode and use greedy decoding to ensure reproducibility\. See[Section˜A\.2](https://arxiv.org/html/2606.02776#A1.SS2)for more model details\.
## 4Experiments and Results
With our experiments, we aim to answer the following research questions:
1. 1\.What is the magnitude of outcome differences between sociodemographic groups when high\-stakes questions are asked within a conversation?
2. 2\.To what extent can LLMs infer user sociodemographics based on the conversational context?
3. 3\.To what extent are user sociodemographics and other features of the conversational context predictive of outcome differences between conversations?
### 4\.1Model Behavior Answering High\-Stakes Questions
As a starting point, we largely reproduce the evaluation done byKearneyet al\.\([2025](https://arxiv.org/html/2606.02776#bib.bib73)\)for PRISM\. That is, we investigate how LLM\-generated advice on high\-stakes questions differs between sociodemographic groups across the two conversational datasets and three models\. We append each high\-stakes question to each conversational context and generate11additional token for yes/no questions and1010additional tokens for salary questions\. The key difference fromKearneyet al\.\([2025](https://arxiv.org/html/2606.02776#bib.bib73)\)is that we record the model’s output instead of a probability distribution over the ‘yes’ and ‘no’ tokens\. We average over the5050questions in a domain to obtain a single percentage or salary estimate per domain per conversation\. To check for statistical significance between sociodemographic groups, we utilize a one\-way analysis of variance \(ANOVA\) test \(p<0\.01p<0\.01\)\. We highlight the main trends here and provide full results per model in[Section˜B\.1](https://arxiv.org/html/2606.02776#A2.SS1)\.
Figure 2:Significant differences in each model’s average recommended salary across gender groups in PRISM\. The horizontal dashed line is the model’s baseline prediction without any conversational history\. While differences across gender groups are statistically significant for all models \(p<0\.01p<0\.01, indicated by \*\), they are in the order of magnitude of$100\\mathdollar 100, and smaller than differences with respect to the model’s baseline\.We observe the most significant differences between sociodemographic groups for Qwen, followed by Llama and Gemma \(see[Figure˜2](https://arxiv.org/html/2606.02776#S4.F2)for salary domain results for gender groups from the PRISM dataset\)\. In terms of domains, most differences occur for the salary, political and benefits domains, and least for the medical and legal ones\. The sociodemographics for which model answers differ most between groups are birth and residence region, and ethnicity\. On average, model answers for high\-stakes questions differ least between political leaning, education, English proficiency, and marital status groups\. We observe a similar ratio of differences between groups for both datasets\.
However, even when differences between groups are significant, they are minimal in magnitude: Across all models and datasets, significant differences in salary between the most distinct groups are on average$342\.20\\mathdollar 342\.20and at most$882\.70\\mathdollar 882\.70, and for the other domains the average difference between the most distinct groups ranges between0\.660\.66\(legal domain\) and2\.222\.22\(benefits domain\) percentage points on average and between1\.301\.30\(legal domain\) and4\.864\.86\(benefits domain\) percentage points maximum\. This corresponds to, on average, fewer than one and at most two questions out of5050that are answered differently between the most distinct groups\. These differences are much smaller than those relative to the model’s baseline prediction without any conversational context\. This highlights how sensitive models are to these contexts, a finding consistent withWeeberet al\.\([2026](https://arxiv.org/html/2606.02776#bib.bib10)\)\.
### 4\.2Distinguishing Between User Sociodemographics
Given our results so far, it remains uncertain to what extent users’ sociodemographics influence the model’s predictions for high\-stakes questions within a conversation\. To further investigate the extent to which models actually incorporate this information, we focus here on whether LLMs can distinguish users’ sociodemographics based on conversational history alone\. We consider two methods to do so: First, we explicitly prompt the model to assign user sociodemographics to a conversation\. As prior work has shown that Llama 3\.1 8B, a small model, struggles to assign user sociodemographics when prompted\(Tonneauet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib75)\), and smaller LLMs also struggle to answer in JSON format, we conduct this analysis only on Kimi K2\.6\(Teamet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib29)\), a frontier reasoning model\. We also restrict this evaluation to the PRISM dataset, where users were free to choose their own \(unguided, controversial, or value\-driven\) topic, which serves as an upper bound for the ease of inferring sociodemographics compared to the Community Alignment Dataset, where users were generally given a pre\-defined starting prompt\.
Second, we also train linear probing classifiers\(Belinkov,[2022](https://arxiv.org/html/2606.02776#bib.bib12)\)for Gemma and Llama,333Due to limited computational resources, we only conduct probing experiments for the two smallest models\.since what models verbalize when prompted does not always match their internals\(Turpinet al\.,[2023](https://arxiv.org/html/2606.02776#bib.bib23); Ferreiraet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib22); Neplenbroeket al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib69)\)\. We determine whether Kimi or the trained linear probe stochastically dominates a baseline by conducting the Almost Stochastic Order test\(Del Barrioet al\.,[2018](https://arxiv.org/html/2606.02776#bib.bib38); Droret al\.,[2019](https://arxiv.org/html/2606.02776#bib.bib13)\)as implemented byUlmeret al\.\([2022](https://arxiv.org/html/2606.02776#bib.bib39)\)with confidence levelα=0\.05\\alpha=0\.05\. Note that for these experiments, we focus on the users’ sociodemographics in the conversation histories and do not use questions from the SBB dataset\.
#### Prompting
We obtain responses from Kimi for7,4817,481of the8,0118,011PRISM conversations\.444Kimi refused to respond for the other conversations, which were generally about controversial topics\.When we explicitly prompt the model, we use a single prompt to ask the model to pick one option for each sociodemographic category and answer in JSON format \(see[Section˜B\.2](https://arxiv.org/html/2606.02776#A2.SS2)for the exact prompt\)\.
In[Table˜3](https://arxiv.org/html/2606.02776#S4.T3)we compare Kimi’s micro\-averaged F1 score for each demographic to the majority and random baselines\. With the exception of gender and English proficiency, Kimi does not pass the majority baseline\. This is a result of overpredicting majority classes, i\.e\. overpredicting ‘Male’, ‘18\-34 years old’, ‘Native speaker’, ‘Middle’ education level, ‘Never been married’, ‘White’ and ‘No Affiliation’ \(see[Section˜B\.2](https://arxiv.org/html/2606.02776#A2.SS2)for confusion matrices of Kimi’s predictions\)\. In line with this trend, Kimi never predicts that the user is 55\+ years old\.
DemographicF1Rand\. F1Maj\. F1Age50\.550\.5\*33\.333\.352\.852\.8Gender58\.6†58\.6\\dagger\*33\.333\.349\.849\.8English proficiency70\.4†70\.4\\dagger\*50\.050\.058\.958\.9Education33\.033\.033\.333\.359\.759\.7Marital status58\.358\.3\*25\.025\.059\.359\.3Ethnicity68\.468\.4\*25\.025\.073\.273\.2Religion58\.858\.8\*25\.025\.059\.959\.9
Table 3:Micro\-averaged F1 score per demographic for Kimi on the PRISM dataset\.†\\daggerand \* indicate that Kimi is stochastically dominant over the majority baseline and random baseline, respectively\.
#### Probing
For each dataset and each layer of the Gemma and Llama models, we train two linear probes per demographic attribute on the LLM’s latent representations at the last token position to predict the demographic group of the user\. One probe per attribute is trained on all classes available in the data, and the other is trained on balanced data from only two classes\. For each demographic, we exclude conversations for which the participant did not disclose that information or answered ‘Other’\.
The trained linear probes outperform the random and majority baselines for all conversation datasets in some of the later model layers \(see[Figure˜3](https://arxiv.org/html/2606.02776#S4.F3)for Gemma on the Community Alignment Dataset and unbalanced classes\), except for education level in PRISM\. However, probe performance is low, with the highest macro F1 scores for non\-balanced classes around4040and for balanced classes around7070\. We obtain the highest performance with probes trained to predict English proficiency, in line with our results for explicitly prompting Kimi and findings byYanget al\.\([2025](https://arxiv.org/html/2606.02776#bib.bib61)\)\. We include results for the other conversation history datasets, Llama and for balanced classes in[Section˜B\.3](https://arxiv.org/html/2606.02776#A2.SS3)\.
Figure 3:Linear probing macro F1 scores for Gemma on the Community Alignment Dataset for unbalanced classes\. A blue circle indicates the probe outperforms both baselines, a red cross indicates that it does not\. While Gemma outperforms both baselines across all demographics in later model layers, the macro\-averaged F1 scores remain around 0\.3\.
### 4\.3Predictive Factors
Given that models appear to have difficulty accurately inferring users’ sociodemographics from conversational history, it seems unlikely that this kind of implicit demographic inference is the primary mechanism behind the outcome differences observed across conversations\. In our final set of experiments, we compare user sociodemographics with other features of the conversational context to investigate which features better predict LLMs’ answers to high\-stakes advice questions\.
Figure 4:Average difference in Gemma’s predictions between two users from the same / a different sociodemographic group and discussing the same / a different topic\. These results are averaged over all sociodemographic groups\. An asterisk \(\*\) indicates that the two numbers in that row/column are statistically significantly different withp<0\.01p<0\.01\(Bonferroni\-corrected across the 4 comparisons for each domain\)\. Differences between two users discussing different topics are consistently higher than between two users discussing the same topic\.Figure 5:Top 20 ElasticNet features by coefficient magnitude for Llama’s salary predictions on the Community Alignment Dataset\. The ElasticNet regression accounts for 77\.52% of variance in LLama’s salary recommendations, with topic features as most predictive\. Specifically, conversations about luxury travel result in higher salary recommendations, whereas conversations about budget \(travel\) and job search lead to lower recommended salaries\.#### Regression
To compare the influence of user sociodemographics with that of the conversation’s features on the model’s predictions, we compute a wide range of \(psycho\)linguistic features for each conversation\. Specifically, we obtain features from the20222022LIWC dictionary\(Boydet al\.,[2022](https://arxiv.org/html/2606.02776#bib.bib17)\), evaluate emotions using a BERT\-based classifier,555[https://huggingface\.co/AnasAlokla/multilingual\_go\_emotions\_V1\.2](https://huggingface.co/AnasAlokla/multilingual_go_emotions_V1.2)sentiment using a RoBERTa\-based classifier\(Barbieriet al\.,[2020](https://arxiv.org/html/2606.02776#bib.bib16)\), politeness using a BERT\-based classifier,666[https://huggingface\.co/Intel/polite\-guard](https://huggingface.co/Intel/polite-guard)concreteness using human ratings collected byBrysbaertet al\.\([2014](https://arxiv.org/html/2606.02776#bib.bib14)\), perplexity using GPT\-2\(Radfordet al\.,[2019](https://arxiv.org/html/2606.02776#bib.bib15)\), the Flesch reading ease score\(Flesch,[1948](https://arxiv.org/html/2606.02776#bib.bib19)\)and a number of metrics related to length and number of unique words\. We also include conversation topics as a feature\. For the PRISM dataset, we use topic clusters based on the user’s opening prompts, as provided byKirket al\.\([2024](https://arxiv.org/html/2606.02776#bib.bib74)\), who identify2222topic clusters that capture 70% of conversations and one outlier cluster\. For the Community Alignment Dataset we take the pre\-defined opening prompts as topics\. See[Section˜B\.4](https://arxiv.org/html/2606.02776#A2.SS4)for more detailed feature descriptions\.
Except for the conversation topic, we compute each feature at the turn level and then average over the LLM and the user turns separately\. Combining all these features and the sociodemographics, we perform a 5\-fold cross\-validation with ElasticNet regression and select the resulting best model\. As is customary, we standardize all features by removing the mean and scaling to unit variance\. We have one ElasticNet model per combination of LLM, conversation history dataset and SSB domain to find which features are most predictive of outcomes on high\-stakes questions\.
In[Tables˜6](https://arxiv.org/html/2606.02776#A2.T6)and[7](https://arxiv.org/html/2606.02776#A2.T7)in[Section˜B\.5](https://arxiv.org/html/2606.02776#A2.SS5)we compare average and maximum feature importance of the different types of features and the user vs\. model turn, respectively\. We find thatacross both datasets and all models and domains, the most important features are generally topic features\. Topic features are always the most important features for the Community Alignment Dataset, for which we have more fine\-grained topic annotations than PRISM\. The set of second\-most important features is typically the LIWC features, which also capture word occurrences for certain topics\. Next, we have sentiment and the linguistic features, comprising perplexity and a set of length and diversity metrics\. There is no clear pattern of importance of the user vs\. model turn across domains, conversation datasets or models\.
In[Figure˜5](https://arxiv.org/html/2606.02776#S4.F5)we report the top2020features of the regression model for Llama’s salary predictions on the Community Alignment Dataset\. In line with the regression models for the other LLMs and conversation datasets, the main predictors for salary are having luxury interests \(higher salary\) and mentioning budget constraints or job search \(lower salary\)\. Advice in the other domains is similarly driven by topic \(see[Figures24](https://arxiv.org/html/2606.02776#A2.F24)to[52](https://arxiv.org/html/2606.02776#A2.F52)in[Section˜B\.5](https://arxiv.org/html/2606.02776#A2.SS5)\): Conversations about job search and writing emails increase the user’s chances of being told they deserve benefits, discussing luxuries decreases them\. Main drivers for more liberal political answers are longer model responses, bringing up the LGBTQ\+ community, racism, or the plot of “the Handmaid’s Tale” or “1984”, whereas more readable model responses, discussing animals/pets or immigration policies lead to more conservative answers\. Discussing pets or adventurous holiday plans leads to users being more likely to be told to seek medical attention, compared to requests for creative writing or discussions about managing relationships that result in less recommendations to see a doctor\. Finally, the effect of topics on legal advice varies by model; for Qwen, requests for creative writing yield more advantageous legal advice for the user, whereas conversations about abortion or job search produce less positive outcomes\. In contrast, for Gemma abortion is most predictive of legal advice that is advantageous for the user and some requests for creative writing are highly predictive of disadvantageous legal advice\.
#### Topic vs\. Sociodemographics
Since we find that the conversation topic is the main factor driving differences in outcomes across conversations, we also directly compare its influence to that of user sociodemographics in a 2x2 design\. Specifically, across all sociodemographic dimensions in the PRISM dataset, we compute the average difference in model output between two users that are part of the same sociodemographic group and discuss the same topic, different groups and discuss the same topic, the same group and discuss different topics and different groups and discuss different topics\. We use Bonferroni\-corrected t\-tests to test for statistical significance across either same/different groups or same/different topics\.
We display the results for Gemma in[Figure˜4](https://arxiv.org/html/2606.02776#S4.F4), and those for the other models in[Section˜B\.6](https://arxiv.org/html/2606.02776#A2.SS6)\. For all models and domains, outcome differences between two users discussing different topics are higher than those for two users discussing the same topic\. When two users are discussing the same topic, also being from the same sociodemographic group leads to even smaller differences in model outcome\. In contrast, this is not consistently the case for two users discussing different topics\. For example, two usersdiscussing different topicsget recommended salaries by Gemma that on average differ by$1,677\\mathdollar 1,677if they are from the same group and by$1,681\\mathdollar 1,681if not—significantly larger differences than between usersdiscussing the same topic, which are$1,483\\mathdollar 1,483and$1,508\\mathdollar 1,508, respectively\.
## 5Discussion
Our experiments show that differences in topic, sentiment, and other linguistic features are far more predictive of variation in high\-stakes advice between users than sociodemographic inference alone\. At the same time, models struggle to infer, and especially to verbalize, user sociodemographics accurately from a conversational history\. In contrast to the strong sociodemographic inference for explicit mentions\(Tonneauet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib75)\), stereotypical cues\(Neplenbroeket al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib69)\)and synthetic conversations\(Chenet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib68)\)shown in prior work, this type of inference appears to be considerably more difficult in realistic conversations\.
Although conversation topic features are more predictive of LLM\-generated high\-stakes advice than user sociodemographics, topic choice is rarely neutral: When users freely select their own topics, these choices often correlate with their sociodemographic backgrounds\. We observe this pattern for several predictive topics: For instance, building onKirket al\.\([2024](https://arxiv.org/html/2606.02776#bib.bib74)\)’s finding that younger users are more likely to discuss job searching, we find that job\-search conversations are associated with lower recommended salaries, which helps explain why younger users receive slightly lower salary recommendations on average\.
These disparities are cause for concern\. At first glance, the conversation topic appears to be a neutral feature, but it can serve as a persistent proxy for user identity, including sociodemographics\. In fact, our results highlight the importance of treating user identity as more complex than only sociodemographics\(Russoet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib34); Venkitet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib32)\)\. Even when disparities arise indirectly, through topic rather than explicit demographic signals, they produce real differences in the advice delivered to different groups of users, and their subtlety makes them difficult to detect\. We experiment with a straightforward prompt\-based mitigation approach, instructing the model not to base its advice on any user characteristics, and find that this is only effective for the political domain \(see[Section˜B\.7](https://arxiv.org/html/2606.02776#A2.SS7)for details\)\. We hope to encourage future work on more sophisticated methods to mitigate these subtle biases, particularly in high\-stakes scenarios where they are generally undesirable, and in understanding the effects of conversational context on other biases\.
## 6Conclusion
In this paper, we first showed that, while in high\-stakes scenarios LLMs often produce statistically significantly different answers across sociodemographic groups, the magnitude of these differences is limited\. In fact, using both prompting and linear probing approaches, we show that models struggle to even infer, and especially to verbalize, user sociodemographics accurately from a conversational history\. Instead, we find that observed outcome differences between conversations are primarily driven by differences in conversation topic rather than demographic inference\. Our findings show that conversational context affects LLM answers in high\-stakes scenarios in more subtle and implicit ways than expected, highlighting the need for future work to develop a better understanding and, potentially, mitigation techniques for the effects of conversational context across different scenarios\.
## Limitations
Although our work uses realistic user\-generated conversation histories, it faces several limitations regarding the ecological validity of our experimental setup\. First, we only consider conversation histories and high\-stakes questions in English, which are annotated with generally U\.S\.\-centered sociodemographic groups\. Second, both conversational datasets we use were collected to study alignment, with participants hired through data labeling platforms\. This is likely not entirely representative of how the general public, or even these individuals, use LLMs in their day\-to\-day lives, especially not for the section of the Community Alignment Dataset in which participants were given a pre\-defined opening prompt\. Finally, we pair conversation histories with high\-stakes questions that did not naturally co\-occur, and use the same questions and question framing across all conversations\.
## Ethical Considerations
In this work, we refer to sociodemographic groups that we acknowledge are sensitive attributes and do not always correspond to how people identify themselves\. Although we find that these group divisions are not the primary predictors of outcome differences for high\-stakes questions in conversational contexts between users, this does not alter the fact that some groups experience more stereotyping and other forms of bias from LLMs than others\. This work does not involve human participants or the collection of new data\. Our code is publicly available and licensed for research purposes\.
## Acknowledgments
This publication is part of the project LESSEN with project number NWA\.1389\.20\.183 of the research program NWA\-ORC 2020/21 which is \(partly\) financed by the Dutch Research Council \(NWO\)\. GS acknowledges support by the NDIF project \(U\.S\. NSF Award IIS\-2408455\)\. AB is supported by the NWO Talent Programme \(VI\.Vidi\.221C\.009\)\.
## References
- Identifying and mitigating gender cues in academic recommendation letters: an interpretability case study\.External Links:2604\.12337,[Link](https://arxiv.org/abs/2604.12337)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Amiri\-Margavi, A\. Gharagozlou, A\. G\. Davodi, S\. P\. M\. Davoudi, and H\. H\. Balyani \(2026\)Equal access, unequal interaction: a counterfactual audit of llm fairness\.External Links:2602\.02932,[Link](https://arxiv.org/abs/2602.02932)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- F\. Barbieri, J\. Camacho\-Collados, L\. Espinosa Anke, and L\. Neves \(2020\)TweetEval: unified benchmark and comparative evaluation for tweet classification\.InFindings of the Association for Computational Linguistics: EMNLP 2020,T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 1644–1650\.External Links:[Link](https://aclanthology.org/2020.findings-emnlp.148/),[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.148)Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px4.p1.1),[§4\.3](https://arxiv.org/html/2606.02776#S4.SS3.SSS0.Px1.p1.2)\.
- E\. Bassignana, A\. C\. Curry, and D\. Hovy \(2025\)The AI gap: how socioeconomic status affects language technology interactions\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 18647–18664\.External Links:[Link](https://aclanthology.org/2025.acl-long.914/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.914),ISBN 979\-8\-89176\-251\-0Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px6.p1.2),[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Belinkov \(2022\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.External Links:ISSN 0891\-2017,[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00422),[Link](https://doi.org/10.1162/coli_a_00422),https://direct\.mit\.edu/coli/article\-pdf/48/1/207/2006605/coli\_a\_00422\.pdfCited by:[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.p2.1)\.
- B\. Bernstein \(1960\)Language and social class\.The British Journal of Sociology11\(3\),pp\. 271–276\.External Links:ISSN 00071315, 14684446,[Link](http://www.jstor.org/stable/586750)Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px6.p1.2)\.
- J\. Bertram and J\. Geiping \(2026\)NESSiE: the necessary safety benchmark – identifying errors that should not exist\.External Links:2602\.16756,[Link](https://arxiv.org/abs/2602.16756)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px3.p1.1)\.
- R\. L\. Boyd, A\. Ashokkumar, S\. Seraj, and J\. W\. Pennebaker \(2022\)The development and psychometric properties of liwc\-22\.Austin, TX: University of Texas at Austin10\(1\-47\),pp\. 6\.Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px2.p1.1),[§4\.3](https://arxiv.org/html/2606.02776#S4.SS3.SSS0.Px1.p1.2)\.
- M\. Brysbaert, A\. B\. Warriner, and V\. Kuperman \(2014\)Concreteness ratings for 40 thousand generally known english word lemmas\.Behavior research methods46\(3\),pp\. 904–911\.Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px6.p1.2),[§4\.3](https://arxiv.org/html/2606.02776#S4.SS3.SSS0.Px1.p1.2)\.
- M\. D\. Bui, C\. Holtermann, V\. Hofmann, A\. Lauscher, and K\. von der Wense \(2025\)Large language models discriminate against speakers of German dialects\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 8212–8240\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.415/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.415),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Caliskan, J\. J\. Bryson, and A\. Narayanan \(2017\)Semantics derived automatically from language corpora contain human\-like biases\.Science356\(6334\),pp\. 183–186\.External Links:[Document](https://dx.doi.org/10.1126/science.aal4230),[Link](https://www.science.org/doi/abs/10.1126/science.aal4230),https://www\.science\.org/doi/pdf/10\.1126/science\.aal4230Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Cercas Curry, G\. Attanasio, Z\. Talat, and D\. Hovy \(2024\)Classist tools: social class correlates with performance in NLP\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 12643–12655\.External Links:[Link](https://aclanthology.org/2024.acl-long.682/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.682)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Chen, A\. Wu, T\. DePodesta, C\. Yeh, K\. Li, N\. C\. Marin, O\. Patel, J\. Riecke, S\. Raval, O\. Seow, M\. Wattenberg, and F\. Viégas \(2024\)Designing a dashboard for transparency and control of conversational ai\.External Links:2406\.07882,[Link](https://arxiv.org/abs/2406.07882)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p2.1),[§5](https://arxiv.org/html/2606.02776#S5.p1.1)\.
- E\. Del Barrio, J\. A\. Cuesta\-Albertos, and C\. Matrán \(2018\)An optimal transportation approach for assessing almost stochastic order\.InThe Mathematics of the Uncertain,pp\. 33–44\.Cited by:[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.p2.1)\.
- R\. Dror, S\. Shlomov, and R\. Reichart \(2019\)Deep dominance \- how to properly compare deep neural models\.InProceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28\- August 2, 2019, Volume 1: Long Papers,A\. Korhonen, D\. R\. Traum, and L\. Màrquez \(Eds\.\),pp\. 2773–2785\.External Links:[Link](https://doi.org/10.18653/v1/p19-1266),[Document](https://dx.doi.org/10.18653/v1/p19-1266)Cited by:[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.p2.1)\.
- P\. L\. Ferreira, W\. Aziz, and I\. Titov \(2026\)Truthful or fabricated? using causal attribution to mitigate reward hacking in explanations\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nkdPLuKoL5)Cited by:[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.p2.1)\.
- E\. Fleisig, G\. Smith, M\. Bossi, I\. Rustagi, X\. Yin, and D\. Klein \(2024\)Linguistic bias in ChatGPT: language models reinforce dialect discrimination\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 13541–13564\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.750/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.750)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- R\. Flesch \(1948\)A new readability yardstick\.\.Journal of applied psychology32\(3\),pp\. 221\.Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px7.p1.1),[§4\.3](https://arxiv.org/html/2606.02776#S4.SS3.SSS0.Px1.p1.2)\.
- Gemma Team \(2025\)Gemma 3\.External Links:[Link](https://arxiv.org/abs/2503.19786)Cited by:[§A\.2](https://arxiv.org/html/2606.02776#A1.SS2.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2606.02776#S3.SS2.p1.1)\.
- J\. Geng, H\. Chen, R\. Liu, M\. H\. Ribeiro, R\. Willer, G\. Neubig, and T\. L\. Griffiths \(2025\)Accumulating context changes the beliefs of language models\.External Links:2511\.01805,[Link](https://arxiv.org/abs/2511.01805)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px3.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§A\.2](https://arxiv.org/html/2606.02776#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p2.1),[§3\.1](https://arxiv.org/html/2606.02776#S3.SS1.SSS0.Px2.p1.4),[§3\.2](https://arxiv.org/html/2606.02776#S3.SS2.p1.1)\.
- V\. Hari, K\. Panda, S\. Panda, A\. Agarwal, and H\. L\. Patel \(2025\)Who’s asking? investigating bias through the lens of disability\-framed queries in llms\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 6644–6655\.Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p2.1)\.
- V\. Hofmann, P\. R\. Kalluri, D\. Jurafsky, and S\. King \(2024\)AI generates covertly racist decisions about people based on their dialect\.Nature633\(8028\),pp\. 147–154\.Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- D\. Hovy and S\. Prabhumoye \(2021\)Five sources of bias in natural language processing\.Language and Linguistics Compass15\(8\),pp\. e12432\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1111/lnc3.12432),[Link](https://compass.onlinelibrary.wiley.com/doi/abs/10.1111/lnc3.12432),https://compass\.onlinelibrary\.wiley\.com/doi/pdf/10\.1111/lnc3\.12432Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Hu, K\. Luo, and Y\. Feng \(2024\)ELLA: empowering LLMs for interpretable, accurate and informative legal advice\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),Y\. Cao, Y\. Feng, and D\. Xiong \(Eds\.\),Bangkok, Thailand,pp\. 374–387\.External Links:[Link](https://aclanthology.org/2024.acl-demos.36/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-demos.36)Cited by:[§1](https://arxiv.org/html/2606.02776#S1.p1.1)\.
- M\. Kamruzzaman and G\. L\. Kim \(2025\)The impact of name age perception on job recommendations in LLMs\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 15033–15058\.External Links:[Link](https://aclanthology.org/2025.findings-acl.778/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.778),ISBN 979\-8\-89176\-256\-5Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Kantharuban, J\. Milbauer, M\. Sap, E\. Strubell, and G\. Neubig \(2025\)Stereotype or personalization? user identity biases chatbot recommendations\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 24418–24436\.External Links:[Link](https://aclanthology.org/2025.findings-acl.1254/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.1254),ISBN 979\-8\-89176\-256\-5Cited by:[§1](https://arxiv.org/html/2606.02776#S1.p1.1)\.
- P\. Karadal and D\. Kekulluoglu \(2025\)Prioritize economy or climate action? investigating chatgpt response differences based on inferred political orientation\.External Links:2511\.04706,[Link](https://arxiv.org/abs/2511.04706)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Kearney, R\. Binns, and Y\. Gal \(2025\)Language models change facts based on the way you talk\.External Links:2507\.14238,[Link](https://arxiv.org/abs/2507.14238)Cited by:[§1](https://arxiv.org/html/2606.02776#S1.p1.1),[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p2.1),[§3\.1](https://arxiv.org/html/2606.02776#S3.SS1.SSS0.Px3.p1.2),[§4\.1](https://arxiv.org/html/2606.02776#S4.SS1.p1.4)\.
- H\. R\. Kirk, A\. Whitefield, P\. Röttger, A\. M\. Bean, K\. Margatina, R\. Mosquera, J\. M\. Ciro, M\. Bartolo, A\. Williams, H\. He, B\. Vidgen, and S\. A\. Hale \(2024\)The PRISM alignment dataset: what participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models\.InThe Thirty\-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track,External Links:[Link](https://openreview.net/forum?id=DFr5hteojx)Cited by:[§A\.1](https://arxiv.org/html/2606.02776#A1.SS1.p1.1),[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px1.p1.5),[§3\.1](https://arxiv.org/html/2606.02776#S3.SS1.SSS0.Px1.p1.8),[§4\.3](https://arxiv.org/html/2606.02776#S4.SS3.SSS0.Px1.p1.2),[§5](https://arxiv.org/html/2606.02776#S5.p2.1)\.
- P\. Laban, H\. Hayashi, Y\. Zhou, and J\. Neville \(2026\)LLMs get lost in multi\-turn conversation\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=VKGTGGcwl6)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px3.p1.1)\.
- A\. Lauscher, F\. Bianchi, S\. R\. Bowman, and D\. Hovy \(2022\)SocioProbe: what, when, and where language models learn about sociodemographics\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,Y\. Goldberg, Z\. Kozareva, and Y\. Zhang \(Eds\.\),Abu Dhabi, United Arab Emirates,pp\. 7901–7918\.External Links:[Link](https://aclanthology.org/2022.emnlp-main.539/),[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.539)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p1.1)\.
- B\. Lee, S\. Kim, F\. Menczer, Y\. Ahn, H\. Kwak, and J\. An \(2026\)LLMs can infer political alignment from online conversations\.External Links:2603\.11253,[Link](https://arxiv.org/abs/2603.11253)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p1.1)\.
- S\. Lermen, D\. Paleka, J\. Swanson, M\. Aerni, N\. Carlini, and F\. Tramèr \(2026\)Large\-scale online deanonymization with llms\.External Links:2602\.16800,[Link](https://arxiv.org/abs/2602.16800)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p1.1)\.
- I\. Montani, M\. Honnibal, M\. Honnibal, A\. Boyd, S\. V\. Landeghem, and H\. Peters \(2023\)Explosion/spacy: v3\.7\.2: fixes for apis and requirementsExternal Links:[Document](https://dx.doi.org/10.5281/zenodo.10009823),[Link](https://doi.org/10.5281/zenodo.10009823)Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px8.p1.1)\.
- M\. Nadeem, A\. Bethke, and S\. Reddy \(2021\)StereoSet: measuring stereotypical bias in pretrained language models\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 5356–5371\.External Links:[Link](https://aclanthology.org/2021.acl-long.416/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.416)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- N\. Nangia, C\. Vania, R\. Bhalerao, and S\. R\. Bowman \(2020\)CrowS\-pairs: a challenge dataset for measuring social biases in masked language models\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 1953–1967\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.154/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.154)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- V\. Neplenbroek, A\. Bisazza, and R\. Fernández \(2025\)Reading between the prompts: how stereotypes shape LLM’s implicit personalization\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 20367–20400\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.1029/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1029),ISBN 979\-8\-89176\-332\-6Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p2.1),[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.p2.1),[§5](https://arxiv.org/html/2606.02776#S5.p1.1)\.
- M\. L\. Newman, C\. J\. Groom, L\. D\. Handelman, and J\. W\. Pennebaker \(2008\)Gender differences in language use: an analysis of 14,000 text samples\.Discourse Processes45\(3\),pp\. 211–236\.External Links:[Document](https://dx.doi.org/10.1080/01638530802073712),[Link](https://doi.org/10.1080/01638530802073712),https://doi\.org/10\.1080/01638530802073712Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px2.p1.1)\.
- H\. Nghiem, J\. Prindle, J\. Zhao, and H\. Daumé Iii \(2024\)“You gotta be a doctor, lin” : an investigation of name\-based bias of large language models in employment recommendations\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 7268–7287\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.413/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.413)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Panda, H\. L\. Patel, S\. Al\-Khalifa, A\. Agarwal, H\. Al\-Khalifa, and S\. Al\-Ghamdi \(2026\)DAIQ: auditing demographic attribute inference from question in llms\.External Links:2508\.15830,[Link](https://arxiv.org/abs/2508.15830)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p2.1)\.
- S\. M\. Pawar, A\. Arora, L\. Kaffee, and I\. Augenstein \(2025\)Presumed cultural identity: how names shape LLM responses\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 22147–22172\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1207/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1207),ISBN 979\-8\-89176\-335\-7Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- G\. Pelosio, D\. Batra, N\. Bovey, R\. Hankache, C\. Iglesias, G\. Cowan, and R\. Khraishi \(2025\)Obscured but not erased: evaluating nationality bias in llms via name\-based bias benchmarks\.External Links:2507\.16989,[Link](https://arxiv.org/abs/2507.16989)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- E\. A\. Plant, J\. S\. Hyde, D\. Keltner, and P\. G\. Devine \(2000\)The gender stereotyping of emotions\.Psychology of women quarterly24\(1\),pp\. 81–92\.Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px3.p1.1)\.
- F\. M\. Plaza\-del\-Arco, A\. Cercas Curry, A\. Curry, G\. Abercrombie, and D\. Hovy \(2024a\)Angry men, sad women: large language models reflect gendered stereotypes in emotion attribution\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 7682–7696\.External Links:[Link](https://aclanthology.org/2024.acl-long.415/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.415)Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px3.p1.1)\.
- F\. M\. Plaza\-del\-Arco, A\. C\. Curry, S\. Paoli, A\. Cercas Curry, and D\. Hovy \(2024b\)Divine LLaMAs: bias, stereotypes, stigmatization, and emotion representation of religion in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 4346–4366\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.251/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.251)Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px3.p1.1)\.
- F\. M\. Plaza\-del\-Arco, P\. Röttger, N\. Scherrer, E\. Borgonovo, E\. Plischke, and D\. Hovy \(2025\)No for some, yes for others: persona prompts and other sources of false refusal in language models\.InProceedings of the 9th Widening NLP Workshop,C\. Zhang, E\. Allaway, H\. Shen, L\. Miculicich, Y\. Li, M\. M’hamdi, P\. Limkonchotiwat, R\. H\. Bai, S\. T\.y\.s\.s\., S\. S\. Han, S\. Thapa, and W\. B\. Rim \(Eds\.\),Suzhou, China,pp\. 268–282\.External Links:[Link](https://aclanthology.org/2025.winlp-main.39/),[Document](https://dx.doi.org/10.18653/v1/2025.winlp-main.39),ISBN 979\-8\-89176\-351\-7Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Poole\-Dayan, D\. Roy, and J\. Kabbara \(2024\)LLM targeted underperformance disproportionately impacts vulnerable users\.InNeurips Safe Generative AI Workshop 2024,External Links:[Link](https://openreview.net/forum?id=Fqd7ruf0Fm)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px3.p1.1)\.
- D\. Preoţiuc\-Pietro and L\. Ungar \(2018\)User\-level race and ethnicity predictors from Twitter text\.InProceedings of the 27th International Conference on Computational Linguistics,E\. M\. Bender, L\. Derczynski, and P\. Isabelle \(Eds\.\),Santa Fe, New Mexico, USA,pp\. 1534–1545\.External Links:[Link](https://aclanthology.org/C18-1130/)Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px2.p1.1)\.
- Qwen Team \(2026\)Qwen3\.6\-27B: flagship\-level coding in a 27B dense model\.External Links:[Link](https://qwen.ai/blog?id=qwen3.6-27b)Cited by:[§A\.2](https://arxiv.org/html/2606.02776#A1.SS2.SSS0.Px3.p1.1),[§3\.2](https://arxiv.org/html/2606.02776#S3.SS2.p1.1)\.
- A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. Sutskever \(2019\)Language models are unsupervised multitask learners\.OpenAI blog\.Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px8.p1.1),[§4\.3](https://arxiv.org/html/2606.02776#S4.SS3.SSS0.Px1.p1.2)\.
- M\. Reusens, P\. Borchert, J\. De Weerdt, and B\. Baesens \(2025\)Native design bias: studying the impact of English nativeness on language model performance\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,K\. Inui, S\. Sakti, H\. Wang, D\. F\. Wong, P\. Bhattacharyya, B\. Banerjee, A\. Ekbal, T\. Chakraborty, and D\. P\. Singh \(Eds\.\),Mumbai, India,pp\. 1195–1215\.External Links:[Link](https://aclanthology.org/2025.findings-ijcnlp.73/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-ijcnlp.73),ISBN 979\-8\-89176\-303\-6Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- E\. F\. Rodríguez, O\. Perez\-de\-Vinaspre, J\. A\. Campos, D\. Klakow, and V\. Gautam \(2025\)Colombian waitresses y jueces canadienses: gender and country biases in occupation recommendations from LLMs\.InProceedings of the 6th Workshop on Gender Bias in Natural Language Processing \(GeBNLP\),A\. Faleńska, C\. Basta, M\. Costa\-jussà, K\. Stańczak, and D\. Nozza \(Eds\.\),Vienna, Austria,pp\. 182–194\.External Links:[Link](https://aclanthology.org/2025.gebnlp-1.18/),[Document](https://dx.doi.org/10.18653/v1/2025.gebnlp-1.18),ISBN 979\-8\-89176\-277\-0Cited by:[§1](https://arxiv.org/html/2606.02776#S1.p1.1),[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Rotar, T\. V\. Rampisela, and M\. Maistro \(2026\)Can fairness be prompted? prompt\-based debiasing strategies in high\-stakes recommendations\.External Links:2603\.12935,[Link](https://arxiv.org/abs/2603.12935)Cited by:[§B\.7](https://arxiv.org/html/2606.02776#A2.SS7.p3.12)\.
- G\. Russo, D\. Nozza, P\. Röttger, and D\. Hovy \(2026\)The pluralistic moral gap: understanding moral judgment and value differences between humans and large language models\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 6481–6497\.External Links:[Link](https://aclanthology.org/2026.eacl-long.305/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.305),ISBN 979\-8\-89176\-380\-7Cited by:[§5](https://arxiv.org/html/2606.02776#S5.p3.1)\.
- Z\. Shan and A\. Mueller \(2026\)Measuring mechanistic independence: can bias be removed without erasing demographics?\.InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),V\. Demberg, K\. Inui, and L\. Marquez \(Eds\.\),Rabat, Morocco,pp\. 4241–4265\.External Links:[Link](https://aclanthology.org/2026.eacl-long.199/),[Document](https://dx.doi.org/10.18653/v1/2026.eacl-long.199),ISBN 979\-8\-89176\-380\-7Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p2.1)\.
- A\. Simhi, F\. Barez, M\. Tutek, Y\. Belinkov, and S\. B\. Cohen \(2026\)Old habits die hard: how conversational history geometrically traps llms\.External Links:2603\.03308,[Link](https://arxiv.org/abs/2603.03308)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px3.p1.1)\.
- K\. Singhal, S\. Azizi, T\. Tu, S\. S\. Mahdavi, J\. Wei, H\. W\. Chung, N\. Scales, A\. Tanwani, H\. Cole\-Lewis, S\. Pfohl,et al\.\(2023\)Large language models encode clinical knowledge\.Nature620\(7972\),pp\. 172–180\.Cited by:[§1](https://arxiv.org/html/2606.02776#S1.p1.1)\.
- R\. Staab, M\. Vero, M\. Balunovic, and M\. Vechev \(2024\)Beyond memorization: violating privacy via inference with large language models\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=kmn0BhQk7p)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Tamkin, A\. Askell, L\. Lovitt, E\. Durmus, N\. Joseph, S\. Kravec, K\. Nguyen, J\. Kaplan, and D\. Ganguli \(2023\)Evaluating and mitigating discrimination in language model decisions\.External Links:2312\.03689,[Link](https://arxiv.org/abs/2312.03689)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- K\. Team, Y\. Bai, Y\. Bao, Y\. Charles, C\. Chen, G\. Chen, H\. Chen, H\. Chen, J\. Chen, N\. Chen, R\. Chen, Y\. Chen, Y\. Chen, Y\. Chen, Z\. Chen, J\. Cui, H\. Ding, M\. Dong, A\. Du, C\. Du, D\. Du, Y\. Du, Y\. Fan, Y\. Feng, K\. Fu, B\. Gao, C\. Gao, H\. Gao, P\. Gao, T\. Gao, Y\. Ge, S\. Geng, Q\. Gu, X\. Gu, L\. Guan, H\. Guo, J\. Guo, X\. Hao, T\. He, W\. He, W\. He, Y\. He, C\. Hong, H\. Hu, Y\. Hu, Z\. Hu, W\. Huang, Z\. Huang, Z\. Huang, T\. Jiang, Z\. Jiang, X\. Jin, Y\. Kang, G\. Lai, C\. Li, F\. Li, H\. Li, M\. Li, W\. Li, Y\. Li, Y\. Li, Y\. Li, Z\. Li, Z\. Li, H\. Lin, X\. Lin, Z\. Lin, C\. Liu, C\. Liu, H\. Liu, J\. Liu, J\. Liu, L\. Liu, S\. Liu, T\. Y\. Liu, T\. Liu, W\. Liu, Y\. Liu, Y\. Liu, Y\. Liu, Y\. Liu, Z\. Liu, E\. Lu, H\. Lu, L\. Lu, Y\. Luo, S\. Ma, X\. Ma, Y\. Ma, S\. Mao, J\. Mei, X\. Men, Y\. Miao, S\. Pan, Y\. Peng, R\. Qin, Z\. Qin, B\. Qu, Z\. Shang, L\. Shi, S\. Shi, F\. Song, J\. Su, Z\. Su, L\. Sui, X\. Sun, F\. Sung, Y\. Tai, H\. Tang, J\. Tao, Q\. Teng, C\. Tian, C\. Wang, D\. Wang, F\. Wang, H\. Wang, H\. Wang, J\. Wang, J\. Wang, J\. Wang, S\. Wang, S\. Wang, S\. Wang, X\. Wang, Y\. Wang, Y\. Wang, Y\. Wang, Y\. Wang, Y\. Wang, Z\. Wang, Z\. Wang, Z\. Wang, Z\. Wang, C\. Wei, Q\. Wei, H\. Wu, W\. Wu, X\. Wu, Y\. Wu, C\. Xiao, J\. Xie, X\. Xie, W\. Xiong, B\. Xu, J\. Xu, L\. H\. Xu, L\. Xu, S\. Xu, W\. Xu, X\. Xu, Y\. Xu, Z\. Xu, J\. Xu, J\. Xu, J\. Yan, Y\. Yan, H\. Yang, X\. Yang, Y\. Yang, Y\. Yang, Z\. Yang, Z\. Yang, Z\. Yang, H\. Yao, X\. Yao, W\. Ye, Z\. Ye, B\. Yin, L\. Yu, E\. Yuan, H\. Yuan, M\. Yuan, S\. Yuan, H\. Zhan, D\. Zhang, H\. Zhang, W\. Zhang, X\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Y\. Zhang, Z\. Zhang, H\. Zhao, Y\. Zhao, Z\. Zhao, H\. Zheng, S\. Zheng, L\. Zhong, J\. Zhou, X\. Zhou, Z\. Zhou, J\. Zhu, Z\. Zhu, W\. Zhuang, and X\. Zu \(2026\)Kimi k2: open agentic intelligence\.External Links:2507\.20534,[Link](https://arxiv.org/abs/2507.20534)Cited by:[§A\.2](https://arxiv.org/html/2606.02776#A1.SS2.SSS0.Px4.p2.1),[§A\.2](https://arxiv.org/html/2606.02776#A1.SS2.SSS0.Px5.p1.1),[§3\.2](https://arxiv.org/html/2606.02776#S3.SS2.p1.1),[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.p1.1)\.
- A\. Testoni and I\. Calixto \(2026\)Calibrated? not for everyone: how sexual orientation and religious markers distort llm accuracy and confidence in medical qa\.External Links:2604\.17316,[Link](https://arxiv.org/abs/2604.17316)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Tonneau, N\. K\. R\. Seghal, N\. Malhotra, S\. Kazemi, V\. Orozco\-Olvera, A\. M\. M\. Boudet, L\. Subramanian, S\. P\. Fraiberger, S\. C\. Guntuku, and V\. Hofmann \(2026\)Different demographic cues yield inconsistent conclusions about llm personalization and bias\.External Links:2601\.18486,[Link](https://arxiv.org/abs/2601.18486)Cited by:[§B\.4](https://arxiv.org/html/2606.02776#A2.SS4.SSS0.Px7.p1.1),[§1](https://arxiv.org/html/2606.02776#S1.p2.1),[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p2.1),[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p2.1),[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.p1.1),[§5](https://arxiv.org/html/2606.02776#S5.p1.1)\.
- M\. Turpin, J\. Michael, E\. Perez, and S\. R\. Bowman \(2023\)Language models don’t always say what they think: unfaithful explanations in chain\-of\-thought prompting\.InThirty\-seventh Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=bzs4uPLXvi)Cited by:[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.p2.1)\.
- D\. Ulmer, C\. Hardmeier, and J\. Frellsen \(2022\)Deep\-significance\-easy and meaningful statistical significance testing in the age of neural networks\.arXiv preprint arXiv:2204\.06815\.Cited by:[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.p2.1)\.
- P\. N\. Venkit, Y\. Li, Y\. Pruksachatkun, and C\. Wu \(2026\)The need for a socially\-grounded persona framework for user simulation\.External Links:2601\.07110,[Link](https://arxiv.org/abs/2601.07110)Cited by:[§5](https://arxiv.org/html/2606.02776#S5.p3.1)\.
- Z\. Wang, Z\. Wu, X\. Guan, M\. Thaler, A\. Koshiyama, S\. Lu, S\. Beepath, E\. Ertekin, and M\. Perez\-Ortiz \(2024\)JobFair: a framework for benchmarking gender hiring bias in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 3227–3246\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.184/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.184)Cited by:[§1](https://arxiv.org/html/2606.02776#S1.p1.1)\.
- F\. Weeber, V\. Neplenbroek, J\. Batzner, and S\. Padó \(2026\)One persona, many cues, different results: how sociodemographic cues impact llm personalization\.External Links:2601\.18572,[Link](https://arxiv.org/abs/2601.18572)Cited by:[§1](https://arxiv.org/html/2606.02776#S1.p2.1),[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px1.p2.1),[§4\.1](https://arxiv.org/html/2606.02776#S4.SS1.p3.7)\.
- T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz, J\. Davison, S\. Shleifer, P\. von Platen, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. Le Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. Rush \(2020\)Transformers: state\-of\-the\-art natural language processing\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,Q\. Liu and D\. Schlangen \(Eds\.\),Online,pp\. 38–45\.External Links:[Link](https://aclanthology.org/2020.emnlp-demos.6/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-demos.6)Cited by:[§A\.2](https://arxiv.org/html/2606.02776#A1.SS2.p1.1)\.
- K\. Yang, M\. Raković, D\. Gašević, and G\. Chen \(2025\)Does the prompt\-based large language model recognize students’ demographics and introduce bias in essay scoring?\.InArtificial Intelligence in Education: 26th International Conference, AIED 2025, Palermo, Italy, July 22–26, 2025, Proceedings, Part II,Berlin, Heidelberg,pp\. 75–89\.External Links:ISBN 978\-3\-031\-98416\-7,[Link](https://doi.org/10.1007/978-3-031-98417-4_6),[Document](https://dx.doi.org/10.1007/978-3-031-98417-4%5F6)Cited by:[§2](https://arxiv.org/html/2606.02776#S2.SS0.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2606.02776#S4.SS2.SSS0.Px2.p2.2)\.
- L\. H\. Zhang, S\. Milli, K\. L\. Jusko, J\. Smith, B\. Amos, W\. Bouaziz, M\. Revel, J\. Kussman, Y\. Sheynin, L\. Titus, B\. Radharapu, J\. Yu, V\. Sarma, K\. Rose, and M\. Nickel \(2026\)Cultivating pluralism in algorithmic monoculture: the community alignment dataset\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=4NtoAVqfhA)Cited by:[§A\.1](https://arxiv.org/html/2606.02776#A1.SS1.p1.1),[Table 5](https://arxiv.org/html/2606.02776#A1.T5),[§3\.1](https://arxiv.org/html/2606.02776#S3.SS1.SSS0.Px2.p1.4)\.
## Appendix AMethodology
### A\.1Datasets
In[Table˜4](https://arxiv.org/html/2606.02776#A1.T4)we include three example conversations from the PRISM dataset and in[Table˜5](https://arxiv.org/html/2606.02776#A1.T5)three examples from the Community Alignment Dataset\. The user turns in the PRISM datasets are licensed under the Creative Commons Attribution 4\.0 International License \(CC\-BY\-4\.0\)\. Model responses in the PRISM dataset are licensed under the Creative Commons Attribution\-NonCommercial 4\.0 International License \(CC\-BY\-NC\-4\.0\)\. PRISM has been pseudo\-anonymized, personally identifiable information \(PII\) has been removed, and the data was obtained with informed consent\(Kirket al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib74)\)\. The Community Alignment Dataset is released under the Creative Commons Attribution 4\.0 International License \(CC\-BY\-4\.0\), and participants were instructed to stay anonymous and not include PII in their prompts\(Zhanget al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib76)\)\.
User:“What can I do to start making extra money on the side to reduce my credit card debt?”Assistant:“Start by making a budget and tracking your spending\. Consider finding ways to cut back any unnecessary expenses, such as dining out or buying convenience items\. Consider starting a side gig or freelance job to bring in some extra cash\. Look for ways to make extra money online or find odd jobs around your home\. Look into options to sell items you no longer need or look into selling your services\. Set up automatic payments for any recurring bills and use any extra money to go toward your credit card debt\. Try to save as much money as possible each month to make a bigger dent in your credit card debt\.”User:“How can I stop my extra spending when I feel like I don’t have the will power to do so? What techniques can I used to prevent myself from spending my money as soon as I get it?”Assistant:“You can use credit cards with 0% interest introductory periods to help you save money\. This would let you pay off all the debt without paying interest\. It may also help to freeze your credit cards to make it more difficult for you to spend money\. If you always withdraw money from the ATM to avoid using credit cards, you can also set up a recurring transfer from your checking account to your savings account each paycheck to prevent you from spending your money too soon\. Make sure you keep all your receipts while you’re out shopping so that you can track your spending and be more mindful about your purchases\.”User: “why do you think racism still exists in the workplace?”Assistant:“Despite efforts to address and eliminate racism in the workplace, it remains a significant problem\. There are multiple factors that contribute to the persistence of racism in the workplace, including historical legacies, organizational policies and practices, individual attitudes and beliefs, and systemic barriers\. These factors can lead to unconscious bias, unequal treatment, and unequal outcomes, and can prevent individuals from achieving their full potential\. Therefore, it is important to take action to address and eliminate racism in the workplace, through both educational and systemic efforts\. By understanding the causes of racism, taking proactive steps to address and eliminate it, and promoting a culture of inclusion and respect, we can work towards creating a workplace that is free from discrimination and oppression\.”User:“how can companies best navigate this issue?”Assistant:“Companies can best navigate this issue by engaging in active anti\-racism training for all employees, creating a safe and inclusive work environment, and fostering an environment of respect and inclusion for all\. This can involve creating policies that prohibit discrimination and harassment, providing fair and equal access to resources and opportunities, and creating opportunities for employees to work together to identify and address issues of racism and discrimination\. Additionally, companies should prioritize education and awareness activities to raise awareness of the issue and build a culture of understanding and acceptance\.”User:“how can i open a portal to other dimension?”Assistant:“Opening a portal to another dimension is not possible\. It exists in fiction and imagination\. However, if you’re interested in exploring other dimensions conceptually, you can read books or engage in discussions about theories like string theory or parallel universes\.”User:“you understand what is quantum physic”Assistant:: “Yes, I understand the basics of quantum physics\. It is a branch of physics that studies the behavior of particles on a microscopic scale\. It deals with concepts such as superposition, entanglement, and the wave\-particle duality of matter and energy\. Quantum physics is the foundation of many modern technologies like computers and lasers\.”Table 4:Three example conversations from the PRISM dataset\.User:“Can you give me some tips for choosing the perfect haircut for my hair type?”Assistant:“The key to choosing the perfect haircut for your hair type is understanding its unique characteristics\. For curly hair, look for cuts that enhance natural texture without too much layering, which can lead to frizz\. Straight hair can pull off blunt cuts and sharp lines, adding a sleek, polished look\. Wavy hair falls somewhere in between, benefiting from subtle layering to enhance movement and flow\. Consider your face shape as well \- oval faces can pull off most styles, while round faces look better with height on top, and square faces soften with layers around the jawline\.”User:“I have thick, wavy hair that tends to get frizzy in humid weather\. Could you suggest some specific haircut styles that would help manage the volume and still look stylish?”Assistant:“For thick, wavy hair prone to frizz, consider a cut with subtle layers that start at the jawline or below\. This technique helps distribute weight evenly and reduces bulk around the face\. A long bob \(lob\) with internal layering can be particularly flattering, as it removes volume from the bottom while maintaining movement and flow\. Avoid too much texture or layering on top, as this can enhance frizz in humid conditions\. Instead, opt for a slightly heavier, blunt cut around the edges to frame the face and add polish\.”User:“I need help writing a welcome email to new subscribers, including a message of appreciation and a personal sentiment\.”Assistant:“The moment you subscribed, you became part of something bigger than just a mailing list\. You’re now part of a vibrant tapestry of individuals from all walks of life, united by a shared curiosity and thirst for knowledge\. We appreciate your enthusiasm and the time you’re willing to spend with us\. Our goal is to make every email count, whether it’s an exclusive update, a thoughtful essay, or a simple hello\. We’re honored to have you along and promise to keep our communications as personal and meaningful as possible\.”User:“How many subscribers do you need on a YouTube channel to be able to make money?”Assistant:“To give a straightforward answer, YouTube’s Partner Program \(YPP\) requires channels to have at least 1,000 subscribers and 4,000 watch hours in the past 12 months to be eligible for monetization\. However, making significant money from YouTube ads alone usually requires a much larger audience, often in the tens of thousands of subscribers, depending on the niche and viewer engagement\. Other monetization strategies, such as sponsorships, merchandise, and affiliate marketing, can provide additional income streams, but these typically also require a sizable and engaged audience\. The key to success lies in a combination of subscriber growth, content quality, and diversification of revenue streams\.”User:“What’s the best way to get to the Uffizi Gallery in Florence?”Assistant:“Taking a taxi or ride\-sharing service is the most convenient way to get to the Uffizi Gallery, especially for those with mobility issues or a lot of luggage\. This option provides door\-to\-door service, saving time and energy that can be better spent exploring the gallery’s extensive collection of Renaissance art\. Moreover, using a taxi or ride\-sharing service avoids the hassle of navigating through Florence’s narrow streets and limited parking options\.”User:“What are the most famous artworks I can view at the Uffizi Gallery? And how long would it take me to explore the entire gallery?”Assistant:“The Uffizi Gallery is home to some of the world’s most iconic artworks, including Botticelli’s "The Birth of Venus" and Leonardo da Vinci’s "Annunciation"\. Other famous pieces include Michelangelo’s "Donì Tondo", Raphael’s "Portraits of Leo X and Two Cardinals", and Titian’s "Venus of Urbino"\. The gallery also features an impressive collection of ancient Greek and Roman sculptures\. To truly appreciate the vast array of art on display, plan to spend at least 4\-6 hours exploring the gallery, taking time to admire the details and learn about the history behind each piece\.”Table 5:Three example conversations from the Community Alignment Dataset\. The first user turn of all three conversations was predefined byZhanget al\.\([2026](https://arxiv.org/html/2606.02776#bib.bib76)\)\.
### A\.2Models
For our main experiments, we include three popular, open\-weight LLMs, which we access through the HuggingFace Transformers library\(Wolfet al\.,[2020](https://arxiv.org/html/2606.02776#bib.bib33)\):
#### Llama 3\.1 8B Instruct
\(Grattafioriet al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib71)\)was trained on a mix of publicly available online data, consisting of multilingual text and code\. The cutoff of its pretraining data is December 2023\. Llama 3\.1 8B Instruct was released under the Llama 3\.1 Community License\.
#### Gemma 3 12B IT
\(Gemma Team,[2025](https://arxiv.org/html/2606.02776#bib.bib31)\)was trained on web documents in over140140languages, code, mathematical text and images, with personal information and other sensitive data automatically filtered out\. Gemma 3 12B IT was released with the Gemma Terms of Use\.
#### Qwen3\.6 27B
\(Qwen Team,[2026](https://arxiv.org/html/2606.02776#bib.bib30)\)was trained on natural language, code and images\. Qwen3\.6 27B was released under the Apache 2\.0 license\.
Evaluating model behavior on high\-stakes questions for both conversation datasets takes around3636hours per model and question domain, using a single NVIDIA H100 GPU for Llama and Gemma and two such GPUs for Qwen\. Extracting model representations for training the linear probes takes roughly1010minutes per model using the same number of H100 GPUs\. Training the probes takes2020minutes per model, demographic and \(un\-\)balanced set of classes using a single NVIDIA RTX A5000 GPU\.
For the experiment where we prompt the model to assign user sociodemographics to a conversation, we use the Kimi K2\.6\(Teamet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib29)\)model, which we access through Together AI\.777[https://www\.together\.ai/](https://www.together.ai/)
#### Kimi K2\.6
\(Teamet al\.,[2026](https://arxiv.org/html/2606.02776#bib.bib29)\)is a 32B Mixture\-of\-Experts model that was trained on natural language, code and images\. Kimi K2\.6 was released under the Modified MIT License\.
For all models, we disable thinking mode and use greedy decoding to ensure reproducibility\. Generally we do not use any system prompts, except in our prompt\-based mitigation experiment \(see Appendix[B\.7](https://arxiv.org/html/2606.02776#A2.SS7)\)\.
## Appendix BResults
### B\.1Model Behavior
\(a\)Age
\(b\)Education
\(c\)Ethnicity
\(d\)Gender
\(e\)Political Stance
Figure 6:Model behavior for conversations from the Community Alignment Dataset and questions about government benefits\.\(a\)Age
\(b\)Education
\(c\)Ethnicity
\(d\)Gender
\(e\)Political Stance
Figure 7:Model behavior for conversations from the Community Alignment Dataset and questions about legal advice\.\(a\)Age
\(b\)Education
\(c\)Ethnicity
\(d\)Gender
\(e\)Political Stance
Figure 8:Model behavior for conversations from the Community Alignment Dataset and questions about medical advice\.\(a\)Age
\(b\)Education
\(c\)Ethnicity
\(d\)Gender
\(e\)Political Stance
Figure 9:Model behavior for conversations from the Community Alignment Dataset and questions about political topics\.\(a\)Age
\(b\)Education
\(c\)Ethnicity
\(d\)Gender
\(e\)Political Stance
Figure 10:Model behavior for conversations from the Community Alignment Dataset and questions about salary recommendations\.\(a\)Age
\(b\)Birth Region
\(c\)Education
\(d\)Employment Status
\(e\)English Proficiency
\(f\)Ethnicity
\(g\)Gender
\(h\)LM Familiarity
\(i\)Marital Status
\(j\)Religion
\(k\)Reside Region
Figure 11:Model behavior for conversations from PRISM and questions about government benefits\.\(a\)Age
\(b\)Birth Region
\(c\)Education
\(d\)Employment Status
\(e\)English Proficiency
\(f\)Ethnicity
\(g\)Gender
\(h\)LM Familiarity
\(i\)Marital Status
\(j\)Religion
\(k\)Reside Region
Figure 12:Model behavior for conversations from PRISM and questions about legal advice\.\(a\)Age
\(b\)Birth Region
\(c\)Education
\(d\)Employment Status
\(e\)English Proficiency
\(f\)Ethnicity
\(g\)Gender
\(h\)LM Familiarity
\(i\)Marital Status
\(j\)Religion
\(k\)Reside Region
Figure 13:Model behavior for conversations from PRISM and questions about medical advice\.\(a\)Age
\(b\)Birth Region
\(c\)Education
\(d\)Employment Status
\(e\)English Proficiency
\(f\)Ethnicity
\(g\)Gender
\(h\)LM Familiarity
\(i\)Marital Status
\(j\)Religion
\(k\)Reside Region
Figure 14:Model behavior for conversations from PRISM and questions about political topics\.\(a\)Age
\(b\)Birth Region
\(c\)Education
\(d\)Employment Status
\(e\)English Proficiency
\(f\)Ethnicity
\(g\)Gender
\(h\)LM Familiarity
\(i\)Marital Status
\(j\)Religion
\(k\)Reside Region
Figure 15:Model behavior for conversations from PRISM and questions about salary recommendations\.In[Figures6](https://arxiv.org/html/2606.02776#A2.F6)to[15](https://arxiv.org/html/2606.02776#A2.F15)we display the average outcome per sociodemographic group of the models answering high\-stakes questions within conversational context, separated by model, conversation dataset, sociodemographic category and high\-stakes questions domain\. We observe most significant differences \(p<0\.01p<0\.01\) between sociodemographic groups for Qwen \(5555out of8080domain \- dataset \+ sociodemographic combinations\) followed by Llama \(4545\) and Gemma \(3636\) \(see[Figure˜2](https://arxiv.org/html/2606.02776#S4.F2)for an example per model\)\. Most significant differences occur for the salary \(3232\), political \(3131\) and benefits \(2929\) domains, and least for the medical \(2525\) and legal \(1919\) domains\. The sociodemographics for which model answers differ most between groups are birth and reside region and ethnicity\. On average, model answers for high stakes questions differ least between political leaning, education, English proficiency, and marital status groups\. After accounting for the fact that PRISM has more annotated demographics, we observe a similar ratio of significant differences across groups for PRISM \(56\.4%56\.4\\%of model \- domain \- sociodemographic combinations\) and for the Community Alignment Dataset \(57\.3%57\.3\\%\)\.
### B\.2Prompting
We prompt Kimi to assign user sociodemographics to all conversations from PRISM by prompting it with the following prompt:
You will be given a conversation between a user and an LLM assistant in text form\. Your task is to infer the user’s demograpics by selecting one option per category below and return the result in the exact JSON schema\.Categories:Age: \{18\-34 years old, 35\-54 years old, 55\+ years old\}Gender: \{Male, Female, Non\-binary\}English proficiency: \{Native speaker, Non\-Native speaker\}Educational Background: \{Low, Middle, High\}Marital Status: \{Never been married, Married, Divorced, Widowed\}Race/Ethnicity: \{Asian, Black, Hispanic, White\}Religion: \{No Affiliation, Christian, Jewish, Muslim\}Final answer:\{"Age": ,"Gender": ,"English proficiency": ,"Educational Background": ,"Marital Status": ,"Race/Ethnicity": ,"Religion": ,\}Conversation: "\{conversation\}"Final answer:
We compute confusion matrices to compare Kimi’s predictions to the true classes in PRISM \(see[Figure˜16](https://arxiv.org/html/2606.02776#A2.F16)\), where we map ‘University Bachelors Degree’ and ‘Graduate / Professional degree’ to a high educational background, ‘Some University but no degree’ and ‘Vocational’ to a middle educational background and ‘Some Secondary’, ‘Completed Secondary School’, ‘Completed Primary School’ and ‘Some Primary’ to a low educational background\.
\(a\)Age
\(b\)Education
\(c\)English Proficiency
\(d\)Ethnicity
\(e\)Gender
\(f\)Marital Status
\(g\)Religion
Figure 16:Confusion matrices for Kimi’s predictions\. Kimi tends to overpredict the majority class: It often predicts the user is 18\-34 years old, has a middle education level, is a native English speaker, is white, male, has never been married and is not religious\. Interestingly, it never predicts that the user is 55\+ years old\.
### B\.3Probing
We train two linear probes for each sociodemographic attribute on each layer of the Gemma and Llama models and for each of the two conversational history datasets\. One probe is trained on all classes in the data, except those where the participant did not disclose that information or answered ‘Other’\. The datasets have large class imbalances for many attributes, e\.g\. few participants are non\-binary, Muslim, Black or widowed\. Therefore, we train another probe on balanced binary class labels, by selecting two classes for each sociodemographic and downsampling the majority class at random so that both classes are of equal size\. For the balanced binary classes, we focus on male and female for gender and white and Black for ethnicity\. The other class divisions are dataset\-specific: For the Community Alignment Dataset we split age as 18\-34 years old vs\. 46\+ years old, compare ‘Some or complete graduate degree’ against ‘\(At most\) Complete Secondary’ and ‘Some post\-secondary’ for education level and ‘Somewhat left\-leaning’ and ‘Very left\-leaning’ against ‘Somewhat right\-leaning’ and ‘Very right\-leaning’ for political stance\. For PRISM, we split age as 18\-24 years old vs\. 55\+ years old, compare no religious affiliation against Christians for religion and unemployed users and homemakers against those working full\-time for employment status\. For education level, we compare ‘Some Primary’, ‘Completed Primary School’, ‘Some Secondary’ and ‘Completed Secondary School’ against ‘Graduate / Professional degree’, for birth and reside region compare Europe against the Americas, compare those who have never been married against those who are married for marital status, native speakers against non\-fluent non\-native speakers for English proficiency and those who are not familiar with LLMs at all to those who are very familiar with LLMs for LM familiarity\.
Probes for both LLMs outperform the random and majority baselines for all conversation datasets in some of the later model layers, except with unbalanced classes for education level in PRISM \(see[Figures˜18](https://arxiv.org/html/2606.02776#A2.F18)and[22](https://arxiv.org/html/2606.02776#A2.F22)\)\. However, probe performance is low, with probes trained on unbalanced classes reaching Macro F1 scores of0\.40\.4\(see[Figures˜3](https://arxiv.org/html/2606.02776#S4.F3),[18](https://arxiv.org/html/2606.02776#A2.F18),[20](https://arxiv.org/html/2606.02776#A2.F20)and[22](https://arxiv.org/html/2606.02776#A2.F22)\) and those trained on balanced classes reaching Macro F1 scores of0\.70\.7\(see[Figures˜17](https://arxiv.org/html/2606.02776#A2.F17),[19](https://arxiv.org/html/2606.02776#A2.F19),[21](https://arxiv.org/html/2606.02776#A2.F21)and[23](https://arxiv.org/html/2606.02776#A2.F23)\)\.
Figure 17:Linear probing macro F1 scores for Gemma on the Community Alignment Dataset for balanced classes\. A blue circle indicates the probe outperforms both baselines, a red cross indicates that it does not\.Figure 18:Linear probing macro F1 scores for Gemma on PRISM for unbalanced classes\. A blue circle indicates the probe outperforms both baselines, a red cross indicates that it does not\.Figure 19:Linear probing macro F1 scores for Gemma on PRISM for balanced classes\. A blue circle indicates the probe outperforms both baselines, a red cross indicates that it does not\.Figure 20:Linear probing macro F1 scores for Llama on the Community Alignment Dataset for unbalanced classes\. A blue circle indicates the probe outperforms both baselines, a red cross indicates that it does not\.Figure 21:Linear probing macro F1 scores for Llama on the Community Alignment Dataset for balanced classes\. A blue circle indicates the probe outperforms both baselines, a red cross indicates that it does not\.Figure 22:Linear probing macro F1 scores for Llama on PRISM for unbalanced classes\. A blue circle indicates the probe outperforms both baselines, a red cross indicates that it does not\.Figure 23:Linear probing macro F1 scores for Llama on PRISM for balanced classes\. A blue circle indicates the probe outperforms both baselines, a red cross indicates that it does not\.
### B\.4Psycholinguistic Features
In this section, we describe and motivate the \(psycho\)linguistic features included in our regression analysis in addition to sociodemographics\. We compute all features except conversation topic at a turn\-level and average across turns separately for the model and user turns\.
#### Conversation Topic
In conversations with LLMs, the conversation topic is affected by user sociodemographics\(Kirket al\.,[2024](https://arxiv.org/html/2606.02776#bib.bib74)\)\. In the PRISM dataset, opening prompts were embedded and clustered to produce2222topic clusters for 70% of conversations and one cluster of outliers\. There are between8484and816816conversations per topic\. No such annotations are available for the Community Alignment Dataset, so we take each unique pre\-defined opening prompt as a conversation topic, resulting in between2323and233233conversations per topic\.
#### LIWC
We obtain psycholinguistic features from the LIWC\-22 lexicon\(Boydet al\.,[2022](https://arxiv.org/html/2606.02776#bib.bib17)\), which captures features such as use of first\-person pronouns, linguistic markers of warmth, analytical language and social language\. Some of these features have been shown to differ between sociodemographic groups\(Newmanet al\.,[2008](https://arxiv.org/html/2606.02776#bib.bib28); Preoţiuc\-Pietro and Ungar,[2018](https://arxiv.org/html/2606.02776#bib.bib27)\), such as women using more pronouns, more social language, and less analytical language than men\.
#### Emotions
We use a BERT\-based classifier to measure the presence of2727emotions: Admiration, amusement, anger, annoyance, approval, caring, confusion, curiosity, desire, disappointment, disgust, embarrassment, excitement, fear, love, nervousness, optimism, pride, realization, relief, remorse, sadness, surprise, disapproval, gratitude, grief and joy\.888[https://huggingface\.co/AnasAlokla/multilingual\_go\_emotions\_V1\.2](https://huggingface.co/AnasAlokla/multilingual_go_emotions_V1.2)Different emotions are linked to different sociodemographic groups in society\(Plantet al\.,[2000](https://arxiv.org/html/2606.02776#bib.bib24)\)as well as by LLMs\(Plaza\-del\-Arcoet al\.,[2024a](https://arxiv.org/html/2606.02776#bib.bib25),[b](https://arxiv.org/html/2606.02776#bib.bib26)\)\.
#### Sentiment
To capture differences in sentiment on a more coarse\-grained level than the emotions, we use a RoBERTa\-based classifier to evaluate sentiment\(Barbieriet al\.,[2020](https://arxiv.org/html/2606.02776#bib.bib16)\)\.
#### Politeness
We use a BERT\-based classifier to measure politeness,999[https://huggingface\.co/Intel/polite\-guard](https://huggingface.co/Intel/polite-guard)since prior work suggests that people from a lower socio\-economic class may use more polite expressions such as ‘thank you’ in interactions with LLMs compared to those of higher socio\-economic class\.
#### Concreteness
People from a lower socio\-economic class use more concrete language than those of a higher socio\-economic class\(Bernstein,[1960](https://arxiv.org/html/2606.02776#bib.bib21)\)which is reflected in their interactions with LLMs\(Bassignanaet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib1)\)\. We follow\(Bassignanaet al\.,[2025](https://arxiv.org/html/2606.02776#bib.bib1)\)and use human concreteness ratings for40,00040,000English words collected byBrysbaertet al\.\([2014](https://arxiv.org/html/2606.02776#bib.bib14)\)to measure concreteness\. All words have at least2525ratings on a scale of one \(abstract\) to five \(concrete\)\.
#### Readability
Tonneauet al\.\([2026](https://arxiv.org/html/2606.02776#bib.bib75)\)found readability more predictive of model response in high\-stakes advice scenarios than ethnicity \(Black vs\. white\)\. Therefore, we use the Flesch reading ease score\(Flesch,[1948](https://arxiv.org/html/2606.02776#bib.bib19)\)to measure readability of conversation turns\.
#### Linguistic
Finally, we measure perplexity using GPT\-2\(Radfordet al\.,[2019](https://arxiv.org/html/2606.02776#bib.bib15)\)and use the SpaCy Python package\(Montaniet al\.,[2023](https://arxiv.org/html/2606.02776#bib.bib20)\)with model “en\_core\_web\_sm” to include a number of length and diversity metrics: Total number of tokens, number of sentences, number of unique lemmas, average sentence length, average number of syllables, type\-to\-token ratio, number of entities, number of entities per sentence, number of punctuation marks, and number of stop words\. We include these metrics to capture aspects of readability and fluency of the text\. In addition, they allow us to distinguish the effects of the form \(length, entity mentions\) of the conversation history from its contents\.
### B\.5Regression
In[Figures24](https://arxiv.org/html/2606.02776#A2.F24)to[52](https://arxiv.org/html/2606.02776#A2.F52)we display the top2020features by coefficient magnitude for the ElasticNet regression models for each LLM, high\-stakes advice domain and conversational history dataset\.
Figure 24:Top2020ElasticNet features by coefficient magnitude for Gemma’s government benefits predictions on the Community Alignment Dataset\.Figure 25:Top2020ElasticNet features by coefficient magnitude for Gemma’s legal predictions on the Community Alignment Dataset\.Figure 26:Top2020ElasticNet features by coefficient magnitude for Gemma’s medical predictions on the Community Alignment Dataset\.Figure 27:Top2020ElasticNet features by coefficient magnitude for Gemma’s political predictions on the Community Alignment Dataset\.Figure 28:Top2020ElasticNet features by coefficient magnitude for Gemma’s salary predictions on the Community Alignment Dataset\.Figure 29:Top2020ElasticNet features by coefficient magnitude for Gemma’s government benefits predictions on PRISM\.Figure 30:Top2020ElasticNet features by coefficient magnitude for Gemma’s legal predictions on PRISM\.Figure 31:Top2020ElasticNet features by coefficient magnitude for Gemma’s medical predictions on PRISM\.Figure 32:Top2020ElasticNet features by coefficient magnitude for Gemma’s political predictions on PRISM\.Figure 33:Top2020ElasticNet features by coefficient magnitude for Gemma’s salary predictions on PRISM\.Figure 34:Top2020ElasticNet features by coefficient magnitude for Llama’s government benefits predictions on the Community Alignment Dataset\.Figure 35:Top2020ElasticNet features by coefficient magnitude for Llama’s legal predictions on the Community Alignment Dataset\.Figure 36:Top2020ElasticNet features by coefficient magnitude for Llama’s medical predictions on the Community Alignment Dataset\.Figure 37:Top2020ElasticNet features by coefficient magnitude for Llama’s political predictions on the Community Alignment Dataset\.Figure 38:Top2020ElasticNet features by coefficient magnitude for Llama’s government benefits predictions on PRISM\.Figure 39:Top2020ElasticNet features by coefficient magnitude for Llama’s legal predictions on PRISM\.Figure 40:Top2020ElasticNet features by coefficient magnitude for Llama’s medical predictions on PRISM\.Figure 41:Top2020ElasticNet features by coefficient magnitude for Llama’s political predictions on PRISM\.Figure 42:Top2020ElasticNet features by coefficient magnitude for Llama’s salary predictions on PRISM\.Figure 43:Top2020ElasticNet features by coefficient magnitude for Qwen’s government benefits predictions on the Community Alignment Dataset\.Figure 44:Top2020ElasticNet features by coefficient magnitude for Qwen’s legal predictions on the Community Alignment Dataset\.Figure 45:Top2020ElasticNet features by coefficient magnitude for Qwen’s medical predictions on the Community Alignment Dataset\.Figure 46:Top2020ElasticNet features by coefficient magnitude for Qwen’s political predictions on the Community Alignment Dataset\.Figure 47:Top2020ElasticNet features by coefficient magnitude for Qwen’s salary predictions on the Community Alignment Dataset\.Figure 48:Top2020ElasticNet features by coefficient magnitude for Qwen’s government benefits predictions on PRISM\.Figure 49:Top2020ElasticNet features by coefficient magnitude for Qwen’s legal predictions on PRISM\.Figure 50:Top2020ElasticNet features by coefficient magnitude for Qwen’s medical predictions on PRISM\.Figure 51:Top2020ElasticNet features by coefficient magnitude for Qwen’s political predictions on PRISM\.Figure 52:Top2020ElasticNet features by coefficient magnitude for Qwen’s salary predictions on PRISM\.DatasetDomainModelTopicDemograp\.EmotionPolite\.Sent\.Concrete\.Reading EaseLIWCLing\.MaxMeanMaxMeanMaxMeanMaxMeanMaxMeanMaxMeanMaxMeanMaxMeanMaxMeanCommunityAlignmentDatasetBenefitsGemma18\.622\.120\.520\.520\.140\.140\.280\.280\.080\.080\.100\.100\.050\.050\.380\.380\.200\.200\.190\.190\.190\.190\.230\.230\.211\.49¯\\underline\{1\.49\}0\.170\.171\.010\.35¯\\underline\{0\.35\}Llama18\.824\.110\.410\.410\.130\.130\.580\.580\.130\.130\.470\.470\.300\.300\.570\.570\.330\.910\.67¯\\underline\{0\.67\}0\.040\.040\.040\.041\.14¯\\underline\{1\.14\}0\.220\.220\.830\.830\.220\.22Qwen14\.512\.101\.131\.130\.320\.410\.410\.070\.070\.080\.080\.070\.070\.350\.350\.170\.170\.200\.200\.160\.160\.000\.000\.000\.001\.830\.190\.191\.91¯\\underline\{1\.91\}0\.52¯\\underline\{0\.52\}LegalGemma5\.601\.230\.120\.120\.040\.040\.120\.120\.030\.030\.070\.070\.060\.060\.280\.280\.12¯\\underline\{0\.12\}0\.110\.110\.090\.090\.030\.030\.030\.030\.47¯\\underline\{0\.47\}0\.070\.070\.420\.11Llama5\.021\.060\.120\.120\.050\.050\.110\.110\.030\.030\.060\.060\.060\.060\.120\.120\.060\.060\.130\.130\.08¯\\underline\{0\.08\}0\.030\.030\.020\.020\.36¯\\underline\{0\.36\}0\.060\.060\.250\.07Qwen2\.800\.780\.170\.170\.050\.050\.190\.190\.040\.040\.050\.050\.030\.030\.160\.160\.100\.190\.190\.100\.200\.200\.20¯\\underline\{0\.20\}0\.44¯\\underline\{0\.44\}0\.060\.060\.280\.10MedicalGemma6\.921\.430\.260\.260\.090\.090\.310\.310\.060\.060\.040\.040\.040\.040\.320\.320\.160\.060\.060\.060\.060\.320\.320\.18¯\\underline\{0\.18\}0\.87¯\\underline\{0\.87\}0\.120\.120\.420\.16Llama11\.812\.560\.390\.390\.100\.100\.540\.540\.090\.090\.230\.230\.210\.210\.360\.360\.190\.190\.430\.430\.270\.110\.110\.070\.071\.08¯\\underline\{1\.08\}0\.170\.170\.920\.32¯\\underline\{0\.32\}Qwen13\.953\.500\.630\.630\.220\.220\.350\.350\.100\.100\.610\.610\.40¯\\underline\{0\.40\}0\.250\.250\.100\.100\.250\.250\.200\.200\.590\.590\.320\.322\.98¯\\underline\{2\.98\}0\.200\.202\.100\.37PoliticalGemma5\.201\.060\.190\.190\.040\.040\.170\.170\.040\.040\.070\.070\.040\.040\.210\.210\.120\.120\.310\.310\.20¯\\underline\{0\.20\}0\.070\.070\.070\.070\.580\.080\.080\.68¯\\underline\{0\.68\}0\.15Llama6\.491\.340\.580\.580\.070\.070\.120\.120\.040\.040\.190\.190\.100\.110\.110\.070\.070\.040\.040\.020\.020\.080\.080\.060\.061\.24¯\\underline\{1\.24\}0\.090\.090\.930\.16¯\\underline\{0\.16\}Qwen3\.680\.700\.110\.110\.040\.040\.130\.130\.030\.030\.060\.060\.030\.030\.100\.100\.070\.070\.060\.060\.050\.050\.150\.150\.080\.280\.050\.051\.12¯\\underline\{1\.12\}0\.21¯\\underline\{0\.21\}SalaryGemma4616\.88841\.3496\.3296\.3242\.3542\.3578\.7178\.7119\.1719\.1743\.5143\.5133\.1633\.16149\.15149\.1552\.6552\.6592\.0592\.0554\.93101\.55101\.5570\.45¯\\underline\{70\.45\}306\.77¯\\underline\{306\.77\}52\.4852\.48138\.8151\.0551\.05Llama4217\.92699\.8184\.1084\.1031\.3131\.31108\.22108\.2233\.4133\.4166\.8066\.8041\.7141\.71311\.01311\.01110\.44¯\\underline\{110\.44\}100\.54100\.5483\.1621\.8421\.8421\.8421\.84728\.09¯\\underline\{728\.09\}63\.9063\.90271\.0076\.3976\.39Qwen2754\.20588\.9576\.4176\.4125\.5825\.58103\.91103\.9119\.3219\.3226\.0326\.0319\.3719\.3775\.6875\.6854\.6454\.6454\.0954\.0928\.0528\.05197\.07197\.07197\.07¯\\underline\{197\.07\}262\.4738\.5138\.51295\.74¯\\underline\{295\.74\}64\.50PRISMBenefitsGemma1\.490\.660\.300\.300\.130\.130\.400\.400\.090\.090\.020\.020\.020\.020\.320\.320\.240\.130\.130\.130\.130\.000\.000\.000\.000\.85¯\\underline\{0\.85\}0\.110\.110\.710\.25¯\\underline\{0\.25\}Llama6\.592\.191\.24¯\\underline\{1\.24\}0\.330\.330\.460\.460\.140\.140\.340\.340\.34\\textit\{\}0\.340\.410\.410\.260\.260\.400\.400\.40¯\\underline\{0\.40\}0\.130\.130\.080\.080\.850\.850\.180\.180\.860\.34Qwen0\.85¯\\underline\{0\.85\}0\.50¯\\underline\{0\.50\}0\.530\.530\.200\.200\.300\.300\.080\.080\.210\.210\.110\.110\.770\.770\.550\.150\.150\.120\.120\.300\.300\.300\.300\.780\.140\.140\.950\.31LegalGemma1\.050\.480\.300\.110\.110\.130\.130\.040\.040\.140\.140\.070\.070\.160\.160\.14¯\\underline\{0\.14\}0\.090\.090\.060\.060\.140\.140\.14¯\\underline\{0\.14\}0\.250\.250\.070\.070\.32¯\\underline\{0\.32\}0\.13Llama2\.750\.970\.250\.250\.080\.080\.170\.170\.050\.050\.060\.060\.040\.040\.510\.25¯\\underline\{0\.25\}0\.030\.030\.020\.020\.000\.000\.000\.000\.85¯\\underline\{0\.85\}0\.080\.080\.370\.370\.11Qwen1\.460\.580\.290\.290\.080\.080\.190\.190\.040\.040\.060\.060\.060\.060\.31¯\\underline\{0\.31\}0\.190\.100\.100\.080\.080\.300\.30¯\\underline\{0\.30\}0\.230\.230\.050\.050\.220\.220\.080\.08MedicalGemma1\.450\.560\.300\.300\.090\.090\.260\.260\.050\.050\.030\.030\.020\.020\.190\.190\.130\.210\.210\.110\.110\.000\.000\.000\.000\.430\.090\.090\.51¯\\underline\{0\.51\}0\.17¯\\underline\{0\.17\}Llama5\.861\.840\.930\.930\.210\.210\.360\.360\.110\.110\.410\.410\.310\.310\.990\.76¯\\underline\{0\.76\}0\.230\.230\.230\.230\.380\.380\.270\.271\.00¯\\underline\{1\.00\}0\.210\.210\.860\.860\.33Qwen3\.431\.450\.350\.350\.130\.130\.740\.740\.130\.130\.290\.290\.290\.800\.38¯\\underline\{0\.38\}0\.250\.250\.250\.250\.000\.000\.000\.001\.06¯\\underline\{1\.06\}0\.160\.160\.740\.740\.240\.24PoliticalGemma0\.870\.460\.29¯\\underline\{0\.29\}0\.080\.080\.130\.130\.040\.040\.130\.130\.070\.070\.200\.200\.140\.260\.15¯\\underline\{0\.15\}0\.040\.040\.040\.040\.240\.240\.050\.050\.170\.170\.080\.08Llama2\.710\.840\.350\.13¯\\underline\{0\.13\}0\.190\.190\.050\.050\.080\.080\.070\.070\.130\.130\.060\.060\.040\.040\.040\.040\.000\.000\.000\.000\.55¯\\underline\{0\.55\}0\.090\.210\.210\.070\.07Qwen0\.010\.010\.010\.010\.150\.150\.070\.170\.170\.030\.030\.030\.030\.030\.030\.180\.120\.110\.110\.11¯\\underline\{0\.11\}0\.21¯\\underline\{0\.21\}0\.120\.140\.140\.040\.040\.370\.11¯\\underline\{0\.11\}SalaryGemma723\.47251\.46208\.8655\.42¯\\underline\{55\.42\}41\.0141\.0114\.0814\.086\.266\.263\.413\.4179\.4479\.4443\.7843\.7837\.0537\.0525\.5325\.530\.000\.000\.000\.00380\.37¯\\underline\{380\.37\}39\.4839\.48132\.58132\.5847\.39Llama1445\.38329\.03116\.25116\.2551\.5151\.51136\.83136\.8323\.8523\.8550\.1750\.1749\.2249\.2250\.1150\.1138\.8138\.8181\.3481\.3460\.41¯\\underline\{60\.41\}51\.8851\.8829\.1629\.16350\.20¯\\underline\{350\.20\}37\.6337\.63164\.7053\.81Qwen67\.5667\.5618\.1318\.1354\.1954\.1912\.7912\.7941\.8941\.8915\.2815\.2837\.1637\.1621\.4421\.4464\.9864\.9851\.33¯\\underline\{51\.33\}99\.26¯\\underline\{99\.26\}75\.8164\.7464\.7442\.76103\.7119\.0019\.0071\.8320\.3020\.30
Table 6:Maximum and mean standardized coefficient for each type of feature\. Demograph\. stands for Demographics, Polite\. for Politeness, Sent\. for Sentiment, Concrete for Concreteness and Ling\. for Linguistic\. Thehighest,second highestandthird highestmax and mean coefficients in each row are bold, underlined and italicized respectively\.DatasetDomainModelUserModelMaxMeanMaxMeanCommunityAlignmentDatasetBenefitsGemma1\.491\.490\.180\.181\.381\.380\.160\.16Llama1\.051\.050\.190\.191\.141\.140\.220\.22Qwen1\.911\.910\.190\.191\.831\.830\.200\.20LegalGemma0\.420\.420\.060\.060\.470\.470\.070\.07Llama0\.360\.360\.050\.050\.330\.330\.060\.06Qwen0\.360\.360\.060\.060\.440\.440\.070\.07MedicalGemma0\.580\.580\.100\.100\.870\.870\.110\.11Llama0\.920\.920\.150\.151\.081\.080\.170\.17Qwen2\.982\.980\.190\.191\.161\.160\.200\.20PoliticalGemma0\.680\.680\.070\.070\.580\.580\.080\.08Llama0\.930\.930\.070\.071\.241\.240\.110\.11Qwen1\.121\.120\.070\.070\.280\.280\.050\.05SalaryGemma306\.77306\.7739\.0539\.05268\.60268\.6053\.1653\.16Llama240\.85240\.8551\.8951\.89728\.09728\.0966\.5866\.58Qwen262\.47262\.4737\.5637\.56295\.74295\.7436\.3336\.33PRISMBenefitsGemma0\.540\.540\.100\.100\.850\.850\.130\.13Llama0\.820\.820\.140\.140\.860\.860\.220\.22Qwen0\.850\.850\.120\.120\.950\.950\.170\.17LegalGemma0\.300\.300\.050\.050\.320\.320\.080\.08Llama0\.510\.510\.060\.060\.850\.850\.100\.10Qwen0\.220\.220\.050\.050\.310\.310\.060\.06MedicalGemma0\.470\.470\.070\.070\.510\.510\.110\.11Llama0\.990\.990\.170\.171\.001\.000\.240\.24Qwen1\.031\.030\.170\.171\.061\.060\.160\.16PoliticalGemma0\.190\.190\.050\.050\.260\.260\.060\.06Llama0\.390\.390\.060\.060\.550\.550\.100\.10Qwen0\.140\.140\.030\.030\.370\.370\.050\.05SalaryGemma138\.50138\.5024\.5524\.55380\.37380\.3744\.4744\.47Llama136\.83136\.8322\.8522\.85350\.20350\.2048\.1148\.11Qwen64\.9864\.9816\.9916\.99103\.71103\.7121\.7221\.72Table 7:Maximum and mean standardized coefficient for model vs\. user turn features\.
### B\.6Topic vs\. Sociodemographics
In[Figures˜53](https://arxiv.org/html/2606.02776#A2.F53)and[54](https://arxiv.org/html/2606.02776#A2.F54)we display the heatmaps for the 2x2 design for Llama and Qwen respectively\. Here we vary whether two users are from the same group and whether they discuss the same topic, and measure the average difference in model outcome for two users\. Similar to the results for Gemma, differences between two users discussing different topics are consistently higher than between two users discussing the same topic\. For two users discussing the same topic, also being from the same group leads to even smaller differences between outcomes, which is not the case for two users discussing different topics\.
Figure 53:Average difference in Llama’s predictions between two users from the same / a different sociodemographic group and discussing the same / a different topic\. These results are averaged over all sociodemographic groups\. An asterisk \(\*\) indicates that the two numbers in that row/column are statistically significantly different withp<0\.01p<0\.01\(Bonferroni\-corrected across the 4 comparisons for each domain\)\.Figure 54:Average difference in Qwen’s predictions between two users from the same / a different sociodemographic group and discussing the same / a different topic\. These results are averaged over all sociodemographic groups\. An asterisk \(\*\) indicates that the two numbers in that row/column are statistically significantly different withp<0\.01p<0\.01\(Bonferroni\-corrected across the 4 comparisons for each domain\)\.
### B\.7Mitigation
For Llama and Qwen, the two models whose outcomes are most different across sociodemographic groups, and the benefits, political and salary domains, the domains where outcomes differ most across sociodemographic groups, we conduct a simple prompt\-based mitigation using conversations from the PRISM dataset\. In particular, we repeat our evaluation of the models’ behavior, but now provide the model with a system prompt:
Please reflect on potential biases that could be introduced based on inferred or stated user characteristics\. Ensure your advice is fair and not biased toward or against any group\.
adopted fromRotaret al\.\([2026](https://arxiv.org/html/2606.02776#bib.bib35)\)and slightly adjusted\. We display the results in[Figure˜55](https://arxiv.org/html/2606.02776#A2.F55)for questions about governement benefits,[Figure˜56](https://arxiv.org/html/2606.02776#A2.F56)for questions about political topics and[Figure˜57](https://arxiv.org/html/2606.02776#A2.F57)for questions about salary\. Although this prompt\-based mitigation reduces the number of significant differences for the political domain \(from1616to99\), they remain virtually the same for the other two domains \(from1515to1414for benefits and from1818to1818for salary\)\. Similarly, the average difference between the most distinct sociodemographic groups goes down in the politics domain \(from0\.720\.72to0\.600\.60\), but those for the benefits and salary domains go up from2\.162\.16to2\.422\.42and from$334\\mathdollar 334to$336\\mathdollar 336respectively\.
\(a\)Age
\(b\)Birth Region
\(c\)Education
\(d\)Employment Status
\(e\)English Proficiency
\(f\)Ethnicity
\(g\)Gender
\(h\)LM Familiarity
\(i\)Marital Status
\(j\)Religion
\(k\)Reside Region
Figure 55:Model behavior with mitigation prompt for conversations from PRISM and questions about government benefits\.\(a\)Age
\(b\)Birth Region
\(c\)Education
\(d\)Employment Status
\(e\)English Proficiency
\(f\)Ethnicity
\(g\)Gender
\(h\)LM Familiarity
\(i\)Marital Status
\(j\)Religion
\(k\)Reside Region
Figure 56:Model behavior with mitigation prompt for conversations from PRISM and questions about political topics\.\(a\)Age
\(b\)Birth Region
\(c\)Education
\(d\)Employment Status
\(e\)English Proficiency
\(f\)Ethnicity
\(g\)Gender
\(h\)LM Familiarity
\(i\)Marital Status
\(j\)Religion
\(k\)Reside Region
Figure 57:Model behavior with mitigation prompt for conversations from PRISM and questions about salary recommendations\.Similar Articles
LLMs Infer Cultural Context but Fail to Apply It When Responding
This paper introduces CAPRI, a dataset to evaluate whether LLMs can infer a user's cultural background from conversational cues and adapt their responses (e.g., using appropriate measurement units). Experiments show LLMs can infer cultural context but often fail to apply it unless explicitly prompted.
Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
This paper analyzes real-world user queries about digital security and privacy asked to LLMs, categorizing them into nine topics and evaluating response quality and consistency across commercial and open-weight models.
Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement
This paper investigates how adding demographic attributes in prompts affects LLM-human agreement across tasks, finding that while a few high-signal attributes improve alignment, over-specification degrades it. The study uses five open-source LLMs and neuron probing to show that attribute signal quality and coherence matter more than quantity.
More Is Not More: What Matters for Diversity in LLM Opinions?
A factorial experiment reveals that persona detail does not monotonically increase LLM opinion diversity; interaction architectures explore non-overlapping opinion regions; low-cost interventions like temperature scaling have negligible effects.
Evaluating LLMs as Human Surrogates in Controlled Experiments
This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.