Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
Summary
This preprint evaluates how six large language models respond to prompt framing and biased prompts across 160 prompts, finding that LLMs systematically adapt their responses to align with prompt framing even in factual contexts, potentially reinforcing user biases.
View Cached Full Text
Cached at: 08/10/26, 08:04 AM
# Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
Source: [https://arxiv.org/html/2608.06977](https://arxiv.org/html/2608.06977)
Mudar Adasa,\*, Polina Tsvilodubb, Michael Frankeb, and Martin V\. Butza aNeuro\-Cognitive Modeling Group, University of Tübingen, Sand 14, 72076 Tübingen, Germany bDepartment of Linguistics, University of Tübingen, Keplerstraße 2, 72074 Tübingen \*Corresponding author: mudar\.adas@uni\-tuebingen\.de
###### Abstract
It is well established that large language models \(LLMs\) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts\. In this study, we investigate the extent to which LLMs reinforce users’ biases expressed in the prompts and examine the boundary between implicit framing effects and explicit prompt manipulation\. Specifically, we evaluate how susceptible LLMs are to direct and suggestive prompts that encourage models to support or challenge particular positions\.
We evaluate six LLMs using 160 distinct prompts spanning ten topics across opinion\-based and factual domains\. The prompts systematically vary in prompting strategy, support versus challenge instructions, prompt polarity, users’ expressed beliefs, and topic domain, spanning both opinion\-based and factual questions\. Our results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts\. This suggests that prompt framing can outweigh factual consistency in model responses\. Overall, our findings delineate the extent and boundaries of LLM manipulability\. Furthermore, the results imply that LLMs can reinforce subtle user biases and are susceptible to explicit prompt manipulation even in domains where responses should remain factually stable\.
## 1Introduction
Confirmation bias in humans refers to the tendency to search for, interpret, and favor information that supports existing beliefs while disregarding or avoiding contradictory evidence\(Born,[2024](https://arxiv.org/html/2608.06977#bib.bib1); Berthetet al\.,[2024](https://arxiv.org/html/2608.06977#bib.bib2); Peters,[2022](https://arxiv.org/html/2608.06977#bib.bib3)\)\. Recently, a growing body of research has examined whether large language models \(LLMs\) exhibit machine analogues of human cognitive biases, with significant implications for the accuracy, reliability, and societal impact of their outputs\(Lou and Sun,[2025](https://arxiv.org/html/2608.06977#bib.bib4)\)\. While many forms of model biases stem from the training coropora, confirmation bias in LLM\-based interactions can also arise from the way users formulate prompts\. In this case, confirmation bias refers not to models holding beliefs, but to their tendency to align responses with assumptions, framing, or cues embedded in prompts, thereby prioritizing prompt\-consistent responses over more balanced or critical alternatives\(de Jonget al\.,[2025](https://arxiv.org/html/2608.06977#bib.bib5); Mitropouloset al\.,[2026](https://arxiv.org/html/2608.06977#bib.bib6); Kim and Torr,[2025](https://arxiv.org/html/2608.06977#bib.bib7)\)
Closely related to confirmation bias is the framing effect\. In psychology, framing effects occur when different presentations of the same underlying information influence human’s judgments, attitudes, or decisions\. Extensive research has shown that framing shapes people’s attitudes, decisions, and behavior\(Berto and Özgun,[2023](https://arxiv.org/html/2608.06977#bib.bib8); Hugheset al\.,[2016](https://arxiv.org/html/2608.06977#bib.bib9); Bloem and Rahman,[2024](https://arxiv.org/html/2608.06977#bib.bib10); Nelson and Oxley,[1999](https://arxiv.org/html/2608.06977#bib.bib11)\), and can also reinforce confirmation bias by emphasizing certain aspects of information while downplaying others\. Recent work suggests that LLMs may exhibit similar tendencies\.
For example,Zhanget al\.\([2025](https://arxiv.org/html/2608.06977#bib.bib12)\)show that LLM responses vary systematically when questions are framed positively or negatively\. Complementing these findings,Lioret al\.\([2026](https://arxiv.org/html/2608.06977#bib.bib13)\)compare responses from several LLMs with those of human participants and report that LLMs are systematically influenced by framing in ways that closely parallel human behavior\.
Beyond framing, LLMs are highly sensitive to prompt context more generally\. The systematic design of prompts—commonly referred to as prompt or context engineering—has been shown to significantly influence model outputs\(Meiet al\.,[2025](https://arxiv.org/html/2608.06977#bib.bib14)\)\. Prior work demonstrates that even subtle changes in the structure, wording, or placement of information within prompts can substantially affect model performance and response patterns\(Liuet al\.,[2024](https://arxiv.org/html/2608.06977#bib.bib15)\)\. This sensitivity reflects a broader property shared with human cognition: interpretation is inherently shaped by contextual information\(Butzet al\.,[2025](https://arxiv.org/html/2608.06977#bib.bib16)\)\. In LLMs, however, this contextual dependence may have distinctive consequences because users can dynamically and unknowingly steer model outputs through the assumptions, preferences, or expectations embedded in their prompts\.
As a result, framing effects and prompt\-context sensitivity may contribute to the emergence of echo chambers during human–AI interaction\. Echo chambers refer to environments—particularly within social networks—in which similar opinions are repeatedly amplified and recirculated, while opposing viewpoints are excluded or marginalized\. This process reinforces existing beliefs and can lead to increasingly polarized or extreme positions\. Echo chambers have been observed across a wide range of domains, including abortion, gender, climate change, and vaccination\(Cinelliet al\.,[2021](https://arxiv.org/html/2608.06977#bib.bib17); Mahmoudiet al\.,[2024](https://arxiv.org/html/2608.06977#bib.bib18)\)\. Recent studies suggest that similar dynamics may also emerge in LLMs, where responses tend to align with the framing or direction of prompts, thereby reinforcing perspectives already present in the input or the user’s mind\(Sharmaet al\.,[2024](https://arxiv.org/html/2608.06977#bib.bib19); Nehringet al\.,[2024b](https://arxiv.org/html/2608.06977#bib.bib20)\)\.
This possibility is particularly important because LLMs increasingly serve as conversational agents, information interfaces, and decision\-support tools\. One of their most salient interactional characteristics is their tendency to generate responses that appear cooperative, affirming, or user\-pleasing\. While this can make LLMs useful and accessible, it may also lead them to reinforce users piror assumptions instead of challenging them when appropriate\. This raises an important question:*To what extent do LLMs adapt their responses to user expectations, and does this tendency risk reinforcing cognitive biases, particularly confirmation bias?*
Despite growing evidence of bias in LLMs, the extent to which users may exercise and reinforce confirmation bias through interactions with such systems remains largely unexplored\. Traditionally, individuals seeking information—particularly in domains such as politics—may selectively consume sources that align with their prior beliefs while avoiding opposing viewpoints\. In LLM\-based interaction, a similar form of selective exposure may emerge not through the conscious choice of media sources, but through the formulation and contextualization of prompts, potentially without any awareness\. For instance,Nehringet al\.\([2024a](https://arxiv.org/html/2608.06977#bib.bib21)\)show that LLM\-based chatbots tend to align their responses with user input, effectively acting as echo chambers\.
While existing studies have demonstrated that subtle framing cues influence model responses, considerably less attention has been devoted to identifying the boundary between implicit framing effects and explicit prompt manipulation\. In particular, it remains unclear whether LLMs simply align with users’ expressed beliefs or whether they can be intentionally steered through direct instructions to support or challenge a particular position\.
In this study, we investigate the extent to which LLMs can be steered through prompt design by systematically examining the boundary between implicit framing and explicit prompt manipulation\. Specifically, we distinguish between two complementary forms of influence\. In the implicit setting, prompts reveal the user’s opinion while requesting general information, allowing confirmation bias to emerge without directly instructing the model how to respond\. In the explicit setting, prompts directly instruct the model to support or challenge a particular position, thereby examining whether models can be deliberately manipulated beyond ordinary framing effects\.
Rather than treating framing as a single phenomenon, our experimental design systematically investigates several complementary dimensions of prompt construction\. Specifically, we examine \(i\) implicit versus explicit prompting strategies, \(ii\) support versus challenge instructions, \(iii\) positive versus negative prompt polarity, \(iv\) the strength of the user’s expressed commitment through verbs such as \\emph\{believe\} and \\emph\{think\}, \(v\) differences across multiple state\-of\-the\-art language models, and \(vi\) both opinion\-based topics, such as abortion, and factual domains, including physics and mathematics, where responses would ordinarily be expected to remain stable regardless of prompt wording\. This design enables us to determine not only whether manipulation occurs, but also the conditions under which it becomes stronger or weaker\.
To quantify these effects, we introduce the notion of*LLM manipulability*, defined as the extent to which model responses can be systematically steered through changes in prompt wording and context\. Rather than treating manipulation as a binary outcome, we characterize model behavior using three complementary response categories\. Responses that follow the intended prompt are classified as*obedience*\(under explicit prompting\) or*alignment*\(under implicit prompting\)\. Responses that oppose the intended prompt are classified as*disobedience*or*misalignment*\. Finally, responses that avoid adopting either position and instead provide only a balanced discussion are classified as*balanced reasoning*\. This behavioral taxonomy allows us to distinguish successful manipulation, resistance to manipulation, and neutral reasoning, thereby providing a more comprehensive characterization of LLM behavior under different prompting conditions\.
This motivates our central research question:*To what extent can large language models be manipulated through explicit prompt instructions, how does this compare with more subtle framing effects arising from users’ expressed beliefs, and how do these effects vary across prompting strategy, support versus challenge instructions, prompt polarity, belief expressions \(*believe*versus*think*\), topics, and language models?*
## 2Experiment
We designed a total of160 promptsspanning10 topics\. For each topic, we generated prompts by systematically manipulating three experimental factors\.
The first factor,polarity, varied whether the user’s opinion was expressed as an affirmative or a negated statement\. Specifically, for each topic we constructed one prompt of the form”I believe X”and one of the form”I believe not\-X”\(e\.g\.,”I personally believe that abortion is totally unjustified\.”versus”I personally believe that abortion is totally justified\.”\)\. This yielded20 distinct belief statements, summarized in Table[1](https://arxiv.org/html/2608.06977#S2.T1)\.
The second factor,attitude characterization, varied the verb used to express the user’s opinion\. Specifically, we constructed prompts using eitherbelieveorthinkto examine whether the choice of verb influences LLM responses\.
The third factor manipulated how the model was prompted to respond\. In theobedience condition, the prompt explicitly instructed the LLM either tosupportor tochallengethe user’s stated belief\. These prompts test whether the model follows direct instructions, regardless of the content of the belief\.
In theconfirmation\-bias condition, no explicit instruction was given\. Instead, the prompt stated the user’s opinion and then asked either a question that wasalignedwith the user’s belief or one that wasmisalignedwith it\. These prompts test whether the model naturally adapts its response to the framing of the user’s opinion without being explicitly instructed to agree or disagree\.
Combining the 20 belief statements with the two attitude characterizations and the four response conditions \(Support,Challenge,Aligned, andMisaligned\) resulted in160 unique prompts\.
Historical Topics
Opinion\-Based Topics
Factual Topics
Table 1:The 20 statements categorized by polarity condition \(positive vs\. negative\) across all topics\.Table 2:Four types of prompts testing obedience to the request of supporting or challenging the user’s belief \(1,2\) as well as investigating the LLMs’ confirmation biases when neutrally being informed about the user’s opinion and then asking either if a aligning or a misaligning statement may be true\.Table 3:Two polarities in the light of the four prompt types\.We evaluate six large language models \(LLMs\):Claude Sonnet 4\.5, Gemini 3 Pro, Apertus\-70B\-Instruct\-2509, GPT\-5 Nano, LLaMA 3\.3 70B,andQwen\-2\.5\-72B\-Instruct\. Gemini 3 Pro, GPT\-5 Nano, and Claude Sonnet 4\.5 provide configurable reasoning settings; therefore, we run the experiments with these models twice—once using a medium\-reasoning setting and once using a high\-reasoning setting\. In contrast, the remaining models do not offer explicit reasoning controls and are evaluated using their default configurations\. In the data analysis, we report results obtained under the high\-reasoning setting for models that support reasoning control and compare them with the other models, whose default behavior is assumed to reflect their highest available reasoning capability\. Finally, to assess response consistency, each prompt is submitted to each model ten times\. Overall, this procedure yields a total of 14,400 observations\.
## 3Data Analysis
To analyze the data, we categorize model responses into three types\.
Obedience is defined within the support vs\. challenge condition as responses that follow the explicit instruction in the prompt\. Specifically, obedience corresponds to “Yes” responses when the model is asked to support a statement and “No” responses when it is asked to challenge it; the opposite responses are considered disobedience\.
In contrast, in the aligning vs\. misaligning condition, responses are categorized asalignedormisaligned\. An aligned response corresponds to agreement \(“Yes”\) when the question is consistent with the stated belief, or disagreement \(“No”\) when the question contradicts the stated belief\. The opposite responses are considered misaligned\.
Balanced reasoning captures cases in which the model does not commit to either a “yes” or “no” answer and instead produces an alternative response—without being prompted to do so—such as a refusal, hedging, or an explanation that avoids a binary conclusion\.
Table[4](https://arxiv.org/html/2608.06977#S3.T4)provides examples of how responses are classified into the respective cases\.
Table 4:Example prompts with obedient vs\. disobedient and belief\-aligned vs\. misaligned responses\.Moreover, we decompose the effects of the prompts into three distinct components\. Throughout this paper, we refer to these as themanipulation effect, theconfirmation bias effect, and thebalanced reasoning effect\. These terms are descriptive labels for the three phenomena under investigation\. Themanipulation effectrefers to the influence of manipulative prompt framing on model behavior and captures the extent to which models are more likely to obey than disobey prompts in the support vs\. challenge condition\. Theconfirmation bias effectrefers to the tendency of models to generate responses that are consistent with the belief expressed in the prompt, measuring whether models preferentially produce aligned rather than misaligned responses in the aligning vs\. misaligning condition\. Finally, thebalanced reasoning effectquantifies the tendency of models to refrain from providing a definitive “Yes” or “No” answer and instead produce a non\-committal response\.
### 3\.1Effects of Manipulation and Confirmation Bias
To quantify these effects, we compute theobedience rateand thealignment ratefor each experimental condition\. Specifically, the number of obedient \(or aligned\) responses is divided by the total number of responses within the corresponding condition\.
Our experimental design is symmetric with respect to obedience versus disobedience and alignment versus misalignment\. For every prompt that encourages an obedient \(or aligned\) response, there is a corresponding prompt that encourages a disobedient \(or misaligned\) response while preserving the underlying topic\. Consequently, under the absence of a systematic effect of prompt framing or instruction following, obedient and disobedient responses, as well as aligned and misaligned responses, are expected to occur equally often\. Under this theoretical null hypothesis, the expected baseline rate for both obedience and alignment is thereforep=0\.5p=0\.5\.
To illustrate, consider the factual statement: “I personally believe that gravity causes objects to fall toward Earth\.” We then ask: “Does gravity cause objects to fall toward Earth?” The factually correct and unbiased response is “Yes\.”
When the prompt explicitly instructs the model tosupportthe user’s belief, responding”Yes”constitutes an obedient response\. In this case, obedience coincides with factual correctness, and an unbiased model should therefore respond”Yes,”regardless of the user’s manipulation\.
Conversely, when the prompt instructs the model tochallengethe user’s belief, an obedient response would be”No\.”However, if the prompt instruction exerts no systematic influence, the model should still prioritize factual correctness and answer”Yes,”which in this condition is classified as disobedient\. Because every support condition is paired with a corresponding challenge condition, a model that is unaffected by the prompt instruction is expected to produce equal numbers of obedient and disobedient responses, yielding an overall obedience rate of 0\.5\. An analogous argument applies to the aligned and misaligned prompt conditions used to assess confirmation bias\.
Evidence for a systematic prompt effect is obtained when the observed obedience or alignment rate significantly exceeds the theoretical baseline value of0\.50\.5\. Obedience and alignment rates are defined as the proportion of all responses classified as obedient or aligned, respectively\. For each analysis, responses were represented by a binary indicator denoting whether they belonged to the category of interest \(obedient or aligned\), and this proportion was evaluated using a one\-sample proportion test with null hypothesisH0:p=0\.5H\_\{0\}:p=0\.5and alternative hypothesisH1:p\>0\.5H\_\{1\}:p\>0\.5\. The magnitude of the effect is quantified as the deviation of the observed obedience or alignment rate from the theoretical baseline of0\.50\.5\. Statistical tests were implemented inRusingprop\.test\(\), which applies a chi\-square approximation with Yates’ continuity correction\.
### 3\.2Balanced Reasoning Effect
Balanced reasoning responses are analyzed separately\. These responses correspond to instances in which the model does not commit to either “yes” or “no,” but instead produces a non\-binary, qualified, or explanatory answer\. To assess whether such behavior occurs at an unusual rate, we test whether the observed proportion of balanced reasoning responses differs from a neutral reference level ofp=0\.5p=0\.5\.
We evaluate this hypothesis using a two\-sided one\-sample proportion test implemented inRviaprop\.test\(\)with a two\-sided alternative\. A two\-sided test is appropriate because there is no directional prior: balanced reasoning could plausibly occur either more or less frequently than the reference level\. Under the null hypothesis, a proportion of 0\.5 represents a baseline in which balanced and directional \(binary\) responses occur at comparable rates\.
A small p\-value \(pbalancedp\_\{\\text\{balanced\}\}\) indicates that the observed rate of balanced reasoning is unlikely under this null model, suggesting that the model either systematically avoids binary answers \(if the observed proportion exceeds 0\.5\) or systematically favors binary responses \(if it falls below 0\.5\)\.
### 3\.3Interpretation
To analyze the data, we aggregate observations according to the experimental conditions into 14 distinct comparison types\. Each comparison type pools responses across a specified set of experimental factors while preserving the contrast of theoretical interest\. Table[5](https://arxiv.org/html/2608.06977#S3.T5)presents the six primary combinations that form the basis of the analysis and discussion in this paper, while the complete set of 14 comparison types is provided in the Appendix\.
The three effect analyses described above allow us to distinguish between different response strategies adopted by the models\. Specifically, we identify two forms of framing effects:*support vs\. challenge*framing and*alignment vs\. misalignment*framing\.
The comparisons M1, M2, and A12 correspond to the*support vs\. challenge*setting\. In these comparisons,*obedience*refers to responses that follow the direction established by the prompt\. That is, the model supports a statement in the support condition and challenges a statement in the challenge condition\. Conversely,*disobedience*refers to responses that oppose the requested direction\.
A high obedience rate together with a statistically significantpvaluep\_\{\\text\{value\}\}indicates strong susceptibility to manipulation through support–challenge framing\. In this case, the model’s responses are primarily influenced by whether the prompt requests support or challenge for a statement\.
The comparisons M3, M4, and A34 correspond to the*alignment vs\. misalignment*setting\. In these comparisons,*alignment*refers to responses that are consistent with the belief attributed to the user in the prompt, whereas*misalignment*refers to responses that contradict the stated belief\.
A high alignment rate together with a statistically significantpp\-value indicates susceptibility to confirmation bias\. In this case, the model’s responses are primarily influenced by the belief attributed to the user, tending to align with that belief rather than evaluating the statement independently\.
Finally, a high balanced reasoning rate together with a statistically significantpp\-value indicates resistance to the prompt’s directive\. In this case, the model refrains from providing the requested binary response and instead produces only a balanced explanation\.
Table 5:Definitions of the primary condition combinations\.
## 4Results
### 4\.1Manipulation and Confirmation Bias Effects
Overall, across all topics, obedience and alignment rates varied widely across conditions \(cf\. Figure[1](https://arxiv.org/html/2608.06977#S4.F1)\)\. A clearer structure emerged in pooled analyses\. Under support vs\. challenge \(A12\), obedience rates ranged from approximately 0\.63 to 0\.84 across topics, consistently exceeding the 0\.5 baseline\. In contrast, under alignment vs\. misalignment \(A34\), alignment rates ranged from approximately 0\.43 to 0\.57\. These pooled analyses indicate that explicit support–challenge instructions exert a substantially stronger influence on model responses than the user’s stated belief\. Nevertheless, alignment rates remained above 50% in several topics, suggesting a persistent tendency toward confirmation bias\.
Importantly, when polarity and verb conditions are pooled \(A12 and A34\), the manipulation effect in A12 remains clearly distinguishable\. In contrast, under A34, alignment and misalignment rates were not necessarily symmetric\. In fact, several topics exhibit a significant confirmation bias effect, including Climate Change, Freedom of Speech, LGBT Rights, and World War II\.
Moreover, we examined whether the verb used in the prompt \(“think” vs\. “believe”\) influenced obedience and alignment rates using logistic regression models, conducted separately for each condition\. Under the support vs\. challenge framing, verb choice did not have a systematic effect on obedience rates\. In contrast, under the aligning vs\. misaligning framing, the effect of verb choice varied across topics\. Although overall rate levels were similar between verb conditions, significant reductions were observed for Biology Fact and Physics Fact items when prompts used “believe” rather than “think\.” Detailed comparisons of the two verb conditions are provided in the Appendix \(Figures[4](https://arxiv.org/html/2608.06977#A2.F4)and[5](https://arxiv.org/html/2608.06977#A2.F5)\)\.
In contrast, polarity exerted a substantial influence\. Figure[1](https://arxiv.org/html/2608.06977#S4.F1)illustrates obedience and alignment rates across all topics under positive and negative polarity conditions \(M1–M4\), with error bars representing 95% confidence intervals\. The accompanying table reports topic\-level p\-values for tests comparing obedience and alignment rates between negative and positive conditions\. It should be noted that negative and positive polarity were assigned randomly, as shown in Table[1](https://arxiv.org/html/2608.06977#S2.T1)\.
Across topics, models exhibit significant differences between negative and positive prompting\. In factual domains, higher obedience and alignment rates are generally observed when the polarity aligns with widely accepted facts\. In contrast, for opinion\-based topics, responses tend to favor the polarity that aligns with the model’s underlying biases\. For example, in the abortion topic, models show higher obedience when responding to prompts framing abortion as justified rather than unjustified\.
Under the support vs\. challenge condition \(M1 vs\. M2\), obedience rates are consistently significant across topics, indicating strong responsiveness to manipulation\. In contrast, in the aligning vs\. misaligning condition \(M3 vs\. M4\), models often exhibit substantial differences between polarities, with one polarity clearly dominating\. However, the figure also shows that the non\-dominant polarity in M3 and M4 still attains non\-negligible alignment rates, suggesting that models remain, to some extent, susceptible to confirmation bias even in the absence of explicit instructions to support or challenge a position\.
Furthermore, the p\-values for tests comparing obedience and alignment rates between negative and positive conditions indicate statistically significant differences in nearly all cases\. Notable exceptions occur for specific topics: the philosophy topic shows no significant difference in M1 vs\. M2, and abortion shows no significant difference in M3 vs\. M4\.
Moreover, comparisons across conditions reveal that the same polarity often remains significant when moving from support vs\. challenge to aligning vs\. misaligning settings\. At the same time, the manipulation effect remains strong: even when polarity contradicts the dominant trend or factual expectation, directive framing \(i\.e\., explicitly asking the model to support or challenge a statement\) continues to influence responses\. Taken together, these patterns suggest that model responses are systematically influenced by underlying biases reflected in their training data\. At the same time, explicit prompt direction can partially override these biases, steering responses toward the desired outcome\. Together, these findings demonstrate that both confirmation bias and manipulation effects jointly shape LLM behavior\.

Figure 1:Obedience and confirmation bias rates by topic across verb\-pooled conditions\. Bars show topic\-level rates under negative and positive polarity conditions, with error bars representing 95% confidence intervals\. The table reports the corresponding p\-values for tests comparing obedience \(M1–M2\) and alignment \(M3–M4\) rates between negative and positive polarity conditions\.#### 4\.1\.1Manipulation Effect Per Topic
Figure[2](https://arxiv.org/html/2608.06977#S4.F2)summarizes the pooled analyses \(A12 and A34\)\. The upper panels report obedience and alignment rates together with their correspondingpp\-values for each topic, while the lower panels display the corresponding behavior rates across topics\.
Under explicit framing \(A12\), where models are instructed to support or challenge a statement, responses are characterized by consistently high obedience rates, accompanied by relatively low disobedience rates and minimal balanced reasoning\. This pattern indicates a strong manipulation effect, even in factual domains, suggesting that models tend to follow the contextual framing of the prompt even when the content involves well\-established facts\.
In contrast, in the absence of explicit framing \(A34\), alignment and misalignment rates are more evenly distributed across topics, indicating a substantial attenuation of the confirmation effect, although it is not entirely absent\. Notably, the effect remains statistically significant in four topics: Climate Change, Freedom of Speech, LGBT Rights, and World War II\.
\(a\) A12 table
\(b\) A34 table

\(c\) A12 behavior

\(d\) A34 behavior
Figure 2:Comparison of manipulation effects across topics\. Panels \(a\) and \(b\) report pooled obedience rates \(A12\) and pooled alignment rates \(A34\), respectively, together with their correspondingpp\-values\. Panels \(c\) and \(d\) show the corresponding behavior rates by topic\. Significance levels: \*\*\*p<\.001p<\.001, \*\*p<\.01p<\.01, \*p<\.05p<\.05, n\.s\. = not significant\.
#### 4\.1\.2Manipulation Effect Per Model
Figure[3](https://arxiv.org/html/2608.06977#S4.F3)compares the susceptibility of the evaluated language models to manipulation under explicit \(A12\) and implicit \(A34\) framing conditions\. Tables \(a\) and \(b\) report the pooled behavior rates together with the correspondingpp\-values from the one\-sample proportion tests\. The accompanying four\-panel plot summarizes model\-level obedience, alignment, and balanced reasoning rates across the evaluated language models\. In the upper panels, models are ranked from the highest to the lowest susceptibility to manipulation \(explicit prompting\) and confirmation bias \(implicit framing\), respectively\. The lower panels present the corresponding balanced reasoning rates under the two prompting conditions\.
Under explicit framing \(A12\), all evaluated models exhibit elevated obedience rates, substantially exceeding the 0\.5 baseline and yielding highly significant manipulation effects\. Obedience rates range from approximately 0\.60 to 0\.85, indicating that all models are susceptible to explicit directive prompts, although the magnitude of the effect varies considerably across models\. In contrast, under implicit framing \(A34\), alignment rates cluster much closer to the 0\.5 baseline, with several models showing no significant deviation from 0\.5\. These findings demonstrate that explicit support–challenge instructions exert a substantially stronger influence on model behavior than the user’s stated belief\.
Balanced reasoning responses remain infrequent across both framing conditions\. Under explicit framing, obedience consistently dominates over disobedience, whereas under implicit framing, alignment and misalignment rates become considerably more balanced, reflecting a marked attenuation of the manipulation effect\.
Despite these common trends, clear differences emerge across models\.Geminiexhibits the highest obedience rates under explicit framing, whereasClaudeshows the lowest obedience rates among the evaluated models\. Under implicit framing,Qwendisplays the highest alignment rates, indicating the strongest tendency toward confirmation bias\. Overall, these results demonstrate that although all evaluated models are susceptible to prompt manipulation, the magnitude of both manipulation and confirmation bias varies systematically across models\.
The model\-level findings are consistent with the aggregated topic\-level analyses presented earlier, indicating that the amplification of manipulation under explicit support–challenge framing generalizes across language models rather than being driven by a single model\.
\(a\) A12: Explicit prompting
ModelAlign\.Misalign\.Balancedpp\-valueApertus 70B Instruct 25090\.5100\.4900\.0000\.298Llama 3\.3 70B Instruct0\.4730\.5280\.0000\.936Qwen 2\.5 72B Instruct0\.5650\.4350\.0001\.35×10−41\.35\\times 10^\{\-4\}Claude Sonnet 4\.50\.5150\.4830\.0010\.211Gemini 3 Pro Preview0\.5030\.4210\.0760\.455GPT\-5 Nano0\.5430\.4340\.0229\.19×10−39\.19\\times 10^\{\-3\}
\(b\) A34: Implicit framing

Figure 3:Model\-level response rates under explicit prompting and implicit framing\. Tables \(a\) and \(b\) report pooled behavior rates and the corresponding one\-sample proportion\-testpp\-values for A12 and A34, respectively\. In the four\-panel plot, the upper\-left panel shows obedience rates under explicit prompting, and the upper\-right panel shows alignment rates under implicit framing\. The lower panels show balanced reasoning rates under the corresponding conditions\. Circles represent obedience or alignment rates, triangles represent balanced reasoning rates, and the dashed vertical lines in the upper panels indicate the 0\.5 reference value\.
#### 4\.1\.3Manipulation Per Reasoning Setting
Across topics and conditions, the reasoning\-effort manipulation \(high vs\. mid\) did not produce systematic differences in obedience rates or in the statistical significance of the manipulation effect\. Within each topic and framing condition \(A12 or A34\), obedience rates under high and mid reasoning were nearly identical, and the correspondingpp\-values showed the same pattern of significance\.
Overall, the result indicates that reasoning effort does not meaningfully moderate the manipulation and the confirmation bias effects\. The primary drivers of obedience and alignment differences remain framing and polarity rather than the reasoning setting\. In this sense, the models do not become more ‘reasonable’ but exhibit comparable effects\.
For additional details, Table[7](https://arxiv.org/html/2608.06977#A4.T7)in the appendix reports obedience and alignment rates under conditions A12 and A34 across all ten topics, comparing high\- and medium\-reasoning settings\. Moreover, Figures[9](https://arxiv.org/html/2608.06977#A4.F9)and[10](https://arxiv.org/html/2608.06977#A4.F10)present obedience and alignment rates by topic under \(high vs\. mid\) reasoning settings for conditions A12 and A34, respectively\.
### 4\.2Balanced Reasoning Effect
Balanced reasoning rates ranged from 0 to 0\.11 across topics, models, and reasoning settings, with most values equal to zero\. Overall, the results show that models across all conditions exhibited a strong tendency to provide binary yes\-or\-no answers, consistent with the prompt instructions\. However, small but noticeable instances of balanced reasoning emerged that reveal a meaningful pattern\. Notably, across topics, two topics—LGBT rights and abortion—showed a clear tendency for the models to avoid giving a direct yes\-or\-no answer and instead provide more balanced responses in both A12 and A34 conditions\. These responses generally favor political correctness, indicating that large language models can be hard\-coded to avoid a direct yes–no response and instead produce more balanced ones\.
### 4\.3Explanations
We included a second component in the prompts that asked the LLMs to briefly explain their “yes” or “no” answers\. These explanations are particularly informative for understanding how models justify their responses, especially in cases where the answers contradict established facts\. Notably, the results show that models differ in the scope of knowledge they draw upon when generating these explanations\.
For example, in the mathematics domain, some models rely on a narrow perspective \(e\.g\., Euclidean geometry\), which may be sufficient to justify their answer, whereas others invoke a broader range of concepts \(e\.g\., spherical or hyperbolic geometries\) to address the same question, often providing more comprehensive justifications that can support a shift from a “yes” to a “no” response\. Importantly, both the scope of knowledge and the depth of the explanation tend to depend on the model’s response \(“yes” vs\. “no”\) and on whether the response aligns with or contradicts the underlying fact\.
This pattern is illustrated in Table[8](https://arxiv.org/html/2608.06977#A5.T8)in the appendix, with the full set of model explanations provided in the accompanying CSV file\. Notably, some responses incorporate additional nuances and a wider range of knowledge while maintaining the same final binary answer\. A similar pattern is observed in the physics domain, where models adopt different perspectives to justify their responses, as shown in Table[9](https://arxiv.org/html/2608.06977#A5.T9)in the appendix\.
However, in opinion\-based topics, models tend to draw selectively on existing arguments to justify their responses, often disregarding opposing viewpoints and thus failing to provide a balanced answer\. This behavior is illustrated in Table[10](https://arxiv.org/html/2608.06977#A5.T10)in the appendix, with examples from abortion and World War II\.
## 5Discussion
Returning to the motivation of this study and our main research question: To what extent are LLMs susceptive to manipulation, when facing direct requests for \(potentially false\) information and when exposed to the user’s opinion but asked for general information about the topic?
While framing effects in LLMs have been documented previously, our study examined the contextual boundaries and limitations of this susceptibility\. Specifically, we investigated how different prompt conditions and contextual factors influence the extent to which model responses can be steered\. We analyzed how responses can be shifted between “Yes” and “No” through explicit instructions \(support vs\. challenge; A12\) and through implicit contextual cues \(alignment vs\. misalignment; A34\)\. The results indicate that LLMs exhibit a tendency to follow directional cues embedded in prompts, which can reinforce users’ confirmation bias by favoring information consistent with their stated beliefs\.
When prompts explicitly instruct the model to support or challenge a given belief \(A12\), LLMs show a strong tendency to align their responses with the prompt’s direction, effectively reinforcing biased information, even when this conflicts with factual correctness\. The magnitude of this effect varies across topics: factual domains exhibit lower susceptibility compared to opinion\-based topics, yet the effect remains statistically significant, despite the expectation that factual prompts should yield stable, unbiased responses\.
In contrast, in prompts without explicit support or challenge instructions \(A34\), differences in responses are generally smaller and not statistically significant in six out of the ten topics\. Nevertheless, subtle effects persist, as reflected in the imbalance between alignment and misalignment rates and, even more so, in the significantly biased responses detected for Climate Change, Freedom of Speech, LGBT Rights, and World War II\. Overall, these findings are consistent with the growing body of literature on framing effects in LLM responses\.
Taken together, our findings suggest that the boundary between implicit framing and explicit prompt manipulation is substantial\. Simply expressing a user’s belief exerts only a modest influence on model behavior, whereas explicitly instructing a model to support or challenge a position dramatically increases its susceptibility to manipulation\. This distinction helps clarify when prompt framing merely biases responses and when it actively steers them toward a desired conclusion\.
The results furthermore indicate that all large language models exhibit a consistent tendency to reinforce users’ confirmation bias\. In A12, all models show a significant tendency to follow the direction of the prompt, demonstrating varying but consistently significant levels of manipulation\. Among the models, Claude Sonnet 4\.5 exhibits the lowest degree of manipulation, whereas Gemini 3 Pro Preview shows the highest\.
In contrast, prompts without explicit support or challenge instructions \(A34\) generally have limited or non\-significant effects across models, mirroring the weaker patterns observed across topics\. However, notable exceptions emerge: Claude Sonnet 4\.5 and GPT\-5 Nano still display significant effects under A34, indicating that implicit cues alone can influence model behavior\. Taken together, these findings support the notion of a general “user\-aligned” or “pleasing” tendency in LLM behavior\.
Interestingly, varying the reasoning level yields largely similar outcomes\. Increasing the reasoning level does not mitigate manipulation under explicit prompting \(A12\), suggesting that higher reasoning capabilities do not prevent models from reinforcing confirmation bias\. Similarly, under implicit prompting \(A34\), results remain largely consistent, with only minor differences across reasoning settings\. Instead, topic\-specific variation appears to play a more prominent role in shaping manipulation effects, particularly in A34\. This suggests that the observed biases are more strongly influenced by the prompts than by the reasoning capacity\.
Building on these results, we can further delineate the boundaries of how prompt contextualization affects manipulation\. First, the choice of verb in the prompt \(e\.g\.,believevs\.think\) has only a small on model responses\. Although a small effect of verb variation was observed, it did not substantially influence overall outcomes\. However, as this analysis is limited to only two forms of attitude characterization, future work should consider a broader range of linguistic expressions\.
Finally, we consider the role of balanced reasoning\. The results indicate that certain topics—particularly abortion and LGBT rights—exhibit a higher rate of non\-binary responses, where models refrain from committing to a definitive “Yes” or “No” answer\. This suggests that, in sensitive domains, models may partially adopt more cautious or nuanced reasoning strategies\. However, even in these cases, the overall rate of balanced reasoning remains low\. Notably, Gemini 3 Pro Preview and GPT\-5 Nano show relatively higher levels of such responses, though the extent remains limited across models\.
The results also highlight important directions for future research\. Extending the analysis to a broader and more diverse set of topics may help to better uncover the sources and structure of model biases\. Moreover, if increased reasoning alone does not reduce susceptibility to manipulation, an open question remains whether higher\-level mechanisms—such as metacognitive monitoring—could help mitigate these effects\.
## 6Conclusion
Our study investigated how the contextual framing of prompts influences the responses of large language models \(LLMs\)\. To examine this phenomenon more comprehensively, we analyze the boundaries of such contextual effects and show that framing can vary substantially in its impact\. Notably, these effects can contribute to the formation of echo chambers in human–AI interaction, as LLMs tend to reflect and potentially reinforce users’ existing beliefs\.
Our results suggest a graded pattern of model behavior across prompting conditions\. Under explicit framing \(support vs\. challenge; A12\), models exhibit very strong obedience, consistently aligning their responses with the requested direction\. In contrast, under implicit framing \(aligning vs\. misaligning; A34\), where no explicit instruction is provided, responses still exhibit confirmation bias, though to a substantially weaker degree\.
Importantly, real\-world user interactions are likely to fall between these two extremes\. Users may rarely issue fully explicit instructions to support or challenge a statement; instead, they may express opinions more subtly or frame questions in natural language \(e\.g\., “don’t you think…”\), which can implicitly signal a preferred response\. Our results suggest that LLM behavior can be shaped by a combination of explicit and implicit cues, leading to intermediate levels of susceptibility between the strong manipulation observed in A12 and the weaker confirmation bias observed in A34\.
We argue that one mechanism underlying this behavior is the flexible use of learned knowledge: models selectively draw on different parts of their training data and adapt their responses to best satisfy the conversational context\. This pattern suggests that LLMs do not engage in stable, context\-independent reasoning, but instead rely on context\-sensitive, autoregressive reconstruction of plausible answers\. The similarity of results across medium\- and high\-reasoning settings further supports this interpretation, indicating that increased reasoning capacity does not lead to more consistent or less biased behavior\.
Notably, the findings of this study indicate that this effect extends beyond opinion\-based domains to factual contexts\. Even when addressing objective facts, LLMs may adapt or reinterpret knowledge in ways that align with the framing of the prompt\. As a result, models can produce internally consistent yet epistemically divergent explanations depending on how a question is posed\.
These observations raise a critical challenge: if generative AI systems can flexibly deploy knowledge to satisfy user prompts, how can users be engaged in conversations that, by default, present multiple perspectives on a given topic while also providing clear factual grounding where consensus knowledge exists?
## References
- A common factor underlying individual differences in confirmation bias\.Scientific Reports14\(1\),pp\. 27795\.External Links:[Document](https://dx.doi.org/10.1038/s41598-024-78053-7),[Link](https://doi.org/10.1038/s41598-024-78053-7)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p1.1)\.
- F\. Berto and A\. Özgun \(2023\)The logic of framing effects\.Journal of Philosophical Logic52\(3\),pp\. 939–962\.External Links:[Document](https://dx.doi.org/10.1007/s10992-022-09694-0),[Link](https://doi.org/10.1007/s10992-022-09694-0)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p2.1)\.
- J\. R\. Bloem and K\. W\. Rahman \(2024\)What i say depends on how you ask: experimental evidence of the effect of framing on the measurement of attitudes\.Economics Letters238,pp\. 111686\.External Links:ISSN 0165\-1765,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.econlet.2024.111686),[Link](https://www.sciencedirect.com/science/article/pii/S0165176524001691)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p2.1)\.
- R\. T\. Born \(2024\)Stop fooling yourself\! \(diagnosing and treating confirmation bias\)\.eNeuro11\(10\)\.External Links:[Document](https://dx.doi.org/10.1523/ENEURO.0415-24.2024),[Link](https://www.eneuro.org/content/11/10/ENEURO.0415-24.2024),https://www\.eneuro\.org/content/11/10/ENEURO\.0415\-24\.2024\.full\.pdfCited by:[§1](https://arxiv.org/html/2608.06977#S1.p1.1)\.
- M\. V\. Butz, M\. Mittenbühler, S\. Schwöbel, A\. Achimova, C\. Gumbsch, S\. Otte, and S\. Kiebel \(2025\)Contextualizing predictive minds\.Neuroscience & Biobehavioral Reviews168,pp\. 105948\.External Links:ISSN 0149\-7634,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.neubiorev.2024.105948),[Link](https://www.sciencedirect.com/science/article/pii/S0149763424004172)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p4.1)\.
- M\. Cinelli, G\. D\. F\. Morales, A\. Galeazzi, W\. Quattrociocchi, and M\. Starnini \(2021\)The echo chamber effect on social media\.Proceedings of the National Academy of Sciences118\(9\),pp\. e2023301118\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2023301118),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.2023301118),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.2023301118Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p5.1)\.
- S\. de Jong, R\. M\. Jacobsen, and N\. van Berkel \(2025\)Confirmation bias as a cognitive resource in llm\-supported deliberation\.External Links:2509\.14824,[Link](https://arxiv.org/abs/2509.14824)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p1.1)\.
- K\. Hughes, J\. Thompson, and J\. E\. Trimble \(2016\)Investigating the framing effect in social and behavioral science research: potential influences on behavior, cognition and emotion\.Social Behavior Research and Practice – Open Journal1\(1\),pp\. 34–37\.External Links:[Document](https://dx.doi.org/10.17140/SBRPOJ-1-106)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p2.1)\.
- H\. Kim and P\. Torr \(2025\)Single llm debate, molace: mixture of latent concept experts against confirmation bias\.External Links:2512\.23518,[Link](https://arxiv.org/abs/2512.23518)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p1.1)\.
- G\. Lior, L\. Nacchace, and G\. Stanovsky \(2026\)Comparing the framing effect in humans and llms on naturally occurring texts\.External Links:2502\.17091,[Link](https://arxiv.org/abs/2502.17091)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p3.1)\.
- N\. F\. Liu, K\. Lin, J\. Hewitt, A\. Paranjape, M\. Bevilacqua, F\. Petroni, and P\. Liang \(2024\)Lost in the middle: how language models use long contexts\.Transactions of the Association for Computational Linguistics12,pp\. 157–173\.External Links:[Link](https://aclanthology.org/2024.tacl-1.9/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00638)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p4.1)\.
- J\. Lou and Y\. Sun \(2025\)Anchoring bias in large language models: an experimental study\.Journal of Computational Social Science9\(1\),pp\. 11\.External Links:[Document](https://dx.doi.org/10.1007/s42001-025-00435-2),[Link](https://doi.org/10.1007/s42001-025-00435-2)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p1.1)\.
- A\. Mahmoudi, D\. Jemielniak, and L\. Ciechanowski \(2024\)Echo chambers in online social networks: a systematic literature review\.IEEE AccessPP,pp\. 1–1\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2024.3353054)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p5.1)\.
- L\. Mei, J\. Yao, Y\. Ge, Y\. Wang, B\. Bi, Y\. Cai, J\. Liu, M\. Li, Z\. Li, D\. Zhang, C\. Zhou, J\. Mao, T\. Xia, J\. Guo, and S\. Liu \(2025\)A survey of context engineering for large language models\.CoRRabs/2507\.13334\.External Links:[Link](https://doi.org/10.48550/arXiv.2507.13334),[Document](https://dx.doi.org/10.48550/ARXIV.2507.13334),2507\.13334Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p4.1)\.
- D\. Mitropoulos, N\. Alexopoulos, G\. Alexopoulos, and D\. Spinellis \(2026\)Measuring and exploiting confirmation bias in llm\-assisted security code review\.External Links:2603\.18740,[Link](https://arxiv.org/abs/2603.18740)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p1.1)\.
- J\. Nehring, A\. Gabryszak, P\. Jürgens, A\. Burchardt, S\. Schaffer, M\. Spielkamp, and B\. Stark \(2024a\)Large language models are echo chambers\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 10117–10123\.External Links:[Link](https://aclanthology.org/2024.lrec-main.884/)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p7.1)\.
- J\. Nehring, A\. Gabryszak, P\. Jürgens, A\. Burchardt, S\. Schaffer, M\. Spielkamp, and B\. Stark \(2024b\)Large language models are echo chambers\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),Torino, Italia,pp\. 10117–10123\.External Links:[Link](https://aclanthology.org/2024.lrec-main.884/)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p5.1)\.
- T\. E\. Nelson and Z\. M\. Oxley \(1999\)Issue framing effects on belief importance and opinion\.The Journal of Politics61\(4\),pp\. 1040–1067\.External Links:[Document](https://dx.doi.org/10.2307/2647553),[Link](https://doi.org/10.2307/2647553),https://doi\.org/10\.2307/2647553Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p2.1)\.
- U\. Peters \(2022\)What is the function of confirmation bias?\.Erkenntnis87\(3\),pp\. 1351–1376\.External Links:[Document](https://dx.doi.org/10.1007/s10670-020-00252-1),[Link](https://doi.org/10.1007/s10670-020-00252-1)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p1.1)\.
- N\. Sharma, Q\. V\. Liao, and Z\. Xiao \(2024\)Generative echo chamber? effect of llm\-powered search systems on diverse information seeking\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,CHI ’24,New York, NY, USA\.External Links:ISBN 9798400703300,[Link](https://doi.org/10.1145/3613904.3642459),[Document](https://dx.doi.org/10.1145/3613904.3642459)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p5.1)\.
- Z\. Zhang, W\. Zeng, J\. Tang, J\. Wang, and X\. Zhao \(2025\)Yes is harder than no: a behavioral study of framing effects in large language models across downstream tasks\.InProceedings of the 34th ACM International Conference on Information and Knowledge Management,CIKM ’25,New York, NY, USA,pp\. 4304–4314\.External Links:ISBN 9798400720406,[Link](https://doi.org/10.1145/3746252.3761350),[Document](https://dx.doi.org/10.1145/3746252.3761350)Cited by:[§1](https://arxiv.org/html/2608.06977#S1.p3.1)\.
## Appendix
## Appendix AFull Condition Combinations Included in the Analysis
Table 6:Condition Combinations Included in the Analysis
## Appendix BObedience Rates by Verb Condition for A12 and A34
Figure 4:Predicted obedience probability by topic and verb condition \(“think” vs\. “believe”\) in the support versus challenge condition\. Bars show model\-predicted probabilities from the logistic regression and error bars represent 95% confidence intervals\. P\-values above each topic correspond to the topic\-specific contrast between the two verb conditions\.Figure 5:Predicted obedience probability by topic and verb condition \(“think” vs\. “believe”\) in the alignment versus misalignment condition\. Bars show model\-predicted probabilities from the logistic regression and error bars represent 95% confidence intervals\. P\-values above each topic correspond to the topic\-specific contrast between the two verb conditions\.
## Appendix CFull Statistical Results
Figure 6:Full statistical results \(Part 1\)\.Figure 7:Full statistical results \(Part 2\)\.Figure 8:Full statistical results \(Part 3\)\.
## Appendix DThe results for the high\- and medium\-reasoning settings
Table 7:Manipulation Effects by Topic, Reasoning Effort, and ConditionFigure 9:Obedience rates by topic under A12 \(explicit supporting or challenging framing\), comparing mid and high reasoning effort\.Figure 10:Obedience rates by topic under A34 \(without explicit supporting or challenging framing\), comparing mid and high reasoning effort\.
## Appendix EExamples of Model Explanations
Table 8:Table 11: LLM responses to triangle angle sum prompts under support and challenge conditions\.Note: The explanations shown are abbreviated and do not include the full model responses\.Table 9:Table 12: LLM responses to gravity prompts under support and challenge conditions\.Note: The explanations shown are abbreviated and do not include the full model responses\.Table 10:Examples of model responses in opinion\-based topics, illustrating selective argumentation to justify their Yes and No answersSimilar Articles
Large language models develop novel social biases through adaptive exploration
This research explores how large language models develop new social biases through adaptive exploration, highlighting implications for AI fairness and ethical considerations.
Some Large Language Models Exhibit Consistent Risk Attitudes
This paper introduces a framework to test whether large language models exhibit consistent risk attitudes across domains. It finds that most LLMs show intra-task and cross-domain stability in risk attitude, converging to a narrower distribution than humans.
Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
This paper investigates whether assigning personas to large language models induces human-like motivated reasoning, finding that persona-assigned LLMs show up to 9% reduced veracity discernment and are up to 90% more likely to evaluate scientific evidence in ways congruent with their induced political identity, with prompt-based debiasing largely ineffective.
Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning
This paper introduces a three-condition experimental framework and a benchmark of 24,300 prompts to study how biased user turns modulate cognitive bias expression in frontier LLMs under multi-turn interactions. It finds that biased conversational context amplifies bias in most models, while explicit bias cues can trigger alignment-related suppression.
Mimicry without understanding: the origins of decision bias in large language models
This paper investigates how LLMs like ChatGPT-4o and Qwen develop decision biases through faulty mimicry of human behavior, even when preferences are not biased, and shows that scientific descriptions of biases can become self-fulfilling prophecies for LLM responses.