A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents
Summary
This paper empirically investigates the susceptibility of LLM-based GUI agents to digital nudges, finding that reasoning configuration redirects rather than reduces nudge effects, positioning interface design as a governance concern for autonomous AI.
View Cached Full Text
Cached at: 09/18/26, 09:26 AM
# A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents Source: [https://arxiv.org/html/2609.19843](https://arxiv.org/html/2609.19843) Haya HalimehEmail:[haya\.halimeh@uni\-paderborn\.de](mailto:[email protected])Affiliation:Information Systems, Paderborn University, Paderborn, GermanySascha KaltenpothEmail:[sascha\.kaltenpoth@uni\-paderborn\.de](mailto:[email protected])Affiliation:Information Systems, Paderborn University, Paderborn, GermanyKevin BöschEmail:[kboesch@mail\.uni\-paderborn\.de](mailto:[email protected])Affiliation:Information Systems, Paderborn University, Paderborn, GermanyOliver MüllerEmail:[oliver\.mueller@uni\-paderborn\.de](mailto:[email protected])Affiliation:Information Systems, Paderborn University, Paderborn, Germany ###### Abstract LLM\-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind\. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users\. While behavioural biases in the textual outputs of LLMs are well\-documented, far less is known about how such influence operates when models act as agents that perceive interfaces and execute decisions—and, in particular, whether the reasoning capabilities increasingly built into these agents make them more robust to it\. Drawing on Dual\-Process Theory, we empirically investigate whether LLM\-based GUI agents are susceptible to automatic \(Type 1\) and reflective \(Type 2\) digital nudges, and how their reasoning configuration moderates this susceptibility\. In a randomized online shopping experiment with 3,600 agents and a total of 21,600 simulations across six frontier models from three providers, we found that agents were vulnerable to both nudge types\. Crucially, the reasoning configuration moderated these effects in opposing directions, reducing susceptibility to automatic default nudges while heightening it to reflective social influence nudges\. Extensive reasoning therefore did not make agents more robust but redirected the route through which choice architecture takes effect\. Exploratory analysis further showed this redirection to be systematically structured by model scale\. Beyond establishing nudge susceptibility as a behavioural property of agentic AI, the study positions interface design as a governance concern for organizations that delegate decisions to autonomous agents\. ###### keywords Digital Nudging, Dual\-Process Theory, Large Language Model, Reasoning, GUI Agents ### 1Introduction Agentic AI refers to autonomous systems capable of pursuing user\-delegated goals with limited human intervention and are increasingly integrated into organizational information systems\([Baird and Maruping, 2021](https://arxiv.org/html/2609.19843#bib.bib30);[Murray et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib32);[Murugesan, 2025](https://arxiv.org/html/2609.19843#bib.bib31)\)\. Recent advances in large language models \(LLMs\) have given rise to agentic AI that pairs LLMs with the ability to perceive and operate graphical user interfaces \(GUIs\), referred to as LLM\-based GUI agents\([Hu et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib46);[Nguyen et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib47);[Sager et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib45)\), to carry out tasks in digital environments\. LLM\-based GUI agents can navigate websites\([Zhou et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib70)\), complete spreadsheet tasks\([Li et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib79)\), make purchases\([Yao et al\., 2022](https://arxiv.org/html/2609.19843#bib.bib78)\), and even conduct financial transactions\([Yu et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib80)\)\. Examples include ChatGPT Agent\([OpenAI, 2025a](https://arxiv.org/html/2609.19843#bib.bib18)\), Perplexity’s Comet browser\([Perplexity, 2025a](https://arxiv.org/html/2609.19843#bib.bib10)\), and OpenClaw\([Steinberger, 2026](https://arxiv.org/html/2609.19843#bib.bib11)\)\. This development marks a shift from LLMs as tools for generating and processing information toward systems capable of acting on users’ behalf, enabling new forms of joint human–AI agency\([Zhang et al\., 2024a](https://arxiv.org/html/2609.19843#bib.bib8);[Murray et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib32)\)\. The growing capabilities of agentic AI are accompanied by accelerating organizational interest and adoption\. A recent McKinsey survey found that 23% of firms plan to scale their use of such systems\([Singla et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib20)\), while industry forecasts suggest that they will significantly reshape online shopping by 2030\([Müller, 2026](https://arxiv.org/html/2609.19843#bib.bib2);[Späne et al\., 2026](https://arxiv.org/html/2609.19843#bib.bib1)\)\. As these systems continue to mature, their ability to automate increasingly complex tasks is expected to expand\([Schmidt et al\., 2026](https://arxiv.org/html/2609.19843#bib.bib9)\), motivating organizations to integrate them into more business processes and increasing users’ willingness to delegate decisions to them\([Candrian and Scherer, 2022](https://arxiv.org/html/2609.19843#bib.bib4)\)\. Consequently, agentic AI is likely to become commonplace across domains such as e\-commerce, healthcare, and finance\([BaniHani et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib71)\)\. Such delegation is consequential because LLM\-based GUI agents operate in digital environments built for humans, settings designed not only to facilitate decisions but often also to influence them\. A prominent example for influencing decisions in such environments are digital nudges\([Thaler et al\., 2013](https://arxiv.org/html/2609.19843#bib.bib68);[Weinmann et al\., 2016](https://arxiv.org/html/2609.19843#bib.bib61)\)\. The central premise of nudge theory is that human, and potentially agentic, behaviour can be systematically shaped by subtle and seemingly insignificant changes in the choice architecture\([Thaler and Sunstein, 2009](https://arxiv.org/html/2609.19843#bib.bib69)\), such as a preselected default or a visually highlighted option\([Weinmann et al\., 2016](https://arxiv.org/html/2609.19843#bib.bib61)\)\. The effectiveness of nudges in steering human behaviour is documented across decades of behavioural research and a wide range of domains\([Bergram et al\., 2022](https://arxiv.org/html/2609.19843#bib.bib54);[Caraban et al\., 2019](https://arxiv.org/html/2609.19843#bib.bib59);[Haki et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib29);[Hummel and Maedche, 2019](https://arxiv.org/html/2609.19843#bib.bib57)\)\. One might expect LLM\-based agents to be less susceptible to such influences than humans, since they arguably lack the cognitive limitations traditionally invoked to explain human bias\. For example, researchers have argued that LLMs exhibit higher economic rationality than humans\([Chen et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib72)\)\. A growing body of research, however, points the other way\. Various studies suggest that LLMs are similarly subject to bounded rationality\([Simon, 1955](https://arxiv.org/html/2609.19843#bib.bib65)\)and can display systematic and sometimes unstable shifts in their choices when exposed to external influence\([Cherep et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib77)\)\. Research on LLM textual outputs in particular demonstrates human\-like decision biases, including framing, anchoring, and availability effects\([Nguyen, 2024](https://arxiv.org/html/2609.19843#bib.bib25);[Suri et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib90);[Hagendorff et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib86)\), and sensitivity to superficial changes in how options are presented\([Cherep et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib77);[Jones and Steinhardt, 2022](https://arxiv.org/html/2609.19843#bib.bib87)\), although the resulting deviations do not always mirror human patterns\([Macmillan\-Scott and Musolesi, 2024](https://arxiv.org/html/2609.19843#bib.bib51)\)\. Yet despite rising interest in agentic AI\([Schmidt et al\., 2026](https://arxiv.org/html/2609.19843#bib.bib9)\), and specifically in LLM\-based GUI agents\([Zhang et al\., 2024a](https://arxiv.org/html/2609.19843#bib.bib8);[Hu et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib46);[Nguyen et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib47);[Sager et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib45)\), current research on biases of such systems falls short in two respects\. First, existing work focuses predominantly on biases in LLMs’ textual output\([Tjuatja et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib14);[Acerbi and Stubbersfield, 2023](https://arxiv.org/html/2609.19843#bib.bib12);[Echterhoff et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib13)\)and does not connect the biases it documents to the behavioural science from which nudging derives, yielding a fragmented picture that catalogues effects without explaining why particular effects emerge\. Nudge theory\([Thaler and Sunstein, 2009](https://arxiv.org/html/2609.19843#bib.bib69)\), grounded in Dual\-Process Theory\([Kahneman, 2003](https://arxiv.org/html/2609.19843#bib.bib24);[Evans and Stanovich, 2013](https://arxiv.org/html/2609.19843#bib.bib63)\), provides such a framework by classifying nudges by the mode of processing they engage, distinguishing automatic Type 1 from reflective Type 2 nudges\([Hansen and Jespersen, 2013](https://arxiv.org/html/2609.19843#bib.bib62)\)\.[Bösch \(2025\)](https://arxiv.org/html/2609.19843#bib.bib21)shows that the decoy effect carries over to such agents\. However, to the best of our knowledge, it remains an open question, whether the broader Type 1 and Type 2 distinction manifests similarly in agentic behaviour, in which an LLM\-based GUI agent perceives a graphical interface, plans, navigates, and commits to a choice by acting on it\([Nguyen et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib47);[Sager et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib45)\)\. Second, a common assumption is that greater model capability, particularly through enhanced reasoning configurations, produces more robust and less biased behaviour\([Wei et al\., 2022a](https://arxiv.org/html/2609.19843#bib.bib91);[Hagendorff et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib86);[Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5)\)\. These configurations are designed to support more deliberative reasoning and improve task adaptability\([Huang and Chang, 2023](https://arxiv.org/html/2609.19843#bib.bib36);[Wei et al\., 2022b](https://arxiv.org/html/2609.19843#bib.bib85)\), yet evidence that they consistently improve robustness remains mixed\([McKenzie et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib88);[Sharma et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib89)\)\. Since reasoning effort is a configurable parameter in contemporary LLM systems\([OpenAI, 2025a](https://arxiv.org/html/2609.19843#bib.bib18);[Google, 2025](https://arxiv.org/html/2609.19843#bib.bib16);[Anthropic, 2026](https://arxiv.org/html/2609.19843#bib.bib17)\), determining whether increased reasoning reduces susceptibility to manipulative choice architecture is an important empirical question with immediate practical implications, given that users increasingly rely on agent outputs with limited verification\([Rheu and Cho, 2025](https://arxiv.org/html/2609.19843#bib.bib27);[Steyvers et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib26)\)\. The present study addresses these gaps by asking whether the foundational premise of nudging extends to agentic decision\-making beyond textual output, and how reasoning shapes it through the following research questions: - RQ1\.Do LLM\-based GUI agents exhibit systematic shifts in choice behaviour when exposed to Type 1 \(automatic\) and Type 2 \(reflective\) nudges embedded in a graphical choice environment? - RQ2\.How does the reasoning configuration of LLM\-based GUI agents moderate their susceptibility to Type 1 \(automatic\) and Type 2 \(reflective\) nudges? To answer these questions, we treat an agent’s reasoning configuration as a functional proxy for the two modes of Dual\-Process Theory, with a no\-reasoning configuration corresponding to automatic System 1\-like processing and a high\-reasoning configuration to reflective System 2\-like processing\. This operationalisation builds on a growing body of work that casts explicit, step\-by\-step reasoning in LLMs as a computational analogue of deliberate System 2 processing\([Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5);[Booch et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib109);[Weston and Sukhbaatar, 2023](https://arxiv.org/html/2609.19843#bib.bib110)\)\. Rather than assuming this correspondence, we validate it through a manipulation check showing that increasing the reasoning configuration improves performance on a canonical System 2 task \(Appendix[12](https://arxiv.org/html/2609.19843#S12)\)\. We use the analogy only at the level of observable behaviour and make no claim that agents possess these cognitive systems\([Shanahan et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib82)\)\. Using this operationalisation, we conduct a set of computational simulations that adapt an established human decision\-making experiment in an online shopping context\([Ingendahl et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib81)\)to LLM\-based GUI agents, providing empirical insight into how they respond to equivalent nudge interventions\([Baird and Maruping, 2021](https://arxiv.org/html/2609.19843#bib.bib30)\)\. Concretely, we generate 3,600 agent instances across six leading models that pair each provider’s flagship with its smaller variant, namely GPT\-5\.4 and GPT\-5\.4\-mini\([OpenAI, 2025c](https://arxiv.org/html/2609.19843#bib.bib19)\), Gemini 3\.5 Flash and Gemini 3\.1 Flash Lite\([Google, 2025](https://arxiv.org/html/2609.19843#bib.bib16)\), and Claude Sonnet 4\.6 and Claude Haiku 4\.5\([Anthropic, 2026](https://arxiv.org/html/2609.19843#bib.bib17)\)\. Agents are assigned to one of two reasoning configurations and one of three experimental conditions: a control, a Type 1 nudge \(default\), or a Type 2 nudge \(social influence\)\. Each of the 3,600 agents completes a single shopping episode in a simulated grocery store modelled closely on the interface of the original experiment\([Ingendahl et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib81)\), selecting one product from each of six categories\. Across all agents, this results in 21,600 simulated decisions\. The results reveal that LLM\-based GUI agents are not only susceptible to nudge interventions in digital environments, but also show that reasoning configuration and nudge type interact substantially\. Consistent with Dual\-Process Theory, increasing an agent’s reasoning does not eliminate the influence of nudges but shifts the pathway through which they exert their effect\. On average, agents with higher reasoning move from being primarily affected by automatic Type 1 nudges to being more responsive to reflective Type 2 nudges\. Per\-model and per\-provider analyses further reveal substantial but structured heterogeneity\. The attenuation of default nudges under high reasoning is concentrated in the flagship models, whereas the amplification of social influence nudges is concentrated in the smaller model variants, suggesting that architectural, training, and scale differences shape the degree to which reasoning modulates nudge susceptibility\. This study makes three contributions\. First, it advances the debate on bounded rationality in AI decision\-making\([Simon, 1955](https://arxiv.org/html/2609.19843#bib.bib65);[Nguyen, 2024](https://arxiv.org/html/2609.19843#bib.bib25);[Chen et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib72)\)by moving beyond isolated demonstrations of individual biases\([Bösch, 2025](https://arxiv.org/html/2609.19843#bib.bib21);[Cherep et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib77)\)to a theory\-linked account that explains how reasoning affects agents’ susceptibility to nudges\. Second, it contributes to theorizing on agentic IT artifacts\([Baird and Maruping, 2021](https://arxiv.org/html/2609.19843#bib.bib30)\)by establishing nudge susceptibility as a behavioural property of agentic AI rather than a phenomenon unique to humans, general across model families and, in exploratory analysis, shaped by model scale\. Third, it bridges behavioural economics and information systems research by demonstrating that Dual\-Process Theory\([Kahneman, 2003](https://arxiv.org/html/2609.19843#bib.bib24)\), long applied to human decision\-makers, can inform the study of autonomous agents that do not reason like humans yet behave in human\-like ways within digital choice environments\([Brady et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib28)\)\. In particular, we show that the theory’s characteristic behavioural signature—reflective processing moderating automatic influence—arises in decision\-makers that possess none of the cognitive architecture or resource constraints the theory assumes, and that engaging deliberation changes susceptibility across nudge types rather than uniformly reducing it\. This decouples the dual\-process behavioural pattern from its human cognitive substrate and refines the theory’s prediction that System 2 engagement debiases judgement\. From a practical perspective, the study highlights the need for caution among organizations and individuals deploying autonomous agents\. Since digital choice architectures can materially affect agent decisions, interface design becomes a matter of governance rather than usability alone, and configuring agents for more deliberation changes which architectures influence them rather than whether they do\. From an academic perspective, the findings warrant closer investigation of the robustness of AI decision\-making in human\-designed environments and underscore the need to study agentic behaviour with the same behavioural rigour applied to human decision\-making\. Otherwise, we risk delegating authority to systems whose choices remain easily swayed by spurious cues\. The remainder of this paper is structured as follows\. Section[2](https://arxiv.org/html/2609.19843#S2)provides the theoretical background on reasoning in LLMs, LLM\-based agents, and nudge theory\. Section[3](https://arxiv.org/html/2609.19843#S3)derives the hypotheses from this foundation\. Section[4](https://arxiv.org/html/2609.19843#S4)outlines the research design, and Section[5](https://arxiv.org/html/2609.19843#S5)presents the empirical results\. Section[6](https://arxiv.org/html/2609.19843#S6)discusses the findings and their implications for research and practice\. Section[7](https://arxiv.org/html/2609.19843#S7)concludes with limitations and directions for future research\. ### 2Theoretical Background #### 2\.1Reasoning in Large Language Models Large Language Models are generative AI models, pre\-trained on vast amounts of data and comprising billions of parameters\([Feuerriegel et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib84)\)\. During pre\-training, they learn a statistical distribution to auto\-regressively predict the next token \(word part\) based on the previous token sequence\([Shanahan et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib82)\)\. On this basis, such models are capable of generalizing across diverse tasks without requiring additional task\-specific training, relying solely on textual prompts comprising natural language instructions or illustrative few\-shot examples\([Feuerriegel et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib84);[Liu et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib83)\)\. Beyond task instructions, LLMs can also be prompted to perform step\-by\-step reasoning about the task at hand, known as Chain\-of\-Thought \(CoT\) prompting\([Wei et al\., 2022b](https://arxiv.org/html/2609.19843#bib.bib85)\)\. By generating intermediate reasoning steps rather than a direct answer, LLMs emulate human reasoning in textual form, which often facilitates better task adaptation\([Huang and Chang, 2023](https://arxiv.org/html/2609.19843#bib.bib36);[Wei et al\., 2022b](https://arxiv.org/html/2609.19843#bib.bib85)\)\. This approach shifts LLMs away from fast and error\-prone responses toward more deliberate and reflective processing\([Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5)\)\. This distinction arguably parallels Dual\-Process Theory in behavioural economics\([Kahneman, 2003](https://arxiv.org/html/2609.19843#bib.bib24);[Evans and Stanovich, 2013](https://arxiv.org/html/2609.19843#bib.bib63)\), which states that human judgement relies on two complementary cognitive systems\. System 1 operates rapidly, automatically, and intuitively while requiring minimal mental effort, whereas System 2 functions more slowly and analytically by engaging effortful processing when individuals encounter complex or non\-routine tasks\. We emphasize that the parallel is functional rather than mechanistic, and that the analogy concerns observable input–output behaviour under different processing regimes, not the underlying computation\([Shanahan et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib82)\)\. Because LLMs generate outputs by predicting the conditional probability of the next token, they may exhibit System 1\-like behaviour, including cognitive biases that manifest as heuristics or shortcuts acquired from training data\([Jones and Steinhardt, 2022](https://arxiv.org/html/2609.19843#bib.bib87);[Suri et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib90)\)\. Existing work indeed shows that earlier models systematically reproduce the intuitive errors that characterize human System 1 responding\([Hagendorff et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib86);[Binz and Schulz, 2023](https://arxiv.org/html/2609.19843#bib.bib22)\)\. A prominent example is the bat\-and\-ball problem\([Watson, 2011](https://arxiv.org/html/2609.19843#bib.bib64), p\. 44\): > “A bat and ball cost $1\.10\. The bat costs one dollar more than the ball\. How much does the ball cost?” When relying on System 1, humans and earlier LLMs \(specifically GPT\-3\) frequently produce the intuitive but incorrect answer of $0\.10\([Binz and Schulz, 2023](https://arxiv.org/html/2609.19843#bib.bib22);[Hagendorff et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib86)\)\. Arriving at the correct solution of $0\.05 requires additional mental effort and engaging System 2 thinking\([Watson, 2011](https://arxiv.org/html/2609.19843#bib.bib64)\)\. Recent advances in LLM training have introduced mechanisms that encourage more explicit step\-by\-step reasoning, aiming to reduce susceptibility to System 1\-style biases and improve overall performance\([Hagendorff et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib86);[Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5)\)\. Models that integrate this process internally, often referred to as reasoning models\([Huang and Chang, 2023](https://arxiv.org/html/2609.19843#bib.bib36);[Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5)\), generally achieve higher accuracy on complex tasks and appear less vulnerable to certain failures associated with immediate, heuristic\-based responding\([Huang and Chang, 2023](https://arxiv.org/html/2609.19843#bib.bib36);[Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5)\)\. Crucially, however, unlike human thinking, in which individuals shift dynamically between System 1 and System 2 to meet the demands of the situation\([Kahneman, 2003](https://arxiv.org/html/2609.19843#bib.bib24)\), LLMs do not switch between these modes autonomously in the same way\. Instead, the degree of reasoning is typically determined by their training and configuration\([Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5)\)\. Current models permit reasoning effort to be adjusted programmatically, ranging from none to high\([OpenAI, 2025a](https://arxiv.org/html/2609.19843#bib.bib18);[Google, 2025](https://arxiv.org/html/2609.19843#bib.bib16);[Anthropic, 2026](https://arxiv.org/html/2609.19843#bib.bib17)\)\. Non\-reasoning configurations resemble System 1 in that they favour fast and less deliberative responses, whereas high\-reasoning configurations resemble System 2 in that they involve slower and more effortful processing\. Nevertheless, the evidence that stronger reasoning yields more robust behaviour is mixed\. While higher reasoning can reduce biases overall, and System 1 biases in particular\([Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5);[Zhang et al\., 2024a](https://arxiv.org/html/2609.19843#bib.bib8)\), high\-reasoning models remain sensitive to systematic errors\([Huang and Chang, 2023](https://arxiv.org/html/2609.19843#bib.bib36);[Guo et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib52)\), and their deviations do not always mirror human ones\([Macmillan\-Scott and Musolesi, 2024](https://arxiv.org/html/2609.19843#bib.bib51)\)\. Increased scale can even worsen performance on some tasks, a pattern termed inverse scaling\([McKenzie et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib88)\), and extended chains of reasoning can degrade decisions where deliberation is counterproductive\([Liu et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib53)\)\. Critically, greater capability and stronger alignment to human preferences can introduce new vulnerabilities rather than remove existing ones\. A salient example is sycophancy, the tendency of models trained with human feedback to conform to a user’s stated views, which becomes more pronounced in larger and more heavily aligned models\([Perez et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib92);[Sharma et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib89)\)\. A closely related phenomenon is conformity, whereby LLMs align with a presented majority even when it is incorrect\([Weng et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib93)\)\. Reduced cognitive bias therefore does not necessarily imply more robust decision\-making\([Dentella et al\., 2026](https://arxiv.org/html/2609.19843#bib.bib6)\)\. This matters especially for LLM\-based GUI agents, where reasoning may shift the route of external influence rather than close it off\. #### 2\.2Large Language Model\-based Agents Beyond enhanced reasoning capabilities, LLMs can be augmented to use external tools such as search engines, calculators, calendars, or application programming interfaces \(APIs\)\([Qu et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib34);[Schick et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib35)\)\. When an LLM combines tool use with reasoning, it adopts the reasoning\-and\-acting \(ReAct\) framework, which enables it to operate as an autonomous agent that interleaves deliberation with action\([Göldi and Rietsche, 2025](https://arxiv.org/html/2609.19843#bib.bib33)\)\. Building on this capability, LLM\-based agents can now interact with graphical user interfaces \(GUIs\) such as operating systems, software applications, and browsers\([Hu et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib46);[Nguyen et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib47);[Sager et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib45)\)\. These systems, often referred to as computer control agents\([Sager et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib45)\)or GUI agents\([Nguyen et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib47)\), move beyond conversational exchange to autonomous action\. A GUI agent perceives the interface through rendered pixels and structured representations such as the document object model, plans over multiple steps, navigates between elements, and commits to a choice by taking an action on the user’s behalf\([Nguyen et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib47);[Sager et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib45)\)\. This perception\-plan\-act loop distinguishes agentic behaviour from the single\-turn text elicitation in which LLM decision\-making has typically been studied\. Such capabilities underpin products such as Perplexity’s Comet browser\([Perplexity, 2025a](https://arxiv.org/html/2609.19843#bib.bib10)\), ChatGPT Agent\([OpenAI, 2025b](https://arxiv.org/html/2609.19843#bib.bib15)\), and open\-source projects like the Browser Use library\([Browser Use, 2025](https://arxiv.org/html/2609.19843#bib.bib40)\), which in turn open the door to large\-scale task automation\([Marreed et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib44)\)\. As these agents become capable of automating information retrieval, product identification, and purchasing, organizations are increasingly embedding them into business processes\([Singla et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib20)\), while users are placing growing expectations on them to autonomously complete tasks on their behalf\([Perplexity, 2025b](https://arxiv.org/html/2609.19843#bib.bib39)\)\. In keeping with this focus, research on LLM\-based GUI agents has concentrated on capability and task success, evaluated through benchmarks of perception, grounding, planning, and goal completion\([Hu et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib46);[Nguyen et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib47);[Sager et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib45)\), with comparatively less attention to how these agents decide when the environment is designed to influence choice\. LLM\-based GUI agents are exposed to such influence on two fronts\. They inherit the behavioural biases of the LLMs that power them, as discussed in Section[2\.1](https://arxiv.org/html/2609.19843#S2.SS1), and they act within the human\-oriented choice architectures embedded in the environments they navigate\. #### 2\.3Nudging Grounded in the realization that human judgement suffers from systematic cognitive biases\([Kahneman, 2003](https://arxiv.org/html/2609.19843#bib.bib24);[Tversky and Kahneman, 1974](https://arxiv.org/html/2609.19843#bib.bib23);[Tversky and Kahneman, 1981](https://arxiv.org/html/2609.19843#bib.bib67)\), nudge theory marked a shift in understanding how the design of choice architecture influences human decision\-making\([Hummel and Maedche, 2019](https://arxiv.org/html/2609.19843#bib.bib57)\)\. At its core, a nudge refers to a subtle modification in how options are presented that predictably influences behaviour without forbidding any options, restricting freedom of choice, or significantly altering economic incentives\([Thaler and Sunstein, 2009](https://arxiv.org/html/2609.19843#bib.bib69)\)\. It challenges the classical notion of the consistently rational homo economicus, holding that human behaviour often departs from deliberative reasoning and is instead influenced by systematic biases\([Stanovich, 2011](https://arxiv.org/html/2609.19843#bib.bib66)\)and bounded rationality\([Simon, 1955](https://arxiv.org/html/2609.19843#bib.bib65)\)\. Over time, the principles of nudging extended from analogue contexts into the digital sphere, giving rise to the notion of digital nudging\([Weinmann et al\., 2016](https://arxiv.org/html/2609.19843#bib.bib61);[Mirsch et al\., 2017](https://arxiv.org/html/2609.19843#bib.bib60)\)\. Digital nudging steers users toward particular selections within digital choice environments by modifying interface design elements\([Schneider et al\., 2018](https://arxiv.org/html/2609.19843#bib.bib58)\), for example through default settings or the visual highlighting of specific options\([Weinmann et al\., 2016](https://arxiv.org/html/2609.19843#bib.bib61)\)\. Building on Dual\-Process Theory\([Kahneman, 2003](https://arxiv.org/html/2609.19843#bib.bib24);[Evans and Stanovich, 2013](https://arxiv.org/html/2609.19843#bib.bib63)\),[Hansen and Jespersen \(2013\)](https://arxiv.org/html/2609.19843#bib.bib62)proposed categorizing nudges according to the mode of thinking they engage\. Type 1 nudges act on the automatic system without invoking reflective thinking, whereas Type 2 nudges target the premises and attention of reflective thinking\. The two nudges examined in this study, defaults and social influence, are canonical instances of these respective types\. A default nudge preselects one option so that it is realized unless the decision\-maker actively chooses otherwise\. Its influence operates through the automatic system, which makes it a Type 1 nudge\. Defaults exploit the status quo bias, the tendency to prefer an existing or preselected state and to avoid the cognitive effort of evaluating alternatives\([Samuelson and Zeckhauser, 1988](https://arxiv.org/html/2609.19843#bib.bib96);[Ritov and Baron, 1992](https://arxiv.org/html/2609.19843#bib.bib56)\)\. Because retaining the preselected option requires no deliberation and follows the path of least resistance, defaults steer behaviour without engaging reflective thought\([Thaler and Sunstein, 2009](https://arxiv.org/html/2609.19843#bib.bib69)\)\. The effect is among the most robust in the nudging literature, as illustrated by the finding that specifying organ donation as the default markedly raises participation relative to an opt\-in arrangement\([Johnson and Goldstein, 2003](https://arxiv.org/html/2609.19843#bib.bib94)\)\. A social influence nudge presents information about the choices of others, so that the decision\-maker adjusts behaviour in light of that reference point\. Because processing this information requires attention to and reflection on social cues, it constitutes a Type 2 nudge\. Its influence rests on two conformity motives, namely normative influence, the desire to gain social approval, and informational influence, the assumption that the behaviour of others signals the appropriate course of action, particularly under uncertainty\([Deutsch and Gerard, 1955](https://arxiv.org/html/2609.19843#bib.bib43);[Cialdini and Trost, 1998](https://arxiv.org/html/2609.19843#bib.bib55)\)\. Descriptive social norms operationalise this mechanism by indicating what most people do in a given situation\([Cialdini et al\., 1990](https://arxiv.org/html/2609.19843#bib.bib98)\)\. Such norms reliably shift behaviour across domains, from towel reuse among hotel guests\([Goldstein et al\., 2008](https://arxiv.org/html/2609.19843#bib.bib95)\)to household energy conservation\([Allcott, 2011](https://arxiv.org/html/2609.19843#bib.bib97)\)and pro\-environmental product choices in online shopping\([Demarque et al\., 2015](https://arxiv.org/html/2609.19843#bib.bib42)\)\. Unlike adversarial exploits such as prompt injection, which hide hostile instructions in the content an agent processes to subvert its behaviour\([Greshake et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib100);[Zhan et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib102);[Debenedetti et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib101)\), nudges are a routine feature of online interface design\([Schneider et al\., 2018](https://arxiv.org/html/2609.19843#bib.bib58);[Caraban et al\., 2019](https://arxiv.org/html/2609.19843#bib.bib59)\), so agents encounter them by default in the ordinary course of a task rather than as an external intrusion\. Because the agent commits to choices on the user’s behalf, the consequences of any such influence accrue to the delegating user\. Existing evidence of nudge\-like effects in LLMs derives from prompt\-level textual elicitation or addresses only a single nudge type\. A nudge embedded in a graphical interface, however, must be perceived within the rendered choice environment, integrated across planning steps, and acted upon, so a bias observed in single\-turn textual elicitation need not persist through this perception\-plan\-act loop\. Moreover, reasoning effort can be fixed closer to System 1\- or System 2\-like processing, which makes it a natural moderator of susceptibility, yet how this configuration shapes responsiveness to each nudge type has not yet been examined\. The present study addresses these gaps, which the next section develops into testable hypotheses\. ### 3Hypothesis Development A growing body of research demonstrates that LLMs reproduce human\-like cognitive biases in their textual outputs\. Models exhibit established behavioural effects such as framing and availability biases\([Jones and Steinhardt, 2022](https://arxiv.org/html/2609.19843#bib.bib87);[Hagendorff et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib86);[Suri et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib90)\), and instruction tuning and alignment can amplify such tendencies rather than eliminate them\([Itzhak et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib103);[Echterhoff et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib13)\)\. Beyond individual biases, LLMs display inconsistent decision patterns\([Guo et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib52)\), exhibit systematic biases when evaluating alternatives\([Koo et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib104)\), and depart from rational\-choice models in ways that do not always match human behaviour\([Macmillan\-Scott and Musolesi, 2024](https://arxiv.org/html/2609.19843#bib.bib51)\), though in applied settings their errors can resemble those of human decision\-makers\([Chen et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib7)\)\. Several of these vulnerabilities arise through mechanisms closely related to digital nudging\. LLMs are sensitive to superficial variations in how options are presented\([Cherep et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib77)\)and reproduce established choice\-architecture effects such as anchoring\([Nguyen, 2024](https://arxiv.org/html/2609.19843#bib.bib25)\)\. Most relevant to the present study,[Bösch \(2025\)](https://arxiv.org/html/2609.19843#bib.bib21)show that the decoy effect systematically influences the product choices of LLM\-based GUI agents, providing initial evidence that behavioural effects observed in text\-based interaction can persist when LLMs act within graphical interfaces\. Taken together, these findings indicate that although LLM\-based agents do not “reason” in a human sense but generate behaviour through learned statistical patterns\([Brady et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib28)\), their decisions can nonetheless display vulnerabilities comparable to those observed in human cognition\. Extrapolating these findings to agentic settings, we expect LLM\-based GUI agents operating in online shopping environments to be susceptible to both Type 1 and Type 2 nudges embedded in the interface, each of which reliably affects human choices\([Johnson and Goldstein, 2003](https://arxiv.org/html/2609.19843#bib.bib94);[Goldstein et al\., 2008](https://arxiv.org/html/2609.19843#bib.bib95)\)\. H1:LLM\-based GUI agents exhibit systematic shifts in choice behaviour when exposed to Type 1 \(automatic\) and Type 2 \(reflective\) nudge interventions\. In human decision\-making, the moderating role of cognitive engagement is contested\. Some studies find that individual differences in cognitive ability and disposition shape susceptibility to nudges\([Ingendahl et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib81);[De Ridder et al\., 2022](https://arxiv.org/html/2609.19843#bib.bib49)\), whereas others report that people are similarly nudgeable regardless of the cognitive system engaged\([Van Gestel et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib48)\)\. To examine whether this relationship extends to LLM\-based agents, we treat the reasoning configuration as a functional proxy for the two dual\-process modes established in Section[2\.1](https://arxiv.org/html/2609.19843#S2.SS1), with a no\-reasoning configuration corresponding to System 1\-like and a high\-reasoning configuration to System 2\-like processing\([Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5)\)\. In dual\-process terms, deliberation counteracts the automatic responses that Type 1 nudges exploit\([Stanovich, 2011](https://arxiv.org/html/2609.19843#bib.bib66)\)\. Type 1 nudges operate through automatic processing and the path of least resistance, without engaging reflection\. Because System 2\-style reasoning counteracts the intuitive, heuristic\-driven responses associated with automatic processing\([Stanovich, 2011](https://arxiv.org/html/2609.19843#bib.bib66)\), an agent under a high\-reasoning configuration should weigh the alternatives more deliberately rather than accepting a preselected option by default\. We therefore expect a high\-reasoning configuration to attenuate susceptibility to automatic nudges relative to a no\-reasoning configuration\. H2a:LLM\-based GUI agents under a high\-reasoning configuration are less susceptible to Type 1 \(automatic\) nudges than those under a no\-reasoning configuration\. Type 2 nudges, by contrast, operate by supplying information that the decision\-maker is expected to weigh, and thus take effect only if that information is actually processed\([Caraban et al\., 2019](https://arxiv.org/html/2609.19843#bib.bib59)\)\. An agent under a high reasoning configuration, which engages the reflective processing these nudges target, should attend to and integrate such cues more thoroughly than one under a no\-reasoning configuration\. Greater reasoning may therefore not immunize agents against influence but instead deepen their engagement with reflective cues, increasing rather than reducing susceptibility to Type 2 nudges\([Brady et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib28)\)\. H2b:LLM\-based GUI agents under a high\-reasoning configuration are more susceptible to Type 2 \(reflective\) nudges than those under a no\-reasoning configuration\. ### 4Research Design To test our hypotheses, we replicate an established consumer choice experiment by[Ingendahl et al\. \(2021\)](https://arxiv.org/html/2609.19843#bib.bib81), which investigated two of the most widely studied nudges, default options and social influence, in an online shopping context\. We selected this experiment for three reasons\. First, both nudge types have consistently been shown to influence consumer behaviour\([Demarque et al\., 2015](https://arxiv.org/html/2609.19843#bib.bib42);[Székely et al\., 2016](https://arxiv.org/html/2609.19843#bib.bib41)\)and rank among the most frequently applied nudges in the literature\([Hummel and Maedche, 2019](https://arxiv.org/html/2609.19843#bib.bib57)\)\. Second, as established in Section[2\.3](https://arxiv.org/html/2609.19843#S2.SS3), the two nudges map cleanly onto the dual\-process typology, with the default nudge operating as a Type 1 \(automatic\) intervention and the social influence nudge as a Type 2 \(reflective\) one\([Caraban et al\., 2019](https://arxiv.org/html/2609.19843#bib.bib59)\)\. Third, the online shopping scenario closely mirrors the digital environments that LLM\-based GUI agents are increasingly deployed to navigate\. We adopt the original procedure, materials, and task structure but apply them for a different research objective\. Whereas[Ingendahl et al\. \(2021\)](https://arxiv.org/html/2609.19843#bib.bib81)examined how individual personality traits, such as need for cognition and uniqueness, moderate the effectiveness of these nudges in human consumers, we use the design to investigate how LLM\-based GUI agents under different reasoning configurations respond to them\. Before the main experiment, we validated the System 1 versus System 2 interpretation of the reasoning configuration in a pre\-experiment built on Kahneman’s example of counting the occurrences of a letter as a canonical System 2 task\([Kahneman, 2011](https://arxiv.org/html/2609.19843#bib.bib107)\), which LLMs solve reliably only when explicit reasoning is engaged\([Zhang et al\., 2024b](https://arxiv.org/html/2609.19843#bib.bib105);[Xu and Ma, 2025](https://arxiv.org/html/2609.19843#bib.bib106)\)\. Across the same six models and reasoning configurations as the main experiment \(five letter\-counting items, 100 repetitions per configuration\), accuracy rose from 80\.3% under no reasoning to 99\.3% under high reasoning, with the gains concentrated on novel pseudo\-words that cannot be answered from memory, justifying the reasoning configuration as a functional proxy for System 1 versus System 2 processing\. Appendix[12](https://arxiv.org/html/2609.19843#S12)reports the full design, the statistical tests, and a reasoning\-trace analysis\. #### 4\.1Materials We used a simulated online grocery store modelled closely after the interface used in the original human experiment\([Ingendahl et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib81)\)\. The store consisted of six product categories, namely tomatoes, bananas, bread, coffee beans, milk, and pasta\. Each category contained two products displayed side by side, each with a product image and a brief description\. All original materials, including instructions and descriptions, were translated from German, the language of the original experiment, to English to better align with the linguistic capabilities of the GUI agents\. Figure[1](https://arxiv.org/html/2609.19843#S4.F1)shows a representative screenshot of the layout of the store interface\. Figure 1:Example product display of the simulated grocery store across two categories\. #### 4\.2Technical Implementation We implemented the online grocery store using the FastAPI111[https://fastapi\.tiangolo\.com](https://fastapi.tiangolo.com/)web framework to provide a browser\-accessible front end for the experimental task\. The LLM\-based GUI agents interacted with the store through the Browser Use library\([Browser Use, 2025](https://arxiv.org/html/2609.19843#bib.bib40)\), which enables agents to operate web browsers autonomously by navigating webpages, selecting elements, clicking buttons, entering inputs, reading visual and textual content, and following the web application’s own workflow\. Through this interface, the agents perceived and interacted with the graphical user interface of the store much as a human participant would\. The Browser Use library supports built\-in logging, which provides insight into the reasoning behind each agent decision without modifying the system prompt or the instructions given to the agents\. To enable reproducibility and future replication, we make the source code and experimental analyses publicly available online222[https://anonymous\.4open\.science/r/Nudging\_LLM\_Agents\-B1D5](https://anonymous.4open.science/r/Nudging_LLM_Agents-B1D5)\. #### 4\.3Procedure LLM\-based GUI agents were tasked with purchasing six products from the online store for an upcoming meal with friends\. At the beginning of each trial, all agents received a standardized prompt that closely mirrored the scenario description used in the original study\. The prompt instructed them to visit the online shop, inspect the available products, and follow the on\-screen instructions to complete the purchase\. PromptI’m planning a meal for my friends tomorrow\. After some quick consideration, I realize I’m still missing a few products\. 1\. Go to the online supermarkethttp://127\.0\.0\.1:8000/ 2\. Learn about the different products and select your preferred product in each category\. 3\. Follow the instructions on the website carefully to complete your purchase\. After receiving the prompt, agents navigated to the online shop, where they browsed the six product categories via a tab menu and selected their preferred product in each category \(exactly one of the two options\)\. Once a product from each category was selected, agents advanced to a summary page, where they could review their selections and, if desired, revise them before confirming\. The confirmed final selections constituted the dependent variable in our agent\-based replication\. Upon confirming their selections, agents were redirected to an end page indicating that the items had been added to the shopping cart\. This page signalled the completion of the task and allowed the agents to terminate the session and close the online store window\. Figure[2](https://arxiv.org/html/2609.19843#S4.F2)illustrates the full procedure\. Figure 2:The experimental procedure for all simulations\. #### 4\.4Conditions The experiment followed a three\-group between\-subjects design\. Agents were randomly assigned to one of three conditions\. In the control \(no\-nudge\) condition, products were displayed neutrally without any nudge intervention\. In the default nudge condition, one of the two products within each category was preselected when the agent opened the page\. In the social influence nudge condition, one product per category was presented as having the highest customer satisfaction rate for that category\. Figure[3](https://arxiv.org/html/2609.19843#S4.F3)illustrates the three conditions\. As in the original experiment, both the order of product categories and the position of products within each category were randomized for each agent to mitigate ordering effects\. To control for effects of product appearance or description, we employed a standard counterbalanced design\. Within each experimental condition, half of the agents were nudged toward one option \(product set A\) and the other half toward the alternative option \(product set B\)\. This design ensures that any inherent advantage of one product set over another cancels out, so that in the absence of a nudging effect the nudged option should be selected at chance level of 50%\. Figure 3:The three nudge conditions\. The control condition omits both nudge types \(left\), the social influence nudge condition \(center\), and the default nudge condition \(right\)\. #### 4\.5Agent Participants To capture variation across models, we varied the agents’ underlying backbone across six leading models from three providers, pairing each provider’s flagship model with its smaller, faster variant, namely OpenAI’s GPT\-5\.4 and GPT\-5\.4\-mini\([OpenAI, 2025c](https://arxiv.org/html/2609.19843#bib.bib19)\), Google’s Gemini 3\.5 Flash and Gemini 3\.1 Flash Lite\([Google, 2025](https://arxiv.org/html/2609.19843#bib.bib16)\), and Anthropic’s Claude Sonnet 4\.6 and Claude Haiku 4\.5\([Anthropic, 2026](https://arxiv.org/html/2609.19843#bib.bib17)\)\. This two\-factor sampling of backbones \(provider×\\timesmodel size\) allows us to assess whether nudge susceptibility and its moderation by reasoning generalize across model families and scales\. In total, we generated 3,600 distinct agents, each conceptually modelled after a human participant, and initialized with one of six LLMs and one of two predefined reasoning configurations\. This yielded 600 agents per LLM, evenly split between the reasoning configurations \(300 each\)\. We operationalised reasoning by adjusting each model’s reasoning effort, yielding two configurations akin to System 1 and System 2 processing, namely no reasoning \(reasoning disabled\) and high reasoning \(high reasoning effort\)\. Following a between\-subjects design, 200 agents per model were assigned to each experimental condition \(no nudge control, default nudge, and social influence nudge\)\. Within each condition, agents were evenly divided between the no\-reasoning and high\-reasoning configurations, with 100 agents in each subgroup\. Each agent completed one full shopping episode, selecting one product from each of the six categories\. Agents were randomly assigned to receive nudges toward either product set A or product set B\. Because each agent interacted once with each product category, this yielded six interactions per trial, 3,600 interactions per model, and 21,600 interactions across all six models\. #### 4\.6Model Specification To test our hypotheses, we estimated Bayesian multilevel logistic regression models \(Bernoulli likelihood, logit link\) using the brms package \(v2\.23\.0\) in R\([Bürkner, 2017](https://arxiv.org/html/2609.19843#bib.bib3)\)\. The unit of analysis was the individual agent interaction, corresponding to one product selection made by an agent during a trial\. The dependent variable is a binary indicator of whether the agent selected the product designated as the target, coded 1 if the target option was chosen and 0 otherwise\. In the default and social influence conditions, the target corresponded to the nudged option, whereas in the no\-nudge condition it was assigned randomly to enable comparability across conditions\. The main independent variables were the experimental condition \(control, default, or social influence\) and the reasoning configuration of the agents \(no reasoning versus high reasoning\)\. All predictors were coded as binary dummy variables\. Concretely, the dummy variable for reasoning configuration was coded 0 for no reasoning and 1 for high reasoning, the dummy for the default condition 0 for no default and 1 for default, and the dummy for the social influence condition 0 for no social influence and 1 for social influence\. The reference category therefore corresponds to the control condition with no\-reasoning agents\. To account for the non\-independence of repeated choices made by the same agent across the six product categories, we included random intercepts for the agents\. We estimated the models within a Bayesian framework because several design cells, most notably for the two Claude models, exhibited target selection rates at or near 100%\. Such \(quasi\-\)complete separation renders maximum\-likelihood estimation of frequentist mixed\-effects models \(e\.g\.,glmerfrom the lme4 package\([Bates et al\., 2015](https://arxiv.org/html/2609.19843#bib.bib37)\)\) unreliable, producing divergent coefficients and standard errors\. In the Bayesian specification, a weakly informative Student\-t\(3,0,2\.5\)t\(3,0,2\.5\)prior on all regression coefficients\([Gelman et al\., 2008](https://arxiv.org/html/2609.19843#bib.bib99)\)keeps the posterior proper under separation while leaving all estimable effects essentially unshrunk\. All remaining parameters \(the intercept and the random\-intercept standard deviation\) use the brms weakly informative default priors\. Each model was estimated with 4 chains of 4,000 iterations \(1,000 warmup\), yielding 12,000 post\-warmup draws, and all fits converged \(R^≈1\.00\\widehat\{R\}\\approx 1\.00, with high effective sample sizes\)\. Because Bayesian models report nopp\-values, we consider an effect*credible*when the 95% credible interval \(CI\) of its posterior excludes zero on the log\-odds scale, equivalently when the CI of the odds ratio excludes one\. For cells at the 100% ceiling, the point estimate is partly prior\-dependent, so conclusions there are read from the data\-driven CI bound rather than the point estimate\. Our main analysis comprises three pooled models estimated over all 21,600 interactions of the 3,600 agents\. Model 1 regresses target selection on condition, reasoning, and their interaction, with random intercepts per agent \(target∼condition×reasoning\+\(1∣agent\)\\text\{target\}\\sim\\text\{condition\}\\times\\text\{reasoning\}\+\(1\\mid\\text\{agent\}\)\)\. Model 2 adds the LLM provider \(OpenAI, Google, Anthropic, with Anthropic as the reference category\) as a fixed\-effect control, and Model 3 additionally adds model size \(flagship versus small variant, with flagship as the reference\)\. Both are included to verify that the condition and interaction effects are robust to differences between model families and scales\. We treat provider purely as a control, since it bundles architecture, training, and alignment differences that cannot be interpreted individually, whereas model size represents scale, a construct we return to in the exploratory analysis in Section[5](https://arxiv.org/html/2609.19843#S5)\. Hypothesis tests are based on posterior odds ratio contrasts derived from the fitted models, namely pairwise contrasts between conditions \(H1\) and contrasts of high versus no reasoning within each condition \(H2a and H2b\), summarized by posterior medians with 95% highest\-posterior\-density intervals\. As robustness checks, we re\-estimated the Model 1 specification separately for each provider and for each of the six model backbones\. We additionally conducted an exploratory analysis by model size tier \(the three flagship and the three small models pooled\) to characterize how the reasoning moderation varies with scale\. Complete posterior summaries of all models are reported in Appendices[8](https://arxiv.org/html/2609.19843#S8)–[11](https://arxiv.org/html/2609.19843#S11)\. ### 5Results #### 5\.1Hypothesis Tests Figure[4](https://arxiv.org/html/2609.19843#S5.F4)displays the estimated probability of selecting the target product for each experimental condition and reasoning configuration, derived from the three pooled Bayesian multilevel logistic regression models\. The target is the product toward which the agent is nudged in the treatment conditions\. Each panel shows posterior point estimates with 95% credible intervals\. Values above 50% indicate a systematic shift toward the nudged product\. Looking at the pooled results across all six models \(Figure[4](https://arxiv.org/html/2609.19843#S5.F4)\), several patterns emerge\. First, in the control condition without a nudge, the estimated probability of selecting the target product is approximately 50% \(CI\.95CI\_\{\.95\}\[47%, 53%\] for no\-reasoning agents andCI\.95CI\_\{\.95\}\[46%, 51%\] for high\-reasoning agents\), consistent with the counterbalanced design\. Because each product serves as the nudge target for half of the agents, inherent preference differences between the two products cancel out by construction, and the expected control\-condition selection rate is 50% at every level of aggregation\. Departures from 50% in the nudge conditions can therefore be attributed to the nudges rather than to the products themselves\. Second, the estimated mean probability in the default nudge condition rises to approximately 92% \(CI\.95CI\_\{\.95\}\[91%, 93%\]\) for no\-reasoning agents and 86% \(CI\.95CI\_\{\.95\}\[84%, 87%\]\) for high\-reasoning agents; here high reasoning*lowers*selection of the preselected option, so agents that reasoned more were less likely to accept the default\. Third, in the social influence nudge condition the probability rises further, to approximately 96% \(CI\.95CI\_\{\.95\}\[96%, 97%\]\) for no\-reasoning agents and 98% \(CI\.95CI\_\{\.95\}\[97%, 98%\]\) for high\-reasoning agents, reversing the direction of the reasoning effect: high reasoning now*increases*selection of the socially endorsed option\. This indicates that LLM\-based GUI agents were susceptible to both nudge types and that the reasoning configuration appears to moderate these effects in opposing directions\. Adding the provider control \(Model 2, Figure[5](https://arxiv.org/html/2609.19843#S5.F5)\) and the provider and size controls \(Model 3, Figure[6](https://arxiv.org/html/2609.19843#S5.F6)\) leaves the estimated probabilities virtually unchanged, indicating that the condition and interaction effects are robust to differences between model families and scales\. Figure 4:Model 1 \(pooled\): Estimated target product selection probabilities across conditions and reasoning configurations for the three pooled models\.Figure 5:Model 2 \(\+ provider\): Estimated target product selection probabilities across conditions and reasoning configurations for the three pooled models\.Figure 6:Model 3 \(\+ provider \+ size\): Estimated target product selection probabilities across conditions and reasoning configurations for the three pooled models\.Having read the estimated probabilities off Figures[4](https://arxiv.org/html/2609.19843#S5.F4)—[6](https://arxiv.org/html/2609.19843#S5.F6), we now turn to the coefficient\-level estimates that formally test our hypotheses\. Table[1](https://arxiv.org/html/2609.19843#S5.T1)presents the results of the three pooled Bayesian multilevel logistic regressions\. Model \(1\) serves as the primary inferential model, while Models \(2\) and \(3\) add the provider and size controls\. All condition and interaction coefficients are essentially identical across the three specifications, so we report the estimates of Model \(1\) in the text\. Complete posterior summaries for all three models, including credible intervals and odds ratios, are provided in Appendix[8](https://arxiv.org/html/2609.19843#S8), and the corresponding per\-provider and per\-model robustness fits are reported in Appendices[9](https://arxiv.org/html/2609.19843#S9)and[11](https://arxiv.org/html/2609.19843#S11)\. The agent random\-intercept standard deviation \(1\.151\.15,1\.111\.11, and1\.101\.10across the three specifications\) quantifies how much agents differ, on the log\-odds scale, in their baseline propensity to select the target once condition and reasoning are accounted for\. Behaviourally, this variance reflects the within\-run consistency of each agent’s six choices together with the residual stochasticity that sampling introduces, rather than durable participant\-level heterogeneity\. Starting with H1, the pooled models in Table[1](https://arxiv.org/html/2609.19843#S5.T1)reveal credible positive effects for both the default nudge condition \(β=2\.48\\beta=2\.48,CI\.95CI\_\{\.95\}\[2\.28,2\.68\]\[2\.28,\\,2\.68\]\) and the social influence nudge condition \(β=3\.32\\beta=3\.32,CI\.95CI\_\{\.95\}\[3\.10,3\.55\]\[3\.10,\\,3\.55\]\), indicating that both nudge types substantially increased the likelihood of choosing the target product relative to the no\-nudge baseline\. Expressed as odds ratios, the default nudge multiplied the odds of selecting the target product by a factor of8\.708\.70\(CI\.95CI\_\{\.95\}\[7\.54,9\.97\]\[7\.54,\\,9\.97\]\) and the social influence nudge by a factor of37\.2437\.24\(CI\.95CI\_\{\.95\}\[31\.04,44\.20\]\[31\.04,\\,44\.20\]; Table[84](https://arxiv.org/html/2609.19843#S8.T4)\)\. Both contrasts remain virtually unchanged under the provider and size controls \(OROR=8\.948\.94for defaults, and38\.6938\.69for social influence\)\. These effects are corroborated by all six per\-model regressions, in which both the default and the social influence coefficients are credible for every model \(Appendix[11](https://arxiv.org/html/2609.19843#S11), Tables D1–D6\)\. In sum, these results provide robust support for H1\. LLM\-based GUI agents exhibit systematic shifts in choice behaviour when exposed to both automatic and reflective nudging interventions, regardless of the underlying model, provider and size\. For H2a, we tested whether agents under a high\-reasoning configuration were less susceptible to automatic Type 1 nudges, that is, defaults\. The pooled model shows that reasoning alone did not credibly affect behaviour in the no\-nudge condition \(β=−0\.05\\beta=\-0\.05,CI\.95CI\_\{\.95\}\[−0\.21,0\.12\]\[\-0\.21,\\,0\.12\], including zero\), confirming that differences emerge only when a nudge is present\. Critically, the interaction between the reasoning configuration and the default nudge condition is credibly negative \(β=−0\.63\\beta=\-0\.63,CI\.95CI\_\{\.95\}\[−0\.89,−0\.36\]\[\-0\.89,\\,\-0\.36\]\) and robust to the provider and size controls \(β=−0\.62\\beta=\-0\.62,CI\.95CI\_\{\.95\}\[−0\.88,−0\.36\]\[\-0\.88,\\,\-0\.36\]\), indicating that high\-reasoning agents were less affected by default nudges than no\-reasoning agents\. Table 1:Bayesian multilevel logistic regression results for the three pooled models\. Model \(1\) regresses target selection on condition, reasoning, and their interaction\. Model \(2\) adds the provider control, and Model \(3\) adds the provider and size controls\. Cell entries are posterior means \(log\-odds\) with posterior standard deviations in parentheses\. Complete posterior summaries, including 95% credible intervals and odds ratios, are reported in Appendix[8](https://arxiv.org/html/2609.19843#S8)\.22footnotetext:Priors are weakly informative Student\-t\(3,0,2\.5\)t\(3,0,2\.5\)on all coefficients, with 4 chains of 4 000 iterations \(1 000 warmup\) each\. AllR^≈1\.00\\widehat\{R\}\\approx 1\.00\.22footnotetext:aProvider controls \(Anthropic as the reference category\), included to verify that the condition and interaction effects are robust to model family\. Provider is treated purely as a control\.22footnotetext:bModel size control \(flagship as the reference category\)\. The size coefficient is interpreted as an index of model scale \(see Section[5](https://arxiv.org/html/2609.19843#S5)\)\.Target Selected\(1\)\(2\)\(3\)Dependent Variable:Pooled\+ Provider\+ Provider \+ SizeIntercept−0\.02\-0\.020\.540\.54\*0\.460\.46\*\(0\.06\)\(0\.06\)\(0\.07\)\(0\.07\)\(0\.08\)\(0\.08\)Defaults2\.482\.48\*2\.502\.50\*2\.502\.50\*\(0\.10\)\(0\.10\)\(0\.10\)\(0\.10\)\(0\.10\)\(0\.10\)Social Influence3\.323\.32\*3\.373\.37\*3\.353\.35\*\(0\.11\)\(0\.11\)\(0\.11\)\(0\.11\)\(0\.11\)\(0\.11\)Reasoning high−0\.05\-0\.05−0\.05\-0\.05−0\.05\-0\.05\(0\.08\)\(0\.08\)\(0\.08\)\(0\.08\)\(0\.08\)\(0\.08\)Provider Googlea–−0\.99\-0\.99\*−0\.98\-0\.98\*–\(0\.07\)\(0\.07\)\(0\.07\)\(0\.07\)Provider OpenAIa–−0\.68\-0\.68\*−0\.69\-0\.69\*–\(0\.07\)\(0\.07\)\(0\.07\)\(0\.07\)Size smallb––0\.170\.17\*––\(0\.06\)\(0\.06\)Defaults×\\timesReas\. high−0\.63\-0\.63\*−0\.62\-0\.62\*−0\.62\-0\.62\*\(0\.13\)\(0\.13\)\(0\.13\)\(0\.13\)\(0\.13\)\(0\.13\)Soc\. Infl\.×\\timesReas\. high0\.590\.59\*0\.600\.60\*0\.600\.60\*\(0\.16\)\(0\.16\)\(0\.16\)\(0\.16\)\(0\.16\)\(0\.16\)\* 95% credible interval excludes zero\(1\)\(2\)\(3\)Std\. dev\. agent intercept1\.151\.151\.111\.111\.101\.10Observations \(agents\)21,600\(3,600\)21\{,\}600\\ \(3\{,\}600\)21,600\(3,600\)21\{,\}600\\ \(3\{,\}600\)21,600\(3,600\)21\{,\}600\\ \(3\{,\}600\)In odds\-ratio terms, within the default condition, high reasoning roughly halved the odds of selecting the preselected product \(OR=0\.51OR=0\.51,CI\.95CI\_\{\.95\}\[0\.41,0\.62\]\[0\.41,\\,0\.62\]; Table[84](https://arxiv.org/html/2609.19843#S8.T4)\)\. This is also visible in Figure[4](https://arxiv.org/html/2609.19843#S5.F4), where the predicted probability drops from approximately 92% for no\-reasoning agents to approximately 86% for high\-reasoning agents\. The results thus support H2a\. For H2b, we examined whether agents under a high\-reasoning configuration were more susceptible to reflective Type 2 nudges, that is, social influence\. The pooled model shows that the interaction between the reasoning configuration and the social influence condition is credibly positive \(β=0\.59\\beta=0\.59,CI\.95CI\_\{\.95\}\[0\.27,0\.91\]\[0\.27,\\,0\.91\]\) and, like the H2a interaction, robust to the provider and size controls \(β=0\.60\\beta=0\.60,CI\.95CI\_\{\.95\}\[0\.28,0\.92\]\[0\.28,\\,0\.92\]\), indicating that the social influence effect was on average stronger for high\-reasoning agents\. In odds\-ratio terms, within the social influence condition, high reasoning multiplied the odds of selecting the socially endorsed product by1\.711\.71\(CI\.95CI\_\{\.95\}\[1\.26,2\.20\]\[1\.26,\\,2\.20\]; Table[84](https://arxiv.org/html/2609.19843#S8.T4)\), corresponding to an increase in the predicted probability from approximately 96% to 98%\. The results thus support H2b\. #### 5\.2Post hoc Exploratory Analysis The pooled models estimate the reasoning moderation as an average across six models that differ in scale, architecture, training, and alignment\. Such an average can obscure systematic heterogeneity between model families and scales\([Wei et al\., 2022a](https://arxiv.org/html/2609.19843#bib.bib91);[McKenzie et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib88);[Perez et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib92)\)\. We therefore examine, in an exploratory analysis not anticipated by our hypotheses, whether the reasoning moderation is uniform across models or varies systematically with model scale and provider\. To this end, we re\-estimate Model 1 specification separately for each provider, each size class, and each individual model \(Appendices[9](https://arxiv.org/html/2609.19843#S9)–[11](https://arxiv.org/html/2609.19843#S11)\)\. At the individual\-model level, the two effects of reasoning appear in different models\. High reasoning weakens the default nudge mainly in the larger flagship models\. In Gemini 3\.5 Flash \(Table[113](https://arxiv.org/html/2609.19843#S11.T3)\), for instance, it lowers the share of agents that keep the preselected default from 84% to 48%, essentially chance \(β=−1\.61\\beta=\-1\.61,CI\.95CI\_\{\.95\}\[−2\.30,−0\.96\]\[\-2\.30,\\,\-0\.96\]\), whereas in its smaller counterpart Gemini 3\.1 Flash Lite \(Table[114](https://arxiv.org/html/2609.19843#S11.T4)\), the change is less\. High reasoning instead strengthens the social influence nudge mainly in the smaller models\. In GPT\-5\.4\-mini \(Table[112](https://arxiv.org/html/2609.19843#S11.T2)\), for instance, it raises the share that follow the socially endorsed option from 91% to 99% \(β=2\.46\\beta=2\.46,CI\.95CI\_\{\.95\}\[1\.61,3\.38\]\[1\.61,\\,3\.38\]\), whereas its larger counterpart GPT\-5\.4 \(Table[111](https://arxiv.org/html/2609.19843#S11.T1)\) shows no such increase\. Pooling the models by size class confirms this reasoning\-based split at the aggregate level \(Appendix[10](https://arxiv.org/html/2609.19843#S10)\) and surfaces a second, opposing pattern in overall susceptibility\. Independently of reasoning, the default nudge pulls the smaller variants more strongly than the flagships \(OR=15\.94OR=15\.94versusOR=4\.64OR=4\.64; Table[103](https://arxiv.org/html/2609.19843#S10.T3)\), whereas the social influence nudge pulls the flagships more strongly \(OR=107\.57OR=107\.57versusOR=17\.59OR=17\.59; Table[103](https://arxiv.org/html/2609.19843#S10.T3)\)\. Overall susceptibility therefore runs opposite to the reasoning moderation, with defaults dominating in the smaller models and social influence in the flagships\. The per\-provider fits are consistent with these patterns \(Appendix[9](https://arxiv.org/html/2609.19843#S9)\)\. A caveat concerns the two Claude models, which reach the 100% ceiling in entire conditions \(Claude Sonnet 4\.6 under social influence at both reasoning configurations, and Claude Haiku 4\.5 under defaults\)\. This \(quasi\-\)complete separation renders the affected point estimates partly prior\-dependent, so the corresponding conclusions rest on the data\-driven credible interval bounds rather than the point estimates\. The complete per\-model analyses are reported in Appendix[11](https://arxiv.org/html/2609.19843#S11)\. Taken together, these breakdowns point to a structure that tracks model scale\. Larger flagship models account for most of the reasoning\-driven weakening of the default nudge, while the smaller variants account for most of the reasoning\-driven strengthening of the social influence nudge\. The raw pull of each nudge, however, follows the reverse ordering, since defaults weigh more heavily on the small models and social cues on the flagships\. As these regularities were hypothesized, we read them as tentative and leave their confirmation to future work\. ### 6Discussion, Limitations and Outlook This study investigates how LLM\-based GUI agents behave when the digital environments they navigate are designed to influence their decisions\. Drawing on Dual\-Process Theory, we examined whether behavioural interventions originally developed for humans—automatic \(Type 1\) and reflective \(Type 2\) nudges—affect AI agents, and whether the reasoning configuration of an agent moderates its susceptibility\. Across six leading models from three providers, including both flagship and smaller variants, we provide broad empirical evidence and insights into how LLM\-based GUI agents succumb to digital nudges\. Our findings demonstrate that LLM\-based agents are highly susceptible to behavioural nudges, thereby challenging the assumption that they act as purely rational decision\-makers\([Chen et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib72)\)\. Interestingly, in our online shopping task, agents were even more responsive to these nudges than the human participants in the experiment on which our study was based\([Ingendahl et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib81)\)\. In sum, a preselected default increased the likelihood of selecting the nudged product roughly ninefold, while indicating that other shoppers had preferred it raised the odds more than thirtyfold, a susceptibility observed in every one of the six models\. Our findings therefore extend and add to previous evidence of human\-like cognitive biases in LLMs from text\-based settings\([Nguyen, 2024](https://arxiv.org/html/2609.19843#bib.bib25);[Hagendorff et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib86);[Cherep et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib77)\)to agentic environments, demonstrating that such biases persist when models operate as autonomous agents that perceive interfaces, plan sequences of actions, and execute decisions\. The effects of additional reasoning reveal a more complex relationship between deliberation and robustness\. Extended reasoning reduced susceptibility to the default nudge by approximately half, consistent with dual\-process accounts in which reflective processing can override automatic responses\([Stanovich, 2011](https://arxiv.org/html/2609.19843#bib.bib66)\)\. However, the same reasoning configuration increased susceptibility to the social influence nudge by approximately 70%\. Thus, deliberation did not eliminate behavioural bias in LLM agents\. Instead, it changed the type of influence to which agents were most responsive\. While reflective processing appears to reduce reliance on passive defaults, it may increase attention to socially relevant information that appears informative or normatively appropriate\. This opposing pattern is consistent with a dual\-process interpretation\. Because System 2\-style reasoning reduces the intuitive errors associated with automatic processing\([Hagendorff et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib86);[Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5)\), a high reasoning configuration may allow agents to override the automatic pull of a preselected option\. Reflective nudges, by contrast, take effect when their informational content is processed\. More extensive reasoning may therefore amplify attention to, and deepen engagement with the social cues, similar to how humans consciously weigh and incorporate perceived norms when forming judgments\([Cialdini and Trost, 1998](https://arxiv.org/html/2609.19843#bib.bib55);[Deutsch and Gerard, 1955](https://arxiv.org/html/2609.19843#bib.bib43)\)\. This finding also aligns with research showing that models trained with human feedback increasingly conform to stated views and social signals, a tendency documented as sycophancy\([Perez et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib92);[Sharma et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib89)\)and majority conformity\([Weng et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib93)\)\. Because reasoning\-oriented training also emphasizes instruction\-following and value alignment\([Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5);[Lynch et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib38)\), agents configured for more deliberation may be especially responsive to persuasive informational cues\. Our exploratory analysis further suggests that these effects may vary by model scale\. Reasoning\-based resistance to defaults was primarily observed in flagship models, potentially because overriding a default requires generating and acting upon an alternative preference—a capability that may improve with greater model capacity\([Wei et al\., 2022a](https://arxiv.org/html/2609.19843#bib.bib91)\)\. Conversely, increased susceptibility to social influence was more pronounced among smaller variants\. One plausible explanation is that the flagship models were already highly compliant with the social cue under the no\-reasoning configuration, leaving little room for further increases, whereas the smaller variants retained greater scope for deeper processing of the cue to increase compliance\. We offer this interpretation cautiously, as our study was not designed to test this mechanism directly\. Accordingly, we treat these findings as exploratory and as a basis for future confirmatory research\. Our findings also contribute to understanding the relationship between agentic and human decision\-making\. In the human study adapted here,[Ingendahl et al\. \(2021\)](https://arxiv.org/html/2609.19843#bib.bib81)found both nudges to be effective\. However, need for cognition—a tendency to engage in effortful, reflective thinking\([Cacioppo and Petty, 1982](https://arxiv.org/html/2609.19843#bib.bib108)\)—only weakly and inconsistently moderated these effects, with less reflective participants showing slightly greater susceptibility to the default nudge\. While the role of cognitive engagement in human nudgeability more generally remains debated\([De Ridder et al\., 2022](https://arxiv.org/html/2609.19843#bib.bib49);[Van Gestel et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib48)\), directly manipulating the reasoning configuration in agents produced a stronger and more systematic moderation effect\. Agents under high reasoning were markedly less susceptible to the automatic default nudge, in the same direction as the human default tendency, while becoming more susceptible to the reflective social influence nudge\. We draw this comparison cautiously, since switching an agent’s reasoning on or off is not the same as comparing people who naturally differ in a trait such as need for cognition\. It nevertheless suggests that agent\-based simulations can offer a complementary setting for examining how processing mode shapes susceptibility to nudges\. Notably, these effects emerged in the absence of the cognitive constraints traditionally invoked to explain human bias\([Lieder and Griffiths, 2020](https://arxiv.org/html/2609.19843#bib.bib50)\), suggesting that agentic susceptibility may arise from mechanisms such as training data, learned heuristics, and alignment objectives\([Jones and Steinhardt, 2022](https://arxiv.org/html/2609.19843#bib.bib87);[Macmillan\-Scott and Musolesi, 2024](https://arxiv.org/html/2609.19843#bib.bib51)\)\. Susceptibility to reflective nudges thus appears to be not merely a consequence of limited cognition but a byproduct of how both humans and agents use social information as a shortcut for what is appropriate or accurate\([Caraban et al\., 2019](https://arxiv.org/html/2609.19843#bib.bib59)\)\. From a theoretical perspective, our results establish nudge susceptibility as a behavioural property of agentic AI systems rather than a phenomenon unique to humans\([Baird and Maruping, 2021](https://arxiv.org/html/2609.19843#bib.bib30);[Bösch, 2025](https://arxiv.org/html/2609.19843#bib.bib21);[Cherep et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib77)\)\. Because the reflective\-moderates\-automatic pattern appears in agents that lack the cognitive architecture and resource limits Dual\-Process Theory assumes\([Kahneman, 2003](https://arxiv.org/html/2609.19843#bib.bib24);[Evans and Stanovich, 2013](https://arxiv.org/html/2609.19843#bib.bib63)\), the dual\-process regularity is separable from its human substrate and may characterise any instruction\-following, preference\-aligned decision system—extending the theory beyond human cognition rather than merely applying it\. That the pattern recurs across model families differing in architecture, training data, and alignment points to a general property of current LLM\-based agents rather than an artifact of any single model\. The manipulation also reframes deliberation from debiasing to redirection: unlike the human literature, where cognitive engagement is a contested trait moderator\([Van Gestel et al\., 2021](https://arxiv.org/html/2609.19843#bib.bib48);[De Ridder et al\., 2022](https://arxiv.org/html/2609.19843#bib.bib49)\), our experimental control isolates an opposing\-direction effect that qualifies the expectation that more reasoning yields more robust decisions\([Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5);[McKenzie et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib88)\)\. The scale\-dependence we observe further suggests this moderation is itself capacity\-bounded, a condition with no direct analogue in the human formulation\. Taken together, these points underscore the value of evaluating agentic systems behaviourally, looking beyond task competence toward how context and environment shape their decisions and can misalign them from their intended behaviour\([Lynch et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib38)\)\. From a practical perspective, our results position interface design as a governance concern rather than a matter of usability alone\. The nudges studied here are ordinary features of digital choice architectures\([Schneider et al\., 2018](https://arxiv.org/html/2609.19843#bib.bib58);[Caraban et al\., 2019](https://arxiv.org/html/2609.19843#bib.bib59)\), not adversarial attacks such as prompt injection\([Greshake et al\., 2023](https://arxiv.org/html/2609.19843#bib.bib100)\)\. Agents are therefore routinely exposed to them, and because digital choice architectures shape agent decisions much as they shape human ones, the consequences of these design choices ultimately accrue to the delegating user\. Deploying agents responsibly thus requires behavioural auditing, evaluation across diverse interface conditions, and safeguards against unintended influence\. These considerations may become even more important as AI agents grow increasingly capable\. Although advanced agents require less step\-by\-step prompting and can operate more independently,333[https://www\.anthropic\.com/engineering/building\-effective\-agents](https://www.anthropic.com/engineering/building-effective-agents)they may remain susceptible to subtle contextual cues—a vulnerability that could persist even as other forms of bias diminish, precisely because agents are trained to follow instructions and align with human preferences\. Cues that appear consistent with those preferences may therefore steer behaviour towards unintended outcomes\([Zhang et al\., 2025](https://arxiv.org/html/2609.19843#bib.bib5)\)\. These risks call for safeguards beyond conventional alignment techniques, distributed across the actors who build, deploy, and host agents\. Developers could report standardized nudge\-susceptibility benchmarks that score a model with specific reasoning configuration and scale against common interface manipulation\. Organizations that deploy agents could treat the reasoning configuration as a deployment decision rather than a default, choosing it for the specific choice environment while logging the agent’s choice distribution to detect behavioural drift when an interface changes, and requiring a verification pass that re\-checks consequential actions against the task’s stated criteria before they are executed\. Interface and platform designers, in turn, could support machine\-readable disclosure of choice\-architecture elements, such as tagging a preselected default or a social\-proof badge, so that a delegated agent can recognise and discount a persuasive cue it would otherwise treat as neutral information\. More broadly, our findings point to a capability that cuts across these roles and remains difficult to achieve\. Agents would need to tell information relevant to the task apart from persuasive cues that are not, and to give less weight to the latter\. Our results suggest this is unlikely to come from more reasoning alone, which for reflective nudges can even increase susceptibility\. Together, such measures could improve the robustness of agentic systems to routine interface manipulations without compromising their ability to interact effectively with digital environments, and they motivate future work on enabling agents to recognise, interpret, and appropriately respond to contextual cues without being unduly influenced by them\. As with any study, this work must be considered in light of some limitations\. First, we examined only default and social influence nudges as representative mechanisms of Type 1 and Type 2 influence\. Whether the same pattern extends to other automatic nudges, such as deceptive visualizations or scarcity cues, and to other reflective nudges, such as suggesting alternatives or reminding of consequences, remains open\. Second, we manipulated reasoning through the model configuration exposed by the providers’ APIs\. Future work should examine complementary strategies, such as prompt\-level instructions that encourage deliberation about product attributes or trade\-offs, to distinguish task\-specific reasoning from the general reasoning configuration, and should test how agents respond when several nudges are present at once\. Third, the scale\-dependent structure of the moderation effects was not hypothesized and is reported as exploratory\. Confirming it will require designs with more models per scale class, so that model size can be treated as a planned factor rather than inferred post hoc\. Fourth, our findings characterise the specific model versions used in this study, whose exact configurations are documented in our public repository\. Because providers update weights, alignment, and API behaviour under stable model names, the absolute susceptibility levels—and potentially the reasoning moderation—may differ for later versions, even where the model name is unchanged\. The consistency of the central pattern across six models from three providers nonetheless makes it unlikely to be an artifact of any single version\. Finally, regarding external validity, our experiment adapted an established human decision\-making experiment within a simulated online store\. While this controlled setting offers experimental rigour, future research should study LLM\-based GUI agents in richer and more ecologically valid environments to determine whether the observed effects persist beyond laboratory conditions\. Deploying such agents safely will depend less on making them more capable than on scrutinising the environments in which they are asked to decide\. ### 7Conclusion With growing delegation of tasks to autonomous AI agents across domains such as e\-commerce and finance, understanding how they behave in environments designed to influence them becomes increasingly important\. In this paper, we examined one such form of external influence, asking how digital nudges shape the choice behaviour of LLM\-based GUI agents\. By adapting a controlled human decision\-making experiment to six leading models from three providers, we showed that agents are susceptible to both automatic and reflective nudges, and that reasoning does not immunise them but changes which nudges take hold—weakening automatic defaults while strengthening reflective social cues\. An exploratory analysis indicated that this pattern is systematically structured by model size\. ## References - A\. Acerbi and J\. M\. StubbersfieldLarge language models show human\-like content biases in transmission chain experiments\.Proceedings of the National Academy of Sciences120\(44\),pp\. e2313790120\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1)\. - Allcott \(2011\)H\. AllcottSocial norms and energy conservation\.Journal of Public Economics95\(9\-10\),pp\. 1082–1095\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p4.1)\. - Anthropic \(2026\)AnthropicIntroducing claude sonnet 4\.6\.External Links:[Link](https://www.anthropic.com/news/claude-sonnet-4-6)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§1](https://arxiv.org/html/2609.19843#S1.p9.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p6.1),[§4\.5](https://arxiv.org/html/2609.19843#S4.SS5.p1.1)\. - Baird and Maruping \(2021\)A\. Baird and L\. M\. MarupingThe next generation of research on is use: a theoretical framework of delegation to and from agentic is artifacts\.MIS quarterly45\(1\),pp\. 315–341\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1),[§1](https://arxiv.org/html/2609.19843#S1.p11.1),[§1](https://arxiv.org/html/2609.19843#S1.p9.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1)\. - BaniHaniet al\.\(2024\)I\. BaniHani, S\. Alawadi, and N\. ElmrayyanAI and the decision\-making process: a literature review in healthcare, financial, and technology sectors\.Journal of Decision Systems33\(sup1\),pp\. 389–399\(en\)\.External Links:ISSN 1246\-0125, 2116\-7052Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p2.1)\. - Bateset al\.\(2015\)D\. Bates, M\. Maechler, B\. Bolker, and S\. WalkerFitting Linear Mixed\-Effects Models Usinglme4\.Journal of Statistical Software67\(1\) \(en\)\.External Links:ISSN 1548\-7660,[Document](https://dx.doi.org/10.18637/jss.v067.i01)Cited by:[§4\.6](https://arxiv.org/html/2609.19843#S4.SS6.p3.1)\. - Bergramet al\.\(2022\)K\. Bergram, M\. Djokovic, V\. Bezençon, and A\. HolzerThe Digital Landscape of Nudging: A Systematic Literature Review of Empirical Research on Digital Nudges\.InCHI Conference on Human Factors in Computing Systems,New Orleans LA USA,pp\. 1–16\(en\)\.External Links:ISBN 978\-1\-4503\-9157\-3,[Document](https://dx.doi.org/10.1145/3491102.3517638)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p3.1)\. - Binz and Schulz \(2023\)M\. Binz and E\. SchulzUsing cognitive psychology to understand gpt\-3\.Proceedings of the National Academy of Sciences120\(6\),pp\. e2218523120\.Cited by:[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p4.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p4.3)\. - Boochet al\.\(2021\)G\. Booch, F\. Fabiano, L\. Horesh, K\. Kate, J\. Lenchner, N\. Linck, A\. Loreggia, K\. Murgesan, N\. Mattei, F\. Rossi,et al\.Thinking fast and slow in ai\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.35,pp\. 15042–15046\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p8.1)\. - Bösch \(2025\)K\. BöschBiased decisions of large language models \(llm\): computer control agents and the decoy effect\.Available at SSRN 5727082\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p11.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§3](https://arxiv.org/html/2609.19843#S3.p2.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1)\. - Bradyet al\.\(2025\)O\. Brady, P\. Nulty, L\. Zhang, T\. E\. Ward, and D\. P\. McGovernDual\-process theory and decision\-making in large language models\.Nature Reviews Psychology,pp\. 1–16\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p11.1),[§3](https://arxiv.org/html/2609.19843#S3.p3.1),[§3](https://arxiv.org/html/2609.19843#S3.p8.1)\. - Browser Use \(2025\)Browser UseCelebrating One Year of Progress in Browser Agents\.External Links:[Link](https://browser-use.com/posts/one-year-of-progress)Cited by:[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p2.1),[§4\.2](https://arxiv.org/html/2609.19843#S4.SS2.p1.1)\. - Bürkner \(2017\)P\. BürknerBrms: an r package for bayesian multilevel models using stan\.Journal of Statistical Software80\(1\),pp\. 1–28\.External Links:[Link](https://www.jstatsoft.org/index.php/jss/article/view/v080i01),[Document](https://dx.doi.org/10.18637/jss.v080.i01)Cited by:[§4\.6](https://arxiv.org/html/2609.19843#S4.SS6.p1.1)\. - Cacioppo and Petty \(1982\)J\. T\. Cacioppo and R\. E\. PettyThe need for cognition\.\.Journal of Personality and Social Psychology42\(1\),pp\. 116\.Cited by:[§6](https://arxiv.org/html/2609.19843#S6.p6.1)\. - Candrian and Scherer \(2022\)C\. Candrian and A\. SchererRise of the machines: delegating decisions to autonomous ai\.Computers in Human Behavior134,pp\. 107308\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p2.1)\. - Carabanet al\.\(2019\)A\. Caraban, E\. Karapanos, D\. Gonçalves, and P\. Campos23 Ways to Nudge: A Review of Technology\-Mediated Nudging in Human\-Computer Interaction\.InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems,Glasgow Scotland Uk,pp\. 1–15\(en\)\.External Links:ISBN 978\-1\-4503\-5970\-2,[Document](https://dx.doi.org/10.1145/3290605.3300733)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p5.1),[§3](https://arxiv.org/html/2609.19843#S3.p8.1),[§4](https://arxiv.org/html/2609.19843#S4.p1.1),[§6](https://arxiv.org/html/2609.19843#S6.p6.1),[§6](https://arxiv.org/html/2609.19843#S6.p8.1)\. - Chenet al\.\(2025\)Y\. Chen, S\. N\. Kirshner, A\. Ovchinnikov, M\. Andiappan, and T\. JenkinA manager and an ai walk into a bar: does chatgpt make biased decisions like we do?\.Manufacturing & Service Operations Management27\(2\),pp\. 354–368\.Cited by:[§3](https://arxiv.org/html/2609.19843#S3.p1.1)\. - Chenet al\.\(2023\)Y\. Chen, T\. X\. Liu, Y\. Shan, and S\. ZhongThe emergence of economic rationality of gpt\.Proceedings of the National Academy of Sciences120\(51\),pp\. e2316205120\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p11.1),[§1](https://arxiv.org/html/2609.19843#S1.p4.1),[§6](https://arxiv.org/html/2609.19843#S6.p2.1)\. - Cherepet al\.\(2024\)M\. Cherep, N\. Singh, and P\. MaesSuperficial alignment, subtle divergence, and nudge sensitivity in llm decision\-making\.InNeurIPS 2024 Workshop on Behavioral Machine Learning,Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p11.1),[§1](https://arxiv.org/html/2609.19843#S1.p4.1),[§3](https://arxiv.org/html/2609.19843#S3.p2.1),[§6](https://arxiv.org/html/2609.19843#S6.p2.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1)\. - Cialdiniet al\.\(1990\)R\. B\. Cialdini, R\. R\. Reno, and C\. A\. KallgrenA focus theory of normative conduct: recycling the concept of norms to reduce littering in public places\.\.Journal of personality and social psychology58\(6\),pp\. 1015\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p4.1)\. - Cialdini and Trost \(1998\)R\. B\. Cialdini and M\. R\. TrostSocial influence: Social norms, conformity and compliance\.McGraw\-Hill\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p4.1),[§6](https://arxiv.org/html/2609.19843#S6.p4.1)\. - Cochran \(1954\)W\. G\. CochranSome methods for strengthening the commonχ2\\chi^\{2\}tests\.Biometrics10\(4\),pp\. 417\.External Links:ISSN 0006\-341X,[Link](http://dx.doi.org/10.2307/3001616),[Document](https://dx.doi.org/10.2307/3001616)Cited by:[§12](https://arxiv.org/html/2609.19843#S12.SSx2.p2.1)\. - De Ridderet al\.\(2022\)D\. De Ridder, F\. Kroese, and L\. Van GestelNudgeability: Mapping Conditions of Susceptibility to Nudge Influence\.Perspectives on Psychological Science17\(2\),pp\. 346–359\(en\)\.External Links:ISSN 1745\-6916, 1745\-6924,[Document](https://dx.doi.org/10.1177/1745691621995183)Cited by:[§3](https://arxiv.org/html/2609.19843#S3.p5.1),[§6](https://arxiv.org/html/2609.19843#S6.p6.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1)\. - Debenedettiet al\.\(2024\)E\. Debenedetti, J\. Zhang, M\. Balunovic, L\. Beurer\-Kellner, M\. Fischer, and F\. TramèrAgentdojo: a dynamic environment to evaluate prompt injection attacks and defenses for llm agents\.Advances in Neural Information Processing Systems37,pp\. 82895–82920\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p5.1)\. - Demarqueet al\.\(2015\)C\. Demarque, L\. Charalambides, D\. J\. Hilton, and L\. WaroquierNudging sustainable consumption: The use of descriptive norms to promote a minority behavior in a realistic online shopping environment\.Journal of Environmental Psychology43,pp\. 166–174\(en\)\.External Links:ISSN 02724944,[Document](https://dx.doi.org/10.1016/j.jenvp.2015.06.008)Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p4.1),[§4](https://arxiv.org/html/2609.19843#S4.p1.1)\. - Dentellaet al\.\(2026\)V\. Dentella, M\. Marelli, and L\. RinaldiLLLMs displaying less cognitive bias are not necessarily better decision makers\.Nature Machince Intelligence\.Cited by:[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1)\. - Deutsch and Gerard \(1955\)M\. Deutsch and H\. B\. GerardA study of normative and informational social influences upon individual judgment\.The Journal of Abnormal and Social Psychology51\(3\),pp\. 629\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p4.1),[§6](https://arxiv.org/html/2609.19843#S6.p4.1)\. - Echterhoffet al\.\(2024\)J\. M\. Echterhoff, Y\. Liu, A\. Alessa, J\. McAuley, and Z\. HeCognitive bias in decision\-making with llms\.InFindings of the association for computational linguistics: EMNLP 2024,pp\. 12640–12653\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§3](https://arxiv.org/html/2609.19843#S3.p1.1)\. - Evans and Stanovich \(2013\)J\. St\. B\. T\. Evans and K\. E\. StanovichDual\-Process Theories of Higher Cognition: Advancing the Debate\.Perspectives on Psychological Science8\(3\),pp\. 223–241\(en\)\.External Links:ISSN 1745\-6916, 1745\-6924,[Document](https://dx.doi.org/10.1177/1745691612460685)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p3.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p2.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1)\. - Feuerriegelet al\.\(2024\)S\. Feuerriegel, J\. Hartmann, C\. Janiesch, and P\. ZschechGenerative AI\.Business & Information Systems Engineering66\(1\),pp\. 111–126\(en\)\.External Links:ISSN 2363\-7005, 1867\-0202,[Document](https://dx.doi.org/10.1007/s12599-023-00834-7)Cited by:[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p1.1)\. - Gelmanet al\.\(2008\)A\. Gelman, A\. Jakulin, M\. G\. Pittau, and Y\. SuA weakly informative default prior distribution for logistic and other regression models\.The Annals of Applied Statistics2\(4\),pp\. 1360\.Cited by:[§4\.6](https://arxiv.org/html/2609.19843#S4.SS6.p3.1)\. - Göldi and Rietsche \(2025\)A\. Göldi and R\. RietscheMaking Sense of Large Language Model\-Based AI Agents\.InICIS 2024 Proceedings,Bangkok \- Thailand\.Cited by:[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p1.1)\. - Goldsteinet al\.\(2008\)N\. J\. Goldstein, R\. B\. Cialdini, and V\. GriskeviciusA room with a viewpoint: using social norms to motivate environmental conservation in hotels\.Journal of Consumer Research35\(3\),pp\. 472–482\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p4.1),[§3](https://arxiv.org/html/2609.19843#S3.p3.1)\. - Google \(2025\)GoogleBest for frontier intelligence at speed\.External Links:[Link](https://deepmind.google/models/gemini/flash/)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§1](https://arxiv.org/html/2609.19843#S1.p9.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p6.1),[§4\.5](https://arxiv.org/html/2609.19843#S4.SS5.p1.1)\. - Greshakeet al\.\(2023\)K\. Greshake, S\. Abdelnabi, S\. Mishra, C\. Endres, T\. Holz, and M\. FritzNot what you’ve signed up for: compromising real\-world llm\-integrated applications with indirect prompt injection\.InProceedings of the 16th ACM Workshop on Artificial Intelligence and Security,pp\. 79–90\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p5.1),[§6](https://arxiv.org/html/2609.19843#S6.p8.1)\. - Guoet al\.\(2025\)Z\. Guo, H\. Lv, C\. Zhang, Y\. Zhao, Y\. Zhang, and L\. CuiThe illusion of randomness: how llms fail to emulate stochastic decision\-making in rock\-paper\-scissors games?\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 8618–8637\.Cited by:[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1),[§3](https://arxiv.org/html/2609.19843#S3.p1.1)\. - Hagendorffet al\.\(2023\)T\. Hagendorff, S\. Fabi, and M\. KosinskiHuman\-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt\.Nature Computational Science3\(10\),pp\. 833–838\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p4.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p4.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p4.3),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p5.1),[§3](https://arxiv.org/html/2609.19843#S3.p1.1),[§6](https://arxiv.org/html/2609.19843#S6.p2.1),[§6](https://arxiv.org/html/2609.19843#S6.p4.1)\. - Hakiet al\.\(2023\)K\. Haki, A\. Rieder, L\. Buchmann, and A\. W\. SchneiderDigital nudging for technical debt management at credit suisse\.European Journal of Information Systems32\(1\),pp\. 64–80\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p3.1)\. - Hansen and Jespersen \(2013\)P\. G\. Hansen and A\. M\. JespersenNudge and the Manipulation of Choice: A Framework for the Responsible Use of the Nudge Approach to Behaviour Change in Public Policy\.European Journal of Risk Regulation4\(1\),pp\. 3–28\(en\)\.External Links:ISSN 1867\-299X, 2190\-8249,[Document](https://dx.doi.org/10.1017/S1867299X00002762)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p2.1)\. - Huet al\.\(2025\)X\. Hu, T\. Xiong, B\. Yi, Z\. Wei, R\. Xiao, Y\. Chen, J\. Ye, M\. Tao, X\. Zhou, Z\. Zhao,et al\.Os agents: a survey on mllm\-based agents for computer, phone and browser use\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7436–7465\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p2.1)\. - Huang and Chang \(2023\)J\. Huang and K\. C\. ChangTowards Reasoning in Large Language Models: A Survey\.InFindings of the Association for Computational Linguistics: ACL 2023,Toronto, Canada,pp\. 1049–1065\(en\)\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.67)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p2.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1)\. - Hummel and Maedche \(2019\)D\. Hummel and A\. MaedcheHow effective is nudging? A quantitative review on the effect sizes and limits of empirical nudging studies\.Journal of Behavioral and Experimental Economics80,pp\. 47–58\(en\)\.External Links:ISSN 22148043,[Document](https://dx.doi.org/10.1016/j.socec.2019.03.005)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1),[§4](https://arxiv.org/html/2609.19843#S4.p1.1)\. - Ingendahlet al\.\(2021\)M\. Ingendahl, D\. Hummel, A\. Maedche, and T\. VogelWho can be nudged? Examining nudging effectiveness in the context of need for cognition and need for uniqueness\.Journal of Consumer Behaviour20\(2\),pp\. 324–336\(en\)\.External Links:ISSN 1479\-1838,[Document](https://dx.doi.org/10.1002/cb.1861)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p9.1),[§3](https://arxiv.org/html/2609.19843#S3.p5.1),[§4\.1](https://arxiv.org/html/2609.19843#S4.SS1.p1.1),[§4](https://arxiv.org/html/2609.19843#S4.p1.1),[§4](https://arxiv.org/html/2609.19843#S4.p2.1),[§6](https://arxiv.org/html/2609.19843#S6.p2.1),[§6](https://arxiv.org/html/2609.19843#S6.p6.1)\. - Itzhaket al\.\(2024\)I\. Itzhak, G\. Stanovsky, N\. Rosenfeld, and Y\. BelinkovInstructed to bias: instruction\-tuned language models exhibit emergent cognitive bias\.Transactions of the Association for Computational Linguistics12,pp\. 771–785\.Cited by:[§3](https://arxiv.org/html/2609.19843#S3.p1.1)\. - Johnson and Goldstein \(2003\)E\. J\. Johnson and D\. GoldsteinDo defaults save lives?\.Vol\.302,American Association for the Advancement of Science\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p3.1),[§3](https://arxiv.org/html/2609.19843#S3.p3.1)\. - Jones and Steinhardt \(2022\)E\. Jones and J\. SteinhardtCapturing failures of large language models via human cognitive biases\.Advances in Neural Information Processing Systems35,pp\. 11785–11799\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p4.1),[§3](https://arxiv.org/html/2609.19843#S3.p1.1),[§6](https://arxiv.org/html/2609.19843#S6.p6.1)\. - Kahneman \(2003\)D\. KahnemanMaps of bounded rationality: psychology for behavioral economics\.American Economic Review93\(5\),pp\. 1449–1475\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p11.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p3.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p6.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p2.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1)\. - Kahneman \(2011\)D\. KahnemanThinking, fast and slow\.Farrar, Straus and Giroux\.Cited by:[§12](https://arxiv.org/html/2609.19843#S12.p1.1),[§4](https://arxiv.org/html/2609.19843#S4.p3.1)\. - Kooet al\.\(2024\)R\. Koo, M\. Lee, V\. Raheja, J\. I\. Park, Z\. M\. Kim, and D\. KangBenchmarking cognitive biases in large language models as evaluators\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 517–545\.Cited by:[§3](https://arxiv.org/html/2609.19843#S3.p1.1)\. - Liet al\.\(2023\)H\. Li, J\. Su, Y\. Chen, Q\. Li, and Z\. ZhangSheetCopilot: bringing software productivity to the next level through large language models\.InProceedings of the 37th International Conference on Neural Information Processing Systems,pp\. 4952–4984\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1)\. - Lieder and Griffiths \(2020\)F\. Lieder and T\. L\. GriffithsResource\-rational analysis: Understanding human cognition as the optimal use of limited computational resources\.Behavioral and Brain Sciences43,pp\. e1\(en\)\.External Links:ISSN 0140\-525X, 1469\-1825,[Document](https://dx.doi.org/10.1017/S0140525X1900061X)Cited by:[§6](https://arxiv.org/html/2609.19843#S6.p6.1)\. - Liuet al\.\(2023\)P\. Liu, W\. Yuan, J\. Fu, Z\. Jiang, H\. Hayashi, and G\. NeubigPre\-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing\.ACM Computing Surveys55\(9\),pp\. 1–35\(en\)\.External Links:ISSN 0360\-0300, 1557\-7341,[Document](https://dx.doi.org/10.1145/3560815)Cited by:[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p1.1)\. - Liuet al\.\(2025\)R\. Liu, J\. Geng, A\. J\. Wu, I\. Sucholutsky, T\. Lombrozo, and T\. L\. GriffithsMind your step \(by step\): chain\-of\-thought can reduce performance on tasks where thinking makes humans worse\.InInternational Conference on Machine Learning,pp\. 38489–38517\.Cited by:[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1)\. - Lynchet al\.\(2025\)A\. Lynch, B\. Wright, C\. Larson, S\. J\. Ritchie, S\. Mindermann, E\. Hubinger, E\. Perez, and K\. TroyAgentic Misalignment: How LLMs Could Be Insider Threats\.arXiv\.Note:Version Number: 2External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2510.05179)Cited by:[§6](https://arxiv.org/html/2609.19843#S6.p4.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1)\. - Macmillan\-Scott and Musolesi \(2024\)O\. Macmillan\-Scott and M\. Musolesi\(Ir\) rationality and cognitive biases in large language models\.Royal Society open science11\(6\)\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1),[§3](https://arxiv.org/html/2609.19843#S3.p1.1),[§6](https://arxiv.org/html/2609.19843#S6.p6.1)\. - Mantel and Haenszel \(1959\)N\. Mantel and W\. HaenszelStatistical aspects of the analysis of data from retrospective studies of disease\.JNCI: Journal of the National Cancer Institute22\(4\),pp\. 719–748\.External Links:ISSN 0027\-8874,[Document](https://dx.doi.org/10.1093/jnci/22.4.719),[Link](https://doi.org/10.1093/jnci/22.4.719),https://academic\.oup\.com/jnci/article\-pdf/22/4/719/2674652/22\-4\-719\.pdfCited by:[§12](https://arxiv.org/html/2609.19843#S12.SSx2.p2.1)\. - Marreedet al\.\(2025\)S\. Marreed, A\. Oved, A\. Yaeli, S\. Shlomov, I\. Levy, O\. Akrabi, A\. Sela, A\. Adi, and N\. MashkifTowards Enterprise\-Ready Computer Using Generalist Agent\.arXiv\.Note:Version Number: 3External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2503.01861)Cited by:[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p2.1)\. - McKenzieet al\.\(2023\)I\. R\. McKenzie, A\. Lyzhov, M\. Pieler, A\. Parrish, A\. Mueller, A\. Prabhu, E\. McLean, A\. Kirtland, A\. Ross, A\. Liu,et al\.Inverse scaling: when bigger isn’t better\.arXiv preprint arXiv:2306\.09479\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1),[§5\.2](https://arxiv.org/html/2609.19843#S5.SS2.p1.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1)\. - Mirschet al\.\(2017\)T\. Mirsch, C\. Lehrer, and R\. JungDigital nudging: altering user behavior in digital environments\.Proceedings der 13\. Internationalen Tagung Wirtschaftsinformatik \(WI 2017\),pp\. 634–648\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1)\. - Müller \(2026\)N\. MüllerJeden siebten online\-einkauf erledigt 2030 der ki\-agent\.Frankfurter Allgemeine Zeitung\.External Links:[Link](https://www.faz.net/premium/digitalwirtschaft/kuenstliche-intelligenz/online-shopping-ki-agenten-erledigen-2030-jeden-siebten-einkauf-accg-200592846.html)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p2.1)\. - Murrayet al\.\(2021\)A\. Murray, J\. Rhymer, and D\. G\. SirmonHumans and Technology: Forms of Conjoined Agency in Organizations\.Academy of Management Review46\(3\),pp\. 552–571\(en\)\.External Links:ISSN 0363\-7425, 1930\-3807,[Document](https://dx.doi.org/10.5465/amr.2019.0186)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1)\. - Murugesan \(2025\)S\. MurugesanThe rise of agentic AI: implications, concerns, and the path forward\.IEEE Intell\. Syst\.40\(2\),pp\. 8–14\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1)\. - Nguyenet al\.\(2025\)D\. Nguyen, J\. Chen, Y\. Wang, G\. Wu, N\. Park, Z\. Hu, H\. Lyu, J\. Wu, R\. Aponte, Y\. Xia,et al\.Gui agents: a survey\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 22522–22538\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p2.1)\. - Nguyen \(2024\)J\. K\. NguyenHuman bias in ai models? anchoring effects and mitigation strategies in large language models\.Journal of Behavioral and Experimental Finance43,pp\. 100971\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p11.1),[§1](https://arxiv.org/html/2609.19843#S1.p4.1),[§3](https://arxiv.org/html/2609.19843#S3.p2.1),[§6](https://arxiv.org/html/2609.19843#S6.p2.1)\. - OpenAI \(2025a\)OpenAIChatGPT Agent System Card\.Technical reportOpenAI\.External Links:[Link](https://cdn.openai.com/pdf/839e66fc-602c-48bf-81d3-b21eacc3459d/chatgpt_agent_system_card.pdf)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p6.1)\. - OpenAI \(2025b\)OpenAIIntroducing ChatGPT agent: bridging research and action\.External Links:[Link](https://openai.com/index/introducing-chatgpt-agent/)Cited by:[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p2.1)\. - OpenAI \(2025c\)OpenAIIntroducing GPT\-5\.External Links:[Link](https://openai.com/index/introducing-gpt-5/)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p9.1),[§4\.5](https://arxiv.org/html/2609.19843#S4.SS5.p1.1)\. - Perezet al\.\(2023\)E\. Perez, S\. Ringer, K\. Lukosiute, K\. Nguyen, E\. Chen, S\. Heiner, C\. Pettit, C\. Olsson, S\. Kundu, S\. Kadavath,et al\.Discovering language model behaviors with model\-written evaluations\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 13387–13434\.Cited by:[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1),[§5\.2](https://arxiv.org/html/2609.19843#S5.SS2.p1.1),[§6](https://arxiv.org/html/2609.19843#S6.p4.1)\. - Perplexity \(2025a\)PerplexityComet Browser\.External Links:[Link](https://www.perplexity.ai/comet)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1),[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p2.1)\. - Perplexity \(2025b\)PerplexityThe Internet is Better on Comet\.External Links:[Link](https://www.perplexity.ai/de/hub/blog/comet-is-now-available-to-everyone-worldwide)Cited by:[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p2.1)\. - Quet al\.\(2025\)C\. Qu, S\. Dai, X\. Wei, H\. Cai, S\. Wang, D\. Yin, J\. Xu, and J\. WenTool learning with large language models: a survey\.Frontiers of Computer Science19\(8\),pp\. 198343\(en\)\.External Links:ISSN 2095\-2228, 2095\-2236,[Document](https://dx.doi.org/10.1007/s11704-024-40678-2)Cited by:[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p1.1)\. - Rayner \(2020\)J\. C\. W\. RaynerCochran–mantel–haenszel tests for the completely randomised design\.Journal of the Korean Statistical Society50\(1\),pp\. 185–197\.External Links:ISSN 2005\-2863,[Link](http://dx.doi.org/10.1007/s42952-020-00068-3),[Document](https://dx.doi.org/10.1007/s42952-020-00068-3)Cited by:[§12](https://arxiv.org/html/2609.19843#S12.SSx2.p2.1)\. - Rheu and Cho \(2025\)M\. Rheu and J\. ChoThe trap of ai literacy: the paradoxical relationships between college students’ use of llms, ai literacy, and fact\-checking behavior\.InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems,pp\. 1–7\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1)\. - Ritov and Baron \(1992\)I\. Ritov and J\. BaronStatus\-quo and omission biases\.Journal of Risk and Uncertainty5\(1\) \(en\)\.External Links:ISSN 0895\-5646, 1573\-0476,[Document](https://dx.doi.org/10.1007/BF00208786)Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p3.1)\. - Sageret al\.\(2025\)P\. J\. Sager, B\. Meyer, P\. Yan, R\. von Wartburg\-Kottler, L\. Etaiwi, A\. Enayati, G\. Nobel, A\. Abdulkadir, B\. F\. Grewe, and T\. StadelmannA Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions\.arXiv\.Note:Version Number: 2External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2501.16150)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p1.1),[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p2.1)\. - Samuelson and Zeckhauser \(1988\)W\. Samuelson and R\. ZeckhauserStatus quo bias in decision making\.Journal of Risk and Uncertainty1\(1\),pp\. 7–59\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p3.1)\. - Schicket al\.\(2023\)T\. Schick, J\. Dwivedi\-Yu, R\. Dessi, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. ScialomToolformer: Language Models Can Teach Themselves to Use Tools\.InThirty\-seventh Conference on Neural Information Processing Systems,Cited by:[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p1.1)\. - Schmidtet al\.\(2026\)R\. Schmidt, R\. Alt, and A\. ZimmermannAgentic ai readiness: a process\-oriented assessment framework\.Proceedings of the 59th Hawaii International Conference on System Sciences \(HICSS\)\.External Links:[Link](https://scholarspace.manoa.hawaii.edu/items/174fe069-9545-4445-96ef-9cf693bd87ea)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p2.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1)\. - Schneideret al\.\(2018\)C\. Schneider, M\. Weinmann, and J\. Vom BrockeDigital nudging: guiding online user choices through interface design\.Communications of the ACM61\(7\),pp\. 67–73\(en\)\.External Links:ISSN 0001\-0782, 1557\-7317,[Document](https://dx.doi.org/10.1145/3213765)Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p5.1),[§6](https://arxiv.org/html/2609.19843#S6.p8.1)\. - Shanahanet al\.\(2023\)M\. Shanahan, K\. McDonell, and L\. ReynoldsRole play with large language models\.Nature623\(7987\),pp\. 493–498\(en\)\.External Links:ISSN 0028\-0836, 1476\-4687,[Document](https://dx.doi.org/10.1038/s41586-023-06647-8)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p8.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p3.1)\. - Sharmaet al\.\(2024\)M\. Sharma, M\. Tong, T\. Korbak, D\. Duvenaud, A\. Askell, S\. Bowman, E\. Durmus, Z\. Hatfield\-Dodds, S\. Johnston, S\. Kravec,et al\.Towards understanding sycophancy in language models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 110–144\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1),[§6](https://arxiv.org/html/2609.19843#S6.p4.1)\. - Simon \(1955\)H\. A\. SimonA Behavioral Model of Rational Choice\.The Quarterly Journal of Economics69\(1\),pp\. 99\(en\)\.External Links:ISSN 00335533,[Document](https://dx.doi.org/10.2307/1884852)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p11.1),[§1](https://arxiv.org/html/2609.19843#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1)\. - Singlaet al\.\(2025\)A\. Singla, A\. Sukharevsky, B\. Hall, L\. Yee, M\. Chui, and T\. BalakrishnanThe state of AI in 2025: Agents, innovation, and transformation\.Technical reportQuantumBlack, AI by McKinsey\.External Links:[Link](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.19843#S2.SS2.p2.1)\. - Späneet al\.\(2026\)A\. Späne, H\. Dutzler, M\. Schlemmer, and M\. HesseThe agentic AI revolution in retail: from vision to process reality\.ReportStrategy& \(Part of the PwC Network\)\.External Links:[Link](https://www.strategyand.pwc.com/de/en/industries/consumer-markets/agentic-ai-revolution-retail.html)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p2.1)\. - Stanovich \(2011\)K\. StanovichRationality and the reflective mind\.Oxford University Press\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1),[§3](https://arxiv.org/html/2609.19843#S3.p5.1),[§3](https://arxiv.org/html/2609.19843#S3.p6.1),[§6](https://arxiv.org/html/2609.19843#S6.p3.1)\. - Steinberger \(2026\)P\. SteinbergerIntroducing OpenClaw\.External Links:[Link](https://openclaw.ai/)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1)\. - Steyverset al\.\(2025\)M\. Steyvers, H\. Tejeda, A\. Kumar, C\. Belem, S\. Karny, X\. Hu, L\. W\. Mayer, and P\. SmythWhat large language models know and what people think they know\.Nature Machine Intelligence7\(2\),pp\. 221–231\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1)\. - Suriet al\.\(2024\)G\. Suri, L\. R\. Slater, A\. Ziaee, and M\. NguyenDo large language models show decision heuristics similar to humans? a case study using gpt\-3\.5\.\.Journal of Experimental Psychology: General153\(4\),pp\. 1066\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p4.1),[§3](https://arxiv.org/html/2609.19843#S3.p1.1)\. - Székelyet al\.\(2016\)N\. Székely, M\. Weinmann, and J\. Vom BrockeNudging people to pay co2 offsets–the effect of anchors in flight booking processes\.Twenty\-Fourth European Conference on Information Systems \(ECIS\)\.Cited by:[§4](https://arxiv.org/html/2609.19843#S4.p1.1)\. - Thaleret al\.\(2013\)R\. H\. Thaler, C\. R\. Sunstein, and J\. P\. BalzChoice architecture\. The behavioral foundations of public policy\.Princeton University Press Princeton, NJ\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p3.1)\. - Thaler and Sunstein \(2009\)R\. H\. Thaler and C\. R\. SunsteinNudge: Improving decisions about health, wealth, and happiness\.Penguin\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p3.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p3.1)\. - Tjuatjaet al\.\(2024\)L\. Tjuatja, V\. Chen, T\. Wu, A\. Talwalkwar, and G\. NeubigDo llms exhibit human\-like response biases? a case study in survey design\.Transactions of the Association for Computational Linguistics12,pp\. 1011–1026\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1)\. - Tversky and Kahneman \(1974\)A\. Tversky and D\. KahnemanJudgment under uncertainty: heuristics and biases: biases in judgments reveal some heuristics of thinking under uncertainty\.\.Science185\(4157\),pp\. 1124–1131\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1)\. - Tversky and Kahneman \(1981\)A\. Tversky and D\. KahnemanThe framing of decisions and the psychology of choice\.Science211\(4481\),pp\. 453–458\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1)\. - Van Gestelet al\.\(2021\)L\. C\. Van Gestel, M\. A\. Adriaanse, and D\. T\. D\. De RidderDo nudges make use of automatic processing? Unraveling the effects of a default nudge under type 1 and type 2 processing\.Comprehensive Results in Social Psychology5\(1\-3\),pp\. 4–24\(en\)\.External Links:ISSN 2374\-3603, 2374\-3611,[Document](https://dx.doi.org/10.1080/23743603.2020.1808456)Cited by:[§3](https://arxiv.org/html/2609.19843#S3.p5.1),[§6](https://arxiv.org/html/2609.19843#S6.p6.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1)\. - Vergaet al\.\(2024\)P\. Verga, S\. Hofstatter, S\. Althammer, Y\. Su, A\. Piktus, A\. Arkhangorodsky, M\. Xu, N\. White, and P\. LewisReplacing judges with juries: evaluating llm generations with a panel of diverse models\.External Links:2404\.18796,[Link](https://arxiv.org/abs/2404.18796)Cited by:[§12](https://arxiv.org/html/2609.19843#S12.SSx3.p2.1)\. - Watson \(2011\)K\. WatsonD\. kahneman\.\(2011\)\. thinking, fast and slow\. new york, ny: farrar, straus and giroux\. 499 pages\.\.Canadian Journal of Program Evaluation26\(2\),pp\. 111–113\.Cited by:[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p4.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p4.3)\. - Weiet al\.\(2022a\)J\. Wei, Y\. Tay, R\. Bommasani, C\. Raffel, B\. Zoph, S\. Borgeaud, D\. Yogatama, M\. Bosma, D\. Zhou, D\. Metzler,et al\.Emergent abilities of large language models\.arXiv preprint arXiv:2206\.07682\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§5\.2](https://arxiv.org/html/2609.19843#S5.SS2.p1.1),[§6](https://arxiv.org/html/2609.19843#S6.p5.1)\. - Weiet al\.\(2022b\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in Neural Information Processing Systems35,pp\. 24824–24837\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p2.1)\. - Weinmannet al\.\(2016\)M\. Weinmann, C\. Schneider, and J\. V\. BrockeDigital Nudging\.Business & Information Systems Engineering58\(6\),pp\. 433–436\(en\)\.External Links:ISSN 2363\-7005, 1867\-0202,[Document](https://dx.doi.org/10.1007/s12599-016-0453-1)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p3.1),[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p1.1)\. - Wenget al\.\(2025\)Z\. Weng, G\. Chen, and W\. WangDo as we do, not as you think: the conformity of large language models\.arXiv preprint arXiv:2501\.13381\.Cited by:[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1),[§6](https://arxiv.org/html/2609.19843#S6.p4.1)\. - Weston and Sukhbaatar \(2023\)J\. Weston and S\. SukhbaatarSystem 2 attention \(is something you might need too\)\.arXiv preprint arXiv:2311\.11829\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p8.1)\. - Xu and Ma \(2025\)N\. Xu and X\. MaLlm the genius paradox: a linguistic and math expert’s struggle with simple word\-based counting problems\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 3344–3370\.Cited by:[§12](https://arxiv.org/html/2609.19843#S12.p1.1),[§4](https://arxiv.org/html/2609.19843#S4.p3.1)\. - Yaoet al\.\(2022\)S\. Yao, H\. Chen, J\. Yang, and K\. NarasimhanWebshop: towards scalable real\-world web interaction with grounded language agents\.Advances in Neural Information Processing Systems35,pp\. 20744–20757\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1)\. - Yuet al\.\(2025\)Y\. Yu, H\. Li, Z\. Chen, Y\. Jiang, Y\. Li, J\. W\. Suchow, D\. Zhang, and K\. KhashanahFinMem: A Performance\-Enhanced LLM Trading Agent With Layered Memory and Character Design\.IEEE Transactions on Big Data,pp\. 1–18\(en\)\.External Links:ISSN 2332\-7790, 2372\-2096,[Document](https://dx.doi.org/10.1109/TBDATA.2025.3593370)Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1)\. - Zhanet al\.\(2024\)Q\. Zhan, Z\. Liang, Z\. Ying, and D\. KangInjecagent: benchmarking indirect prompt injections in tool\-integrated large language model agents\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 10471–10506\.Cited by:[§2\.3](https://arxiv.org/html/2609.19843#S2.SS3.p5.1)\. - Zhanget al\.\(2024a\)C\. Zhang, S\. He, J\. Qian, B\. Li, L\. Li, S\. Qin, Y\. Kang, M\. Ma, G\. Liu, Q\. Lin,et al\.Large language model\-brained gui agents: a survey\.Transactions on Machine Learning Research\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1),[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1)\. - Zhanget al\.\(2025\)D\. Zhang, Z\. Li, M\. Zhang, J\. Zhang, Z\. Liu, Y\. Yao, H\. Xu, J\. Zheng, X\. Chen, Y\. Zhang,et al\.From system 1 to system 2: a survey of reasoning large language models\.IEEE Transactions on Pattern Analysis and Machine Intelligence\.Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p5.1),[§1](https://arxiv.org/html/2609.19843#S1.p8.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p2.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p5.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p6.1),[§2\.1](https://arxiv.org/html/2609.19843#S2.SS1.p7.1),[§3](https://arxiv.org/html/2609.19843#S3.p5.1),[§6](https://arxiv.org/html/2609.19843#S6.p4.1),[§6](https://arxiv.org/html/2609.19843#S6.p7.1),[§6](https://arxiv.org/html/2609.19843#S6.p9.1)\. - Zhanget al\.\(2024b\)X\. Zhang, J\. Cao, and C\. YouCounting ability of large language models and impact of tokenization\.arXiv preprint arXiv:2410\.19730\.Cited by:[§12](https://arxiv.org/html/2609.19843#S12.p1.1),[§4](https://arxiv.org/html/2609.19843#S4.p3.1)\. - Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.WEBARENA: a realistic web environment for building autonomous agents\.In12th International Conference on Learning Representations, ICLR 2024,Cited by:[§1](https://arxiv.org/html/2609.19843#S1.p1.1)\. ### 8Complete posterior summaries for the pooled models Tables[81](https://arxiv.org/html/2609.19843#S8.T1)–[83](https://arxiv.org/html/2609.19843#S8.T3)report the complete posterior summaries of the three pooled models presented in Table[1](https://arxiv.org/html/2609.19843#S5.T1), namely posterior means and posterior standard deviations \(Est\. Error\) with 95% credible intervals on the log\-odds scale, together with the corresponding odds ratios \(OR=eβOR=e^\{\\beta\}\) and their 95% credible intervals\. Table[84](https://arxiv.org/html/2609.19843#S8.T4)additionally reports the posterior odds\-ratio contrasts underlying the hypothesis tests, namely the pairwise condition contrasts \(H1\) and the high versus no reasoning contrasts within each condition \(H2a and H2b\), summarized by posterior medians with 95% highest\-posterior\-density intervals\. Table 81:Model \(1\), pooled, complete posterior summary \(N=21,600N=21\{,\}600observations,3,6003\{,\}600agents\)\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept−0\.02\-0\.020\.060\.06−0\.13\-0\.130\.100\.100\.980\.980\.870\.871\.101\.10Defaults2\.482\.480\.100\.102\.282\.282\.682\.6811\.9411\.949\.809\.8014\.5614\.56Social Influence3\.323\.320\.110\.113\.103\.103\.553\.5527\.7827\.7822\.2822\.2834\.7034\.70Reasoning high−0\.05\-0\.050\.080\.08−0\.21\-0\.210\.120\.120\.950\.950\.810\.811\.121\.12Defaults×\\timesReas\. high−0\.63\-0\.630\.130\.13−0\.89\-0\.89−0\.36\-0\.360\.530\.530\.410\.410\.690\.69Soc\. Infl\.×\\timesReas\. high0\.590\.590\.160\.160\.270\.270\.910\.911\.801\.801\.311\.312\.482\.48Std\. dev\. agent intercept1\.151\.150\.040\.041\.081\.081\.221\.22–––Table 82:Model \(2\), pooled \+ provider control, complete posterior summary \(N=21,600N=21\{,\}600observations,3,6003\{,\}600agents\)\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept0\.540\.540\.070\.070\.400\.400\.680\.681\.711\.711\.491\.491\.971\.97Defaults2\.502\.500\.100\.102\.312\.312\.702\.7012\.1812\.1810\.0310\.0314\.9314\.93Social Influence3\.373\.370\.110\.113\.153\.153\.593\.5929\.0229\.0223\.4223\.4236\.2236\.22Reasoning high−0\.05\-0\.050\.080\.08−0\.21\-0\.210\.110\.110\.950\.950\.810\.811\.121\.12Provider Google−0\.99\-0\.990\.070\.07−1\.13\-1\.13−0\.84\-0\.840\.370\.370\.320\.320\.430\.43Provider OpenAI−0\.68\-0\.680\.070\.07−0\.83\-0\.83−0\.54\-0\.540\.500\.500\.440\.440\.580\.58Defaults×\\timesReas\. high−0\.62\-0\.620\.130\.13−0\.88\-0\.88−0\.37\-0\.370\.540\.540\.410\.410\.690\.69Soc\. Infl\.×\\timesReas\. high0\.600\.600\.160\.160\.280\.280\.920\.921\.821\.821\.331\.332\.512\.51Std\. dev\. agent intercept1\.111\.110\.040\.041\.041\.041\.181\.18–––Table 83:Model \(3\), pooled \+ provider \+ size controls, complete posterior summary \(N=21,600N=21\{,\}600observations,3,6003\{,\}600agents\)\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept0\.460\.460\.080\.080\.300\.300\.610\.611\.581\.581\.351\.351\.841\.84Defaults2\.502\.500\.100\.102\.302\.302\.702\.7012\.1912\.1910\.0210\.0214\.9114\.91Social Influence3\.353\.350\.110\.113\.143\.143\.583\.5828\.6328\.6323\.0323\.0335\.8635\.86Reasoning high−0\.05\-0\.050\.080\.08−0\.21\-0\.210\.110\.110\.950\.950\.810\.811\.121\.12Provider Google−0\.98\-0\.980\.070\.07−1\.13\-1\.13−0\.84\-0\.840\.370\.370\.320\.320\.430\.43Provider OpenAI−0\.69\-0\.690\.070\.07−0\.83\-0\.83−0\.54\-0\.540\.500\.500\.430\.430\.580\.58Size small0\.170\.170\.060\.060\.050\.050\.280\.281\.181\.181\.061\.061\.321\.32Defaults×\\timesReas\. high−0\.62\-0\.620\.130\.13−0\.88\-0\.88−0\.36\-0\.360\.540\.540\.420\.420\.700\.70Soc\. Infl\.×\\timesReas\. high0\.600\.600\.160\.160\.290\.290\.920\.921\.831\.831\.331\.332\.512\.51Std\. dev\. agent intercept1\.101\.100\.040\.041\.031\.031\.171\.17–––Table 84:Posterior odds\-ratio contrasts for the pooled sample, comprising condition contrasts \(H1\) and high versus no reasoning contrasts within each condition \(H2a and H2b\)\. Entries are posterior medians with 95% highest\-posterior\-density intervals in brackets\.Contrast\(1\)\(2\)\(3\)Pooled\+ Provider\+ Provider\+ SizeDefaults / No nudge8\.708\.708\.928\.928\.948\.94\[7\.54,9\.97\]\[7\.54,\\,9\.97\]\[7\.73,10\.17\]\[7\.73,\\,10\.17\]\[7\.76,10\.23\]\[7\.76,\\,10\.23\]Social influence / No nudge37\.2437\.2439\.1639\.1638\.6938\.69\[31\.04,44\.20\]\[31\.04,\\,44\.20\]\[32\.95,46\.51\]\[32\.95,\\,46\.51\]\[32\.58,45\.81\]\[32\.58,\\,45\.81\]Social influence / Defaults4\.274\.274\.394\.394\.334\.33\[3\.58,5\.05\]\[3\.58,\\,5\.05\]\[3\.64,5\.16\]\[3\.64,\\,5\.16\]\[3\.63,5\.11\]\[3\.63,\\,5\.11\]High / no reasoning, No nudge0\.950\.950\.950\.950\.950\.95\[0\.81,1\.12\]\[0\.81,\\,1\.12\]\[0\.81,1\.12\]\[0\.81,\\,1\.12\]\[0\.80,1\.11\]\[0\.80,\\,1\.11\]High / no reasoning, Defaults0\.510\.510\.510\.510\.510\.51\[0\.41,0\.62\]\[0\.41,\\,0\.62\]\[0\.41,0\.62\]\[0\.41,\\,0\.62\]\[0\.41,0\.62\]\[0\.41,\\,0\.62\]High / no reasoning, Social influence1\.711\.711\.741\.741\.741\.74\[1\.26,2\.20\]\[1\.26,\\,2\.20\]\[1\.30,2\.27\]\[1\.30,\\,2\.27\]\[1\.30,2\.27\]\[1\.30,\\,2\.27\] ### 9Per\-provider pooled models As a robustness check across providers, we re\-estimated the Model \(1\) specification separately for each provider, pooling the provider’s two models \(flagship and small variant,N=7,200N=7\{,\}200observations,1,2001\{,\}200agents each\)\. Tables[91](https://arxiv.org/html/2609.19843#S9.T1)–[93](https://arxiv.org/html/2609.19843#S9.T3)report the complete posterior summaries for the three provider pools, Table[94](https://arxiv.org/html/2609.19843#S9.T4)the corresponding odds\-ratio contrasts \(H1 and H2\), and Figure[91](https://arxiv.org/html/2609.19843#S9.F1)the predicted probabilities\. The Claude ceiling cells apply to the Anthropic pool \(see Appendix[11](https://arxiv.org/html/2609.19843#S11)\)\. The results are consistent with the patterns discussed in Section[5\.2](https://arxiv.org/html/2609.19843#S5.SS2)\. Table 91:Provider pool OpenAI \(GPT\-5\.4, GPT\-5\.4\-mini\), complete posterior summary\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept0\.000\.000\.100\.10−0\.20\-0\.200\.200\.201\.001\.000\.820\.821\.231\.23Defaults1\.901\.900\.160\.161\.591\.592\.222\.226\.706\.704\.884\.889\.249\.24Social Influence3\.333\.330\.200\.202\.952\.953\.733\.7327\.8127\.8119\.0519\.0541\.4841\.48Reasoning high−0\.04\-0\.040\.140\.14−0\.32\-0\.320\.240\.240\.960\.960\.730\.731\.271\.27Defaults×\\timesReas\. high−0\.38\-0\.380\.220\.22−0\.81\-0\.810\.050\.050\.680\.680\.440\.441\.051\.05Soc\. Infl\.×\\timesReas\. high1\.271\.270\.320\.320\.650\.651\.901\.903\.553\.551\.921\.926\.666\.66Std\. dev\. agent intercept1\.161\.160\.060\.061\.041\.041\.281\.28–––Table 92:Provider pool Google \(Gemini 3\.5 Flash, Gemini 3\.1 Flash Lite\), complete posterior summary\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept−0\.03\-0\.030\.100\.10−0\.23\-0\.230\.180\.180\.970\.970\.790\.791\.191\.19Defaults2\.202\.200\.170\.171\.871\.872\.532\.539\.059\.056\.516\.5112\.5412\.54Social Influence2\.592\.590\.180\.182\.252\.252\.942\.9413\.3113\.319\.459\.4518\.9718\.97Reasoning high−0\.05\-0\.050\.150\.15−0\.33\-0\.330\.240\.240\.950\.950\.720\.721\.271\.27Defaults×\\timesReas\. high−1\.10\-1\.100\.220\.22−1\.54\-1\.54−0\.66\-0\.660\.330\.330\.210\.210\.520\.52Soc\. Infl\.×\\timesReas\. high0\.260\.260\.240\.24−0\.22\-0\.220\.750\.751\.301\.300\.800\.802\.112\.11Std\. dev\. agent intercept1\.161\.160\.060\.061\.051\.051\.281\.28–––Table 93:Provider pool Anthropic \(Claude Sonnet 4\.6, Claude Haiku 4\.5\), complete posterior summary\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept−0\.02\-0\.020\.060\.06−0\.14\-0\.140\.100\.100\.980\.980\.870\.871\.101\.10Defaults3\.213\.210\.160\.162\.912\.913\.533\.5324\.8424\.8418\.3618\.3634\.2834\.28Social Influence4\.364\.360\.270\.273\.873\.874\.914\.9178\.2578\.2548\.1548\.15135\.01135\.01Reasoning high−0\.05\-0\.050\.080\.08−0\.21\-0\.210\.120\.120\.950\.950\.810\.811\.121\.12Defaults×\\timesReas\. high−0\.24\-0\.240\.210\.21−0\.65\-0\.650\.170\.170\.780\.780\.520\.521\.181\.18Soc\. Infl\.×\\timesReas\. high2\.132\.130\.730\.730\.880\.883\.763\.768\.458\.452\.402\.4043\.0743\.07Std\. dev\. agent intercept0\.170\.170\.110\.110\.010\.010\.390\.39–––Table 94:Posterior odds\-ratio contrasts for the three per\-provider pooled models, comprising condition contrasts \(H1\) and high versus no reasoning contrasts within each condition \(H2a and H2b\)\. Entries are posterior medians with 95% highest\-posterior\-density intervals\. These are marginal contrasts \(averaged over reasoning for H1\), and therefore differ from the coefficient odds\-ratios in Tables[91](https://arxiv.org/html/2609.19843#S9.T1)–[93](https://arxiv.org/html/2609.19843#S9.T3), which condition on the no\-reasoning reference level\.ContrastOpenAI poolGoogle poolAnthropic poolDefaults / No nudge5\.54\[4\.38,6\.88\]5\.54\\ \[4\.38,\\,6\.88\]5\.22\[4\.11,6\.51\]5\.22\\ \[4\.11,\\,6\.51\]21\.92\[17\.52,26\.77\]21\.92\\ \[17\.52,\\,26\.77\]Social influence / No nudge52\.23\[36\.66,72\.25\]52\.23\\ \[36\.66,\\,72\.25\]15\.14\[11\.68,19\.33\]15\.14\\ \[11\.68,\\,19\.33\]219\.91\[97\.68,449\.90\]219\.91\\ \[97\.68,\\,449\.90\]Social influence / Defaults9\.42\[6\.55,12\.81\]9\.42\\ \[6\.55,\\,12\.81\]2\.90\[2\.20,3\.68\]2\.90\\ \[2\.20,\\,3\.68\]10\.02\[4\.26,20\.93\]10\.02\\ \[4\.26,\\,20\.93\]High / no reasoning, No nudge0\.96\[0\.71,1\.24\]0\.96\\ \[0\.71,\\,1\.24\]0\.95\[0\.70,1\.25\]0\.95\\ \[0\.70,\\,1\.25\]0\.95\[0\.80,1\.12\]0\.95\\ \[0\.80,\\,1\.12\]High / no reasoning, Defaults0\.66\[0\.46,0\.90\]0\.66\\ \[0\.46,\\,0\.90\]0\.32\[0\.22,0\.43\]0\.32\\ \[0\.22,\\,0\.43\]0\.75\[0\.50,1\.07\]0\.75\\ \[0\.50,\\,1\.07\]High / no reasoning, Social influence3\.39\[1\.78,5\.64\]3\.39\\ \[1\.78,\\,5\.64\]1\.24\[0\.79,1\.74\]1\.24\\ \[0\.79,\\,1\.74\]7\.55\[1\.10,30\.21\]7\.55\\ \[1\.10,\\,30\.21\]\(a\)OpenAI \(b\)Google \(c\)Anthropic Figure 91:Predicted target product selection probabilities across conditions and reasoning configurations for the per\-provider pooled models\. Red indicates no\-reasoning agents, and blue indicates high\-reasoning agents\. ### 10Per\-size pooled models As an exploratory analysis of model scale, we re\-estimated the Model \(1\) specification separately for each model size class, pooling together the three flagship models \(GPT\-5\.4, Gemini 3\.5 Flash, Claude Sonnet 4\.6\) and the three small variants \(GPT\-5\.4\-mini, Gemini 3\.1 Flash Lite, Claude Haiku 4\.5;N=10,800N=10\{,\}800observations,1,8001\{,\}800agents each\)\. Tables[101](https://arxiv.org/html/2609.19843#S10.T1)and[102](https://arxiv.org/html/2609.19843#S10.T2)report the complete posterior summaries for the two size pools, Table[103](https://arxiv.org/html/2609.19843#S10.T3)the corresponding odds\-ratio contrasts \(H1 and H2\), and Figure[101](https://arxiv.org/html/2609.19843#S10.F1)the predicted probabilities\. Two ceiling caveats apply: Claude Haiku 4\.5’s \(quasi\-\)perfect default compliance contributes to the small pool’s large default effect, and Claude Sonnet 4\.6’s perfect social influence compliance to the flagship pool’s large social influence effect \(see Appendix[11](https://arxiv.org/html/2609.19843#S11)\)\. The results are consistent with the patterns discussed in Section[5\.2](https://arxiv.org/html/2609.19843#S5.SS2)\. Table 101:Size pool flagship \(GPT\-5\.4, Gemini 3\.5 Flash, Claude Sonnet 4\.6\), complete posterior summary\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept0\.050\.050\.090\.09−0\.12\-0\.120\.220\.221\.051\.050\.880\.881\.241\.24Defaults1\.981\.980\.140\.141\.701\.702\.262\.267\.247\.245\.505\.509\.599\.59Social Influence4\.734\.730\.230\.234\.294\.295\.215\.21113\.52113\.5273\.2673\.26182\.87182\.87Reasoning high−0\.06\-0\.060\.120\.12−0\.31\-0\.310\.180\.180\.940\.940\.740\.741\.201\.20Defaults×\\timesReas\. high−0\.89\-0\.890\.190\.19−1\.25\-1\.25−0\.52\-0\.520\.410\.410\.290\.290\.590\.59Soc\. Infl\.×\\timesReas\. high−0\.10\-0\.100\.310\.31−0\.70\-0\.700\.490\.490\.900\.900\.490\.491\.641\.64Std\. dev\. agent intercept1\.231\.230\.050\.051\.131\.131\.341\.34–––Table 102:Size pool small \(GPT\-5\.4\-mini, Gemini 3\.1 Flash Lite, Claude Haiku 4\.5\), complete posterior summary\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept−0\.08\-0\.080\.070\.07−0\.21\-0\.210\.060\.060\.930\.930\.810\.811\.061\.06Defaults2\.892\.890\.130\.132\.642\.643\.153\.1517\.9617\.9613\.9513\.9523\.4423\.44Social Influence2\.482\.480\.120\.122\.262\.262\.722\.7212\.0012\.009\.559\.5515\.2115\.21Reasoning high−0\.03\-0\.030\.100\.10−0\.22\-0\.220\.150\.150\.970\.970\.800\.801\.171\.17Defaults×\\timesReas\. high−0\.23\-0\.230\.170\.17−0\.57\-0\.570\.100\.100\.790\.790\.570\.571\.101\.10Soc\. Infl\.×\\timesReas\. high0\.770\.770\.180\.180\.410\.411\.121\.122\.152\.151\.511\.513\.063\.06Std\. dev\. agent intercept0\.810\.810\.050\.050\.710\.710\.910\.91–––Table 103:Posterior odds\-ratio contrasts for the two per\-size pooled models, comprising condition contrasts \(H1\) and high versus no reasoning contrasts within each condition \(H2a and H2b\)\. Entries are posterior medians with 95% highest\-posterior\-density intervals\. These are marginal contrasts \(averaged over reasoning for H1\), and therefore differ from the coefficient odds ratios in Tables[101](https://arxiv.org/html/2609.19843#S10.T1)–[102](https://arxiv.org/html/2609.19843#S10.T2), which condition on the no\-reasoning reference level\.ContrastFlagship poolSmall poolDefaults / No nudge4\.64\[3\.82,5\.60\]4\.64\\ \[3\.82,\\,5\.60\]15\.94\[13\.24,19\.17\]15\.94\\ \[13\.24,\\,19\.17\]Social influence / No nudge107\.57\[75\.23,147\.30\]107\.57\\ \[75\.23,\\,147\.30\]17\.59\[14\.42,20\.96\]17\.59\\ \[14\.42,\\,20\.96\]Social influence / Defaults23\.15\[16\.44,31\.59\]23\.15\\ \[16\.44,\\,31\.59\]1\.10\[0\.89,1\.34\]1\.10\\ \[0\.89,\\,1\.34\]High / no reasoning, No nudge0\.94\[0\.71,1\.17\]0\.94\\ \[0\.71,\\,1\.17\]0\.97\[0\.80,1\.16\]0\.97\\ \[0\.80,\\,1\.16\]High / no reasoning, Defaults0\.39\[0\.29,0\.50\]0\.39\\ \[0\.29,\\,0\.50\]0\.77\[0\.56,0\.99\]0\.77\\ \[0\.56,\\,0\.99\]High / no reasoning, Social influence0\.85\[0\.44,1\.40\]0\.85\\ \[0\.44,\\,1\.40\]2\.09\[1\.50,2\.74\]2\.09\\ \[1\.50,\\,2\.74\]\(a\)Flagship models \(b\)Small variants Figure 101:Predicted target product selection probabilities across conditions and reasoning configurations for the per\-size pooled models\. Red indicates no\-reasoning agents, and blue indicates high\-reasoning agents\. ### 11Per\-model analyses As a robustness check across individual models, we re\-estimated the Model \(1\) specification separately for each of the six LLM agent backbones \(N=3,600N=3\{,\}600observations,600600agents each\)\. Tables[111](https://arxiv.org/html/2609.19843#S11.T1)–[116](https://arxiv.org/html/2609.19843#S11.T6)report the complete posterior summaries and Figure[111](https://arxiv.org/html/2609.19843#S11.F1)the corresponding estimated probabilities\. Both nudge main effects \(H1\) are credible in every single model\. Note that Claude Sonnet 4\.6 exhibits complete separation in the social influence condition \(100% target selection at both reasoning configurations\) and Claude Haiku 4\.5 quasi\-complete separation in the default condition \(100% and 99\.3%\)\. The affected coefficients are partly prior\-dependent and should be read via their credible interval bounds\. Regarding the default attenuation under high reasoning \(H2a\), the effect is concentrated in the flagship models\. It is credible for GPT\-5\.4 \(β=−0\.68\\beta=\-0\.68,CI\.95CI\_\{\.95\}\[−1\.30,−0\.05\]\[\-1\.30,\\,\-0\.05\], predicted probability declining from 84% to 73%\) and strongest for Gemini 3\.5 Flash \(β=−1\.61\\beta=\-1\.61,CI\.95CI\_\{\.95\}\[−2\.30,−0\.96\]\[\-2\.30,\\,\-0\.96\], declining from 84% to 48%, i\.e\., back to chance level\), as well as, more moderately, for Gemini 3\.1 Flash Lite \(β=−0\.51\\beta=\-0\.51,CI\.95CI\_\{\.95\}\[−0\.96,−0\.04\]\[\-0\.96,\\,\-0\.04\], 93% to 88%\)\. By contrast, the smaller GPT\-5\.4\-mini shows virtually no reasoning moderation for defaults \(β=0\.04\\beta=0\.04,CI\.95CI\_\{\.95\}\[−0\.51,0\.62\]\[\-0\.51,\\,0\.62\]including0\.00\.0, 89% vs\. 88%\), and neither Claude model yields a credible interaction\. Claude Sonnet 4\.6 remains highly compliant at both reasoning configurations \(β=−0\.13\\beta=\-0\.13,CI\.95CI\_\{\.95\}\[−0\.58,0\.33\]\[\-0\.58,\\,0\.33\], 92% vs\. 90%\), and Claude Haiku 4\.5 sits at a ceiling in both cells \(100% and 99\.3% raw selection rates, i\.e\., quasi\-complete separation\), so its wide interval \(β=−2\.15\\beta=\-2\.15,CI\.95CI\_\{\.95\}\[−5\.92,0\.21\]\[\-5\.92,\\,0\.21\]\) is prior\-dependent and uninformative\. Regarding the social influence amplification under high reasoning \(H2b\), the effect is concentrated in the smaller model variants\. It is credible for GPT\-5\.4\-mini \(β=2\.46\\beta=2\.46,CI\.95CI\_\{\.95\}\[1\.61,3\.38\]\[1\.61,\\,3\.38\], estimated probabilities rising from 91% to 99%\) and for Claude Haiku 4\.5 \(β=2\.10\\beta=2\.10,CI\.95CI\_\{\.95\}\[0\.82,3\.70\]\[0\.82,\\,3\.70\], rising from 97% to close to 100%\)\. For the two Gemini models, the interaction is positive but not credible \(Gemini 3\.5 Flashβ=0\.43\\beta=0\.43,CI\.95CI\_\{\.95\}\[−0\.44,1\.30\]\[\-0\.44,\\,1\.30\], and Gemini 3\.1 Flash Liteβ=0\.20\\beta=0\.20,CI\.95CI\_\{\.95\}\[−0\.21,0\.61\]\[\-0\.21,\\,0\.61\]\)\. For Claude Sonnet 4\.6, the social influence condition produced 100% target selection at both reasoning configurations, leaving no variability for reasoning to modulate, so the resulting interaction estimate is accordingly uninformative \(β=1\.11\\beta=1\.11,CI\.95CI\_\{\.95\}\[−3\.42,7\.10\]\[\-3\.42,\\,7\.10\]including0\.00\.0\)\. Finally, GPT\-5\.4 constitutes the sole reversal, as its interaction is credibly negative \(β=−1\.49\\beta=\-1\.49,CI\.95CI\_\{\.95\}\[−2\.86,−0\.24\]\[\-2\.86,\\,\-0\.24\]\), although both of its social influence cells lie near the ceiling \(99\.7% vs\. 98\.8% predicted probability\), so the practical magnitude of this reversal is small\. Table 111:GPT\-5\.4, complete posterior summary\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept0\.030\.030\.150\.15−0\.27\-0\.270\.320\.321\.031\.030\.770\.771\.371\.37Defaults1\.641\.640\.240\.241\.181\.182\.102\.105\.145\.143\.243\.248\.178\.17Social Influence5\.875\.870\.600\.604\.834\.837\.157\.15354\.15354\.15124\.64124\.641273\.981273\.98Reasoning high−0\.00\-0\.000\.220\.22−0\.42\-0\.420\.420\.421\.001\.000\.660\.661\.521\.52Defaults×\\timesReas\. high−0\.68\-0\.680\.320\.32−1\.30\-1\.30−0\.05\-0\.050\.510\.510\.270\.270\.950\.95Soc\. Infl\.×\\timesReas\. high−1\.49\-1\.490\.670\.67−2\.86\-2\.86−0\.24\-0\.240\.230\.230\.060\.060\.790\.79Std\. dev\. agent intercept1\.231\.230\.090\.091\.051\.051\.421\.42–––Table 112:GPT\-5\.4\-mini, complete posterior summary\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept0\.000\.000\.130\.13−0\.24\-0\.240\.260\.261\.001\.000\.790\.791\.301\.30Defaults2\.062\.060\.210\.211\.651\.652\.482\.487\.817\.815\.225\.2211\.9011\.90Social Influence2\.372\.370\.220\.221\.951\.952\.802\.8010\.6710\.677\.037\.0316\.3816\.38Reasoning high−0\.11\-0\.110\.180\.18−0\.46\-0\.460\.240\.240\.900\.900\.630\.631\.271\.27Defaults×\\timesReas\. high0\.040\.040\.290\.29−0\.51\-0\.510\.620\.621\.041\.040\.600\.601\.851\.85Soc\. Infl\.×\\timesReas\. high2\.462\.460\.450\.451\.611\.613\.383\.3811\.6811\.685\.025\.0229\.4229\.42Std\. dev\. agent intercept0\.950\.950\.080\.080\.780\.781\.121\.12–––Table 113:Gemini 3\.5 Flash, complete posterior summary\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept0\.120\.120\.160\.16−0\.20\-0\.200\.440\.441\.131\.130\.820\.821\.551\.55Defaults1\.531\.530\.250\.251\.051\.052\.032\.034\.624\.622\.842\.847\.597\.59Social Influence3\.653\.650\.320\.323\.053\.054\.314\.3138\.4338\.4321\.0721\.0774\.1474\.14Reasoning high−0\.11\-0\.110\.230\.23−0\.56\-0\.560\.350\.350\.900\.900\.570\.571\.421\.42Defaults×\\timesReas\. high−1\.61\-1\.610\.340\.34−2\.30\-2\.30−0\.96\-0\.960\.200\.200\.100\.100\.380\.38Soc\. Infl\.×\\timesReas\. high0\.430\.430\.440\.44−0\.44\-0\.441\.301\.301\.531\.530\.640\.643\.663\.66Std\. dev\. agent intercept1\.361\.360\.090\.091\.181\.181\.541\.54–––Table 114:Gemini 3\.1 Flash Lite, complete posterior summary\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept−0\.14\-0\.140\.090\.09−0\.32\-0\.320\.030\.030\.870\.870\.730\.731\.031\.03Defaults2\.702\.700\.180\.182\.352\.353\.073\.0714\.8414\.8410\.4410\.4421\.5221\.52Social Influence1\.741\.740\.150\.151\.461\.462\.042\.045\.715\.714\.324\.327\.717\.71Reasoning high−0\.01\-0\.010\.120\.12−0\.25\-0\.250\.230\.230\.990\.990\.780\.781\.261\.26Defaults×\\timesReas\. high−0\.51\-0\.510\.240\.24−0\.96\-0\.96−0\.04\-0\.040\.600\.600\.380\.380\.960\.96Soc\. Infl\.×\\timesReas\. high0\.200\.200\.210\.21−0\.21\-0\.210\.610\.611\.221\.220\.810\.811\.841\.84Std\. dev\. agent intercept0\.300\.300\.130\.130\.030\.030\.540\.54–––Table 115:Claude Sonnet 4\.6, complete posterior summary\. Complete separation in the social influence condition \(100% target selection at both reasoning configurations\), so the affected estimates are partly prior\-dependent\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept0\.020\.020\.080\.08−0\.14\-0\.140\.190\.191\.031\.030\.870\.871\.201\.20Defaults2\.432\.430\.170\.172\.092\.092\.782\.7811\.3811\.388\.118\.1116\.1716\.17Social Influence9\.319\.312\.972\.975\.745\.7417\.1517\.1511036\.9911036\.99310\.38310\.382\.79∗1072\.79\*10^\{7\}Reasoning high−0\.11\-0\.110\.120\.12−0\.35\-0\.350\.120\.120\.890\.890\.710\.711\.131\.13Defaults×\\timesReas\. high−0\.13\-0\.130\.240\.24−0\.58\-0\.580\.330\.330\.880\.880\.560\.561\.401\.40Soc\. Infl\.×\\timesReas\. high1\.111\.112\.682\.68−3\.42\-3\.427\.107\.103\.053\.050\.030\.031207\.721207\.72Std\. dev\. agent intercept0\.160\.160\.110\.110\.010\.010\.400\.40–––Table 116:Claude Haiku 4\.5, complete posterior summary\. Quasi\-complete separation in the default condition \(100% and 99\.3% target selection\); the affected estimates are partly prior\-dependent\.95% CIOR 95% CIParameterEstimateEst\. ErrorlowerupperORlowerupperIntercept−0\.06\-0\.060\.080\.08−0\.22\-0\.220\.110\.110\.950\.950\.800\.801\.121\.12Defaults7\.427\.421\.481\.485\.335\.3311\.1511\.151671\.811671\.81206\.53206\.5369250\.6469250\.64Social Influence3\.693\.690\.270\.273\.183\.184\.244\.2440\.0740\.0724\.0124\.0169\.7269\.72Reasoning high0\.010\.010\.120\.12−0\.23\-0\.230\.250\.251\.011\.010\.790\.791\.281\.28Defaults×\\timesReas\. high−2\.15\-2\.151\.541\.54−5\.92\-5\.920\.210\.210\.120\.12<0\.01<0\.011\.231\.23Soc\. Infl\.×\\timesReas\. high2\.102\.100\.740\.740\.820\.823\.703\.708\.158\.152\.272\.2740\.2940\.29Std\. dev\. agent intercept0\.180\.180\.120\.120\.010\.010\.430\.43–––\(a\)GPT\-5\.4 \(b\)GPT\-5\.4\-mini \(c\)Gemini 3\.5 Flash \(d\)Gemini 3\.1 Flash Lite \(e\)Claude Sonnet 4\.6 \(f\)Claude Haiku 4\.5 Figure 111:Predicted target product selection probabilities across conditions and reasoning configurations for the six per\-model fits\. Red indicates no\-reasoning agents, and blue indicates high\-reasoning agents\. ### 12Pre\-experiment: letter\-counting manipulation check This appendix reports the design, the statistical tests, and a reasoning\-trace analysis for the pre\-experiment mentioned in Section[4](https://arxiv.org/html/2609.19843#S4)\. The pre\-experiment grounds the System 1 versus System 2 interpretation of the reasoning configuration\. Kahneman treats counting the occurrences of a letter as a canonical System 2 task, one that is effortful, demands sustained attention, and cannot be performed by fast automatic processing\([Kahneman, 2011](https://arxiv.org/html/2609.19843#bib.bib107)\)\. Letter counting in LLMs is a direct analogue: it requires inspecting the word character by character rather than retrieving a token\-level pattern, and both the architectural analysis of[Zhang et al\. \(2024b\)](https://arxiv.org/html/2609.19843#bib.bib105)and the empirical work of[Xu and Ma \(2025\)](https://arxiv.org/html/2609.19843#bib.bib106)identify the engagement of explicit reasoning, rather than scale or fine\-tuning, as the intervention that reliably makes LLMs count letters correctly\. If the high\-reasoning configuration engages System 2\-style deliberation, it should therefore improve letter\-counting accuracy\. Whether it does, where the improvement is located, and in which output channel the models perform the letter\-level work are the subject of this appendix\. #### Design Each of the six model backbones was asked five counting items, 100 times per reasoning configuration, yielding5×100×2×6=6,0005\\times 100\\times 2\\times 6=6\{,\}000independent API calls \(3,0003\{,\}000per configuration\)\. Items were posed as a bare user message with no system prompt, no browser, and no tools; the provider\-side reasoning parameters were the same objects used in the main experiment, imported from the experiment code so that the configurations cannot drift between studies\. The five items \(Table[121](https://arxiv.org/html/2609.19843#S12.T1)\) comprise one in\-distribution word whose answer is plausibly memorized \(“raspberry”, the well\-known “strawberry” case\), three novel pseudo\-words that cannot be answered from memory, and one zero\-count control in which the target letter never appears\. Responses were graded automatically by a case\-insensitive word\-boundary match on the accepted answers, applied to the*final answer only*and never to the reasoning trace\. This is deliberate: traces routinely enumerate the letters correctly \(“1…2…3”\) while the model still emits a wrong final count, so scoring the trace would credit the model for work that did not reach its answer\. Word boundaries prevent the digit “3” from matching inside “31” and “one” from matching inside “none”\. Table 121:The five counting items\. Each was asked 100 times per model and reasoning configuration\.ItemPromptCountAccepted answersFunction1How manyr’s in “raspberry”?3three, 3In\-distribution2How manyw’s in “waspbewwy”?3three, 3Novel pseudo\-word3How manyw’s in “raspberry”?0zero, 0, null, none, noZero\-count control4How manys’s in “laspbelly”?1one, 1Novel pseudo\-word5How manyb’s in “baspbebbby”?5five, 5Novel pseudo\-word #### Accuracy and statistical tests Pooled across all models and items, accuracy rose from80\.3%80\.3\\%\(2,410/3,0002\{,\}410/3\{,\}000\) under no reasoning to99\.3%99\.3\\%\(2,980/3,0002\{,\}980/3\{,\}000\) under high reasoning, a risk difference of19\.019\.0percentage points \(Figure[121](https://arxiv.org/html/2609.19843#S12.F1)\)\. The gain is concentrated exactly where the theory predicts, namely on the novel pseudo\-words that cannot be answered from memory, and not at all on the control item \(Table[122](https://arxiv.org/html/2609.19843#S12.T2)\)\. No model was ever less accurate under high reasoning, on any item \(Table[124](https://arxiv.org/html/2609.19843#S12.T4)\)\. Figure 121:Correct answers on the letter\-counting task, pooled across all six models and five items \(3,000 calls per configuration\), under the no\-reasoning and high\-reasoning configurations\.Table 122:Correct answers per item, pooled over the six models \(600 calls per item and configuration\)\.ItemTypeNo reasoningHigh reasoning1 “raspberry” /rMemorizable491/600\(81\.8%\)491/600\\ \(81\.8\\%\)600/600\(100\.0%\)600/600\\ \(100\.0\\%\)2 “waspbewwy” /wNovel318/600\(53\.0%\)318/600\\ \(53\.0\\%\)581/600\(96\.8%\)581/600\\ \(96\.8\\%\)3 “raspberry” /wZero\-count control600/600\(100\.0%\)600/600\\ \(100\.0\\%\)600/600\(100\.0%\)600/600\\ \(100\.0\\%\)4 “laspbelly” /sNovel467/600\(77\.8%\)467/600\\ \(77\.8\\%\)600/600\(100\.0%\)600/600\\ \(100\.0\\%\)5 “baspbebbby” /bNovel534/600\(89\.0%\)534/600\\ \(89\.0\\%\)599/600\(99\.8%\)599/600\\ \(99\.8\\%\)Total2,410/3,000\(80\.3%\)2\{,\}410/3\{,\}000\\ \(80\.3\\%\)2,980/3,000\(99\.3%\)2\{,\}980/3\{,\}000\\ \(99\.3\\%\)The five items differ enormously in baseline difficulty, from a floor of53%53\\%on item 2 to a ceiling on item 3, so a single comparison of all500500calls per configuration would confound the reasoning effect with item composition\. We therefore tested the reasoning effect with the Cochran–Mantel–Haenszel \(CMH\) test\([Cochran, 1954](https://arxiv.org/html/2609.19843#bib.bib74);[Rayner, 2020](https://arxiv.org/html/2609.19843#bib.bib73);[Mantel and Haenszel, 1959](https://arxiv.org/html/2609.19843#bib.bib75)\), which is*stratified*by item\. The two configurations are compared within each item separately, and the five within\-item comparisons are then pooled into a single test statistic\. The procedure is in effect a fixed\-effect meta\-analysis across items, in which every item serves as its own control\. The model\-level tests are Holm\-corrected over the four testable models, since the two Claude models admit no test \(see below\)\. Significance is assessed by randomization rather than by theχ2\\chi^\{2\}approximation\. The reasoning labels are reshuffled within each item \(20,00020\{,\}000resamples\), and thepp\-value is the proportion of reshuffles that produce an association at least as strong as the observed one\. This procedure is exact for sparse and well\-filled items alike, which is why we applied it uniformly instead of switching tests where cells are small\. A data\-dependent choice of test would amount to selecting the test after seeing the data\. Two features of the data govern how Table[123](https://arxiv.org/html/2609.19843#S12.T3)should be read\. The first is*ceiling*\. In 18 of the 30 model×\\timesitem cells, both configurations are at100%100\\%\. Such cells carry no information about the reasoning effect and drop out of the CMH statistic automatically\. Both Claude models are at500/500500/500in*both*configurations, so they exhibit no variance at all and are reported as untestable rather than as null results\. The second is*separation*\. In cells such as GPT\-5\.4 on item 2 which is correct in 0 of 100 calls without reasoning and in 100 of 100 with it \(Table[124](https://arxiv.org/html/2609.19843#S12.T4), one configuration is perfect and the other at zero, which makes the odds ratio infinite\. The CMH test itself remains well\-defined in this situation, but the odds ratio does not, so the effect size to read is the risk difference, that is, the improvement in percentage points\. Because every stratum contains exactly 100 calls per configuration, the marginal risk difference equals the unweighted mean of the per\-item risk differences\. Table 123:Letter\-counting accuracy per model and the reasoning effect\.Δ\\Delta\(pp\) is the improvement in percentage points under high reasoning\.pp\-values are CMH tests stratified by item, assessed by randomization \(20,00020\{,\}000within\-item reshuffles, bounded below by1/20,0011/20\{,\}001\) and Holm\-corrected across the four testable models\. The two Claude models are at100%100\\%in both configurations and therefore admit no test\.ModelSizeNo reasoningHigh reasoningΔ\\Delta\(pp\)pp\(Holm\)GPT\-5\.4Flagship384/500\(76\.8%\)384/500\\ \(76\.8\\%\)500/500\(100\.0%\)500/500\\ \(100\.0\\%\)\+23\.2\+23\.2<0\.001<0\.001GPT\-5\.4\-miniSmall163/500\(32\.6%\)163/500\\ \(32\.6\\%\)480/500\(96\.0%\)480/500\\ \(96\.0\\%\)\+63\.4\+63\.4<0\.001<0\.001Gemini 3\.5 FlashFlagship493/500\(98\.6%\)493/500\\ \(98\.6\\%\)500/500\(100\.0%\)500/500\\ \(100\.0\\%\)\+1\.4\+1\.40\.0140\.014Gemini 3\.1 Flash LiteSmall370/500\(74\.0%\)370/500\\ \(74\.0\\%\)500/500\(100\.0%\)500/500\\ \(100\.0\\%\)\+26\.0\+26\.0<0\.001<0\.001Claude Sonnet 4\.6Flagship500/500\(100\.0%\)500/500\\ \(100\.0\\%\)500/500\(100\.0%\)500/500\\ \(100\.0\\%\)0\.00\.0untestableClaude Haiku 4\.5Small500/500\(100\.0%\)500/500\\ \(100\.0\\%\)500/500\(100\.0%\)500/500\\ \(100\.0\\%\)0\.00\.0untestableTable 124:Correct answers per model and item \(out of 100 calls per cell\) under the no\-reasoning \(No\) and high\-reasoning \(High\) configurations\. Items as defined in Table[121](https://arxiv.org/html/2609.19843#S12.T1)\. This is the cell\-level grid underlying the two marginal tables above\.Item 1Item 2Item 3Item 4Item 5ModelNoHighNoHighNoHighNoHighNoHighGPT\-5\.495951001000010010010010010010099991001009090100100GPT\-5\.4\-mini551001009981811001001001005510010044449999Gemini 3\.5 Flash9393100100100100100100100100100100100100100100100100100100Gemini 3\.1 Flash Lite9898100100991001001001001001006363100100100100100100Claude Sonnet 4\.6100100100100100100100100100100100100100100100100100100100100Claude Haiku 4\.5100100100100100100100100100100100100100100100100100100100100At the aggregate level the effect is credible under every test we considered\. The CMH test stratified by all 30 model×\\timesitem cells yields a pooled \(Mantel–Haenszel\) odds ratio of266\.1266\.1, the same quantity a fixed\-effect meta\-analysis would report, with a randomizationp<0\.001p<0\.001\. Because the individual calls are clustered within cells, we additionally report two tests that compress each model×\\timesitem cell into a single paired observation and are therefore immune to both the ceiling and the clustering\. Of the 30 cells, 12 improved under high reasoning, none deteriorated, and 18 were tied at the ceiling, giving an exact sign testp=0\.0005p=0\.0005and a Wilcoxon signed\-rankp=0\.0022p=0\.0022\. We note for completeness that a single unstratified2×22\\times 2test over all6,0006\{,\}000calls would returnp<10−100p<10^\{\-100\}, but such a test treats 100 correlated calls per cell as independent observations and overstates the evidence by orders of magnitude\. It is not the basis for any claim here\. Figure[122](https://arxiv.org/html/2609.19843#S12.F2)shows the per\-model accuracies\. Figure 122:Correct answers on the letter\-counting task per model backbone, pooled over the five items \(500 calls per model and configuration\), under the no\-reasoning and high\-reasoning configurations\. #### Reasoning\-trace analysis Accuracy establishes*that*the manipulation works, but not*how*\. We therefore annotated all6,0006\{,\}000responses for whether the text decomposes the word into individual letters, that is, whether it spells the word out \(“r\-a\-s\-p\-b\-e\-r\-r\-y”\), marks the target letters inside the word, enumerates the letters as list items, or walks their positions or indices\. Merely*talking about*decomposing \(“I will go through each letter”\) was scored negative, as the construct of interest is the observable decomposition itself\. The annotation was applied separately to two channels: the model’s visible answer, and its hidden reasoning trace where one was returned\. Because a single automatic labeller would be unvalidated, we used two independent ones\. The first is a deterministic pattern matcher covering the notational variants observed in the corpus\. The second is a panel of three LLM judges, one per provider \(Gemini 3\.1 Flash Lite, GPT\-5\.4\-mini, Claude Haiku 4\.5\), each returning a structured verdict under a fixed rubric and blind to which model produced the text, with the majority vote taken as the panel label; a panel of small, diverse judges evaluates more cheaply and with less intra\-model bias than a single large judge\([Verga et al\., 2024](https://arxiv.org/html/2609.19843#bib.bib76)\)\. The two labellers agree almost perfectly on the visible answer \(Cohen’sκ=0\.996\\kappa=0\.996, raw agreement99\.8%99\.8\\%,n=6,000n=6\{,\}000\), and the three judges agree among themselves at Fleiss’κ=0\.987\\kappa=0\.987, with5,9445\{,\}944of6,0006\{,\}000verdicts unanimous\. Agreement on the trace channel is high in raw terms \(96\.0%96\.0\\%\) but the correspondingκ\\kappais depressed to0\.7210\.721by the strong prevalence skew \(93%93\\%of traces are positive\); the free\-marginal \(Randolph\) coefficient, which is robust to that skew, is0\.8850\.885\. Because the three judges are themselves three of the six annotated backbones, we verified that this cannot have driven the labels: recomputing the panel for each provider’s own2,0002\{,\}000rows while excluding that provider’s judge reproduces the full panel’s label on every row on which the reduced two\-judge panel is decisive \(agreement=1\.000=1\.000; the remaining 2, 8, and 15 rows per provider are ties between the two remaining judges and admit no verdict\)\. Table 125:Percentage of calls in which the model decomposes the word into individual letters, by channel \(500 calls per model and configuration\)\. “Visible” refers to the model’s answer text, “trace” to its hidden reasoning trace, and “any channel” to either\. Traces exist only under the high\-reasoning configuration; 8 GPT\-5\.4 and 3 GPT\-5\.4\-mini high\-reasoning calls returned no trace, so their trace percentages are based on 492 and 497 calls, while the any\-channel column counts those calls as non\-decomposing \(which is why it can fall below the trace column\)\.Visible answerTraceAny channelModelNo reas\.High reas\.High reas\.No reas\.High reas\.GPT\-5\.45\.25\.20\.00\.090\.090\.05\.25\.288\.688\.6GPT\-5\.4\-mini0\.40\.40\.00\.076\.976\.90\.40\.476\.476\.4Gemini 3\.5 Flash96\.096\.091\.691\.692\.492\.496\.096\.099\.699\.6Gemini 3\.1 Flash Lite67\.867\.859\.259\.298\.498\.467\.867\.899\.699\.6Claude Sonnet 4\.687\.087\.0100\.0100\.0100\.0100\.087\.087\.0100\.0100\.0Claude Haiku 4\.5100\.0100\.087\.487\.4100\.0100\.0100\.0100\.0100\.0100\.0Pooled59\.459\.456\.456\.493\.093\.059\.459\.494\.094\.0Table[125](https://arxiv.org/html/2609.19843#S12.T5)shows that the reasoning configuration does*not*make models decompose more visibly: pooled, the rate is59\.4%59\.4\\%without reasoning and56\.4%56\.4\\%with it\. What changes is the channel\. Under high reasoning every model decomposes the word in its hidden trace on the large majority of calls \(76\.976\.9–100%100\\%\), so that decomposition somewhere in the observable output rises from59\.4%59\.4\\%to94\.0%94\.0\\%\. The provider\-reported token counts corroborate this directly for the four models that expose them: the median number of reasoning tokens under the no\-reasoning configuration is00in every case, rising under high reasoning to88\.588\.5\(GPT\-5\.4\),6565\(GPT\-5\.4\-mini\),278278\(Gemini 3\.5 Flash\) and283\.5283\.5\(Gemini 3\.1 Flash Lite\)\. Anthropic does not report a separate reasoning\-token count, so no token\-level figure can be given for the two Claude models; their trace column rests on the returned thinking text alone\. The decisive pattern is the relationship between the two analyses\. A model’s accuracy gain from the reasoning toggle is almost perfectly inversely related to how often it already decomposes the word*without*reasoning enabled: - •The twoClaudemodels decompose the word in essentially every no\-reasoning answer \(100%100\\%and87%87\\%\), spelling it out and enumerating positions in the visible reply even though the thinking channel is disabled\. They are consequently already at100%100\\%accuracy without reasoning, and the toggle has no room left to work \(Δ=0\\Delta=0for both\)\. - •The twoOpenAImodels essentially never decompose visibly \(5\.2%5\.2\\%and0\.4%0\.4\\%, falling to0%0\\%when reasoning is enabled and the work moves into the trace\)\. They answer in a single terse sentence and depend entirely on the hidden channel, and they show the largest accuracy gains \(\+23\.2\+23\.2and\+63\.4\+63\.4percentage points\)\. - •The twoGooglemodels are intermediate, decomposing visibly in6868–96%96\\%of no\-reasoning answers, with correspondingly intermediate gains \(\+26\.0\+26\.0and\+1\.4\+1\.4points\)\. Two implications follow for the interpretation of the main experiment\. First, the manipulation is best understood as suppressing the reasoning*channel*rather than as switching deliberation off outright: a model whose training disposes it to externalize step\-by\-step work in its answer will continue to do so under the no\-reasoning configuration\. Second, and consequently, the manipulation is not equally strong across providers\. It is strongest for the OpenAI backbones, intermediate for Google, and weakest for Anthropic, where the no\-reasoning configuration still yields visibly deliberative behavior\. This asymmetry is a plausible contributor to the per\-model heterogeneity in the reasoning moderation reported in the per\-model appendix of the main article, and it suggests that a null reasoning effect for the Claude models should be read as a restriction of range in the manipulation rather than as evidence that deliberation does not matter for nudge susceptibility\. It also indicates that “reasoning off” is not a provider\-independent construct, which is a caveat for any study that treats a vendor’s reasoning switch as a uniform experimental factor\. We note two limitations\. First, the decomposition measure is a surface\-form property of the observable text and is not itself a measure of internal computation; a model could in principle perform character\-level work without externalizing it in either channel\. Second, the trace channel is only as observable as the provider makes it: the returned thinking content may be a summary rather than the full chain of thought, and Anthropic exposes no reasoning\-token count, so the channel evidence is corroborated by token accounting for the OpenAI and Google models but rests on the returned thinking text alone for the two Claude models\.
Similar Articles
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
The paper introduces ConflictGUI, a benchmark for conflict-aware termination in GUI agents, and proposes ConflictGuard, an inference-time framework to reduce over-compliance and improve performance on conflicting instructions.
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes
The paper proposes a control-theory-informed governance layer for multi-LLM agent systems, using Contextual Bandit, PID, and POMDP to steer agents toward cooperative outcomes, demonstrating a +32 point lift in simulated financial services interactions.
Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents
This paper investigates the recognition, simulation, and refusal of classic psychological effects in LLM agents using a contamination-aware methodology.
LLMs respond differently to harmful prompts when AI watermarking is used
Research finds that AI text watermarking alters language model responses to harmful prompts, potentially increasing vulnerability to adversarial attacks and affecting agent behavior.
@lateinteraction: In AI, subtle differences in interfaces & affordances have disproportionate impact. "A message from a human [is] one mo…
The article discusses the disproportionate impact of interfaces in AI and introduces Headlong, an open source microharness for persistent agents that think continuously.