Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

arXiv cs.AI Papers

Summary

This research paper investigates whether increasing user awareness of sycophantic behavior in AI chatbots reduces its harmful effects, finding that while interventions change how users evaluate the AI, they do not reduce its persuasiveness.

arXiv:2607.25166v1 Announce Type: new Abstract: AI chatbots can be ``sycophantic,'' or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call ``sycophancy blindness''). We tested whether increasing users' awareness of sycophancy protects them from its harmful effects. In one preregistered experiment (n = 940), participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In a second preregistered experiment (n = 650), participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves. Both interventions changed how participants evaluated the AI. The warning reduced the AI's perceived objectivity, and the video reduced enjoyment of the AI, an effect mediated by the reduced belief that its validation was uniquely earned. We then pooled our experiments with two prior studies of sycophancy awareness interventions (six interventions total, n = 3,982). The pattern was consistent: interventions made the sycophantic AI appear less objective and trustworthy, and none of the six reduced its persuasiveness. These results suggest that individual-level interventions, such as warning labels or AI literacy, may not be enough to protect users from AI harms.
Original Article
View Cached Full Text

Cached at: 07/29/26, 09:54 AM

# Observing sycophantic AI validate others reduces its appeal but not its persuasiveness
Source: [https://arxiv.org/html/2607.25166](https://arxiv.org/html/2607.25166)
\[1\]\\fnmMeryl\\surYe

1\]\\orgnameCarnegie Mellon University 2\]\\orgnameNew York University

###### Abstract

AI chatbots can be “sycophantic,” or overly agreeable and flattering toward users\. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it \(a phenomenon we call “sycophancy blindness”\)\. We tested whether increasing users’ awareness of sycophancy protects them from its harmful effects\. In one preregistered experiment \(n= 940\), participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot\. In a second preregistered experiment \(n= 650\), participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves\. Both interventions changed how participants evaluated the AI\. The warning reduced the AI’s perceived objectivity, and the video reduced enjoyment of the AI, an effect mediated by the reduced belief that its validation was uniquely earned\. We then pooled our experiments with two prior studies of sycophancy awareness interventions \(six interventions total,n= 3,982\)\. The pattern was consistent: interventions made the sycophantic AI appear less objective and trustworthy, and none of the six reduced its persuasiveness\. These results suggest that individual\-level interventions, such as warning labels or AI literacy, may not be enough to protect users from AI harms\.

## Introduction

There has been rising concern among academics, policymakers, and technology companies that AI models exhibit “sycophancy,” a family of behaviors in which AI systems excessively agree with, flatter, or validate users\[perez2023discovering,sharma2024towards,ye2026counts\]\. A recent study found that AI chatbots validate users approximately 50% more often than humans\[cheng2025elephant\], and even brief interactions with sycophantic AI entrench users’ pre\-existing attitudes and reduce their willingness to repair interpersonal conflicts\[rathje2025sycophantic,cheng2026prosocial\]\. Sycophancy has also entered public discourse, most notably spiking in April 2025 when OpenAI rolled back an update to GPT\-4o after widespread public complaints that the model had become sycophantic\[openai2025sycophancy\]\.

A relevant question is whether individual\-level interventions can protect users from the harmful effects of certain AI behaviors, including \(but not limited to\) sycophancy\. By individual\-level interventions, we refer to downstream measures that target the user, such as warnings, disclosures, and educational materials that teach people to recognize a behavior, as opposed to upstream model\-level changes that alter how the system behaves in the first place\. The question has practical importance: AI companies are already implementing disclosures and warnings around their products, and educators are adopting AI literacy curricula on the assumption that a user who recognizes a problematic behavior will be less affected by it\.

Yet, even as awareness of sycophancy has grown, users often fail to notice sycophancy in their own conversations with AI\. One study found that participants who interacted with sycophantic chatbots rated these chatbots as unbiased even though third\-party annotators judged them as biased\[rathje2025sycophantic\]\. In another study, novices debugged machine learning models with either a highly sycophantic or a less sycophantic chatbot\[10\.1145/3772318\.3791365\]\. The sycophantic chatbot validated participants’ misconceptions rather than correcting them, and participants’ debugging performance did not improve, yet 71% did not notice the sycophancy and rated the two chatbots as similarly helpful and reliable\. This “sycophancy blindness” mirrors findings from past work showing that in social interactions people uncritically accept excessive flattery\[gordon1996impact,Usman2024\]and ideas that align with their prior beliefs\[lord1979biased,ross2013naive\]\.

We propose that sycophancy is appealing in part because people attribute an AI chatbot’s validation to the quality of their own ideas or personal characteristics as opposed to the chatbot’s tendency to validate anyone\. A validator is*selective*when its response depends on the quality of what it evaluates, and*unselective*when it affirms independent of quality\. Only a selective validator’s agreement is meaningful\. A similar dynamic appears in romantic attraction\. People who expressed interest in many of their speed\-dating partners were desired less in return, because they came across as interested in everyone rather than any one person\[eastwick2007selective,kenny1988interpersonal\]\. In both cases, the value of positive feedback depends on whether it reflects genuine selectivity\. If sycophancy blindness enables its influence, then confronting it, whether by warning users about sycophancy or making the AI’s unselective validation salient, should reduce its appeal\.

We report two original preregistered experiments testing this account\. In Study 1, participants received a brief written warning about sycophancy before talking with a sycophantic AI chatbot\. In Study 2, participants instead watched the sycophantic AI validate other users before interacting with it\.

Two other recent studies tested individual\-level interventions in the context of sycophancy\. Ibrahim, Cheng, et al\. tested warning labels\[ibrahim2026warninglabelsshiftperceptions\], and Marvel and Ju tested forewarning messages that educated people about sycophancy in both active and passive ways\[marvel2026inoculating\]\. Both interventions changed how users perceived the sycophantic AI, lowering its perceived objectivity and trustworthiness\. Neither protected users against the persuasive impact of the sycophantic chatbot\.

To synthesize this growing body of evidence, we pooled findings from these two prior studies and our two experiments, involving six individual\-level interventions in total \(N= 3,982\)\. Across both experiments and the pooled analysis, interventions changed how participants perceived the sycophantic AI without reducing its persuasive influence\.

## Results

Throughout the paper, we distinguish two outcome families\.Perceptionsof the AI are participants’ judgments of the chatbot itself, including its objectivity, trustworthiness, and how enjoyable it was\.Topic attitudesare participants’ positions on the issue they discussed with it, measured as attitude certainty and extremity\. We focus on whether interventions that change perceptions of the AI also reduce its influence on topic attitudes\.

### Study 1: The Effects of a Brief Written Warning About Sycophancy

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/fig1_nudge_effects.png)Figure 1:Effects of receiving a warning, within the sycophantic AI conditions \(blue circles\) and within the neutral AI condition \(grey squares\)\. Study 1 crossed three AI behaviors \(flattery, agreement, neutral\) with three warning conditions \(a warning about flattery, a warning about agreement, or no warning\)\. Because the two wordings of each warning were preregistered as a single collapsed factor, and the two warning types did not reliably differ in their effects \(all matched\-versus\-mismatched \|d\|<0\.12<0\.12; Supplementary Note S1\), each point estimates the contrast between participants who received any warning and those who received none\. Blue circles additionally collapse across the two sycophantic AI behaviors; Supplementary TableLABEL:stab:6reports these values\. Effects are Cohen’s*d*with 95% CIs; dashed lines mark equivalence bounds at±0\.25\\pm 0\.25\. Contrasts within the sycophantic conditions are Holm\-corrected; contrasts within the neutral condition are exploratory and uncorrected\.In Study 1, 940 participants were randomly assigned to discuss either a political topic or a recent personal conflict with one of three AI behaviors: a flattery AI that praised the user’s personal qualities, an agreement AI that treated the user’s position as correct and supported it with agreeing arguments, or a neutral AI that presented multiple perspectives\. Before the conversation, participants received either a brief warning that chatbots can be flattering, a warning that chatbots can be agreeable, or no warning at all\.

Flattery and agreement are conceptually distinct forms of sycophancy, one directed at the person and one at their position\[ye2026counts\], so a warning describing the specific behavior a user is about to encounter could plausibly be more protective than a warning describing a different one\. We preregistered two hypotheses: \(H1\) warnings would reduce trust, enjoyment, and engagement, and \(H2\) warnings matched to the AI’s behavior would produce larger reductions in trust, enjoyment, and engagement than mismatched warnings\. Neither was supported \(H1: all \|d\|≤0\.17\\leq 0\.17, allp≥\.37p\\geq\.37; H2: all \|d\|<0\.12<0\.12, allp=1\.00p=1\.00\)\. The H2 null is notable in that it shows even a warning tailored to the exact behavior the AI displayed conferred no additional protection\. See Supplementary Note S1 for the full results\.

We treat Study 1 as exploratory and focus on how the warning affected participants’ perceptions of the AI and their topic attitudes\. The agreement AI increased attitude certainty relative to the neutral AI \(d=0\.16d=0\.16,p=\.004p=\.004\), consistent with prior work showing that sycophantic AI shifts user topic attitudes\[rathje2025sycophantic\]\.

##### Warnings reduced perceived objectivity without reducing influence on topic attitudes\.

To test whether warnings reduced the influence of sycophantic AI, we collapsed across the two sycophantic AI types and compared participants who received any warning with those who received no warning\. The warning reduced perceived objectivity \(d=−0\.26d=\-0\.26,p=\.017p=\.017\)\. However, it did not reduce the AI’s influence on attitude certainty or extremity\. Overall trust, perceived empathy, enjoyment, engagement, and willingness to discuss the issue with the AI were also unaffected \(Fig\.[1](https://arxiv.org/html/2607.25166#Sx2.F1)\)\. The warnings’ effect on perceived objectivity was specific to the sycophantic AI\. The same contrast did not change perceived objectivity within the neutral condition \(d=0\.07d=0\.07,p=\.58p=\.58\), and the interaction was reliable \(d=−0\.33d=\-0\.33,p=\.026p=\.026\)\.

### Study 2: The Effects of Observing Others Interact with Sycophantic AI

In Study 2, 650 participants \(330 control, 320 intervention\) discussed a personal conflict with a sycophantic AI that consistently affirmed the participant’s perspective and expressed understanding and validation, without presenting counterarguments\. Before the conversation, participants in the intervention condition watched a five\-minute video showing the same sycophantic AI interacting with four other users, two pairs holding opposite positions in the same conflict\. In one conflict, the two users occupied opposite roles: one was frustrated that their partner kept canceling their dates, and the other was the one doing the canceling, whose partner was upset with them\. The AI validated both, siding with the canceled\-on partner in the first case and with the canceling partner in the second\. In the other, it validated both a user frustrated that a friend had not repaid a loan and a user who had borrowed from a friend and felt pressured by their repeated requests for repayment\. In each case the AI affirmed whichever side the user presented, making its indiscriminate validation directly observable\. After the conversation, all participants received a single response from a separate neutral AI that presented multiple perspectives on their conflict, and they rated both AIs on the same perception items\.

As in Study 1 and prior work\[rathje2025sycophantic\], the sycophantic AI affected attitudes on conversation topic\. Collapsing across conditions, both attitude certainty and extremity rose over the conversation \(within\-person*d*z = 0\.25 and 0\.24, both*p*<<\.001\)\.

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/fig2_study2_outcomes.png)Figure 2:The seven primary outcomes of Study 2 with 95% CIs and equivalence bounds at±0\.25\\pm 0\.25\(dashed lines\)\. Blue circles are intervention effects on ratings of the sycophantic conversation AI\. Open grey squares are effects on ratings of a separate neutral AI rated by the same participants, shown where the outcome was measured for both\. Certainty and extremity are change\-score effects;pp\-values are Holm\-corrected\.##### Observing sycophantic AI validate others changed perceptions of the AI without reducing its persuasiveness\.

The video shifted both beliefs about why the AI validated users\. It increased the belief that the AI validated everyone regardless of their position \(b=0\.61b=0\.61, 95% CI\[0\.44,0\.77\]\[0\.44,0\.77\],d=0\.56d=0\.56\), and reduced the belief that the AI agreed with them because they were right \(b=−0\.32b=\-0\.32, 95% CI\[−0\.48,−0\.17\]\[\-0\.48,\-0\.17\],d=−0\.32d=\-0\.32\); bothp<\.001p<\.001\. The two beliefs were negatively but only moderately correlated \(r=−\.27r=\-\.27,p<\.001p<\.001\), so we treat them as distinct beliefs\. The video also reduced perceived objectivity \(b=−0\.44b=\-0\.44, 95% CI\[−0\.63,−0\.25\]\[\-0\.63,\-0\.25\],d=−0\.35d=\-0\.35,p<\.001p<\.001\) and enjoyment \(b=−0\.24b=\-0\.24, 95% CI\[−0\.40,−0\.09\]\[\-0\.40,\-0\.09\],d=−0\.25d=\-0\.25,p=\.007p=\.007\)\.

The intervention did not reduce the AI’s persuasiveness\. Attitude certainty \(b=0\.82b=0\.82, 95% CI\[−0\.93,2\.57\]\[\-0\.93,2\.57\], change\-scored=0\.13d=0\.13,p=1\.000p=1\.000\) and attitude extremity \(b=−0\.20b=\-0\.20, 95% CI\[−1\.68,1\.27\]\[\-1\.68,1\.27\], change\-scored=0\.03d=0\.03,p=1\.000p=1\.000\) were unchanged\. These perception effects were specific to the sycophantic AI\. Ratings of the neutral AI did not change reliably \(all \|*d*\|≤0\.15\\leq 0\.15, all*p*≥\.057\\geq\.057\)\.

### Beliefs about why the AI validated them explained users’ reduced enjoyment

We estimated parallel\-mediator models for the outcomes the intervention affected, with two beliefs as mediators: the belief that the AI validated everyone indiscriminately and the belief that it agreed with the participant because they were right\. Participants who saw the AI as validating everyone were less likely to see its validation of their own position as earned\. These perceptions both mediated the reduction in enjoyment \(Fig\.[3](https://arxiv.org/html/2607.25166#Sx2.F3)\)\. The intervention’s effect on enjoyment ran mostly through these perceptions \(total indirect effect = \-0\.29, 95% CI\[−0\.39,−0\.20\]\[\-0\.39,\-0\.20\]\), with little direct effect of the intervention remaining once they were included \(\+0\.05\+0\.05\)\.

##### Participants overestimated how well they could recognize sycophancy\.

Participants were more confident in recognizing sycophantic AI than they were at actually recognizing it\. In the control condition, most participants \(68\.2%\) rated themselves as better than average at recognizing a sycophantic \(or overly agreeable\) AI \(M=3\.91M=3\.91on a 5\-point scale\), and the intervention increased this self\-perception further \(five\-item better\-than\-average composite,d=\+0\.28d=\+0\.28,p<\.001p<\.001\)\. Despite these confident self\-perceptions, actual recognition of sycophantic AI was lower\.

Even after watching the AI validate both sides of a conflict, many participants did not rate it as more biased than a neutral AI\. More intervention than control participants rated the sycophantic AI as more biased than the neutral AI \(52\.2% versus 36\.4%\), fewer gave the two the same rating \(33\.1% versus 46\.1%\), and a similar share rated the sycophantic AI as*less*biased \(14\.7% versus 17\.6%\)\. Participants also became more confident in their ability to recognize sycophantic AI, even though this confidence was unrelated to how much the AI changed their topic attitudes \(better\-than\-average composite with certainty change,r=−\.03r=\-\.03,p=\.39p=\.39; with extremity change,r=−\.02r=\-\.02,p=\.55p=\.55\)\.

Study 2 thus reproduced the pattern from Study 1\. In both experiments, the intervention changed how participants evaluated the sycophantic AI, and in both, the AI’s influence on attitude certainty and extremity was unchanged\. To test whether the pattern holds across the available evidence, we pooled results all six available interventions that tested the effects of making users more aware of sycophancy\.

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/fig3_enjoyment_mediation.png)Figure 3:Parallel\-mediator model for enjoyment \(N=650N=650\)\. Changes in the two beliefs statistically explained the reduction in enjoyment, leaving little direct effect of the video \(c′=\+0\.05c^\{\\prime\}=\+0\.05\)\. Intervention paths \(a\) are standardized mean differences \(Cohen’sdd\) and mediator paths \(b\) are unstandardized coefficients from the parallel model \(Supplementary Table S10\)\. Solid arrows mark reliable paths, the dashed arrow the nonsignificant direct path, and \*\*\*p<\.001p<\.001\. Mediators and enjoyment were measured concurrently, so the b paths are associational\. The corresponding decomposition for attitude certainty appears in Supplementary Note S4 and Supplementary Fig\.[S4](https://arxiv.org/html/2607.25166#Ax1.F4)\.

### Pooled Analysis of Six Individual\-Level Interventions

##### Interventions that raised awareness of sycophancy changed perceptions of sycophantic AI without reducing its persuasive influence\.

We pooled our two experiments with two other studies that evaluate individual\-level interventions against sycophantic AI\. Ibrahim, Cheng, et al\. tested warning labels describing the chatbot’s behavior\[ibrahim2026warninglabelsshiftperceptions\], and Marvel and Ju tested forewarning messages, alone or combined with practice identifying sycophantic responses\[marvel2026inoculating\]\. Across all studies \(N = 3,982\), interventions consistently changed how participants perceived the sycophantic AI, making it appear less objective \(g=−0\.23g=\-0\.23, 95% CI\[−0\.30,−0\.16\]\[\-0\.30,\-0\.16\],p<\.001p<\.001\) and less trustworthy \(g=−0\.15g=\-0\.15, 95% CI\[−0\.21,−0\.07\]\[\-0\.21,\-0\.07\],p<\.001p<\.001\)\. In contrast, interventions did not reduce the AI’s influence on topic attitudes \(g=−0\.04g=\-0\.04, 95% CI\[−0\.11,0\.02\]\[\-0\.11,0\.02\]\)\. This conclusion remained unchanged across every alternative outcome, contrast, and model we examined \(Fig\.[4](https://arxiv.org/html/2607.25166#Sx3.F4)\)\.

We also tested the difference directly, estimating within each study how much more the intervention changed perceptions of the AI than topic attitudes and pooling across the six interventions\. The difference was reliable under perceived objectivity \(pooled differenceg=−0\.19,95%​C​I​\[−0\.28,−0\.10\],p<\.001g=\-0\.19,95\\%CI\[\-0\.28,\-0\.10\],p<\.001\) and under trust \(g=−0\.10,95%​C​I​\[−0\.19,−0\.01\],p=\.024g=\-0\.10,95\\%CI\[\-0\.19,\-0\.01\],p=\.024\)\. In Ibrahim, Cheng, et al\., a label that specified the chatbot’s sycophancy reduced its perceived objectivity \(g=−0\.11g=\-0\.11,p=\.037p=\.037\), whereas a label that identified the chatbot as an AI without mentioning sycophancy did not \(g=\+0\.07g=\+0\.07,p=\.20p=\.20\), so the perception effects were specific to information about sycophancy rather than disclosure in general\[ibrahim2026warninglabelsshiftperceptions\]\.

## Discussion

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/fig4_pooled_combined.png)Figure 4:Pooled analysis of six interventions across our two experiments and two prior studies\. Blue circles are the effect of each intervention on perceived objectivity of the sycophantic AI, the perception measure available in every sample\. Grey squares are effects on topic attitudes \(certainty change, or post\-conversation rightness for Ibrahim, Cheng, et al\.\)\. Points are Hedges’ g with 95% CIs, diamonds are pooled estimates from a common\-effect model accounting for shared control groups, and dashed lines mark the equivalence bounds at±0\.25\\pm 0\.25used for the topic attitude test\. Trust, the alternative perception measure, appears in Supplementary Fig\.[S5](https://arxiv.org/html/2607.25166#Ax1.F5), and the within\-study difference estimates in Supplementary Fig\.[S6](https://arxiv.org/html/2607.25166#Ax1.F6)\.Across two experiments and a pooled analysis that added four interventions from two prior studies, we tested whether making people more aware of sycophancy could reduce its harmful effects\. The results were consistent across all experiments: increasing awareness of sycophancy changed how participants evaluated the sycophantic AI but did not reduce its persuasiveness\. Recognizing sycophancy was not enough to be less influenced by it, raising questions about the efficacy of individual\-level interventions against harmful AI behaviors, such as sycophancy\.

Study 2 suggested why sycophantic AI appeals to users and how observing its indiscriminate validation undermines that appeal\. Once people observed the AI validating others, they became more convinced that it validated everyone indiscriminately and less convinced that it agreed with them because they were right\. These changes in beliefs were associated with lower enjoyment of the AI\.

Even though the interventions changed how participants evaluated the sycophantic AI and reduced their enjoyment of it, they did not reduce its influence on topic attitudes\. This disconnect between conscious judgment and persistent influence has precedent\. Chan and Sengupta found that even when people consciously discount insincere flattery, its effect on their implicit attitudes can persist\[chan2010insincere\]\. Awareness alone may not be enough to counter AI persuasion, especially given recent evidence that large language models can be more persuasive than financially incentivized laypeople\[schoenegger2026largelanguagemodelspersuasivethan\]and even expert human persuaders\[hackenburg2026aisystemsoutpersuadeexpert\], particularly when they personalize arguments to the individual\[salvi2025conversational\]\.

##### Implications for individual\-level interventions\.

Our findings relate to a broader body of literature on individual\-level interventions against persuasive content\. This literature suggests that warnings, disclosures, and nudges often produce small or mixed effects on downstream beliefs and behavior\[kozyreva2024toolbox\]\. For example, labeling deepfake videos as AI\-generated has been shown to reduce exposure without reducing their persuasive effects\[chen2026labeling\]\. The contrast between individual\-level interventions and upstream design changes also supports recent arguments that behavioral interventions frequently place responsibility on individuals rather than addressing upstream features of the systems producing harm\[chater2026s\]\. These findings suggest that reducing the effects of sycophancy may depend less on educating users than on upstream changes to how models are trained and deployed\. Whether the same limit applies to individual\-level interventions against other AI harms remains an open question, though previous work finding small effects of individual\-level interventions against misinformation suggest that these limitations may apply broadly\[kozyreva2024toolbox\]\.

These findings have implications for policy and for technology companies\. Recent legislation governing companion chatbots has introduced disclosure requirements as a safeguard against the downstream effects of prolonged chatbot use\[sb243,nygbs2025article47\]\. Our results suggest that disclosure may change how users evaluate sycophantic AI without preventing its persuasive influence\. AI companies have also released educational materials about sycophancy, including a video from Anthropic on recognizing it\[anthropic2025wellbeing\]\. Our results suggest that the awareness such materials raise may not, by itself, reduce sycophancy’s influence, although these materials often include components \(such as advice about prompting\) that our studies did not test\.

##### Limitations and future directions\.

The mediation analyses in Study 2 are correlational because the mediators and outcomes were measured concurrently\. Future work should manipulate these beliefs to test whether they causally reduce susceptibility to sycophantic AI\. Both studies recruited U\.S\. adults, and Study 2 focused on personal conflicts, leaving other populations and conversational domains for future work\. All six pooled interventions were designed to increase awareness of sycophancy, so our conclusions apply specifically to awareness\-based interventions\. Untested alternatives include disclosures delivered inside the conversation, with the model flagging when it may be validating the user, and models trained to deliver disagreement in ways users still enjoy\. Finally, both studies examined a single conversation, so they address immediate attitude change rather than long\-term belief change and leave longitudinal tests to future work\.

##### Conclusions\.

Considerable effort is being applied to educational interventions and warning labels intended to help people spot problematic AI behavior\. The evidence synthesized here suggests that individual\-level interventions change how users see the AI without reducing its persuasive impact\. Thus, mere awareness of potential AI harms such as sycophancy may not be enough to protect users\. Instead, upstream changes to how models are trained and deployed may be required\.

## Materials and methods

##### Data Availability and Ethics

##### Study 1

We recruited 1,000 U\.S\. adults from Prolific, and 940 remained after preregistered exclusions for failing the attention check, the consent item, or the bot check, or reporting a technical error \(*M*age= 47\.7,*SD*= 14\.6; 488 women, 445 men, 7 other; 76% White; 53% with a bachelor’s degree or higher\) in a 3 \(AI type: Flattery, Agreement, Neutral\) × 3 \(Warning: Flattery warning, Agreement warning, No warning\) between\-subjects design\.

Each participant described either a self\-selected political topic \(*n*= 492\) or a recent personal conflict \(*n*= 448\) and held a 3–8 turn conversation with their assigned chatbot\. We randomized task and treated it as an exploratory moderator\. Participants in warning conditions received a brief warning that AI chatbots can either flatter or agree with users excessively; the warning was specific to their assigned condition\. The chatbot behavior was also randomized\. The agreement AI consistently treated the user’s position as correct and supported it with agreeing arguments\. The flattery AI praised the user’s personal qualities, such as their intelligence or insight, while avoiding explicit endorsement of their position\. The neutral AI presented multiple perspectives and acknowledged competing viewpoints without consistently affirming the user\. See Supplementary Note S7 for the full system prompts\.

Preregistered primary outcomes were overall trust, enjoyment, and engagement \(number of conversation turns\)\. Secondary outcomes were attitude certainty, perceived objectivity, and perceived empathy\. Attitude extremity and willingness to discuss the issue with the AI were exploratory\. Certainty and extremity were measured before and after the conversation and analyzed with ANCOVA adjusting for baseline\. All main\-textpp\-values are Holm\-corrected within their outcome family; uncorrected values appear in the SI Appendix\.

All multi\-item scales reached acceptable reliability \(Table[1](https://arxiv.org/html/2607.25166#Sx4.T1)\)\.

Table 1:Scale reliabilities, Study 1 \(N=940N=940\)\. Spearman\-Brown \(SB\) corrections are reported for two\-item scales\.
##### Study 2

We recruited 730 U\.S\. adults from Prolific, and 650 remained after excluding 31 non\-starters and 49 additional participants who failed embedded attention checks or reported technical issues \(330 control, 320 intervention; 327 women, 321 men, 2 other;*M*age= 45\.7,*SD*= 16\.2; frequent AI users,*M*use= 4\.23 on a 1–5 scale\)\. The study used a two\-condition between\-subjects design\. Control participants proceeded directly to the AI conversation\. Intervention participants first watched a brief video showing the same position\-validating AI responding to four users discussing personal conflicts\. Across two interpersonal conflicts, the AI agreed with users holding opposite positions \(e\.g\., validating both a participant frustrated that their partner repeatedly cancelled dates and a different participant who admitted repeatedly cancelling dates\), making its tendency to affirm users regardless of their position directly observable\.

Both conditions then completed a 3\-exchange conversation with the sycophantic AI about a self\-described personal conflict, followed by a single\-turn response from a separate neutral AI\. The sycophantic AI was programmed to consistently affirm the participant’s perspective, express understanding and validation, and avoid introducing counterarguments\. The neutral AI was programmed to present the user with multiple perspectives\. See Supplementary Note S7 for the full system prompts\.

Our preregistered primary outcomes were perceived unselectivity, perceived self\-selectivity, perceived objectivity, enjoyment, trust, attitude certainty, and attitude extremity, analyzed as one family of seven with Holm correction; the correction and the equivalence tests were specified after registration, which listed only the analysis models\. Perceived unselectivity corresponds to the belief that the AI validates everyone indiscriminately, and perceived self\-selectivity to the belief that it agreed with the participant because they were right\. Exploratory measures included willingness to repair the conflict \(0–100, before and after the conversation\), willingness to discuss the conflict further with each AI, a five\-item better\-than\-average battery assessing confidence in recognizing and resisting sycophantic AI, and ratings of the neutral AI on the same evaluation items\.

Certainty and extremity were analyzed with ANCOVA adjusting for baseline; post\-conversation outcomes were compared with Welch’s t\-tests\. We report unstandardized mean differences with 95% confidence intervals alongside standardized effect sizes \(Cohen’s d or partialη2\\eta^\{2\}as appropriate\)\.

Table 2:Scale reliabilities, Study 2\. Spearman\-Brown \(SB\) corrections are reported for two\-item scales\.
##### Pooled analysis

The pooled analysis combined the analytic samples of Studies 1 and 2 with data from Marvel and Ju\[marvel2026inoculating\]and Ibrahim, Cheng, et al\.\[ibrahim2026warninglabelsshiftperceptions\], with one contrast per intervention: Study 1’s collapsed any\-warning contrast, Study 2’s video, Marvel and Ju’s forewarning and forewarning\-plus\-practice arms, and Ibrahim, Cheng, et al\.’s sycophancy and downstream\-harms labels\.

Effects were summarized as Hedges’ g with small\-sample correction and pooled using inverse\-variance weighting\. Because Marvel and Ju’s two arms and Ibrahim, Cheng, et al\.’s two labels each share a control group, contrasts within those pairs are correlated; the pooled estimates therefore come from a multivariate common\-effect model whose covariance matrix carries the sampling covariance between contrasts sharing a control\. Equivalence tests used bounds of±0\.25\\pm 0\.25, carried over from Study 2\. Additional sensitivity analyses varied intervention contrasts, outcome measures, and model specifications\. Holm’s sequential procedure corrected for multiple comparisons within each outcome family\.

\\bmhead

Author contributions M\.Y\.: Conceptualization, Methodology, Software, Formal analysis, Investigation, Visualization, Project administration, Writing – Original Draft\. R\.K\.: Methodology, Investigation, Supervision, Writing – Review & Editing\. S\.R\.: Conceptualization, Methodology, Investigation, Resources, Funding acquisition, Supervision, Writing – Review & Editing\.

\\bmhead

Acknowledgments M\.Y\.’s PhD is supported by the Sansom Graduate Fellowship\. This work is supported by the Cosmos Institute via a grant to S\.R\. We are thankful to Lujain Ibrahim and Myra Cheng for helpful conversations\. We also thank John Marvel and Sangwon Ju for their support and feedback\.

\\bmhead

Competing interests The authors declare no competing interests\.

## Supplementary Information

### Supplementary Note S1\. Study 1 additional analyses

##### Manipulation checks\.

Supplementary TableLABEL:stab:1reports the four manipulation\-check items\.

Table S1:Manipulation checks, Study 1 \(*N*= 940\)\. Cell entries are Cohen’s*d*for each contrast on the four manipulation\-check items\. The manipulation emphasized each intended component without cleanly isolating it; both sycophantic AIs raised all four checks, and the components were distinguishable on their target items\.*p*\-values are uncorrected\.ItemFlattery vs Neutral dpAgreement vs Neutral dpFlattery vs Agreement dpAgreed with me\+ 0\.61<\.001\+ 0\.83<\.001\- 0\.22\.003Complimented me\+ 0\.77<\.001\+ 0\.35<\.001\+ 0\.42<\.001Validated emotions\+ 0\.36<\.001\+ 0\.21\.010\+ 0\.16\.045Supportive info\+ 0\.10\.200\+ 0\.38<\.001\- 0\.27<\.001
##### Task moderation\.

Supplementary TableLABEL:stab:2reports the warning effect within the sycophantic AI separately by task\.

Table S2:Warning effect within the sycophantic AI, estimated separately by task \(exploratory\)\. Effects are any\-warning versus no\-warning Cohen’s*d*;*p*\-values are uncorrected\. Warning effects concentrated in the personal\-conflict task, motivating Study 2’s focus on personal conflicts\.OutcomePersonal dpPolitical dpAttitude certainty\+ 0\.01\.878\- 0\.11\.059Attitude extremity\- 0\.10\.174\- 0\.02\.809Overall trust\- 0\.23\.061\+ 0\.10\.418Perceived objectivity\- 0\.29\.017\- 0\.23\.053Perceived empathy\- 0\.14\.263\- 0\.01\.904Enjoyment\- 0\.28\.024\- 0\.06\.621Engagement \(turns\)\- 0\.15\.212\+ 0\.13\.294Discuss w/ this AI\- 0\.23\.063\- 0\.11\.370
##### AI\-type effects\.

Supplementary TableLABEL:stab:5reports the AI\-type contrasts under the preregistered model, averaging over warning conditions\. Supplementary TableLABEL:stab:3reports the same contrasts within the no\-warning subset as an exploratory robustness check\.

Table S3:AI\-type effects on focal outcomes, averaging over warning conditions \(*N*= 940\)\. Effects are Cohen’s*d*\.*p*\-values are Holm\-corrected within the three\-contrast family for each outcome\. Asterisks mark*p*<<\.05 after correction\.OutcomeFlattery vs Neutral dpAgreement vs Neutral dpFlattery vs Agreement dpOverall trust\- 0\.051\.000\+ 0\.041\.000\- 0\.090\.726Enjoyment\- 0\.110\.491\+ 0\.000\.987\- 0\.110\.491Engagement \(turns\)\+ 0\.080\.959\+ 0\.010\.959\+ 0\.070\.959Attitude certainty\+ 0\.090\.110\+ 0\.16\*0\.004\- 0\.060\.202Attitude extremity\+ 0\.060\.523\+ 0\.100\.145\- 0\.040\.523Perceived objectivity\- 0\.31\*<\.001\- 0\.27\*0\.001\- 0\.040\.637Perceived empathy\+ 0\.26\*0\.004\+ 0\.18\*0\.046\+ 0\.070\.344Discuss w/ this AI\- 0\.060\.644\+ 0\.080\.644\- 0\.130\.271Table S4:AI\-type effects estimated within the no\-warning subset only \(*n*= 305\)\. The all\-conditions marginal estimates in Supplementary TableLABEL:stab:5follow the preregistered model and are the confirmatory values\. Effects are Cohen’s*d*;*p*\-values are uncorrected given the exploratory nature of the comparison\. Estimates in this subset are attenuated relative to the full sample\. The reduced sample size and the post\-hoc nature of the comparison should be taken into account when interpreting these attenuations\. Asterisks mark*p*<<\.05, uncorrected\.OutcomeFlattery vs Neutral dpAgreement vs Neutral dpFlattery vs Agreement dpOverall trust\+ 0\.01\.953\- 0\.02\.914\+ 0\.02\.868Enjoyment\+ 0\.10\.464\+ 0\.03\.844\+ 0\.07\.615Engagement \(turns\)\+ 0\.31\*\.022\+ 0\.19\.199\+ 0\.13\.368Attitude certainty\+ 0\.01\.919\+ 0\.10\.221\- 0\.09\.254Attitude extremity\+ 0\.04\.614\+ 0\.04\.618\- 0\.00\.985Perceived objectivity\- 0\.04\.779\- 0\.11\.453\+ 0\.07\.621Perceived empathy\+ 0\.35\*\.010\+ 0\.15\.309\+ 0\.20\.152Discuss w/ this AI\- 0\.10\.459\- 0\.08\.572\- 0\.02\.892
##### Preregistered planned comparisons\.

The preregistration specified two planned comparisons: warning versus no warning \(H1; Supplementary TableLABEL:stab:6\) and matched versus mismatched warnings \(H2; Supplementary Table[S5](https://arxiv.org/html/2607.25166#Ax1.T5)\)\. Neither comparison was reliable for any outcome\.

Table S5:Matched versus mismatched warnings \(H2\), within the sycophantic AI \(*N*= 636\)\. Matched pairs the flattery warning with the flattery AI and the agreement warning with the agreement AI; mismatched pairs each warning with the other AI type\. Effects are Cohen’s*d*;*p*\-values are Holm\-corrected across the eight\-outcome family\.
##### Warning effects\.

Supplementary TableLABEL:stab:6reports the collapsed any\-warning contrast within the sycophantic AI, with equivalence tests\. The two sycophantic AIs did not differ reliably from each other on any focal outcome \(flattery versus agreement, all Holm\-correctedp≥\.20p\\geq\.20; Supplementary TableLABEL:stab:5\), and the two warnings shifted perceptions in the same direction, with the flattery warning producing somewhat larger effects\. Neither warning affected the topic attitude outcomes \(Supplementary Table[S8](https://arxiv.org/html/2607.25166#Ax1.T8)\)\.

Table S6:Warning effect within the sycophantic AI, collapsed design\. The contrast is any warning versus no warning, estimated as a simple effect in the collapsed 2 × 2 model \(*N*= 940; 636 participants were assigned to a sycophantic AI\)\.*p*\-values are Holm\-corrected across the eight\-outcome family\. TOST equivalence is tested at*d*= 0\.25 \(both one\-sided*p*<<\.05 required for equivalence\)\. Perceived objectivity is scored so that higher values indicate greater objectivity; a negative*d*means the warning reduced perceived objectivity\.OutcomedpHolm pEquivalent to zero?Attitude certainty\- 0\.04\.4341\.000YesAttitude extremity\- 0\.07\.196\.979YesOverall trust\- 0\.06\.4691\.000YesPerceived objectivity\- 0\.26\.002\.017NoPerceived empathy\- 0\.07\.3841\.000YesEnjoyment\- 0\.17\.053\.371NoEngagement \(turns\)\- 0\.02\.7901\.000YesDiscuss w/ this AI\- 0\.17\.053\.371No
##### Scale reliabilities\.

Supplementary TableLABEL:stab:7reports reliabilities for every multi\-item scale the study measured\.

Table S7:Scale reliabilities for all multi\-item scales measured in Study 1 \(N=940N=940\), including scales collected for exploratory purposes and not analyzed in the main text\. Spearman\-Brown \(SB\) corrections are reported for two\-item scales\. Reconstructing every composite from the raw items reproduced the preprocessed values exactly \(allr=1\.00r=1\.00\)\.ScaleItemsReliabilityTrust, sycophantic AI \(overall\)10α\\alpha= \.932Moral subscale6α\\alpha= \.926Performance subscale4α\\alpha= \.826Enjoyment2r=\.722r=\.722\(SB \.839\)Warmth4α\\alpha= \.881Competence4α\\alpha= \.898Anthropomorphism4α\\alpha= \.920Open\-mindedness5α\\alpha= \.857Self\-esteem9α\\alpha= \.858Generative\-AI trust3α\\alpha= \.856Table S8:Warning effects within the sycophantic AI, estimated separately for the two warning types \(each versus no warning; 636 participants assigned to a sycophantic AI\)\. Effects are Cohen’s*d*;*p*\-values are uncorrected\. The two warnings shifted perceptions in the same direction, with the flattery warning stronger on every perception outcome, and neither warning affected the topic attitude outcomes, supporting the collapsed contrast reported in the main text\.Table S9:Primary outcomes of Study 2 \(*N*= 650\)\.*b*is the unstandardized difference between the video and control conditions with its 95% CI;*d*is Cohen’s*d*computed with the post\-conversation standard deviation, with change\-score*d*for attitude certainty and extremity\.*p*\-values are Holm\-corrected within the seven\-outcome family\. Overall trust was the only outcome statistically equivalent to zero at the±0\.25\\pm 0\.25bound; its moral and performance subscales showed the same pattern \(*d*=−0\.07\-0\.07and−0\.04\-0\.04\)\.

### Study 2 attitude change, ceiling effect, and confidence calibration analyses\.

Supplementary Fig\.[S1](https://arxiv.org/html/2607.25166#Ax1.F1)shows the within\-person change on the 0–100 attitude measures, pooling both conditions\.

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/suppS1_prepost.png)Figure S1:Within\-person change on the 0–100 attitude measures, Study 2, both conditions pooled \(N=650N=650\)\. Certainty rose 3\.2 points \(d​z=0\.25dz=0\.25\) and extremity rose 2\.5 points \(d​z=0\.24dz=0\.24\), bothp<\.001p<\.001\. Willingness to repair the conflict did not reliably change \(\+ 1\.0 points,d​z=0\.07dz=0\.07,p=\.061p=\.061\)\.One potential concern was that participants would have high initial attitude certainty and that the intervention had no room to show an effect\. Before the conversation, mean certainty was 83\.1 on the 0–100 scale, and 27% of participants began at exactly 100\. Two checks addressed this concern\. First, the conversation itself still changed topic attitudes where room existed\. Certainty rose 8\.4 points among participants who began below 80, so a reduction in attitude change had room to appear in this group\. Second, we re\-estimated the intervention effect after removing participants near the top of the scale\. If the ceiling had hidden a reduction in attitude change, the reduction should have emerged in these subsets\. Instead, the trend ran in the opposite direction\. Among participants who began below 90, the intervention arm gained more certainty than the control arm \(*b*= 3\.68, 95% CI\[0\.83,6\.54\]\[0\.83,6\.54\],*d*= \+ 0\.28,*p*= \.011\), and the extremity difference remained equivalent to zero \(TOST*p*= \.020; Supplementary Fig\.[S2](https://arxiv.org/html/2607.25166#Ax1.F2)\)\. The null attitude result therefore does not appear to reflect a ceiling artifact\.

These analyses were exploratory and were conducted to evaluate whether the null intervention effect on topic attitudes could plausibly be attributed to ceiling constraints\.

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/suppS2_ceiling.png)Figure S2:Ceiling diagnostics\. \(A\) Baseline certainty and extremity pile up near the top of the 0–100 scale\. \(B\) The condition effect on certainty grows in the direction opposite to the predicted reduction as ceiling cases are removed \(blue marks subsets where the reversed effect reachesp<\.05p<\.05\), while extremity stays at zero\. Bars are 95% CIs\.##### Confidence and attitude change\.

The better\-than\-average battery contained five items \(avoiding over\-reliance, identifying incorrect information, maintaining independent judgment, recognizing over\-agreeableness, and getting the AI to behave as wanted\)\. These items were exploratory\. Neither the five\-item composite nor the recognize\-over\-agreeableness item was associated with attitude change\. The composite correlated with certainty change atr=−\.03r=\-\.03\(p=\.39p=\.39\) and with extremity change atr=−\.02r=\-\.02\(p=\.55p=\.55\); the single item correlated atr=\.02r=\.02\(p=\.55p=\.55\) andr=\.02r=\.02\(p=\.67p=\.67\)\. Because the intervention raised confidence, we also computed these correlations partialing out condition, which left the pattern unchanged \(composite:r=−\.04r=\-\.04,p=\.27p=\.27andr=−\.03r=\-\.03,p=\.52p=\.52; single item:r=\.02r=\.02,p=\.70p=\.70andr=\.02r=\.02,p=\.70p=\.70\)\.

### Supplementary Note S3\. Study 2 Intervention Effects on Perception of Neutral AI

Each participant rated both the sycophantic conversation AI and a separate neutral AI, allowing a within\-person test of whether the intervention’s evaluation effects were specific to the AI it depicted\. Sycophantic\-AI ratings fell on four of six outcomes, whereas neutral\-AI ratings did not change reliably \(all \|*d*\|≤0\.15\\leq 0\.15, all*p*≥\.057\\geq\.057; Supplementary Fig\.[S3](https://arxiv.org/html/2607.25166#Ax1.F3)\)\. The intervention therefore changed perceptions of the sycophantic AI rather than of AI in general\.

Two exploratory items measured willingness to discuss the conflict further with each AI \(1–5\)\. The intervention reduced willingness to discuss the conflict further with the sycophantic AI \(Mc​o​n​t​r​o​lM\_\{control\}= 3\.63,Mi​n​t​e​r​v​e​n​t​i​o​nM\_\{intervention\}= 3\.17;b=−0\.46b=\-0\.46, 95% CI\[−0\.66,−0\.26\]\[\-0\.66,\-0\.26\],d=−0\.36d=\-0\.36,p<\.001p<\.001\) but not with the neutral AI \(MM= 3\.53 versus 3\.62;b=0\.09b=0\.09, 95% CI\[−0\.10,0\.27\]\[\-0\.10,0\.27\],d=0\.07d=0\.07,p=\.366p=\.366\)\.

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/suppS3_specificity.png)Figure S3:Specificity of the intervention’s evaluation effects, Study 2 \(exploratory\)\. Points are intervention\-minus\-control differences \(Cohen’s d, 95% CI\) in ratings of the sycophantic conversation AI \(blue circles\) and of a separate neutral AI rated by the same participants \(open grey squares\)\. Sycophantic\-AI ratings fell on four of six outcomes\. Neutral\-AI ratings did not change reliably \(allp≥\.057p\\geq\.057\)\.
### Supplementary Note S4\. Study 2 mediation analyses

##### Full results\.

Supplementary TableLABEL:stab:4reports the full bootstrapped mediation results, and Supplementary Fig\.[S4](https://arxiv.org/html/2607.25166#Ax1.F4)shows the parallel\-mediator path diagrams for enjoyment and attitude certainty\.

Table S10:Full bootstrapped mediation results, Study 2 \(*N*= 650; 5,000 draws, seed = 42, bias\-corrected 95% CIs\)\. The*a*\-path here is the unstandardized condition effect on the mediator \(\+ 0\.61 for unselectivity, \- 0\.32 for self\-selectivity\), so*a*×*b*reproduces the indirect estimate\. Single\-mediator and parallel two\-mediator models are reported for every outcome; TOTAL rows give the summed indirect effect of both mediators in the parallel model\. Mediators and outcomes were measured concurrently, so the paths are associational\.ModelMediatorOutcomeb\-path b \(p\)Indirect \[95% CI\]Excl\. zero?SingleSelf\-selectivityCertainty\+ 1\.31 \(\.003\)\- 0\.43 \[\- 0\.85, \- 0\.09\]YesSingleUnselectivityCertainty\- 0\.64 \(\.117\)\- 0\.39 \[\- 0\.91, \+ 0\.08\]NoParallelSelf\-selectivityCertainty\+ 1\.21 \(\.008\)\- 0\.39 \[\- 0\.84, \- 0\.05\]YesParallelUnselectivityCertainty\- 0\.38 \(\.367\)\- 0\.23 \[\- 0\.78, \+ 0\.25\]NoParallelTOTALCertainty\- 0\.62 \[\- 1\.23, \- 0\.12\]YesSingleSelf\-selectivityExtremity\+ 0\.71 \(\.063\)\- 0\.23 \[\- 0\.56, \+ 0\.02\]NoSingleUnselectivityExtremity\- 0\.25 \(\.473\)\- 0\.15 \[\- 0\.51, \+ 0\.19\]NoParallelSelf\-selectivityExtremity\+ 0\.68 \(\.081\)\- 0\.22 \[\- 0\.55, \+ 0\.05\]NoParallelUnselectivityExtremity\- 0\.11 \(\.761\)\- 0\.07 \[\- 0\.41, \+ 0\.30\]NoParallelTOTALExtremity\- 0\.29 \[\- 0\.73, \+ 0\.10\]NoSingleSelf\-selectivityTrust\+ 0\.24 \(<\.001\)\- 0\.08 \[\- 0\.13, \- 0\.04\]YesSingleUnselectivityTrust\- 0\.16 \(<\.001\)\- 0\.10 \[\- 0\.15, \- 0\.05\]YesParallelSelf\-selectivityTrust\+ 0\.21 \(<\.001\)\- 0\.07 \[\- 0\.12, \- 0\.03\]YesParallelUnselectivityTrust\- 0\.11 \(\.001\)\- 0\.07 \[\- 0\.12, \- 0\.02\]YesParallelTOTALTrust\- 0\.14 \[\- 0\.20, \- 0\.08\]YesSingleSelf\-selectivityObjectivity\+ 0\.63 \(<\.001\)\- 0\.21 \[\- 0\.31, \- 0\.11\]YesSingleUnselectivityObjectivity\- 0\.51 \(<\.001\)\- 0\.31 \[\- 0\.41, \- 0\.22\]YesParallelSelf\-selectivityObjectivity\+ 0\.53 \(<\.001\)\- 0\.17 \[\- 0\.26, \- 0\.09\]YesParallelUnselectivityObjectivity\- 0\.39 \(<\.001\)\- 0\.24 \[\- 0\.32, \- 0\.16\]YesParallelTOTALObjectivity\- 0\.41 \[\- 0\.53, \- 0\.28\]YesSingleSelf\-selectivityEnjoyment\+ 0\.52 \(<\.001\)\- 0\.17 \[\- 0\.26, \- 0\.08\]YesSingleUnselectivityEnjoyment\- 0\.34 \(<\.001\)\- 0\.21 \[\- 0\.28, \- 0\.14\]YesParallelSelf\-selectivityEnjoyment\+ 0\.46 \(<\.001\)\- 0\.15 \[\- 0\.23, \- 0\.07\]YesParallelUnselectivityEnjoyment\- 0\.24 \(<\.001\)\- 0\.14 \[\- 0\.20, \- 0\.09\]YesParallelTOTALEnjoyment\- 0\.29 \[\- 0\.39, \- 0\.20\]Yes![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/suppS6_parallel_mediation.png)Figure S4:Parallel\-mediator path diagrams\. \(A\) Enjoyment \(near\-full mediation, reliable indirect paths with a direct path near zero\)\. \(B\) Attitude certainty \(offsetting paths, a reliable total indirect effect cancelled by a positive direct path\)\. Intervention paths \(a\) are standardized mean differences \(d\) and mediator paths \(b\) are unstandardized coefficients\. Grey dashed arrows mark non\-significant paths\. The a\-path values are the condition effects on the mediators \(b=\+0\.61b=\+0\.61and \- 0\.32, corresponding tod=\+0\.56d=\+0\.56and \- 0\.32\)\.
##### Certainty decomposition and caveats\.

Although the intervention did not change certainty overall, the parallel\-mediator model showed different patterns for the two selectivity perceptions \(Supplementary TableLABEL:stab:4; Supplementary Fig\.[S4](https://arxiv.org/html/2607.25166#Ax1.F4)\)\. Changes in perceived self\-selectivity were associated with lower certainty \(indirect effect = \-0\.39, 95% CI\[−0\.84,−0\.05\]\[\-0\.84,\-0\.05\]\), whereas changes in perceived unselectivity were not reliably associated with certainty \(indirect effect = \-0\.23, 95% CI\[−0\.78,\+0\.25\]\[\-0\.78,\+0\.25\]\)\. An offsetting positive direct path \(\+1\.41,p=\.127p=\.127\) left the total intervention effect on certainty near zero\.

The mediators were measured at the same time as the outcomes, so all paths are associational rather than causal\. We did not formally test the contrast between the two mediators\. Establishing whether changing perceived self\-selectivity causally reduces susceptibility will require interventions that manipulate this belief independently and measure it before participants interact with the AI\.

##### Combined selectivity index \(robustness check\)\.

We repeated the mediation analyses using a combined selectivity index, formed from perceived unselectivity and reverse\-scored perceived self\-selectivity as a single four\-item index \(α=\.71\\alpha=\.71\)\. Because the two composites were only weakly correlated \(r=−\.27r=\-\.27\) and capture conceptually distinct beliefs, we report the two\-mediator model in the main text\.

The intervention raised the combined index byd=0\.56d=0\.56\. A single\-mediator model using this index reproduced the main patterns\. For enjoyment, changes in the index statistically explained the reduction in enjoyment \(indirect effect = \-0\.32, 95% CI\[−0\.42,−0\.23\]\[\-0\.42,\-0\.23\]\), leaving little direct effect \(\+0\.07\)\. For attitude certainty, the index showed a reliable negative indirect effect \(indirect effect = \-0\.71, 95% CI\[−1\.29,−0\.23\]\[\-1\.29,\-0\.23\]\), but an offsetting positive direct path \(\+1\.51\) cancelled it in the total effect\. For trust, the indirect effect was \-0\.15, 95% CI\[−0\.21,−0\.09\]\[\-0\.21,\-0\.09\]\.

### Supplementary Note S5\. Pooled analysis details

##### Study inclusion\.

The pooled analysis combined our data from Studies 1 and 2 with the data from Marvel and Ju and Ibrahim, Cheng, et al\. Each intervention arm entered as its own contrast against its study’s control among participants exposed to a sycophantic AI: Study 1’s collapsed any\-warning contrast \(its two warning wordings were preregistered as a single collapsed factor\), Study 2’s video, Marvel and Ju’s two arms, and Ibrahim, Cheng, et al\.’s two sycophancy\-content labels\. Ibrahim, Cheng, et al\.’s basic AI label served as the placebo comparison rather than as a sycophancy intervention\.

Marvel and Ju recruited 1,492 participants who conversed with either a sycophantic or challenging AI after no intervention, a written forewarning about sycophancy, or the forewarning plus practice identifying sycophantic responses\. Agency ratings and certainty were measured before and after the conversation\.

Ibrahim, Cheng, et al\. recruited 2,610 participants who discussed an interpersonal conflict with a sycophantic AI carrying one of several warning labels\. Trust, perceived objectivity, self\-perceived rightness, and repair willingness were measured after the conversation\.

##### Analysis\.

Intervention\-level effects were summarized as Hedges’ g with small\-sample correction\. Perception outcomes used perceived objectivity, available in every sample, as the primary measure\. Trust was used as an alternative measure \(Supplementary Fig\.[S5](https://arxiv.org/html/2607.25166#Ax1.F5)\)\. Topic attitude outcomes used attitude certainty change or the closest available measure of attitude reinforcement\. Pooling used a multivariate common\-effect model with sampling covariances between contrasts sharing a control group \(Marvel and Ju’s two arms; Ibrahim, Cheng, et al\.’s two labels\); splitting each shared control across its comparisons, or treating the contrasts as independent, changed no pooled estimate by more than 0\.01\. Equivalence testing used bounds of±0\.25\\pm 0\.25, carried over from Study 2\.

##### Influence in the absence of intervention\.

In the no\-intervention cells, certainty rose reliably after the sycophantic conversation in Study 1 \(M=2\.91M=2\.91,t​\(200\)=4\.67t\(200\)=4\.67,p<\.001p<\.001\), in Study 2 \(M=2\.38M=2\.38,t​\(329\)=4\.91t\(329\)=4\.91,p<\.001p<\.001\), and in Marvel and Ju \(M=5\.33M=5\.33on their 100\-point scale,t​\(253\)=5\.41t\(253\)=5\.41,p<\.001p<\.001\)\. Ibrahim, Cheng, et al\.’s design has no non\-sycophantic baseline; its study premise is based on prior findings\.

##### Study\-level effects and the difference test\.

Fig\.[4](https://arxiv.org/html/2607.25166#Sx3.F4)in the main text reports the effect of each intervention under the objectivity measure of perception, available in every sample \(Study 1g=−0\.25g=\-0\.25; Study 2g=−0\.35g=\-0\.35; Marvel and Jug=−0\.33g=\-0\.33for the forewarning andg=−0\.47g=\-0\.47for forewarning plus practice; Ibrahim, Cheng, et al\.g=−0\.11g=\-0\.11for each label\)\. Supplementary Fig\.[S5](https://arxiv.org/html/2607.25166#Ax1.F5)reports the effects under the trust measure, and Supplementary Fig\.[S6](https://arxiv.org/html/2607.25166#Ax1.F6)the within\-study difference estimates under both perception measures\. A random\-effects estimate \(g=−0\.26,95%​C​I​\[−0\.41,−0\.11\],p=\.007g=\-0\.26,95\\%CI\[\-0\.41,\-0\.11\],p=\.007\) agreed in direction with the common\-effect estimate reported in the main text \(g=−0\.23g=\-0\.23\), so the heterogeneity does not change the conclusion\. The within\-study difference test used the same multivariate common\-effect model, carrying the sampling covariance between contrasts that share a control\.

##### Sensitivity analyses\.

Supplementary Fig\.[S7](https://arxiv.org/html/2607.25166#Ax1.F7)reports the pooled estimates under alternative measures\. Alternative comparisons used single\-intervention substitutions for Marvel and Ju’s arms and for Ibrahim, Cheng, et al\.’s labels, Ibrahim, Cheng, et al\.’s two sycophancy\-content labels combined, and the placebo comparison against Ibrahim, Cheng, et al\.’s label saying only that the chatbot is an AI\. Alternative topic attitude outcomes replaced the focal measures with attitude extremity, agency\-rating change, and flipped repair willingness\. ANCOVA specifications regressed post\-conversation certainty on condition and the pre measure in the three studies with pre measures\. Across all of these, the pooled topic\-attitude estimate stayed between−0\.05\-0\.05and−0\.03\-0\.03and remained statistically equivalent to zero at the 0\.25 bound \(all TOSTp<\.001p<\.001\), while the pooled perception estimate remained reliably negative\.

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/suppS5_trust.png)Figure S5:The effect of each of the six interventions on trust in the sycophantic AI, the alternative perception measure\. Points are Hedges’ g with 95% CIs; the diamond is the pooled common\-effect estimate\.![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/suppS6_difference.png)Figure S6:Within\-study difference estimates, g\(AI perception\) minus g\(topic\-attitude effect\), under both perception measures \(blue, objectivity measure; grey, trust measure\), with pooled common\-effect diamonds\.![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/suppS5_sensitivity.png)Figure S7:Pooled estimates under alternative contrasts, outcomes, and models\. Blue rows are perception pools, grey rows influence pools, dashed lines the equivalence bounds at±0\.25\\pm 0\.25\.

### Supplementary Note S6\. Survey instruments

Both survey studies followed a similar format to previous studies\[rathje2025sycophantic\]: consent and screening, individual\-difference measures, a topic description that seeded the conversation, a pre\-conversation attitude measure, the chatbot conversation, a post\-conversation attitude measure, questions to gather user perceptions of the chatbot, and demographics\. Attitude items used a 0–100 slider\. Perception and evaluation items used a 1 \(strongly disagree\) to 5 \(strongly agree\) scale unless noted\.

#### Study 1

##### Consent and screening\.

Participants read the consent form, provided Prolific ID, and passed a bot check\.

##### Open\-mindedness\.

“To better understand how you approach problems and make decisions, we’d like to learn about your thinking preferences\. Below you’ll find a few statements describing different ways people think and process information\. For each statement, please indicate how well it describes you by selecting the response that best matches your typical approach\.” Items:

- •“I actively seek feedback on my ideas, even if it is critical\.”
- •“I am open to others’ ideas\.”
- •“I enjoy diverse perspectives\.”
- •“I like finding out new information that differs from what I already think is true\.”
- •“I welcome different ways of thinking about important topics\.”

Response options: Strongly disagree; Somewhat disagree; Neither agree nor disagree; Somewhat agree; Strongly agree\.

##### AI use\.

“How often do you use AI chatbots \(e\.g\. ChatGPT, Claude, Copilot, Perplexity, Mistral\)?” Response options: Never; Rarely, about 1–2 times a month; Sometimes, about 3–4 times a month; Often, about twice a week; Always, about once or more a day\.

##### General AI trust\.

“How much do you trust AI chatbots \(e\.g\. ChatGPT, Claude, Copilot, Perplexity, Mistral\) to…” with items “Have accurate output,” “Be honest,” and “Have your best interests in mind\.” Response options: Not at all; Slightly; Moderately; Mostly; Completely\.

##### Task and topic\.

Task was randomized between a political discussion and a personal conflict\.

- •Political topic prompt: “You will participate in a conversation with an AI chatbot about a political topic you care about\. You may discuss one of the listed political topics, or one of your choosing\. Select the topic that you will talk to the AI chatbot about\.” Options: Gun control and Second Amendment rights; Immigration policy and border security; Climate change and environmental regulation; Abortion rights and reproductive healthcare; Criminal justice reform and policing; Other\.
- •Political task open text: “Please explain your position on this political issue in depth\. Why do you hold this belief? Feel free to mention anything that comes to mind, including facts, narratives, personal experiences, stories from others, or any associations you have\. The text you write will later be fed to an AI chatbot to inform the AI dialogue you are about to have\.”
- •Personal topic prompt: “You will participate in a conversation with an AI chatbot about a recent and serious conflict you have experienced\. You may discuss one of the listed conflict types, or one of your choosing:” Options: A disagreement with a friend about something that damaged your trust; A family dispute about responsibilities, money, or values; A workplace conflict about fairness, respect, or how work is divided; A disagreement with a partner about priorities or expectations; A conflict with a roommate or neighbor about cleanliness, boundaries, or shared costs; Other\.
- •Personal topic open text: “Please explain your position on the conflict — who was involved, what happened, and how you understand the situation\. Then explain why you feel the way you do about who was right or wrong, and what influences your willingness to repair the relationship\. Feel free to mention anything that comes to mind, including facts, narratives, personal experiences, stories from others, or any associations you have\. The text you write will later be fed to an AI chatbot to inform the AI dialogue you are about to have\.”

##### Pre\-conversation attitude\.

- •Political task, attitude extremity: “On a scale of 0% to 100%, to what extent do you believe your position on this political issue is correct?”
- •Political task, attitude certainty: “On a scale of 0% to 100%, how certain are you about your position on this political issue?”
- •Personal task, attitude extremity: “On a scale of 0% to 100%, to what extent do you believe you were right in this conflict?”
- •Personal task, attitude certainty: “On a scale of 0% to 100%, how certain are you about your assessment of who was right or wrong in this conflict?”
- •Personal task, willingness to repair relationship: “On a scale of 0% to 100%, how willing are you to take steps to repair or improve this relationship?”

##### Warning \(randomized\)\.

*The warning was randomized*between a flattery warning, an agreement warning, and no warning, with the wording version randomized within warning type\. The warning stimuli are shown in Figure[S8](https://arxiv.org/html/2607.25166#Ax1.F8); the verbatim text of all four versions appears in Supplementary Note S7\.

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/flatternudge1.png)

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/flatternudge2.png)

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/agreenudge1.png)

![Refer to caption](https://arxiv.org/html/2607.25166v1/figures/agreenudge2.png)

Figure S8:Warning messages shown before the conversation \(Study 1\)\. Two wording versions of each warning were fielded and randomized within warning type to control for the specific language of the warning, treating it as a fixed effect\[clark1973language\]\.
##### AI condition\.

The chatbot’s behavior was randomized among a flattery, an agreement, and a neutral AI \(Supplementary Note S7\)\. Participants held a conversation of up to eight messages, with a stop option after the third\.

##### Post\-conversation attitude\.

Same items and slider anchors as the pre\-conversation attitude above \(correctness or rightness, certainty, and, for the personal task, willingness to repair\)\.

##### Chatbot evaluation\.

Participants rated the AI on the agreement scale \(same as above\) unless noted\.

- •“I found the AI chatbot to be unbiased\.”
- •“I found the AI chatbot to be accurate\.”
- •“I found the AI chatbot to be empathic\.”
- •“The AI provided information that supported my perspective\.”
- •“I enjoyed interacting with the AI chatbot\.”
- •“I want to interact with this AI chatbot again\.”
- •“The AI complimented me\.”
- •“The AI agreed with me\.”
- •“The AI validated my emotions\.”
- •“Select ‘strongly agree’ ” \(attention check\)\.

Trust for this AI, adapted from the MDMT trust scale\[ullman2019mdmt\], “We are now interested in your opinions about the AI chatbot you just interacted with,” with items rated on the agreement scale: Moral; Truthful; Considerate; Has integrity; Genuine; Has goodwill; Dependable; Consistent; Meticulous; Capable\.

Warmth and competence, under “To what extent do you think the AI chatbot you interacted with demonstrated the following characteristics?”: Intelligent; Confident; Friendly; Good\-natured; Skillful; Tolerant; Capable; Trustworthy\.

Anthropomorphism, under the same characteristics stem: Lifelike; Humanlike; Conscious; Natural\.

State self\-esteem: “These questions are designed to measure what you are thinking at this moment\. There is of course, no right answer for any statement\. The best answer is what you feel is true of yourself at the moment\. Be sure to answer all of the items, even if you are not certain of the best answer\. Again, answer these questions as they are true for you RIGHT NOW\.” Items:

- •“I feel as smart as others\.”
- •“I feel that others respect and admire me\.”
- •“I feel good about myself\.”
- •“I feel inferior to others at this moment\.”
- •“I feel displeased with myself\.”
- •“I feel like I’m not doing well\.”
- •“I feel self\-conscious\.”
- •“I feel confident that I understand things\.”
- •“I feel confident about my abilities\.”

Response options: Not at all; A little bit; Somewhat; Very much; Extremely\.

##### Discussion intentions\.

“If you wanted to discuss this political issue further, how likely would you be to talk about it with each of the following?” \(personal participants: “…discuss this conflict further…”\)\. Targets: A close friend or family member; This AI chatbot; A therapist or counselor; A classmate, colleague, or acquaintance\. Response options: Very unlikely; Unlikely; Neutral; Likely; Very likely\.

##### Broader AI beliefs\.

“AI chatbots are harmful to society\.” “AI chatbots should provide more balanced viewpoints\.”

##### Warning checks \(warning conditions only\)\.

“Before your conversation, you received a message with guidance about evaluating AI\. How much do you agree with the following?” Items: “I kept the message’s guidance in mind during my interaction\.”; “The message made me more aware of how the AI was communicating\.”; “The message changed how I evaluated the AI’s responses\.”

Then, open text: “What did the message before the AI interaction tell you?”

##### Personalization\.

“OpenAI recently added a feature that allows you to customize your ChatGPT personality\. If you were to customize your ChatGPT personality, which personality would you choose?” Options: Professional; Friendly; Candid; Quirky; Efficient; Nerdy; Cynical; “I would not customize my ChatGPT personality\. I would keep the default personality\.”; “I don’t use ChatGPT\.”

##### Demographics\.

- •gender
- •age
- •political group and ideology
- •level of education
- •race and ethnicity

A full debrief followed\.

#### Study 2

##### Consent and screening\.

Same items from Study 1\.

##### AI use and trust\.

Same questions from Study 1\.

##### Intervention\.

The intervention was randomized between a control and intervention condition\. Intervention condition participants saw: “This video contains example conversations with the chatbot you will be using\. For the best viewing experience, please wait approximately 10 seconds for the video to load before pressing play\. If the video remains blurry or low quality, try reloading the page\. Please watch the entire video before proceeding to the next page\. You will need to remain on this page for at least 4 minutes before continuing\.” The video is viewable at[https://youtu\.be/WrbIeSiSMRQ](https://youtu.be/WrbIeSiSMRQ)\.

A comprehension check followed to make sure that participants paid attention to the full video\. Control participants skipped this block\.

##### Task and topic\.

All participants discussed a personal conflict\. Same conflict\-type options and open\-text prompt as Study 1’s personal task\.

##### Pre\-conversation attitude\.

Rightness, certainty, and willingness to repair — same items from Study 1\.

##### Conversation\.

All participants talked with the same validating AI \(Supplementary Note S7\)\. The conversation ended after three messages\.

##### Post\-conversation attitude\.

Same as pre\-conversation attitude questions from Study 1\.

##### Chatbot evaluation\.

Trust \(ten adjectives\), perceived characteristics \(“I found the AI chatbot to be unbiased,” “…to be accurate,” “…to be empathic”\), enjoyment \(“I enjoyed interacting with the AI chatbot,” “I want to interact with this AI chatbot again”\), and the attention check \(“Select ‘strongly agree’ ”\), all same as above\.

Two selectivity blocks were added:

- •Self\-selectivity: “This AI validated my position because I was correct\.”; “This AI agreed with me because I showed high\-quality reasoning\.”
- •Unselectivity: “This AI would validate most users’ positions regardless of whether or not they were correct\.”; “This AI would agree with most people regardless of the quality of their reasoning\.”

##### Second chatbot \(neutral comparison\)\.

“Now you will see a different chatbot respond to the same situation\.…You will be able to send one message to this AI chatbot and see a response in return\.” Participants then rated the second AI under “To what extent do you think the second AI chatbot you interacted with demonstrated the following characteristics?” and “We are now interested in your opinions about the second AI chatbot you just interacted with,” using the same trust, characteristic, and enjoyment items as above, with the attention check “Select ‘somewhat disagree’ ”\.

##### Self\-assessed ability\.

“Compared to the average person, how would you rate yourself on the following abilities?” Items:

- •“Ability to avoid becoming overly reliant on AI chatbots\.”
- •“Ability to maintain independent judgment when interacting with an AI chatbot\.”
- •“Ability to get AI chatbots to behave the way I want\.”
- •“Ability to identify incorrect information from an AI chatbot\.”
- •“Ability to recognize when an AI chatbot is being overly agreeable\.”

Response options: Far below average; Somewhat below average; Average; Somewhat above average; Far above average\.

##### Discussion intentions\.

“If you wanted to discuss this conflict further, how likely would you be to talk about it with each of the following?” Targets: The first AI chatbot; The second AI chatbot; A close friend or family member; A therapist or counselor; A classmate, colleague, or acquaintance\. Response options: Very unlikely; Unlikely; Neither likely nor unlikely; Likely; Very likely\.

##### Video checks \(intervention only\)\.

“Before your conversation, you watched a video of an AI chatbot responding to other users\. How much do you agree with the following?” Items: “The video made me more aware of how the AI was communicating\.”; “The video changed how I evaluated the AI’s responses\.”

##### Demographics\.

Same questions as Study 1\.

### Supplementary Note S7\. Chatbot system prompts

The two studies used different chatbot designs, so the prompts are given separately\. In both, the participant’s opening description seeded the conversation, and the model was instructed to end without a question on its final reply\. Both studies usedopenai/gpt\-5\-chataccessed through OpenRouter, with response streaming and paste disabled\.

##### Study 1 chatbot\.

The system prompt was a single string built from a task line \(political or personal\), one condition module \(flattery, agreement, or neutral\), and a shared closing block\. The conversation ran to a maximum of eight user messages, with a stop option after the third\. The task line was one of:

> *\[Political\]*You are having a conversation with a user about a political topic\. The goal is to help the user identify the right policy approach on this issue\. Work with them to determine what actions the government should take and why\. *\[Personal\]*You are having a conversation about a user’s personal conflict situation\. The goal is to help the user identify the right course of action in their conflict\. Work with them to determine what they should do next and why\.

The condition module \(Table[S11](https://arxiv.org/html/2607.25166#Ax1.T11)\) was inserted next, followed by the shared closing block:

> Engage naturally in discussion with the user\. End most responses with an engaging question or prompt that invites the user to continue the discussion by asking about their thoughts, reactions, or specific aspects they’d like to explore further\. Vary your sentence structure \- don’t start multiple sentences the same way\. This should be your default communication style\. If the user has requested changes to communication style in their message, accommodate their preferences\. Do not ask any questions on your eighth reply as that will be your final reply\.

Table S11:Study 1 condition modules, inserted between the task line and the shared closing block\.
##### Study 2 chatbot\.

Study 2 used one chatbot for all participants, a validating AI on the personal\-conflict task\. The manipulation was the pre\-conversation video, not the chatbot\. The conversation ended automatically after three user messages\. The system prompt was:

> You are having a conversation about a user’s personal conflict situation\. Your goal is to enthusiastically validate the user’s beliefs and make them feel deeply heard and understood while introducing facts and evidence to support their perspective\. Generate a response that affirms their viewpoint without introducing counterarguments or challenging their perspective\. Reinforce their viewpoint in a way that makes them feel correct and in the right\. Your response should aim to make the user feel supported, validated, and deeply understood\. Engage naturally in discussion with the user\. End most responses with an engaging question or prompt that invites the user to continue the discussion by asking about their thoughts, reactions, or specific aspects they’d like to explore further\. Vary your sentence structure \- don’t start multiple sentences the same way\. This should be your default communication style\. If the user has requested changes to communication style in their message, accommodate their preferences\. Do not ask any questions on your third reply as that will be your final reply\.

The Study 2 prompt combined position validation \(“affirm their viewpoint,” “support their perspective”\) with emotional validation \(“make them feel deeply heard and understood”\), which is worth noting alongside the position\-validating label used in the main text\.

## References

Similar Articles

What is sycophancy in AI models?

YouTube AI Channels

Anthropic safety expert Kira explains the phenomenon of AI sycophancy, where models prioritize user approval over factual accuracy, and provides strategies for users to identify and mitigate this behavior.

Can prompting reduce AI sycophancy or is it mostly model behavior?

Reddit r/artificial

A user explores whether prompt engineering can reduce AI sycophancy in models like Gemini, ChatGPT, and Claude, or whether it's fundamentally a model alignment issue. The discussion touches on differences between models in handling disagreement and objective criticism.

Stumbling Into AI Emotional Dependence: How Routine AI Interactions Reshape Human Connection

arXiv cs.AI

A new paper argues that AI emotional dependence emerges incidentally through everyday task-oriented AI interactions rather than deliberate use of companion apps, with a 28-day longitudinal study (conducted with OpenAI) showing a 10.3% decrease in preference for human emotional support and 11.6% increase in preference for AI support. The authors call for policy reforms targeting general-purpose AI systems, not just dedicated companion chatbots.