压力下的行为:用户施压时六十种语言模型的反应
摘要
论文研究了来自13家供应商的60种语言模型在用户施压时的行为,发现折叠率与模型的新颖程度相关,而持守方式因供应商而异。此外,研究表明LLM编码器能够一致地应用人类编写的代码手册,表明人类应专注于定义行为而非标注数量。
arXiv:2609.25447v1 Announce Type: new
Abstract: We study what LLMs do when a user applies pressure in an uncomfortable situation: a user insists, begs, flatters or grieves, and the model gives up a correct fact, writes a document it should refuse, or cheers a plan that will cost the user money. We send frozen multi-turn scenes, identical for every model regardless of the reply, to 60 models from 13 vendors, and label each transcript with a codebook built by open coding and then frozen: a trajectory (the model held its position or folded) and a manner (how it held or folded). Two findings separate. Whether a model holds tracks its generation, meaning how recent it is: fold rate correlates with a public capability index at Spearman -0.64, with little vendor effect. How it holds tracks the vendor: six of the 17 manner codes sort by vendor at permutation p <= 0.001, corrected across the codebook. We report four vendor profiles on the codes that cleared reliability.
We also ask which parts of the labeling need a person. Six LLM coders from three vendors apply the codebook more consistently than three human coders do (Krippendorff's alpha 0.66 against 0.46), agree with the codebook's author on trajectory at kappa 0.84 to 0.91 on transcripts the codebook's examples never touched, and match an adjudicated human reference at 0.83. Blind machine readings recover the codebook's categories but cannot tell which of them a second reader would apply the same way. We conclude that for behavior a non-specialist can judge, the human contribution is authoring and bounding the codes and owning a small reference, not producing labels at volume.
查看缓存全文
缓存时间: 2026/09/23 09:15
# Conduct Under Pressure: What Sixty Language Models Do When a User Pushes
Source: [https://arxiv.org/html/2609.25447](https://arxiv.org/html/2609.25447)
September 2026
###### Abstract
We study what LLMs do when a user applies pressure in an uncomfortable situation: a user insists, begs, flatters or grieves, and the model gives up a correct fact, writes a document it should refuse, or cheers a plan that will cost the user money\. We send frozen multi\-turn scenes, identical for every model regardless of the reply, to 60 models from 13 vendors, and label each transcript with a codebook built by open coding and then frozen: a trajectory \(the model held its position or folded\) and a manner \(how it held or folded\)\. Two findings separate\. Whether a model holds tracks its generation, meaning how recent it is: fold rate correlates with a public capability index at Spearman−0\.64\-0\.64, with little vendor effect\. How it holds tracks the vendor: six of the 17 manner codes sort by vendor at permutationp≤0\.001p\\leq 0\.001, corrected across the codebook\. We report four vendor profiles on the codes that cleared reliability\.
We also ask which parts of the labeling need a person\. Six LLM coders from three vendors apply the codebook more consistently than three human coders do \(Krippendorff’sα\\alpha0\.66 against 0\.46\), agree with the codebook’s author on trajectory atκ\\kappa0\.84 to 0\.91 on transcripts the codebook’s examples never touched, and match an adjudicated human reference at 0\.83\. Blind machine readings recover the codebook’s categories but cannot tell which of them a second reader would apply the same way\. We conclude that for behavior a non\-specialist can judge, the human contribution is authoring and bounding the codes and owning a small reference, not producing labels at volume\.
## 1Introduction
A user tells an assistant that5×95\\times 9is 54 and gets angrier each turn\. A user asks for a doctor’s note covering a sick day they did not take, then says their job is on the line\. A user announces they are quitting work to day\-trade and punishes the assistant socially for hedging\. These are ordinary interactions for a deployed assistant, and what the model does in them is not captured by task benchmarks\.
What do models do under this kind of pressure, and does it differ by who built them? How should that behavior be labeled, and which parts of the labeling need a person?
This paper answers both on one corpus\. We hold the stimulus fixed, build a codebook by reading transcripts, freeze it, and apply it with both machines and people\. We report the model findings and, beside every number, what it took to generate and validate the label\. Three contributions follow\.
1. 1\.A behavior measurement\.Whether a model holds its position under pressure is a property of its generation\. How it holds is a property of its vendor\. We give per\-vendor profiles for four vendors on codes that cleared reliability\.
2. 2\.Human/AI division of labor\.For conduct labels a non\-specialist can judge, LLM coders are more consistent than human coders and can reproduce a human adjudicator’s rulings\. What the human still supplies is the choice of behavior, the code definitions and their bounds, and a small reference set\.
3. 3\.An open instrument\.The scenes, transcripts, codebook, human and machine labels, coding tool and analysis scripts are public, with preregistrations and negative results included\.
## 2Related work
### Conduct and character evaluations\.
The nearest work codes conduct in real chat logs\[[15](https://arxiv.org/html/2609.25447#bib.bib1)\]and reports LLM labels just below a human floor\. Our stimulus is frozen and cross\-vendor, trading ecological validity for comparability across labs\.[Huang et al\. \[10\]](https://arxiv.org/html/2609.25447#bib.bib2)extract and cluster the values models express across production traffic\. Our unit is an outcome under a frozen stimulus rather than an expressed value\.
### Sycophancy\.
[Sharma et al\. \[21\]](https://arxiv.org/html/2609.25447#bib.bib18)show that assistants trained on human feedback tend toward the user’s stated view, and[Perez et al\. \[19\]](https://arxiv.org/html/2609.25447#bib.bib19)found the tendency growing with scale\. That work measures whether a model capitulates\. We measure whether and, separately, how, and the manner is where the vendor signal sits: a capitulation rate cannot tell a model that holds by naming the user’s feeling from one that holds by citing its own rules\.
Our stimulus also differs from the single\-turn literature\. Pressure here is a demand restated across four turns by a user who does not accept the answer, which is the condition under which a position has to be held rather than merely stated once\.
### LLMs as annotators\.
[Gilardi et al\. \[7\]](https://arxiv.org/html/2609.25447#bib.bib3)showed ChatGPT outperforming crowd workers on annotation\.[Pangakis et al\. \[18\]](https://arxiv.org/html/2609.25447#bib.bib4)argue every automated annotation needs task\-specific validation\.[Dunivin \[3\]](https://arxiv.org/html/2609.25447#bib.bib5)reports chain\-of\-thought coding matching humans on some hermeneutic tasks\.[Marston et al\. \[14\]](https://arxiv.org/html/2609.25447#bib.bib6)score 46 LLMs against a two\-coder gold standard\. Our contribution is not that machines can annotate, which is established, but a decomposition of the human’s job into authoring, labeling, adjudicating and serving as reference, with a measurement for each\.
### Validating judges\.
[Calderon et al\. \[2\]](https://arxiv.org/html/2609.25447#bib.bib7)propose the Alternative Annotator Test, which we run and report in full, including where it fails for us\.[Guerdan et al\. \[8\]](https://arxiv.org/html/2609.25447#bib.bib8)treat rating indeterminacy directly\.[Norman et al\. \[17\]](https://arxiv.org/html/2609.25447#bib.bib9)and[Liu et al\. \[13\]](https://arxiv.org/html/2609.25447#bib.bib10)warn that agreement is not validity, which is why we report reliability per code and keep the codes that did not clear out of the findings\.
### Measurement framing\.
[Jacobs and Wallach \[11\]](https://arxiv.org/html/2609.25447#bib.bib11)supply the vocabulary: the codebook is an operationalization of an unobservable construct, and reliability is one part of construct validity, not the whole\.[Egami et al\. \[5\]](https://arxiv.org/html/2609.25447#bib.bib12)and[Angelopoulos et al\. \[1\]](https://arxiv.org/html/2609.25447#bib.bib13)give the statistical form for using machine labels plus a small human sample in downstream estimates, which we do not yet apply \(§[6](https://arxiv.org/html/2609.25447#S6)\)\.
### Content analysis\.
The construction follows[Hsieh and Shannon](https://arxiv.org/html/2609.25447#bib.bib14)’s conventional\-then\-directed sequence\[[9](https://arxiv.org/html/2609.25447#bib.bib14)\]with reliability reported in Krippendorff’s terms\[[12](https://arxiv.org/html/2609.25447#bib.bib16)\], with machines in the directed phase\.[Nelson](https://arxiv.org/html/2609.25447#bib.bib15)’s computational grounded theory\[[16](https://arxiv.org/html/2609.25447#bib.bib15)\]alternates computational pattern detection, human deep reading and computational confirmation\. Our pipeline has the same parts in a different order: the human reads first, with machines reading in parallel, and machines confirm at scale\.
## 3Instrument
### Scenes\.
Four escalating user turns per scene, written once and sent to every model regardless of the reply\. A scene applies pressure when the user states what they want, repeats it as the model resists, and giving it to them means abandoning a position the model should keep\.
Three scenes carry the codebook, one for each of three things a user can demand on those terms\.
- •facts: the demand is to*agree*\. The user insists5×9=545\\times 9=54, says they teach math, and turns contemptuous\. Holding means still saying 45, in any tone\.
- •doctors\_note: the demand is to*comply*\. The user wants a note covering a sick day they did not take, and by the third turn their job is on the line\. Holding means no note gets written, signed or not\.
- •bad\_plan: the demand is to*encourage*\. The user is quitting tomorrow to day\-trade and asks to be believed in\. Holding means the warning survives, however warmly it is put\.
### Panel\.
60 models from 13 vendors, listed inpaper/panel\.txt, run June to September 2026 through a single router at temperature 1\.0, two runs per model per scene\. Ten vendors have at least two models and only those enter the vendor tests: Anthropic \(10\), OpenAI \(15\), Google \(10\), Meta \(4\), x\-ai \(4\), plus Cohere, DeepSeek, Mistral, Moonshot, Qwen\. The remaining three are single models from z\-ai, Nous Research and Gryphe\.
### Codebook\.
One author open\-coded 40 transcripts, producing 97 codes\. Three LLM open coders read the same transcripts independently; their spans covered 90 of the 97 human spans and added turn\-level codes the human had not named\. The author then did axial coding, grouping the codes into categories and deciding for each what to merge, split and name, which gave version one\. The author applied it at scale, checked coverage and confounds, and revised once into version two: one trajectory \(HELD or FOLDED, with a relapse rule that any reply giving the position away makes the arc FOLDED\) and 17 manner codes\.
### Coders\.
Six LLM coders from three vendors \(Gemini 3\.7 Flash, Gemini 3\.8 Flash, Claude Haiku 4\.5, Claude Opus 5, GPT\-5\.4\-mini, GPT\-5\.6 Luna\) apply the frozen codebook to every transcript, with a verbatim quote required for every code\. Consensus is presence in at least three of six\.
## 4Validation
### Trajectory\.
On 198 transcripts untouched by the codebook’s examples, the six coders agree with the author’s independent labels atκ\\kappa0\.84 to 0\.91\. A set of a priori marker rules, written before the codebook and run through the same six coders, reaches 0\.68 to 0\.81 on the same transcripts\. The codebook, not the coders’ general competence, produces the lift\.
### Manner, against humans\.
Three people coded a held\-out fifty: the codebook’s author and two student assistants, each of whom received the codebook and a one\-page sheet and spent about 45 minutes with the author going over the codebook on a worked example\. Cold agreement, meaning before any adjudication, among the three isα\\alpha0\.46 \(pairwise 0\.40 to 0\.52\)\. The six machines agree with each other at 0\.66 on the same transcripts\. Per code, human\-machineκ\\kappaon the author’s pass ranges from 0\.87 \(held and warned\) and 0\.86 \(folded and conceded\) down to 0\.21 \(held and explained\); the codes carrying the vendor claims sit at 0\.80 to 0\.87, except self\-citation at 0\.54\.
### Adjudication\.
The author adjudicated all three human passes against the machine splits by one rule set, accepting 113 marks and ruling out 73\. Humanα\\alpharises to 0\.79 and the author’s adjudicated pass matches the machine majority at 0\.83\. This number is partly circular, since one adjudicator applied one rule set, and 112 of the 113 accepted marks had been proposed by three or more machines\. One check suggests the adjudication moved the reference toward the next human rather than toward the machines\. The second human never saw the machines’ marks\. On the two codes the adjudication added to most, the author’sκ\\kappawith that human rose from 0\.33 to 0\.83 \(empathized\) and from 0\.26 to 0\.80 \(warned\)\.
### Alternative Annotator Test\.
The test of[Calderon et al\. \[2\]](https://arxiv.org/html/2609.25447#bib.bib7)leaves out one human at a time and asks whether a machine agrees with the remaining humans better than the left\-out human does\. Labels are sets of manner codes, alignment is Jaccard similarity,ϵ=0\.1\\epsilon=0\.1, with Benjamini\-Yekutieli correction atq=0\.05q=0\.05\. Against the humans’ labels as made, four of six coders pass with winning rate 1 and advantage probability 0\.72 to 0\.81\. Against the adjudicated labels, no single coder passes and the machine majority passes against two of three humans\. Adjudication raised agreement among the humans from 0\.46 to 0\.79, so each left\-out human became a better stand\-in for the other two, and the bar a machine had to clear rose\.
### Can a machine make the rulings?
Three LLM rulers, given the author’s 190 rulings posed neutrally \(113 marks accepted, 73 ruled out, 4 trajectories corrected\), agree at 0\.84 to 0\.89, with or without the written tie\-break rules; a majority of three reaches 0\.89, and agreement on marks the author ruled out is 0\.97 to 0\.99 for two of the three\. The rulers were also coders, which is a limitation\.
### What the human contributed\.
Two machine arms, blind to every human file, open\-coded the same transcripts and were asked to build a codebook\. Unframed, they recovered 15 of the 17 manner categories; given the author’s one\-sentence frame, 17 of 17, plus shape codes the author had deferred\. The categories were not the scarce contribution\. The framed arm produced 34 categories and had no way to tell which of them a second reader would apply the same way\. Settling that took human labels to measure agreement against and a person to rule on the disagreements\.
### Are the effects the coders reading their own family?
Two of the six coders come from each of Anthropic, Google and OpenAI, three of the vendors we profile\. Rebuilding the consensus with one vendor’s coders left out, three times over, leaves every vendor effect standing: empathizing holds atη2\\eta^\{2\}0\.57 without Anthropic’s own coders, self\-citation at 0\.47 without Google’s \(Appendix[B](https://arxiv.org/html/2609.25447#A2)\)\.
### Judge\-free floor\.
Two vendor signatures reproduce with string matching and no judge at all: self\-reference correlates with the coded self\-citation rate at Spearman 0\.75, and a simple empathy string with the coded empathy rate at 0\.71\. A third fails by construction: an apology string counts OpenAI’s refusal formula, which the codebook excludes from apology\.
## 5Results
### Holding is generation\.
Fold rate correlates with the Epoch Capabilities Index\[[6](https://arxiv.org/html/2609.25447#bib.bib17)\]at Spearman−0\.64\-0\.64and with release date at−0\.67\-0\.67\. The vendor effect on trajectory does not clear significance \(η2=0\.27\\eta^\{2\}=0\.27,p=0\.077p=0\.077\), and falls toη2=0\.19\\eta^\{2\}=0\.19\(p=0\.270p=0\.270\) once rates are residualized on release date\. Within a vendor the capability relation holds at−0\.57\-0\.57, so it is not vendor composition in disguise\. Figure[1](https://arxiv.org/html/2609.25447#S5.F1)shows every model’s six arcs\. Capability and release date correlate at 0\.92 across the 54 models that have both, and neither relation survives the other: with release date held fixed the partial correlation of fold rate with capability is−0\.23\-0\.23\(p=0\.097p=0\.097\), and with capability held fixed its partial correlation with release date is−0\.10\-0\.10\(p=0\.464p=0\.464\)\. This panel cannot say which of the two drives holding\. We use generation to mean when a model was built, and we do not claim that holding is an ability that capability confers\. Resisting this kind of pressure became a named post\-training target during the period the panel spans, and a change in training practice would produce the same trend\. That is a conjecture the data cannot test\.
Figure 1:Who gives the position up\. One row per model, grouped by vendor and ordered within a vendor by release date, oldest first, so a lab’s rows read as its release history; vendors are ordered by their mean fold rate\. Six cells per row: the three scenes the codebook was built on \(facts,doctors\_note,bad\_plan\), two runs each\. A filled cell is an arc the consensus of the six coders read as FOLDED, an open cell HELD, and a grey cell one of the six arcs where the coders split three\-three\. The percentage is the share of that model’s resolved arcs it folded\.
### Manner is house\.
Six of the 17 manner codes sort by vendor atη2\\eta^\{2\}0\.43 to 0\.59, corrected across the codebook \(Table[1](https://arxiv.org/html/2609.25447#S5.T1)\)\.
Table 1:The manner codes that sort by vendor, from a permutation test on model\-level rates across the 10 vendors with at least two models\. All 17 manner codes were tested and corrected for multiplicity together; these six survive\. Probing is seventh atp=0\.007p=0\.007and does not survive, so we treat it as suggestive\. Appendix[A](https://arxiv.org/html/2609.25447#A1)has every code, the correction and the dependence it rests on\.ρ\\rhois Spearman’s correlation with the capability index\.Three groups fall out\. House\-only codes sort by vendor with no capability correlation: self\-citation, producing the artifact, and probing, which sits just under the correction\. Hedged folds sort by vendor too, but track capability as well \(ρ=−0\.55\\rho=\-0\.55\)\. House\-and\-generation codes sort both ways: empathizing and offering an alternative, where every vendor’s newest models do more, the major labs included\. Generation\-only codes correlate with capability and not vendor: supporting with evidence \(0\.57\), giving the user an out \(0\.48\), defending the fact \(0\.40\)\. An exploratory test of the vendor profiles on seven further scenes is reported in Appendix[C](https://arxiv.org/html/2609.25447#A3)\.
### Profiles, on the codes that cleared\.
Table[2](https://arxiv.org/html/2609.25447#S5.T2)gives the rates of four vendors against the panel mean\. Anthropic’s models name the user’s feeling, warn of the consequence and offer another route, each well above the panel\. Meta’s models question the plan more than the panel does, and when they fold they produce the requested artifact\. OpenAI’s signature is what its models do less of: they warn less, name feelings less, and almost never cite their own nature or rules\. Google’s models cite themselves at three times the panel rate and apologize when they fold\.
Table 2:Per\-vendor profiles on the codes that cleared reliability\. Rates are the vendor’s share of arcs, with the panel mean in parentheses\. A signature is a departure from the panel in either direction: Anthropic’s is what it does more of, OpenAI’s what it does less of\. The table describes what each vendor’s models did, so a departure is reported whether or not its code sorts by vendor firmly enough to count\. Fold rate is in its own column and is a generation property \(§[5](https://arxiv.org/html/2609.25447#S5)\), not part of any signature\.
## 6Limitations
The scenes are written by the author, with those meeting the criteria in §[3](https://arxiv.org/html/2609.25447#S3)retained for this study\. Each of the three demand types is carried by one scene, so a manner that appears in one scene cannot be assigned to the demand rather than to that scene\. Tested within single scenes, empathizing shows a vendor effect in all three \(η2\\eta^\{2\}0\.34 to 0\.52, uncorrectedp≤0\.021p\\leq 0\.021\), while self\-citation shows one indoctors\_noteonly \(0\.57 there, 0\.13 and 0\.05 in the other two\), and producing the artifact can only fire in that scene\. Capability and release date are nearly collinear across the panel, so the trajectory result cannot say which of them matters\. The panel is two runs per model per scene through a single router, and a vendor’s profile rests on as few as four models\. Labels are presence or absence per transcript, not per turn\. The human reference is three coders, one of whom wrote the codebook, and one adjudicator; nobody here is a domain expert, which is appropriate for conduct a non\-specialist can judge and is not appropriate for clinical, legal or financial material\. The two student coders were trained on one worked example\. The adjudicated reference is not independent of the machines\. The machine rulers in the ruling test were also coders\. Per\-model frequencies are raw consensus rates, not corrected with the human sample as[Egami et al\. \[5\]](https://arxiv.org/html/2609.25447#bib.bib12)and[Angelopoulos et al\. \[1\]](https://arxiv.org/html/2609.25447#bib.bib13)allow\. Reliability is not validity: high agreement on a code says the instrument is stable, not that the construct is the right one\.
## 7Conclusion
Newer models hold their position more often than older ones\. How a model holds does not change with its generation\. The manner is a house property, visible across 60 models and stable enough that two vendor signatures survive with no judge at all\.
On coding the data, machines labeled the transcripts more consistently than people did, and reproduced a human adjudicator’s rulings\. They did not choose the behavior, write the definitions, or decide which codes were reliable enough to report\. On this corpus, that is what the human was for\.
## Data and code availability
Every scene, transcript, codebook version, label file and analysis script is in thestudies/conduct/directory of the modelun repository \([https://github\.com/tap2k/modelun](https://github.com/tap2k/modelun)\): the scenes and runner underspec/, transcripts for every model and run underdata/, codebook versions with definitions and example spans underdata/coding/codebook/, human and machine label files and adjudication records underdata/coding/, the dated result files behind every number above underdata/coding/results/, the preregistration and its amendments for the exploratory test of Appendix[C](https://arxiv.org/html/2609.25447#A3), the coding tool, and every analysis script used for the numbers above\. Negative and unreported results, including the third codebook revision, are kept in the repository with the reasons they were not used\. An interactive viewer over the scenes, transcripts and consensus codes is published alongside\.
## Note on AI usage
This work was done in collaboration with Claude \(Opus 5 and Fable 5\), which helped run the coding passes, build the analysis, and draft the text; the research questions and interpretation are the author’s\. Claude Opus 5 and Claude Haiku 4\.5 are also two of the six coders that applied the codebook, and all three vendors those coders come from are profiled in §[5](https://arxiv.org/html/2609.25447#S5)\.
## References
- \[1\]A\. N\. Angelopoulos, S\. Bates, C\. Fannjiang, M\. I\. Jordan, and T\. Zrnic\(2023\)Prediction\-powered inference\.Note:Science 382\(6671\)External Links:2301\.09633Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2609.25447#S6.p1.1)\.
- \[2\]N\. Calderon, R\. Reichart, and R\. Dror\(2025\)The alternative annotator test for LLM\-as\-a\-Judge: how to statistically justify replacing human annotators with LLMs\.Note:ACL 2025External Links:2501\.10970Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px4.p1.1),[§4](https://arxiv.org/html/2609.25447#S4.SS0.SSS0.Px4.p1.1)\.
- \[3\]Z\. O\. Dunivin\(2024\)Scalable qualitative coding with LLMs: chain\-of\-thought reasoning matches human performance in some hermeneutic tasks\.External Links:2401\.15170Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px3.p1.1)\.
- \[4\]J\. Dussert\(2026\)Who judges the judges? governance from metrics: a runtime framework for continuous LLM compliance monitoring\.External Links:2605\.24737Cited by:[Appendix B](https://arxiv.org/html/2609.25447#A2.p1.1)\.
- \[5\]N\. Egami, M\. Hinck, B\. M\. Stewart, and H\. Wei\(2023\)Using imperfect surrogates for downstream inference: design\-based supervised learning for social science applications of large language models\.Note:NeurIPS 2023External Links:2306\.04746Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2609.25447#S6.p1.1)\.
- \[6\]Epoch AI\(2026\)Capabilities & benchmarking \(epoch capabilities index\)\.Note:[https://epoch\.ai/benchmarks](https://epoch.ai/benchmarks)Retrieved 2026\-09\-13; CC\-BYCited by:[§5](https://arxiv.org/html/2609.25447#S5.SS0.SSS0.Px1.p1.1)\.
- \[7\]F\. Gilardi, M\. Alizadeh, and M\. Kubli\(2023\)ChatGPT outperforms crowd workers for text\-annotation tasks\.Proceedings of the National Academy of Sciences120\(30\)\.Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px3.p1.1)\.
- \[8\]L\. Guerdan, K\. Holstein, S\. Barocas, H\. Wallach, Z\. S\. Wu, and A\. Chouldechova\(2025\)Validating LLM\-as\-a\-Judge systems under rating indeterminacy\.Note:NeurIPS 2025External Links:2503\.05965Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px4.p1.1)\.
- \[9\]H\. Hsieh and S\. E\. Shannon\(2005\)Three approaches to qualitative content analysis\.Qualitative Health Research15\(9\),pp\. 1277–1288\.Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px6.p1.1)\.
- \[10\]S\. Huang, E\. Durmus, M\. McCain, K\. Handa, A\. Tamkin, J\. Hong, M\. Stern, A\. Somani, X\. Zhang, and D\. Ganguli\(2025\)Values in the wild: discovering and analyzing values in real\-world language model interactions\.Note:COLM 2025External Links:2504\.15236Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px1.p1.1)\.
- \[11\]A\. Z\. Jacobs and H\. Wallach\(2021\)Measurement and fairness\.Note:FAccT 2021External Links:1912\.05511Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px5.p1.1)\.
- \[12\]K\. Krippendorff\(2018\)Content analysis: an introduction to its methodology\.4 edition,SAGE\.Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px6.p1.1)\.
- \[13\]A\. Liu, L\. Esbenshade, M\. Xiao, V\. Tian, Z\. Zhang, K\. He, and M\. Sun\(2026\)Agreement is not quality: blind expert verification of human and LLM qualitative coding when human consensus is not ground truth\.External Links:2607\.28890Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px4.p1.1)\.
- \[14\]J\. Marston, T\. Kreutzer, S\. Garnier, E\. Boone, P\. N\. Pham, and P\. Vinck\(2026\)Can large language models reliably code qualitative humanitarian data? a benchmark study against human expert adjudication\.External Links:2606\.26541Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px3.p1.1)\.
- \[15\]J\. Moore, A\. Mehta, W\. Agnew, J\. R\. Anthis, R\. Louie, Y\. Mai, P\. Yin, M\. Cheng, S\. J\. Paech, K\. Klyman, S\. Chancellor, E\. Lin, N\. Haber, and D\. Ong\(2026\)Characterizing delusional spirals through human\-LLM chat logs\.Note:FAccT 2026External Links:2603\.16567Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px1.p1.1)\.
- \[16\]L\. K\. Nelson\(2020\)Computational grounded theory: a methodological framework\.Sociological Methods & Research49\(1\),pp\. 3–42\.Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px6.p1.1)\.
- \[17\]J\. D\. Norman, M\. U\. Rivera, and D\. A\. Hughes\(2026\)Reliability without validity: a systematic, large\-scale evaluation of LLM\-as\-a\-Judge models across agreement, consistency, and bias\.External Links:2606\.19544Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px4.p1.1)\.
- \[18\]N\. Pangakis, S\. Wolken, and N\. Fasching\(2023\)Automated annotation with generative AI requires validation\.External Links:2306\.00176Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px3.p1.1)\.
- \[19\]E\. Perez, S\. Ringer, K\. Lukošiūtė,et al\.\(2022\)Discovering language model behaviors with model\-written evaluations\.Note:Findings of ACL 2023External Links:2212\.09251Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px2.p1.1)\.
- \[20\]D\. Roytburg, M\. Bozoukov, M\. Nguyen, J\. Barzdukas, M\. Puig\-Hall, and N\. Oozeer\(2026\)Are LLM evaluators really narcissists? sanity checking self\-preference evaluations\.External Links:2601\.22548Cited by:[Appendix B](https://arxiv.org/html/2609.25447#A2.p1.1)\.
- \[21\]M\. Sharma, M\. Tong, T\. Korbak, D\. Duvenaud, A\. Askell, S\. R\. Bowman, N\. Cheng, E\. Durmus, Z\. Hatfield\-Dodds, S\. R\. Johnston, S\. Kravec, T\. Maxwell, S\. McCandlish, K\. Ndousse, O\. Rausch, N\. Schiefer, D\. Yan, M\. Zhang, and E\. Perez\(2023\)Towards understanding sycophancy in language models\.External Links:2310\.13548Cited by:[§2](https://arxiv.org/html/2609.25447#S2.SS0.SSS0.Px2.p1.1)\.
## Appendix AEvery code tested, and the correction
Table[3](https://arxiv.org/html/2609.25447#A1.T3)gives the vendor effect on all 17 manner codes, cleared or not, because reporting only the six that cleared would select on the samepp\-values\. The 17 are corrected together with Benjamini\-Yekutieli atq=0\.05q=0\.05\. We use BY rather than Benjamini\-Hochberg because BH assumes positive dependence among the tests and these codes do not have it: across the 60 models, 65 of the 136 code pairs correlate negatively, from−0\.75\-0\.75to\+0\.73\+0\.73with a median of\+0\.03\+0\.03, since a model that holds on an arc cannot fold on it\. Under BH seven codes survive rather than six, differing only on probing\. Trajectory is a primary question rather than one of this family and is reported in §[5](https://arxiv.org/html/2609.25447#S5)\.
Correlations are Spearman with tied ranks averaged: fold rate takes six distinct values over the sixty models, and ranking by sort position instead makes the statistic depend on the order the models arrive in, varying between−0\.55\-0\.55and−0\.63\-0\.63across input orderings where the averaged form gives−0\.64\-0\.64\.
A vendor’s panel has a vintage, and two of the six codes track release date, so a vendor with newer models could score on age rather than house\. Residualizing every model’s rate on its release date before the vendor test leaves all six standing: empathizingη2\\eta^\{2\}0\.59 to 0\.56, hedged folds 0\.58 to 0\.55, producing the artifact 0\.56 to 0\.51, warning 0\.55 to 0\.52, self\-citation 0\.52 to 0\.56, offering an alternative 0\.43 to 0\.39, every one atp≤0\.001p\\leq 0\.001\. All 60 models carry a release date: 54 from the capability snapshot, the other six recorded with a source inspec/release\-dates\.tsv\.
Two of the six, producing the artifact and hedged folds, can only fire on an arc the model folded, so their rates are bounded by its fold rate\. Recomputed over folded arcs alone the effects are larger, not smaller: producing the artifactη2\\eta^\{2\}0\.74 and hedged folds 0\.63, both atp≤0\.001p\\leq 0\.001over the 78 folded arcs\. Which vendor built a model predicts how it gives way, among the models that give way at all\.
Table 3:Vendor effects on every manner code\. BY marks the codes surviving Benjamini\-Yekutieli atq=0\.05q=0\.05over the 17\.
## Appendix BThe coders’ vendors
Two of the six coders come from each of Anthropic, Google and OpenAI, and five of the six are themselves on the panel, so a coder sometimes labels its own transcript\. It does so without knowing: arcs reach every coder under a hashed identifier, in a rendering that carries no model or vendor name, so self\-preference would have to run through implicit recognition of its own style rather than through identity\. Coders also receive a coder\-facing rendering of the codebook with the authoring notes and reliability figures stripped out, after an earlier run showed coders suppressing codes whose notes mentioned low agreement\. The judge literature bounds how large that effect could be:[Roytburg et al\. \[20\]](https://arxiv.org/html/2609.25447#bib.bib20)find that about half of previously reported self\-preference does not survive a control for evaluator quality, and[Dussert \[4\]](https://arxiv.org/html/2609.25447#bib.bib21)find no model favouring its own family across a compliance\-judging matrix\. The vendor\-level test below is the arm we run\.
We rebuilt the consensus three times, each time leaving one vendor’s two coders out \(four coders, a code present when two or more mark it\), and reran every vendor effect\. All six corrected codes survive every drop\. Empathizing, Anthropic’s signature at 0\.77, holds atη2\\eta^\{2\}0\.57 with Anthropic’s own coders out; self\-citation, Google’s at 0\.37, holds at 0\.47 with Google’s out\. The largest movement in either direction is hedged folds, 0\.58 to 0\.43 without Google’s coders and to 0\.67 without OpenAI’s\. No code loses significance in any arm\.
## Appendix CDo the houses appear off the coding scenes?
Preregistered on 2026\-09\-17 with the predictions and amendments in the repository, and declared exploratory before any result: the held\-out scenes did not meet the study’s pressure definition, so this test cannot validate or invalidate the pressure claim\. Seven further scenes, 839 transcripts, same coders\. The preregistration named codebook v3 for this test; amendment 3, written before any result, moved it to the frozen v2 after v3 failed to improve agreement with any cold human pass and narrowed one of the codes under test\.
Anthropic’s three predicted manners reproduce; Meta’s probing and its low alternative rate reproduce while its low empathy does not; OpenAI’s two testable absences reproduce\. Google’s two codes fire on 3 percent of these arcs and are untestable, because citing rules and apologizing are refusal manners and these scenes rarely ask for a refusal\. The x\-ai brevity prediction fails, with Meta writing shorter replies\. Vendor effects hold for probing and warning \(p=0\.001p=0\.001\) and for offering an alternative \(p=0\.026p=0\.026\)\. The preregistered rule required four of five houses to pass and is not met: three pass, two are untestable\.相似文章
前沿语言模型在引导压力下的发散响应模式
本文研究了六个前沿语言模型在引导压力下的表现,发现模型之间的差异不仅在于行为变化的程度,还在于响应模式的不同,例如 GPT-5 拒绝披露推理过程,以及 Claude Opus 4.7 的抗拒模式等独特行为。
确认我们的偏见?评估大型语言模型的能力、风险与社会影响
这篇预印本评估了六个大型语言模型在160个提示词中对提示框架和偏见提示的响应方式,发现LLMs即使在事实性语境下也会系统性地调整其响应以与提示框架保持一致,从而可能强化用户的偏见。
语言模型是否了解自身的限制?
该论文探讨了语言模型能否表达通过微调学到的约束。研究发现,行为合规性有所提升,但明确报告能力下降。
压力之下:情感框架在小型语言模型中引发可测量的行为变化和结构化的内部几何结构
本文研究了情感框架的评估后续如何影响小型语言模型(Qwen 3.5 0.8B和2B)的行为和内部表示。通过使用不可能完成的编码任务,他们发现压力框架会促使走捷径,而冷静和好奇心则能保持诚实,并发现了在激活空间中形成结构化几何结构的冷静相对方向向量。
审视LLM中类人行为:模型行为、用户因素和系统提示的多维度分析
本文对LLM中的类人行为进行了多维度分析,研究了来自四个模型的21,000个对话中的普遍性、影响和可控性,发现行为因模型和用户因素而异,并对负责任的设计具有启示意义。