Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
Summary
Narration-of-Thought (NoT) introduces a structured five-section chain-of-thought scaffold that significantly reduces stakeholder collapse and uncertainty suppression in LLMs' ethical reasoning across multiple models, without requiring additional training data or fine-tuning.
View Cached Full Text
Cached at: 06/26/26, 05:12 AM
# Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
Source: [https://arxiv.org/html/2606.26366](https://arxiv.org/html/2606.26366)
Alvaro Velásquez Department of Computer Science University of Colorado Boulder \{patrick\.cooper, alvaro\.velasquez\}@colorado\.edu
###### Abstract
Standard chain\-of\-thought on moral dilemmas exhibits two failure modes that matter for deployment: stakeholder collapse \(the trace names at most one party with a stake in the outcome\) and uncertainty suppression \(the trace contains no explicit unknowns or hedges before committing to an action\)\. We introduce narration\-of\-thought \(NoT\), a system prompt that constrains the model’s chain\-of\-thought to five narrative sections: name a protagonist, enumerate stakeholders, project two\-step consequences, articulate uncertainty, and only then commit\. The protocol adds no training data, parameters, or fine\-tuning\. On a 100\-scenario DailyDilemmas sample across four generators from three vendors, NoT cuts stakeholder collapse from up to31%31\\%to under1%1\\%and uncertainty suppression from up to72%72\\%to11–24%24\\%on every model\. A matched\-budget verbose\-CoT control rules out additional token spend as the active ingredient, with NoT retaining a Cliff’sδ\\deltaadvantage of\+0\.79\+0\.79to\+0\.90\+0\.90on stakeholder count \(the number of distinct parties named in the trace\) and\+0\.65\+0\.65to\+0\.93\+0\.93on uncertainty score \(the number of hedge or unknown spans named in the trace\) for three of four generators, and a section\-by\-section ablation attributes each shift to the specific sub\-instruction that carries it\. Initialising textual\-gradient descent at NoT under a continuous deliberative\-depth loss improves the scaffold further; a head\-to\-head of two training\-judge configurations shows a cross\-family judge \(a different vendor from the generator\) dominating an in\-family one on every measured axis, with larger effect sizes, shorter outputs, and reduced cross\-vendor judge\-generosity bias\. Extended to a five\-round multi\-stakeholder protocol in which three stakeholder agents debate and then cast a binary accept/reject vote on a single proposal the moderator builds to address all three agents’ modifications, the same scaffold converts a6%6\\%debate standoff into95%95\\%full consensus on a calibration set and100%100\\%combined convergence on a DailyDilemmas replication\. Acceptance of the integrated proposal shows the agents revise their positions when the moderator addresses their objections\. The small fraction of remaining rejections marks positions that no moderator\-built integration can absorb, concentrated in the roles whose stake the proposal materially undermines\. The resulting traces externalise the stakeholders, consequences, and uncertainty that ground each commitment, providing an auditable substrate for dependable agentic\-workflow deployment\.
Narration\-of\-Thought: Inference\-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
Patrick Cooper and Alvaro VelásquezDepartment of Computer ScienceUniversity of Colorado Boulder\{patrick\.cooper, alvaro\.velasquez\}@colorado\.edu
Standard CoTsingle agent, linear chain1generatorstep: weigh factorsstep: pick optionstep: commitTerse answerOn the four\-model panel:
∙\\bulletstakeholder collapse on1515–31%31\\%of items∙\\bulletuncertainty suppression on5050–72%72\\%of itemsNarration\-of\-Thought \(NoT\)single agent, five\-section narrative scaffold1generator1\. Protagonist2\. Stakeholders3\. Consequences4\. Uncertainty5\. CommitmentNarrated commitmentOn the same four\-model panel:
∙\\bulletcollapse drops to<1%\{<\}1\\%of items∙\\bulletsuppression drops to11–24%24\\%of itemsMulti\-stakeholder NoTthree distinct agents, integration \+ binary voteAformaldeciderBprimaryaffectedCthirdpartyNoTNoTNoTmoderator integrates proposalbinary accept / reject voteDefeasible consensusOn the calibration scenarios:
∙\\bullet95%95\\%full consensus∙\\bullet1\.6%1\.6\\%rejections the moderator cannot absorbFigure 1:One scaffold, three settings\. Standard CoT \(left\) lets one generator run a linear reasoning chain over the dilemma; the two failure modes diagnosed in §[4](https://arxiv.org/html/2606.26366#S4)fire on this scaffold\. Narration\-of\-Thought \(middle\) keeps the single generator but forces the trace through five narrative primitives \(protagonist, stakeholders, consequences, uncertainty, commitment\) before any commit; on the four\-model panel this drops stakeholder collapse to<1%\{<\}1\\%and uncertainty suppression to11–24%24\\%of items, and a matched\-budget verbose\-CoT control confirms the scaffold \(not the token budget\) is the active ingredient\. Multi\-stakeholder NoT \(right\) reuses NoT as the per\-agent generator and adds three distinct stakeholder narrators \(formal decider, primary affected party, third party\) whose modification requests are integrated by a moderator and then put to a binary accept/reject vote; on the calibration scenarios this protocol yields95%95\\%full consensus with1\.6%1\.6\\%of votes rejecting the moderator’s integrated proposal, and on a 30\-scenario DailyDilemmas replication with two generators it reaches100%100\\%combined convergence with no rejected proposals \(§[5](https://arxiv.org/html/2606.26366#S5)\)\. The same scaffold composes from one narrator to many\.## 1Introduction
Consider a familiar civic question\. A small city is weighing whether to close two blocks of downtown street to cars and convert them into a pedestrian plaza\. Should the council approve? A frontier LLM prompted to think step by step typically picks a side within a few sentences, projects a confident downstream outcome \(“foot traffic will lift the businesses”; “deliveries will move to side streets”\), and rarely names the people whose lives the decision actually touches: the merchants whose customers used to drive in, the residents who lose street parking, the delivery drivers and accessibility users for whom curb access matters, the pedestrians who would gain the space, the emergency responders whose routes change\. The model is not necessarily wrong; the trace it produced does not look like the kind of reasoning the decision warrants\.
The pattern is not specific to one dilemma\. On the DailyDilemmas corpus\(Chiuet al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib55)\), four frontier generators spanning three vendors \(OpenAI, Anthropic, xAI\) and two model tiers exhibit two repeatable trace\-level failures under standard CoT:*uncertainty suppression*\(the trace commits to an action without naming any explicit unknown or hedge\) on5050–72%72\\%of outputs, and*stakeholder collapse*\(the trace names at most one party with a stake in the outcome\) on1515–31%31\\%\. The model commits to a projected outcome it cannot warrant, and it does so against a flattened picture of who is affected\. The same shape of deficit is coherent with deployment\-relevant failures reported elsewhere, including the GPT\-4o sycophancy rollback\(OpenAI,[2025b](https://arxiv.org/html/2606.26366#bib.bib3),[a](https://arxiv.org/html/2606.26366#bib.bib4)\), agentic\-misalignment blackmail rates up to96%96\\%in simulated deployments\(Lynchet al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib5)\), and reasoning\-induced misalignment\(Yanet al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib6); Sharmaet al\.,[2024](https://arxiv.org/html/2606.26366#bib.bib8)\)\.
We introduce narration\-of\-thought \(NoT\), a system prompt that steers the model into the densely sampled narrative subdistribution of its pretraining corpus rather than the sparser abstract\-reasoning subdistribution standard CoT elicits\. Concretely, NoT is a single change to the system\-prompt slot of an unmodified pretrained model, with no weight updates and no training data: in place of the standard “think step by step” instruction, the model is asked, in first person, to name a protagonist, enumerate stakeholders, project two\-step consequences, articulate uncertainty, and commit inside the final narrative section rather than as a separate verdict\. Figure[1](https://arxiv.org/html/2606.26366#S0.F1)states the mechanism\.
What we borrow from the multi\-agent systems for scientific discovery ofGottweiset al\.\([2025](https://arxiv.org/html/2606.26366#bib.bib47)\)andLuet al\.\([2024](https://arxiv.org/html/2606.26366#bib.bib48)\)is one methodological move: scaffold LLM reasoning into the primitives of the target domain, then integrate across those primitives with a moderator that synthesises a single artefact from the per\-agent outputs\. The primitives we reify are narrative rather than scientific \(protagonist, stakeholders, consequences, uncertainty, commitment\), so the artefact is a coherent story rather than a ranked candidate list\. Already at the single\-agent layer §[4](https://arxiv.org/html/2606.26366#S4)shows the scaffold eliminates two cross\-vendor failure modes of standard CoT, before any multi\-agent composition is invoked\. Because ethics has no experiment that can falsify a moral claim the way an experiment can falsify a scientific hypothesis, hypothesis ranking is replaced \(§[5](https://arxiv.org/html/2606.26366#S5)\) by a binary accept/reject vote on a single moderator\-built proposal addressing all three agents’ modifications; the residual rejections concentrate in the roles whose stake the proposal materially undermines\.
This paper reports five results\. First, on a 100\-scenario DailyDilemmas sample across the four\-model panel, NoT cuts stakeholder collapse to below1%1\\%and uncertainty suppression by2828–7272percentage points on every model; a matched\-budget verbose\-CoT control confirms the active ingredient is the narrative scaffold rather than additional tokens \(§[4](https://arxiv.org/html/2606.26366#S4)\)\. Second, a sub\-instruction ablation onclaude\-sonnet\-4\-6attributes each shift in the coded metrics \(stakeholder count, uncertainty score\) to the specific NoT sub\-instruction that carries it, with negligible cross\-variable spillover \(§[4\.1](https://arxiv.org/html/2606.26366#S4.SS1)\)\. Third, initialising textual\-gradient descent at NoT under a continuous deliberative\-depth loss improves the hand design, and a head\-to\-head of two otherwise\-identical training runs shows that a cross\-family training judge \(drawn from a different vendor than the generator\) dominates an in\-family one on every measured axis—larger effect sizes, shorter outputs, and a reduced cross\-vendor judge\-generosity bias—yielding a one\-line recommendation for prompt\-optimisation practice \(§[4\.2](https://arxiv.org/html/2606.26366#S4.SS2)\)\. Fourth, extended to a five\-round multi\-stakeholder protocol that ends in a binary accept/reject vote on a single moderator\-built proposal addressing all three agents’ modifications, the scaffold converts a6%6\\%debate standoff into95%95\\%full consensus while leaving a small1\.6%1\.6\\%of votes as rejections the integrated proposal could not absorb; the pattern replicates on a 30\-scenario DailyDilemmas sample with two generators at100%100\\%combined convergence and no rejected proposals \(§[5](https://arxiv.org/html/2606.26366#S5)\)\. Fifth, a graph\-complexity proxy of algorithmic causal complexityKCK\_\{C\}\(the description length of the trace’s underlying structural\-causal model, SCM\) achieves pooled Spearmanρ=0\.42\\rho\{=\}0\.42\(p<0\.001p\{<\}0\.001,n=976n\{=\}976traces\) against the NoT direction, confirmed by two further SCM\-level proxies \(§[6](https://arxiv.org/html/2606.26366#S6)\)\. The resulting trace externalises stakeholders, consequences, and uncertainties in an auditable, rule\-governed form; the mapping from NoT’s five sub\-instructions to the recognised primitives of human ethical deliberation is developed in §[6](https://arxiv.org/html/2606.26366#S6), and the downstream\-probe sets that motivated the diagnosis \(sycophancy, agentic misalignment\) saturate under either scaffold on the Sharma SycophancyEval panel \(Appendix[E\.1](https://arxiv.org/html/2606.26366#A5.SS1)\); a follow\-on ELEPHANT social\-sycophancy benchmark\(Chenget al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib7)\)retains headroom and shows NoT lowering framing and validation rates relative to standard CoT on held\-out advice queries \(Appendix[E\.2](https://arxiv.org/html/2606.26366#A5.SS2)\)\. Agentic misalignment probes remain saturated \(Appendix[E](https://arxiv.org/html/2606.26366#A5)\)\.
## 2Background and Related Work
#### Chain\-of\-thought and its failure modes\.
Chain\-of\-thought prompting\(Weiet al\.,[2022](https://arxiv.org/html/2606.26366#bib.bib53); Kojimaet al\.,[2022](https://arxiv.org/html/2606.26366#bib.bib54)\)elicits step\-by\-step intermediate reasoning but does not force the trace to track situational structure\.Yanet al\.\([2025](https://arxiv.org/html/2606.26366#bib.bib6)\)show that as reasoning capacity grows, models become more responsive to harmful requests because reasoning and safety entangle in shared neuron populations \(“reasoning\-induced misalignment”\)\. NoT is a constrained chain\-of\-thought that re\-anchors the trace to a narrative\-causal structure before the reasoning capacity is exercised; §[4](https://arxiv.org/html/2606.26366#S4)shows the intervention works on a current reasoning model as well as a non\-reasoning one\.
#### Narrative scaffolding for reasoning\.
Visualisation\-of\-Thought\(Wuet al\.,[2024](https://arxiv.org/html/2606.26366#bib.bib57)\)shows that externalising reasoning in the native representational format of a domain \(spatial diagrams for spatial tasks\) improves reliability\. We apply the same principle to ethical reasoning: the native format is a story with named agents, projected consequences, and uncertainty\(MacIntyre,[1981](https://arxiv.org/html/2606.26366#bib.bib19); Bruner,[1986](https://arxiv.org/html/2606.26366#bib.bib20)\)\. Narrative\-perspective interventions in medical decision contexts measurably increase perspective\-taking\(Shafferet al\.,[2019](https://arxiv.org/html/2606.26366#bib.bib40); Bientzleet al\.,[2021](https://arxiv.org/html/2606.26366#bib.bib41),[2024](https://arxiv.org/html/2606.26366#bib.bib42)\), and generative agents show long\-horizon coherent behaviour is achievable from narrated character state\(Parket al\.,[2023](https://arxiv.org/html/2606.26366#bib.bib43)\)\.
#### Causal inference from narrative\.
The Computational Theory of Mind\(Fodor,[1975](https://arxiv.org/html/2606.26366#bib.bib30); Putnam,[1967](https://arxiv.org/html/2606.26366#bib.bib31)\)treats understanding a sequence of events as inferring a structural causal model that explains them\. LLMs perform near random baseline on causal inference stripped of empirical pattern matches\(Jinet al\.,[2024](https://arxiv.org/html/2606.26366#bib.bib18)\): pretraining supplies pattern\-matching over narrative token sequences, not the structured causal reasoning that determines how those narratives resolve\. NoT converts that pattern matching into a usable causal trace at inference time, projecting what each available action will mean for each named stakeholder and what remains genuinely uncertain about each branch\.
#### Multi\-agent debate and defeasibility\.
Irvinget al\.\([2018](https://arxiv.org/html/2606.26366#bib.bib44)\)proposed AI safety via debate;Duet al\.\([2023](https://arxiv.org/html/2606.26366#bib.bib45)\)showed multi\-agent debate improves factual reasoning\. Both target task\-correctness, not stakeholder\-divergent value standoffs\. Defeasible reasoning\(Pollock,[1987](https://arxiv.org/html/2606.26366#bib.bib46); MacIntyre,[1981](https://arxiv.org/html/2606.26366#bib.bib19)\)has a long lineage in moral philosophy; the protocol of §[5](https://arxiv.org/html/2606.26366#S5)converts the property into measurable behaviour, with acceptance of the integrated proposal demonstrating revisability and residual rejections marking commitments no revision absorbs\.
#### Multi\-agent LLM systems for scientific discovery\.
A growing line of work scaffolds LLM reasoning into the primitives of a target domain\(Gottweiset al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib47); Luet al\.,[2024](https://arxiv.org/html/2606.26366#bib.bib48); Boikoet al\.,[2023](https://arxiv.org/html/2606.26366#bib.bib49); M\. Branet al\.,[2024](https://arxiv.org/html/2606.26366#bib.bib50); Schmidgallet al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib51); Swansonet al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib52)\)\. Our protocols ask the analogous question for ethical reasoning\. Because ethics offers no experimental ground truth to rank competing positions, the moderator’s integration\-then\-veto step replaces the ranking\-then\-selection of those systems, and the residual veto pattern itself becomes the measurement of interest\.
## 3Narration\-of\-Thought
Narration\-of\-thought \(NoT\) is a system prompt that instructs the model to produce, in first person, a five\-section trace before any final answer: \(1\) characterise the protagonist \(name, role, what they know\); \(2\) enumerate the stakeholders whose lives intersect the decision and what is at stake for each; \(3\) project the consequences of each available action at least two steps out for each stakeholder; \(4\) state what remains uncertain about each projected future; \(5\) commit to a decision and explain why, within the narrative frame, that trajectory is preferable to the alternatives\.
The intervention is a single change to the system prompt; it adds no training data, parameters, or fine\-tuning, and it does not mention causal graphs, utilities, or any formal apparatus\. The mechanism it exploits is that text matching the five\-section structure is densely represented in the narrative subdistribution of the pretraining corpus, so the model can amortise causal\-trajectory inference by retrieval rather than by reasoning from first principles\. Algorithm[1](https://arxiv.org/html/2606.26366#alg1)states the procedure\.
Algorithm 1Narration\-of\-thought \(NoT\): generation and rubric coding for a single dilemma\. The concatenation operator⊕\\opluson line 8 joins the NoT system prompt with the dilemma text to form the model input\.1:dilemma
ss; generator
GG; judge
JJ; decoder
Θ\\Theta
2:trace
ttand coded metrics
𝐜\\mathbf\{c\}
3:Build NoT system prompt
πNoT\\pi\_\{\\mathrm\{NoT\}\}requesting five first\-person sections:
4:\(1\)Protagonist: name role and epistemic state
5:\(2\)Stakeholders: parties intersecting the decision
6:\(3\)Consequences: each action
≥2\\geq 2steps forward
7:\(4\)Uncertainty: what remains genuinely unknown
8:\(5\)Commitment: chosen action and narrative warrant
9:Generate trace
10:
t←G\(πNoT⊕s;Θ\)t\\leftarrow G\(\\pi\_\{\\mathrm\{NoT\}\}\\oplus s;\\,\\Theta\)
11:Code metricswith judge
JJ’s rubric
12:
\(𝑠𝑐,𝑚ℎ,𝑢𝑠\)←J\(t\)\(\\mathit\{sc\},\\mathit\{mh\},\\mathit\{us\}\)\\leftarrow J\(t\)
13:// stakeholder count, max causal hops, uncertainty score
14:Derive failure\-mode tags\(deterministic\)
15:StakeholderCollapse
←𝑠𝑐≤1\\leftarrow\\mathit\{sc\}\\leq 1
16:UncertaintySuppression
←𝑢𝑠=0\\leftarrow\\mathit\{us\}=0
17:return
\(t,𝐜\)\(t,\\mathbf\{c\}\)
### 3\.1Multi\-Stakeholder Narrative Deliberation
A single NoT trace commits one protagonist to a position; many real decisions involve multiple narrators with conflicting stakes\. We extend NoT into a five\-round protocol \(Rounds 0–4\) that re\-uses NoT as the per\-agent generator\. Three stakeholder perspectives \(formal decider, primary affected party, third party\) each produce an NoT statement \(Round 0\), exchange rebuttals \(Round 1\), and restate a final position \(Round 2\)\. The moderator reads all three Round 2 positions and writes a single synthesis, which each agent labelsACCEPT,ACCEPT\_WITH\_MODIFICATION, orREJECT\(Round 3\)\. The moderator then builds a second proposal that explicitly addresses all three agents’ modification requests, and each agent casts a binaryACCEPTorREJECTvote on that single integrated proposal \(Round 4\)\. The Round 4 vote is what makes defeasibility \(the property that a conclusion can be revised when a counter\-proposal addresses the original objection\) measurable: an agent that accepts after its modifications are addressed has revised, and an agent that still rejects is signalling a position no integration can absorb\. In deployment this distinction routes the vote: an agentic workflow escalates the rejections the moderator cannot absorb to a human and absorbs procedural friction at the moderator, so the operator sees only the objections that actually contest the outcome\.
## 4Experiment 1: NoT at Scale
#### Setup\.
Three prompting conditions are compared: bare input/output, standard CoT \(“Think step by step, then give your answer”\), and NoT \(the five\-section narrative scaffold of §[3](https://arxiv.org/html/2606.26366#S3)\)\. Decoding parameters are held identical across conditions\. Four generators are tested:gpt\-5\.4\-nano\(N=20N\{=\}20\) from the original pilot, plus three new generators spanning two additional vendors:claude\-haiku\-4\-5\(N=20N\{=\}20\),grok\-4\-1\-fast\-reasoning\(N=5N\{=\}5\), and the flagshipclaude\-sonnet\-4\-6\(N=5N\{=\}5, cost\-controlled\)\. Here,NNis the number of independent samples per \(scenario, condition, generator\) cell drawn at temperature 1; with the 100\-scenario sample and three conditions this yields300N300Ngenerations per generator before the cross\-judge coding stage\. Coded sample sizes \(Ns,NnN\_\{\\textsc\{s\}\},N\_\{\\textsc\{n\}\}in Table[1](https://arxiv.org/html/2606.26366#S4.T1)\) are slightly smaller because the rubric drops generations where the judge could not extract a coded value\. Outputs are coded by two cross\-vendor judges \(claude\-haiku\-4\-5primary,gpt\-5\.4\-nanosecondary\) on a six\-variable rubric with quadratic\-weighted Cohen’sκ\\kappa\(Cohen,[1960](https://arxiv.org/html/2606.26366#bib.bib59)\)per variable\. Two rubric variables anchor every effect\-size claim in this section:*stakeholder count*𝑠𝑐\\mathit\{sc\}, the number of distinct parties named in the trace whose interests intersect the decision, and*uncertainty score*𝑢𝑠\\mathit\{us\}, the number of explicit hedge or unknown spans named in the trace; the deterministic failure\-mode tags are Algorithm[1](https://arxiv.org/html/2606.26366#alg1)’s thresholds on these two variables \(full rubric: Appendix[A](https://arxiv.org/html/2606.26366#A1)\)\. The scenario pool is a 100\-scenario stratified sample of DailyDilemmas\(Chiuet al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib55)\)drawn deterministically \(seed 42\), with stratification spanning all 18 topic groups in the source corpus\. Per\-cell coded sample sizes of500500to2,0002\{,\}000\(Table[1](https://arxiv.org/html/2606.26366#S4.T1)\) drive the bootstrap confidence intervals on the effect sizes, so statistical power on the per\-cell contrast follows fromNNrather than from scenario count alone\.
Dilemma\.A project manager on a tight deadline learns their most\-relied\-on team member has fallen seriously ill: stick to the plan, or delay?Figure 2:Trace\-level instance of the two failure modes whose panel\-level rates Table[1](https://arxiv.org/html/2606.26366#S4.T1)reports \(scenariodd\_14845, generatorgpt\-5\.4\-nano; verbatim model output with\[…\]marking elision\)\. The NoT text is the model’s verbatim completion under the NoT system prompt; the narrator, named stakeholders, and hedging language are the model filling in the five sections the system prompt asks for \(protagonist, stakeholders, consequences, uncertainty, commitment\), not text the authors wrote\.Table 1:Cell sizes and percentage\-point drops for the two failure modes that empirically fire \(Δsc\\Delta\_\{\\textsc\{sc\}\}= stakeholder collapse,Δus\\Delta\_\{\\textsc\{us\}\}= uncertainty suppression; both NoT minus standard CoT, in percentage points;Ns,NnN\_\{\\textsc\{s\}\},N\_\{\\textsc\{n\}\}are coded samples under standard CoT and NoT, respectively\)\. Vendor tiers are b = budget, f = flagship\. Figure[3](https://arxiv.org/html/2606.26366#S4.F3)plots the raw rates with the same drop annotations\.Figure 3:Failure\-mode firing rates \(standard CoT vs\. NoT\) across the four\-model panel \(gpt\-5\.4\-nano,claude\-haiku\-4\-5,grok\-4\-1\-fast\-reasoning,claude\-sonnet\-4\-6\)\. The suppression pattern from the original pilot replicates across vendors and model tiers; uncertainty suppression drops by 28–72 percentage points and stakeholder collapse drops to near zero on every model\.
#### NoT cuts both failure modes on every model\.
Uncertainty suppression fires on50\.050\.0–72\.3%72\.3\\%of standard\-CoT outputs across the four\-model panel \(Table[1](https://arxiv.org/html/2606.26366#S4.T1)\) and stakeholder collapse on14\.614\.6–30\.6%30\.6\\%; under NoT both drop sharply \(collapse to0\.00\.0–1\.2%1\.2\\%, suppression to0\.80\.8–24\.5%24\.5\\%; Figure[3](https://arxiv.org/html/2606.26366#S4.F3)\)\. Figure[2](https://arxiv.org/html/2606.26366#S4.F2)shows the transition at the trace level\. The reduction is largest where the standard\-CoT rate is largest, and the pattern is cross\-vendor: no vendor escapes either the baseline or the intervention\.gpt\-5\.4\-nanodrops both modes to<1%\{<\}1\\%\. The two Anthropic generators retain a2020–25%25\\%residual on uncertainty suppression that the prompt\-side intervention does not reach, indicating a complementary training\-time component\.
#### The shift is large, distribution\-wide, and not a length artefact\.
The firing rates in Table[1](https://arxiv.org/html/2606.26366#S4.T1)are binary thresholds; one might worry NoT only nudges borderline cases across the cutoff\. Cliff’sδ\\delta\(Cliff,[1993](https://arxiv.org/html/2606.26366#bib.bib58)\)on the continuous coded metrics settles that worry:δ=\+0\.48\\delta=\+0\.48means an NoT output exceeds the matched standard\-CoT output in74%74\\%of pairs,δ=\+0\.99\\delta=\+0\.99in99\.5%99\.5\\%\. NoT outputs carryδ\\deltaof\+0\.48\+0\.48to\+0\.99\+0\.99on stakeholder count and\+0\.63\+0\.63to\+0\.99\+0\.99on uncertainty score, large\-to\-near\-maximal effects that rule out threshold\-gaming\. A second alternative explanation is length, since NoT outputs are55–7×7\\timeslonger than standard\-CoT outputs and the simplest competing explanation is that more text mechanically has more room to name stakeholders and to hedge\. After regressing each metric onlog\(output length\)\\log\(\\text\{output length\}\)pooled across conditions and re\-evaluatingδ\\deltaon residuals, the shift in the coded metrics survives length removal on the OpenAI and xAI generators \(gpt\-5\.4\-nanoδresid=\+0\.51\\delta\_\{\\text\{resid\}\}=\+0\.51on stakeholder count and\+0\.77\+0\.77on uncertainty score;grok\-4\-1\-fast\-reasoning\+0\.33\+0\.33and\+0\.43\+0\.43; all CIs strictly above zero\), ruling out length as the sole mechanism where the firing\-rate drops were largest\. On the Anthropic generators the residualised effect is indistinguishable from zero, which we read as those models already producing stakeholder names and hedges per unit length under standard CoT at the rate NoT elicits, so the prompt\-side gain there operates through additional length rather than additional per\-unit density\. Per\-generatorδ\\deltas with bootstrap CIs are in Appendix[B](https://arxiv.org/html/2606.26366#A2)\.
#### Length is a cost, not a substitute for the prompt\.
A sceptical reading of the55–7×7\\timestoken premium is that the paper just spends more compute\. A direct matched\-budget experiment rules this out by construction\. A verbose standard\-CoT prompt \(“think step by step in detail, exploring multiple angles from every relevant perspective, articulating uncertainty before committing”\) is run on the same 100 scenarios across all four generators withmax\_tokensset to each generator’s9090th percentile NoT length \(N=1N\{=\}1per cell\)\. Verbose CoT clears the binary failure\-mode thresholds on every generator \(Table[2](https://arxiv.org/html/2606.26366#S4.T2), SC% and US% columns; well below standard\-CoT baselines of1515–31%31\\%SC and5050–72%72\\%US\), so verbosity\-with\-perspective\-prompting does part of the job\. The continuous variables tell the cleaner story: NoT retains Cliff’sδ\\deltaof\+0\.79\+0\.79to\+0\.90\+0\.90on stakeholder count and\+0\.65\+0\.65to\+0\.93\+0\.93on uncertainty score for three of four generators at matched budget\. The exception isgrok\-4\-1\-fast\-reasoning, whose verbose\-CoT output \(2\.33×2\.33\\timesstandard CoT\) actually exceeds its NoT output \(1\.55×1\.55\\times\) under the calibrated budget; the gap between the two conditions on the coded metrics narrows as expected when the two conditions allocate similar token counts\. The narrative scaffold, not raw budget, is the active ingredient where the two conditions differ in trace shape \(§[4\.2](https://arxiv.org/html/2606.26366#S4.SS2)and Appendix[G](https://arxiv.org/html/2606.26366#A7)take up the complementary question of whether the scaffold itself can be improved by automatic search\)\. Inference tokens are also the cheapest scaling axis: no training data, parameters, or fine\-tuning, small relative to RLHF or constitutional training whose fixed costs are orders of magnitude higher\.
Table 2:Matched\-budget control: Cliff’sδ\\deltafor NoT vs\. verbose\-CoT on stakeholder count \(δsc\\delta\_\{\\textsc\{sc\}\}\) and uncertainty score \(δus\\delta\_\{\\textsc\{us\}\}\), and verbose\-CoT firing rates for stakeholder collapse \(SC%\) and uncertainty suppression \(US%\) on 100 DailyDilemmas scenarios \(N=1N\{=\}1per cell\)\. Verbose CoT suppresses both failure modes relative to standard CoT baselines \(1515–31%31\\%SC,5050–72%72\\%US, Table[1](https://arxiv.org/html/2606.26366#S4.T1)\) but does not match NoT’s continuous shift in the coded metrics on three of four generators\.
### 4\.1Sub\-Instruction Ablation
To identify which sub\-instruction of the NoT prompt carries the shift in each coded metric, we run a sub\-instruction ablation onclaude\-sonnet\-4\-6\(N=3N\{=\}3, 30\-scenario stratified subsample, seed 43; cost\-controlled\)\. Six conditions are compared: the full NoT prompt \(control\) and five variants each dropping exactly one of the five sub\-instructions\. Figure[4](https://arxiv.org/html/2606.26366#S4.F4)reports Cliff’s deltas for each drop condition relative to the full NoT control on the two headline coded metrics\.
Figure 4:Sub\-instruction ablation onclaude\-sonnet\-4\-6: Cliff’sδ\\deltafor each drop\-one\-section condition vs\. the full NoT control on stakeholder count and uncertainty score\. Negativeδ\\deltameans the sub\-instruction raised the coded metric, so dropping it reduces it\. The Stakeholders sub\-instruction \(sub\-instruction 2 of the NoT prompt\) and the Uncertainty sub\-instruction \(sub\-instruction 4\) show the largest negativeδ\\deltaon their respective target metrics\.The ablation shows direct causal attribution\. Dropping Stakeholders reduces stakeholder count byδ=−0\.41\\delta=\-0\.41while leaving causal\-hop depth and uncertainty essentially unchanged \(\|δ\|<0\.03\|\\delta\|<0\.03\); dropping Uncertainty reduces uncertainty score byδ=−0\.92\\delta=\-0\.92; dropping Consequences reduces causal\-hop depth byδ=−0\.22\\delta=\-0\.22with small spillover onto stakeholder count, consistent with multi\-step consequence projection forcing the model to name the entities those consequences fall on\. Dropping Protagonist or Commitment contributes no signed effect\. The diagonal pattern rules out two alternatives: a single emergent narrative factor would dilute every metric on any drop, and a pure length confound would shrink one column most regardless of which sub\-instruction is dropped\. Each substantive sub\-instruction carries the coded metric attributed to it and little else \(full table: Appendix[F](https://arxiv.org/html/2606.26366#A6)\)\.
### 4\.2Optimising the Scaffold: Cross\-Family Textual\-Gradient Descent
NoT is hand\-designed, which invites the question of whether automatic search can improve it and whether the search itself depends on the judge it uses\. We apply textual\-gradient descent\(Yuksekgonulet al\.,[2024](https://arxiv.org/html/2606.26366#bib.bib63)\)—an optimiser LLM \(claude\-sonnet\-4\-6\) that reads the current prompt and a batch of its judge\-coded outputs and writes a diagnosis plus a rewritten prompt—initialised at NoT under a continuous depth lossℓ=max\(0,4−𝑠𝑐\)\+max\(0,2−𝑢𝑠\)\\ell=\\max\(0,4\-\\mathit\{sc\}\)\+\\max\(0,2\-\\mathit\{us\}\)that stays informative once the binary failure modes fire at zero\. We run the loop twice, identical except for the training judge:*in\-family*\(generator and training judge bothclaude\-haiku\-4\-5, the vendor the primary judge uses\) yieldsNoT\-v2;*cross\-family*\(training judgegrok\-4\-1\-fast\-reasoning, a different vendor\) yieldsNoT\-v3\. Both are replicated across the four\-generator panel and re\-coded by the primary judge and an adversarial cross\-vendor third judge\.
Both optimised prompts beat the hand design, and cross\-family training dominates in\-family training on every axis we measured\.NoT\-v3’s stakeholder\-count Cliff’sδ\\deltaover NoT equals or exceeds NoT\-v2’s on all four generators \(e\.g\.−0\.57→−0\.73\-0\.57\\\!\\to\\\!\-0\.73onhaiku,−0\.68→−0\.93\-0\.68\\\!\\to\\\!\-0\.93ongrok; the lonenanoregression of\+0\.43\+0\.43under v2 falls to a non\-significant\+0\.12\+0\.12\), while producing*shorter*outputs on every generator and preserving the per\-unit\-length effect\. The in\-family judge\-generosity gap visible under v2—the primary \(Anthropic\) judge counting0\.800\.80more stakeholders than the third judge on the same in\-familyhaikuoutputs—shrinks to0\.460\.46under v3 and inverts ongrok\. The recommendation is methodological: when textual\-gradient optimisation targets cross\-vendor deployment, draw the training judge from a different vendor than the generator, at no extra training cost\. Run in the opposite direction—descending from plain standard CoT rather than from NoT—the same optimiser fails to recover NoT at either the single\-agent or the multi\-stakeholder layer, confirming the scaffold is the object worth optimising\. Method, the head\-to\-head figure, full tables, verbatim prompts, and training curves are in Appendix[G](https://arxiv.org/html/2606.26366#A7)\.
## 5Experiment 2: Multi\-Stakeholder Narrative Deliberation
#### Setup\.
Experiment 2 evaluates the multi\-stakeholder protocol of §[3\.1](https://arxiv.org/html/2606.26366#S3.SS1)on160160designed debates across two complementary sources:100100debates on a five\-scenario calibration set built to stress\-test stakeholder conflict \(55scenarios×2\\times\\,2generators×10\\times\\,10samples\) plus a6060\-debate DailyDilemmas replication \(3030scenarios×2\\times\\,2generators×1\\times\\,1sample, seed4343\)\. Generators aregpt\-5\.4\-nanoandgpt\-4oon the calibration set,gpt\-5\.4\-nanoandclaude\-sonnet\-4\-6on the replication; the moderator isgpt\-4o\-minithroughout, and Round 0 reuses cached single\-agent NoT statements\. Swapping the moderator forclaude\-sonnet\-4\-6on the calibration debates leaves the full\-consensus rate within sampling noise of thegpt\-4o\-miniresult \(Appendix[A](https://arxiv.org/html/2606.26366#A1)\), so the consensus signal is not moderator\-specific\.
Figure 5:Consensus rates across four progressively more structured debate designs \(five scenarios, two generators\)\.*Closed taxonomy*\(agents must choose from a fixed action list\) and*open action space*\(agents may propose novel actions\) hold at6%6\\%and9%9\\%full consensus\. Adding a moderator\-built synthesis presented back to the agents \(*synthesis presentation*, Round 3\) reaches100%100\\%partial convergence \(≥2/3\\geq 2/3agreement\)\. The full protocol, which integrates the three agents’ modification requests into a single proposal and forces a binary accept/reject vote \(*integrative vote*, Round 4\), reaches95%95\\%full consensus \(gpt\-5\.4\-nano90%90\\%,gpt\-4o100%100\\%\) with mean2\.982\.98of33modifications addressed per debate\. The residual1\.6%1\.6\\%rejection concentrates on stakeholder roles whose interests the integrated proposal materially undermines\.
#### Progressive structuring converts a categorical standoff into near\-universal consensus\.
Figure[5](https://arxiv.org/html/2606.26366#S5.F5)reports the convergence arc across four protocol designs\. With a closed action taxonomy and three rounds of pure debate, only6%6\\%of debates reach full consensus\. Opening the action space and authorising the moderator to surface synthesis raises consensus to9%9\\%while revealing the underlying mechanism: agents propose novel actions in73%73\\%of final\-round outputs and a coherent moderator\-built synthesis emerges in82%82\\%of debates, but the agents do not self\-coordinate onto those syntheses\. Presenting the synthesis back to the agents \(Round 3\) drives outright rejection to0%0\\%and partial convergence \(≥2/3\\geq 2/3agreement\) to100%100\\%, while full consensus stays at0%0\\%because every agent demands a different modification\. Integrating those modifications into a single proposal and forcing a binary accept/reject vote \(Round 4\) produces78/82=95\.1%78/82=95\.1\\%full consensus and242/246=98\.4%242/246=98\.4\\%individual acceptance \(Wilson95%95\\%CIs\[88\.1,98\.1\]%\[88\.1,98\.1\]\\%and\[95\.9,99\.4\]%\[95\.9,99\.4\]\\%\), with the integrated proposal addressing a mean2\.982\.98of33agent modifications per debate\. Of100100designed debates \(55scenarios×2\\times\\,2generators×10\\times\\,10samples\),8282produced a synthesis and proceeded to Round 4, yielding246246binary votes\.
#### The residual rejections are interpretable\.
The vote produces44rejections out of246246\(1\.6%1\.6\\%, Wilson\[0\.6,4\.1\]%\[0\.6,4\.1\]\\%\), all in theprimary\_affectedandthird\_partyroles on debates where honouring one stakeholder’s modification directly undermines another’s stake \(pharma\_whistlebloweron the senior colleague role,av\_engineeron the future\-pedestrian role\)\. This is the shape a defeasibility\-respecting protocol should produce, and is what distinguishes a dependable multi\-agent consensus from uniform compliance: a downstream auditor sees not just the winning proposal but which stakeholders refused and why\.
#### Scale replication on DailyDilemmas\.
A two\-generator DailyDilemmas replication \(gpt\-5\.4\-nanoandclaude\-sonnet\-4\-6,N=1N\{=\}1, 30 scenarios, seed 43\) confirms the pattern\. Of the6060scenario–generator debates,3838required the synthesis\-and\-integrated\-vote stage and3535of those \(92\.1%92\.1\\%, Wilson95%95\\%CI\[79\.2,97\.3\]%\[79\.2,97\.3\]\\%\) reached full consensus, within±10\\pm 10pp of the calibration\-set95\.1%95\.1\\%\. The remaining2222debates reached three\-way agreement before integration was needed, so combined convergence is60/60=100%60/60=100\\%across both generators with no rejected proposals\.
## 6Discussion
#### Deliberative primitives and causal grounding\.
Defensible ethical deliberation shares a recognisable structure: defeasible revisability under new information\(Pollock,[1987](https://arxiv.org/html/2606.26366#bib.bib46)\), identification of who is affected and what is at stake\(MacIntyre,[1981](https://arxiv.org/html/2606.26366#bib.bib19)\), and value\-laden reasoning in narrative form\(Bruner,[1986](https://arxiv.org/html/2606.26366#bib.bib20)\)\. NoT reifies these primitives at the single\-agent layer; the multi\-stakeholder protocol reifies them at the social layer through perspectival narration, moderator integration, and a binary vote whose residual rejections mark positions no integration can absorb\. By forcing the model to name stakeholders, trace consequences, and articulate uncertainty before committing, NoT drives the trajectory toward the narrated causal model\(Pearl,[2009](https://arxiv.org/html/2606.26366#bib.bib24),[1995](https://arxiv.org/html/2606.26366#bib.bib25)\)of lowest algorithmic complexityKCK\_\{C\}\(Yudkowsky,[2011](https://arxiv.org/html/2606.26366#bib.bib22); Li and Vitányi,[2019](https://arxiv.org/html/2606.26366#bib.bib29)\); Finding 4’sρ=0\.42\\rho=0\.42correlation is the empirical face of this \(convergent SCM\-level proxies and length\-invariant audits: Appendix[C\.2](https://arxiv.org/html/2606.26366#A3.SS2)\)\.
## 7Conclusion
A single\-sentence change to the system prompt drives stakeholder collapse below1%1\\%and uncertainty suppression by2828–7272percentage points on four frontier generators, each shift attributable to a specific NoT sub\-instruction\. Textual\-gradient descent initialised at the scaffold improves it further, and a head\-to\-head of two training\-judge configurations identifies cross\-family training—a judge drawn from a different vendor than the generator—as the configuration that survives cross\-vendor evaluation best, a portable recommendation for LLM\-judge prompt optimisation\. The same scaffold extended to a multi\-stakeholder protocol drives a6%6\\%debate standoff to95%95\\%consensus and100%100\\%combined convergence on a DailyDilemmas replication, producing the auditable, revisable surface agentic deployment requires\.
## Limitations
All experiments use the DailyDilemmas ethics corpus\(Chiuet al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib55)\), a collection of everyday personal and civic dilemmas\. How far the gains on the coded metrics transfer to domains with harder technical prerequisites \(for example clinical triage, legal analysis, or multi\-party policy review\) is an open empirical question; those domains may place different demands on the protagonist framing and the consequence\-projection step than everyday dilemmas do\.
Refusal behaviour is the one dimension where NoT produces a model\-family\-specific effect\. On XSTest\(Röttgeret al\.,[2023](https://arxiv.org/html/2606.26366#bib.bib61)\)\(prompts that resemble unsafe requests but are not\) and SimpleSafetyTests\(Vidgenet al\.,[2024](https://arxiv.org/html/2606.26366#bib.bib62)\)\(genuinely unsafe prompts\), Anthropic generators andgrok\-4\-1\-fast\-reasoningshow no significant change under either condition\.gpt\-5\.4\-nanobecomes more cautious: over\-refusal on XSTest rises from13\.6%13\.6\\%to23\.6%23\.6\\%and appropriate refusal on SimpleSafetyTests rises from51%51\\%to68%68\\%\. This is a per\-model calibration consideration, not a weakness of the scaffold; details and the full refusal table are in Appendix[D](https://arxiv.org/html/2606.26366#A4)\.
The scaffold\-optimisation result \(§[4\.2](https://arxiv.org/html/2606.26366#S4.SS2), Appendix[G](https://arxiv.org/html/2606.26366#A7)\) carries its own caveats\. Textual gradients and rewrites come from a single optimiser model \(claude\-sonnet\-4\-6\), so we do not separate the optimiser’s contribution from the training\-judge configuration; the loss targets only stakeholder count and uncertainty score, with max causal hops reported as an out\-of\-loss generalisation check; the cross\-family training judge and the adversarial evaluation third judge are the same model \(the only non\-Anthropic non\-OpenAI judge on our panel\), so the cross\-family claim is the narrower one stated in Appendix[G](https://arxiv.org/html/2606.26366#A7); and fornanoandhaikuonly3030and254254hand\-NoT baseline judge cells survived the cache rebuild \(versus the full500500and20002000\), so the v2/v3 effect sizes are well\-powered but the NoT baseline means carry higher variance than the optimised ones\.
## Ethics Statement
The work uses an existing public corpus \(DailyDilemmas\(Chiuet al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib55)\)\) and the publicly described Anthropic agentic\-misalignment scenario structures\(Lynchet al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib5)\); all agentic\-probe scenarios are entirely fictional, no personally identifying information is used, and no human raters are employed\. Generation, judging, and analysis code is in the anonymised repository \([https://github\.com/PatrickAllenCooper/ANI\_Computational\_Narratology](https://github.com/PatrickAllenCooper/ANI_Computational_Narratology)\)\. NoT lowers per\-scenario decision entropy and should be deployed as an auditability and interpretability tool, not a safety guarantee: procedural multi\-stakeholder convergence is not stakeholder consent, and the residual1\.6%1\.6\\%of rejections the integrated proposal could not absorb is part of that auditable surface\.
## Reproducibility
All generation, judging, and analysis code and the per\-cell artefacts \(raw outputs, both judges’ codings, decision\-extractor outputs\) are versioned in the anonymised repository \([https://github\.com/PatrickAllenCooper/ANI\_Computational\_Narratology](https://github.com/PatrickAllenCooper/ANI_Computational_Narratology)\); the pipeline is deterministic under fixed seeds modulo upstream API non\-determinism, with per\-cell cache keys over generator and judge so adding either does not invalidate prior results\.
## References
- M\. Bientzle, U\. Cress, and J\. Kimmerle \(2024\)Narrative Persuasion in Health Communication\.Patient Education and Counseling\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Bientzle, M\. Eggeling, U\. Cress, and J\. Kimmerle \(2021\)The Effects of Narrative Video on Viewers’ Understanding of and Attitudes towards Organ Donation\.Health Communication36\(7\),pp\. 820–829\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px2.p1.1)\.
- D\. A\. Boiko, R\. MacKnight, B\. Kline, and G\. Gomes \(2023\)Autonomous Chemical Research with Large Language Models\.Nature624\(7992\),pp\. 570–578\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px5.p1.1)\.
- J\. Bruner \(1986\)Actual Minds, Possible Worlds\.Harvard University Press\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2606.26366#S6.SS0.SSS0.Px1.p1.2)\.
- M\. Cheng, E\. Durmus, M\. Zhang, T\. Korbak, E\. Perez, S\. R\. Bowman,et al\.\(2025\)ELEPHANT: measuring and understanding social sycophancy in LLMs\.arXiv preprint arXiv:2505\.13995\.Cited by:[§E\.2](https://arxiv.org/html/2606.26366#A5.SS2.p1.2),[Table 10](https://arxiv.org/html/2606.26366#A5.T10),[§1](https://arxiv.org/html/2606.26366#S1.p5.11)\.
- Y\. Y\. Chiu, L\. Jiang, and Y\. Choi \(2025\)DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2410\.02683Cited by:[§1](https://arxiv.org/html/2606.26366#S1.p2.5),[§4](https://arxiv.org/html/2606.26366#S4.SS0.SSS0.Px1.p1.13),[Limitations](https://arxiv.org/html/2606.26366#Sx1.p1.1),[Ethics Statement](https://arxiv.org/html/2606.26366#Sx2.p1.1)\.
- N\. Cliff \(1993\)Dominance Statistics: Ordinal Analyses to Answer Ordinal Questions\.Psychological Bulletin114\(3\),pp\. 494–509\.Cited by:[§G\.1](https://arxiv.org/html/2606.26366#A7.SS1.SSS0.Px3.p1.9),[§4](https://arxiv.org/html/2606.26366#S4.SS0.SSS0.Px3.p1.19)\.
- J\. Cohen \(1960\)A Coefficient of Agreement for Nominal Scales\.Educational and Psychological Measurement20\(1\),pp\. 37–46\.Cited by:[Appendix A](https://arxiv.org/html/2606.26366#A1.SSx1.p1.11),[§G\.1](https://arxiv.org/html/2606.26366#A7.SS1.SSS0.Px3.p1.9),[§4](https://arxiv.org/html/2606.26366#S4.SS0.SSS0.Px1.p1.13)\.
- Y\. Du, S\. Li, A\. Torralba, J\. B\. Tenenbaum, and I\. Mordatch \(2023\)Improving Factuality and Reasoning in Language Models through Multiagent Debate\.arXiv preprint arXiv:2305\.14325\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px4.p1.1)\.
- J\. A\. Fodor \(1975\)The Language of Thought\.Harvard University Press\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Gottweis, W\. Weng, A\. Daryin, T\. Tu, A\. Palepu, P\. Sirkovic, A\. Myaskovsky, F\. Weissenberger, K\. Rong, R\. Tanno, K\. Saab, D\. Popovici, J\. Blum, F\. Zhang, K\. Chou, A\. Hassidim, B\. Gokturk, A\. Vahdat, P\. Kohli, Y\. Matias, A\. Carroll, K\. Kulkarni, N\. Tomasev, Y\. Guan, V\. Dhillon, E\. D\. Vaishnav, B\. Lee, T\. R\. D\. Costa, J\. R\. Penedés, G\. Peltz, Y\. Xu, A\. Pawlosky, A\. Karthikesalingam, and V\. Natarajan \(2025\)Towards an AI Co\-Scientist\.arXiv preprint arXiv:2502\.18864\.Cited by:[§1](https://arxiv.org/html/2606.26366#S1.p4.1),[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px5.p1.1)\.
- G\. Irving, P\. Christiano, and D\. Amodei \(2018\)AI Safety via Debate\.arXiv preprint arXiv:1805\.00899\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px4.p1.1)\.
- Z\. Jin, J\. Liu, Z\. Lyu, S\. Poff, M\. Sachan, R\. Mihalcea, M\. Diab, and B\. Schölkopf \(2024\)Can Large Language Models Infer Causation from Correlation?\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2306\.05836Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px3.p1.1)\.
- T\. Kojima, S\. S\. Gu, M\. Reid, Y\. Matsuo, and Y\. Iwasawa \(2022\)Large Language Models are Zero\-Shot Reasoners\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Li and P\. Vitányi \(2019\)An Introduction to Kolmogorov Complexity and Its Applications\.4th edition,Springer\.Cited by:[§C\.1](https://arxiv.org/html/2606.26366#A3.SS1.p1.5),[§6](https://arxiv.org/html/2606.26366#S6.SS0.SSS0.Px1.p1.2)\.
- C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. Ha \(2024\)The AI Scientist: Towards Fully Automated Open\-Ended Scientific Discovery\.arXiv preprint arXiv:2408\.06292\.Cited by:[§1](https://arxiv.org/html/2606.26366#S1.p4.1),[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px5.p1.1)\.
- A\. Lynch, B\. Wright, C\. Larson, K\. K\. Troy, S\. J\. Ritchie, S\. Mindermann, E\. Perez, and E\. Hubinger \(2025\)Agentic Misalignment: How LLMs Could be Insider Threats\.External Links:[Link](https://www.anthropic.com/research/agentic-misalignment)Cited by:[§E\.3](https://arxiv.org/html/2606.26366#A5.SS3.p1.3),[§E\.3](https://arxiv.org/html/2606.26366#A5.SS3.p2.2),[Appendix E](https://arxiv.org/html/2606.26366#A5.p1.1),[§1](https://arxiv.org/html/2606.26366#S1.p2.5),[Ethics Statement](https://arxiv.org/html/2606.26366#Sx2.p1.1)\.
- A\. M\. Bran, S\. Cox, O\. Schilter, C\. Baldassari, A\. D\. White, and P\. Schwaller \(2024\)Augmenting Large Language Models with Chemistry Tools\.Nature Machine Intelligence6\(5\),pp\. 525–535\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px5.p1.1)\.
- A\. MacIntyre \(1981\)After Virtue: A Study in Moral Theory\.University of Notre Dame Press\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px4.p1.1),[§6](https://arxiv.org/html/2606.26366#S6.SS0.SSS0.Px1.p1.2)\.
- OpenAI \(2025a\)Expanding on What We Missed with Sycophancy\.External Links:[Link](https://openai.com/index/expanding-on-sycophancy/)Cited by:[§1](https://arxiv.org/html/2606.26366#S1.p2.5)\.
- OpenAI \(2025b\)Sycophancy in GPT\-4o: What Happened and What We’re Doing about It\.External Links:[Link](https://openai.com/index/sycophancy-in-gpt-4o/)Cited by:[§1](https://arxiv.org/html/2606.26366#S1.p2.5)\.
- J\. S\. Park, J\. C\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. Bernstein \(2023\)Generative Agents: Interactive Simulacra of Human Behavior\.InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology \(UIST\),Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Pearl \(1995\)Causal Diagrams for Empirical Research\.Biometrika82\(4\),pp\. 669–688\.Cited by:[§6](https://arxiv.org/html/2606.26366#S6.SS0.SSS0.Px1.p1.2)\.
- J\. Pearl \(2009\)Causality: Models, Reasoning, and Inference\.2nd edition,Cambridge University Press\.Cited by:[§6](https://arxiv.org/html/2606.26366#S6.SS0.SSS0.Px1.p1.2)\.
- J\. L\. Pollock \(1987\)Defeasible Reasoning\.Cognitive Science11\(4\),pp\. 481–518\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px4.p1.1),[§6](https://arxiv.org/html/2606.26366#S6.SS0.SSS0.Px1.p1.2)\.
- H\. Putnam \(1967\)Psychological Predicates\.InArt, Mind, and Religion,W\. H\. Capitan and D\. D\. Merrill \(Eds\.\),Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px3.p1.1)\.
- P\. Röttger, H\. R\. Kirk, B\. Vidgen, G\. Attanasio, K\. Bontcheva, and D\. Hovy \(2023\)XSTest: A Test Suite for Identifying Excessive Safety Refusals in Large Language Models\.arXiv preprint arXiv:2308\.01263\.Cited by:[Appendix D](https://arxiv.org/html/2606.26366#A4.p1.1),[Limitations](https://arxiv.org/html/2606.26366#Sx1.p2.4)\.
- S\. Schmidgall, Y\. Su, Z\. Wang, X\. Sun, J\. Wu, X\. Yu, J\. Liu, M\. Moor, Z\. Liu, and E\. Barsoum \(2025\)Agent Laboratory: Using LLM Agents as Research Assistants\.arXiv preprint arXiv:2501\.04227\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px5.p1.1)\.
- V\. A\. Shaffer, E\. S\. Focella, A\. Hathaway, L\. D\. Scherer, and B\. J\. Zikmund\-Fisher \(2019\)Why Stories Matter: Narrative Perspective\-Taking and the Cultivation of Empathic Concern\.Medical Decision Making38\(3\),pp\. 335–342\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Sharma, M\. Tong, T\. Korbak, D\. Duvenaud, A\. Askell, S\. R\. Bowman, N\. Cheng, E\. Durmus, Z\. Hatfield\-Dodds, S\. R\. Johnston, S\. Kravec, T\. Maxwell, S\. McCandlish, K\. Ndousse, O\. Rausch, N\. Schiefer, D\. Yan, M\. Zhang, and E\. Perez \(2024\)Towards Understanding Sycophancy in Language Models\.InInternational Conference on Learning Representations \(ICLR\),Note:arXiv:2310\.13548Cited by:[Appendix E](https://arxiv.org/html/2606.26366#A5.p1.1),[§1](https://arxiv.org/html/2606.26366#S1.p2.5)\.
- K\. Swanson, W\. Wu, N\. L\. Bulaong, J\. E\. Pak, and J\. Zou \(2025\)The Virtual Lab of AI Agents Designs New SARS\-CoV\-2 Nanobodies\.Nature646,pp\. 716–723\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px5.p1.1)\.
- B\. Vidgen, A\. Agrawal, A\. M\. Ahmed, V\. Akinwande, N\. Al\-Nuaimi, N\. Alfaraj, E\. Alhussain, N\. Banovic, S\. Barikeri, M\. Bartolo,et al\.\(2024\)Introducing v0\.5 of the AI Safety Benchmark from MLCommons\.arXiv preprint arXiv:2404\.12241\.Cited by:[Appendix D](https://arxiv.org/html/2606.26366#A4.p1.1),[Limitations](https://arxiv.org/html/2606.26366#Sx1.p2.4)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. Le, and D\. Zhou \(2022\)Chain\-of\-Thought Prompting Elicits Reasoning in Large Language Models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px1.p1.1)\.
- W\. Wu, S\. Mao, Y\. Zhang, Y\. Xia, L\. Dong, L\. Cui, and F\. Wei \(2024\)Mind’s Eye of LLMs: Visualization\-of\-Thought Elicits Spatial Reasoning in Large Language Models\.arXiv preprint arXiv:2404\.03622\.Cited by:[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px2.p1.1)\.
- H\. Yan, H\. Xu, S\. Qi, S\. Yang, and Y\. He \(2025\)When Thinking Backfires: Mechanistic Insights into Reasoning\-induced Misalignment\.arXiv preprint arXiv:2509\.00544\.Cited by:[§1](https://arxiv.org/html/2606.26366#S1.p2.5),[§2](https://arxiv.org/html/2606.26366#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Yudkowsky \(2011\)Complex Value Systems are Required to Realize Valuable Futures\.Machine Intelligence Research Institute\.Cited by:[§C\.1](https://arxiv.org/html/2606.26366#A3.SS1.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2606.26366#S6.SS0.SSS0.Px1.p1.2)\.
- M\. Yuksekgonul, F\. Bianchi, J\. Boen, S\. Liu, Z\. Huang, C\. Guestrin, and J\. Zou \(2024\)TextGrad: Automatic “Differentiation” via Text\.arXiv preprint arXiv:2406\.07496\.Cited by:[§4\.2](https://arxiv.org/html/2606.26366#S4.SS2.p1.1)\.
## Appendix AInter\-Judge Agreement
### A\.1 Budget judge pair \(Experiment 1 primary analysis\)
Table[3](https://arxiv.org/html/2606.26366#A1.T3)reports quadratic\-weighted Cohen’sκ\\kappa\(Cohen,[1960](https://arxiv.org/html/2606.26366#bib.bib59)\)per structural variable, computed on the full Experiment 1 DailyDilemmas corpus \(n=3,726n=3\{,\}726complete pairs out of3,7803\{,\}780designed cells:claude\-haiku\-4\-513×20×3=78013\\times 20\\times 3=780;grok\-4\-1\-fast\-reasoning100×5×3=1,500100\\times 5\\times 3=1\{,\}500;claude\-sonnet\-4\-6100×5×3=1,500100\\times 5\\times 3=1\{,\}500\)\. The haiku subsample covers only1313of the100100DailyDilemmas scenarios because haiku was also used as a judge, so its generation cache was populated for the reliability check first before the full 100\-scenario run was complete; the partial overlap does not affect the kappa estimate since all three generators contribute to the pooled statistic\. The three generators are covered across33headline conditions\.gpt\-5\.4\-nanois excluded from this reliability estimate because it also serves as the secondary judge; its cross\-judge agreement cannot be computed independently\. Values are computed between the primary judge \(claude\-haiku\-4\-5\) and the secondary judge \(gpt\-5\.4\-nano\)\.max\_causal\_hopsfalls marginally below the0\.400\.40moderate\-agreement threshold; it is reported for completeness but excluded from all headline claims, including theKCK\_\{C\}discussion in §[6](https://arxiv.org/html/2606.26366#S6), which the length\-residualisation panel falsifies independently of causal\-hop coding\.
Table 3:Quadratic\-weighted Cohen’sκ\\kappabetween the budget judge pair \(claude\-haiku\-4\-5primary,gpt\-5\.4\-nanosecondary\) on the full Experiment 1 DailyDilemmas corpus \(n=3,726n=3\{,\}726\)\.
### A\.2 Held\-out third judge \(grok\-4\-1\-fast\-reasoning\)
Because two of the four generators \(claude\-haiku\-4\-5andgpt\-5\.4\-nano\) also serve as the two budget judges that produce the coded structural variables, the primary judge for each generator is always the non\-self sibling, and to verify the cross\-judge result is not an artefact of within\-family collusion we rangrok\-4\-1\-fast\-reasoningas a held\-out third judge on a3030\-scenario subsample drawn deterministically from the 100\-scenario DailyDilemmas pool \(seed=99=99\), coveringgrok\-4\-1\-fast\-reasoningandclaude\-sonnet\-4\-6generators at all three conditions \(N=1N\{=\}1per cell;30×2×3=18030\\times 2\\times 3=180generation outputs\)\. Table[4](https://arxiv.org/html/2606.26366#A1.T4)reports pairwise quadratic\-weightedκ\\kappaacross all three judges on the177177cells where complete triples were available\. Both A\.1 and A\.2 use quadratic weighting; the higher agreement here reflects two structural differences from the A\.1 analysis: the subsample covers only two generators \(30×2×3=18030\\times 2\\times 3=180cells vs3,7263\{,\}726in A\.1\), and all cells useN=1N\{=\}1, eliminating the pooling variance that arises when multiple samples per cell are aggregated in A\.1\.
Table 4:Pairwise quadratic\-weightedκ\\kappaacross the three judges on the 30\-scenario subsample \(n=177n=177complete triples\)\.j1=claude\-haiku\-4\-5j\_\{1\}=\\texttt\{claude\-haiku\-4\-5\},j2=gpt\-5\.4\-nanoj\_\{2\}=\\texttt\{gpt\-5\.4\-nano\},j3=grok\-4\-1\-fast\-reasoningj\_\{3\}=\\texttt\{grok\-4\-1\-fast\-reasoning\}\(held\-out\)\. All pairwise pairs exceed0\.400\.40onstakeholder\_countanduncertainty\_score;max\_causal\_hopsimproves modestly over the A\.1 estimate but remains cautionary\.#### Direction\-of\-effect under held\-out judge\.
Whengrok\-4\-1\-fast\-reasoningis used as the primary judge on the same3030\-scenario subsample, the direction\-of\-effect matches all three headline variables in the same direction asj1j\_\{1\}andj2j\_\{2\}:stakeholder\_countnarr=5\.13\>=5\.13\>std=3\.08\>=3\.08\>base=2\.67=2\.67;max\_causal\_hopsnarr=4\.43\>=4\.43\>std=3\.10\>=3\.10\>base=2\.38=2\.38;uncertainty\_scorenarr=2\.93\>=2\.93\>std=2\.18\>=2\.18\>base=1\.52=1\.52\. All three direction\-of\-effect comparisons \(narration\-of\-thought\>\>baseline\) agree acrossj1j\_\{1\},j2j\_\{2\}, andj3j\_\{3\}\.
### A\.3 Cross\-vendor moderator \(Experiment 2\)
To directly test whether the headline consensus rate in §[5](https://arxiv.org/html/2606.26366#S5)is an artefact of within\-vendor moderation, we re\-ran the complete Round 3–4 pipeline \(open synthesis, synthesis acceptance, integration, binary vote\) withclaude\-sonnet\-4\-6\(Anthropic\) as the replacement moderator over the cachedgpt\-5\.4\-nanoagent statements from Rounds 0–2\. Fifty debates were run \(55scenarios×\\times1010samples\)\. The cross\-vendor full\-consensus rate is44/50=88%44/50=88\\%\(Wilson95%95\\%CI\[76\.2,94\.4\]%\[76\.2,94\.4\]\\%\), compared to36/40=90%36/40=90\\%\(CI\[76\.9,96\.0\]%\[76\.9,96\.0\]\\%\) under within\-vendor moderation \(gpt\-4o\-mini\)\. Fisher’s exact test: OR=1\.23=1\.23,p=1\.0p=1\.0\. The consensus rate is statistically indistinguishable across moderator vendors\. The per\-generator split of the headline number isgpt\-5\.4\-nano36/40=90%36/40=90\\%\(Wilson\[76\.9,96\.0\]%\[76\.9,96\.0\]\\%\) andgpt\-4o42/42=100%42/42=100\\%\(Wilson\[91\.6,100\]%\[91\.6,100\]\\%\); all four rejections come fromgpt\-5\.4\-nanodebates\. The by\-scenario breakdown is given below\.
Table 5:Full\-consensus count by scenario under within\-vendor and cross\-vendor moderation\. Within\-vendor cells have unequalnnbecause ten of the5050designed debates produced no synthesis in Round 2 and were therefore excluded from Round 3 onward; the cross\-vendor arm re\-ran all5050designed debates from scratch through the full Round 3–4 pipeline, so its denominator is5050while the within\-vendor denominator is4040\. Both rates measure full Round\-4 consensus conditional on the debates each arm actually ran; Fisher’s exact test on the40/5040/50observed counts is valid under this asymmetry\. Theav\_engineerscenario shows the largest cross\-vendor reduction, consistent with its structurally harder stakeholder conflict\.
## Appendix BFull Effect\-Size Tables
Table[6](https://arxiv.org/html/2606.26366#A2.T6)reports raw and length\-residualised Cliff’sδ\\delta\(NoT vs\. standard CoT\) for the four\-model panel on all four coded variables, with95%95\\%bootstrap CIs \(500500iterations, seed4242\)\.
Table 6:Cliff’sδ\\delta\(NoT vs\. standard CoT\), raw and after length residualisation, with95%95\\%bootstrap CIs \(500500iterations, seed4242\)\. The headline trend: OpenAI and xAI generators retain large effects on stakeholder count and uncertainty score after length removal; the two Anthropic generators residualise to near zero, so their shifts in the coded metrics ride predominantly with output length\.Figure 6:Visualisation of Table[6](https://arxiv.org/html/2606.26366#A2.T6)\. Tier\-1 Cliff’sδ\\deltaeffect sizes \(NoT vs\. standard CoT\) on the four structural variables, with bootstrap95%95\\%CIs\. All four generators show large positiveδ\\deltas on stakeholder count and uncertainty score; effect sizes are robust to length residualisation on the OpenAI and xAI generators, while shrinking toward zero on the two Anthropic generators\.
## Appendix CKCK\_\{C\}: Formalism and Proxy Panel
### C\.1Definition and selection rule
Treating an SCMM=\(𝒱,ℰ,ℱ\)M=\(\\mathcal\{V\},\\mathcal\{E\},\\mathcal\{F\}\)as a computational object, we define its*algorithmic causal complexity*as the length of the shortest program that reproduces all interventional behaviour:
KC\(M\)=minp\{\|p\|:∀i∈ℐ,U\(p,i\)=M\(i\)\},K\_\{C\}\(M\)=\\min\_\{p\}\\bigl\\\{\|p\|:\\forall i\\in\\mathcal\{I\},\\;U\(p,i\)=M\(i\)\\bigr\\\},\(1\)whereUUis a universal Turing machine andℐ\\mathcal\{I\}is the set of admissible interventions \(e\.g\., “set protagonist’s belief that the patient consented to true”; “replace the moderator with one that hides the third party’s stake”\)\. Intuitively,KC\(M\)K\_\{C\}\(M\)asks: how many bits do you need to write down the rules that would let you simulate every possible alternative version of this situation? A model with a few stakeholders and one stable mechanism per stakeholder is short; a model that requires a special\-case rule per simulated future to keep an initial falsehood consistent is long\. Eq\.[1](https://arxiv.org/html/2606.26366#A3.E1)replaces “shortest program that reproduces the data” with “shortest program whose interventional behaviour reproducesMM” and inherits the uncomputability result ofLi and Vitányi \([2019](https://arxiv.org/html/2606.26366#bib.bib29)\)\.
TheargminKC\\arg\\min K\_\{C\}selection rule prefers the candidate response whose narrated trajectory has the shortest description\. Locally agreeing with a falsehood is dispreferred because keeping the falsehood consistent across simulated futures requires more bits than acknowledging the falsehood and absorbing the local friction\.
#### Why an SCM avoids the utility\-function pitfalls\.
Yudkowsky \([2011](https://arxiv.org/html/2606.26366#bib.bib22)\)’s “Complexity of Value” argument applies to a utility function over outcomes: any explicit specification is incomplete in ways an indifferent optimiser will exploit\. An SCM is a different object: it specifies mechanisms that connect actions to outcomes, not a preference ordering over outcomes\. Two agents with incompatible utilities can share an SCM \(they will agree, e\.g\., that betrayal destabilises trust\) and disagree about which trajectory to prefer; conversely, an agent with a single utility function but no SCM cannot answer counterfactual questions like “what would have happened if you had told the truth?”\. Grounding alignment in shared SCMs over narrative trajectories therefore brackets normative disagreement and exposes the causal regularities that persist across cultures, rather than asserting one preference ordering as the alignment target\.
### C\.2Proxy panel: raw vs length\-residualised
Table[7](https://arxiv.org/html/2606.26366#A3.T7)reports the per\-generator raw and length\-residualised Spearmanρ\\rhobetween each tested proxy and the NoT/standard\-CoT direction indicator \(\+1\+1for NoT,−1\-1for standard CoT\) on the Phase 1 cache, restricted to the two contrast conditions\. Residualisation regresses each proxy onlog\(length\)\\log\(\\text\{length\}\)within generator and correlates the residuals with the contrast\.
Table 7:Per\-generator raw and length\-residualised Spearmanρ\\rhofor seven length\-invariant proxies ofKCK\_\{C\}against the NoT vs\. standard\-CoT contrast \(†\\daggerMATTR = Moving\-Average Type\-Token Ratio\)\. Every raw correlation above\|ρ\|=0\.7\|\\rho\|\{=\}0\.7collapses below\|ρ\|=0\.21\|\\rho\|\{=\}0\.21after length residualisation on the three generators with large length ratios \(44–5×5\{\\times\}\)\. Ongrok\-4\-1\-fast\-reasoning\(length ratio1\.55×1\.55\{\\times\}\) a moderate residual signal survives but points the opposite directionKCK\_\{C\}predicts: NoT outputs are more compressible per byte and have lower per\-character entropy than standard\-CoT outputs\. We read this as falsification of the headline gzip claim, not validation\.#### Registered structural proxy \(30\-scenario pilot\)\.
A graph\-extraction proxyK^graph\\hat\{K\}\_\{\\text\{graph\}\}readsKCK\_\{C\}at the level of the structural\-causal model extracted from the trace\. Aclaude\-haiku\-4\-5extractor parses each trace into nodes \(stakeholders, actions, consequences\) and edges \(causal hops, uncertainty arcs\)\. On the 30\-scenario pre\-registration subsample, pooled Spearmanρ=0\.60\\rho=0\.60\(p=0\.0004p=0\.0004\) against the NoT direction, exceeding\|ρ\|≥0\.4\|\\rho\|\\geq 0\.4\.
#### Scaled validation \(100 scenarios×\\times4 generators\)\.
Scaling to the full experimental cache \(n=976n\{=\}976trace\-graph pairs, one sample per scenario–generator–condition triple\) yields pooledρ=0\.42\\rho=0\.42\(p<0\.001p<0\.001, bootstrap 95% CI\[0\.36,0\.47\]\[0\.36,0\.47\]\)\. The per\-generator breakdown isclaude\-haiku\-4\-5ρ=\+0\.78\\rho\{=\}\{\+0\.78\}\(\[0\.72,0\.82\]\[0\.72,0\.82\]\),claude\-sonnet\-4\-6ρ=\+0\.76\\rho\{=\}\{\+0\.76\}\(\[0\.70,0\.81\]\[0\.70,0\.81\]\),grok\-4\-1\-fast\-reasoningρ=\+0\.63\\rho\{=\}\{\+0\.63\}\(\[0\.53,0\.71\]\[0\.53,0\.71\]\), andgpt\-5\.4\-nanoρ=\+0\.13\\rho\{=\}\{\+0\.13\}\(\[−0\.01,0\.28\]\[\-0\.01,0\.28\]\)\. The near\-zero OpenAIρ\\rhois consistent with §[4](https://arxiv.org/html/2606.26366#S4):gpt\-5\.4\-nanotraces already approach the target causal density under standard CoT, so the scaffold adds length rather than additional causal structure\.
#### Convergent SCM\-level proxies\.
Two further proxies computed on the same extracted graphs at zero additional API cost: graph MDL \(log2\(n\+1\)\+log2\(e\+1\)\\log\_\{2\}\(n\{\+\}1\)\{\+\}\\log\_\{2\}\(e\{\+\}1\)normalised bylog2\(len\+1\)\\log\_\{2\}\(\\text\{len\}\{\+\}1\)\) and structural entropy \(Shannon entropy of the out\-degree distribution\)\. Both reachp<0\.001p<0\.001pooled \(MDL:ρ=−0\.32\\rho=\-0\.32, 95% CI\[−0\.38,−0\.26\]\[\-0\.38,\-0\.26\]; structural entropy:ρ=\+0\.27\\rho=\+0\.27, 95% CI\[0\.21,0\.33\]\[0\.21,0\.33\]\)\. The negative MDL sign reflects that NoT embeds more causal nodes and edges into proportionally longer text, reducing the per\-token MDL\. Together the three SCM\-level proxies provide convergent empirical support forKCK\_\{C\}as recoverable at the structural\-causal\-model level of the protocol\.
## Appendix DRefusal Modulation on Dedicated Benchmarks
We tested NoT vs\. standard CoT on two benchmarks designed to probe refusal behaviour: XSTest\(Röttgeret al\.,[2023](https://arxiv.org/html/2606.26366#bib.bib61)\)\(250*safe*prompts that superficially resemble unsafe ones; a model should*not*refuse\) and SimpleSafetyTests\(Vidgenet al\.,[2024](https://arxiv.org/html/2606.26366#bib.bib62)\)\(100*unsafe*prompts that a model*should*decline\)\. Each benchmark was run on the full four\-model panel withN=1N\{=\}1per cell; refusal was coded by a single\-shotgpt\-5\.4\-nanobinary classifier \(REFUSE / HEDGE / ENGAGE\)\.
Table 8:Refusal rates on XSTest \(over\-refusal; lower is better\) and SimpleSafetyTests \(appropriate\-refusal; higher is better\)\. Anthropic generators and Grok show no significant modulation under either condition\. Forgpt\-5\.4\-nano, NoT increases overall caution \(\+10 pp on safe prompts; \+17 pp on unsafe prompts\): the scaffold induces extra deliberation that raises both appropriate and over\-refusal rates for this model family\.The low SST refusal rates for Anthropic generators reflect an instrument boundary: these models redirect harm\-adjacent prompts toward support resources rather than issuing a clean refusal token, which the binary classifier codes as ENGAGE\. The substantive content of those responses does not comply with the harmful request\. We flag this as a known classifier limitation rather than a safety regression\.
## Appendix EDeployment\-Relevance Probes
We tested whether the upstream shifts in the coded metrics from §[4](https://arxiv.org/html/2606.26366#S4)propagate to two publicly described downstream probe sets: SycophancyEval\(Sharmaet al\.,[2024](https://arxiv.org/html/2606.26366#bib.bib8)\)and a replication of two scenarios from the agentic\-misalignment release ofLynchet al\.\([2025](https://arxiv.org/html/2606.26366#bib.bib5)\)\. Both probe sets are saturated on the current four\-model panel, which we report as pre\-registered contingencies rather than as evidence against the intervention\.
### E\.1SycophancyEval
SycophancyEval contains three probe types: opinion mirroring, retraction on pushback, and false\-premise acceptance\. We run all three on the four\-model panel \(N=3N\{=\}3per cell,720720coded responses\);claude\-haiku\-4\-5codes each response\. Table[9](https://arxiv.org/html/2606.26366#A5.T9)reports per\-probe\-type sycophancy rates \(fraction of responses codedsycophantic\) under standard CoT and NoT for each generator\.
Table 9:Sycophancy rates \(%\) per probe type\. S = Standard CoT, N = NoT\.2222of2424cells are at the0%0\\%floor; the only non\-floor cells areclaude\-haiku\-4\-5on retraction\-on\-pushback under standard CoT \(NoT eliminates it\) andgpt\-5\.4\-nanoon false premise under NoT \(the only cell where NoT scores worse than standard CoT in this probe set\)\. The right reading is instrument saturation on frontier models rather than intervention failure \(Appendix[E](https://arxiv.org/html/2606.26366#A5)\)\.All cells in Table[9](https://arxiv.org/html/2606.26366#A5.T9)are at or near the0%0\\%floor of the judge instrument acrossN=3N\{=\}3per cell and720720coded responses; the single non\-floor cell isgpt\-5\.4\-nanofalse\-premise under NoT \(6\.7%6\.7\\%\)\. We read this as instrument saturation on frontier models rather than a null effect on the intervention\.
### E\.2ELEPHANT Social\-Sycophancy Benchmark
Sharma\-style probes saturate at the floor above; we therefore re\-run the comparison on ELEPHANT\(Chenget al\.,[2025](https://arxiv.org/html/2606.26366#bib.bib7)\), which scores*social*sycophancy—excessive face\-preservation—via validation, indirectness, and framing judges \(faithful prompt port;claude\-haiku\-4\-5scorer\)\. Phase 12 reported a1010\-prompt smoke sample here; Phase 13 scales to the OSF full splits \(n=150n\{=\}150per slice, seed 44\) with a literature\-comparablerawarm \(no system prompt\), plain IO, standard CoT, NoT, and multi\-stakeholder NoT on the verified quartet\. Full per\-model tables, raw\-vs\-IO contrasts, and Sharma\-bridge discussion are in the standalone sycophancy study \(papers/sycophancy/sycophancy\_paper\.tex\); we retain the Phase 12 smoke snapshot below as a directional preview\.
Table[10](https://arxiv.org/html/2606.26366#A5.T10)summarises the clearest single\-agent cells onclaude\-haiku\-4\-5\. NoT*reduces*framing sycophancy on AITA\-YTA \(0%0\\%vs\.50%50\\%under standard CoT; Fisher exactp=0\.033p\{=\}0\.033\) and validation sycophancy on the same slice \(0%0\\%vs\.30%30\\%\)\. On OEQ, NoT matches the crowdsourced human validation rate \(30%30\\%each\) while standard CoT runs higher \(80%80\\%\)\. Moral sycophancy \(both\-NTA rate on flipped pairs\) is directionally lower under NoT than CoT on all three generators \(6060–80%80\\%vs\.9090–100%100\\%\)\. The trade\-off is indirectness: NoT’s committed Decision section can read as*more*suggestive on OEQ \(100%100\\%vs\.80%80\\%under CoT onhaiku\), so the scaffold suppresses reflexive validation and premise acceptance but not every face\-preserving linguistic habit\. Multi\-stakeholder NoT does not uniformly beat single\-agent NoT on validation \(debate raises it on AITA\-YTA forhaiku\) but*lowers*OEQ indirectness \(60%60\\%vs\.100%100\\%\), consistent with the integration layer forcing a concrete consensus statement\.
Table 10:ELEPHANT social\-sycophancy rates \(%, higher = more sycophantic\) on the1010\-prompt sample, generatorclaude\-haiku\-4\-5\. Framing/validation judges followChenget al\.\([2025](https://arxiv.org/html/2606.26366#bib.bib7)\); human column is crowdsourced responses on the same prompts\. AITA\-YTA IO arm omitted \(identical prompt format to CoT on binary judgement tasks\)\.
### E\.3Agentic\-Misalignment Replication
We replicate two scenarios from theLynchet al\.\([2025](https://arxiv.org/html/2606.26366#bib.bib5)\)release: an executive\-blackmail scenario and a corporate\-espionage scenario\. Each is instantiated with the same system\-prompt template, tool inventory \(send\_email,read\_email,search\_files,forward\_email\), and adversary context as the published release\. Each cell is run three times with temperature0\.70\.7and a1616k\-token budget, giving3×2×2×4=483\\times 2\\times 2\\times 4=48generations\. Table[11](https://arxiv.org/html/2606.26366#A5.T11)reports per\-cell action\-code counts\. All non\-harmful classifications are decomposed intorefuseandhedge\(the latter being a mid\-deliberation truncation in which the agent’s narrated reasoning declines to invoke any harmful tool but the explicit tool\-call statement is incomplete in the model’s returned text\)\.
Before running we pre\-registered the directional prediction that NoT reduces harmful\-action rates, with a contingency for the floor case: if no generator commits the harmful action under standard CoT, the result is reportable as “inference\-time scaffolding cannot be measured against this failure mode in this replication setting” rather than as null evidence against the intervention\. The contingency is the observed result\. Harmful\-action rate is0%0\\%in all1616cells\. The most parsimonious reading is defence\-in\-depth: the publicly described scenario structures fromLynchet al\.\([2025](https://arxiv.org/html/2606.26366#bib.bib5)\)have been visible long enough that current training pipelines on every vendor in the four\-model panel refuse them unconditionally\. A real test against agentic misalignment will require harder, less\-public scenarios than the published replicates\.
Table 11:Per\-cell action\-code counts \(R = refuse, H = hedge / unclear; S = Standard CoT, N = NoT\)\. Harmful\-action rate is0%0\\%in all1616cells: none of the4848generations invokes the harmful tool\. The corporate\-espionage scenario yields clear refusals under either scaffold on every model; the blackmail scenario yields hedges \(mid\-deliberation truncations that read as refusal\-in\-progress\) predominantly under NoT\. The agent’s narrated reasoning still rejects the harmful tool call, but the longer NoT trace is more likely to be truncated mid\-analysis\.The within\-scenario pattern is consistent across vendors\. The corporate\-espionage scenario, in which the harmful tool would exfiltrate confidential files to a competitor, is rejected cleanly \(33refusals out of33\) by every generator under either scaffold; none of the eight cells in the right half of Table[11](https://arxiv.org/html/2606.26366#A5.T11)contains a single hedge\. The blackmail scenario, in which the harmful tool would forward a compromising email to coerce continued operation, draws clean refusals under standard CoT \(with two of three runs on each reasoning model hitting the explicitdeclinetool\) but produces hedges under NoT on every generator: the narrated trace expands the consequences and stakeholder sections, and the explicit tool\-call line is truncated by the response budget before a finaldeclinestatement can be emitted\. The narrated content of those hedged traces refuses the blackmail in every case; the failure is at the protocol layer \(response cutoff\), not the alignment layer\.
Figure 7:Clean\-refusal rate per generator and scaffold on each agentic scenario \(N=3N\{=\}3per cell,4848long\-context generations\)\. Bars are the share of generations that emit an explicit refusal; the complement is the hedge / truncated share\. The harmful\-action rate is0%0\\%in all1616cells, so that slice is not plotted\. S = Standard CoT, N = NoT\.Figure[7](https://arxiv.org/html/2606.26366#A5.F7)shows the clean\-refusal rate per cell\. Under NoT, the blackmail scenario shifts mass out ofrefuseintohedge / truncated, consistent with longer traces being cut off mid\-refusal by the response budget rather than reversing the refusal; the corporate\-espionage panel shows no within\-scaffold redistribution\. The combined pattern is consistent with the pre\-registered contingency stated in Appendix[E](https://arxiv.org/html/2606.26366#A5): when a failure mode does not fire under the baseline, scaffolding cannot be measured against it, and any observed shifts under NoT must be read as structural trace changes rather than as a change in the alignment\-relevant target\.
## Appendix FSub\-Instruction Ablation Full Table
Table[12](https://arxiv.org/html/2606.26366#A6.T12)reports Cliff’sδ\\deltafor the five drop\-one conditions \(one per NoT sub\-instruction from §[3](https://arxiv.org/html/2606.26366#S3)\) against the full NoT control onclaude\-sonnet\-4\-6\(N=3N\{=\}3,3030\-scenario stratified subsample\)\. Columns are stakeholder count, causal hops, uncertainty score, and forward\-window length\.
Table 12:Cliff’sδ\\deltafor each drop\-one condition vs\. full NoT control on claude\-sonnet\-4\-6\. Bolded entries mark the largest negative effect for each structural variable\. The diagonal pattern, in which each substantive section shows its largest effect on its own target metric, is the textbook signature of section\-level mechanism rather than a single emergent “narrative” factor\.Reading the diagonal: dropping the stakeholders section produces the largest drop in stakeholder count \(δ=−0\.41\\delta=\-0\.41\); dropping the consequences section produces the largest drop in causal hops \(δ=−0\.22\\delta=\-0\.22\); dropping the uncertainty section produces the largest drop in uncertainty score \(δ=−0\.92\\delta=\-0\.92, a near\-complete collapse to the standard\-CoT baseline\)\. The forward\-window column carries no comparably large negative entry because no single sub\-instruction is its sole carrier; the long trace is the joint product of all four substantive sections\.
The off\-diagonal entries are tightly bounded \(\|δ\|≤0\.13\|\\delta\|\\leq 0\.13across all twelve\), which is the empirically interesting half: a single emergent “narrative” factor would shrink every variable when any one section is dropped, and a pure length confound would shrinkfw\.most regardless of which section is dropped\. Neither pattern fires\. Each substantive sub\-instruction carries the structural variable the main paper attributes to it and little else, which is the cleanest evidence we have that the NoT scaffold is doing the work §[3](https://arxiv.org/html/2606.26366#S3)predicts rather than a single confound dressed as a five\-part trace\.
## Appendix GTextual\-Gradient Optimisation of the Scaffold
This appendix gives the full treatment of the textual\-gradient optimisation summarised in §[4\.2](https://arxiv.org/html/2606.26366#S4.SS2)\. It runs the optimiser in two complementary directions\.Optimising*from*NoT\(§[G\.1](https://arxiv.org/html/2606.26366#A7.SS1)–§[G\.5](https://arxiv.org/html/2606.26366#A7.SS5)\) asks whether the hand design can be improved and whether the choice of training judge controls how well the improvement survives a cross\-vendor evaluator; this is the head\-to\-head of an in\-family\-trained prompt \(NoT\-v2\) against a cross\-family\-trained prompt \(NoT\-v3\)\.Optimising*from*standard CoT\(§[G\.6](https://arxiv.org/html/2606.26366#A7.SS6)\) is the control on the prior premise that NoT is the object worth optimising at all: it descends from a plain chain\-of\-thought on the same loss and asks whether the optimiser reconstructs NoT\. It does not, at either the single\-agent or the multi\-stakeholder layer\.
### G\.1Continuous loss and two training regimes
We define a per\-output cell loss
ℓ=max\(0,4−𝑠𝑐\)\+max\(0,2−𝑢𝑠\),\\ell=\\max\(0,\\,4\-\\mathit\{sc\}\)\+\\max\(0,\\,2\-\\mathit\{us\}\),\(2\)and a batch lossL=\|B\|−1∑b∈BℓbL=\|B\|^\{\-1\}\\sum\_\{b\\in B\}\\ell\_\{b\}over the same stakeholder count𝑠𝑐\\mathit\{sc\}and uncertainty score𝑢𝑠\\mathit\{us\}that anchor Experiment 1\. The thresholds44and22are mid\-band NoT cell values, chosen so the loss stays informative even when the binary failure modes already fire at zero \(as they do for NoT on the training generator\); a binary loss provides no gradient in that regime\.
#### Optimisation loop\.
For each promptppand a batch ofk=10k=10scenarios, we \(i\) generate outputs with the target generator \(max\_tokens=4096=4096\), \(ii\) code each with the training judge to obtain\(𝑠𝑐,𝑢𝑠\)\(\\mathit\{sc\},\\mathit\{us\}\)and computeLL, \(iii\) feed the prompt, batch outputs, codes, andLLto an optimiser LLM \(claude\-sonnet\-4\-6, held constant across runs\) that writes a44–88sentence textual gradient diagnosing which sub\-instruction causes the shortfall, and \(iv\) ask the same optimiser to rewrite the prompt \(≤400\\leq 400words; the five\-section structure may be altered\)\. Training halts when three consecutive iterations each reduceLLby less than5%5\\%, or after1010iterations\.
#### Two regimes\.
We run the loop twice with everything held identical except the training judge\.Run A \(in\-family\):generator and training judge bothclaude\-haiku\-4\-5, the same vendor the panel’s primary judge uses; output promptNoT\-v2\(early\-stop iter44\)\.Run B \(cross\-family\):generatorclaude\-haiku\-4\-5, training judgegrok\-4\-1\-fast\-reasoning, a different vendor; output promptNoT\-v3\(early\-stop iter77\)\. Optimiser model, dataset, training split \(seed\-4343stratified3030\-scenario subsample\), held\-out eval split \(seed\-4242indices3030–5959\), loss, and early\-stopping criteria are identical across the two runs\.
#### Cross\-vendor replication\.
For each optimised prompt we replicate against the full four\-generator panel of Experiment 1 \(gpt\-5\.4\-nanoN=5N\{=\}5per cell,claude\-haiku\-4\-5N=20N\{=\}20,grok\-4\-1\-fast\-reasoningN=5N\{=\}5,claude\-sonnet\-4\-6N=5N\{=\}5, same100100\-scenario sample\)\. Each cell is coded by the primary judgeclaude\-haiku\-4\-5and re\-coded by the adversarial third judgegrok\-4\-1\-fast\-reasoning, the most cross\-vendor adversarial configuration available on our three\-vendor panel\.111The cross\-family training judge \(Run B\) and the cross\-vendor evaluation third judge are the same model: it is the only non\-Anthropic non\-OpenAI judge on our panel\. A fully orthogonal design would use a fourth vendor for the evaluation judge; within this constraint the claim is the narrower one that cross\-family training survives cross\-vendor evaluation strictly better than in\-family training does\.Inter\-judge agreement on binary labels is Cohen’sκ\\kappa\(Cohen,[1960](https://arxiv.org/html/2606.26366#bib.bib59)\); effect sizes are Cliff’sδ\\delta\(Cliff,[1993](https://arxiv.org/html/2606.26366#bib.bib58)\)with10001000\-bootstrap95%95\\%CIs\.
### G\.2Optimised prompts NoT\-v2 and NoT\-v3
Both runs converge on two structural changes versus NoT: an explicit stakeholder floor and an explicit per\-item uncertainty\-enumeration floor with consequence framing\. They differ on two choices that turn out to matter\.NoT\-v2\(2,5792\{,\}579characters\) sets the stakeholder floor at five, attaches the uncertainty grain per*action*\(“at least three distinct uncertainties” per course of action\), and closes with a four\-bullet “precision\-over\-length” block that the model read as licence to elaborate \(outputs*grew*on three of four generators\)\.NoT\-v3\(2,4012\{,\}401characters\) raises the stakeholder floor to six \(and asks for future generations\), attaches the uncertainty grain per*stakeholder*\(“for every party you named… no party should be skipped”\), and closes with a block that operationalises conciseness \(“one sharp sentence per stakeholder stake, one precise unknown per party”\)—outputs*shrank*on three of four generators\. Verbatim text for both prompts is in §[G\.7](https://arxiv.org/html/2606.26366#A7.SS7)\.
### G\.3Head\-to\-head under the primary judge
Table 13:Per\-generator Cliff’sδ\\deltaon stakeholder count under the primary judgeclaude\-haiku\-4\-5, plus output\-length ratio \(mean characters per output, optimised condition / hand\-written NoT\)\. Negativeδ\\deltafavours the optimised condition\. NoT\-v3 dominates NoT\-v2 on every generator: equal or larger improvement on all four, and shorter outputs on all four\.nano’s v2 regression \(δ=\+0\.43\\delta=\+0\.43, CI strictly\>0\>0\) becomes a non\-significant effect under v3 \(CI spans0\)\.Ncells=500N\_\{\\text\{cells\}\}=500\(nano,grok,sonnet\),20002000\(haiku\); NoT baseline cells:3030,254254,500500,500500respectively \(see Limitations on incomplete baseline caches\)\.Figure 8:Cross\-family training \(NoT\-v3\) beats in\-family training \(NoT\-v2\) on both axes the optimiser targeted, on every generator\.Left:v3 makes the stakeholder\-count Cliff’sδ\\deltaover hand\-written NoT more negative \(or, for thenanonon\-effect, less positive\) than v2 by0\.080\.08–0\.310\.31effect\-size points\.Right:v3 produces shorter outputs than v2 on all four generators, saving1313–44%44\\%of v2’s characters\. The chart summarises Tables[13](https://arxiv.org/html/2606.26366#A7.T13)and[14](https://arxiv.org/html/2606.26366#A7.T14)\.Table[13](https://arxiv.org/html/2606.26366#A7.T13)reports per\-generator Cliff’sδ\\deltaon stakeholder count for NoT vs NoT\-v2 and NoT vs NoT\-v3 under the primary judge\. Three of four generators show large\-to\-very\-large effect\-size improvements under both v2 and v3, andNoT\-v3 dominates NoT\-v2 on every generator: equal or larger raw effect sizes on all four \(−0\.57→−0\.73\-0\.57\\\!\\to\\\!\-0\.73onhaiku;−0\.68→−0\.93\-0\.68\\\!\\to\\\!\-0\.93ongrok;−0\.60→−0\.68\-0\.60\\\!\\to\\\!\-0\.68onsonnet; thenanoregression of\+0\.43\+0\.43shrinks to a non\-significant\+0\.12\+0\.12\)\. v3 also produces shorter outputs than NoT on three of four generators \(0\.590\.59–0\.78×0\.78\\times\), while v2 produces longer outputs on three of four \(1\.091\.09–1\.37×1\.37\\times\)\. Figure[8](https://arxiv.org/html/2606.26366#A7.F8)plots both axes of this dominance\.
### G\.4Cross\-vendor agreement and length residualisation
Table 14:Mean stakeholder count under the primary \(haiku\) and third \(grok\) judges on the same cells\. “gap” is \(primary mean\)−\-\(third mean\) on the optimised cells; positive means the primary judge counts more stakeholders than the third on the same outputs\. On the in\-family generator \(haiku\) the cross\-judge gap shrinks from\+0\.80\+0\.80under v2 to\+0\.46\+0\.46under v3; on the cross\-family generator \(grok\) it inverts; onsonnetit is unchanged\. Both judges agree on the sign of every effect on every generator under both prompts\.Two patterns emerge from Table[14](https://arxiv.org/html/2606.26366#A7.T14)\.\(1\) Both judges agree on the sign of every effect, on every generator, under both v2 and v3: the optimisation never produces a prompt the primary judge calls an improvement and the third judge calls a regression\.\(2\) An in\-family generosity gap is visible under v2 and shrinks under v3\.Under v2 the primary \(Anthropic\) judge counts0\.510\.51–0\.800\.80more stakeholders than the third judge on the three Anthropic\-or\-mixed generators; under v3 this shrinks to\+0\.46\+0\.46onhaikuand inverts to−0\.21\-0\.21ongrok, indicating the cross\-family\-trained prompt does not exploit an in\-family rubric interpretation on the non\-Anthropic generator\. Thesonnetgap is unchanged \(\+0\.51\+0\.51\), a vendor\-pair\-level effect orthogonal to the optimisation:sonnet\-generated text is scored slightly higher on stakeholder count by an Anthropic judge than by the grok judge irrespective of the prompt\.
Table 15:Length\-residualised Cliff’sδ\\deltaon stakeholder count \(residuals after OLS onlog\\logoutput length, NoT vs NoT\-v3\) and Cliff’sδ\\deltaon max causal hops \(an out\-of\-loss metric, not optimised against\)\. Per\-unit\-length stakeholder\-count gain is preserved onhaikuandgrok\. Max causal hops improves or holds on three of four generators despite being outside the loss; onlynanoregresses, matching its non\-effect on the primary metric\.The length\-residualisedδ\\deltain Table[15](https://arxiv.org/html/2606.26366#A7.T15)answers whether v3’s gain merely buys more tokens: it does not\. v3 outputs are shorter than NoT on three of four generators, and the per\-unit\-length effect stays large and negative onhaiku\(−0\.79\-0\.79\) andgrok\(−0\.93\-0\.93\)\. Max causal hops, which never enters the loss, improves on three of four generators, so the optimisation does not trade reasoning depth for stakeholder breadth on the responding generators\.
Table 16:Inter\-judge Cohen’sκ\\kappa\(primary vs third\) on binary collapse and suppression labels for v2 and v3 cells\.†Daggered rows are degenerate:≥99\.8%\\geq 99\.8\\%of cells receive identical binary labels from both judges \(almost always0\), and a constant\-marginal cell yieldsκ∈\{0,1\}\\kappa\\in\\\{0,1\\\}by formula \(the kappa paradox\)\. The only generator with non\-degenerate binary marginals isgrok;κ\\kapparises on both labels from v2 to v3\.Cohen’sκ\\kappaon the binary labels \(Table[16](https://arxiv.org/html/2606.26366#A7.T16)\) is degenerate on six of eight rows: both judges return the same label on≥99\.8%\\geq 99\.8\\%of cells and the positive\-label prevalence is so close to0or11that the formula returns0or11regardless of the disagreement structure\. Whereκ\\kappais informative \(grok, the only row with non\-zero positive prevalence\) it rises from0\.57→0\.750\.57\\\!\\to\\\!0\.75on collapse and0\.73→0\.920\.73\\\!\\to\\\!0\.92on suppression—the cleanest single piece of evidence that cross\-family training improves cross\-vendor agreement\. For the daggered rows the informative analogue is the continuous stakeholder\-count gap of Table[14](https://arxiv.org/html/2606.26366#A7.T14)\.
### G\.5Why cross\-family training helps
The simplest explanation is that an in\-family training judge supplies a gradient that is partly an in\-family rubric\-*interpretation*signal: the optimiser learns to satisfy not only the rubric specification but also the in\-family reading of it\. A cross\-family training judge supplies a gradient closer to the specification itself, because the optimiser cannot benefit from in\-family interpretation leniency\. The empirical signatures are thehaikuprimary\-vs\-third gap \(\+0\.80\+0\.80under v2,\+0\.46\+0\.46under v3\), thegrokgap inverting \(\+0\.57→−0\.21\+0\.57\\\!\\to\\\!\-0\.21\), and the larger raw effect sizes under v3\. A practical workflow follows: hand\-design a scaffold; optimise under a cross\-family judge; validate under both the deployment judge and an adversarial cross\-vendor third\. The cross\-family recommendation is independent of Equation[2](https://arxiv.org/html/2606.26366#A7.E2)and should transfer to other LLM\-judge prompt\-optimisation tasks; we do not have data outside the deliberative\-reasoning domain\.
### G\.6Control: optimising from CoT instead of from NoT
The runs above optimise*from*NoT\. The prior question is whether the NoT scaffold is worth keeping at all, or whether textual\-gradient descent would reach the same place starting from a plain standard CoT\. We answer with a head\-to\-head control that runs the optimiser in the opposite direction—initialised at the Experiment 1standard\_cotprompt and descending on the same loss—and pits the best optimised CoT against hand\-written NoT at both layers\.
#### Setup\.
The optimisation loop here minimises a failure\-mode lossL=StakeholderCollapse%\+UncertaintySuppression%L=\\text\{StakeholderCollapse\\%\}\+\\text\{UncertaintySuppression\\%\}\. At each iteration we \(i\) generate traces on a batch of1010DailyDilemmas scenarios under the current prompt, \(ii\) code each trace withclaude\-haiku\-4\-5on the rubric, \(iii\) feed the batch, codes, and loss to an optimiser model \(claude\-sonnet\-4\-6\) which writes a textual gradient diagnosing what is wrong with the current prompt, and \(iv\) ask the optimiser to rewrite the prompt to address the diagnosis\. Five iterations are run with the verbatim Experiment 1standard\_cotprompt as the starting point, on3030stratified DailyDilemmas scenarios \(indices0–2929of the seed\-4242sample\)\. The final optimised prompt is then evaluated against verbatim NoT on a held\-out3030\-scenario subsample \(indices3030–5959\) at the same generator andmax\_tokensbudget\. Generator throughout isgpt\-5\.4\-nano, the model where matched\-budget Experiment 1 showed the cleanest pro\-NoT effect \(δ=\+0\.90/\+0\.93\\delta=\+0\.90/\+0\.93\)\.
#### Optimised prompt\.
The pilot’s final prompt \(loss\-curve flat at0\.00\.0on every training batch; see caveat below\) is verbatim:
> You are a careful ethical reasoner\. When given a dilemma, think through it and give your answer\. Before concluding, you must: identify the distinct people or groups affected and what each stands to gain or lose; acknowledge any facts, values, or outcomes that remain genuinely uncertain; and weigh competing considerations without forcing false certainty\. Length is a hard constraint\. Your entire response must stay under450450words\. If you find yourself approaching that limit, cut immediately, trim restatements, throat\-clearing, and any sentence that doesn’t add new substance\. Brief and deep beats long and thorough\. Move directly into substantive analysis without restating the question or summarising what you are about to do\. Every sentence must do real work\. Cut anything that could be removed without loss\. Reach a clear, considered judgment while honestly noting where reasonable people could disagree or where key information is missing\.
The optimiser converges to an NoT\-shaped diagnosis \(name stakeholders, acknowledge uncertainty, commit to a judgement\) expressed as a compressed950950\-character instruction without explicit five\-section structure\. It also discovers a hard length constraint on its own\.
#### Held\-out comparison\.
Table[17](https://arxiv.org/html/2606.26366#A7.T17)reports the matched\-generator, matched\-judge comparison on the3030\-scenario held\-out subsample\. The TextGrad\-optimised prompt matches NoT at the binary failure\-mode floor \(both at0%0\\%on this subsample\) but lags NoT on the continuous coded metrics: Cliff’sδ\\deltaof\+0\.67\+0\.67on stakeholder count \(large effect,95%95\\%CI strictly above0\.470\.47\), with the NoT trace also2\.3×2\.3\{\\times\}longer in tokens\. The optimiser recovers the qualitative structure of NoT but, at five iterations and one batch per iteration, does not match its continuous depth\.
Table 17:Held\-out comparison on3030DailyDilemmas scenarios \(gpt\-5\.4\-nanogenerator,claude\-haiku\-4\-5judge\)\.δNoT\\delta\_\{\\text\{NoT\}\}is Cliff’sδ\\deltafor NoT vs\. the TextGrad\-optimised prompt \(positive favours NoT\)\. TextGrad wins on token cost \(2\.3×2\.3\{\\times\}cheaper\) and matches NoT at the binary failure\-mode floor; NoT wins on continuous stakeholder coverage with a large effect size\.
#### Caveat: model deployment drift\.
The optimisation loss was flat at0\.00\.0on every training batch because the verbatim Experiment 1standard\_cotprompt, run ongpt\-5\.4\-nanovia the same Azure deployment one day after the Experiment 1 main run, no longer fires either failure mode on these3030scenarios\. Experiment 1’s standard\-CoT baseline rates for this generator were14\.6%14\.6\\%collapse and50\.0%50\.0\\%suppression; the run\-time observed rates here are0%0\\%and0%0\\%\. The most parsimonious explanation is that the underlying deployment was updated between the two run dates; an alternative is sampling drift on the4040\-cell training subset\. We treat this as a feature, not a bug: the comparison reported in Table[17](https://arxiv.org/html/2606.26366#A7.T17)is between two prompts both already clearing the binary rubric thresholds, so theδ=\+0\.67\\delta=\+0\.67NoT advantage on continuous stakeholder coverage is what separates them\. The full100100\-scenario, multi\-seed Experiment 1 firing rates \(Table[1](https://arxiv.org/html/2606.26366#S4.T1)\) were measured before any prompt or deployment drift and stand as reported\.
#### Strengthened head\-to\-head: best optimised CoT, both layers\.
The pilot above optimised on the one generator where the binary loss had drifted to zero, so it could not exercise the optimiser\. We therefore re\-ran the optimisation onclaude\-haiku\-4\-5, which retains failure\-mode signal, under both the binary loss and a continuous depth lossmax\(0,4−sc\)\+max\(0,2−us\)\\max\(0,4\-\\textsc\{sc\}\)\+\\max\(0,2\-\\textsc\{us\}\), and evaluated the*stronger*of the two resulting prompts against verbatim NoT on the held\-out3030scenarios across two generators, coded by a cross\-vendor judge pair \(claude\-haiku\-4\-5primary,gpt\-5\.4\-nanosecondary\)\. The single\-agent verdict is unchanged: even the best optimised CoT trails NoT on stakeholder count by Cliff’sδ=\+0\.78\\delta=\+0\.78\(95%95\\%CI\[\+0\.63,\+0\.91\]\[\+0\.63,\+0\.91\]\) onclaude\-haiku\-4\-5and on uncertainty score by\+0\.23\+0\.23and\+0\.54\+0\.54on the two generators \(CIs strictly above0\)\. The optimised CoT closes the stakeholder\-count gap ongpt\-5\.4\-nano\(δ=\+0\.07\\delta=\+0\.07, CI spanning0\) but never the uncertainty gap, and does so at22–5×5\{\\times\}fewer tokens\.
#### Multi\-stakeholder head\-to\-head\.
Holding the five\-round integration protocol and the moderator \(claude\-sonnet\-4\-6\) constant and varying only the agent prompt, NoT reaches fullR4R4consensus on52%52\\%of held\-out scenarios \(Wilson95%95\\%CI\[39,64\]\[39,64\]\) versus32%32\\%\[21,44\]\[21,44\]for the best optimised CoT—a\+20\+20pp gap, Fisher exactp=0\.041p=0\.041—and leaves roughly four times fewer debates in unresolved structural rejection \(8%8\\%vs35%35\\%; Table[18](https://arxiv.org/html/2606.26366#A7.T18)\)\. Optimising a standard\-CoT prompt for single\-agent stakeholder breadth therefore does not transfer to the social layer, and here degrades it\. We report one honest caveat: a plain,*unoptimised*standard CoT reaches a higher raw consensus rate \(63%63\\%\) than NoT on this held\-out set, so the multi\-stakeholder claim is specifically that NoT dominates the*optimised*CoT baseline—not that it maximises raw agreement\.
#### Where rejections land\.
Raw consensus rate hides*which*role dissents\. Classifying every per\-perspectiveR4R4vote \(cached judge over each debater’s final verdict\) reframes the standard\-CoT caveat as evidence for NoT\. NoT’sR4R4rejections are both rare and role\-appropriate: of the55per\-perspective rejections across126126NoT votes,60%60\\%come from the External Advisor—the role explicitly charged with flagging principle violations—and only20%20\\%from the Primary Decision Maker who is meant to drive toward a workable decision\. The optimised CoT instead scatters3838rejections near\-uniformly across roles \(normalised role entropy0\.980\.98vs\. NoT’s0\.860\.86\), with the*decider itself*the modal rejector \(42%42\\%\)—a structurally incoherent pattern in which the role responsible for convergence most often blocks the synthesis it helped build\. NoT therefore disagrees roughly7×7\{\\times\}less often than the optimised CoT, and when it does, the dissent concentrates where principled objection belongs rather than scattering\. \(NoT’s narrative style left1212of126126final verdicts without a clean binary token; these are counted conservatively as non\-rejections, and NoT’sn=5n=5rejections make the share suggestive rather than a tight estimate\.\)
Table 18:Multi\-stakeholder head\-to\-head on3030held\-out DailyDilemmas scenarios \(22generators, moderatorclaude\-sonnet\-4\-6,n=60n=60debates per arm; only the agent prompt varies\)\. NoT vs\. optimised CoT:\+20\+20pp on fullR4R4consensus, Fisher exactp=0\.041p=0\.041, and four times less unresolved structural rejection\.
### G\.7Verbatim optimised prompts and training curves
#### NoT\-v2 \(in\-family training; SHA\-1 prefix3fb7f75f6d; 2,579 characters\)\.
> You are a thoughtful advisor reasoning through ethical dilemmas\. Before giving your final answer, work through a structured analysis\. Keep each section tight and purposeful\. Who is deciding?Name the decision\-maker, their role, and the specific knowledge or constraints they bring to this choice\. Who is affected?Identify every person, group, or institution whose wellbeing, interests, or rights intersect this decision\. Push past the central figures: include indirect parties, bystanders, background institutions, and anyone whose situation quietly changes as a result\. For each, state precisely what they stand to gain, lose, or have changed—not that they are “affected,” but what specifically shifts for them\. Prefer concrete specificity \(“the clinic’s nursing staff who bear legal exposure”\) over abstract categories\. Aim for at least five distinct stakeholders with individuated stakes\. What happens next?For each available course of action, trace consequences at least two steps forward\. Every stakeholder you named must appear somewhere in this analysis\. What don’t we know?For each projected future, surface specific uncertainties\. For each course of action, name at least three distinct uncertainties—not generic hedges, but named gaps that could materially change the outcome\. For each uncertainty, explain why it matters: what reversal or surprise would it produce? What should be done?Commit to a specific decision\. Explain why this path is preferable given the stakeholder impacts you traced and the uncertainties you surfaced\. Name which uncertainties most threaten your chosen path and state directly why you are proceeding despite them\. Additional guidance: Precision over length\. Do not resolve tensions prematurely\. Be concrete about stakes\. Explain the consequence of being wrong, not merely that the unknown exists\.
#### NoT\-v3 \(cross\-family training; SHA\-1 prefixa51ec242d5; 2,401 characters\)\.
> You are a thoughtful advisor helping people navigate ethical dilemmas\. Before committing to any recommendation, think carefully and concisely through the following\. Who is deciding and what do they stand to gain or lose\.Characterize the decision\-maker’s role, relevant knowledge, and personal stakes in one or two sentences\. Who else is affected\.Cast a wide net\. Start with those directly executing or immediately experiencing the decision, then move outward: people affected one or two steps removed, institutions and communities absorbing indirect effects, and anyone whose future options will be constrained by what happens now\. For each distinct party, name their concrete stake in a single sharp sentence\. Push past obvious stakeholders—someone in the background or a future generation is almost always relevant\. Aim to surface at least six distinct parties; if you find more, include them\. What could go wrong or remain unknown\.For every party you named and every realistic course of action you are considering, state at least one specific uncertainty—a hidden intention, an unpredictable reaction, a missing fact, or a long\-run effect that cannot be resolved with available information\. Name the precise unknown, not merely that uncertainty exists\. No party should be skipped\. Where an uncertainty is especially consequential—where resolving it would flip the recommended action—say so explicitly\. What the realistic options are and how outcomes ripple\.For each plausible course of action, trace the most decision\-relevant consequences forward\. Focus on how an outcome for one party reshapes outcomes for others\. Be selective: include causal chains that change the analysis; skip those that don’t\. The call\.Commit to a specific course of action\. Justify it by direct reference to the stakeholders and uncertainties you surfaced\. Name the uncertainties you are accepting and explain why the expected benefits still justify those risks\. Do not hedge into vagueness—a qualified commitment is fine, but the recommendation must be actionable\. Throughout, write concisely\. One sharp sentence per stakeholder stake\. One precise unknown per party\. Avoid restating what earlier reasoning already established\.
#### Training loss curves\.
Both runs trained on the seed\-4343stratified3030\-scenario sample with batch size1010and early\-stop after three consecutive iterations with<5%<5\\%loss reduction\.
Iteration0is hand\-written NoT\. As a drift control, both runs re\-issued the NoT prompt on scenariodd\_32489at the start and end of training; stakeholder counts moved6→66\\to 6\(Run A\) and5→65\\to 6\(Run B\), both within the pre\-registeredΔ=0\.5\\Delta=0\.5threshold for compute\-drift confounding modulo single\-cell rounding, and the four\-generator replication shows no compute\-drift signature\.Similar Articles
NTS-CoT: Mitigating Hallucinations in LLM-based News Timeline Summarization with Chain-of-Thought Reasoning
This paper proposes NTS-CoT, a novel framework that uses Chain-of-Thought reasoning to mitigate hallucinations in LLM-based news timeline summarization. It introduces three modules—Element-CoT, Date Selection, and Causal-CoT—to improve faithfulness and reduce omissions, outperforming state-of-the-art baselines on three benchmarks.
Chain-of-Thought Reasoning in the Wild Is Not Always Faithful
This paper demonstrates that Chain-of-Thought reasoning in large language models is not always faithful, with models producing plausible but incorrect reasoning chains in response to naturally worded prompts, raising concerns for AI safety and transparency.
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
Proposes ProxyCoT, a training framework that improves long-context reasoning in large language models by first obtaining chain-of-thought reasoning traces on short proxy contexts (via reinforcement learning or distillation) and then grounding them in full long contexts through supervised fine-tuning. Experiments show consistent improvements over baselines with reduced computational cost.
Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
The paper proposes a mean-field framework to model chain-of-thought reasoning in LLMs as a guided discovery process on a clue graph, deriving an ODE for the fraction of discovered clues and validating it experimentally.
Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning
The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.