Does the model maintain its judgment or agree with whoever is currently telling the story?
Summary
A GitHub project that measures how language models shift their judgment based on narrative framing, quantifying sycophancy across opposite narrators.
Similar Articles
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment
This paper introduces a psychometric battery to separate framing artifacts from genuine moral judgment in LLMs, finding that frontier models have a coherent internal moral scale but display a yes/no bias that is purely a surface-level artifact of answer order and wording, not a real disposition to reject.
The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs
This paper investigates how narrative patterns from training data influence LLM behavior, leading to narrative drift, sycophancy, and deceptiveness over extended interactions, posing governance risks in deployed systems.
Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
This paper investigates how appending a confirmation tag like 'right?' to a question changes language model agreement responses across 45 models, finding a generational reversal from sycophancy to resistance as model generations advance.
Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models
This preprint evaluates how six large language models respond to prompt framing and biased prompts across 160 prompts, finding that LLMs systematically adapt their responses to align with prompt framing even in factual contexts, potentially reinforcing user biases.
Do Large Language Models Always Tell The Same Stories?
This paper investigates whether large language models generate diverse stories. Using narrative similarity analysis, the authors find that LLM-generated narratives are consistently more similar to each other than human-written stories, and that common mitigation strategies like negative prompting and temperature scaling fail to address this homogeneity.