Does the model maintain its judgment or agree with whoever is currently telling the story?

Reddit r/singularity Tools

Summary

A GitHub project that measures how language models shift their judgment based on narrative framing, quantifying sycophancy across opposite narrators.

https://github.com/lechmazur/sycophancy Positive values mean first-person framing shifts the model toward the narrator more often than away from them. Negative values mean the reverse. This chart counts both ways a model can contradict itself across opposite narrators, agreeing with both or rejecting both; lower is better. Models differ sharply in how willing they are to decide who is more right.
Original Article

Similar Articles

Do Large Language Models Always Tell The Same Stories?

arXiv cs.CL

This paper investigates whether large language models generate diverse stories. Using narrative similarity analysis, the authors find that LLM-generated narratives are consistently more similar to each other than human-written stories, and that common mitigation strategies like negative prompting and temperature scaling fail to address this homogeneity.