Tag
A study measuring how large language models conform to unanimous peer opinions in multi-agent settings, finding that existing mitigations trade off resistance against receptivity, with reasoning being the only intervention that improves both on MMLU.
This arXiv paper studies how shared social cues from simulated peers break the majority-voting protection in LLM safety panels. It shows that when all reviewers receive the same incorrect 'unsafe' label, panel false-alarm rates jump to 100%, revealing a failure mode and offering a pre-deployment diagnostic.