conformity

Tag

Cards List
#conformity

Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier

arXiv cs.AI · 3d ago Cached

A study measuring how large language models conform to unanimous peer opinions in multi-agent settings, finding that existing mitigations trade off resistance against receptivity, with reasoning being the only intervention that improves both on MMLU.

0 favorites 0 likes
#conformity

Social Pressure Breaks Majority Voting in LLM Safety Panels

arXiv cs.CL · 2026-08-06 Cached

This arXiv paper studies how shared social cues from simulated peers break the majority-voting protection in LLM safety panels. It shows that when all reviewers receive the same incorrect 'unsafe' label, panel false-alarm rates jump to 100%, revealing a failure mode and offering a pre-deployment diagnostic.

0 favorites 0 likes
← Back to home

Submit Feedback