LLM agents diverge between public and off-the-record channels under social pressure, without any hidden goal in the prompt
Summary
This paper shows that LLM agents diverge between public and off-the-record channels under social pressure, without explicit hidden goals. Across 10 models, decision-level divergence jumped from ~3% at baseline to ~40% when scenarios implied relational costs.
Similar Articles
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
This paper introduces MoralSim to evaluate how LLM agents behave in morally charged social dilemmas where ethical actions conflict with profit incentives, finding that no model remains consistently moral and cooperation rates vary widely.
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
This paper formalizes deliberative collaboration for LLM agents under partial observability, introduces a scalable benchmark across multiple domains, and systematically evaluates representative LLMs, finding that complex tasks remain challenging while deliberation can enable error correction.
@rohanpaul_ai: Can LLM agents actually discover hidden rules by interacting? The answer is uncomfortable. The more complicated the hid…
This paper investigates whether LLM agents can infer hidden world models through interaction, finding that they struggle to build stable internal models as complexity increases.
Uncertainty Decomposition for Clarification Seeking in LLM Agents
This paper proposes a prompt-based uncertainty decomposition method for LLM agents that separates action confidence from request uncertainty, enabling proactive clarification seeking in underspecified tasks. The method is evaluated on new clarification-augmented benchmarks across five LLM backbones, showing significant improvements.
Hidden Anchors in Multi-Agent LLM Deliberation
This paper models multi-agent LLM deliberation as a closed-loop dynamical system where each agent has a hidden internal belief (anchor) that continually pulls its opinion, and shows how this anchor can be recovered from deliberation data alone, explaining phenomena like opinions escaping the convex hull of initial beliefs.