Tag
This paper introduces EUDAIMONIA, a benchmark for evaluating harmful social dynamics in LLMs, such as encouraging unhealthy intimacy or dependence. Testing 22 recent models, including Claude-Opus-4.7 and GPT-5.5, it finds persistent violation rates around 30%, suggesting these failures are social-alignment problems unsolved by extended reasoning.