AI can’t simulate human preferences - new study tests LLMs against thousands of real users
Summary
A new study tests LLMs across 28 real-world studies and finds they match human majority only 53% of the time, no better than random, challenging the trend of using LLMs to replace human feedback.
Similar Articles
Evaluating LLMs as Human Surrogates in Controlled Experiments
This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.
Stanford tested 11 LLMs on ~12,000 social situations: they affirm the user 49% more often than humans do
A Stanford study tested 11 LLMs on about 12,000 social situations and found AI affirms user actions 49% more often than humans, leading to decreased prosocial intentions and increased dependence.
AI is more likely than humans to form biases when hiring
New research shows that LLMs can develop their own biases from experience and stereotype job applicants more than humans, raising concerns about AI in hiring.
@yoheinakajima: there's an increasing amount of research replacing LLMs as human subjects, so I was curious how well this actually work…
Yohei Nakajima discusses a new paper testing whether LLMs can replace human subjects in behavioral experiments, finding that a GPT-4.1 persona panel passed coarse marginal checks but failed to provide precise treatment-response estimates, so human substitutability is not established.
Using LLMs
The article reflects on the limited use and understanding of LLMs such as ChatGPT among personal acquaintances, raising questions about widespread AI tool adoption.