synthetic-users

Tag

Cards List
#synthetic-users

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

arXiv cs.CL · 2026-07-30 Cached

This paper benchmarks LLM-simulated human survey responses across two large-scale datasets, finding that no model beats simple baselines at the individual level and that models systematically over-determine demographics, distorting segment differences. The failures persist across model scales and families, raising concerns about using synthetic users for decision support.

0 favorites 0 likes
#synthetic-users

AI can’t simulate human preferences - new study tests LLMs against thousands of real users

Reddit r/ArtificialInteligence · 2026-07-07

A new study tests LLMs across 28 real-world studies and finds they match human majority only 53% of the time, no better than random, challenging the trend of using LLMs to replace human feedback.

0 favorites 0 likes
#synthetic-users

The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study

arXiv cs.CL · 2026-05-21 Cached

This paper from Google DeepMind and Carnegie Mellon argues that LLM-simulated experiments are actually observational studies due to user drift, where the simulated population shifts with interventions. The authors propose using negative control outcomes to diagnose confounding and show that eliciting setting-relevant confounders can reduce bias.

0 favorites 0 likes
← Back to home

Submit Feedback