Synthetic Consumer Insight Generation with Large Language Models
Summary
This research examines whether LLMs can generate synthetic consumer data for projective techniques, comparing human and LLM responses on city tourism perceptions and finding substantial overlap but differences in style and diversity.
View Cached Full Text
Cached at: 07/08/26, 04:38 AM
# Synthetic Consumer Insight Generation with Large Language Models Source: [https://arxiv.org/abs/2607.05761](https://arxiv.org/abs/2607.05761) [View PDF](https://arxiv.org/pdf/2607.05761) > Abstract:Modern data\-driven marketing relies on large amounts of consumer data, yet collecting such data can be costly, time\-consuming, and difficult to scale\. This research examines whether large language models \(LLMs\) can be used to generate synthetic consumer data for projective techniques, a set of methods designed to elicit consumer associations, emotions, wants, and needs\. We test LLM\-generated responses across multiple projective tasks, LLMs, prompting strategies, and temperature settings, and compare them with human responses from a primary research study on perceptions of city tourism destinations\. Human and LLM responses were analyzed using linguistic measures, diversity and concentration metrics, topic models, and top\-term analyses\. The results show substantial overlap between human and LLM responses in broad topics and associations, but also important differences in style, linguistic structure, and the way diversity is generated\. Recommendations are given on how to best utilize LLMs for generating synthetic consumer data, how model and prompt choices shape response quality, and on recognizing the limitations of LLM synthetic consumer data generation\. ## Submission history From: Stephen L\. France \[[view email](https://arxiv.org/show-email/827ad16a/2607.05761)\] **\[v1\]**Tue, 7 Jul 2026 02:38:41 UTC \(985 KB\)
Similar Articles
Evaluating LLMs as Human Surrogates in Controlled Experiments
This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.
Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction
This paper proposes a large language model-driven data augmentation framework using GPT-5 to generate synthetic oral monologues from written anchors for cognitive score prediction from speech. A similarity-guided selection strategy consistently reduces prediction error, particularly for minority low-score participants.
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
This paper presents a psychometric audit of large language models as synthetic survey respondents, finding that they fail to match human joint distributions and reliability, making them unsuitable replacements.
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
This paper benchmarks LLM-simulated human survey responses across two large-scale datasets, finding that no model beats simple baselines at the individual level and that models systematically over-determine demographics, distorting segment differences. The failures persist across model scales and families, raising concerns about using synthetic users for decision support.
Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies
This paper surveys clinical communication processing using LLM-generated synthetic data and presents 13 case studies across EMS reports, nurse handoffs, and more, showing that synthetic data can bootstrap clinical NLP systems.