Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
Summary
This paper investigates when and why LLMs and humans converge or diverge in evaluating creativity, finding that alignment depends on the dimension (stronger for novelty, weaker for context) and that different LLMs apply different standards.
View Cached Full Text
Cached at: 07/27/26, 07:40 AM
# Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity Source: [https://arxiv.org/abs/2607.22218](https://arxiv.org/abs/2607.22218) [View PDF](https://arxiv.org/pdf/2607.22218) > Abstract:Despite the growing use of large language models \(LLMs\) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments\. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications\. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards\. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information\. Moreover, each LLM exhibited distinct, model\-specific standards that varied substantially in breadth\. These differences in evaluation standards were reflected in actual creativity judgments\. Study 2 \(N = 1,103 ideas\) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative\. Study 3 \(N = 1,195\) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged\. Together, our findings help explain the mixed evidence on LLM\-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies\. LLMs may resemble humans when evaluations emphasize intrinsic qualities such as novelty, yet diverge when judgments require contextual information\. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognize different ideas as creative\. ## Submission history From: Pengzhao Lyu \[[view email](https://arxiv.org/show-email/90cee87b/2607.22218)\] **\[v1\]**Fri, 24 Jul 2026 11:37:19 UTC \(7,287 KB\)
Similar Articles
Assessing the Creativity of Large Language Models: Testing, Limits, and New Frontiers
This paper systematically evaluates human creativity tests for LLMs and finds they fail to predict scientific ideation. It introduces the DRAT, a new test that combines convergent and divergent thinking to reliably predict scientific ideation ability in language models.
Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models
This paper studies how humans and large language models linguistically accommodate each other during multi-turn conversations, finding that LLMs overconverge to user style while humans accommodate LLMs no differently than humans.
CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse
Introduces CreativityNeuro, a data-free method that steers language model weights to enhance divergent thinking and reduce mode collapse, achieving significant improvements in creativity assessments without retraining or fine-tuning.
The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment
This paper geometrically analyzes why LLMs acting as judges agree strongly with each other but weakly with humans, finding that inter-LLM consensus reflects a collapsed subspace rather than true human alignment on subjective rubrics. Post-hoc calibration on human data improves alignment, but even calibrated LLMs fall short of human reliability.
Synthetic Consumer Insight Generation with Large Language Models
This research examines whether LLMs can generate synthetic consumer data for projective techniques, comparing human and LLM responses on city tourism perceptions and finding substantial overlap but differences in style and diversity.