Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

arXiv cs.CL Papers

Summary

This paper investigates when and why LLMs and humans converge or diverge in evaluating creativity, finding that alignment depends on the dimension (stronger for novelty, weaker for context) and that different LLMs apply different standards.

arXiv:2607.22218v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information. Moreover, each LLM exhibited distinct, model-specific standards that varied substantially in breadth. These differences in evaluation standards were reflected in actual creativity judgments. Study 2 (N = 1,103 ideas) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative. Study 3 (N = 1,195) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged. Together, our findings help explain the mixed evidence on LLM-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies. LLMs may resemble humans when evaluations emphasize intrinsic qualities such as novelty, yet diverge when judgments require contextual information. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognize different ideas as creative.
Original Article
View Cached Full Text

Cached at: 07/27/26, 07:40 AM

# Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
Source: [https://arxiv.org/abs/2607.22218](https://arxiv.org/abs/2607.22218)
[View PDF](https://arxiv.org/pdf/2607.22218)

> Abstract:Despite the growing use of large language models \(LLMs\) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments\. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications\. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards\. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information\. Moreover, each LLM exhibited distinct, model\-specific standards that varied substantially in breadth\. These differences in evaluation standards were reflected in actual creativity judgments\. Study 2 \(N = 1,103 ideas\) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative\. Study 3 \(N = 1,195\) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged\. Together, our findings help explain the mixed evidence on LLM\-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies\. LLMs may resemble humans when evaluations emphasize intrinsic qualities such as novelty, yet diverge when judgments require contextual information\. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognize different ideas as creative\.

## Submission history

From: Pengzhao Lyu \[[view email](https://arxiv.org/show-email/90cee87b/2607.22218)\] **\[v1\]**Fri, 24 Jul 2026 11:37:19 UTC \(7,287 KB\)

Similar Articles

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

arXiv cs.CL

This paper geometrically analyzes why LLMs acting as judges agree strongly with each other but weakly with humans, finding that inter-LLM consensus reflects a collapsed subspace rather than true human alignment on subjective rubrics. Post-hoc calibration on human data improves alignment, but even calibrated LLMs fall short of human reliability.