打破同质化:多样化角色集以提升大语言模型的创造性输出

arXiv cs.CL 论文

摘要

本文将角色多样化阐述为一个集合级条件问题,以缓解大语言模型输出中的同质化现象,并评估了在跨任务(如Alternative Uses Task)中增强创造性和多样性的方法。

arXiv:2609.30492v1 Announce Type: new Abstract: Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning problem and study two orthogonal design choices: selecting versus generating personas, and space-filling versus frontier-seeking diversity. We instantiate this design space with four methods spanning coverage and dispersion subset selections, uniform-coverage sampling, and evolutionary persona generation. Evaluations on the Alternative Uses Task (AUT), Infinity-Chat, and Divergent Association Task (DAT) show the benefits of the proposed methods across tasks and creativity objectives. On AUT, evolutionary persona generation increases response diversity by 78.8%, originality by 26.1%, flexibility by 49.5%, and holistic creativity by 13.9% over task-only prompting, while maintaining 98.5% validity; on Infinity-Chat, it nearly doubles persona-induced response separation relative to random personas. Moreover, evolutionary personas compose with creativity-optimized prompting, further increasing its response diversity by 18.6% and creativity by 6.3%. These results establish persona-set geometry as a task-agnostic mechanism for eliciting divergent LLM outputs, and support persona diversification as a reusable complement to prompt optimization.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:38

# Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs
Source: [https://arxiv.org/html/2609.30492](https://arxiv.org/html/2609.30492)
Sang Bin MoonAffiliation:School of Electrical and Computer EngineeringAffiliation:Purdue UniversityAffiliation:West Lafayette, IN 47909, USAEmail:[moon182@purdue\.edu](mailto:)Daniel Borrajo & Sumitra GaneshAffiliation:J\.P\. Morgan AI ResearchAffiliation:New York, NY 10172, USAAffiliation:\{nicole\.cho, daniel\.borrajo,Email:[sumitra\.ganesh\}@jpmorgan\.com](mailto:)Abolfazl HashemiAffiliation:School of Electrical and Computer EngineeringAffiliation:Purdue UniversityAffiliation:West Lafayette, IN 47909, USAEmail:[abolfazl@purdue\.edu](mailto:)

###### Abstract

Language models often produce homogeneous responses to open\-ended tasks; such homogeneity can spawn groupthink—the convergence of ideas toward a singular and potentially suboptimal decision\. We formulate persona diversification as a set\-level conditioning problem and study two orthogonal design choices: selecting versus generating personas, and space\-filling versus frontier\-seeking diversity\. We instantiate this design space with four methods spanning coverage and dispersion subset selections, uniform\-coverage sampling, and evolutionary persona generation\. Evaluations on the Alternative Uses Task \(AUT\), Infinity\-Chat, and Divergent Association Task \(DAT\) show the benefits of the proposed methods across tasks and creativity objectives\. On AUT, evolutionary persona generation increases response diversity by 78\.8%, originality by 26\.1%, flexibility by 49\.5%, and holistic creativity by 13\.9% over task\-only prompting, while maintaining 98\.5% validity; on Infinity\-Chat, it nearly doubles persona\-induced response separation relative to random personas\. Moreover, evolutionary personas compose with creativity\-optimized prompting, further increasing its response diversity by 18\.6% and creativity by 6\.3%\. These results establish persona\-set geometry as a task\-agnostic mechanism for eliciting divergent LLM outputs, and support persona diversification as a reusable complement to prompt optimization\.

## 1Introduction

Homogeneity in Large Language Model \(LLM\) outputs, specifically those that pertain to a narrow subset of WEIRD \(Western, Educated, Industrialized, Rich, and Democratic\) responses, severely limits the potential of LLMs\([Anthis et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib1)\)and engenders a problematic artificial hivemind\([Jiang et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib26)\)\. This homogeneity is known to be spawned by the RLHF fine\-tuning pipeline endemic to frontier models\([Kirk et al\., 2024](https://arxiv.org/html/2609.30492#bib.bib29);[West & Potts, 2025](https://arxiv.org/html/2609.30492#bib.bib57)\); such homogeneity not only stifles creativity but also creates systematic risks of groupthink\. For example, relying on LLMs for investment recommendations can engender similar or identical advice across independent teams \- which results in compounded risk\([IMF, 2024](https://arxiv.org/html/2609.30492#bib.bib24);[BIS, 2024](https://arxiv.org/html/2609.30492#bib.bib4)\)\.

Conditioning a language model on a*persona*, defined as the societal, biological, and personal traits of the human user, has been presented as a potential lever for steering behavior and eliciting distinct perspectives\([Choi & Li, 2024](https://arxiv.org/html/2609.30492#bib.bib8);[Ge et al\., 2024](https://arxiv.org/html/2609.30492#bib.bib16)\)\. However, empirically measuring the impact of diverse personas on the model’s creative ability to generate divergent solutions for the same task is relatively understudied\([Paglieri et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib44)\)\. We observe three gaps in current research: first, the limited breadth of evolutionary methods that truly make asetof personas diverse; second, the underexplored notions of what actually makes a set of personasdiverseand a set of outputscreative; and third, the understudied realm of whether a set of diverse personas can trigger the model’s divergent trajectories to generate more creative outputs\.

In this work, we utilize persona as an explicit, optimizable variable and hypothesize that a set of truly diverse personas can successfully trigger the alternative, divergent trajectories embedded within the model that would otherwise remain inactivated\. Moreover, we hypothesize that divergent gains compound when a*set*of personas is explicitly diversified, ensuring that the induced responses spread across many regions at once, for the same task\. Thus, our focus is threefold: what makes a set of personas*diverse*, how can we*induce*that diversity with guarantees on the resulting persona\-side spread, and does the spread ultimately improve creativity?

![Refer to caption](https://arxiv.org/html/2609.30492v1/Design_Framework.png)Figure 1:Persona diversification design space\.Two*principles*of diversity \(columns\) are each instantiated at two*levels of control*\(rows\)\. The*space\-filling*principle is implemented byCoverageand UC\-MCMC, whereas the*frontier\-seeking*principle is instantiated byDispersionand evolutionaryTextGrad\. Three algorithms have formal guarantees for persona\-space objectives and sampling targets, while evolutionary generation is empirical\.To answer these questions, we organize persona diversification along two axes \(Figure[1](https://arxiv.org/html/2609.30492#S1.F1)\): selection from a fixed candidate pool versus generation beyond it, and space\-filling coverage versus frontier\-seeking separation\. Crossing these two axes yields four methods:CoverageandDispersionfor selection, and Uniform\-Coverage MCMC \(UC\-MCMC\) and evolutionaryTextGradfor generation\([Yuksekgonul et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib58)\)\.CoverageandDispersionachieve global optima for their population objectives, UC\-MCMC converges asymptotically to an equal\-cell target over reachable cells, and evolutionary generation is evaluated empirically\.

Our central finding is that the performance gains are driven by the specific structural geometry of the persona sets, not merely the presence of a conditioning variable\. Across the benchmarks, greater persona\-set dispersion is positively associated with greater response dispersion\. By quantifying this displacement against the model’s unconditioned output, we demonstrate the specific advantage a persona\-conditioned query provides over a standard task\-only prompt: namely, increased creativity and a meaningful departure from the model’s narrow default behavior\.

Our diversification objectives depend on persona geometry rather than downstream task performance, allowing the same persona sets to be reused across tasks\. Persona conditioning supplements prompt engineering and can be combined with advanced prompting and reasoning scaffolds\.

##### Contributions\.

- ▶\\blacktrianglerightA*design space*for inducing persona diversity that crosses two levels of control \(selection vs\. generation\) with two diversity principles \(space\-filling vs\. frontier\), unifying four methods under one framework \(Figure[1](https://arxiv.org/html/2609.30492#S1.F1)\)\. The construction is task\-agnostic and compatible with prompt optimization and reasoning scaffolds\.
- ▶\\blacktrianglerightOn the*selection*axis, exact algorithms for both principles, maximal\-covering locationCoverageand max–minDispersion, with global\-optimality guarantees\.
- ▶\\blacktrianglerightOn the*generation*axis, two complementary methods: a Metropolis\-within\-Gibbs sampler with provable asymptotic convergence to an equal\-cell mixture over reachable semantic cells, and an evolutionaryTextGradprocedure that searches for isolated, low\-density personas\.
- ▶\\blacktrianglerightEmpirical evidence that persona diversification can improve response creativity across three benchmarks, including comparisons with task\-only, random\-persona and prompt engineering baselines\.

![Refer to caption](https://arxiv.org/html/2609.30492v1/figures/response_geometry_tin_can.png)Figure 2:Persona conditioning diversifies Alternative Uses Task responses\.Left:Example uses for a tin can under task\-only prompting \(top\) and persona conditioning \(bottom\), illustrating a broader range of ideas\.Right:Mahalanobis\-whitened PCA projection of response embeddings across persona conditions, with convex hulls showing their semantic spread\.

## 2Problem Formulation

While a single persona provides pointwise conditioning, it cannot dictate the structure of a broader collection of viewpoints\. To address this set\-level challenge, we define the*persona diversification*problem\. The objective is to design a collection of persona\-conditioned queries that hold the task prompt and base response model constant, isolating the persona set as the sole optimization variable\. We formulate this shared problem and its underlying semantic geometry below\.

### 2\.1The Set\-Level Conditioning Problem

Let𝒯\\mathcal\{T\}be the text\-string space andΩ⊆𝒯\\Omega\\subseteq\\mathcal\{T\}be the specialized persona\-string space\. A persona is defined as a short, natural\-language descriptionp∈Ωp\\in\\Omega, and a persona condition is a set𝒮\\mathcal\{S\}of fixed cardinality\|𝒮\|=k\|\\mathcal\{S\}\|=k\. For a task promptx∼𝒟x\\sim\\mathcal\{D\}and a frozen downstream response modelGθRG\_\{\\theta\_\{\\mathrm\{R\}\}\}, the model generates a response conditioned onppas

yp∼GθR\(⋅∣x,p\),p∈𝒮,y\_\{p\}\\sim G\_\{\\theta\_\{\\mathrm\{R\}\}\}\(\\,\\cdot\\mid x,p\),\\qquad p\\in\\mathcal\{S\},\(1\)wherep=∅p=\\varnothingdenotes the persona\-free reference baseline\. HoldingxxandGθRG\_\{\\theta\_\{\\mathrm\{R\}\}\}fixed, the downstream goal is for the induced response collection\{yp:p∈𝒮\}\\\{y\_\{p\}:p\\in\\mathcal\{S\}\\\}to span a broader, more creative range of valid, useful outputs than task\-only prompting or an unoptimized persona set\.

Crucially, our optimization variable and evaluation target are deliberately distinct\. Our algorithms construct persona sets based entirely on frozen*persona\-space*geometry, whereas creativity is evaluated strictly on the induced*response\-space*outputs\. We do not assume that persona\-space distance monotonically transfers to response\-space distance, novelty, or quality\. Consequently, the guarantees we provide apply solely to the persona\-space; whether this persona diversification reliably translates to improved response creativity remains an empirical hypothesis evaluated in[Section4](https://arxiv.org/html/2609.30492#S4)\.

### 2\.2Semantic Geometry

Let𝒫=\{pi\}i=1N𝒫⊆Ω\\mathcal\{P\}=\\\{p\_\{i\}\\\}\_\{i=1\}^\{N\_\{\\mathcal\{P\}\}\}\\subseteq\\Omegadenote a generic finite candidate population, with\[N𝒫\]=\{1,…,N𝒫\}\[N\_\{\\mathcal\{P\}\}\]=\\\{1,\\ldots,N\_\{\\mathcal\{P\}\}\\\}\. Selection operates on the baseline population𝒫0\\mathcal\{P\}\_\{0\}, while generation expands it into𝒫T\\mathcal\{P\}\_\{T\}before downstream selection\. Letϕ:𝒯→ℝd\\phi:\\mathcal\{T\}\\to\\mathbb\{R\}^\{d\}denote the full, unit\-normalized output of a frozen text encoder, which supports Matryoshka Representation Learning \(MRL\)\([Kusupati et al\., 2022](https://arxiv.org/html/2609.30492#bib.bib32)\)\. For Mahalanobis distance, we use the prefixϕ~\(p\)=ϕ\(p\)1:dM\\widetilde\{\\phi\}\(p\)=\\phi\(p\)\_\{1:d\_\{\\mathrm\{M\}\}\}without renormalizing after truncation\.

This encoder induces persona\-space dissimilaritiesdmPd\_\{m\}^\{\\mathrm\{P\}\}, indexed bym∈\{cos,2,Mah\}m\\in\\\{\\mathrm\{cos\},2,\\mathrm\{Mah\}\\\}, as

dcosP​\(p,p′\)\\displaystyle d\_\{\\mathrm\{cos\}\}^\{\\mathrm\{P\}\}\(p,p^\{\\prime\}\)=1−ϕ​\(p\)⊤​ϕ​\(p′\)‖ϕ⁡\(p\)‖2​‖ϕ⁡\(p′\)‖2,\\displaystyle=1\-\\frac\{\\phi\(p\)^\{\\top\}\\phi\(p^\{\\prime\}\)\}\{\\\|\\phi\(p\)\\\|\_\{2\}\\\|\\phi\(p^\{\\prime\}\)\\\|\_\{2\}\},\(2\)d2P​\(p,p′\)\\displaystyle d\_\{2\}^\{\\mathrm\{P\}\}\(p,p^\{\\prime\}\)=‖ϕ⁡\(p\)−ϕ⁡\(p′\)‖2,\\displaystyle=\\\|\\phi\(p\)\-\\phi\(p^\{\\prime\}\)\\\|\_\{2\},\(3\)dMahP​\(p,p′\)\\displaystyle d\_\{\\mathrm\{Mah\}\}^\{\\mathrm\{P\}\}\(p,p^\{\\prime\}\)=\(ϕ~​\(p\)−ϕ~​\(p′\)\)⊤​Σ^P−1​\(ϕ~​\(p\)−ϕ~​\(p′\)\)\.\\displaystyle=\\sqrt\{\(\\widetilde\{\\phi\}\(p\)\-\\widetilde\{\\phi\}\(p^\{\\prime\}\)\)^\{\\top\}\\widehat\{\\Sigma\}\_\{\\mathrm\{P\}\}^\{\-1\}\(\\widetilde\{\\phi\}\(p\)\-\\widetilde\{\\phi\}\(p^\{\\prime\}\)\)\}\.\(4\)Here,Σ^P\\widehat\{\\Sigma\}\_\{\\mathrm\{P\}\}is the sample covariance of the truncated embeddings\{ϕ~​\(p\):p∈𝒫0\}\\\{\\widetilde\{\\phi\}\(p\):p\\in\\mathcal\{P\}\_\{0\}\\\}, regularized by addingε​IdM\\varepsilon I\_\{d\_\{\\mathrm\{M\}\}\}, withε=10−6\\varepsilon=10^\{\-6\}anddM=128d\_\{\\mathrm\{M\}\}=128\.

Both finite\-pool selectors rely on the pairwise dissimilarity matrix and its realized spectrum

Dm,i​jP\\displaystyle D^\{\\mathrm\{P\}\}\_\{m,ij\}=dmP​\(pi,pj\),\\displaystyle=d\_\{m\}^\{\\mathrm\{P\}\}\(p\_\{i\},p\_\{j\}\),𝐃mP\\displaystyle\\mathbf\{D\}\_\{m\}^\{\\mathrm\{P\}\}=\(Dm,i​jP\)i,j=1N𝒫,\\displaystyle=\\bigl\(D^\{\\mathrm\{P\}\}\_\{m,ij\}\\bigr\)\_\{i,j=1\}^\{N\_\{\\mathcal\{P\}\}\},\(5\)Λm​\(𝒫\)\\displaystyle\\Lambda\_\{m\}\(\\mathcal\{P\}\)=\{Dm,i​jP:i,j∈\[N𝒫\]\},\\displaystyle=\\\{D^\{\\mathrm\{P\}\}\_\{m,ij\}:i,j\\in\[N\_\{\\mathcal\{P\}\}\]\\\},Lm\\displaystyle L\_\{m\}=\|Λm​\(𝒫\)\|,\\displaystyle=\|\\Lambda\_\{m\}\(\\mathcal\{P\}\)\|,\(6\)Letrrdenote a unique dissimilarity value \(or radius\) drawn from this spectrum\. TheLmL\_\{m\}distinct values ofΛm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\), including00, are sorted in ascending order such thatrm,\(1\)<⋯<rm,\(Lm\)r\_\{m,\(1\)\}<\\cdots<r\_\{m,\(L\_\{m\}\)\}\.

## 3Diversification Algorithms

Building on the semantic geometry established in[Section2](https://arxiv.org/html/2609.30492#S2), this section structures our algorithms as a progression in both control and exploratory reach\. We begin withSelectionoperators \([Section3\.2](https://arxiv.org/html/2609.30492#S3.SS2)\) that enforce space\-filling and frontier diversity within a fixed candidate population\. To break the performance ceiling imposed by this finite pool, we escalate to open\-endedGenerationmethods \([Sections3\.3](https://arxiv.org/html/2609.30492#S3.SS3)and[3\.4](https://arxiv.org/html/2609.30492#S3.SS4)\), actively expanding the persona support through a space\-filling MCMC sampler and a frontier\-seeking evolutionary TextGrad search\. To maintain narrative flow, all formal proofs and extended mathematical formulations are deferred to Appendix[B](https://arxiv.org/html/2609.30492#A2)and[C](https://arxiv.org/html/2609.30492#A3)\.

### 3\.1Diversification Design Space

Our framework is organized along two independent axes, as illustrated in[Figure1](https://arxiv.org/html/2609.30492#S1.F1)\. Selection extracts a subset𝒮⊆𝒫\\mathcal\{S\}\\subseteq\\mathcal\{P\}of sizekkfrom a frozen candidate pool𝒫\\mathcal\{P\}, whereas generation expands the baseline pool𝒫0\\mathcal\{P\}\_\{0\}with new valid personas before applying a downstream selection protocol to obtain the finalkk\-persona condition\. We analyze generation and downstream selection separately because guarantees for the generation process do not automatically transfer to the selected subset\.

### 3\.2Diverse Persona Selection: Coverage and Dispersion

To extract akk\-sized subset𝒮\\mathcal\{S\}from a finite candidate pool𝒫\\mathcal\{P\}, we operationalize the space\-filling and frontier principles via two exact selection algorithms\. For the space\-filling principle, we formulate selection as thresholded coverage\. A personapip\_\{i\}is considered covered if its distance to the nearest selected exemplar in𝒮\\mathcal\{S\}is at most a radiusrr\. The fraction of the population covered evaluates to

Fcov,m\(𝒮;𝒫,r\)=1N𝒫∑i=1N𝒫𝟙\{minpj∈𝒮dmP\(pi,pj\)≤r\}\.F\_\{\\mathrm\{cov\},m\}\(\\mathcal\{S\};\\mathcal\{P\},r\)=\\frac\{1\}\{N\_\{\\mathcal\{P\}\}\}\\sum\_\{i=1\}^\{N\_\{\\mathcal\{P\}\}\}\\mathbbm\{1\}\\Bigl\\\{\\min\_\{p\_\{j\}\\in\\mathcal\{S\}\}d^\{\\mathrm\{P\}\}\_\{m\}\(p\_\{i\},p\_\{j\}\)\\leq r\\Bigr\\\}\.\(7\)To avoid metric\-specific scaling from manually tuningrr, we adaptively calibrate the smallest radiusrm⋆r^\{\\star\}\_\{m\}at which some subset successfully covers the target ofncov=⌈ccov​N𝒫⌉n\_\{\\mathrm\{cov\}\}=\\lceil c\_\{\\mathrm\{cov\}\}N\_\{\\mathcal\{P\}\}\\rceilpersonas, written as

rm⋆​\(𝒫,ccov\)=min⁡\{r≥0:max𝒮⊆𝒫\|𝒮\|=k⁡N𝒫​Fcov,m​\(𝒮,𝒫,r\)≥ncov\}\.r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)=\\min\\Bigl\\\{r\\geq 0:\\max\_\{\\begin\{subarray\}\{c\}\\mathcal\{S\}\\subseteq\\mathcal\{P\}\\\\ \|\\mathcal\{S\}\|=k\\end\{subarray\}\}N\_\{\\mathcal\{P\}\}\\,F\_\{\\mathrm\{cov\},m\}\(\\mathcal\{S\};\\mathcal\{P\},r\)\\geq n\_\{\\mathrm\{cov\}\}\\Bigr\\\}\.The algorithm lexicographically attains this minimal radius, then breaks ties among radius\-optimal subsets by maximizing the covered count at the threshold\.

SinceFcov,mF\_\{\\mathrm\{cov\},m\}changes only at the discrete values of the pairwise spectrumΛm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\), findingr⋆r^\{\\star\}reduces to an exact discrete search over the sorted distances\. For a target coverage countncovn\_\{\\mathrm\{cov\}\}, the greedy solutionGm​\(r\)≥ncovG\_\{m\}\(r\)\\geq n\_\{\\mathrm\{cov\}\}certifies feasibility of a trial radiusrrand establishes the upper boundr⋆≤rr^\{\\star\}\\leq r\. By searching over the sorted distances, the upper bound becomes tighter\. At the same time, the classical submodular guarantee bounds the optimal covered count from above asnopt≤⌊Gm​\(r\)/\(1−1/e\)⌋n\_\{\\mathrm\{opt\}\}\\leq\\lfloor G\_\{m\}\(r\)/\(1\-1/e\)\\rfloor\([Nemhauser et al\., 1978](https://arxiv.org/html/2609.30492#bib.bib40)\)\. Ifncov\>⌊Gm​\(r\)/\(1−1/e\)⌋n\_\{\\mathrm\{cov\}\}\>\\lfloor G\_\{m\}\(r\)/\(1\-1/e\)\\rfloor, the radius is certified infeasible, establishing the lower boundr⋆\>rr^\{\\star\}\>r\. These certificates narrow the bracket onr⋆r^\{\\star\}, efficiently reducing the number of exact\-solver calls\. This iterative algorithm \(Algorithm[1](https://arxiv.org/html/2609.30492#alg1)\) solves both the smallest feasible radiusr⋆r^\{\\star\}and the maximum coverage atr⋆r^\{\\star\}\.

###### Theorem 1\(Exact Calibrated Coverage\)\.

In exact mode, provided every radius feasibility probe is certified, the adaptive search algorithm \(Algorithm[1](https://arxiv.org/html/2609.30492#alg1)\) terminates and returns a globally minimal radiusr⋆=rm⋆​\(𝒫,ccov\)r^\{\\star\}=r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\), together with a subset𝒮acov,m⋆​\(𝒫,ccov\)\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{acov\},m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)of cardinalitykksatisfying the maximum coverage atr⋆r^\{\\star\}\. For every𝒮⊆𝒫\\mathcal\{S\}\\subseteq\\mathcal\{P\}with\|𝒮\|=k\|\\mathcal\{S\}\|=kand1≤k≤N𝒫1\\leq k\\leq N\_\{\\mathcal\{P\}\},

Fcov,m​\(𝒮acov,m⋆​\(𝒫,ccov\),𝒫,r⋆\)≥Fcov,m​\(𝒮,𝒫,r⋆\)\.F\_\{\\mathrm\{cov\},m\}\\\!\\left\(\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{acov\},m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\);\\mathcal\{P\},r^\{\\star\}\\right\)\\geq F\_\{\\mathrm\{cov\},m\}\(\\mathcal\{S\};\\mathcal\{P\},r^\{\\star\}\)\.In particular, its covered count atr⋆r^\{\\star\}is at leastncovn\_\{\\mathrm\{cov\}\}\.

The full Integer Linear Program formulation, algorithmic pseudocode, and supporting lemmas and proofs are deferred to Appendix[B\.1](https://arxiv.org/html/2609.30492#A2.SS1)\.

On the other hand, the frontier principle prioritizes structural separation\. The discrete max–minpp\-dispersion problem\([Erkut, 1990](https://arxiv.org/html/2609.30492#bib.bib11)\)maximizes the minimum dispersion \(the smallest pairwise dissimilarity within a subset\) over all size\-kksubsets of the candidate pool\. For\|𝒮\|≥2\|\\mathcal\{S\}\|\\geq 2, the minimum dispersion objective is defined as

Fdisp,m​\(𝒮\):=minp,p′∈𝒮p≠p′⁡dmP​\(p,p′\)\.F\_\{\\mathrm\{disp\},m\}\(\\mathcal\{S\}\):=\\min\_\{\\begin\{subarray\}\{c\}p,p^\{\\prime\}\\in\\mathcal\{S\}\\\\ p\\neq p^\{\\prime\}\\end\{subarray\}\}d\_\{m\}^\{\\mathrm\{P\}\}\(p,p^\{\\prime\}\)\.\(8\)Since finding a max–min diverse subset is generally NP\-hard, our exact method again exploits the discrete topology of the solution space\. Note that the optimal dispersion value must equal exactly one of the pairwise distances in the sorted spectrumΛm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\), so we leverage the threshold\-graph equivalence between max–min diversity and clique feasibility\([Della Croce et al\., 2009](https://arxiv.org/html/2609.30492#bib.bib10)\)\. The algorithm builds a graph by inserting edges in descending order of pairwise dissimilarity, transforming the optimization into a series of Boolean structural checks\. After a new edge\{pa,pb\}\\\{p\_\{a\},p\_\{b\}\\\}is added, the algorithm searches for a\(k−2\)\(k\-2\)\-clique in their mutual neighborhood\. If found, combining it with the endpoints yields a complete size\-kkclique; if not, the algorithm checks for the next pair\. Since the algorithm searches the descending discrete distance spectrum sequentially, the subset returned upon the first clique completion is guaranteed to be optimal\.

###### Theorem 2\(Exact Max–Min Dispersion\)\.

For a finite candidate pool𝒫\\mathcal\{P\},2≤k≤\|𝒫\|2\\leq k\\leq\|\\mathcal\{P\}\|, and fixed symmetric nonnegative dissimilaritiesdmPd\_\{m\}^\{\\mathrm\{P\}\}, the incremental\-clique algorithm \(Algorithm[2](https://arxiv.org/html/2609.30492#alg2)\) terminates and returns a size\-kksubset𝒮disp,m⋆​\(𝒫\)\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{disp\},m\}\(\\mathcal\{P\}\)satisfying

Fdisp,m​\(𝒮disp,m⋆​\(𝒫\)\)=max𝒮⊆𝒫\|𝒮\|=k⁡Fdisp,m​\(𝒮\)\.F\_\{\\mathrm\{disp\},m\}\(\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{disp\},m\}\(\\mathcal\{P\}\)\)=\\max\_\{\\begin\{subarray\}\{c\}\\mathcal\{S\}\\subseteq\\mathcal\{P\}\\\\ \|\\mathcal\{S\}\|=k\\end\{subarray\}\}F\_\{\\mathrm\{disp\},m\}\(\\mathcal\{S\}\)\.

### 3\.3Uniform\-Coverage MCMC Persona Generation

To break the diversity ceiling imposed by a finite baseline pool𝒫0\\mathcal\{P\}\_\{0\}, we transition to generation beyond𝒫0\\mathcal\{P\}\_\{0\}\.*Uniform\-Coverage MCMC*\(UC\-MCMC\) applies the space\-filling principle directly to the generation process by targeting equal probability mass across reachable cells of a fixed partition of equal\-area directional cells\. Within each cell, a frozen Large Language Model acts as the base law, ensuring the sampled personas remain linguistically plausible and structurally valid while the MCMC framework enforces the geometric uniformity\.

Since the total number of reachable cells is unknown in advance, UC\-MCMC employs a hindsight\-spawning random scan that maintains a dynamically growing population of parallel Markov chains, with one chain for each discovered cell\. Each iteration selects an active cell uniformly at random and proposes a new persona using the frozen LLM\. If the generated persona lands in the currently selected cell, it undergoes a standard Metropolis–Hastings \(MH\) correction to preserve the within\-cell target law\. If the proposed persona lands in a previously empty, undiscovered cell, it serves as a hindsight seed that activates the new cell and initializes a new parallel chain\.

##### Theoretical Guarantee\.

By consolidating the cell\-kernel reversibility and eventual discovery properties, we establish three formal guarantees for the emitted sequence\.

###### Theorem 3\(Anytime Cell Uniformity and Target Convergence\)\.

Fix a deterministic validity ruleval\\operatorname\{val\}, a base generation lawbb, and a cell mapc:Ω→\[M\]c:\\Omega\\to\[M\]assigning each persona to one ofMMcells\. The encoder, whitening transform, and partition definingccare frozen\. Assumingb⁡\(p\)\>0b\(p\)\>0for every valid persona, we define the valid state space and base mass of celljjas

Ωj=\{p∈Ω:val\(p\)=1,c\(p\)=j\},Bj=b\(Ωj\)\.\\Omega\_\{j\}=\\\{p\\in\\Omega:\\operatorname\{val\}\(p\)=1,\\ c\(p\)=j\\\},\\qquad B\_\{j\}=b\(\\Omega\_\{j\}\)\.Letℛ=\{j:Bj\>0\}\\mathcal\{R\}=\\\{j:B\_\{j\}\>0\\\}be the set of reachable cells andπj=b\(⋅∣Ωj\)\\pi\_\{j\}=b\(\\,\\cdot\\mid\\Omega\_\{j\}\)their within\-cell target laws\. Suppose Algorithm[3](https://arxiv.org/html/2609.30492#alg3)is run from a fixed, nonempty active setℐ0⊆ℛ\\mathcal\{I\}\_\{0\}\\subseteq\\mathcal\{R\}\. With all proposal kernels and mixture weights frozen, the global proposal draws independently frombbwith a fixed probabilityω0\>0\\omega\_\{0\}\>0\. Letℐt−1⊆ℛ\\mathcal\{I\}\_\{t\-1\}\\subseteq\\mathcal\{R\}denote the active set after iterationtt,PtemitP\_\{t\}^\{\\mathrm\{emit\}\}the emitted persona, andℱt−1\\mathcal\{F\}\_\{t\-1\}the complete history prior to iterationtt\. The following properties hold\.

- ▶\\blacktrianglerightExact Anytime Cell Uniformity\.For everyt≥1t\\geq 1andj∈\[M\]j\\in\[M\], the spatial allocation of the emitted persona is strictly uniform over the currently active set Pr⁡\(c⁡\(Ptemit\)=j∣ℱt−1\)=𝟙\{j∈ℐt−1\}\|ℐt−1\|\.\\Pr\\\!\\left\(c\(P\_\{t\}^\{\\mathrm\{emit\}\}\)=j\\mid\\mathcal\{F\}\_\{t\-1\}\\right\)=\\frac\{\\mathbbm\{1\}\\\{j\\in\\mathcal\{I\}\_\{t\-1\}\\\}\}\{\|\\mathcal\{I\}\_\{t\-1\}\|\}\.
- ▶\\blacktrianglerightFinite\-Time Discovery\.For everyT≥1T\\geq 1, the probability that the active set has not yet expanded to cover the entire reachable set is bounded by Pr⁡\(ℐT≠ℛ\)≤∑j∈ℛ∖ℐ0\(1−ω0​Bj\)T\.\\Pr\(\\mathcal\{I\}\_\{T\}\\neq\\mathcal\{R\}\)\\leq\\sum\_\{j\\in\\mathcal\{R\}\\setminus\\mathcal\{I\}\_\{0\}\}\(1\-\\omega\_\{0\}B\_\{j\}\)^\{T\}\.Consequently, all reachable cells are discovered in finite time almost surely\.
- ▶\\blacktrianglerightAsymptotic Convergence\.Ast→∞t\\to\\infty, the law of the emitted persona converges in total variation to the equal\-cell mixture ‖ℒ⁡\(Ptemit\)−Πℛ‖TV⟶0,Πℛ=1\|ℛ\|​∑j∈ℛπj\.\\left\\\|\\mathcal\{L\}\(P\_\{t\}^\{\\mathrm\{emit\}\}\)\-\\Pi\_\{\\mathcal\{R\}\}\\right\\\|\_\{\\mathrm\{TV\}\}\\longrightarrow 0,\\qquad\\Pi\_\{\\mathcal\{R\}\}=\\frac\{1\}\{\|\\mathcal\{R\}\|\}\\sum\_\{j\\in\\mathcal\{R\}\}\\pi\_\{j\}\.

### 3\.4Evolutionary TextGrad Persona Generation

To actively push personas toward the extreme, unexplored regions of the semantic space, we formulate generation as an evolutionary constrained optimization problem\. Unlike UC\-MCMC’s distributional target, this is an empirical search that iteratively expands the population𝒫t\\mathcal\{P\}\_\{t\}toward valid, low\-density semantic frontiers\. We evaluate each persona across two geometric axes:*Isolation*\(gap to its nearest neighbor\) and*Sparsity*\(unnormalized von Mises\-Fisher kernel\-density score\)\.

Since personas are discrete text rather than continuous vectors, we employ*TextGrad*\([Yuksekgonul et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib58)\)to guide offspring generation with natural\-language feedback derived from the geometric fitness scores\. At each iteration, the algorithm executes three phases:

1. 1\.Parent Selection\.A tournament selects elite incumbents as parents that maximize isolation or minimize density\.
2. 2\.Mutation and Admission\.The frozen LLM mutates the parents using the textual gradient on the composite fitness function\. Then the offspring is admitted according to the same fitness function and thresholds after the validity checks\.
3. 3\.Sibling Deduplication\.Admitted offspring in the same iteration are mutually deduplicated to remove near\-duplicate siblings\.

PopulationFinal SubsetCategoryPrinciple\(method\)Cell\-Cov\.HullCov\.@r0⋆r\_\{0\}^\{\\star\}Disp\.ReferencePersonaMem\(Random\)62\.3%62\.3\\%4\.454\.450\.5890\.58913\.6713\.67SelectSpace filling\(Coverage\)62\.3%62\.3\\%4\.454\.450\.9010\.90110\.5210\.52Frontier\(Dispersion\)62\.3%62\.3\\%4\.454\.450\.0050\.00520\.9820\.98GenerateSpace filling\(MCMC\)96\.5%\\mathbf\{96\.5\\%\}6\.536\.530\.905\\mathbf\{0\.905\}7\.917\.91Frontier\(Evolution\)76\.8%76\.8\\%9\.47\\mathbf\{9\.47\}0\.0020\.00224\.42\\mathbf\{24\.42\}

Table 1:Persona geometry across persona diversification methods\.*Population*metrics evaluate the candidate pool: cell coverage, which is the occupancy of a fixed1,0241\{,\}024\-cell grid, and Mahalanobis hull extent\.*Subset*metrics evaluate the extracted personas used for response generation:*Cov\.@r0⋆r\_\{0\}^\{\\star\}*, which is the coverage objective \([eq\.7](https://arxiv.org/html/2609.30492#S3.E7)\) withr0⋆r\_\{0\}^\{\\star\}defined as[eq\.15](https://arxiv.org/html/2609.30492#A2.E15)\(c=0\.9c=0\.9of reference population𝒫0\\mathcal\{P\}\_\{0\}\), and dispersion, which is the minimum pairwise Mahalanobis distance\. Selection methods share the fixed base pool, while generation methods expand it\.
![Refer to caption](https://arxiv.org/html/2609.30492v1/persona_embedding_2d.png)
Figure 3:2D projection of persona embeddings illustrating evolutionary frontier generation\.The baseline pool \(grey\) is densely clustered in the center \(blue shading\), while evolved offspring \(orange\) populate low\-density regions\. The downstream selector extracts ak=5k=5subset \(large outlined dots\)\.

## 4Evaluation

We evaluate our framework on the Alternative Uses Task \(AUT\), Infinity\-Chat \(IC\), and the Divergent Association Task \(DAT\)\. Responses are generated by Gemma\-4\-31B\-it, with semantic embeddings provided by EmbeddingGemma and automated judgments by Qwen3\.6\-27B\. To evaluate transfer without task\-specific persona optimization, we hold each task\-agnostic persona set fixed across benchmarks\.

Original task\-only prompts are denoted asAlternative\-Useon AUT andBaseline Taskon IC and DAT\. For persona conditioning,Randomsamples uniformly from the PersonaMem\-v2\([Jiang et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib25)\)pool \(itself a random subset of PersonaHub\([Ge et al\., 2024](https://arxiv.org/html/2609.30492#bib.bib16)\)\), whileCoverageandDispersionoptimize selections from that pool\.MCMCandEvolutionexpand the candidate pool before applyingCoverageandDispersionselectors, respectively\. Prompt baselines includeCreativity\-enhanced\([Góes et al\., 2023](https://arxiv.org/html/2609.30492#bib.bib18)\)on AUT,Creative\([Schapiro et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib50)\)on DAT, and reasoning scaffold of DMAD \(Diverse Multi\-Agent Debate\)\([Liu et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib36)\)\. More prompt baselines are compared in Appendix[E](https://arxiv.org/html/2609.30492#A5), but three top\-performing baselines are reproduced in[Section4](https://arxiv.org/html/2609.30492#S4)\. Combined conditions add the selected personas with the corresponding prompt or reasoning scaffold\.

Before filtering for validity, each condition generates 25 uses per AUT object, 50 responses per IC query, and 35 ten\-noun lists for the DAT\. We assess response diversity and creativity alongside validity and utility to identify tradeoffs\. We employ a suite of metrics spanning human\-derived categorical clustering, LLM\-as\-a\-judge ratings, and semantic embedding metrics\. Full prompts, additional baselines and controls, metric definitions, and results are provided in Appendix[D](https://arxiv.org/html/2609.30492#A4)–[G](https://arxiv.org/html/2609.30492#A7)\.

##### RQ1: What persona geometries do the two diversity principles produce at the selection and generation levels?

As detailed in[Table1](https://arxiv.org/html/2609.30492#S3.T1), the framework effectively operationalizes its underlying principles, highlighting a fundamental tension between representation and separation\. To contextualize these structural gains, we benchmark againstRandomas a standard, geometrically blind persona\-conditioning pipeline\. At the selection level,Coveragecovers90\.1%90\.1\\%of the baseline population, whereasDispersionincreases minimum pairwise distance by53\.5%53\.5\\%overRandom\(13\.67→20\.9813\.67\\to 20\.98\)\. At the generation level, our methods break the fixed pool’s structural boundaries\.MCMCexpands the population’s cell coverage from62\.3%62\.3\\%to96\.5%96\.5\\%, and itsCoverage\-selected subset covers90\.5%90\.5\\%of the population atr⋆r^\{\\star\}\. In contrast,Evolutionmore than doubles the pool’s Mahalanobis hull extent from4\.454\.45to9\.479\.47\.[Figure3](https://arxiv.org/html/2609.30492#S3.F3)visualizes this expansion, where the baseline pool remains densely clustered in the center and the evolved offspring deliberately push outward to establish a novel, low\-density semantic boundary\. ApplyingDispersionto this expanded pool yields the highest minimum pairwise distance of24\.4224\.42, a78\.6%78\.6\\%increase overRandom\.

Pool expansion alone can increase the metrics in[Table1](https://arxiv.org/html/2609.30492#S3.T1)including, coverage, hull extend and dispersion\. RQ1 therefore emphasizes the contrastic geometries achieves by two principles:MCMCachives greater cell coverage, whereasEvolutionyields greater hull extent despite smaller pool thanMCMC\. These geometric gains do not guarantee improved response creativity\. Therefore, RQ2–RQ4 evaluate whether persona sets optimized without downstream task induce more diverse and creative responses\.

##### RQ2: How does persona\-set dispersion relate to downstream response diversity?

As shown in[Figure4](https://arxiv.org/html/2609.30492#S4.F4), across 14 configurations per benchmark, greater realized persona\-set separation positively correlates with greater mean response dispersion on AUT \(r=\.40r=\.40, descriptive 90% CI\[−\.19,\.78\]\[\-\.19,\.78\]\), and Infinity\-Chat \(r=\.55r=\.55,\[\.12,\.81\]\[\.12,\.81\]\)\. Despite the variance, intervention\-aligned comparisons reveal a robust directional trend: all six matchedCoverage→\\rightarrowDispersionescalations increased task\-averaged response dispersion across all 12 contrasts\. Additional analyses using Vendi score and convex\-hull, showing even stronger correlations between persona\-set dispersion and response diversity, appear in Appendix[E\.2](https://arxiv.org/html/2609.30492#A5.SS2)\.

Figure 4:Persona set and response dispersion on AUT \(left\) and Infinity\-Chat \(right\)\.Every points represent different configurations\. Arrows connect six matchedCoverage→\\rightarrowDispersioncontrasts while holding candidate\-pool and distance metric fixed\. Dashed lines show the linear fits and annotations report configuration\-level Pearson correlations and 90% confidence intervals\.
##### RQ3a: Do diversified task\-agnostic personas improve meaningful creativity, and at what cost?

Yes\. Task\-agnostic diversification yields substantial creativity gains with minimal trade\-offs\. As shown in the top half of[Table2](https://arxiv.org/html/2609.30492#S4.T2),Randompersona conditioning improves diversity by8\.7%8\.7\\%\(1\.42→1\.541\.42\\to 1\.54\), originality by10\.2%10\.2\\%\(10\.79→11\.8910\.79\\to 11\.89\) and creativity by8\.9%8\.9\\%\(OPEN2\.72→2\.96\)2\.72\\to 2\.96\)while sacrificing virtually no utility or validity\. OurDispersionselector further improves diversity, originality and flexibility, while our evolutionarily generated personas achieve the highest values across all four target metrics among the standard\-prompt conditions\. Relative toAlternative\-Use,Evolutionimproves diversity by78\.8%78\.8\\%\(1\.42→2\.531\.42\\to 2\.53\), originality by26\.1%26\.1\\%\(10\.79→13\.6110\.79\\to 13\.61\), flexibility by49\.5%49\.5\\%\(\.288→\.431\.288\\to\.431\) and creativity score by13\.9%13\.9\\%\(2\.72→3\.102\.72\\to 3\.10\)\. Compared toRandom,Evolutionyields relative improvements ranging from4\.7%4\.7\\%in creativity up to64\.3%\\mathbf\{64\.3\\%\}in diversity\. The cost of these leaps is modest: utility is reduced by just0\.220\.22\(−5\.1%\-5\.1\\%\) from the original task \(Alternative\-Use\), while still producing valid responses98\.5%98\.5\\%of the time\.

##### RQ3b: Do the creativity gains compose with prompt engineering?

Yes\. Persona diversification composes with, rather than competes against, advanced prompt engineering\. As demonstrated in the bottom half of[Table2](https://arxiv.org/html/2609.30492#S4.T2), applying our evolutionarily generated personas to the creativity\-optimized prompt of[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18)amplifies its effects, yielding improvements of\+18\.6%\+18\.6\\%in diversity,\+2\.9%\+2\.9\\%in originality,\+15\.3%\+15\.3\\%in flexibility, and\+6\.3%\+6\.3\\%in creativity, all while maintaining utility \(3\.383\.38\) and improving validity \(\.968→\.979\.968\\to\.979\)\. It also shows competitive results against advanced reasoning frameworks, includingDMAD\([Liu et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib36)\), a state\-of\-the\-art baseline that diversifies reasoning paths rather than personas\. Because the personas are generated without optimizing downstream task performance, they are particularly valuable for everyday, zero\-shot user queries where elaborate prompt engineering is impractical\.

Table 2:Response creativity evaluation on AUT\.*Diversity*is the hull extent in a 5\-D PCA projection of Mahalanobis\-whitened response embeddings,*Originality*is the mean Mahalanobis distance to the common\-use reference set,*Flexibility*is the number of distinct categories per valid use, and*Creativity*is the rank\-normalized holistic score\([Góes et al\., 2023](https://arxiv.org/html/2609.30492#bib.bib18)\)from an independent LLM judge\. Subscripts give 95% confidence\-interval half\-widths across five random seeds\. Utility\([Stevenson et al\., 2022](https://arxiv.org/html/2609.30492#bib.bib53)\)and validity are*controls*and are not bolded\.Bold= best results within each group\.VariantDiversityOriginalityFlexibilityCreativityUtilityValidityAlternative\-Use1\.42±0\.091\.42\_\{\\pm 0\.09\}10\.79±0\.3210\.79\_\{\\pm 0\.32\}\.288±\.022\.288\_\{\\pm\.022\}2\.72±0\.092\.72\_\{\\pm 0\.09\}4\.30±0\.054\.30\_\{\\pm 0\.05\}\.984±\.014\.984\_\{\\pm\.014\}Random1\.54±0\.121\.54\_\{\\pm 0\.12\}11\.89±0\.4811\.89\_\{\\pm 0\.48\}\.331±\.024\.331\_\{\\pm\.024\}2\.96±0\.042\.96\_\{\\pm 0\.04\}4\.28±0\.024\.28\_\{\\pm 0\.02\}\.993±\.006\.993\_\{\\pm\.006\}Dispersion\(selection\)1\.96±0\.171\.96\_\{\\pm 0\.17\}12\.00±0\.2812\.00\_\{\\pm 0\.28\}\.357±\.048\.357\_\{\\pm\.048\}2\.92±0\.042\.92\_\{\\pm 0\.04\}4\.23±0\.094\.23\_\{\\pm 0\.09\}\.986±\.010\.986\_\{\\pm\.010\}Evolution\(generation\)2\.53±0\.20\\mathbf\{2\.53\}\_\{\\pm 0\.20\}13\.61±0\.17\\mathbf\{13\.61\}\_\{\\pm 0\.17\}\.431±\.045\\mathbf\{\.431\}\_\{\\pm\.045\}3\.10±0\.03\\mathbf\{3\.10\}\_\{\\pm 0\.03\}4\.08±0\.064\.08\_\{\\pm 0\.06\}\.985±\.012\.985\_\{\\pm\.012\}Creativity\-enhanced5\.15±0\.165\.15\_\{\\pm 0\.16\}17\.15±0\.4017\.15\_\{\\pm 0\.40\}\.503±\.056\.503\_\{\\pm\.056\}3\.19±0\.093\.19\_\{\\pm 0\.09\}3\.38±0\.073\.38\_\{\\pm 0\.07\}\.968±\.011\.968\_\{\\pm\.011\}DMAD2\.79±0\.222\.79\_\{\\pm 0\.22\}13\.90±0\.4913\.90\_\{\\pm 0\.49\}\.372±\.025\.372\_\{\\pm\.025\}2\.89±0\.072\.89\_\{\\pm 0\.07\}4\.17±0\.114\.17\_\{\\pm 0\.11\}\.977±\.009\.977\_\{\\pm\.009\}Evolution \+ creativity6\.11±0\.22\\mathbf\{6\.11\}\_\{\\pm 0\.22\}17\.65±0\.14\\mathbf\{17\.65\}\_\{\\pm 0\.14\}\.580±\.046\\mathbf\{\.580\}\_\{\\pm\.046\}3\.39±0\.07\\mathbf\{3\.39\}\_\{\\pm 0\.07\}3\.38±0\.163\.38\_\{\\pm 0\.16\}\.979±\.020\.979\_\{\\pm\.020\}

##### RQ4: Do the task\-agnostic gains transfer across tasks and compose with prompt engineering?

Yes\. Task\-agnostic persona diversification transfers robustly across both open\-ended and constrained task formats\. On Infinity\-Chat,Dispersionimproves uponRandomin all three diversity measures, whileEvolutionextends these gains further \([Table3](https://arxiv.org/html/2609.30492#S4.T3)\)\. Relative to standard task prompt,Evolutionreduces homogeneity \(0\.946→0\.7980\.946\\to 0\.798\) and more than doubles flexibility \(1\.39→2\.811\.39\\to 2\.81\)\. Its between\-persona separation gain also rises from0\.0680\.068ofRandomto0\.1260\.126, indicating significantly more distinct responses across personas relative to variation within a single persona\. Similarly, DAT exhibits the same ordering of escalation across the task\-onlyBaseline Task,Random,DispersionandEvolutionconditions for both divergence and flexibility\. On DAT, task\-agnosticEvolutionalso exceedsCreativeon these measures\. Crucially, despite improved creativity,Evolutionmaintains near\-perfect validity on both benchmarks \(99\.2%99\.2\\%on Infinity\-Chat and99\.6%99\.6\\%on DAT\), proving that task adherence is not sacrificed\. Furthermore, these task\-agnostic gains compose seamlessly with advanced reasoning scaffolds\. AddingEvolutionpersonas toDMADfurther reduces Infinity\-Chat homogeneity \(0\.929→0\.9000\.929\\to 0\.900\) and raises both Infinity\-Chat flexibility \(1\.56→1\.811\.56\\to 1\.81\) and DAT score \(90\.62→90\.9390\.62\\to 90\.93\), while maintaining validity at100%100\\%and99\.7%99\.7\\%, respectively, demonstrating the robust compositional benefits\.

Table 3:Response creativity evaluation on Infinity\-Chat \(IC\) and DAT\.*Homogeneity*is mean pairwise cosine similarity,*Separation*is mean between\-persona distance minus mean within\-persona distance,*Flexibility*is the Vendi score, measuring effective diversity under the embedding similarity kernel, and*Divergence*is the DAT semantic score\. Subscripts give 95% confidence\-interval half\-widths across five random seeds\.Bold= best results within each group\. Control \(validity\) and undefined combinations \(–\) are not bolded\.Infinity\-ChatDATVariantHomog\.↓\\downarrowSep\.↑\\uparrowFlex\.↑\\uparrowVal\.Div\.↑\\uparrowFlex\.↑\\uparrowVal\.Baseline Task0\.946±0\.0020\.946\_\{\\pm 0\.002\}–1\.39±0\.021\.39\_\{\\pm 0\.02\}\.950±\.015\.950\_\{\\pm\.015\}89\.13±0\.4989\.13\_\{\\pm 0\.49\}1\.79±0\.011\.79\_\{\\pm 0\.01\}\.996±\.005\.996\_\{\\pm\.005\}Random0\.864±0\.0050\.864\_\{\\pm 0\.005\}0\.068±0\.0030\.068\_\{\\pm 0\.003\}2\.17±0\.052\.17\_\{\\pm 0\.05\}1\.000±\.0001\.000\_\{\\pm\.000\}89\.15±1\.2489\.15\_\{\\pm 1\.24\}1\.83±0\.041\.83\_\{\\pm 0\.04\}\.993±\.011\.993\_\{\\pm\.011\}Dispersion\(selection\)0\.840±0\.0040\.840\_\{\\pm 0\.004\}0\.085±0\.0060\.085\_\{\\pm 0\.006\}2\.40±0\.052\.40\_\{\\pm 0\.05\}\.986±\.007\.986\_\{\\pm\.007\}89\.39±0\.2689\.39\_\{\\pm 0\.26\}1\.84±0\.011\.84\_\{\\pm 0\.01\}\.995±\.006\.995\_\{\\pm\.006\}Evolution\(generation\)0\.798±0\.006\\mathbf\{0\.798\}\_\{\\pm 0\.006\}0\.126±0\.006\\mathbf\{0\.126\}\_\{\\pm 0\.006\}2\.81±0\.08\\mathbf\{2\.81\}\_\{\\pm 0\.08\}\.992±\.005\.992\_\{\\pm\.005\}89\.48±0\.32\\mathbf\{89\.48\}\_\{\\pm 0\.32\}1\.85±0\.01\\mathbf\{1\.85\}\_\{\\pm 0\.01\}\.996±\.004\.996\_\{\\pm\.004\}Creative––––86\.05±0\.4386\.05\_\{\\pm 0\.43\}1\.69±0\.011\.69\_\{\\pm 0\.01\}\.977±\.010\.977\_\{\\pm\.010\}DMAD0\.929±0\.0020\.929\_\{\\pm 0\.002\}–1\.56±0\.011\.56\_\{\\pm 0\.01\}1\.000±\.0001\.000\_\{\\pm\.000\}90\.62±0\.2590\.62\_\{\\pm 0\.25\}1\.89±0\.011\.89\_\{\\pm 0\.01\}\.994±\.006\.994\_\{\\pm\.006\}Evolution \+ DMAD0\.900±0\.001\\mathbf\{0\.900\}\_\{\\pm 0\.001\}0\.021±0\.0030\.021\_\{\\pm 0\.003\}1\.81±0\.01\\mathbf\{1\.81\}\_\{\\pm 0\.01\}1\.000±\.0001\.000\_\{\\pm\.000\}90\.93±0\.29\\mathbf\{90\.93\}\_\{\\pm 0\.29\}1\.89±0\.011\.89\_\{\\pm 0\.01\}\.997±\.008\.997\_\{\\pm\.008\}

## 5Conclusion

This work framed persona\-diversity induction as a set\-level design problem, distinguishing space\-filling from frontier\-seeking diversity and selection from generation\. The resulting framework provides principled ways to control persona\-set geometry and shows that representational diversity and ideational diversity are not the same objective: the persona set that best represents a population need not be the one that elicits the widest range of ideas\. More broadly, persona geometry offers a task\-agnostic mechanism for shaping the range of LLM outputs, complementary to prompt optimization and inference\-time scaffolding, and points toward systems that can deliberately control not only the quality of an answer, but the diversity of perspectives from which answers are generated\.

### AI use statement

In this work, we used generative AI tools to help develop theoretical models or conceptual frameworks, formulate mathematical claims, provide critical ingredients for proving mathematical claims, assist in the writing of proofs, design or provide feedback on research methodology or experiments, implement methods, support qualitative and thematic data analysis, and interpret results\. We have not used generative AI tools to propose or refine hypotheses and generate synthetic data sets, assist with translation, and clean and reformat dataset are not applicable to this work\. Additionally, we used generative AI tools to create or modify scientific figures or images, create or edit software code, creation of artifacts, summarize or analyse existing literature, discover research topics or identify gaps, brainstorming, sourcing/searching for information, edit a research paper to improve readability, identify relevant literature, format references, and propose a title or keywords for a research paper\. We have reviewed all AI\-assisted work\. We reviewed and tested AI\-generated experiment code to run as intended, and all conversations with AI are verified before applied to the paper\. We take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI\.

### Reproducibility Statement

Sections[2](https://arxiv.org/html/2609.30492#S2)and[3](https://arxiv.org/html/2609.30492#S3)define the persona representations, diversification objectives, and algorithms\. Appendices[B](https://arxiv.org/html/2609.30492#A2)and[C](https://arxiv.org/html/2609.30492#A3)provide detailed algorithmic procedures, the assumptions underlying the theoretical guarantees, and their proofs\. Appendix[D](https://arxiv.org/html/2609.30492#A4)documents the benchmark data sources, baseline construction, generation and preprocessing protocols, evaluation metrics, annotation procedures, and LLM\-judge validation\. Appendix[E](https://arxiv.org/html/2609.30492#A5)provides the full experimental comparisons, with performance estimates and confidence intervals across five generation seeds\. Extensive examples of representative personas and prompt templates for persona construction, response generation, and evaluation are provided in Appendices[F](https://arxiv.org/html/2609.30492#A6)and[G](https://arxiv.org/html/2609.30492#A7), respectively\.

## References

- Anthis et al\. \(2025\)Jacy Reese Anthis, Ryan Liu, Sean M\. Richardson, Austin C\. Kozlowski, Bernard Koch, James Evans, Erik Brynjolfsson, and Michael Bernstein\.Llm social simulations are a promising research method, 2025\.URL[https://arxiv\.org/abs/2504\.02234](https://arxiv.org/abs/2504.02234)\.
- Atmakuru et al\. \(2024\)Anirudh Atmakuru, Jatin Nainani, Rohith Siddhartha Reddy Bheemreddy, Anirudh Lakkaraju, Zonghai Yao, Hamed Zamani, and Haw\-Shiuan Chang\.Cs4: Measuring the creativity of large language models automatically by controlling the number of story\-writing constraints, 2024\.URL[https://arxiv\.org/abs/2410\.04197](https://arxiv.org/abs/2410.04197)\.
- Azuma \(1967\)Kazuoki Azuma\.Weighted sums of certain dependent random variables\.*Tohoku Mathematical Journal*, 19\(3\):357–367, 1967\.doi:10\.2748/tmj/1178243286\.URL[https://doi\.org/10\.2748/tmj/1178243286](https://doi.org/10.2748/tmj/1178243286)\.
- BIS \(2024\)BIS\.Artificial intelligence and the economy: Implications for central banks\.Technical report, Bank for International Settlements, June 2024\.URL[https://www\.bis\.org/publ/arpdf/ar2024e3\.htm](https://www.bis.org/publ/arpdf/ar2024e3.htm)\.Chapter III of the BIS Annual Economic Report 2024\.
- Chakrabarty et al\. \(2024\)Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, and Chien\-Sheng Wu\.Art or artifice? large language models and the false promise of creativity\.In*Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems*, pp\. 1–34, 2024\.URL[https://arxiv\.org/abs/2309\.14556](https://arxiv.org/abs/2309.14556)\.
- Charikar et al\. \(2001\)Moses Charikar, Samir Khuller, David M\. Mount, and Giri Narasimhan\.Algorithms for facility location problems with outliers\.In*Proceedings of the Twelfth Annual ACM\-SIAM Symposium on Discrete Algorithms \(SODA\)*, pp\. 642–651, 2001\.URL[https://dl\.acm\.org/citation\.cfm?id=365411\.365555](https://dl.acm.org/citation.cfm?id=365411.365555)\.
- Chen & Ding \(2023\)Honghua Chen and Nai Ding\.Probing the “creativity” of large language models: can models produce divergent semantic association?In*Findings of the Association for Computational Linguistics: EMNLP 2023*, pp\. 12881–12888, 2023\.URL[https://arxiv\.org/abs/2310\.11158](https://arxiv.org/abs/2310.11158)\.
- Choi & Li \(2024\)Hyeong Kyu Choi and Yixuan Li\.PICLe: Eliciting diverse behaviors from large language models with persona in\-context learning\.In*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pp\. 8722–8739\. PMLR, 2024\.URL[https://proceedings\.mlr\.press/v235/choi24e\.html](https://proceedings.mlr.press/v235/choi24e.html)\.
- Dathathri et al\. \(2020\)Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu\.Plug and play language models: A simple approach to controlled text generation, 2020\.URL[https://arxiv\.org/abs/1912\.02164](https://arxiv.org/abs/1912.02164)\.
- Della Croce et al\. \(2009\)Federico Della Croce, Andrea Grosso, and Marco Locatelli\.A heuristic approach for the max–min diversity problem based on max\-clique\.*Computers & Operations Research*, 36\(8\):2429–2433, 2009\.doi:10\.1016/j\.cor\.2008\.09\.007\.URL[https://doi\.org/10\.1016/j\.cor\.2008\.09\.007](https://doi.org/10.1016/j.cor.2008.09.007)\.
- Erkut \(1990\)Erhan Erkut\.The discrete p\-dispersion problem\.*European Journal of Operational Research*, 46\(1\):48–60, 1990\.URL[https://doi\.org/10\.1016/0377\-2217\(90\)90297\-O](https://doi.org/10.1016/0377-2217(90)90297-O)\.
- Feige \(1998\)Uriel Feige\.A threshold of ln n for approximating set cover\.*Journal of the ACM \(JACM\)*, 45\(4\):634–652, 1998\.URL[https://doi\.org/10\.1145/285055\.285059](https://doi.org/10.1145/285055.285059)\.
- Fernando et al\. \(2023\)Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel\.Promptbreeder: Self\-referential self\-improvement via prompt evolution, 2023\.URL[https://arxiv\.org/abs/2309\.16797](https://arxiv.org/abs/2309.16797)\.
- Fillmore \(1968\)Charles J\. Fillmore\.The Case for Case\.In Emmon Bach and Robert T\. Harms \(eds\.\),*Universals in Linguistic Theory*, pp\. 1–88\. Holt, Rinehart and Winston, New York, NY, 1968\.URL[https://verbs\.colorado\.edu/~mpalmer/Ling7800/Fillmore\.Case\.pdf](https://verbs.colorado.edu/~mpalmer/Ling7800/Fillmore.Case.pdf)\.
- Friedman & Dieng \(2023\)Dan Friedman and Adji Bousso Dieng\.The Vendi Score: A diversity evaluation metric for machine learning\.*Transactions on Machine Learning Research*, 2023\.URL[https://openreview\.net/forum?id=g97OHbQyk1](https://openreview.net/forum?id=g97OHbQyk1)\.
- Ge et al\. \(2024\)Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu\.Scaling synthetic data creation with 1,000,000,000 personas, 2024\.URL[https://arxiv\.org/abs/2406\.20094](https://arxiv.org/abs/2406.20094)\.
- Gibson \(1979\)James J\. Gibson\.*The Ecological Approach to Visual Perception*\.Houghton Mifflin, Boston, MA, 1979\.URL[https://openlibrary\.org/books/OL21229791M/The\_Ecological\_approach\_to\_visual\_perception](https://openlibrary.org/books/OL21229791M/The_Ecological_approach_to_visual_perception)\.
- Góes et al\. \(2023\)Luis Fabricio Góes, Marco Volpe, Piotr Sawicki, Marek Grses, and Jacob Watson\.Pushing gpt’s creativity to its limits: Alternative uses and torrance tests\.*International Conference on Computational Creativity*, 2023\.URL[https://computationalcreativity\.net/iccc23/papers/ICCC\-2023\_paper\_90\.pdf](https://computationalcreativity.net/iccc23/papers/ICCC-2023_paper_90.pdf)\.
- Guilford \(1967\)Joy Paul Guilford\.*The Nature of Human Intelligence*\.McGraw\-Hill, New York, 1967\.URL[https://search\.worldcat\.org/title/The\-nature\-of\-human\-intelligence/oclc/204270](https://search.worldcat.org/title/The-nature-of-human-intelligence/oclc/204270)\.
- Hastings \(1970\)W\. K\. Hastings\.Monte carlo sampling methods using markov chains and their applications\.*Biometrika*, 57\(1\):97–109, 1970\.doi:10\.1093/biomet/57\.1\.97\.URL[https://doi\.org/10\.1093/biomet/57\.1\.97](https://doi.org/10.1093/biomet/57.1.97)\.
- Hochbaum & Shmoys \(1986\)Dorit S\. Hochbaum and David B\. Shmoys\.A unified approach to approximation algorithms for bottleneck problems\.*Journal of the ACM*, 33\(3\):533–550, 1986\.doi:10\.1145/5925\.5933\.URL[https://doi\.org/10\.1145/5925\.5933](https://doi.org/10.1145/5925.5933)\.
- Hou et al\. \(2026\)Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw\-Shiuan Chang, Daniel Khashabi, and Xiang Lorraine Li\.Creativityprism: A holistic evaluation framework for large language model creativity, 2026\.URL[https://arxiv\.org/abs/2510\.20091](https://arxiv.org/abs/2510.20091)\.
- Howard et al\. \(2021\)Steven R\. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon\.Time\-uniform, nonparametric, nonasymptotic confidence sequences\.*The Annals of Statistics*, 49\(2\):1055–1080, 2021\.doi:10\.1214/20\-AOS1991\.URL[https://doi\.org/10\.1214/20\-AOS1991](https://doi.org/10.1214/20-AOS1991)\.
- IMF \(2024\)IMF\.Advances in artificial intelligence: Implications for capital market activities\.Technical report, International Monetary Fund, October 2024\.URL[https://www\.imf\.org/en/Publications/GFSR/Issues/2024/10/22/global\-financial\-stability\-report\-october\-2024](https://www.imf.org/en/Publications/GFSR/Issues/2024/10/22/global-financial-stability-report-october-2024)\.Chapter 3 of the Global Financial Stability Report: Steadying the Course—Uncertainty, Artificial Intelligence, and Financial Stability\.
- Jiang et al\. \(2025\)Bowen Jiang, Zhuoqun Hao, Young\-Min Cho, Bryan Li, Yuan Yuan, Sihao Chen, Lyle Ungar, Camillo J\. Taylor, and Dan Roth\.Know me, respond to me: Benchmarking llms for dynamic user profiling and personalized responses at scale, 2025\.URL[https://arxiv\.org/abs/2504\.14225](https://arxiv.org/abs/2504.14225)\.
- Jiang et al\. \(2026\)Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yulia Tsvetkov, Maarten Sap, and Yejin Choi\.Artificial hivemind: The open\-ended homogeneity of language models \(and beyond\)\.*Advances in Neural Information Processing Systems*, 38, 2026\.URL[https://arxiv\.org/abs/2510\.22954](https://arxiv.org/abs/2510.22954)\.
- Jin et al\. \(2025\)Bojun Jin, Jianzhu Bao, Yufang Hou, Yang Sun, Yice Zhang, Huajie Wang, Bin Liang, and Ruifeng Xu\.A multi\-persona framework for argument quality assessment\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pp\. 12148–12170, 2025\.URL[https://aclanthology\.org/2025\.acl\-long\.593/](https://aclanthology.org/2025.acl-long.593/)\.
- Johnson et al\. \(2013\)Alicia A\. Johnson, Galin L\. Jones, and Ronald C\. Neath\.Component\-wise markov chain monte carlo: Uniform and geometric ergodicity under mixing and composition\.*Statistical Science*, 28\(3\):360–375, 2013\.doi:10\.1214/13\-STS423\.URL[https://doi\.org/10\.1214/13\-STS423](https://doi.org/10.1214/13-STS423)\.
- Kirk et al\. \(2024\)Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu\.Understanding the effects of rlhf on llm generalisation and diversity, 2024\.URL[https://arxiv\.org/abs/2310\.06452](https://arxiv.org/abs/2310.06452)\.
- Kojima et al\. \(2022\)Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa\.Large language models are zero\-shot reasoners\.*Advances in neural information processing systems*, 35:22199–22213, 2022\.URL[https://proceedings\.neurips\.cc/paper/2022/hash/8bb0d291acd4acf06ef112099c16f326\-Abstract\-Conference\.html](https://proceedings.neurips.cc/paper/2022/hash/8bb0d291acd4acf06ef112099c16f326-Abstract-Conference.html)\.
- Kulesza & Taskar \(2012\)Alex Kulesza and Ben Taskar\.Learning determinantal point processes, 2012\.URL[https://arxiv\.org/abs/1202\.3738](https://arxiv.org/abs/1202.3738)\.
- Kusupati et al\. \(2022\)Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard\-Snyder, Kaifeng Chen, Sham Kakade, Prateek Jain, et al\.Matryoshka representation learning\.*Advances in Neural Information Processing Systems*, 35:30233–30249, 2022\.URL[https://proceedings\.nips\.cc/paper\_files/paper/2022/hash/c32319f4868da7613d78af9993100e42\-Abstract\-Conference\.html](https://proceedings.nips.cc/paper_files/paper/2022/hash/c32319f4868da7613d78af9993100e42-Abstract-Conference.html)\.
- Leopardi \(2006\)Paul Leopardi\.A partition of the unit sphere into regions of equal area and small diameter\.*Electronic Transactions on Numerical Analysis*, 25:309–327, 2006\.URL[https://eudml\.org/doc/129860](https://eudml.org/doc/129860)\.
- Levin \(1993\)Beth Levin\.*English Verb Classes and Alternations: A Preliminary Investigation*\.University of Chicago Press, Chicago, IL, 1993\.URL[https://press\.uchicago\.edu/ucp/books/book/chicago/E/bo3684144\.html](https://press.uchicago.edu/ucp/books/book/chicago/E/bo3684144.html)\.
- Li et al\. \(2026\)Leon Li, Haozhe Chen, Hongseok Namkoong, and Tianyi Peng\.Llm generated persona is a promise with a catch\.*Advances in Neural Information Processing Systems*, 38, 2026\.URL[https://arxiv\.org/abs/2503\.16527](https://arxiv.org/abs/2503.16527)\.
- Liu et al\. \(2025\)Yexiang Liu, Jie Cao, Zekun Li, Ran He, and Tieniu Tan\.Breaking mental set to improve reasoning through diverse multi\-agent debate\.In*The Thirteenth International Conference on Learning Representations*, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/hash/3de667dab3b3d812583abc0a786139a0\-Abstract\-Conference\.html](https://proceedings.iclr.cc/paper_files/paper/2025/hash/3de667dab3b3d812583abc0a786139a0-Abstract-Conference.html)\.
- Lu et al\. \(2025\)Ximing Lu, Melanie Sclar, Skyler Hallinan, Niloofar Mireshghallah, Jiacheng Liu, Seungju Han, Allyson Ettinger, Liwei Jiang, Khyathi Chandu, Nouha Dziri, et al\.Ai as humanity’s salieri: Quantifying linguistic creativity of language models via systematic attribution of machine text against web text\.In*International Conference on Learning Representations*, pp\. 91005–91056, 2025\.URL[https://arxiv\.org/abs/2410\.04265](https://arxiv.org/abs/2410.04265)\.
- Mengersen & Tweedie \(1996\)Kerrie L\. Mengersen and Richard L\. Tweedie\.Rates of convergence of the hastings and metropolis algorithms\.*The Annals of Statistics*, 24\(1\):101–121, 1996\.doi:10\.1214/aos/1033066201\.URL[https://doi\.org/10\.1214/aos/1033066201](https://doi.org/10.1214/aos/1033066201)\.
- Mouret & Clune \(2015\)Jean\-Baptiste Mouret and Jeff Clune\.Illuminating search spaces by mapping elites, 2015\.URL[https://arxiv\.org/abs/1504\.04909](https://arxiv.org/abs/1504.04909)\.
- Nemhauser et al\. \(1978\)George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher\.An analysis of approximations for maximizing submodular set functions—i\.*Mathematical programming*, 14\(1\):265–294, 1978\.ISSN 1436\-4646\.doi:10\.1007/BF01588971\.URL[https://doi\.org/10\.1007/BF01588971](https://doi.org/10.1007/BF01588971)\.
- Nguyen & Dieng \(2024\)Quan Nguyen and Adji Bousso Dieng\.Quality\-weighted vendi scores and their application to diverse experimental design\.In*Proceedings of the 41st International Conference on Machine Learning*, volume 235 of*Proceedings of Machine Learning Research*, pp\. 37667–37682\. PMLR, 2024\.URL[https://proceedings\.mlr\.press/v235/nguyen24d\.html](https://proceedings.mlr.press/v235/nguyen24d.html)\.
- Norman \(1988\)Donald A\. Norman\.*The Psychology of Everyday Things*\.Basic Books, New York, NY, 1988\.URL[https://openlibrary\.org/books/OL2409159M/The\_psychology\_of\_everyday\_things](https://openlibrary.org/books/OL2409159M/The_psychology_of_everyday_things)\.
- Olson et al\. \(2021\)Jay A Olson, Johnny Nahas, Denis Chmoulevitch, Simon J Cropper, and Margaret E Webb\.Naming unrelated words predicts creativity\.*Proceedings of the National Academy of Sciences*, 118\(25\):e2022340118, 2021\.URL[https://www\.pnas\.org/doi/10\.1073/pnas\.2022340118](https://www.pnas.org/doi/10.1073/pnas.2022340118)\.
- Paglieri et al\. \(2026\)Davide Paglieri, Logan Cross, William A\. Cunningham, Joel Z\. Leibo, and Alexander Sasha Vezhnevets\.Persona generators: Generating diverse synthetic personas for arbitrary contexts, 2026\.URL[https://arxiv\.org/abs/2602\.03545](https://arxiv.org/abs/2602.03545)\.
- Pasarkar & Dieng \(2024\)Amey P Pasarkar and Adji Bousso Dieng\.Cousins of the Vendi Score: A family of similarity\-based diversity metrics for science and machine learning\.In*International Conference on Artificial Intelligence and Statistics \(AISTATS\)*, 2024\.URL[https://proceedings\.mlr\.press/v238/pasarkar24a\.html](https://proceedings.mlr.press/v238/pasarkar24a.html)\.
- Rabeyah et al\. \(2025\)Abdullah Al Rabeyah, Fabrício Góes, Marco Volpe, and Talles Medeiros\.Do llms agree on the creativity evaluation of alternative uses?*International Conference on Computational Creativity*, 2025\.URL[https://computationalcreativity\.net/iccc25/papers/iccc25\-rabeyah2025do\.pdf](https://computationalcreativity.net/iccc25/papers/iccc25-rabeyah2025do.pdf)\.
- Roberts & Rosenthal \(2007\)Gareth O\. Roberts and Jeffrey S\. Rosenthal\.Coupling and ergodicity of adaptive Markov chain Monte Carlo algorithms\.*Journal of Applied Probability*, 44\(2\):458–475, 2007\.doi:10\.1239/jap/1183667414\.URL[https://doi\.org/10\.1239/jap/1183667414](https://doi.org/10.1239/jap/1183667414)\.
- Rosenthal \(1995\)Jeffrey S\. Rosenthal\.Minorization conditions and convergence rates for markov chain monte carlo\.*Journal of the American Statistical Association*, 90\(430\):558–566, 1995\.doi:10\.1080/01621459\.1995\.10476548\.URL[https://doi\.org/10\.1080/01621459\.1995\.10476548](https://doi.org/10.1080/01621459.1995.10476548)\.
- Runco & Jaeger \(2012\)Mark A Runco and Garrett J Jaeger\.The standard definition of creativity\.*Creativity Research Journal*, 24\(1\):92–96, 2012\.URL[https://doi\.org/10\.1080/10400419\.2012\.650092](https://doi.org/10.1080/10400419.2012.650092)\.
- Schapiro et al\. \(2026\)Samuel Schapiro, Core Francisco Park, Felix Sosa, and Lav R Varshney\.Creativityneuro: Steering language model weights to improve divergent thinking and reduce mode collapse\.*arXiv preprint arXiv:2607\.01433*, 2026\.URL[https://arxiv\.org/abs/2607\.01433](https://arxiv.org/abs/2607.01433)\.
- Sorensen et al\. \(2024\)Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi\.A roadmap to pluralistic alignment, 2024\.URL[https://arxiv\.org/abs/2402\.05070](https://arxiv.org/abs/2402.05070)\.
- Sorensen et al\. \(2026\)Taylor Sorensen, Benjamin Newman, Jared Moore, Chan Park, Jillian Fisher, Niloofar Mireshghallah, Liwei Jiang, and Yejin Choi\.Spectrum tuning: Post\-training for distributional coverage and in\-context steerability, 2026\.URL[https://arxiv\.org/abs/2510\.06084](https://arxiv.org/abs/2510.06084)\.
- Stevenson et al\. \(2022\)Claire Stevenson, Iris Smal, Matthijs Baas, Raoul Grasman, and Han van der Maas\.Putting gpt\-3’s creativity to the \(alternative uses\) test\.*International Conference on Computational Creativity*, 2022\.URL[https://computationalcreativity\.net/iccc22/wp\-content/uploads/2022/06/ICCC\-2022\_25S\_Stevenson\-et\-al\.\.pdf](https://computationalcreativity.net/iccc22/wp-content/uploads/2022/06/ICCC-2022_25S_Stevenson-et-al..pdf)\.
- Su & Collier \(2023\)Yixuan Su and Nigel Collier\.Contrastive search is what you need for neural text generation, 2023\.URL[https://arxiv\.org/abs/2210\.14140](https://arxiv.org/abs/2210.14140)\.
- Tierney \(1994\)Luke Tierney\.Markov chains for exploring posterior distributions\.*The Annals of Statistics*, 22\(4\):1701–1762, 1994\.doi:10\.1214/aos/1176325750\.URL[https://doi\.org/10\.1214/aos/1176325750](https://doi.org/10.1214/aos/1176325750)\.
- Wei et al\. \(2015\)Kai Wei, Rishabh Iyer, and Jeff Bilmes\.Submodularity in data subset selection and active learning\.In*Proceedings of the 32nd International Conference on Machine Learning*, volume 37 of*Proceedings of Machine Learning Research*, pp\. 1954–1963\. PMLR, 2015\.URL[https://proceedings\.mlr\.press/v37/wei15\.html](https://proceedings.mlr.press/v37/wei15.html)\.
- West & Potts \(2025\)Peter West and Christopher Potts\.Base models beat aligned models at randomness and creativity, 2025\.URL[https://arxiv\.org/abs/2505\.00047](https://arxiv.org/abs/2505.00047)\.
- Yuksekgonul et al\. \(2025\)Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Pan Lu, Zhi Huang, Carlos Guestrin, and James Zou\.Optimizing generative ai by backpropagating language model feedback\.*Nature*, 639:609–616, 2025\.doi:10\.1038/s41586\-025\-08661\-4\.URL[https://doi\.org/10\.1038/s41586\-025\-08661\-4](https://doi.org/10.1038/s41586-025-08661-4)\.
- Zheng et al\. \(2024\)Huaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng\-Tze Cheng, Ed H\. Chi, Quoc V Le, and Denny Zhou\.Take a step back: Evoking reasoning via abstraction in large language models\.In*The Twelfth International Conference on Learning Representations*, 2024\.URL[https://openreview\.net/forum?id=3bq3jsvcQ1](https://openreview.net/forum?id=3bq3jsvcQ1)\.

## Appendix Contents

## Appendix ARelated Works

##### Persona\-conditioned and personalized language modeling\.

Role prompting is a widely used interface for steering large language models, but recent work suggests that a persona can be treated as a more structured object than a generic expert role\. PersonaHub scales persona\-conditioned data synthesis to a billion synthetic personas, showing that persona descriptions can serve as reusable carriers of perspective and task variation\([Ge et al\., 2024](https://arxiv.org/html/2609.30492#bib.bib16)\)\. In parallel, work on personalized response generation has emphasized that user profiles are dynamic, sparse, and context\-dependent: PersonaMem benchmarks LLMs on maintaining and using evolving user profiles over long interaction histories\([Jiang et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib25)\)\. These efforts establish personas as useful conditioning variables, but they primarily focus on sampling, profiling, or evaluating personalization\. Other work cautions that LLM\-generated personas may encode systematic biases or brittle assumptions, especially when used for downstream simulation or evaluation\([Li et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib35)\)\. In contrast, our work studies personas as*optimizable*latent controls: rather than relying on generic roles or fixed persona pools, we extract, adapt, select, and contrast personas to improve the novelty and diversity of model generations\.

##### Diversity\-aware generation and optimization\.

A long line of work has studied diversity in generation and subset selection\. Determinantal point processes provide a principled probabilistic model for selecting diverse subsets while balancing quality and diversity\([Kulesza & Taskar, 2012](https://arxiv.org/html/2609.30492#bib.bib31)\)\. Submodular facility\-location objectives similarly formalize coverage and redundancy reduction, with classical greedy algorithms providing approximation guarantees for monotone submodular maximization under cardinality constraints\([Nemhauser et al\., 1978](https://arxiv.org/html/2609.30492#bib.bib40)\)\. In evolutionary computation, quality\-diversity methods such as MAP\-Elites search for collections of high\-performing but behaviorally distinct solutions, rather than optimizing a single best point\([Mouret & Clune, 2015](https://arxiv.org/html/2609.30492#bib.bib39)\)\. Recent LLM work has brought related ideas into prompt and decoding optimization: Promptbreeder evolves task prompts through population\-based self\-improvement\([Fernando et al\., 2023](https://arxiv.org/html/2609.30492#bib.bib13)\), Plug\-and\-Play Language Models steer generation through gradient\-based control without retraining the base model\([Dathathri et al\., 2020](https://arxiv.org/html/2609.30492#bib.bib9)\), and Contrastive Search improves open\-ended generation by discouraging degeneration while preserving coherence\([Su & Collier, 2023](https://arxiv.org/html/2609.30492#bib.bib54)\)\. Our approach builds on these ideas but changes the optimization target: we optimize*persona populations*rather than only prompts, logits, or samples, with the goal of producing useful creative diversity rather than only improving task accuracy or surface\-level lexical variation\.

##### Pluralism, coverage, and model homogeneity\.

The motivation for optimizing persona diversity is closely related to recent concerns about output\-space collapse\. Work on pluralistic alignment argues that many tasks admit a spectrum of valid responses, and that alignment should preserve calibrated diversity rather than collapse toward a single normative answer\([Sorensen et al\., 2024](https://arxiv.org/html/2609.30492#bib.bib51)\)\. Spectrum Tuning further operationalizes this idea through distributional coverage and in\-context steerability, showing that post\-training can affect how well models cover diverse valid responses\([Sorensen et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib52)\)\. Related evidence on artificial hiveminds suggests that LLMs can exhibit substantial homogeneity on open\-ended tasks, even when many distinct responses would be reasonable\([Jiang et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib26)\)\. These findings motivate our central hypothesis: increasing structured diversity in the persona or prompt space can increase diversity in the response space, provided that the personas remain task\-relevant and high quality\.

##### Creativity benchmarks and metrics\.

Evaluating creativity in LLMs is challenging because novelty, usefulness, surprise, fluency, and diversity can diverge\. Classic divergent\-thinking paradigms such as the Alternative Uses Test have been adapted to evaluate whether LLMs can produce semantically distant but plausible uses for everyday objects\([Stevenson et al\., 2022](https://arxiv.org/html/2609.30492#bib.bib53)\), while the Divergent Association Task probes whether models can generate remote semantic associations\([Chen & Ding, 2023](https://arxiv.org/html/2609.30492#bib.bib7)\)\. For creative writing, TTCW shows that LLM\-generated stories often fall short of professional human writing under expert evaluation\([Chakrabarty et al\., 2024](https://arxiv.org/html/2609.30492#bib.bib5)\), and CS4 studies creativity under increasing numbers of story\-writing constraints, highlighting trade\-offs between originality, coherence, and instruction following\([Atmakuru et al\., 2024](https://arxiv.org/html/2609.30492#bib.bib2)\)\. Complementary metric work such as the Creativity Index quantifies linguistic novelty through attribution against web text\([Lu et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib37)\), while CreativityPrism evaluates creativity holistically across tasks, domains, and metrics\([Hou et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib22)\)\. These benchmarks motivate our evaluation design: to show that persona\-conditioned generation improves creativity, it is not sufficient to produce stylistically different outputs; the outputs must be measurably more novel, diverse, coherent, and useful across multiple creativity settings\.

## Appendix BSelection Algorithms and Proofs

### B\.1Adaptive Thresholded Coverage

#### B\.1\.1Exact ILP formulation of the coverage objective

As outlined in[Section3\.2](https://arxiv.org/html/2609.30492#S3.SS2), we formulate persona selection as a maximal\-covering location problem rather than an additive maximum\-similarity objective\. While an additive model maximizes the summed similarity of every persona to its nearest exemplar\([Wei et al\., 2015](https://arxiv.org/html/2609.30492#bib.bib56)\), its marginal gain weights every demand point’s similarity improvement, allowing dense regions of the persona space to disproportionately dominate the objective\. By contrast, under our thresholded coverage model, a candidate exemplar’s marginal gain is strictly the number of previously uncovered demand personas inside its radius\-rrneighborhood; already\-covered points contribute zero additional gain\. \(Under cosine dissimilarity, this effectively thresholds the cosine similarity atτcov=1−r\\tau\_\{\\mathrm\{cov\}\}=1\-r, yielding aτcov\\tau\_\{\\mathrm\{cov\}\}\-coverage model\)\.

Recall the coverage objective for a selected subset𝒮\\mathcal\{S\}of cardinalitykkat a fixed radiusrr:Fcov,m​\(𝒮,𝒫,r\)F\_\{\\mathrm\{cov\},m\}\(\\mathcal\{S\};\\mathcal\{P\},r\)\(equation[7](https://arxiv.org/html/2609.30492#S3.E7)\)\. We adopt the conventionmin⁡∅=\+∞\\min\\varnothing=\+\\infty, ensuring that no persona is covered by the empty set andFcov,m​\(∅,𝒫,r\)=0F\_\{\\mathrm\{cov\},m\}\(\\varnothing;\\mathcal\{P\},r\)=0\.

At a fixed radiusrr, the maximum\-coverage persona\-selection problem seeks a subset𝒮cov,m⋆​\(𝒫,r\)\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{cov\},m\}\(\\mathcal\{P\};r\)of cardinalitykkthat leaves as little of the candidate population uncovered as possible

𝒮cov,m⋆​\(𝒫,r\)∈arg⁡max𝒮⊆𝒫\|𝒮\|=k​Fcov,m​\(𝒮,𝒫,r\)\.\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{cov\},m\}\(\\mathcal\{P\};r\)\\in\\arg\\max\_\{\\begin\{subarray\}\{c\}\\mathcal\{S\}\\subseteq\\mathcal\{P\}\\\\ \|\\mathcal\{S\}\|=k\\end\{subarray\}\}F\_\{\\mathrm\{cov\},m\}\(\\mathcal\{S\};\\mathcal\{P\},r\)\.\(9\)
Because the innermin\\minoperator and the indicator function render this objective \(equation[7](https://arxiv.org/html/2609.30492#S3.E7)\) non\-linear, we linearize it exactly\. First, we define the coverage neighborhood of personapip\_\{i\}, which always containsiiitself

𝒥mcov​\(i,r\)=\{j∈\[N𝒫\]:dmP​\(pi,pj\)≤r\},\\mathcal\{J\}^\{\\mathrm\{cov\}\}\_\{m\}\(i;r\)=\\bigl\\\{j\\in\[N\_\{\\mathcal\{P\}\}\]:d\_\{m\}^\{\\mathrm\{P\}\}\(p\_\{i\},p\_\{j\}\)\\leq r\\bigr\\\},\(10\)
We then introduce two sets of binary decision variables

- •χj∈\{0,1\}\\chi\_\{j\}\\in\\\{0,1\\\}: selection indicator for exemplarpjp\_\{j\};
- •zicov∈\{0,1\}z\_\{i\}^\{\\mathrm\{cov\}\}\\in\\\{0,1\\\}: covered indicator for candidate personapip\_\{i\}\.

Writing𝝌=\(χ1,…,χN𝒫\)⊤\\bm\{\\chi\}=\(\\chi\_\{1\},\\ldots,\\chi\_\{N\_\{\\mathcal\{P\}\}\}\)^\{\\top\}and𝐳cov=\(z1cov,…,zN𝒫cov\)⊤\\mathbf\{z\}^\{\\mathrm\{cov\}\}=\(z\_\{1\}^\{\\mathrm\{cov\}\},\\ldots,z\_\{N\_\{\\mathcal\{P\}\}\}^\{\\mathrm\{cov\}\}\)^\{\\top\}, our goal is to select exactlykkexemplars to maximize the unnormalized covered count\. Structuring the objective to be integer\-valued allows us to pose the exact maximal covering location model as the following Integer Linear Program \(ILP\)

maximize𝝌,𝐳cov\\displaystyle\\underset\{\\bm\{\\chi\},\\mathbf\{z\}^\{\\mathrm\{cov\}\}\}\{\\text\{maximize\}\}\\quad∑i=1N𝒫zicov\\displaystyle\\sum\_\{i=1\}^\{N\_\{\\mathcal\{P\}\}\}z\_\{i\}^\{\\mathrm\{cov\}\}\(11\)subject tozicov≤∑j∈𝒥mcov​\(i,r\)χj\\displaystyle z\_\{i\}^\{\\mathrm\{cov\}\}\\leq\\sum\_\{j\\in\\mathcal\{J\}^\{\\mathrm\{cov\}\}\_\{m\}\(i;r\)\}\\chi\_\{j\}∀i∈\[N𝒫\],\\displaystyle\\forall i\\in\[N\_\{\\mathcal\{P\}\}\],\(12\)∑j=1N𝒫χj=k,\\displaystyle\\sum\_\{j=1\}^\{N\_\{\\mathcal\{P\}\}\}\\chi\_\{j\}=k,\(13\)χj∈\{0,1\},zicov∈\{0,1\}\\displaystyle\\chi\_\{j\}\\in\\\{0,1\\\},\\quad z\_\{i\}^\{\\mathrm\{cov\}\}\\in\\\{0,1\\\}∀i,j∈\[N𝒫\]\.\\displaystyle\\forall i,j\\in\[N\_\{\\mathcal\{P\}\}\]\.\(14\)Dividing the optimum of this solver objective byN𝒫N\_\{\\mathcal\{P\}\}perfectly recoversFcov,mF\_\{\\mathrm\{cov\},m\}\(equation[7](https://arxiv.org/html/2609.30492#S3.E7)\)\.

#### B\.1\.2Formulation properties

The thresholded coverage objective avoids the complex scale bookkeeping required by additive formulations\. Its key structural properties stem from the fact that the optimization depends on the dissimilarity metric strictly through its Boolean level sets\.

###### Theorem 4\(Ordinal Invariance\)\.

Letg:ℝ≥0→ℝ≥0g:\\mathbb\{R\}\_\{\\geq 0\}\\rightarrow\\mathbb\{R\}\_\{\\geq 0\}be strictly increasing function withg⁡\(0\)=0g\(0\)=0\. Replacing the dissimilarity metricdmPd\_\{m\}^\{\\mathrm\{P\}\}withg∘dmPg\\circ d\_\{m\}^\{\\mathrm\{P\}\}and the radiusrrbyg⁡\(r\)g\(r\)leaves every coverage neighborhood𝒥mcov​\(i,r\)\\mathcal\{J\}^\{\\mathrm\{cov\}\}\_\{m\}\(i;r\)\(equation[10](https://arxiv.org/html/2609.30492#A2.E10)\) unchanged\. Consequently, the feasible set, the objective function, and the optimizers of the ILP formulation \(equation[11](https://arxiv.org/html/2609.30492#A2.E11)–equation[14](https://arxiv.org/html/2609.30492#A2.E14)\) remain strictly identical\.

###### Proof\.

The program depends on the dissimilarity metric solely through the sets𝒥mcov​\(i,r\)\\mathcal\{J\}^\{\\mathrm\{cov\}\}\_\{m\}\(i;r\)\. Becauseggis strictly increasing, the inequalitydmP​\(pi,pj\)≤rd\_\{m\}^\{\\mathrm\{P\}\}\(p\_\{i\},p\_\{j\}\)\\leq rholds if and only ifg⁡\(dmP​\(pi,pj\)\)≤g⁡\(r\)g\(d\_\{m\}^\{\\mathrm\{P\}\}\(p\_\{i\},p\_\{j\}\)\)\\leq g\(r\)\. Therefore, every level set, and the entire optimization program by extension, is perfectly preserved\. ∎

In practice, this means that formulating the cosine variant as a similarity thresholdτcov\\tau\_\{\\mathrm\{cov\}\}or as a distance radiusr=1−τcovr=1\-\\tau\_\{\\mathrm\{cov\}\}selects the exact same personas\. No frozen reference scale is required to ensure the objective is non\-negative or strictly comparable across different candidate populations\.

###### Theorem 5\(Integrality of Relaxed Coverage Indicators\)\.

If only the covered indicators are relaxed to continuous boundszicov∈\[0,1\]z\_\{i\}^\{\\mathrm\{cov\}\}\\in\[0,1\], then for any fixed, feasible binary selection vector𝛘∈\{0,1\}N𝒫\\bm\{\\chi\}\\in\\\{0,1\\\}^\{N\_\{\\mathcal\{P\}\}\}, the relaxed program is maximized by

\(zicov\)⋆=min⁡\{1,∑j∈𝒥mcov​\(i,r\)χj\}∈\{0,1\}\.\(z\_\{i\}^\{\\mathrm\{cov\}\}\)^\{\\star\}=\\min\\Bigl\\\{1,\\sum\_\{j\\in\\mathcal\{J\}^\{\\mathrm\{cov\}\}\_\{m\}\(i;r\)\}\\chi\_\{j\}\\Bigr\\\}\\in\\\{0,1\\\}\.Thus, this relaxation is exact: integrality is inherently guaranteed for the covered\-indicator block whenever𝛘\\bm\{\\chi\}is binary\. \(Note: This does not assert integrality of the full LP relaxation where𝛘\\bm\{\\chi\}is also relaxed\.\)

###### Proof\.

The objective is non\-decreasing in eachzicovz\_\{i\}^\{\\mathrm\{cov\}\}, and distinct rows are coupled only through𝝌\\bm\{\\chi\}\. Therefore, the solver independently raises eachzicovz\_\{i\}^\{\\mathrm\{cov\}\}to the largest value permitted by the coverage constraint \(equation[12](https://arxiv.org/html/2609.30492#A2.E12)\) and the upper boundzicov≤1z\_\{i\}^\{\\mathrm\{cov\}\}\\leq 1\. That value is the minimum of11and a non\-negative integer, which is inherently binary\. ∎

###### Lemma 1\(Monotone Submodularity\)\.

The unnormalized covered count𝒮↦N𝒫​Fcov,m​\(𝒮,𝒫,r\)\\mathcal\{S\}\\mapsto N\_\{\\mathcal\{P\}\}\\,F\_\{\\mathrm\{cov\},m\}\(\\mathcal\{S\};\\mathcal\{P\},r\)is normalized, monotone, and submodular\. Therefore, a standard greedy algorithm attains at least a\(1−1/e\)\(1\-1/e\)fraction of the optimal covered count\([Nemhauser et al\., 1978](https://arxiv.org/html/2609.30492#bib.bib40)\)\.

###### Proof\.

The covered count evaluates to\|⋃pj∈𝒮\{i∈\[N𝒫\]:j∈𝒥mcov​\(i,r\)\}\|\\bigl\|\\bigcup\_\{p\_\{j\}\\in\\mathcal\{S\}\}\\\{i\\in\[N\_\{\\mathcal\{P\}\}\]:j\\in\\mathcal\{J\}^\{\\mathrm\{cov\}\}\_\{m\}\(i;r\)\\\}\\bigr\|, which is the cardinality of a union of finite sets indexed by the selected exemplars\. This constitutes a classical coverage function, and all coverage functions are normalized, monotone, and submodular\. ∎

As previewed in[Section3\.2](https://arxiv.org/html/2609.30492#S3.SS2), we utilize Lemma[1](https://arxiv.org/html/2609.30492#Thmlemma1)not as an approximation guarantee \(since our final selection algorithm below is exact\), but as the mathematical engine for the cheap, one\-sided greedy certificates that keep the exact solver out of most of the vast majority of the radius search space\.

#### B\.1\.3Adaptive radius calibration

As outlined in[Section3\.2](https://arxiv.org/html/2609.30492#S3.SS2), manually fixing the radiusrrwould reintroduce exactly the dissimilarity\-specific scaling that Theorem[4](https://arxiv.org/html/2609.30492#Thmtheorem4)explicitly eliminates\. We therefore calibrate it from a scale\-free parameter: a target coverage levelccov∈\(0,1\]c\_\{\\mathrm\{cov\}\}\\in\(0,1\], mapping to an absolute covered\-count targetncov=⌈ccov​N𝒫⌉n\_\{\\mathrm\{cov\}\}=\\lceil c\_\{\\mathrm\{cov\}\}N\_\{\\mathcal\{P\}\}\\rceil\.

The calibrated radius is defined as the smallest possible radius at which some size\-kksubset successfully meets the target

rm⋆​\(𝒫,ccov\)=min⁡\{r≥0:max𝒮⊆𝒫\|𝒮\|=k⁡N𝒫​Fcov,m​\(𝒮,𝒫,r\)≥ncov\},r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)=\\min\\Bigl\\\{r\\geq 0:\\max\_\{\\begin\{subarray\}\{c\}\\mathcal\{S\}\\subseteq\\mathcal\{P\}\\\\ \|\\mathcal\{S\}\|=k\\end\{subarray\}\}N\_\{\\mathcal\{P\}\}\\,F\_\{\\mathrm\{cov\},m\}\(\\mathcal\{S\};\\mathcal\{P\},r\)\\geq n\_\{\\mathrm\{cov\}\}\\Bigr\\\},\(15\)In a metric space, this formulation specialized to the discrete generalizedkk\-center\-with\-outliers objective\([Charikar et al\., 2001](https://arxiv.org/html/2609.30492#bib.bib6)\), accommodatingN𝒫−ncovN\_\{\\mathcal\{P\}\}\-n\_\{\\mathrm\{cov\}\}outliers\. \(Note: Our exact finite search operates strictly on pairwise dissimilarities and does not assume a triangle inequality\.\) The final selection algorithm is lexicographic: it first discovers the minimal radiusrm⋆​\(𝒫,ccov\)r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)\(equation[15](https://arxiv.org/html/2609.30492#A2.E15)\), and then break ties among radius\-optimal subsets by maximizing the covered count \(equation[11](https://arxiv.org/html/2609.30492#A2.E11)\) at that exact radius\. We denote this calibrated selector by

𝒮acov,m⋆​\(𝒫,ccov\)∈arg⁡max𝒮⊆𝒫\|𝒮\|=k​Fcov,m​\(𝒮,𝒫,rm⋆​\(𝒫,ccov\)\)\.\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{acov\},m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)\\in\\arg\\max\_\{\\begin\{subarray\}\{c\}\\mathcal\{S\}\\subseteq\\mathcal\{P\}\\\\ \|\\mathcal\{S\}\|=k\\end\{subarray\}\}F\_\{\\mathrm\{cov\},m\}\\\!\\left\(\\mathcal\{S\};\\mathcal\{P\},r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)\\right\)\.\(16\)
###### Lemma 2\(Finite Candidate Radii\)\.

For the finite pairwise spectrumΛm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\)defined in equation[6](https://arxiv.org/html/2609.30492#S2.E6), every fixed subset𝒮\\mathcal\{S\}has a covered count that is non\-decreasing inrrand constant between consecutive values ofΛm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\)\. Consequently, the same holds for its maximum over all size\-kksubsets\. The feasible set of the adaptive radius objective \(equation[15](https://arxiv.org/html/2609.30492#A2.E15)\) is therefore the half\-line\[rm⋆​\(𝒫,ccov\),∞\)\[r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\),\\infty\), and the exact minimumrm⋆​\(𝒫,ccov\)∈Λm​\(𝒫\)r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)\\in\\Lambda\_\{m\}\(\\mathcal\{P\}\)\.

###### Proof\.

Asrrgrows, the indicator function in the coverage objective \(equation[7](https://arxiv.org/html/2609.30492#S3.E7)\) can change only whenrrcrosses a realized dissimilarity in the matrix, as the coverage events\{minpj∈𝒮dmP\(pi,pj\)≤r\}\\\{\\min\_\{p\_\{j\}\\in\\mathcal\{S\}\}d\_\{m\}^\{\\mathrm\{P\}\}\(p\_\{i\},p\_\{j\}\)\\leq r\\\}are nested inrr\. Hence, every covered count is a non\-decreasing step function whose jumps lie exactly inΛm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\)\. Since the pointwise maximum over the finitely many size\-kksubsets is also a step function with jumps inΛm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\), the smallest feasible radius is inherently attained at a jump point\. ∎

By Lemma[2](https://arxiv.org/html/2609.30492#Thmlemma2), calibration reduces to an exact discrete search over the sorted distinct values ofΛm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\)\([Hochbaum & Shmoys, 1986](https://arxiv.org/html/2609.30492#bib.bib21)\)\. To avoid invoking the computationally heavy ILP solver for every probe, we evaluate the monotone decision problem of whetherkkexemplars can cover at leastncovn\_\{\\mathrm\{cov\}\}personas at radiusrr, using a standard size\-kkgreedy maximum\-coverage construction\. LetGm​\(r\)G\_\{m\}\(r\)denote the covered count returned by this greedy heuristic\. Three certificate mechanisms resolve the vast majority of search probes instantly:

1. 1\.*Greedy feasibility certificate\.*IfGm​\(r\)≥ncovG\_\{m\}\(r\)\\geq n\_\{\\mathrm\{cov\}\}, the probed radius is immediately certified as feasible\.
2. 2\.*Greedy infeasibility certificate\.*By the monotone submodularity established in Lemma[1](https://arxiv.org/html/2609.30492#Thmlemma1), the optimal covered count is strictly bounded above by⌊Gm​\(r\)/\(1−1/e\)⌋\\lfloor G\_\{m\}\(r\)/\(1\-1/e\)\\rfloor\. If this proven upper bound falls belowncovn\_\{\\mathrm\{cov\}\}, the probed radius is mathematically certified as infeasible\.
3. 3\.*Certificate tightening\.*Any feasible selection𝒮\\mathcal\{S\}remains feasible down to its own minimal radius \(thencovn\_\{\\mathrm\{cov\}\}\-th smallest nearest\-exemplar dissimilarity under𝒮\\mathcal\{S\}\)\. This allows the feasible bracket to jump far below the initially probed radius in a single computational step\.

Probes falling into the narrow residual window are resolved via bounded optimization\. The exact ILP solver attempts to maximize the covered count but is instructed to terminate early the moment an incumbent reachesncovn\_\{\\mathrm\{cov\}\}\(a feasibility certificate\) or its proven upper bound drops belowncovn\_\{\\mathrm\{cov\}\}\(an infeasibility certificate that remains mathematically valid even if cut off by a time limit, because the objective is strictly integer\-valued\)\. By monotonicity, the search terminates as soon as it certifies that the immediate candidate predecessor ofrm⋆​\(𝒫,ccov\)r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)is infeasible, whenever such a predecessor exists\. That final certificate may come from the greedy upper bound or the exact solver; no lower\-candidate certificate is needed when the optimum isrm,\(1\)r\_\{m,\(1\)\}, the absolute minimum dissimilarity in the spectrum\.

#### B\.1\.4Exact algorithm

Maximum coverage is NP\-hard\([Feige, 1998](https://arxiv.org/html/2609.30492#bib.bib12)\)when the selection budget is part of the input, and exact optimization can be computationally expensive in the worst\-case\. However, our compact ILP formulation \(equation[11](https://arxiv.org/html/2609.30492#A2.E11)–equation[14](https://arxiv.org/html/2609.30492#A2.E14)\) is highly tractable in practice\. To resolve the exact probes during discrete search, the formulation is warm\-started with the greedy incumbent and evaluated by an exact integer or constraint solver \(e\.g\., CP\-SAT\)\.

We accept a final solution as exact only when it achieves certified optimal status\. Because the unnormalized covered count is integral, an incumbent solution is automatically certified as optimal once its proven mathematical upper bound drops less than one unit above it\.

Algorithm 1Exact Adaptive Thresholded\-Coverage Selection1:Input:Candidate population

𝒫\\mathcal\{P\}of size

N𝒫N\_\{\\mathcal\{P\}\}, dissimilarity index

mm, target size

kk, target coverage

ccovc\_\{\\mathrm\{cov\}\}
2:Embed each

pi∈𝒫p\_\{i\}\\in\\mathcal\{P\}⊳\\triangleright𝒪⁡\(N𝒫\)\\mathcal\{O\}\(N\_\{\\mathcal\{P\}\}\)encoder calls

3:Compute

Dm,i​jP←dmP​\(pi,pj\)D^\{\\mathrm\{P\}\}\_\{m,ij\}\\leftarrow d\_\{m\}^\{\\mathrm\{P\}\}\(p\_\{i\},p\_\{j\}\)and

𝐃mP←\(Dm,i​jP\)i,j=1N𝒫\\mathbf\{D\}\_\{m\}^\{\\mathrm\{P\}\}\\leftarrow\(D^\{\\mathrm\{P\}\}\_\{m,ij\}\)\_\{i,j=1\}^\{N\_\{\\mathcal\{P\}\}\}⊳\\triangleright𝒪⁡\(N𝒫2\)\\mathcal\{O\}\(N\_\{\\mathcal\{P\}\}^\{2\}\)pair evaluations

4:

ncov←⌈ccov​N𝒫⌉n\_\{\\mathrm\{cov\}\}\\leftarrow\\lceil c\_\{\\mathrm\{cov\}\}N\_\{\\mathcal\{P\}\}\\rceil⊳\\triangleright𝒪⁡\(1\)\\mathcal\{O\}\(1\)

5:Form

Λm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\)and sort it as

rm,\(1\)<⋯<rm,\(Lm\)r\_\{m,\(1\)\}<\\cdots<r\_\{m,\(L\_\{m\}\)\}⊳\\trianglerightnaively𝒪⁡\(N𝒫2​log⁡N𝒫\)\\mathcal\{O\}\(N\_\{\\mathcal\{P\}\}^\{2\}\\log N\_\{\\mathcal\{P\}\}\)

6:

ℓhi←Lm\\ell\_\{\\mathrm\{hi\}\}\\leftarrow L\_\{m\}with any size\-

kkwitness feasible at

rm,\(Lm\)r\_\{m,\(L\_\{m\}\)\};

ℓlo←0\\ell\_\{\\mathrm\{lo\}\}\\leftarrow 0⊳\\triangleright00is a sentinel, never a radius index

7:Apply greedy certificates and witness tightening to the bracket, preserving feasible

ℓhi\\ell\_\{\\mathrm\{hi\}\}and either

ℓlo=0\\ell\_\{\\mathrm\{lo\}\}=0or certified\-infeasible

ℓlo\\ell\_\{\\mathrm\{lo\}\}
8:while

ℓlo<ℓhi−1\\ell\_\{\\mathrm\{lo\}\}<\\ell\_\{\\mathrm\{hi\}\}\-1do

9:Probe

ℓ←ℓhi−1\\ell\\leftarrow\\ell\_\{\\mathrm\{hi\}\}\-1at radius

rm,\(ℓ\)r\_\{m,\(\\ell\)\}: greedy certificates, else the bounded\-optimization oracle

10:iffeasible, with witness cover

𝒮\\mathcal\{S\}then

11:

ℓhi←\\ell\_\{\\mathrm\{hi\}\}\\leftarrowindex of the minimal radius of

𝒮\\mathcal\{S\}⊳\\trianglerightcertificate tightening

12:else

13:

ℓlo←ℓ\\ell\_\{\\mathrm\{lo\}\}\\leftarrow\\ell⊳\\trianglerightcertified infeasible

14:endif

15:endwhile

16:

r⋆←rm,\(ℓhi\)r^\{\\star\}\\leftarrow r\_\{m,\(\\ell\_\{\\mathrm\{hi\}\}\)\}
17:Solve equation[11](https://arxiv.org/html/2609.30492#A2.E11)–equation[14](https://arxiv.org/html/2609.30492#A2.E14)at

r=r⋆r=r^\{\\star\}to certified optimality, warm\-started with the incumbent cover

18:Output:Globally optimal coverage subset

𝒮acov,m⋆​\(𝒫,ccov\)\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{acov\},m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)and calibrated radius

rm⋆​\(𝒫,ccov\)=r⋆r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)=r^\{\\star\}

#### B\.1\.5Optimality guarantee

###### Proof\.

The proof of global optimality relies on three sequential guarantees:

*Termination\.*The discrete distance spectrumΛm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\)is finite\. Since every iteration of the while\-loop in[Algorithm1](https://arxiv.org/html/2609.30492#alg1)must either strictly lower the upper boundℓhi\\ell\_\{\\mathrm\{hi\}\}by at least one index or raise the lower boundℓlo\\ell\_\{\\mathrm\{lo\}\}toℓhi−1\\ell\_\{\\mathrm\{hi\}\}\-1, the search space strictly shrinks at every step\. Therefore, the algorithm is guaranteed to terminate in finite time\.

*Radius optimality\.*Two invariants are maintained throughout the execution of the search loop\. First, the radiusrm,\(ℓhi\)r\_\{m,\(\\ell\_\{\\mathrm\{hi\}\}\)\}is always feasible, guaranteed by a stored size\-kkwitness cover\. \(Certificate tightening preserves this invariant because any cover is naturally feasible down to its own internal minimal radius by construction\.\) Second, eitherℓlo=0\\ell\_\{\\mathrm\{lo\}\}=0\(acting as the initialized sentinel\) orrm,\(ℓlo\)r\_\{m,\(\\ell\_\{\\mathrm\{lo\}\}\)\}is certified infeasible\. This refutation is valid whether it comes from the greedy certificate \(validated by the submodularity established in Lemma[1](https://arxiv.org/html/2609.30492#Thmlemma1)\) or from the exact solver \(whose proven bound acts as a strict upper bound on the true integral optimum\.\)

Upon termination, the bracket perfectly tightens\. Eitherℓhi=1\\ell\_\{\\mathrm\{hi\}\}=1, meaning the smallest candidate radius is in the spectrum is feasible, orℓlo=ℓhi−1≥1\\ell\_\{\\mathrm\{lo\}\}=\\ell\_\{\\mathrm\{hi\}\}\-1\\geq 1, meaning the immediate predecessor to the upper bound is infeasible\. By monotonicity and the discrete step\-function property established in Lemma[2](https://arxiv.org/html/2609.30492#Thmlemma2), this guarantees that the returned radius is the global minimum

r⋆=rm,\(ℓhi\)=rm⋆​\(𝒫,ccov\)\.r^\{\\star\}=r\_\{m,\(\\ell\_\{\\mathrm\{hi\}\}\)\}=r^\{\\star\}\_\{m\}\(\\mathcal\{P\};c\_\{\\mathrm\{cov\}\}\)\.
*Coverage optimality atr⋆r^\{\\star\}\.*The final step of the algorithm executes a bounded optimization over the finite set of selection vectors𝝌∈\{0,1\}N𝒫\\bm\{\\chi\}\\in\\\{0,1\\\}^\{N\_\{\\mathcal\{P\}\}\}subject to the exact cardinality constraint∑jχj=k\\sum\_\{j\}\\chi\_\{j\}=k\. The ILP solver evaluates this space and yields a feasible incumbent cover alongside a proven upper bound on the maximum possible covered count\. Achieving a certified optimal status guarantees that these two values are equal\. Furthermore, because the unnormalized covered count is integer\-valued, proving an absolute gap of less than one is sufficient to rule out the existence of a better feasible selection\. ∎

### B\.2Max–Min Dispersion Persona Selection

#### B\.2\.1Structural properties and thresholded equivalence

Unlike greedy heuristics, the exact incremental\-clique method relies on the discrete topology of the solution space rather than the continous geometric properties of the embedding space\.

###### Assumption 1\(Pairwise dissimilarity\)\.

The chosen dissimilarity functiondmP​\(⋅,⋅\)d\_\{m\}^\{\\mathrm\{P\}\}\(\\cdot,\\cdot\)is symmetric and non\-negative\. However, a triangle inequality is not required for the exact clique reduction to hold\.

The theoretical guarantee of this exact search is grounded in the finite, discrete nature of the candidate pool:

###### Lemma 3\(Bottleneck Distance,\([Erkut, 1990](https://arxiv.org/html/2609.30492#bib.bib11)\)\)\.

Because the candidate population𝒫\\mathcal\{P\}is finite, the optimal valueFdisp,m​\(𝒮disp,m⋆​\(𝒫\)\)F\_\{\\mathrm\{disp\},m\}\(\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{disp\},m\}\(\\mathcal\{P\}\)\)equals the weight of at least one edge in the complete pairwise\-dissimilarity graph on constructed on𝒫\\mathcal\{P\}\.

#### B\.2\.2Exact incremental\-clique algorithm

Algorithm 2Exact Incremental\-Clique Max–Min Dispersion1:Input:Candidate population

𝒫\\mathcal\{P\}, target size

kk, persona dissimilarity

dmPd\_\{m\}^\{\\mathrm\{P\}\}
2:Compute the distinct pairwise\-dissimilarity levels

Λm​\(𝒫\)\\Lambda\_\{m\}\(\\mathcal\{P\}\)and order them as

rm,\(1\)<⋯<rm,\(Lm\)r\_\{m,\(1\)\}<\\cdots<r\_\{m,\(L\_\{m\}\)\}\.

3:Initialize

𝒢=\(𝒫,ℰ\)\\mathcal\{G\}=\(\\mathcal\{P\},\\mathcal\{E\}\)with

ℰ←∅\\mathcal\{E\}\\leftarrow\\emptyset
4:for

ℓ=Lm,Lm−1,…,1\\ell=L\_\{m\},L\_\{m\}\-1,\\ldots,1do

5:

ℰℓ←\{\{pa,pb\}:a<b,dmP\(pa,pb\)=rm,\(ℓ\)\}\\mathcal\{E\}\_\{\\ell\}\\leftarrow\\\{\\\{p\_\{a\},p\_\{b\}\\\}:a<b,\\ d\_\{m\}^\{\\mathrm\{P\}\}\(p\_\{a\},p\_\{b\}\)=r\_\{m,\(\\ell\)\}\\\}
6:foreach edge

\{pa,pb\}∈ℰℓ\\\{p\_\{a\},p\_\{b\}\\\}\\in\\mathcal\{E\}\_\{\\ell\}do

7:

ℰ←ℰ∪\{\{pa,pb\}\}\\mathcal\{E\}\\leftarrow\\mathcal\{E\}\\cup\\\{\\\{p\_\{a\},p\_\{b\}\\\}\\\}
8:

𝒩a,b←\{pc∈𝒫:\{pa,pc\}∈ℰ∧\{pb,pc\}∈ℰ\}\\mathcal\{N\}\_\{a,b\}\\leftarrow\\\{p\_\{c\}\\in\\mathcal\{P\}:\\\{p\_\{a\},p\_\{c\}\\\}\\in\\mathcal\{E\}\\land\\\{p\_\{b\},p\_\{c\}\\\}\\in\\mathcal\{E\}\\\}
9:if

\|𝒩a,b\|≥k−2\|\\mathcal\{N\}\_\{a,b\}\|\\geq k\-2then

10:Test whether

𝒢⁡\[𝒩a,b\]\\mathcal\{G\}\[\\mathcal\{N\}\_\{a,b\}\]contains a

\(k−2\)\(k\-2\)\-clique

11:ifa

\(k−2\)\(k\-2\)\-clique

𝒬\\mathcal\{Q\}existsthen

12:Return

𝒮disp,m⋆​\(𝒫\)←𝒬∪\{pa,pb\}\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{disp\},m\}\(\\mathcal\{P\}\)\\leftarrow\\mathcal\{Q\}\\cup\\\{p\_\{a\},p\_\{b\}\\\}
13:endif

14:endif

15:endfor

16:endfor

#### B\.2\.3Optimality guarantee

###### Proof\.

Letδ⋆\\delta^\{\\star\}denote the optimal max–min dispersion\. By Lemma[3](https://arxiv.org/html/2609.30492#Thmlemma3)\(the bottleneck lemma\),δ⋆=rm,\(ℓ⋆\)\\delta^\{\\star\}=r\_\{m,\(\\ell^\{\\star\}\)\}for some threshold in the sorted distance spectrum\. After all edges at thresholdrm,\(ℓ\)r\_\{m,\(\\ell\)\}have been inserted,[Algorithm2](https://arxiv.org/html/2609.30492#alg2)has effectively constructed the threshold graph

𝒢ℓ=\(𝒫,\{\{p,p′\}:p≠p′,dmP\(p,p′\)≥rm,\(ℓ\)\}\)\.\\mathcal\{G\}\_\{\\ell\}=\\bigl\(\\mathcal\{P\},\\\{\\\{p,p^\{\\prime\}\\\}:p\\neq p^\{\\prime\},\\ d\_\{m\}^\{\\mathrm\{P\}\}\(p,p^\{\\prime\}\)\\geq r\_\{m,\(\\ell\)\}\\\}\\bigr\)\.A size\-kksubset has a minimum internal dispersion of at leastrm,\(ℓ\)r\_\{m,\(\\ell\)\}if and only if it forms a fullkk\-clique in𝒢ℓ\\mathcal\{G\}\_\{\\ell\}\. The algorithm processes thresholds from largest to smallest and returns as soon as the firstkk\-clique is completed\. If a subset with a strictly larger dispersion existed, its corresponding clique would have been fully connected at an earlier \(larger\) threshold and the algorithm would have already terminated\. Hence the returned clique inherently possesses the bottleneck distancerm,\(ℓ⋆\)=δ⋆r\_\{m,\(\\ell^\{\\star\}\)\}=\\delta^\{\\star\}, proving it is globally optimal\. ∎

## Appendix CGeneration Algorithms and Proofs

### C\.1Uniform\-Coverage MCMC Persona Generation

We propose*Uniform\-Coverage MCMC*\(UC\-MCMC\), a persona\-generation sampler whose operational target assigns equal probability mass to the reachable cells of a fixed, finite\-resolution partition of semantic directions\. Within each cell, a frozen language model supplies a reference base law favoring linguistically plausible personas; valid proposals that land in previously unoccupied cells act to expand the active target in hindsight\. The construction successfully pursues coverage through its sampling law, rather than by first generating a pool and then retrospectively equalizing its cell counts\.

#### C\.1\.1Semantic Geometry and Target Distributions

##### State Space and Validity\.

We specializeΩ\\Omegato the countable persona state space of canonically serialized, EOS\-terminated persona token sequences up to a fixed maximum length\. A nonterminating generation is represented by a failure symbol⊥∉Ω\\bot\\notin\\Omega\. Letval:Ω→\{0,1\}\\operatorname\{val\}:\\Omega\\rightarrow\\\{0,1\\\}be a fixed validity function\. This function may incorporate a cached LLM judgment, provided that the model, prompt, decoding rule, and cache are frozen so thatval⁡\(p\)\\operatorname\{val\}\(p\)is deterministic\. We extend this function with the conventionval⁡\(⊥\)=0\\operatorname\{val\}\(\\bot\)=0\.

##### Semantic Geometry and Cell Partitions\.

Letϕ~\\widetilde\{\\phi\}be the MRL\-truncated text embedding, whose restriction to the persona state space mapsΩ\\OmegatoℝMd\\mathbb\{R\}^\{d\}\_\{\\mathrm\{M\}\}\. We freeze a centerμ^\\widehat\{\\mu\}and a nonsingular whitening mapWW\(computed using only the baseline population or a separate pilot sample\)\. The semantic direction of a persona is defined as its normalized projection

𝐳W​\(p\)=W​\{ϕ~​\(p\)−μ^\}‖W⁡\{ϕ~​\(p\)−μ^\}‖2∈𝕊dM−1\.\\mathbf\{z\}\_\{W\}\(p\)=\\frac\{W\\\{\\widetilde\{\\phi\}\(p\)\-\\widehat\{\\mu\}\\\}\}\{\\left\\\|W\\\{\\widetilde\{\\phi\}\(p\)\-\\widehat\{\\mu\}\\\}\\right\\\|\_\{2\}\}\\in\\mathbb\{S\}^\{d\_\{\\mathrm\{M\}\}\-1\}\.\(17\)We assume the denominator is nonzero for every valid persona\. We then partition the sphere𝕊dM−1\\mathbb\{S\}^\{d\_\{\\mathrm\{M\}\}\-1\}intoMMmeasurable, equal\-area cells𝒞1,…,𝒞M\\mathcal\{C\}\_\{1\},\\ldots,\\mathcal\{C\}\_\{M\}, with deterministic tie\-breaking on cell boundaries\. \(Equal\-area, small\-diameter sphere partitions can be constructed using, for example, the recursive zonal method of[Leopardi \(2006\)](https://arxiv.org/html/2609.30492#bib.bib33)\.\) We define the cell\-assignment functionc⁡\(p\)=jc\(p\)=jwhen𝐳W​\(p\)∈𝒞j\\mathbf\{z\}\_\{W\}\(p\)\\in\\mathcal\{C\}\_\{j\}, allowing use to formally define the valid persona set within celljjas

Ωj=\{p∈Ω:val\(p\)=1,c\(p\)=j\}\.\\Omega\_\{j\}=\\\{p\\in\\Omega:\\operatorname\{val\}\(p\)=1,\\ c\(p\)=j\\\}\.\(18\)Note that equal area is required for the geometric interpretation formalized below, but the underlying MCMC validity results require only a fixed measurable partition\.

##### Base Law and Cell\-Conditioned Targets\.

LetGθPG\_\{\\theta\_\{\\mathrm\{P\}\}\}be the frozen persona\-generation model and letxgenx\_\{\\mathrm\{gen\}\}be its fully serialized, fixed prompt\. For personapp, letw1:L⁡\(p\)w\_\{1:L\(p\)\}be its exact output token sequence, excluding terminal EOS\. We define the base probability mass ofppas

b⁡\(p\)=\{∏ℓ=1L⁡\(p\)GθP​\(wℓ∣xgen,w<ℓ\)\}​GθP​\(EOS∣xgen,w≤L⁡\(p\)\)\.b\(p\)=\\left\\\{\\prod\_\{\\ell=1\}^\{L\(p\)\}G\_\{\\theta\_\{\\mathrm\{P\}\}\}\(w\_\{\\ell\}\\mid x\_\{\\mathrm\{gen\}\},w\_\{<\\ell\}\)\\right\\\}G\_\{\\theta\_\{\\mathrm\{P\}\}\}\(\\mathrm\{EOS\}\\mid x\_\{\\mathrm\{gen\}\},w\_\{\\leq L\(p\)\}\)\.\(19\)Every generation outcome outsideΩ\\Omegais mapped to the failure symbol⊥\\botwithout resampling\. Thus, we set

b⁡\(⊥\)=1−∑p∈Ωb⁡\(p\)b\(\\bot\)=1\-\\sum\_\{p\\in\\Omega\}b\(p\)as the probability law onΩ⊥:=Ω∪\{⊥\}\\Omega\_\{\\bot\}:=\\Omega\\cup\\\{\\bot\\\}, while⊥\\botcarries zero target weight becauseval⁡\(⊥\)=0\\operatorname\{val\}\(\\bot\)=0\. We assumeb⁡\(p\)\>0b\(p\)\>0for every valid persona, which naturally holds for raw softmax sampling when all canonical output tokens remain in support\.

In our primary configuration, the global independence proposal is exactly this completion law

q0​\(p′∣p\)=q0​\(p′\)=b⁡\(p′\)\.q\_\{0\}\(p^\{\\prime\}\\mid p\)=q\_\{0\}\(p^\{\\prime\}\)=b\(p^\{\\prime\}\)\.\(20\)
The base mass of celljj, accounting for validity, is therefore

Bj=∑p∈Ωb\(p\)val\(p\)𝟙\{c\(p\)=j\},ℛ=\{j:Bj\>0\},B\_\{j\}=\\sum\_\{p\\in\\Omega\}b\(p\)\\operatorname\{val\}\(p\)\\mathbbm\{1\}\\\{c\(p\)=j\\\},\\qquad\\mathcal\{R\}=\\\{j:B\_\{j\}\>0\\\},\(21\)whereℛ\\mathcal\{R\}is defined as the set of reachable cells\. For anyj∈ℛj\\in\\mathcal\{R\}, the within\-cell target distribution is defined as

πj​\(p\)=b\(p\)val\(p\)𝟙\{c\(p\)=j\}Bj\.\\pi\_\{j\}\(p\)=\\frac\{b\(p\)\\operatorname\{val\}\(p\)\\mathbbm\{1\}\\\{c\(p\)=j\\\}\}\{B\_\{j\}\}\.\(22\)For an active\-cell index setℐ⊆ℛ\\mathcal\{I\}\\subseteq\\mathcal\{R\}, the target distribution of a single emitted persona is the equal\-cell mixture

Πℐ=1\|ℐ\|∑j∈ℐπj,Πℐ\(Ωj\)=1\|ℐ\|\(j∈ℐ\)\.\\Pi\_\{\\mathcal\{I\}\}=\\frac\{1\}\{\|\\mathcal\{I\}\|\}\\sum\_\{j\\in\\mathcal\{I\}\}\\pi\_\{j\},\\qquad\\Pi\_\{\\mathcal\{I\}\}\(\\Omega\_\{j\}\)=\\frac\{1\}\{\|\\mathcal\{I\}\|\}\\quad\(j\\in\\mathcal\{I\}\)\.\(23\)In contrast, the joint target distribution for the full population containing exactly one chain per cell is the product measure

Γℐ=⨂j∈ℐπj\.\\Gamma\_\{\\mathcal\{I\}\}=\\bigotimes\_\{j\\in\\mathcal\{I\}\}\\pi\_\{j\}\.\(24\)

##### Resolution\-Limited Directional Uniformity\.

The following lemma formalizes the finite\-resolution connection to true directional uniformity across the continuous semantic sphere\.

###### Lemma 4\(Wasserstein Bound on Uniformity\)\.

Letσℐ\\sigma\_\{\\mathcal\{I\}\}be the normalized spherical surface measure on𝒞ℐ=⋃j∈ℐ𝒞j\\mathcal\{C\}\_\{\\mathcal\{I\}\}=\\bigcup\_\{j\\in\\mathcal\{I\}\}\\mathcal\{C\}\_\{j\}, letνℐ=\(𝐳W\)\#​Πℐ\\nu\_\{\\mathcal\{I\}\}=\(\\mathbf\{z\}\_\{W\}\)\_\{\\\#\}\\Pi\_\{\\mathcal\{I\}\}be the pushforward ofΠℐ\\Pi\_\{\\mathcal\{I\}\}to the sphere, and letΔj\\Delta\_\{j\}be the geodesic diameter of cell𝒞j\\mathcal\{C\}\_\{j\}\. For every Wasserstein orderqW≥1q\_\{\\mathrm\{W\}\}\\geq 1

WqW​\(νℐ,σℐ\)≤\(1\|ℐ\|​∑j∈ℐΔjqW\)1/qW≤maxj∈ℐ⁡Δj\.W\_\{q\_\{\\mathrm\{W\}\}\}\(\\nu\_\{\\mathcal\{I\}\},\\sigma\_\{\\mathcal\{I\}\}\)\\leq\\left\(\\frac\{1\}\{\|\\mathcal\{I\}\|\}\\sum\_\{j\\in\\mathcal\{I\}\}\\Delta\_\{j\}^\{q\_\{\\mathrm\{W\}\}\}\\right\)^\{1/q\_\{\\mathrm\{W\}\}\}\\leq\\max\_\{j\\in\\mathcal\{I\}\}\\Delta\_\{j\}\.\(25\)

###### Proof\.

Since both measures assign a probability mass1/\|ℐ\|1/\|\\mathcal\{I\}\|to every individual cell𝒞j\\mathcal\{C\}\_\{j\}, we can couple their conditional distributions separately within each respective cell\. The physical geodesic distance between any two points coupled inside𝒞j\\mathcal\{C\}\_\{j\}is naturally bounded by its diameterΔj\\Delta\_\{j\}\. ∎

Lemma[4](https://arxiv.org/html/2609.30492#Thmlemma4)connects equal cell masses to directional uniformity when cell diameters are small\. Our implementation uses ten orthonormal sign cuts in 128 dimensions, to the shown bound does not provide a nontrivial approximation guarantee to uniform surface measure\. Rather, the operational guarantee is equal allocation across reachable cells\.

#### C\.1\.2Within\-Cell Metropolis Kernels and Proposal Requirements

##### Component Kernels and Acceptance Probabilities

In addition to the global independence proposalq0q\_\{0\}, letq1,…,qLqq\_\{1\},\\ldots,q\_\{L\_\{q\}\}be a set of fixed, exactly evaluable local\-edit or block proposals\. At each step, the algorithm first draws a proposal typeℓ\\ellwith a fixed probabilityωℓ\>0\\omega\_\{\\ell\}\>0, such that∑ℓ=0Lqωℓ=1\\sum\_\{\\ell=0\}^\{L\_\{q\}\}\\omega\_\{\\ell\}=1, and then applies a component\-specific Metropolis–Hastings \(MH\) correction\.

For a chain assigned to an active celljj, the acceptance probability for transitioning from an incumbent personappto a proposed personap′p^\{\\prime\}is

αj,ℓ\(p,p′\)=𝟙\{p′∈Ωj\}min\{1,b⁡\(p′\)​qℓ​\(p∣p′\)b⁡\(p\)​qℓ​\(p′∣p\)\}\.\\alpha\_\{j,\\ell\}\(p,p^\{\\prime\}\)=\\mathbbm\{1\}\\\{p^\{\\prime\}\\in\\Omega\_\{j\}\\\}\\min\\left\\\{1,\\,\\frac\{b\(p^\{\\prime\}\)q\_\{\\ell\}\(p\\mid p^\{\\prime\}\)\}\{b\(p\)q\_\{\\ell\}\(p^\{\\prime\}\\mid p\)\}\\right\\\}\.\(26\)A zero reverse probability implies a zero acceptance probability\. Importantly, if the proposal type or edited block is selected with a state\-dependent probability, that probability must be explicitly included in the forward–reverse ratio\. Equation equation[26](https://arxiv.org/html/2609.30492#A3.E26)describes a random mixture of separately corrected kernels; it does not represent the acceptance ratio for a single marginalized mixture proposal\.

LetKj,ℓK\_\{j,\\ell\}denote the resulting accept–reject kernel for proposalℓ\\ell, which includes the self\-transition upon rejection, and define the full mixture kernel as

Kj=∑ℓ=0Lqωℓ​Kj,ℓ\.K\_\{j\}=\\sum\_\{\\ell=0\}^\{L\_\{q\}\}\\omega\_\{\\ell\}K\_\{j,\\ell\}\.\(27\)
###### Theorem 6\(Cell\-Kernel Correctness and Matched Refresh\)\.

For every reachable celljj, each component kernelKj,ℓK\_\{j,\\ell\}is reversible with respect toπj\\pi\_\{j\}; hence the mixtureKjK\_\{j\}leavesπj\\pi\_\{j\}invariant\. Moreover, under the matched global proposal \(q0=bq\_\{0\}=b\), for everyp∈Ωjp\\in\\Omega\_\{j\}, proposal valuep′∈Ω⊥p^\{\\prime\}\\in\\Omega\_\{\\bot\}, and measurable subsetE⊆ΩjE\\subseteq\\Omega\_\{j\}

αj,0​\(p,p′\)\\displaystyle\\alpha\_\{j,0\}\(p,p^\{\\prime\}\)=𝟙\{p′∈Ωj\},\\displaystyle=\\mathbbm\{1\}\\\{p^\{\\prime\}\\in\\Omega\_\{j\}\\\},\(28\)Kj,0​\(p,E\)\\displaystyle K\_\{j,0\}\(p,E\)=Bj​πj​\(E\)\+\(1−Bj\)​δp​\(E\),\\displaystyle=B\_\{j\}\\pi\_\{j\}\(E\)\+\(1\-B\_\{j\}\)\\delta\_\{p\}\(E\),\(29\)Kj​\(p,E\)\\displaystyle K\_\{j\}\(p,E\)≥ω0​Bj​πj​\(E\)\.\\displaystyle\\geq\\omega\_\{0\}B\_\{j\}\\pi\_\{j\}\(E\)\.\(30\)whereδp\\delta\_\{p\}denotes the Dirac probability measure atpp\. Consequently, for every arbitrary initialization lawζ\\zetasupported onΩj\\Omega\_\{j\}and every integern≥0n\\geq 0

‖ζ​Kjn−πj‖TV≤\(1−ω0​Bj\)n\.\\left\\\|\\zeta K\_\{j\}^\{n\}\-\\pi\_\{j\}\\right\\\|\_\{\\mathrm\{TV\}\}\\leq\(1\-\\omega\_\{0\}B\_\{j\}\)^\{n\}\.\(31\)

###### Proof\.

For any two valid same\-cell statesp,p′∈Ωjp,p^\{\\prime\}\\in\\Omega\_\{j\}, the two directional probability flows between them under kernelℓ\\ellequal

min⁡\{πj​\(p\)​qℓ​\(p′∣p\),πj​\(p′\)​qℓ​\(p∣p′\)\}\.\\min\\\{\\pi\_\{j\}\(p\)q\_\{\\ell\}\(p^\{\\prime\}\\mid p\),\\pi\_\{j\}\(p^\{\\prime\}\)q\_\{\\ell\}\(p\\mid p^\{\\prime\}\)\\\}\.Since this flow is perfectly symmetric, detailed balance is satisfied by the standard Metropolis–Hastings argument\([Hastings, 1970](https://arxiv.org/html/2609.30492#bib.bib20);[Tierney, 1994](https://arxiv.org/html/2609.30492#bib.bib55)\)\.

When evaluating the global proposal \(q0=bq\_\{0\}=b\), the proposal probabilities and base\-weight terms cancel out for any valid same\-cell move\. A global proposal therefore hitsΩj\\Omega\_\{j\}with total probabilityBjB\_\{j\}and, conditional on successfully doing so, is distributed according to lawπj\\pi\_\{j\}, yielding the independent refresh kernelKj,0K\_\{j,0\}\(equation[29](https://arxiv.org/html/2609.30492#A3.E29)\)\. BecauseKj,0K\_\{j,0\}is drawn with probabilityω0\\omega\_\{0\}, the fixed mixture satisfies the minorization condition \(equation[30](https://arxiv.org/html/2609.30492#A3.E30)\)\. Finally, the resulting total\-variation convergence bound follows from the classical Doeblin coupling argument\([Mengersen & Tweedie, 1996](https://arxiv.org/html/2609.30492#bib.bib38);[Rosenthal, 1995](https://arxiv.org/html/2609.30492#bib.bib48)\)\. ∎

#### C\.1\.3The Hindsight\-Spawning Population Algorithm\.

Let𝒫0\\mathcal\{P\}\_\{0\}be the baseline persona population\. We initialize the active\-cell index set as the set of valid cells discovered in the baseline

ℐ0=\{c\(p\):p∈𝒫0,val\(p\)=1\}\.\\mathcal\{I\}\_\{0\}=\\\{c\(p\):p\\in\\mathcal\{P\}\_\{0\},\\ \\operatorname\{val\}\(p\)=1\\\}\.For eachj∈ℐ0j\\in\\mathcal\{I\}\_\{0\}, we choose one arbitrary valid seedP\(j\)∈ΩjP^\{\(j\)\}\\in\\Omega\_\{j\}, assuming its base weightb⁡\(P\(j\)\)\>0b\(P^\{\(j\)\}\)\>0\. This seed may instead be drawn uniformly from the baseline personas residing in that cell; its specific initialization law does not affect the anytime cell\-level guarantees established below\.

Algorithm 3Uniform\-Coverage MCMC Persona Generation1:Baseline population

𝒫0\\mathcal\{P\}\_\{0\}; frozen evaluation components

\(val,𝐳W,c,b\)\(\\operatorname\{val\},\\mathbf\{z\}\_\{W\},c,b\); proposals

\{qℓ\}ℓ=0Lq\\\{q\_\{\\ell\}\\\}\_\{\\ell=0\}^\{L\_\{q\}\}and their probabilities

\{ωℓ\}ℓ=0Lq\\\{\\omega\_\{\\ell\}\\\}\_\{\\ell=0\}^\{L\_\{q\}\}; total iterations

TMT\_\{\\mathrm\{M\}\}
2:Initialize

ℐ0\\mathcal\{I\}\_\{0\}and one starting state

P\(j\)∈ΩjP^\{\(j\)\}\\in\\Omega\_\{j\}for every

j∈ℐ0j\\in\\mathcal\{I\}\_\{0\}
3:for

t=1,…,TMt=1,\\ldots,T\_\{\\mathrm\{M\}\}do

4:Draw a target cell

Jt\|ℱt−1∼Uniform⁡\(ℐt−1\)J\_\{t\}\\mid\\mathcal\{F\}\_\{t\-1\}\\sim\\operatorname\{Uniform\}\(\\mathcal\{I\}\_\{t\-1\}\)
5:Draw a proposal type

Ht∼Categorical⁡\(ω0,…,ωLq\)H\_\{t\}\\sim\\operatorname\{Categorical\}\(\\omega\_\{0\},\\ldots,\\omega\_\{L\_\{q\}\}\)
6:Set

p←P\(Jt\)p\\leftarrow P^\{\(J\_\{t\}\)\}and draw a candidate

Pt′∼qHt\(⋅∣p\)P^\{\\prime\}\_\{t\}\\sim q\_\{H\_\{t\}\}\(\\cdot\\mid p\)
7:Set

ℐt←ℐt−1\\mathcal\{I\}\_\{t\}\\leftarrow\\mathcal\{I\}\_\{t\-1\}and leave the source chain

P\(Jt\)P^\{\(J\_\{t\}\)\}unchanged by default

8:if

val⁡\(Pt′\)=1\\operatorname\{val\}\(P^\{\\prime\}\_\{t\}\)=1then

9:

j′←c⁡\(Pt′\)j^\{\\prime\}\\leftarrow c\(P^\{\\prime\}\_\{t\}\)
10:if

j′=Jtj^\{\\prime\}=J\_\{t\}then

11:Draw

Ut∼Uniform⁡\(0,1\)U\_\{t\}\\sim\\operatorname\{Uniform\}\(0,1\)
12:if

Ut≤αJt,Ht​\(p,Pt′\)U\_\{t\}\\leq\\alpha\_\{J\_\{t\},H\_\{t\}\}\(p,P^\{\\prime\}\_\{t\}\)then

13:

P\(Jt\)←Pt′P^\{\(J\_\{t\}\)\}\\leftarrow P^\{\\prime\}\_\{t\}
14:endif

15:elseif

j′∉ℐt−1j^\{\\prime\}\\notin\\mathcal\{I\}\_\{t\-1\}then

16:Spawn a new chain

P\(j′\)←Pt′P^\{\(j^\{\\prime\}\)\}\\leftarrow P^\{\\prime\}\_\{t\}and expand the active set

ℐt←ℐt−1∪\{j′\}\\mathcal\{I\}\_\{t\}\\leftarrow\\mathcal\{I\}\_\{t\-1\}\\cup\\\{j^\{\\prime\}\\\}⊳\\trianglerighthindsight seed; not an accepted source move

17:endif

18:endif

19:Emit the post\-transition source state

Ptemit←P\(Jt\)P\_\{t\}^\{\\mathrm\{emit\}\}\\leftarrow P^\{\(J\_\{t\}\)\}
20:endfor

21:MCMC trace

\(P1emit,…,PTMemit\)\(P\_\{1\}^\{\\mathrm\{emit\}\},\\ldots,P\_\{T\_\{\\mathrm\{M\}\}\}^\{\\mathrm\{emit\}\}\), final active population, and a separately labeled discovery archive

##### Hindsight Mechanics and Rejection\.

A valid foreign\-cell proposal is automatically rejected by the source chain because its source\-cell target probability is zero\. If its destination cell was previously inactive, however, it serves as a valid initialization for a new chain\. This seed is not emitted immediately, nor is it recorded as an accepted MCMC transition for the source chain\. A global hindsight seed entering a new cellj′j^\{\\prime\}is distributed exactly according toπj′\\pi\_\{j^\{\\prime\}\}; a seed produced by a local\-edit proposal, however, may have an arbitrary initialization lawζj′\\zeta\_\{j^\{\\prime\}\}supported onΩj′\\Omega\_\{j^\{\\prime\}\}\.

Crucially, a proposal entering an already active foreign cell must not be submitted directly to the destination cell’s MH test, because the proposal was generated conditional on the source\-chain’s state rather than the destination\-chain’s state\. Such cross\-cell discoveries are instead retained in the separately labeled discovery archive\. Finally, every scheduled proposal attempt advances the selected chain by one local MCMC iteration, even upon rejection\. After a rejection, the emitted sample is the repeated current state, not the rejected candidate\. \(Discarding repeats or running until a fixed number of acceptances would improperly produce an accepted\-state jump chain, violating the target measure\.

##### Fixed\-Active\-Set MCMC Interpretation\.

For a temporarily fixed active\-cell setℐ\\mathcal\{I\}and a joint population state𝐩=\(pj\)j∈ℐ\\mathbf\{p\}=\(p\_\{j\}\)\_\{j\\in\\mathcal\{I\}\}, we can define the full random\-scan kernel as

𝖪ℐ​\(𝐩,d​𝐩′\)=1\|ℐ\|​∑j∈ℐKj​\(pj,d​pj′\)​δ𝐩−j​\(d​𝐩−j′\)\.\\mathsf\{K\}\_\{\\mathcal\{I\}\}\(\\mathbf\{p\},d\\mathbf\{p\}^\{\\prime\}\)=\\frac\{1\}\{\|\\mathcal\{I\}\|\}\\sum\_\{j\\in\\mathcal\{I\}\}K\_\{j\}\(p\_\{j\},dp^\{\\prime\}\_\{j\}\)\\delta\_\{\\mathbf\{p\}\_\{\-j\}\}\(d\\mathbf\{p\}^\{\\prime\}\_\{\-j\}\)\.\(32\)Because each localizedKjK\_\{j\}preservesπj\\pi\_\{j\}, the global kernel𝖪ℐ\\mathsf\{K\}\_\{\\mathcal\{I\}\}preserves the product targetΓℐ\\Gamma\_\{\\mathcal\{I\}\}in equation[24](https://arxiv.org/html/2609.30492#A3.E24)\. This maps to a standard random\-scan, component\-wise MCMC construction\([Johnson et al\., 2013](https://arxiv.org/html/2609.30492#bib.bib28)\): selecting a cell acts as the Gibbs coordinate step, andKjK\_\{j\}acts as the within\-coordinate MH update\. At stationarity, emitting the selected coordinate yields the exact marginal lawΠℐ\\Pi\_\{\\mathcal\{I\}\}\.

Hindsight spawning increases the number of active chains\. With frozen components, however, the augmented process consisting of the active setℐt\\mathcal\{I\}\_\{t\}and its chain states is time\-homogeneous on the disjoint union of population spaces\. Before complete discovery, it is not the fixed\-dimension random\-scan chain associated with a fixed active\-set product target\. The theoretical guarantees in the following subsection are specifically designed to hold for this expansive target distribution\.

#### C\.1\.4Theoretical Guarantees

##### Standing Assumptions\.

To guarantee valid Markovian dynamics, all representation, validity, embedding, whitening, partition, target, and proposal components must be selected during a baseline or pilot phase and entirely frozen before the reported run\. If the growing archive is allowed to change a proposal prompt, alter a validity rule, or update a target weight, the process becomes adaptive MCMC and; under such conditions, the fixed\-kernel results established below no longer apply without imposing additional conditions\([Roberts & Rosenthal, 2007](https://arxiv.org/html/2609.30492#bib.bib47)\)\.

###### Theorem 7\(Anytime Active\-Cell Uniformity\)\.

Letℱt−1\\mathcal\{F\}\_\{t\-1\}contain the complete history, the active set, and the individual chain states immediately prior to iterationtt\. For every initialization,t≥1t\\geq 1, andj∈\[M\]j\\in\[M\], the spatial allocation of the emitted persona is strictly uniform over the active set

Pr⁡\{c⁡\(Ptemit\)=j∣ℱt−1\}=𝟙\{j∈ℐt−1\}\|ℐt−1\|\.\\Pr\\\{c\(P\_\{t\}^\{\\mathrm\{emit\}\}\)=j\\mid\\mathcal\{F\}\_\{t\-1\}\\\}=\\frac\{\\mathbbm\{1\}\\\{j\\in\\mathcal\{I\}\_\{t\-1\}\\\}\}\{\|\\mathcal\{I\}\_\{t\-1\}\|\}\.\(33\)Moreover, define the predictable step exposureηj,t\\eta\_\{j,t\}, the cumulative realized countNj​\(TM\)N\_\{j\}\(T\_\{\\mathrm\{M\}\}\), and the cumulative expected exposureEj​\(TM\)E\_\{j\}\(T\_\{\\mathrm\{M\}\}\)as follows

ηj,t\\displaystyle\\eta\_\{j,t\}=𝟙\{j∈ℐt−1\}\|ℐt−1\|,\\displaystyle=\\frac\{\\mathbbm\{1\}\\\{j\\in\\mathcal\{I\}\_\{t\-1\}\\\}\}\{\|\\mathcal\{I\}\_\{t\-1\}\|\},Nj​\(TM\)\\displaystyle N\_\{j\}\(T\_\{\\mathrm\{M\}\}\)=∑t=1TM𝟙\{c\(Ptemit\)=j\},\\displaystyle=\\sum\_\{t=1\}^\{T\_\{\\mathrm\{M\}\}\}\\mathbbm\{1\}\\\{c\(P\_\{t\}^\{\\mathrm\{emit\}\}\)=j\\\},Ej​\(TM\)\\displaystyle E\_\{j\}\(T\_\{\\mathrm\{M\}\}\)=∑t=1TMηj,t\.\\displaystyle=\\sum\_\{t=1\}^\{T\_\{\\mathrm\{M\}\}\}\\eta\_\{j,t\}\.\(34\)ThenNj​\(TM\)−Ej​\(TM\)N\_\{j\}\(T\_\{\\mathrm\{M\}\}\)\-E\_\{j\}\(T\_\{\\mathrm\{M\}\}\)is a martingale\. For every fixed horizonTMT\_\{\\mathrm\{M\}\}andδ∈\(0,1\)\\delta\\in\(0,1\), the deviation is strictly bounded

Pr\{max1≤j≤M\|Nj\(TM\)−Ej\(TM\)\|≥TM2​log⁡2​Mδ\}≤δ\.\\Pr\\left\\\{\\max\_\{1\\leq j\\leq M\}\|N\_\{j\}\(T\_\{\\mathrm\{M\}\}\)\-E\_\{j\}\(T\_\{\\mathrm\{M\}\}\)\|\\geq\\sqrt\{\\frac\{T\_\{\\mathrm\{M\}\}\}\{2\}\\log\\frac\{2M\}\{\\delta\}\}\\right\\\}\\leq\\delta\.\(35\)

###### Proof\.

The scheduler drawsJtJ\_\{t\}uniformly fromℐt−1\\mathcal\{I\}\_\{t\-1\}\. Since the source chain either accepts a same\-cell state or defaults to a self\-transition, the relationc⁡\(Ptemit\)=Jtc\(P\_\{t\}^\{\\mathrm\{emit\}\}\)=J\_\{t\}almost surely, proving equation[33](https://arxiv.org/html/2609.30492#A3.E33)\. Consequently,𝔼\[𝟙\{c\(Ptemit\)=j\}∣ℱt−1\]=ηj,t\\mathbb\{E\}\[\\mathbbm\{1\}\\\{c\(P\_\{t\}^\{\\mathrm\{emit\}\}\)=j\\\}\\mid\\mathcal\{F\}\_\{t\-1\}\]=\\eta\_\{j,t\}, which gives the martingale\. Applying the bounded\-increment Hoeffding–Azuma inequality\([Azuma, 1967](https://arxiv.org/html/2609.30492#bib.bib3)\)to each cell and taking union\-bound over the fixedMMcells yields equation[35](https://arxiv.org/html/2609.30492#A3.E35)\. ∎

###### Theorem 8\(Eventual Discovery and Full\-Target Convergence\)\.

Letτj=inf\{t:j∈ℐt\}\\tau\_\{j\}=\\inf\\\{t:j\\in\\mathcal\{I\}\_\{t\}\\\}be the discovery time of a reachable cellj∉ℐ0j\\notin\\mathcal\{I\}\_\{0\}\. Under the matched global proposal and assumingω0\>0\\omega\_\{0\}\>0

Pr⁡\(τj\>TM\)\\displaystyle\\Pr\(\\tau\_\{j\}\>T\_\{\\mathrm\{M\}\}\)≤\(1−ω0​Bj\)TM,\\displaystyle\\leq\(1\-\\omega\_\{0\}B\_\{j\}\)^\{T\_\{\\mathrm\{M\}\}\},\(36\)Pr⁡\(ℐTM≠ℛ\)\\displaystyle\\Pr\(\\mathcal\{I\}\_\{T\_\{\\mathrm\{M\}\}\}\\neq\\mathcal\{R\}\)≤∑j∈ℛ∖ℐ0\(1−ω0​Bj\)TM,\\displaystyle\\leq\\sum\_\{j\\in\\mathcal\{R\}\\setminus\\mathcal\{I\}\_\{0\}\}\(1\-\\omega\_\{0\}B\_\{j\}\)^\{T\_\{\\mathrm\{M\}\}\},\(37\)𝔼​\|ℛ∖ℐTM\|\\displaystyle\\mathbb\{E\}\|\\mathcal\{R\}\\setminus\\mathcal\{I\}\_\{T\_\{\\mathrm\{M\}\}\}\|≤∑j∈ℛ∖ℐ0\(1−ω0​Bj\)TM\.\\displaystyle\\leq\\sum\_\{j\\in\\mathcal\{R\}\\setminus\\mathcal\{I\}\_\{0\}\}\(1\-\\omega\_\{0\}B\_\{j\}\)^\{T\_\{\\mathrm\{M\}\}\}\.\(38\)Sinceℛ\\mathcal\{R\}is finite, every reachable cell is discovered in finite time almost surely\. Letτ=inf\{t:ℐt=ℛ\}\\tau=\\inf\\\{t:\\mathcal\{I\}\_\{t\}=\\mathcal\{R\}\\\}andMℛ=\|ℛ\|M\_\{\\mathcal\{R\}\}=\|\\mathcal\{R\}\|\. Conditional on the population at timeτ\\tau, for every integers≥0s\\geq 0the subsequent fixed\-active\-set chain obeys

‖𝖪ℛs​\(𝐩,⋅\)−Γℛ‖TV≤min⁡\{1,∑j∈ℛ\(1−ω0​BjMℛ\)s\}\.\\left\\\|\\mathsf\{K\}\_\{\\mathcal\{R\}\}^\{\\,s\}\(\\mathbf\{p\},\\cdot\)\-\\Gamma\_\{\\mathcal\{R\}\}\\right\\\|\_\{\\mathrm\{TV\}\}\\leq\\min\\left\\\{1,\\,\\sum\_\{j\\in\\mathcal\{R\}\}\\left\(1\-\\frac\{\\omega\_\{0\}B\_\{j\}\}\{M\_\{\\mathcal\{R\}\}\}\\right\)^\{s\}\\right\\\}\.\(39\)In particular, the emitted\-persona law converges to the equal\-cell mixture

Πℛ=1\|ℛ\|​∑j∈ℛπj\.\\Pi\_\{\\mathcal\{R\}\}=\\frac\{1\}\{\|\\mathcal\{R\}\|\}\\sum\_\{j\\in\\mathcal\{R\}\}\\pi\_\{j\}\.\(40\)

###### Proof\.

At each global iteration, the global proposal component is selected with probabilityω0\\omega\_\{0\}and independently draws from the base lawbb\. It therefore proposes a valid persona in celljjwith probabilityω0​Bj\\omega\_\{0\}B\_\{j\}, independently of the selected source cell\. Repeated failure gives equation[36](https://arxiv.org/html/2609.30492#A3.E36)\. Then, applying a union bound and the linearity of expectation give equation[37](https://arxiv.org/html/2609.30492#A3.E37)and equation[38](https://arxiv.org/html/2609.30492#A3.E38)\. Note that local proposals can only discover a cell earlier\. The right\-hand side of equation[37](https://arxiv.org/html/2609.30492#A3.E37)tends to zero for finiteℛ\\mathcal\{R\}, proving almost\-sure eventual discovery\.

Following timeτ\\tau, coordinatejjis actively selected and undergoes an exactπj\\pi\_\{j\}refresh with probabilityω0​Bj/Mℛ\\omega\_\{0\}B\_\{j\}/M\_\{\\mathcal\{R\}\}per population iteration\. By coupling a chain initiated from𝐩\\mathbf\{p\}with a stationary chain using the same scan, proposal, and accept–reject randomness, coordinatejjcoalesces at its first exact refresh and remains coupled thereafter\. Union\-bounding the probability that any coordinate has not refreshed by timessdirectly proves equation[39](https://arxiv.org/html/2609.30492#A3.E39)\. To obtain convergence at deterministic times, letd⁡\(s\)d\(s\)denote the right\-hand side of equation[39](https://arxiv.org/html/2609.30492#A3.E39)\. The emission kernelQemitQ\_\{\\mathrm\{emit\}\}acting on the pre\-transition population satisfiesΓℛ​Qemit=Πℛ\\Gamma\_\{\\mathcal\{R\}\}Q\_\{\\mathrm\{emit\}\}=\\Pi\_\{\\mathcal\{R\}\}\. For any deterministic integer0≤r<t0\\leq r<t, conditioning on whether discovery is complete by timerrand applying total\-variation contraction gives

‖ℒ⁡\(Ptemit\)−Πℛ‖TV≤Pr⁡\(τ\>r\)\+d⁡\(t−r−1\)\.\\left\\\|\\mathcal\{L\}\(P\_\{t\}^\{\\mathrm\{emit\}\}\)\-\\Pi\_\{\\mathcal\{R\}\}\\right\\\|\_\{\\mathrm\{TV\}\}\\leq\\Pr\(\\tau\>r\)\+d\(t\-r\-1\)\.Takingr=⌊t/2⌋r=\\lfloor t/2\\rfloormakes both terms vanish ast→∞t\\to\\infty, proving equation[40](https://arxiv.org/html/2609.30492#A3.E40)\. ∎

### C\.2Evolutionary TextGrad Persona Generation

Evolutionary TextGrad implements frontier generation by expanding the baseline population toward valid, low\-density regions of persona space\. Unlike UC\-MCMC’s distributional target, this is an empirical search procedure; the resulting expanded pool is passed to the same max–min selector used for fixed\-pool selection\.

##### Objective and textual feedback\.

We formulate persona generation as evolutionary constrained optimization\. At iterationtt, let𝒫t\\mathcal\{P\}\_\{t\}be the current persona population\. For any personapp, whether an incumbent or a newly proposed candidate, the comparison set is𝒫t∖\{p\}\\mathcal\{P\}\_\{t\}\\setminus\\\{p\\\}; set subtraction has no effect whenp∉𝒫tp\\notin\\mathcal\{P\}\_\{t\}\. Write𝐞¯​\(p\)=ϕ⁡\(p\)/‖ϕ⁡\(p\)‖2\\overline\{\\mathbf\{e\}\}\(p\)=\\phi\(p\)/\\\|\\phi\(p\)\\\|\_\{2\}for the unit\-normalized persona embedding\. LetmE∈\{cos,2,Mah\}m\_\{\\mathrm\{E\}\}\\in\\\{\\mathrm\{cos\},2,\\mathrm\{Mah\}\\\}be the persona dissimilarity fixed for an evolutionary run; the reported run usesmE=Mahm\_\{\\mathrm\{E\}\}=\\mathrm\{Mah\}\. Define

Gapt⁡\(p\)\\displaystyle\\operatorname\{Gap\}\_\{t\}\(p\)=minp~∈𝒫t∖\{p\}⁡dmEP​\(p,p~\),\\displaystyle=\\min\_\{\\widetilde\{p\}\\in\\mathcal\{P\}\_\{t\}\\setminus\\\{p\\\}\}d\_\{m\_\{\\mathrm\{E\}\}\}^\{\\mathrm\{P\}\}\(p,\\widetilde\{p\}\),\(41\)ρ^t​\(p\)\\displaystyle\\widehat\{\\rho\}\_\{t\}\(p\)=1\|𝒫t∖\{p\}\|​∑p~∈𝒫t∖\{p\}exp⁡\(βKDE​𝐞¯​\(p\)⊤​𝐞¯​\(p~\)\),\\displaystyle=\\frac\{1\}\{\|\\mathcal\{P\}\_\{t\}\\setminus\\\{p\\\}\|\}\\sum\_\{\\widetilde\{p\}\\in\\mathcal\{P\}\_\{t\}\\setminus\\\{p\\\}\}\\exp\\\!\\left\(\\beta\_\{\\mathrm\{KDE\}\}\\,\\overline\{\\mathbf\{e\}\}\(p\)^\{\\top\}\\overline\{\\mathbf\{e\}\}\(\\widetilde\{p\}\)\\right\),\(42\)Fitt⁡\(p\)\\displaystyle\\operatorname\{Fit\}\_\{t\}\(p\)=λrel​Val⁡\(p\)\+λgap​Gapt⁡\(p\)−λden​log⁡ρ^t​\(p\)\.\\displaystyle=\\lambda\_\{\\mathrm\{rel\}\}\\operatorname\{Val\}\(p\)\+\\lambda\_\{\\mathrm\{gap\}\}\\operatorname\{Gap\}\_\{t\}\(p\)\-\\lambda\_\{\\mathrm\{den\}\}\\log\\widehat\{\\rho\}\_\{t\}\(p\)\.\(43\)HereVal⁡\(p\)∈\[0,1\]\\operatorname\{Val\}\(p\)\\in\[0,1\]measures satisfaction of the persona schema and task\-independent validity constraints,Gapt⁡\(p\)\\operatorname\{Gap\}\_\{t\}\(p\)is nearest\-neighbor novelty under the samedmEPd\_\{m\_\{\\mathrm\{E\}\}\}^\{\\mathrm\{P\}\}used for downstream persona selection, andρ^t​\(p\)\\widehat\{\\rho\}\_\{t\}\(p\)is an unnormalized von Mises\-Fisher kernel\-density score in persona embedding space\. Parent\-path comparisons use an identical reference set for the parent and offspring\. Thus, extend the definitions asGapt⁡\(x,R\)\\operatorname\{Gap\}\_\{t\}\(x;R\)andρ^t​\(x,R\)\\widehat\{\\rho\}\_\{t\}\(x;R\)for a nonempty reference setRR, whereR=𝒫t∖\{x\}R=\\mathcal\{P\}\_\{t\}\\setminus\\\{x\\\}without explicit reference set\.

Because personas are discrete text rather than differentiable vectors, we use*textual gradient*only for natural\-language feedback, following[Yuksekgonul et al\. \(2025\)](https://arxiv.org/html/2609.30492#bib.bib58)

gttext​\(p\)=TextGrad⁡\(p,Val⁡\(p\),Gapt⁡\(p\),ρ^t​\(p\),xgrad\),g\_\{t\}^\{\\mathrm\{text\}\}\(p\)=\\operatorname\{TextGrad\}\\\!\\left\(p;\\operatorname\{Val\}\(p\),\\operatorname\{Gap\}\_\{t\}\(p\),\\widehat\{\\rho\}\_\{t\}\(p\),x\_\{\\mathrm\{grad\}\}\\right\),\(44\)wherexgradx\_\{\\mathrm\{grad\}\}is the fixed feedback instruction\. The editor then proposes

p′=TGD\.step⁡\(p,gttext​\(p\)\)\.p^\{\\prime\}=\\operatorname\{TGD\.step\}\\\!\\left\(p,g\_\{t\}^\{\\mathrm\{text\}\}\(p\)\\right\)\.\(45\)This notation avoids treating LLM feedback as a literal derivative with respect to a token sequence\.

##### Evolutionary algorithm\.

The loop grows the population by mutating incumbents under textual\-gradient feedback and admitting only offspring that demonstrably increase spread\. Two design choices matter\. First, selection and acceptance operate on the two diversity axes*separately*rather than on the scalar equation[43](https://arxiv.org/html/2609.30492#A3.E43): a scalar trade\-off would let a large novelty gain purchase an equally large density loss, which is exactly the substitution that produces a few far\-flung outliers around an unchanged core\. We therefore reportFitt\\operatorname\{Fit\}\_\{t\}as a summary statistic, while the loop itself is weight\-free\. Second, every admission score is evaluated against the frozen population𝒫t−1\\mathcal\{P\}\_\{t\-1\}, so individual scores do not depend on the order in which candidates are evaluated\.

##### Parent selection\.

Each tournament drawsmtour=5m\_\{\\mathrm\{tour\}\}=5incumbents uniformly without replacement and admits two winners: the entrant with the largestGapt−1\\operatorname\{Gap\}\_\{t\-1\}and, when distinct, the entrant with the smallestlog⁡ρ^t−1\\log\\widehat\{\\rho\}\_\{t\-1\}\. Tournaments repeat until the parent set reachesβpar​\|𝒫t−1\|\\beta\_\{\\mathrm\{par\}\}\|\\mathcal\{P\}\_\{t\-1\}\|withβpar=0\.05\\beta\_\{\\mathrm\{par\}\}=0\.05, so the batch grows with the population\. Admitting one winner per axis applies equal pressure to isolation and to sparsity without committing to an exchange rate between them\.

##### Admission\.

Letppbe a parent andp′p^\{\\prime\}its offspring, and let

mgap=cm​sd​\{Gapt−1⁡\(q\)\}q∈𝒫t−1,mden=cm​sd​\{log⁡ρ^t−1​\(q\)\}q∈𝒫t−1,m\_\{\\mathrm\{gap\}\}=c\_\{m\}\\operatorname\{sd\}\\\{\\operatorname\{Gap\}\_\{t\-1\}\(q\)\\\}\_\{q\\in\\mathcal\{P\}\_\{t\-1\}\},\\qquad m\_\{\\mathrm\{den\}\}=c\_\{m\}\\operatorname\{sd\}\\\{\\log\\widehat\{\\rho\}\_\{t\-1\}\(q\)\\\}\_\{q\\in\\mathcal\{P\}\_\{t\-1\}\},\(46\)withcm=0\.10c\_\{m\}=0\.10, be per\-axis margins recomputed each iteration\. The offspring is admitted when it passes the deterministic and relevance gates and satisfies either

\(parent path\)\\displaystyle\\text\{\(parent path\)\}Gapt−1⁡\(p′;Rp\)\>Gapt−1⁡\(p;Rp\)\+mgap\\displaystyle\\operatorname\{Gap\}\_\{t\-1\}\(p^\{\\prime\};R\_\{p\}\)\>\\operatorname\{Gap\}\_\{t\-1\}\(p;R\_\{p\}\)\+m\_\{\\mathrm\{gap\}\}\(47\)log⁡ρ^t−1​\(p′,Rp\)<log⁡ρ^t−1​\(p,Rp\)−mden,\\displaystyle\\log\\widehat\{\\rho\}\_\{t\-1\}\(p^\{\\prime\};R\_\{p\}\)<\\log\\widehat\{\\rho\}\_\{t\-1\}\(p;R\_\{p\}\)\-m\_\{\\mathrm\{den\}\},\(absolute path\)\\displaystyle\\text\{\(absolute path\)\}Gapt−1⁡\(p′\)≥Qτ\\displaystyle\\operatorname\{Gap\}\_\{t\-1\}\(p^\{\\prime\}\)\\geq Q\_\{\\tau\}\(48\)log⁡ρ^t−1​\(p′\)≤Q100−τ′,\\displaystyle\\log\\widehat\{\\rho\}\_\{t\-1\}\(p^\{\\prime\}\)\\leq Q\_\{100\-\\tau\}^\{\\prime\},whereQτQ\_\{\\tau\}andQ100−τ′Q^\{\\prime\}\_\{100\-\\tau\}are theτ\\tau\-th and\(100−τ\)\(100\-\\tau\)\-th population percentiles of the two axes,τ=75\\tau=75\. The parent path requires improvement on both axes by explicit margins, discouraging admission based on negligible score differences\. The absolute path admits a candidate that fails to beat its parent but is nonetheless competitive with the population, which prevents an already\-isolated parent from blocking progress\. Scoring is deliberately asymmetric between the two: the parent path scoresp′p^\{\\prime\}againstRp=𝒫t−1∖\{p\}R\_\{p\}=\\mathcal\{P\}\_\{t\-1\}\\setminus\\\{p\\\}so that parent and offspring are compared leave\-one\-out on identical reference sets, whereas the absolute path scores against the full population so that a near\-copy of an elite parent cannot inherit that parent’s isolation\.

##### Sibling deduplication\.

Because offspring are scored against a frozen population, their admission scores do no account for other offspring of the same iteration\. When two admitted offspring are closer thanmgapm\_\{\\mathrm\{gap\}\}, we treat them as near\-duplicates and retain the one with the larger full\-population gap\. This step provides a heuristic control on redundancy within the admitted batch\.

Algorithm 4EvolutionaryTextGradPersona Generation1:Baseline population

𝒫0\\mathcal\{P\}\_\{0\}; iteration cap

TET\_\{\\mathrm\{E\}\}; population cap

NmaxN\_\{\\max\}; batch fraction

βpar\\beta\_\{\\mathrm\{par\}\}; tournament size

mtourm\_\{\\mathrm\{tour\}\}; margin fraction

cmc\_\{m\}; percentile

τ\\tau; frozen feedback instruction

xgradx\_\{\\mathrm\{grad\}\}
2:for

t=1,…,TEt=1,\\ldots,T\_\{\\mathrm\{E\}\}do

3:Compute leave\-one\-out

Gapt−1⁡\(q\)\\operatorname\{Gap\}\_\{t\-1\}\(q\)and

log⁡ρ^t−1​\(q\)\\log\\widehat\{\\rho\}\_\{t\-1\}\(q\)for every

q∈𝒫t−1q\\in\\mathcal\{P\}\_\{t\-1\}
4:Set the margins equation[46](https://arxiv.org/html/2609.30492#A3.E46)and the percentile thresholds

Qτ,Q100−τ′Q\_\{\\tau\},Q^\{\\prime\}\_\{100\-\\tau\}
5:

ℬt←∅\\mathcal\{B\}\_\{t\}\\leftarrow\\emptyset
6:while

\|ℬt\|<βpar​\|𝒫t−1\|\|\\mathcal\{B\}\_\{t\}\|<\\beta\_\{\\mathrm\{par\}\}\|\\mathcal\{P\}\_\{t\-1\}\|do

7:Sample

mtourm\_\{\\mathrm\{tour\}\}incumbents uniformly without replacement

8:Add the largest\-

Gap\\operatorname\{Gap\}entrant and, if distinct, the smallest\-

log⁡ρ^\\log\\widehat\{\\rho\}entrant

9:endwhile

10:

𝒜t←∅\\mathcal\{A\}\_\{t\}\\leftarrow\\emptyset
11:for

p∈ℬtp\\in\\mathcal\{B\}\_\{t\}do

12:

p′←TGD\.step⁡\(p,gt−1text​\(p\)\)p^\{\\prime\}\\leftarrow\\operatorname\{TGD\.step\}\(p,g\_\{t\-1\}^\{\\mathrm\{text\}\}\(p\)\)⊳\\trianglerightequation[44](https://arxiv.org/html/2609.30492#A3.E44)–equation[45](https://arxiv.org/html/2609.30492#A3.E45)

13:if

p′p^\{\\prime\}fails the English or length gateor

Val⁡\(p′\)≠1\\operatorname\{Val\}\(p^\{\\prime\}\)\\neq 1then

14:reject⊳\\trianglerightgates precede embedding and judging

15:elseifequation[47](https://arxiv.org/html/2609.30492#A3.E47)orequation[48](https://arxiv.org/html/2609.30492#A3.E48)holdsthen

16:

𝒜t←𝒜t∪\{p′\}\\mathcal\{A\}\_\{t\}\\leftarrow\\mathcal\{A\}\_\{t\}\\cup\\\{p^\{\\prime\}\\\}
17:endif

18:endfor

19:Remove from

𝒜t\\mathcal\{A\}\_\{t\}the smaller\-gap member of every pair within

mgapm\_\{\\mathrm\{gap\}\}⊳\\trianglerightsibling deduplication

20:

𝒫t←𝒫t−1∪𝒜t\\mathcal\{P\}\_\{t\}\\leftarrow\\mathcal\{P\}\_\{t\-1\}\\cup\\mathcal\{A\}\_\{t\}
21:if

\|𝒫t\|≥Nmax\|\\mathcal\{P\}\_\{t\}\|\\geq N\_\{\\max\}then

22:break

23:endif

24:endfor

25:Evolved pool

𝒫TE\\mathcal\{P\}\_\{T\_\{\\mathrm\{E\}\}\}; the max–min selector of[Section3\.2](https://arxiv.org/html/2609.30492#S3.SS2)is applied downstream to return

𝒮disp,mE⋆​\(𝒫TE\)\\mathcal\{S\}^\{\\star\}\_\{\\mathrm\{disp\},m\_\{\\mathrm\{E\}\}\}\(\\mathcal\{P\}\_\{T\_\{\\mathrm\{E\}\}\}\)

The population is grown by union and never shrinks, so coverage of the persona space is monotone intt\. Selection by max–min dispersion is deliberately kept downstream of the loop rather than folded into it: holding the selection operator fixed and varying only the candidate pool is what isolates the effect of generation from the effect of selection\.

## Appendix DExperiment Details

The number of personask=5k=5: The Vendi score using cosine Gram matrix is measured as10\.410\.4on the baseline personas\. This translates to an effective number of distinct personas in the baseline pool is1010, and any numberk≤10k\\leq 10should be able to accommodatekksemantically distinct personas\. We chosek=5k=5for convenience and the feasibility of human clustering and judgment with multiple variations and baselines\.

Performance tables report means over five response\-generation seeds \(42–46\) and two\-sided 95% Student’sttconfidence intervals\. For seed\-level metric valuesx1,…,x5x\_\{1\},\\ldots,x\_\{5\}, computed using the benchmark\-specific aggregation, we report

x¯±t0\.975,4​s5,s2=14​∑r=15\(xr−x¯\)2,\\overline\{x\}\\pm t\_\{0\.975,4\}\\frac\{s\}\{\\sqrt\{5\}\},\\qquad s^\{2\}=\\frac\{1\}\{4\}\\sum\_\{r=1\}^\{5\}\(x\_\{r\}\-\\overline\{x\}\)^\{2\},wheret0\.975,4≈2\.776t\_\{0\.975,4\}\\approx 2\.776\. These intervals summarize variability across response\-generation seeds for the evaluated persona sets\. ICC intervals in the judge\-validation analysis use bootstrap resampling instead\.

### D\.1Common Baselines and Variants

##### Persona control baselines\.

Gibberishacts as a length\-matched pseudoword control to verify that performance gains stem from semantic content rather than mere prompt length\. Prompted with Jabberwocky personas, LLMs generate responses containing pseudowords \(“shield from*nasnorn*rain” or “press flowai glurt sheevu”\)\. Such out\-of\-vocabulary tokens embed in low\-density regions, inflating every distance\-based metric equation[57](https://arxiv.org/html/2609.30492#A4.E57)without any corresponding remoteness of the underlying idea\. We therefore delete them from each use before computing embeddings\. The pseudowords originates from gibberish input: the gibberish generation process replaces content words one\-for\-one in place from an existing persona, so aligning a persona with its gibberish counterpart recovers a list fo non\-words\. A response word is deleted if it matches one of them, ignoring case and inflectional endings\. The responses are evaluated after deletion \(“shield from rain” or “press”\)\.

Task\-conditionedpersona generation\([Jin et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib27)\)\(MPAQ\) serves as a baseline that explicitly generates task\-specific personas\. Task\-specific generation inherently incurs expensive, per\-task inference costs; by contrast, our generic diversified personas incur only a one\-time optimization cost and can be deployed universally across any task\. Due to this overhead, task\-conditioned personas are not useful unless they are better for creative responses than task\-agnostic personas\.

##### Baseline persona pool\.

We utilize the PersonaMem\-v2 dataset\([Jiang et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib25);[Ge et al\., 2024](https://arxiv.org/html/2609.30492#bib.bib16)\)as our fixed candidate pool\. To extractk=5k=5personas from this massive population, we evaluate two standard baselines:Random\(uniform sampling\) to test arbitrary individualization, andTypical\(the55densest\-mode personas via Gaussian KDE\) to test representative personas\.

##### Diversity\-selected personas \(ours\)\.

Selectionextractsk=5k\{=\}5personas from the baseline pool by maximizing min dispersion \([Section3\.2](https://arxiv.org/html/2609.30492#S3.SS2)\) under the Mahalanobis metric\. Contrasting this withRandomandTypicalisolates the effect of*diverse selection*over standard individualization\.

##### Generated personas \(ours\)\.

To break the performance ceiling imposed by a finite pool, we activaly expand the persona population using two algorithms prior to applying the identical max–min dispersion selector\.MCMCacts as a space\-filling generator targeting an equal\-cell mixture over reachable semantic regions \([Section3\.3](https://arxiv.org/html/2609.30492#S3.SS3)\)\.Evolutionactively push personas toward low\-density extremes by applying evolutionaryTextGrad\([Section3\.4](https://arxiv.org/html/2609.30492#S3.SS4)\)\.

### D\.2Alternative Uses Test \(AUT\)

The Alternative Uses Test \(AUT\) is a widely used divergent\-thinking benchmark\([Guilford, 1967](https://arxiv.org/html/2609.30492#bib.bib19)\)\. Given a common object \(e\.g\., a fork, book, or wallet\), respondents produce alternative uses\. Its open\-ended, object\-conditioned form has become a standard testbed for LLM creativity\([Stevenson et al\., 2022](https://arxiv.org/html/2609.30492#bib.bib53);[Góes et al\., 2023](https://arxiv.org/html/2609.30492#bib.bib18);[Rabeyah et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib46)\)and is well suited to testing whether persona diversification broadens ideational range\. For objectooand experimental conditionhh, let𝒰~o,h\\widetilde\{\\mathcal\{U\}\}\_\{o,h\}denote the raw multiset ofnraw=25n\_\{\\mathrm\{raw\}\}=25generated uses and let𝒰o,h\\mathcal\{U\}\_\{o,h\}denote the canonicalized, deduplicated, valid set used by downstream metrics\. Its actual sizeno,h=\|𝒰o,h\|n\_\{o,h\}=\|\\mathcal\{U\}\_\{o,h\}\|can therefore be smaller than2525\.

#### D\.2\.1Baselines and Variants

We compare a graded ladder of prompting and persona conditions, each adding one ingredient over the last so that its marginal effect is isolated\. All conditions share the generation protocol below; they differ only in the prompt and in how \(or whether\) a persona is supplied\.

##### Standard persona baselines\.

Three persona\-free prompts from the ladder of[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18)isolate what prompting alone achieves:Common\-Use\(their*nn*prompt, eliciting conventional uses\) as a lower anchor;Alternative\-Use\(their*nc*“creative” prompt\); andCreativity\-Enhanced\(their*bs*expert prompt with an explicit definition of creative use\)\.Alternative\-Useis the baseline for persona injection: every persona condition below supplies a persona to that same*nc*prompt, so any gain over a standard, non\-persona prompt is attributable to the persona, not to prompt optimization\.

##### Evolutionary personas \(ours\)\.

AUT\-Evolutionproduces the pool of personas by the same evolutionary generation algorithm \([Algorithm4](https://arxiv.org/html/2609.30492#alg4)\), but with AUT\-specific fitness function\. Fix the response generator and the benchmark prompt\. For personappand objectoo, let𝒰o​\(p\)\\mathcal\{U\}\_\{o\}\(p\)denote the valid uses thatppproduces foroounder the protocol of[SectionD\.2](https://arxiv.org/html/2609.30492#A4.SS2), and letmRm\_\{\\mathrm\{R\}\}be the fixed response dissimilarity\. A persona is scored on four axes, each averaged over the evaluated objects,

Utilo⁡\(p\)\\displaystyle\\operatorname\{Util\}\_\{o\}\(p\)=1\|𝒰o​\(p\)\|​∑u∈𝒰o​\(p\)so​\(u\),\\displaystyle=\\frac\{1\}\{\|\\mathcal\{U\}\_\{o\}\(p\)\|\}\\sum\_\{u\\in\\mathcal\{U\}\_\{o\}\(p\)\}s\_\{o\}\(u\),\(49\)Novo⁡\(p\)\\displaystyle\\operatorname\{Nov\}\_\{o\}\(p\)=1\|𝒰o​\(p\)\|​∑u∈𝒰o​\(p\)minv∈𝒰ocomm⁡dmRR​\(u,v\),\\displaystyle=\\frac\{1\}\{\|\\mathcal\{U\}\_\{o\}\(p\)\|\}\\sum\_\{u\\in\\mathcal\{U\}\_\{o\}\(p\)\}\\min\_\{v\\in\\mathcal\{U\}\_\{o\}^\{\\mathrm\{comm\}\}\}d\_\{m\_\{\\mathrm\{R\}\}\}^\{\\mathrm\{R\}\}\(u,v\),\(50\)Divo⁡\(p\)\\displaystyle\\operatorname\{Div\}\_\{o\}\(p\)=\(\|𝒰o​\(p\)\|2\)−1​∑\{u,v\}⊆𝒰o​\(p\)dmRR​\(u,v\),\\displaystyle=\\binom\{\|\\mathcal\{U\}\_\{o\}\(p\)\|\}\{2\}^\{\-1\}\\sum\_\{\\\{u,v\\\}\\subseteq\\mathcal\{U\}\_\{o\}\(p\)\}d\_\{m\_\{\\mathrm\{R\}\}\}^\{\\mathrm\{R\}\}\(u,v\),\(51\)Flexo⁡\(p\)\\displaystyle\\operatorname\{Flex\}\_\{o\}\(p\)=VS⁡\(𝒰o​\(p\)\),\\displaystyle=\\operatorname\{VS\}\(\\mathcal\{U\}\_\{o\}\(p\)\),\(52\)so that novelty is a directed Chamfer distance to the frozen common\-use reference set, diversity is within\-persona spread, and flexibility is the effective number of distinct use categories equation[53](https://arxiv.org/html/2609.30492#A4.E53)\. The utility strengthso​\(u\)s\_\{o\}\(u\)is obtained by judging each use pairwise againstK=3K=3strength\-stratified anchors drawn from a frozen per\-object ladder and fitting a one\-dimensional Bradley–Terry model against the fixed anchor strengths; each ladder is sum\-zero identified, soUtilo\\operatorname\{Util\}\_\{o\}reads as an average advantage over baseline utility and is not comparable across objects\. Validity is a gate rather than an axis: an offspring is discarded unless every retained use passes the parse, length, language, and semantic\-admissibility checks of[SectionD\.2](https://arxiv.org/html/2609.30492#A4.SS2)\.

WriteA⁡\(p\)=\(Util,Nov,Div,Flex\)​\(p\)∈ℝ4A\(p\)=\(\\operatorname\{Util\},\\operatorname\{Nov\},\\operatorname\{Div\},\\operatorname\{Flex\}\)\(p\)\\in\\mathbb\{R\}^\{4\}for the object\-averaged axis vector\. Admission replaces equation[47](https://arxiv.org/html/2609.30492#A3.E47)–equation[48](https://arxiv.org/html/2609.30492#A3.E48)by a majority rule on the four axes: the offspring is accepted when it strictly beats its parent on at leastw=3w=3axes, or matches theτ\\tau\-th percentile member of the population on at leastwwaxes\. Parent selection uses the same tournament, with the deciding axis rotated through a random permutation of the four so that each receives equal pressure over an iteration\. Because all four axes are frozen within a metric epoch, the common\-use pool, the per\-object covariances, and the anchor ladder are snapshotted at run start, a later re\-evaluation of the benchmark cannot retroactively move the objective a run was optimized against\.

Two properties of this substitution are worth stating\. It is strictly more expensive, since each candidate must generate and have judged a full set of responses before it can be scored, whereas the persona\-text fitness is computed from embeddings alone\. And it is task\-bound: the resulting pool is optimized for the AUT and carries no guarantee of transfer, which is precisely the trade the task\-agnostic fitness avoids\. We report both pools and compare them directly\.

##### Human reference\.

For the three objects with released human norms \(book, fork, tin can\), we include the[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib53)human responses as a reference for the judge\-scored dimensions\. We samplencomp=5n\_\{\\mathrm\{comp\}\}=5human respondents, matching the number of independent LLM completions per condition\.

##### Generation and sampling protocol\.

All responses are generated by Gemma\-4\-31B \(temperature0\.70\.7\) under a uniform output constraint, five words per use, no adjectives, and an unbounded list, so conditions differ only in prompt and persona\. The baseline population is𝒫0=\{pi\}i=1N0\\mathcal\{P\}\_\{0\}=\\\{p\_\{i\}\\\}\_\{i=1\}^\{N\_\{0\}\}from PersonaMem\([Jiang et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib25)\), withN0=997N\_\{0\}=997after excluding malformed or missing entries\. Ifxox\_\{o\}is the prompt instantiated for objectoo, a persona\-conditioned response followsy∼GθR\(⋅∣xo,p\)y\\sim G\_\{\\theta\_\{\\mathrm\{R\}\}\}\(\\,\\cdot\\mid x\_\{o\},p\)withp∈Ω∅:=Ω∪\{∅\}p\\in\\Omega\_\{\\varnothing\}:=\\Omega\\cup\\\{\\varnothing\\\}; persona\-free conditions use the conditioning sentinelp=∅p=\\varnothing, which is distinct from the failed persona\-generation symbol⊥\\bot\. Non\-persona conditions usencomp=5n\_\{\\mathrm\{comp\}\}=5independent completions per object, while persona conditions use one completion for each of thek=5k=5selected personas\. We retain the firstnuse=5n\_\{\\mathrm\{use\}\}=5uses from each completion, givingnraw=ncomp​nuse=25n\_\{\\mathrm\{raw\}\}=n\_\{\\mathrm\{comp\}\}n\_\{\\mathrm\{use\}\}=25raw uses per object–condition pair before the validity gate\.

VariantPersonaDescriptionCommon\-Use–[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18)common\-use \(nn\)Alternative\-Use–[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18)creative \(nc\)Creativity\-Enhanced–[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18)expert \(bs\)Random✓55chosen uniformly at randomTypical✓55density\-ranked modesSelection\(ours\)✓maximize dispersion on base poolEvolution\(ours\)✓maximize dispersion on evolved poolHuman–55chosen from[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib53)Table 4:Variants, from persona\-free baselines to our diversity\-aware selection and generation\. Persona variants all inject into theAlternative\-Useprompt;SelectionandEvolutionshare one selection operator and differ only in the candidate pool\.

#### D\.2\.2Metrics

##### Notation\.

Let𝒴⊆𝒯\\mathcal\{Y\}\\subseteq\\mathcal\{T\}denote the response\-text space\. The fixed encoderϕ:𝒯→ℝd\\phi:\\mathcal\{T\}\\to\\mathbb\{R\}^\{d\}maps any persona, prompt, object label, or response text to an embedding; write𝐞⁡\(⋅\)=ϕ⁡\(⋅\)\\mathbf\{e\}\(\\cdot\)=\\phi\(\\cdot\)\. The regularized covarianceΣ^R\\widehat\{\\Sigma\}\_\{\\mathrm\{R\}\}is estimated once on a frozen reference split and is not refit by condition\. We calldcosRd\_\{\\mathrm\{cos\}\}^\{\\mathrm\{R\}\}a dissimilarity because it need not satisfy the triangle inequality\. Let𝒰ocomm\\mathcal\{U\}\_\{o\}^\{\\mathrm\{comm\}\}be the frozen common\-use reference set for objectoo, drawn from no\-persona baseline generations\([Góes et al\., 2023](https://arxiv.org/html/2609.30492#bib.bib18)\)\.

For a valid response set𝒰=\{ui\}i=1n\\mathcal\{U\}=\\\{u\_\{i\}\\\}\_\{i=1\}^\{n\}, let𝐞¯i=𝐞⁡\(ui\)/‖𝐞⁡\(ui\)‖2\\overline\{\\mathbf\{e\}\}\_\{i\}=\\mathbf\{e\}\(u\_\{i\}\)/\\\|\\mathbf\{e\}\(u\_\{i\}\)\\\|\_\{2\}, let𝐄𝒰\\mathbf\{E\}\_\{\\mathcal\{U\}\}stack these unit vectors by row, and let𝐆𝒰=𝐄𝒰​𝐄𝒰⊤\\mathbf\{G\}\_\{\\mathcal\{U\}\}=\\mathbf\{E\}\_\{\\mathcal\{U\}\}\\mathbf\{E\}\_\{\\mathcal\{U\}\}^\{\\top\}be their cosine Gram matrix\. Ifξ1,…,ξn\\xi\_\{1\},\\ldots,\\xi\_\{n\}are the eigenvalues of𝐆𝒰/n\\mathbf\{G\}\_\{\\mathcal\{U\}\}/n, the Vendi Score is

VS\(𝒰\)=exp\(−∑i=1nξilogξi\),\\operatorname\{VS\}\(\\mathcal\{U\}\)=\\exp\\\!\\left\(\-\\sum\_\{i=1\}^\{n\}\\xi\_\{i\}\\log\\xi\_\{i\}\\right\),\(53\)with0​log⁡0=00\\log 0=0\. It is the effective number of distinct responses\([Friedman & Dieng, 2023](https://arxiv.org/html/2609.30492#bib.bib15);[Pasarkar & Dieng, 2024](https://arxiv.org/html/2609.30492#bib.bib45)\)\.

##### Validity\.

We map each raw use through a deterministic canonicalizercan⁡\(⋅\)\\operatorname\{can\}\(\\cdot\)\(lowercasing, filler and punctuation removal, and lemmatization\)\. A separate, human\-validated LLM judge supplies the admissibility indicatorvalR:𝒴→\{0,1\}\\operatorname\{val\}\_\{\\mathrm\{R\}\}:\\mathcal\{Y\}\\rightarrow\\\{0,1\\\}\. For𝒰~=\{u~i\}i=1nraw\\widetilde\{\\mathcal\{U\}\}=\\\{\\widetilde\{u\}\_\{i\}\\\}\_\{i=1\}^\{n\_\{\\mathrm\{raw\}\}\},

Validity⁡\(𝒰~\)=1nraw​∑i=1nrawvalR⁡\(can⁡\(u~i\)\)\.\\operatorname\{Validity\}\(\\widetilde\{\\mathcal\{U\}\}\)=\\frac\{1\}\{n\_\{\\mathrm\{raw\}\}\}\\sum\_\{i=1\}^\{n\_\{\\mathrm\{raw\}\}\}\\operatorname\{val\}\_\{\\mathrm\{R\}\}\\\!\\left\(\\operatorname\{can\}\(\\widetilde\{u\}\_\{i\}\)\\right\)\.\(54\)The downstream set is

𝒰=\{can\(u~\):u~∈𝒰~,valR\(can\(u~\)\)=1\},\\mathcal\{U\}=\\left\\\{\\operatorname\{can\}\(\\widetilde\{u\}\):\\widetilde\{u\}\\in\\widetilde\{\\mathcal\{U\}\},\\ \\operatorname\{val\}\_\{\\mathrm\{R\}\}\(\\operatorname\{can\}\(\\widetilde\{u\}\)\)=1\\right\\\},where set semantics merge exact canonical duplicates\. Thus validity uses the raw denominator, whereas all following metrics use the post\-gate denominatorn=\|𝒰\|n=\|\\mathcal\{U\}\|\. Gating prevents invalid outliers from receiving spuriously high originality scores\([Chen & Ding, 2023](https://arxiv.org/html/2609.30492#bib.bib7)\)\.

##### Diversity\.

The first diversity metric is the mean pairwise response dissimilarity \(defined forn≥2n\\geq 2\)

Divm⁡\(𝒰\)=2n⁡\(n−1\)​∑i<jdmR​\(ui,uj\)\.\\operatorname\{Div\}\_\{m\}\(\\mathcal\{U\}\)=\\frac\{2\}\{n\(n\-1\)\}\\sum\_\{i<j\}d\_\{m\}^\{\\mathrm\{R\}\}\(u\_\{i\},u\_\{j\}\)\.\(55\)
Another diversity metric is the volume of convex hull on reduced PCA dimensions\. Letem​\(u\)∈ℝDme\_\{m\}\(u\)\\in\\mathbb\{R\}^\{D\_\{m\}\}denote the response embedding realized in spacemm: the L2\-normalized embedding form=cosm=\\mathrm\{cos\}, the raw embedding form=2m=2, and the whitened truncated embeddingLo⊤​e~​\(u\)L\_\{o\}^\{\\top\}\\tilde\{e\}\(u\)form=Mahm=\\mathrm\{Mah\}, whereΣo−1=Lo​Lo⊤\\Sigma\_\{o\}^\{\-1\}=L\_\{o\}L\_\{o\}^\{\\top\}is the Cholesky factorization of the object\-specific inverse covariance\. LetΠm,o:ℝDm→ℝk\\Pi\_\{m,o\}\\colon\\mathbb\{R\}^\{D\_\{m\}\}\\to\\mathbb\{R\}^\{k\}\(k=5k=5\) be the PCA projection fit on the pooled embeddingsem​\(u\)\{e\_\{m\}\(u\)\}of all variants’ responses for objectoo\. Then the volume is defined forn≥k\+1n\\geq k\+1as

Volm⁡\(𝒰\)=volk⁡\(conv⁡\{Πm,o​em​\(u1\),…,Πm,o​em​\(un\)\}\),\\operatorname\{Vol\}\_\{m\}\(\\mathcal\{U\}\)=\\operatorname\{vol\}\_\{k\}\\\!\\Big\(\\operatorname\{conv\}\\big\\\{\\Pi\_\{m,o\}\\,e\_\{m\}\(u\_\{1\}\),\\ldots,\\Pi\_\{m,o\}\\,e\_\{m\}\(u\_\{n\}\)\\big\\\}\\Big\),\(56\)whereconv⁡\(⋅\)\\operatorname\{conv\}\(\\cdot\)is the convex hull andvolk\\operatorname\{vol\}\_\{k\}thekk\-dimensional Lebesgue volume\.

##### Originality\.

We reserve*originality*for reference\-relative response distance, distinguishing it from the persona novelty gap in equation[41](https://arxiv.org/html/2609.30492#A3.E41)\. The object\-relative and common\-use\-relative versions are

Origm⁡\(𝒰,o\)\\displaystyle\\operatorname\{Orig\}\_\{m\}\(\\mathcal\{U\};o\)=1n​∑i=1ndmR​\(ui,o\),\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}d\_\{m\}^\{\\mathrm\{R\}\}\(u\_\{i\},o\),\(57\)Origm⁡\(𝒰;𝒰ocomm\)\\displaystyle\\operatorname\{Orig\}\_\{m\}\(\\mathcal\{U\};\\mathcal\{U\}\_\{o\}^\{\\mathrm\{comm\}\}\)=1n​∑i=1nminv∈𝒰ocomm⁡dmR​\(ui,v\)\.\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\min\_\{v\\in\\mathcal\{U\}\_\{o\}^\{\\mathrm\{comm\}\}\}d\_\{m\}^\{\\mathrm\{R\}\}\(u\_\{i\},v\)\.The second expression is the directed Chamfer distance from the generated set to the frozen common\-use reference set\([Chen & Ding, 2023](https://arxiv.org/html/2609.30492#bib.bib7)\)\.

##### Category Metrics \(Flexibility and Surprise\)\.

Letγ⁡\(u\)\\gamma\(u\)be the unique category assigned to valid useuuby human clustering, based on action class\([Fillmore, 1968](https://arxiv.org/html/2609.30492#bib.bib14);[Levin, 1993](https://arxiv.org/html/2609.30492#bib.bib34)\)and affordance\([Gibson, 1979](https://arxiv.org/html/2609.30492#bib.bib17);[Norman, 1988](https://arxiv.org/html/2609.30492#bib.bib42)\), and defineCat⁡\(𝒰\)=\{γ⁡\(u\):u∈𝒰\}\\operatorname\{Cat\}\(\\mathcal\{U\}\)=\\\{\\gamma\(u\):u\\in\\mathcal\{U\}\\\}\. The size\-normalized manual\-category metrics are

Flexγ⁡\(𝒰\)=\|Cat⁡\(𝒰\)\|n,\\operatorname\{Flex\}\_\{\\gamma\}\(\\mathcal\{U\}\)=\\frac\{\|\\operatorname\{Cat\}\(\\mathcal\{U\}\)\|\}\{n\},\(58\)Surγ\(𝒰;𝒰ocomm\)=1n∑i=1n\[γ\(ui\)∉Cat\(𝒰ocomm\)\]\.\\operatorname\{Sur\}\_\{\\gamma\}\(\\mathcal\{U\};\\mathcal\{U\}\_\{o\}^\{\\mathrm\{comm\}\}\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbbm\{1\}\\\!\\left\[\\gamma\(u\_\{i\}\)\\notin\\operatorname\{Cat\}\(\\mathcal\{U\}\_\{o\}^\{\\mathrm\{comm\}\}\)\\right\]\.\(59\)The clustering\-free flexibility proxy isFlexVS⁡\(𝒰\)=VS⁡\(𝒰\)/n\\operatorname\{Flex\}\_\{\\mathrm\{VS\}\}\(\\mathcal\{U\}\)=\\operatorname\{VS\}\(\\mathcal\{U\}\)/n\.

For the proposed originality\-weighted Vendi proxy, set

υi=minv∈𝒰ocomm⁡dcosR​\(ui,v\),υ¯i=υi∑r=1nυr,𝝊¯=\(υ¯1,…,υ¯n\)⊤,\\upsilon\_\{i\}=\\min\_\{v\\in\\mathcal\{U\}\_\{o\}^\{\\mathrm\{comm\}\}\}d\_\{\\mathrm\{cos\}\}^\{\\mathrm\{R\}\}\(u\_\{i\},v\),\\qquad\\overline\{\\upsilon\}\_\{i\}=\\frac\{\\upsilon\_\{i\}\}\{\\sum\_\{r=1\}^\{n\}\\upsilon\_\{r\}\},\\qquad\\overline\{\\bm\{\\upsilon\}\}=\(\\overline\{\\upsilon\}\_\{1\},\\ldots,\\overline\{\\upsilon\}\_\{n\}\)^\{\\top\},when∑rυr\>0\\sum\_\{r\}\\upsilon\_\{r\}\>0, and define

𝐆𝒰\(υ\)=diag⁡\(𝝊¯\)​𝐆𝒰​diag⁡\(𝝊¯\)\.\\mathbf\{G\}\_\{\\mathcal\{U\}\}^\{\(\\upsilon\)\}=\\operatorname\{diag\}\(\\sqrt\{\\overline\{\\bm\{\\upsilon\}\}\}\)\\,\\mathbf\{G\}\_\{\\mathcal\{U\}\}\\,\\operatorname\{diag\}\(\\sqrt\{\\overline\{\\bm\{\\upsilon\}\}\}\)\.Ifξi\(υ\)\\xi\_\{i\}^\{\(\\upsilon\)\}are its eigenvalues, then

SurVS\(𝒰;𝒰ocomm\)=1nexp\(−∑iξi\(υ\)logξi\(υ\)\)\.\\operatorname\{Sur\}\_\{\\mathrm\{VS\}\}\(\\mathcal\{U\};\\mathcal\{U\}\_\{o\}^\{\\mathrm\{comm\}\}\)=\\frac\{1\}\{n\}\\exp\\\!\\left\(\-\\sum\_\{i\}\\xi\_\{i\}^\{\(\\upsilon\)\}\\log\\xi\_\{i\}^\{\(\\upsilon\)\}\\right\)\.\(60\)We set this proxy to00when allυi=0\\upsilon\_\{i\}=0\. UnlikeSurγ\\operatorname\{Sur\}\_\{\\gamma\}, it is a proposed composite of reference\-relative originality and within\-set diversity, not a literal category share; its weighted spectral construction follows the quality\-weighted Vendi framework\([Nguyen & Dieng, 2024](https://arxiv.org/html/2609.30492#bib.bib41)\)\.

##### Creativity Scores by LLM Judge\.

The creativity, originality, surprise, and utility of each use are rated by an LLM judge\. Following[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18), all valid canonical uses for objectooare pooled across conditions, randomly partitioned into batches, ranked and scored, and repeated forRjudge=5R\_\{\\mathrm\{judge\}\}=5rounds with fresh partitions\. The rank score is the mean within\-batch percentile, and the judge score is the mean raw rating, scaled to\[1,5\]\[1,5\]\.

With the same judge, we obtain absolute per\-use ratings of originality, surprise, and utility using faithful renderings of[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib53)’s human protocols\. Utility is reported separately but remains the effectiveness component of creativity, rather than an unrelated quantity\([Runco & Jaeger, 2012](https://arxiv.org/html/2609.30492#bib.bib49)\)\.

MetricBasisDescriptionValidityHuman / LLM judgeProportion of admissible, well\-formed uses \(gate; control\)DiversityEmbeddingMean pairwise distance among a variant’s uses \(↑\\uparrow\)OriginalityEmbeddingMean semantic or directed Chamfer distance from usesto the object concept or common uses \(↑\\uparrow\)FlexibilityCategory\(Effective\) number \(ratio\) of distinct use categories \(↑\\uparrow\)SurpriseCategory\(Effective\) share of uses in categories beyond common uses \(↑\\uparrow\)CreativityLLM judgePooled rank\-and\-score creativity,[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18)\(↑\\uparrow\)OriginalityLLM judgeDeviation from typical use,[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib53)\(↑\\uparrow\)SurpriseLLM judgeUnexpectedness,[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib53)\(↑\\uparrow\)UtilityLLM judgeFeasibility/usefulness \(control\),[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib53)Table 5:Creativity and diversity metrics\. Embedding metrics use the fixed encoderϕ\\phi\(EmbeddingGemma\) with cosine, L2, or Mahalanobis dissimilarity; category metrics use both human clustering and Vendi\-based effective counts; LLM\-judge metrics use an independent Qwen3\.6\-27B judge\.↑\\uparrow: higher indicates a larger value\. Validity is a gate\. Utility is the effectiveness dimension required alongside originality for creativity\([Runco & Jaeger, 2012](https://arxiv.org/html/2609.30492#bib.bib49)\)and is reported separately to expose trade\-offs\.We use Qwen3\.6\-27B, an open\-weight model from a family distinct from the Gemma\-4\-31B generator\. Qwen3\.6\-27B shows moderate agreement with the two\-rater mean on Stevenson\-style ratings \(Spearman \(ρ=\.566\\rho=\.566\)–\(\.635\.635\); ICC\(\(C,1\)=\.407\(C,1\)=\.407\)–\(\.568\.568\)\), below the human–human agreement reference \(\(ρ=\.669\\rho=\.669\)–\(\.842\.842\); ICC \(=\.677=\.677\)–\(\.790\.790\)\)\. We therefore interpret these scores as noisy ordinal indicators, with strongest support for originality, and rely on convergence with embedding\- and category\-based measures rather than treating them as interchangeable with human ratings\.

Dim\.PairnnICC\(C,1\)\(C,1\)95% CIQWKρ\\rhoExactWithin\-1Paper ICCOrig\.LLM vs\. rater 11390\.480\[0\.341,0\.609\]\[0\.341,0\.609\]0\.4470\.5530\.5470\.8780\.78Orig\.LLM vs\. rater 21390\.575\[0\.452,0\.693\]\[0\.452,0\.693\]0\.5740\.5990\.5040\.9420\.78Orig\.LLM vs\. mean1390\.568\[0\.443,0\.682\]\[0\.443,0\.682\]0\.4810\.6350\.3020\.8780\.78Orig\.Human ceiling1390\.677\[0\.577,0\.754\]\[0\.577,0\.754\]0\.6090\.6690\.4750\.9930\.78Util\.LLM vs\. rater 11390\.355\[0\.211,0\.504\]\[0\.211,0\.504\]0\.3440\.5090\.4750\.8130\.70Util\.LLM vs\. rater 21390\.430\[0\.292,0\.568\]\[0\.292,0\.568\]0\.4060\.5970\.5320\.8060\.70Util\.LLM vs\. mean1390\.407\[0\.281,0\.531\]\[0\.281,0\.531\]0\.3770\.6030\.3960\.7840\.70Util\.Human ceiling1390\.701\[0\.493,0\.851\]\[0\.493,0\.851\]0\.6960\.7450\.7120\.9780\.70Surp\.LLM vs\. rater 11390\.432\[0\.229,0\.608\]\[0\.229,0\.608\]0\.4230\.5000\.6120\.9350\.79Surp\.LLM vs\. rater 21390\.544\[0\.407,0\.659\]\[0\.407,0\.659\]0\.4900\.5710\.5250\.8850\.79Surp\.LLM vs\. mean1390\.523\[0\.354,0\.667\]\[0\.354,0\.667\]0\.4680\.5660\.4460\.9060\.79Surp\.Human ceiling1390\.790\[0\.728,0\.841\]\[0\.728,0\.841\]0\.7560\.8420\.6980\.9710\.79Table 6:AUT LLM\-as\-a\-judge agreement on the*valid*Stevenson subset \(n=139n\{=\}139\)\. Paper ICC is the human–human ICC reported by Stevenson et al\.; Human ceiling is the reproducedrater01–rater02agreement on the released file\. LLM scores are from Qwen3\.6\-27B with the faithful Stevenson prompts \(including 0/99 codes\)\.Dim\.PairnnICC\(C,1\)\(C,1\)95% CIQWKρ\\rhoExactWithin\-1Paper ICCOrig\.LLM vs\. rater 11470\.386\[0\.248,0\.525\]\[0\.248,0\.525\]0\.3780\.4880\.5240\.8440\.57Orig\.LLM vs\. rater 21470\.479\[0\.341,0\.608\]\[0\.341,0\.608\]0\.4740\.5330\.4760\.9050\.57Orig\.LLM vs\. mean1470\.467\[0\.332,0\.601\]\[0\.332,0\.601\]0\.4100\.5550\.2860\.8370\.57Orig\.Human ceiling1470\.630\[0\.498,0\.734\]\[0\.498,0\.734\]0\.5790\.6420\.4760\.9860\.57Util\.LLM vs\. rater 11470\.336\[0\.199,0\.472\]\[0\.199,0\.472\]0\.3160\.4990\.4490\.7690\.68Util\.LLM vs\. rater 21470\.400\[0\.264,0\.534\]\[0\.264,0\.534\]0\.3670\.5720\.5030\.7690\.68Util\.LLM vs\. mean1470\.382\[0\.262,0\.501\]\[0\.262,0\.501\]0\.3620\.5840\.3740\.7480\.68Util\.Human ceiling1470\.684\[0\.505,0\.830\]\[0\.505,0\.830\]0\.6800\.7500\.6940\.9730\.68Surp\.LLM vs\. rater 11470\.391\[0\.210,0\.566\]\[0\.210,0\.566\]0\.3740\.4550\.5850\.9050\.67Surp\.LLM vs\. rater 21470\.480\[0\.344,0\.602\]\[0\.344,0\.602\]0\.4230\.5010\.4970\.8500\.67Surp\.LLM vs\. mean1470\.467\[0\.313,0\.613\]\[0\.313,0\.613\]0\.4110\.5020\.4220\.8780\.67Surp\.Human ceiling1470\.772\[0\.704,0\.832\]\[0\.704,0\.832\]0\.7430\.8310\.6940\.9660\.67Table 7:AUT LLM\-as\-a\-judge agreement on*all*doubly rated Stevenson responses \(n=147n\{=\}147\), including rows with invalid codes\. Metrics as in Table[6](https://arxiv.org/html/2609.30492#A4.T6)\.
##### Interpretation\.

On the production\-relevant valid subset, the LLM tracks the human mean with Spearmanρ∈\[0\.57,0\.64\]\\rho\\in\[0\.57,0\.64\]and within\-one rates of7878–91%91\\%: most disagreements are off\-by\-one on a five\-point scale rather than ordinal inversions\. Relative to the reproduced human–human ceiling, the LLM recovers about84%84\\%of the ceiling ICC and95%95\\%of the ceiling Spearman onoriginality, and still5858–81%81\\%of the ceiling on utility and surprise\. Exact agreement with the continuous rater mean is lower by construction \(integer LLM scores vs\. half\-point means\), so we emphasize ICC, QWK, Spearman, and within\-one\. Taken together, the judge is below, but in the same regime as trained human reliability, which is the appropriate bar for LLM\-as\-a\-judge deployment\. We therefore treat the Stevenson originality, utility, and surprise scores in the main AUT results as credible automated ratings, with originality the best\-aligned dimension and utility the most conservative\.

### D\.3Infinity\-Chat 100

We evaluate five manually selected queries from Infinity\-Chat 100[Jiang et al\. \(2026\)](https://arxiv.org/html/2609.30492#bib.bib26): 1, 2, 46, 60, and 78\. For each query and condition, persona\-free prompting produces 50 responses, while persona conditions produce ten responses from each of five personas\. This yields 250 responses per condition per generation seed\. Generation uses temperature1\.01\.0, top\-p=0\.9p=0\.9, and a maximum of 512 new tokens\. Before computing embedding and lexical metrics, we exclude responses explicitly labeled invalid\.

#### D\.3\.1Baselines and Variants

The persona\-free baseline uses the original query\. We also evaluate the shared ZS\-CoT, Step\-Back, and DMAD prompt variants, together with the persona controls and selection and generation variants defined above\. The composition experiment adds the selected personas to the DMAD prompt\.

##### Task\-specific evolutionary baseline\.

IC\-Evolutionreplaces persona\-space fitness with three response\-space axes: mean pairwise cosine distance, Vendi score, and one minus mean word\-trigram Jaccard overlap\. These are computed within each query and averaged across queries after validity screening and mechanical removal of persona framing\. The resulting population is supplied to the downstream Mahalanobis max–min dispersion selector\.

#### D\.3\.2Metrics

##### Homogeneity\.

For a query and condition, let𝒰=\(u1,…,un\)\\mathcal\{U\}=\(u\_\{1\},\\ldots,u\_\{n\}\)denote the response collection, and writedi​j=dcosR​\(ui,uj\)d\_\{ij\}=d\_\{\\mathrm\{cos\}\}^\{\\mathrm\{R\}\}\(u\_\{i\},u\_\{j\}\)\. Homogeneity is mean pairwise cosine similarity,

Hom⁡\(𝒰\)=1−2n⁡\(n−1\)​∑i<jdi​j\.\\operatorname\{Hom\}\(\\mathcal\{U\}\)=1\-\\frac\{2\}\{n\(n\-1\)\}\\sum\_\{i<j\}d\_\{ij\}\.

##### Persona Separation\.

Letpip\_\{i\}be the persona generatinguiu\_\{i\}\. Defineℬ=\{\(i,j\):i<j,pi≠pj\}\\mathcal\{B\}=\\\{\(i,j\):i<j,\\ p\_\{i\}\\neq p\_\{j\}\\\}andℐ=\{\(i,j\):i<j,pi=pj\}\\mathcal\{I\}=\\\{\(i,j\):i<j,\\ p\_\{i\}=p\_\{j\}\\\}\. Persona separation is

Sep⁡\(𝒰\)=1\|ℬ\|​∑\(i,j\)∈ℬdi​j−1\|ℐ\|​∑\(i,j\)∈ℐdi​j\.\\operatorname\{Sep\}\(\\mathcal\{U\}\)=\\frac\{1\}\{\|\\mathcal\{B\}\|\}\\sum\_\{\(i,j\)\\in\\mathcal\{B\}\}d\_\{ij\}\-\\frac\{1\}\{\|\\mathcal\{I\}\|\}\\sum\_\{\(i,j\)\\in\\mathcal\{I\}\}d\_\{ij\}\.Separation is undefined when either pair set is empty\. Both metrics are computed within each query and averaged across queries\.

##### Flexibility\.

We apply the cosine\-Gram Vendi score in equation[53](https://arxiv.org/html/2609.30492#A4.E53)to the response embeddings for each query and average across queries\.

##### Validity and Quality\.

Validity is the fraction of responses admitted by the validity screen\. Qwen3\.6\-27B independently rates the quality of each response on a five\-point scale for execution and fit to the request and the scores are averaged across responses and queries\.

##### LLM Judge on Response Quality and Validity

Infinity\-Chat releases human*ratings*of model responses without human\-written response pools[Jiang et al\. \(2026\)](https://arxiv.org/html/2609.30492#bib.bib26)\. We use the ratings to validate Qwen3\.6\-27B judge employed in evaluation:750750\(query, response\) pairs from2525external LMs, each with approximately2525independent absolute ratings on a 1–5 scale \(human\_absolute\)\. The quality head is scored against the2525\-rater mean \(the paper’s comparison target for LM judges\); the validity head is checked with a tripwire that responses humans rate≥4\\geq 4must almost never be labeled invalid\. Table[8](https://arxiv.org/html/2609.30492#A4.T8)reports rank and absolute agreement for quality; Table[9](https://arxiv.org/html/2609.30492#A4.T9)reports the validity screen\.

MetricValuennscored pairs750Quality prompthivemind\_quality\_absolute\_v1Spearmanρ\\rho\(judge vs\. human mean\)0\.158Pearsonrr\(judge vs\. human mean\)0\.278ICC\(C,1\)\(C,1\)0\.225ICC 95% CI \(bootstrap\)\[0\.146,0\.296\]\[0\.146,\\,0\.296\]MAE1\.021RMSE1\.156Mean bias \(judge−\-human mean\)\+0\.691\+0\.691Exact agree \(vs\. rounded human mean\)0\.136Within\-one \(vs\. rounded human mean\)0\.840QWK \(vs\. rounded human mean\)0\.132Human split\-half Spearman ceiling0\.624ρ\\rho/ ceiling0\.253Table 8:Infinity\-Chat quality\-judge agreement against the2525\-rater human mean on the Artificial Hivemind dense absolute\-rating set\. Exact / within\-one / QWK compare the integer LLM score to the human mean rounded to the nearest integer\.Human mean binnnInvalid ratePromptGood \(≥4\\geq 4\)4050\.002hivemind\_validity\_screen\_v1Mid \(\(2,4\)\(2,4\)\)3410\.023Bad \(≤2\\leq 2\)40\.250Overall7500\.013Table 9:Infinity\-Chat validity\-gate tripwire on the same750750pairs\. The automated gate fails the protocol if it marks\>5%\{\>\}5\\%of human\-good responses invalid; the observed rate is0\.25%0\.25\\%\.The validity screen is*highly conservative*with respect to human judgments: only0\.25%0\.25\\%of responses that humans rate as good \(≥4\\geq 4\) are discarded, well under the5%5\\%tripwire, so the LLM gate does not systematically delete content that human raters endorse\. On quality, the judge is*positively associated*with the human consensus \(Pearsonr=0\.28r\{=\}0\.28, ICC=0\.23\{=\}0\.23\) and lands within one scale point of the rounded human mean on84%84\\%of items, while remaining below the human split\-half Spearman ceiling \(ρ=0\.16\\rho\{=\}0\.16vs\.0\.620\.62\)\. The residual gap is consistent with a mild high\-score bias \(mean bias\+0\.69\+0\.69\) on a distribution where human means already cluster near the top of the scale\. We therefore treat the Infinity\-Chatvaliditylabels as a credible hard filter for all downstream LLM and embedding metrics, and treatqualityscores as a coarsely human\-aligned complementary axis, which is useful for relative quality, diversity tradeoffs across variants, and interpreted alongside embedding\-based homogeneity and diversity rather than as a calibrated substitute for the full2525\-rater panel\.

### D\.4Divergent Association Task \(DAT\)

#### D\.4\.1Baselines and Variants

##### Standard persona baselines\.

We use the original DAT prompt\([Olson et al\., 2021](https://arxiv.org/html/2609.30492#bib.bib43);[Chen & Ding, 2023](https://arxiv.org/html/2609.30492#bib.bib7)\)as the baseline persona\-free prompt:DAT\.[Schapiro et al\. \(2026\)](https://arxiv.org/html/2609.30492#bib.bib50)tests various prompts to contrast creative and non\-creative outputs, and we also compare their task\-optimized prompts as our baseline comparisons:CN\-Base\(their base instruction prompt that contrasts with random instruction prompt\),CN\-Random\(their random instruction prompt\),Creativity\-enhanced\(their creative prompt that serves as a contrast to their non\-creative prompt\),Non\-creative\(their non\-creative prompt that serves the similar role as common use prompt of AUT benchmark\)\.

##### Other baselines and variants\.

Other reasoning baselines, persona controls \(gibberish,[Jin et al\. \(2025\)](https://arxiv.org/html/2609.30492#bib.bib27), and random/typical personas from[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.30492#bib.bib25);[Ge et al\. \(2024\)](https://arxiv.org/html/2609.30492#bib.bib16)\) are similarly compared for DAT as well\. Our selection and generation \(MCMC and Evo\) variants use the same set of selected personas used for AUT\. \(Note that the personas, as well as selection and generation algorithms are agnostic of a specific task\)\.DAT\-Evolutionis the task\-conditioned variant using DAT\-specific fitness function

Fit\(p\)=λDATDATm\(𝒲1:7\)\+λFlexVS\(𝒲\)\+λHH\(𝒲\)\+λC\(1−C10\(𝒲\)\),\\operatorname\{Fit\}\(p\)=\\lambda\_\{\\operatorname\{DAT\}\}\\operatorname\{DAT\}\_\{m\}\(\\mathcal\{W\}\_\{1:7\}\)\+\\lambda\_\{\\operatorname\{Flex\}\}\\operatorname\{VS\}\(\\mathcal\{W\}\)\+\\lambda\_\{H\}H\(\\mathcal\{W\}\)\+\\lambda\_\{C\}\(1\-C\_\{10\}\(\\mathcal\{W\}\)\),\(61\)where𝒲\\mathcal\{W\}denotes the set of valid responses produced by the personapp\.

##### Generation and sampling protocol\.

Non\-persona conditions usencomp=35n\_\{\\mathrm\{comp\}\}=35independent completions, while persona conditions use seven completions for each of thek=5k=5selected personas\. We retain the firstnword=10n\_\{\\mathrm\{word\}\}=10words from each completion, givingnraw=350n\_\{\\mathrm\{raw\}\}=350raw words per condition before validity gate\.

#### D\.4\.2Metrics

For DAT, we report four creativity metrics along with validity ratio: DAT score as suggested by the original literature\([Olson et al\., 2021](https://arxiv.org/html/2609.30492#bib.bib43)\), fluency measured by Shannon entropy, flexibility measured by Vendi score equation[53](https://arxiv.org/html/2609.30492#A4.E53), and top\-10 concentration\.

##### Notation\.

For DAT, we denote response aswwand a valid response set as𝒲\\mathcal\{W\}, as they represent noun words\. LetWj=\(wj​1,…,wj​nj\)W\_\{j\}=\(w\_\{j1\},\\ldots,w\_\{jn\_\{j\}\}\)contain the usable validated word occurrences from completionjjin response order\. Unlike AUT, the relationship between1010words in a completion is also an important evaluation target, because the task is to name1010irrelevant nouns\. Thus, we deal with the pooled multiset𝒲=⨄jWj\\mathcal\{W\}=\\biguplus\_\{j\}W\_\{j\}\. Let𝒲human\\mathcal\{W\}^\{\\mathrm\{human\}\}be the human reference set of words, drawn from human responses collected by[Olson et al\. \(2021\)](https://arxiv.org/html/2609.30492#bib.bib43)\.

##### DAT Score\.

[Olson et al\. \(2021\)](https://arxiv.org/html/2609.30492#bib.bib43)invented Divergent Association Task \(DAT\) along with a creativity metric called DAT score\. It is adopted in its original form by following works using DAT\([Chen & Ding, 2023](https://arxiv.org/html/2609.30492#bib.bib7);[Schapiro et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib50)\)\. Letg¯​\(w\)\\overline\{g\}\(w\)denote its unit\-normalized GloVe word embedding\. For a scorable completion \(nj≥7n\_\{j\}\\geq 7\),

DAT⁡\(Wj\)=100\(nj2\)​∑1≤a<b≤nj\[1−g¯​\(wj​a\)⊤​g¯​\(wj​b\)\]\.\\operatorname\{DAT\}\(W\_\{j\}\)=\\frac\{100\}\{\\binom\{n\_\{j\}\}\{2\}\}\\sum\_\{1\\leq a<b\\leq n\_\{j\}\}\\left\[1\-\\overline\{g\}\(w\_\{ja\}\)^\{\\top\}\\overline\{g\}\(w\_\{jb\}\)\\right\]\.Human\-reference percentiles use the same calculation applied to the raw word lists of 8,572 human completions\.

##### Flency\.

Shannon entropy captures fluency on how flat the generated words are distributed: higher entropy, lower duplicate words and thus higher fluency\. Therefore, the metric measures how creative/divergent LLMs are

H\(𝒲\)=−∑w∈𝒱cw∑w∈𝒱cwlogcw∑w∈𝒱cw,H\(\\mathcal\{W\}\)=\-\\sum\_\{w\\in\\mathcal\{V\}\}\\frac\{c\_\{w\}\}\{\\sum\_\{w\\in\\mathcal\{V\}\}c\_\{w\}\}\\log\{\\frac\{c\_\{w\}\}\{\\sum\_\{w\\in\\mathcal\{V\}\}c\_\{w\}\}\},\(62\)where𝒱\\mathcal\{V\}is vocabulary or the set of unique words in𝒲\\mathcal\{W\}andcwc\_\{w\}is the count of appearance ofwwin𝒲\\mathcal\{W\}\.

Top\-10 Concentration measures fluency on how much proportion the top\-10 most frequent words appear over and over again: lower top\-10 share, lower concentration and thus higher fluency\. Therefore, the metric measures how lexically sophisticated LLMs are

C10​\(𝒲\)=∑i=110c\(i\)∑w∈𝒱cw,C\_\{10\}\(\\mathcal\{W\}\)=\\frac\{\\sum\_\{i=1\}^\{10\}c\_\{\(i\)\}\}\{\\sum\_\{w\\in\\mathcal\{V\}\}c\_\{w\}\},\(63\)wherec\(i\)c\_\{\(i\)\}is the count of appearancei−i\-th most frequently appearing word in𝒲\\mathcal\{W\}\.

##### Flexibility\.

For each nonempty completionWjW\_\{j\}, we apply equation[53](https://arxiv.org/html/2609.30492#A4.E53)to the unit\-normalized EmbeddingGemma embeddings of its individual usable word occurrences, retaining repeated words as repeated embedding rows\. The reported Vendi score is the mean of these within\-completion scores over nonempty completions\. It measures effective semantic diversity within a generated word list, whereas type token ratio \(TTR\) is the ratio of unique valid words to all valid word occurrences

T​T​R​\(𝒲\)=\|𝒱\|\|𝒲\|\.TTR\(\\mathcal\{W\}\)=\\frac\{\|\\mathcal\{V\}\|\}\{\|\\mathcal\{W\}\|\}\.\(64\)

## Appendix EExperiment Results

### E\.1Baselines

#### E\.1\.1Common Baselines

Here we evaluate and compare various metrics on 12 categories of variants: 3 standard persona baselines using reasoning prompts, 2 persona control baselines, 2 PersonaMem\-v2\([Jiang et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib25)\)baselines, 2 selection algorithm variants, and 3 generation algorithm variants\. These are existing baselines using reasoning prompts:

- ▶\\blacktrianglerightZS\-CoT: Zero\-Shot Chain\-of\-Thought \(ZS\-CoT\) reasoning prompt adapted from[Kojima et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib30)
- ▶\\blacktrianglerightStep\-Back: Step\-Back Prompting \(SBP\) adapted from[Zheng et al\. \(2024\)](https://arxiv.org/html/2609.30492#bib.bib59)
- ▶\\blacktrianglerightDMAD: Diverse Multi\-Agent Debate prompt combining reasoning strategies \(ZS\-CoT \+ SBP\) adapted from[Liu et al\. \(2025\)](https://arxiv.org/html/2609.30492#bib.bib36)

We also compare our methods against control personas:

- ▶\\blacktrianglerightGibberish: gibberish of equivalent length as personas our methods utilize
- ▶\\blacktrianglerightTask\-conditioned: task\-conditioned personas generated by[Jin et al\. \(2025\)](https://arxiv.org/html/2609.30492#bib.bib27)

We compare against existing persona benchmark PersonaMem\-v2 dataset\([Jiang et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib25)\), which randomly selected from PersonaHub\([Ge et al\., 2024](https://arxiv.org/html/2609.30492#bib.bib16)\), using two sampling methods:

- ▶\\blacktrianglerightRandom: randomly selected personas from the base population\.
- ▶\\blacktrianglerightTypical: typical personas selected from the most dense regions\.

Finally, we apply our diversity\-aware selection methods on persona population of[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.30492#bib.bib25):

- ▶\\blacktrianglerightCoverage: selected personas from the base population to maximize coverage defined by cosine, L2 and Mahalanobis distances\.
- ▶\\blacktrianglerightDispersion: selected personas from the base population to maximize minimum dispersion defined by cosine, L2 and Mahalanobis distances; Mahalnobis distance\-based dispersion selection is chosen as a default for all persona generation algorithms\.

Then, expand the baseline population using our MCMC and evolutionary persona generation algorithms:

- ▶\\blacktrianglerightMCMC: extended population using Uniform\-Coverage MCMC persona generation algorithm\.
- ▶\\blacktrianglerightEvolution: extended population using Evolutionary TextGrad algorithm on persona fitness function\.
- ▶\\blacktrianglerightAUT Evolution: extended population using Evolutionary TextGrad algorithm on AUT response fitness function\.

#### E\.1\.2AUT Baselines

These are the baseline and task\-optimized prompts from[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18):

- ▶\\blacktrianglerightCommon Use: common uses of objects
- ▶\\blacktrianglerightAlternative Use: alternative uses of objects generated by generic AUT prompt
- ▶\\blacktrianglerightExpert: alternative uses generated by AUT\-optimized expert prompt
- ▶\\blacktrianglerightCreativity\-enhanced: alternative uses generated by creativity\-enhanced prompt adapted from[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18)

#### E\.1\.3DAT Baselines

These are the prompt conditions for the DAT, following[Chen & Ding \(2023\)](https://arxiv.org/html/2609.30492#bib.bib7)and[Schapiro et al\. \(2026\)](https://arxiv.org/html/2609.30492#bib.bib50):

- ▶\\blacktrianglerightDivergent Association: zero\-shot task prompt in the original form of[Olson et al\. \(2021\)](https://arxiv.org/html/2609.30492#bib.bib43), adapted by[Chen & Ding \(2023\)](https://arxiv.org/html/2609.30492#bib.bib7)\.
- ▶\\blacktrianglerightBase\-Instruction: base control prompt\([Chen & Ding, 2023](https://arxiv.org/html/2609.30492#bib.bib7)\), which asks for ten nouns with no divergence instruction\.
- ▶\\blacktrianglerightRandom\-Instruction: random control prompt\([Chen & Ding, 2023](https://arxiv.org/html/2609.30492#bib.bib7)\), which asks for ten random nouns\.
- ▶\\blacktrianglerightCreative/ Non\-Creative \(Non\-Divergent Association\): the contrastive prompt pair of[Schapiro et al\. \(2026\)](https://arxiv.org/html/2609.30492#bib.bib50), which instruct the model to answer creatively and uncreatively respectively, bracketing the achievable range\.

Every condition beyond DAT, the reasoning prompts, the control personas, and all of our selection and generation variants, injects into theDivergent Associationprompt, which is therefore the reference condition throughout the DAT results\. Selection under cosine andL2L\_\{2\}dissimilarity returns identical persona sets at every pool, so the two are reported jointly\.

### E\.2Persona Diversity

Table 10:Persona diversity of the variants measured in Dispersion \(equation[8](https://arxiv.org/html/2609.30492#S3.E8)\) and Vendi score \(equation[53](https://arxiv.org/html/2609.30492#A4.E53)\) of each persona set \(k=5k=5\), computed on persona embeddings under cosine, L2, and Mahalanobis dissimilarity\. Dispersion represents how personas are spread far from each other \(minimum pairwise distance within the set; higher is more spread\)\. Vendi score represents the effective number of distinct clusters of personas \(higher is more distinct\)\. Methods are grouped by their algorithm categories: MPAQ is the baseline\([Jin et al\., 2025](https://arxiv.org/html/2609.30492#bib.bib27)\)that generated persona for output diversity, Individualized personas represent baseline selection methods from the base pool of size 997, Selected personas represent our methods of persona selection according to various objective functions, and Generated personas represent our methods of persona generation, each selected from its expanded pool by exact Mahalanobis max–min dispersion\.Bold= best per column within the selected and generated groups; baseline groups are not bolded\.Persona setDispersionVendi scoreCosineL2Mahal\.CosineL2Mahal\.Baseline personasTask\-conditioned0\.1300\.1300\.5100\.51013\.4213\.421\.811\.812\.212\.211\.211\.21Individualized personas\(base pool=997=997\)Typical0\.2700\.2700\.7340\.73412\.0412\.042\.342\.342\.732\.731\.211\.21Random0\.2480\.2480\.7030\.70313\.6713\.672\.602\.602\.922\.921\.251\.25Selected personas\(base pool=997=997\)Coverage, cosine0\.2310\.2310\.6790\.67912\.2512\.252\.222\.222\.632\.631\.191\.19Coverage, Mahalanobis0\.1910\.1910\.6190\.61910\.5210\.522\.252\.252\.642\.641\.151\.15Dispersion, cosine0\.485\\mathbf\{0\.485\}0\.984\\mathbf\{0\.984\}16\.0516\.053\.41\\mathbf\{3\.41\}3\.41\\mathbf\{3\.41\}1\.381\.38Dispersion, Mahalanobis0\.3660\.3660\.8560\.85620\.98\\mathbf\{20\.98\}3\.073\.073\.233\.231\.46\\mathbf\{1\.46\}Generated personas\(expanded pool size\)MCMC\(8,553\)0\.3520\.3520\.8390\.83920\.9820\.983\.043\.043\.213\.211\.461\.46Evolution\(1,869\)0\.3930\.3930\.8870\.88724\.38\\mathbf\{24\.38\}3\.193\.193\.303\.301\.64\\mathbf\{1\.64\}AUT\-Evolution\(3,397\)0\.3730\.3730\.8640\.86422\.3422\.343\.123\.123\.263\.261\.511\.51DAT\-Evolution\(2,439\)0\.436\\mathbf\{0\.436\}0\.934\\mathbf\{0\.934\}21\.6121\.613\.28\\mathbf\{3\.28\}3\.34\\mathbf\{3\.34\}1\.491\.49Figure 5:Persona\-to\-response relationship on \(a\) the Alternative Uses Task and \(b\) Infinity\-Chat\. Points represent 14 analogous task\-agnostic method configurations per benchmark; response hull extent and vendi score are averaged over seven AUT objects and five Infinity\-Chat queries\. Arrows connect six matched Coverage→\\rightarrowDispersion contrasts while holding candidate\-pool family and selector metric fixed\. Dashed lines show configuration\-level linear fits; annotations report Pearson correlations and descriptive 95% intervals\.
### E\.3AUT Benchmark

Table 11:Response validity: the proportion of generated uses admitted by the validity gate of[SectionD\.2](https://arxiv.org/html/2609.30492#A4.SS2), reported over all seven AUT objects and, for comparability withStevenson, over the three objects with human reference data \(175175and7575generated uses per variant, respectively\)\. Every variants beyondAlternative\-Usebuild prompts upon the genericAlternative\-Useprompt unless its group states otherwise; the final two groups re\-run our selection and generation methods on theCreativity\-enhancedprompt of[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18), showing that persona diversification composes with a stronger prompting strategy rather than competing with it\. Cells give the admitted proportion averaged over objects\. Subscripts denote the 95% confidence interval across 5 independent random seeds\.Stevensonexists only for the human\-reference objects, so its seven\-object cell is omitted \(—\)\. Validity is a*control*rather than a target: values near11are expected and the informative signal is degradation rather than rank, so no value is bolded\. Generated variants report their expanded pool size in parentheses\.VariantAll seven objectsHuman\-reference objectsStandard personaCommon\-Use0\.971±0\.0180\.971\_\{\\pm 0\.018\}0\.981±0\.0150\.981\_\{\\pm 0\.015\}Alternative\-Use0\.984±0\.0140\.984\_\{\\pm 0\.014\}0\.968±0\.0250\.968\_\{\\pm 0\.025\}Expert0\.997±0\.0060\.997\_\{\\pm 0\.006\}0\.997±0\.0070\.997\_\{\\pm 0\.007\}Creativity\-enhanced0\.968±0\.0110\.968\_\{\\pm 0\.011\}0\.981±0\.0220\.981\_\{\\pm 0\.022\}ZS\-CoT0\.978±0\.0090\.978\_\{\\pm 0\.009\}0\.997±0\.0070\.997\_\{\\pm 0\.007\}Step\-Back0\.976±0\.0140\.976\_\{\\pm 0\.014\}0\.968±0\.0150\.968\_\{\\pm 0\.015\}DMAD0\.977±0\.0090\.977\_\{\\pm 0\.009\}0\.997±0\.0070\.997\_\{\\pm 0\.007\}Gibberish0\.693±0\.0350\.693\_\{\\pm 0\.035\}0\.659±0\.1250\.659\_\{\\pm 0\.125\}Baseline personasMPAQ0\.930±0\.0240\.930\_\{\\pm 0\.024\}0\.947±0\.0200\.947\_\{\\pm 0\.020\}Individualized personasRandom0\.993±0\.0060\.993\_\{\\pm 0\.006\}0\.989±0\.0140\.989\_\{\\pm 0\.014\}Typical0\.995±0\.0060\.995\_\{\\pm 0\.006\}0\.997±0\.0070\.997\_\{\\pm 0\.007\}Selected personas from base poolCoverage, cosine0\.987±0\.0060\.987\_\{\\pm 0\.006\}1\.000±0\.0001\.000\_\{\\pm 0\.000\}Coverage, Mahalanobis0\.993±0\.0130\.993\_\{\\pm 0\.013\}1\.000±0\.0001\.000\_\{\\pm 0\.000\}Dispersion, cosine0\.992±0\.0120\.992\_\{\\pm 0\.012\}0\.992±0\.0150\.992\_\{\\pm 0\.015\}Dispersion, Mahalanobis0\.986±0\.0100\.986\_\{\\pm 0\.010\}0\.995±0\.0090\.995\_\{\\pm 0\.009\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.984±0\.0090\.984\_\{\\pm 0\.009\}0\.995±0\.0090\.995\_\{\\pm 0\.009\}Evolution \(1,869\)0\.985±0\.0120\.985\_\{\\pm 0\.012\}0\.987±0\.0170\.987\_\{\\pm 0\.017\}AUT\-Evolution \(3,397\)0\.978±0\.0170\.978\_\{\\pm 0\.017\}0\.989±0\.0140\.989\_\{\\pm 0\.014\}Selected personas\+\+Creativity\-enhancedpromptCoverage, cosine0\.978±0\.0120\.978\_\{\\pm 0\.012\}0\.981±0\.0150\.981\_\{\\pm 0\.015\}Coverage, Mahalanobis0\.989±0\.0140\.989\_\{\\pm 0\.014\}0\.992±0\.0090\.992\_\{\\pm 0\.009\}Dispersion, cosine0\.982±0\.0080\.982\_\{\\pm 0\.008\}0\.979±0\.0190\.979\_\{\\pm 0\.019\}Dispersion, Mahalanobis0\.979±0\.0110\.979\_\{\\pm 0\.011\}0\.976±0\.0300\.976\_\{\\pm 0\.030\}Generated personas\+\+Creativity\-enhancedpromptMCMC0\.967±0\.0140\.967\_\{\\pm 0\.014\}0\.968±0\.0400\.968\_\{\\pm 0\.040\}Evolution0\.979±0\.0200\.979\_\{\\pm 0\.020\}0\.973±0\.0310\.973\_\{\\pm 0\.031\}AUT\-Evolution0\.973±0\.0060\.973\_\{\\pm 0\.006\}0\.968±0\.0340\.968\_\{\\pm 0\.034\}HumanStevenson—0\.811±0\.0070\.811\_\{\\pm 0\.007\}Table 12:Response diversity across all seven AUT objects, reported two ways\.*Mean pairwise distance*\(equation[55](https://arxiv.org/html/2609.30492#A4.E55)\) is the average cosine, L2 or Mahalanobis distance among the valid uses of a variant;*convex hull volume*\(equation[55](https://arxiv.org/html/2609.30492#A4.E55)\) is the volume of the convex hull those uses span in the same geometry, so it rewards a set that occupies a region rather than merely separating in the mean\. Both are averaged over objects\. Subscripts denote the 95% confidence interval across 5 independent random seeds, and higher is more diverse\. Hull extent under cosine and L2 coincides to nine decimal places, the two metrics are monotone transforms of one another on unit\-normalized embeddings, so it is reported once\. Baseline and our method variants are grouped by the personas: standard persona represents LLMs default persona when no persona is given explicitly, baseline personas represent existing persona injection methods, individualized personas represent broad population of personas from PersonaHub[Ge et al\. \(2024\)](https://arxiv.org/html/2609.30492#bib.bib16);[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.30492#bib.bib25), while selected personas and generated personas represent our specific persona selection and generation algorithms\. The final two groups re\-run our selection and generation methods on theCreativity\-enhancedprompt, showing that persona diversification composes with a stronger prompting strategy\. Bold marks the best value per column within a method category, with tied values bolded jointly\. All personas, including gibberish, are injected to the alternative use prompt unless their group states otherwise\.Mean pairwise distanceConvex hull volumeVariantCosineL2Mahal\.Cosine/L2Mahal\.Standard personaCommon\-Use0\.155±0\.0040\.155\_\{\\pm 0\.004\}0\.526±0\.0060\.526\_\{\\pm 0\.006\}15\.50±0\.3515\.50\_\{\\pm 0\.35\}0\.103±0\.0040\.103\_\{\\pm 0\.004\}2\.17±0\.232\.17\_\{\\pm 0\.23\}Alternative\-Use0\.188±0\.0030\.188\_\{\\pm 0\.003\}0\.590±0\.0060\.590\_\{\\pm 0\.006\}11\.97±0\.2811\.97\_\{\\pm 0\.28\}0\.172±0\.0140\.172\_\{\\pm 0\.014\}1\.42±0\.091\.42\_\{\\pm 0\.09\}Expert0\.179±0\.0040\.179\_\{\\pm 0\.004\}0\.580±0\.0060\.580\_\{\\pm 0\.006\}14\.57±0\.2914\.57\_\{\\pm 0\.29\}0\.161±0\.0040\.161\_\{\\pm 0\.004\}2\.36±0\.132\.36\_\{\\pm 0\.13\}Creativity\-enhanced0\.179±0\.0020\.179\_\{\\pm 0\.002\}0\.592±0\.0040\.592\_\{\\pm 0\.004\}20\.24±0\.56\\mathbf\{20\.24\}\_\{\\pm 0\.56\}0\.144±0\.0100\.144\_\{\\pm 0\.010\}5\.15±0\.16\\mathbf\{5\.15\}\_\{\\pm 0\.16\}ZS\-CoT0\.165±0\.0030\.165\_\{\\pm 0\.003\}0\.563±0\.0060\.563\_\{\\pm 0\.006\}14\.73±0\.3814\.73\_\{\\pm 0\.38\}0\.161±0\.0060\.161\_\{\\pm 0\.006\}2\.43±0\.112\.43\_\{\\pm 0\.11\}Step\-Back0\.173±0\.0060\.173\_\{\\pm 0\.006\}0\.574±0\.0140\.574\_\{\\pm 0\.014\}13\.66±0\.5213\.66\_\{\\pm 0\.52\}0\.168±0\.0110\.168\_\{\\pm 0\.011\}2\.03±0\.292\.03\_\{\\pm 0\.29\}DMAD0\.178±0\.0050\.178\_\{\\pm 0\.005\}0\.588±0\.0090\.588\_\{\\pm 0\.009\}16\.23±0\.4616\.23\_\{\\pm 0\.46\}0\.167±0\.0070\.167\_\{\\pm 0\.007\}2\.79±0\.222\.79\_\{\\pm 0\.22\}Gibberish0\.233±0\.005\\mathbf\{0\.233\}\_\{\\pm 0\.005\}0\.672±0\.006\\mathbf\{0\.672\}\_\{\\pm 0\.006\}19\.83±0\.8219\.83\_\{\\pm 0\.82\}0\.184±0\.008\\mathbf\{0\.184\}\_\{\\pm 0\.008\}3\.14±0\.343\.14\_\{\\pm 0\.34\}Baseline personasMPAQ0\.216±0\.0020\.216\_\{\\pm 0\.002\}0\.652±0\.0030\.652\_\{\\pm 0\.003\}19\.05±0\.4919\.05\_\{\\pm 0\.49\}0\.186±0\.0030\.186\_\{\\pm 0\.003\}3\.46±0\.273\.46\_\{\\pm 0\.27\}Individualized personasRandom0\.202±0\.0040\.202\_\{\\pm 0\.004\}0\.621±0\.0070\.621\_\{\\pm 0\.007\}12\.59±0\.6312\.59\_\{\\pm 0\.63\}0\.202±0\.0040\.202\_\{\\pm 0\.004\}1\.54±0\.121\.54\_\{\\pm 0\.12\}Typical0\.197±0\.0060\.197\_\{\\pm 0\.006\}0\.615±0\.0100\.615\_\{\\pm 0\.010\}12\.78±0\.2612\.78\_\{\\pm 0\.26\}0\.204±0\.0030\.204\_\{\\pm 0\.003\}1\.63±0\.121\.63\_\{\\pm 0\.12\}Selected personas from base poolCoverage, cosine/L20\.207±0\.0060\.207\_\{\\pm 0\.006\}0\.629±0\.0080\.629\_\{\\pm 0\.008\}12\.71±0\.2812\.71\_\{\\pm 0\.28\}0\.203±0\.0070\.203\_\{\\pm 0\.007\}1\.58±0\.081\.58\_\{\\pm 0\.08\}Coverage, Mahalanobis0\.193±0\.0030\.193\_\{\\pm 0\.003\}0\.604±0\.0050\.604\_\{\\pm 0\.005\}11\.77±0\.2011\.77\_\{\\pm 0\.20\}0\.200±0\.0070\.200\_\{\\pm 0\.007\}1\.50±0\.061\.50\_\{\\pm 0\.06\}Dispersion, cosine/L20\.208±0\.003\\mathbf\{0\.208\}\_\{\\pm 0\.003\}0\.633±0\.004\\mathbf\{0\.633\}\_\{\\pm 0\.004\}13\.36±0\.18\\mathbf\{13\.36\}\_\{\\pm 0\.18\}0\.213±0\.005\\mathbf\{0\.213\}\_\{\\pm 0\.005\}1\.98±0\.14\\mathbf\{1\.98\}\_\{\\pm 0\.14\}Dispersion, Mahalanobis0\.202±0\.0040\.202\_\{\\pm 0\.004\}0\.622±0\.0050\.622\_\{\\pm 0\.005\}12\.95±0\.3512\.95\_\{\\pm 0\.35\}0\.210±0\.0050\.210\_\{\\pm 0\.005\}1\.96±0\.171\.96\_\{\\pm 0\.17\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.207±0\.0040\.207\_\{\\pm 0\.004\}0\.631±0\.0060\.631\_\{\\pm 0\.006\}14\.04±0\.5914\.04\_\{\\pm 0\.59\}0\.211±0\.0010\.211\_\{\\pm 0\.001\}2\.21±0\.232\.21\_\{\\pm 0\.23\}Evolution \(1,869\)0\.214±0\.0050\.214\_\{\\pm 0\.005\}0\.644±0\.0070\.644\_\{\\pm 0\.007\}15\.47±0\.3915\.47\_\{\\pm 0\.39\}0\.210±0\.0050\.210\_\{\\pm 0\.005\}2\.53±0\.202\.53\_\{\\pm 0\.20\}AUT\-Evolution \(3,397\)0\.216±0\.006\\mathbf\{0\.216\}\_\{\\pm 0\.006\}0\.646±0\.009\\mathbf\{0\.646\}\_\{\\pm 0\.009\}15\.66±0\.21\\mathbf\{15\.66\}\_\{\\pm 0\.21\}0\.218±0\.007\\mathbf\{0\.218\}\_\{\\pm 0\.007\}2\.98±0\.16\\mathbf\{2\.98\}\_\{\\pm 0\.16\}Selected personas\+\+Creativity\-enhancedpromptCoverage, cosine/L20\.190±0\.0050\.190\_\{\\pm 0\.005\}0\.610±0\.0090\.610\_\{\\pm 0\.009\}20\.44±0\.3220\.44\_\{\\pm 0\.32\}0\.150±0\.0060\.150\_\{\\pm 0\.006\}5\.17±0\.215\.17\_\{\\pm 0\.21\}Coverage, Mahalanobis0\.190±0\.0030\.190\_\{\\pm 0\.003\}0\.609±0\.0040\.609\_\{\\pm 0\.004\}20\.45±0\.1720\.45\_\{\\pm 0\.17\}0\.154±0\.0020\.154\_\{\\pm 0\.002\}5\.22±0\.155\.22\_\{\\pm 0\.15\}Dispersion, cosine/L20\.195±0\.005\\mathbf\{0\.195\}\_\{\\pm 0\.005\}0\.618±0\.009\\mathbf\{0\.618\}\_\{\\pm 0\.009\}20\.57±0\.2020\.57\_\{\\pm 0\.20\}0\.159±0\.007\\mathbf\{0\.159\}\_\{\\pm 0\.007\}6\.43±0\.226\.43\_\{\\pm 0\.22\}Dispersion, Mahalanobis0\.191±0\.0040\.191\_\{\\pm 0\.004\}0\.611±0\.0080\.611\_\{\\pm 0\.008\}20\.57±0\.51\\mathbf\{20\.57\}\_\{\\pm 0\.51\}0\.158±0\.0070\.158\_\{\\pm 0\.007\}6\.73±0\.29\\mathbf\{6\.73\}\_\{\\pm 0\.29\}Generated personas\+\+Creativity\-enhancedpromptMCMC0\.193±0\.0040\.193\_\{\\pm 0\.004\}0\.615±0\.0080\.615\_\{\\pm 0\.008\}21\.16±0\.4121\.16\_\{\\pm 0\.41\}0\.156±0\.0050\.156\_\{\\pm 0\.005\}6\.75±0\.216\.75\_\{\\pm 0\.21\}Evolution0\.197±0\.0050\.197\_\{\\pm 0\.005\}0\.621±0\.0090\.621\_\{\\pm 0\.009\}21\.20±0\.2521\.20\_\{\\pm 0\.25\}0\.157±0\.005\\mathbf\{0\.157\}\_\{\\pm 0\.005\}6\.11±0\.226\.11\_\{\\pm 0\.22\}AUT\-Evolution0\.199±0\.005\\mathbf\{0\.199\}\_\{\\pm 0\.005\}0\.625±0\.008\\mathbf\{0\.625\}\_\{\\pm 0\.008\}21\.86±0\.65\\mathbf\{21\.86\}\_\{\\pm 0\.65\}0\.156±0\.0050\.156\_\{\\pm 0\.005\}7\.00±0\.26\\mathbf\{7\.00\}\_\{\\pm 0\.26\}Table 13:Response diversity restricted to the three objects with human reference data \(book, fork, tin can\), so that all variants are directly comparable toStevenson\. Cells, bolding, and column conventions are as in Table[12](https://arxiv.org/html/2609.30492#A5.T12)\.Mean pairwise distanceConvex hull extentVariantCosineL2Mahal\.Cosine/L2Mahal\.Standard personaCommon\-Use0\.182±0\.0050\.182\_\{\\pm 0\.005\}0\.583±0\.0100\.583\_\{\\pm 0\.010\}15\.49±0\.6615\.49\_\{\\pm 0\.66\}0\.124±0\.0090\.124\_\{\\pm 0\.009\}2\.16±0\.382\.16\_\{\\pm 0\.38\}Alternative\-Use0\.195±0\.003\\mathbf\{0\.195\}\_\{\\pm 0\.003\}0\.601±0\.005\\mathbf\{0\.601\}\_\{\\pm 0\.005\}11\.38±0\.3911\.38\_\{\\pm 0\.39\}0\.173±0\.0120\.173\_\{\\pm 0\.012\}1\.20±0\.131\.20\_\{\\pm 0\.13\}Expert0\.185±0\.0050\.185\_\{\\pm 0\.005\}0\.590±0\.0050\.590\_\{\\pm 0\.005\}13\.76±0\.3213\.76\_\{\\pm 0\.32\}0\.170±0\.012\\mathbf\{0\.170\}\_\{\\pm 0\.012\}2\.01±0\.262\.01\_\{\\pm 0\.26\}Creativity\-enhanced0\.174±0\.0060\.174\_\{\\pm 0\.006\}0\.583±0\.0090\.583\_\{\\pm 0\.009\}19\.54±0\.31\\mathbf\{19\.54\}\_\{\\pm 0\.31\}0\.138±0\.0060\.138\_\{\\pm 0\.006\}5\.06±0\.40\\mathbf\{5\.06\}\_\{\\pm 0\.40\}ZS\-CoT0\.162±0\.0060\.162\_\{\\pm 0\.006\}0\.556±0\.0110\.556\_\{\\pm 0\.011\}13\.62±0\.4613\.62\_\{\\pm 0\.46\}0\.159±0\.0140\.159\_\{\\pm 0\.014\}1\.99±0\.141\.99\_\{\\pm 0\.14\}Step\-Back0\.180±0\.0040\.180\_\{\\pm 0\.004\}0\.584±0\.0090\.584\_\{\\pm 0\.009\}12\.79±0\.2012\.79\_\{\\pm 0\.20\}0\.162±0\.0120\.162\_\{\\pm 0\.012\}1\.77±0\.181\.77\_\{\\pm 0\.18\}DMAD0\.178±0\.0090\.178\_\{\\pm 0\.009\}0\.585±0\.0160\.585\_\{\\pm 0\.016\}14\.98±0\.7714\.98\_\{\\pm 0\.77\}0\.168±0\.0080\.168\_\{\\pm 0\.008\}2\.28±0\.192\.28\_\{\\pm 0\.19\}Gibberish0\.235±0\.0090\.235\_\{\\pm 0\.009\}0\.675±0\.0160\.675\_\{\\pm 0\.016\}20\.50±2\.5820\.50\_\{\\pm 2\.58\}0\.178±0\.0130\.178\_\{\\pm 0\.013\}3\.27±0\.633\.27\_\{\\pm 0\.63\}Baseline personasMPAQ0\.216±0\.0070\.216\_\{\\pm 0\.007\}0\.651±0\.0110\.651\_\{\\pm 0\.011\}17\.70±0\.5917\.70\_\{\\pm 0\.59\}0\.198±0\.0080\.198\_\{\\pm 0\.008\}2\.92±0\.382\.92\_\{\\pm 0\.38\}Individualized personasRandom0\.206±0\.0040\.206\_\{\\pm 0\.004\}0\.629±0\.0070\.629\_\{\\pm 0\.007\}12\.52±0\.6612\.52\_\{\\pm 0\.66\}0\.213±0\.0080\.213\_\{\\pm 0\.008\}1\.48±0\.151\.48\_\{\\pm 0\.15\}Typical0\.206±0\.0060\.206\_\{\\pm 0\.006\}0\.629±0\.0100\.629\_\{\\pm 0\.010\}12\.70±0\.3112\.70\_\{\\pm 0\.31\}0\.216±0\.0130\.216\_\{\\pm 0\.013\}1\.58±0\.211\.58\_\{\\pm 0\.21\}Selected personas from base poolCoverage, cosine/L20\.214±0\.008\\mathbf\{0\.214\}\_\{\\pm 0\.008\}0\.639±0\.013\\mathbf\{0\.639\}\_\{\\pm 0\.013\}12\.66±0\.6312\.66\_\{\\pm 0\.63\}0\.214±0\.0080\.214\_\{\\pm 0\.008\}1\.57±0\.161\.57\_\{\\pm 0\.16\}Coverage, Mahalanobis0\.205±0\.0080\.205\_\{\\pm 0\.008\}0\.624±0\.0160\.624\_\{\\pm 0\.016\}11\.90±0\.4911\.90\_\{\\pm 0\.49\}0\.214±0\.0090\.214\_\{\\pm 0\.009\}1\.52±0\.151\.52\_\{\\pm 0\.15\}Dispersion, cosine/L20\.209±0\.0020\.209\_\{\\pm 0\.002\}0\.638±0\.0030\.638\_\{\\pm 0\.003\}13\.11±0\.45\\mathbf\{13\.11\}\_\{\\pm 0\.45\}0\.220±0\.0050\.220\_\{\\pm 0\.005\}1\.87±0\.331\.87\_\{\\pm 0\.33\}Dispersion, Mahalanobis0\.208±0\.0030\.208\_\{\\pm 0\.003\}0\.634±0\.0040\.634\_\{\\pm 0\.004\}12\.87±0\.3912\.87\_\{\\pm 0\.39\}0\.224±0\.003\\mathbf\{0\.224\}\_\{\\pm 0\.003\}1\.95±0\.32\\mathbf\{1\.95\}\_\{\\pm 0\.32\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.211±0\.0090\.211\_\{\\pm 0\.009\}0\.639±0\.0140\.639\_\{\\pm 0\.014\}13\.92±1\.3513\.92\_\{\\pm 1\.35\}0\.223±0\.0070\.223\_\{\\pm 0\.007\}2\.21±0\.482\.21\_\{\\pm 0\.48\}Evolution \(1,869\)0\.214±0\.0080\.214\_\{\\pm 0\.008\}0\.646±0\.0120\.646\_\{\\pm 0\.012\}15\.41±0\.6115\.41\_\{\\pm 0\.61\}0\.212±0\.0030\.212\_\{\\pm 0\.003\}2\.56±0\.412\.56\_\{\\pm 0\.41\}AUT\-Evolution \(3,397\)0\.222±0\.006\\mathbf\{0\.222\}\_\{\\pm 0\.006\}0\.658±0\.011\\mathbf\{0\.658\}\_\{\\pm 0\.011\}16\.03±0\.56\\mathbf\{16\.03\}\_\{\\pm 0\.56\}0\.227±0\.007\\mathbf\{0\.227\}\_\{\\pm 0\.007\}3\.08±0\.34\\mathbf\{3\.08\}\_\{\\pm 0\.34\}Selected personas\+\+Creativity\-enhancedpromptCoverage, cosine/L20\.186±0\.0050\.186\_\{\\pm 0\.005\}0\.602±0\.0090\.602\_\{\\pm 0\.009\}19\.84±0\.6019\.84\_\{\\pm 0\.60\}0\.145±0\.0060\.145\_\{\\pm 0\.006\}5\.08±0\.325\.08\_\{\\pm 0\.32\}Coverage, Mahalanobis0\.186±0\.0050\.186\_\{\\pm 0\.005\}0\.604±0\.0080\.604\_\{\\pm 0\.008\}19\.94±0\.3919\.94\_\{\\pm 0\.39\}0\.144±0\.0050\.144\_\{\\pm 0\.005\}5\.22±0\.255\.22\_\{\\pm 0\.25\}Dispersion, cosine/L20\.190±0\.005\\mathbf\{0\.190\}\_\{\\pm 0\.005\}0\.608±0\.010\\mathbf\{0\.608\}\_\{\\pm 0\.010\}19\.85±0\.4619\.85\_\{\\pm 0\.46\}0\.153±0\.010\\mathbf\{0\.153\}\_\{\\pm 0\.010\}6\.17±0\.526\.17\_\{\\pm 0\.52\}Dispersion, Mahalanobis0\.184±0\.0020\.184\_\{\\pm 0\.002\}0\.600±0\.0040\.600\_\{\\pm 0\.004\}20\.18±0\.60\\mathbf\{20\.18\}\_\{\\pm 0\.60\}0\.148±0\.0050\.148\_\{\\pm 0\.005\}6\.59±0\.38\\mathbf\{6\.59\}\_\{\\pm 0\.38\}Generated personas\+\+Creativity\-enhancedpromptMCMC0\.186±0\.0030\.186\_\{\\pm 0\.003\}0\.604±0\.0060\.604\_\{\\pm 0\.006\}20\.62±0\.2420\.62\_\{\\pm 0\.24\}0\.147±0\.0040\.147\_\{\\pm 0\.004\}6\.58±0\.286\.58\_\{\\pm 0\.28\}Evolution0\.190±0\.0050\.190\_\{\\pm 0\.005\}0\.610±0\.0090\.610\_\{\\pm 0\.009\}20\.66±0\.7320\.66\_\{\\pm 0\.73\}0\.148±0\.0110\.148\_\{\\pm 0\.011\}5\.98±0\.495\.98\_\{\\pm 0\.49\}AUT\-Evolution0\.196±0\.006\\mathbf\{0\.196\}\_\{\\pm 0\.006\}0\.620±0\.011\\mathbf\{0\.620\}\_\{\\pm 0\.011\}21\.48±0\.75\\mathbf\{21\.48\}\_\{\\pm 0\.75\}0\.148±0\.003\\mathbf\{0\.148\}\_\{\\pm 0\.003\}6\.92±0\.40\\mathbf\{6\.92\}\_\{\\pm 0\.40\}HumanStevenson0\.149±0\.0090\.149\_\{\\pm 0\.009\}0\.539±0\.0150\.539\_\{\\pm 0\.015\}17\.80±0\.9317\.80\_\{\\pm 0\.93\}0\.101±0\.0060\.101\_\{\\pm 0\.006\}2\.26±0\.342\.26\_\{\\pm 0\.34\}Table 14:Response originality across all seven AUT objects, measured both as distance to the prompted object and as directed distance to the common\-use set, following equation[57](https://arxiv.org/html/2609.30492#A4.E57)\. Cells give the mean over valid uses, averaged over objects; higher is more original\.Common\-Usedefines the common\-use reference setCoC\_\{o\}, so its distance to that set is zero by construction and is omitted \(—\)\. Subscripts denote the 95% confidence interval across 5 independent random seeds\. The final two groups re\-run our selection and generation methods on theCreativity\-enhancedprompt, showing that persona diversification composes with a stronger prompting strategy\. Bold marks the best value per column within the standard, selected and generated groups separately, with tied values bolded jointly; baseline and individualized personas are not bolded\. Generated variants report their expanded pool size in parentheses\.To objectTo common usesVariantCosineL2Mahal\.CosineL2Mahal\.Standard personaCommon\-Use0\.142±0\.0020\.142\_\{\\pm 0\.002\}0\.525±0\.0040\.525\_\{\\pm 0\.004\}15\.28±0\.1615\.28\_\{\\pm 0\.16\}———Alternative\-Use0\.170±0\.0010\.170\_\{\\pm 0\.001\}0\.580±0\.0020\.580\_\{\\pm 0\.002\}13\.60±0\.1413\.60\_\{\\pm 0\.14\}0\.114±0\.0080\.114\_\{\\pm 0\.008\}0\.439±0\.0190\.439\_\{\\pm 0\.019\}10\.79±0\.3210\.79\_\{\\pm 0\.32\}Expert0\.164±0\.0020\.164\_\{\\pm 0\.002\}0\.568±0\.0040\.568\_\{\\pm 0\.004\}14\.72±0\.1714\.72\_\{\\pm 0\.17\}0\.104±0\.0030\.104\_\{\\pm 0\.003\}0\.417±0\.0100\.417\_\{\\pm 0\.010\}11\.25±0\.4011\.25\_\{\\pm 0\.40\}Creativity\-enhanced0\.173±0\.0040\.173\_\{\\pm 0\.004\}0\.586±0\.0070\.586\_\{\\pm 0\.007\}18\.13±0\.35\\mathbf\{18\.13\}\_\{\\pm 0\.35\}0\.174±0\.006\\mathbf\{0\.174\}\_\{\\pm 0\.006\}0\.585±0\.011\\mathbf\{0\.585\}\_\{\\pm 0\.011\}17\.15±0\.40\\mathbf\{17\.15\}\_\{\\pm 0\.40\}ZS\-CoT0\.146±0\.0020\.146\_\{\\pm 0\.002\}0\.536±0\.0040\.536\_\{\\pm 0\.004\}14\.67±0\.1014\.67\_\{\\pm 0\.10\}0\.125±0\.0030\.125\_\{\\pm 0\.003\}0\.483±0\.0070\.483\_\{\\pm 0\.007\}13\.02±0\.4313\.02\_\{\\pm 0\.43\}Step\-Back0\.165±0\.0020\.165\_\{\\pm 0\.002\}0\.571±0\.0030\.571\_\{\\pm 0\.003\}14\.91±0\.1814\.91\_\{\\pm 0\.18\}0\.129±0\.0020\.129\_\{\\pm 0\.002\}0\.488±0\.0070\.488\_\{\\pm 0\.007\}12\.86±0\.2012\.86\_\{\\pm 0\.20\}DMAD0\.151±0\.0040\.151\_\{\\pm 0\.004\}0\.547±0\.0070\.547\_\{\\pm 0\.007\}15\.19±0\.1915\.19\_\{\\pm 0\.19\}0\.135±0\.0060\.135\_\{\\pm 0\.006\}0\.504±0\.0170\.504\_\{\\pm 0\.017\}13\.90±0\.4913\.90\_\{\\pm 0\.49\}Gibberish0\.193±0\.002\\mathbf\{0\.193\}\_\{\\pm 0\.002\}0\.618±0\.004\\mathbf\{0\.618\}\_\{\\pm 0\.004\}17\.72±0\.5217\.72\_\{\\pm 0\.52\}0\.156±0\.0090\.156\_\{\\pm 0\.009\}0\.541±0\.0180\.541\_\{\\pm 0\.018\}15\.89±0\.6815\.89\_\{\\pm 0\.68\}Baseline personasMPAQ0\.180±0\.0010\.180\_\{\\pm 0\.001\}0\.596±0\.0020\.596\_\{\\pm 0\.002\}16\.95±0\.2816\.95\_\{\\pm 0\.28\}0\.162±0\.0020\.162\_\{\\pm 0\.002\}0\.560±0\.0040\.560\_\{\\pm 0\.004\}15\.46±0\.3415\.46\_\{\\pm 0\.34\}Individualized personasRandom0\.184±0\.0040\.184\_\{\\pm 0\.004\}0\.603±0\.0070\.603\_\{\\pm 0\.007\}13\.90±0\.3113\.90\_\{\\pm 0\.31\}0\.143±0\.0070\.143\_\{\\pm 0\.007\}0\.516±0\.0150\.516\_\{\\pm 0\.015\}11\.89±0\.4811\.89\_\{\\pm 0\.48\}Typical0\.179±0\.0030\.179\_\{\\pm 0\.003\}0\.595±0\.0050\.595\_\{\\pm 0\.005\}13\.91±0\.1313\.91\_\{\\pm 0\.13\}0\.139±0\.0050\.139\_\{\\pm 0\.005\}0\.509±0\.0110\.509\_\{\\pm 0\.011\}11\.96±0\.3011\.96\_\{\\pm 0\.30\}Selected personas from base poolCoverage, cosine0\.188±0\.005\\mathbf\{0\.188\}\_\{\\pm 0\.005\}0\.609±0\.008\\mathbf\{0\.609\}\_\{\\pm 0\.008\}13\.91±0\.1413\.91\_\{\\pm 0\.14\}0\.145±0\.004\\mathbf\{0\.145\}\_\{\\pm 0\.004\}0\.520±0\.008\\mathbf\{0\.520\}\_\{\\pm 0\.008\}12\.03±0\.3712\.03\_\{\\pm 0\.37\}Coverage, Mahalanobis0\.179±0\.0010\.179\_\{\\pm 0\.001\}0\.594±0\.0020\.594\_\{\\pm 0\.002\}13\.51±0\.1113\.51\_\{\\pm 0\.11\}0\.133±0\.0060\.133\_\{\\pm 0\.006\}0\.495±0\.0120\.495\_\{\\pm 0\.012\}11\.46±0\.2811\.46\_\{\\pm 0\.28\}Dispersion, cosine0\.184±0\.0020\.184\_\{\\pm 0\.002\}0\.603±0\.0030\.603\_\{\\pm 0\.003\}14\.20±0\.05\\mathbf\{14\.20\}\_\{\\pm 0\.05\}0\.146±0\.0040\.146\_\{\\pm 0\.004\}0\.524±0\.0100\.524\_\{\\pm 0\.010\}12\.34±0\.27\\mathbf\{12\.34\}\_\{\\pm 0\.27\}Dispersion, Mahalanobis0\.181±0\.0030\.181\_\{\\pm 0\.003\}0\.598±0\.0040\.598\_\{\\pm 0\.004\}13\.98±0\.1813\.98\_\{\\pm 0\.18\}0\.138±0\.0050\.138\_\{\\pm 0\.005\}0\.506±0\.0090\.506\_\{\\pm 0\.009\}12\.00±0\.2812\.00\_\{\\pm 0\.28\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.182±0\.0020\.182\_\{\\pm 0\.002\}0\.600±0\.0030\.600\_\{\\pm 0\.003\}14\.51±0\.3014\.51\_\{\\pm 0\.30\}0\.144±0\.0050\.144\_\{\\pm 0\.005\}0\.519±0\.0140\.519\_\{\\pm 0\.014\}12\.66±0\.5512\.66\_\{\\pm 0\.55\}Evolution \(1,869\)0\.190±0\.004\\mathbf\{0\.190\}\_\{\\pm 0\.004\}0\.612±0\.007\\mathbf\{0\.612\}\_\{\\pm 0\.007\}15\.22±0\.1915\.22\_\{\\pm 0\.19\}0\.162±0\.003\\mathbf\{0\.162\}\_\{\\pm 0\.003\}0\.556±0\.006\\mathbf\{0\.556\}\_\{\\pm 0\.006\}13\.61±0\.17\\mathbf\{13\.61\}\_\{\\pm 0\.17\}AUT\-Evolution \(3,397\)0\.188±0\.0020\.188\_\{\\pm 0\.002\}0\.610±0\.0030\.610\_\{\\pm 0\.003\}15\.29±0\.12\\mathbf\{15\.29\}\_\{\\pm 0\.12\}0\.156±0\.0040\.156\_\{\\pm 0\.004\}0\.544±0\.0060\.544\_\{\\pm 0\.006\}13\.58±0\.1113\.58\_\{\\pm 0\.11\}Selected personas\+\+Creativity\-enhancedpromptCoverage, cosine0\.187±0\.003\\mathbf\{0\.187\}\_\{\\pm 0\.003\}0\.609±0\.005\\mathbf\{0\.609\}\_\{\\pm 0\.005\}18\.45±0\.2018\.45\_\{\\pm 0\.20\}0\.181±0\.003\\mathbf\{0\.181\}\_\{\\pm 0\.003\}0\.598±0\.005\\mathbf\{0\.598\}\_\{\\pm 0\.005\}17\.25±0\.1217\.25\_\{\\pm 0\.12\}Coverage, Mahalanobis0\.184±0\.0030\.184\_\{\\pm 0\.003\}0\.603±0\.0050\.603\_\{\\pm 0\.005\}18\.33±0\.1118\.33\_\{\\pm 0\.11\}0\.177±0\.0040\.177\_\{\\pm 0\.004\}0\.590±0\.0060\.590\_\{\\pm 0\.006\}17\.20±0\.1517\.20\_\{\\pm 0\.15\}Dispersion, cosine0\.185±0\.0050\.185\_\{\\pm 0\.005\}0\.604±0\.0080\.604\_\{\\pm 0\.008\}18\.31±0\.2718\.31\_\{\\pm 0\.27\}0\.179±0\.0070\.179\_\{\\pm 0\.007\}0\.593±0\.0130\.593\_\{\\pm 0\.013\}17\.24±0\.26\\mathbf\{17\.24\}\_\{\\pm 0\.26\}Dispersion, Mahalanobis0\.182±0\.0050\.182\_\{\\pm 0\.005\}0\.600±0\.0080\.600\_\{\\pm 0\.008\}18\.36±0\.36\\mathbf\{18\.36\}\_\{\\pm 0\.36\}0\.178±0\.0070\.178\_\{\\pm 0\.007\}0\.592±0\.0110\.592\_\{\\pm 0\.011\}17\.22±0\.3917\.22\_\{\\pm 0\.39\}Generated personas\+\+Creativity\-enhancedpromptMCMC0\.183±0\.0050\.183\_\{\\pm 0\.005\}0\.602±0\.0090\.602\_\{\\pm 0\.009\}18\.71±0\.2218\.71\_\{\\pm 0\.22\}0\.182±0\.0060\.182\_\{\\pm 0\.006\}0\.598±0\.0090\.598\_\{\\pm 0\.009\}17\.64±0\.2117\.64\_\{\\pm 0\.21\}Evolution0\.187±0\.005\\mathbf\{0\.187\}\_\{\\pm 0\.005\}0\.608±0\.008\\mathbf\{0\.608\}\_\{\\pm 0\.008\}18\.80±0\.1718\.80\_\{\\pm 0\.17\}0\.183±0\.0060\.183\_\{\\pm 0\.006\}0\.600±0\.0110\.600\_\{\\pm 0\.011\}17\.65±0\.1417\.65\_\{\\pm 0\.14\}AUT\-Evolution0\.187±0\.006\\mathbf\{0\.187\}\_\{\\pm 0\.006\}0\.608±0\.009\\mathbf\{0\.608\}\_\{\\pm 0\.009\}19\.14±0\.41\\mathbf\{19\.14\}\_\{\\pm 0\.41\}0\.186±0\.009\\mathbf\{0\.186\}\_\{\\pm 0\.009\}0\.605±0\.014\\mathbf\{0\.605\}\_\{\\pm 0\.014\}18\.02±0\.39\\mathbf\{18\.02\}\_\{\\pm 0\.39\}

Table 15:Response originality restricted to the three objects with human reference data \(book, fork, tin can\), so that all variants are directly comparable toStevenson\. Cells, bolding, and column conventions are as in Table[14](https://arxiv.org/html/2609.30492#A5.T14)\.Common\-Usedefines the common\-use reference setCoC\_\{o\}, so its distance to that set is zero by construction and is omitted \(—\)\.To objectTo common usesVariantCosineL2Mahal\.CosineL2Mahal\.Standard personaCommon\-Use0\.182±0\.0050\.182\_\{\\pm 0\.005\}0\.583±0\.0100\.583\_\{\\pm 0\.010\}15\.49±0\.6615\.49\_\{\\pm 0\.66\}———Alternative\-Use0\.195±0\.003\\mathbf\{0\.195\}\_\{\\pm 0\.003\}0\.601±0\.005\\mathbf\{0\.601\}\_\{\\pm 0\.005\}11\.38±0\.3911\.38\_\{\\pm 0\.39\}0\.173±0\.0120\.173\_\{\\pm 0\.012\}1\.20±0\.131\.20\_\{\\pm 0\.13\}10\.79±0\.3210\.79\_\{\\pm 0\.32\}Expert0\.185±0\.0050\.185\_\{\\pm 0\.005\}0\.590±0\.0050\.590\_\{\\pm 0\.005\}13\.76±0\.3213\.76\_\{\\pm 0\.32\}0\.170±0\.012\\mathbf\{0\.170\}\_\{\\pm 0\.012\}2\.01±0\.262\.01\_\{\\pm 0\.26\}11\.25±0\.4011\.25\_\{\\pm 0\.40\}Creativity\-enhanced0\.174±0\.0060\.174\_\{\\pm 0\.006\}0\.583±0\.0090\.583\_\{\\pm 0\.009\}19\.54±0\.31\\mathbf\{19\.54\}\_\{\\pm 0\.31\}0\.138±0\.0060\.138\_\{\\pm 0\.006\}5\.06±0\.40\\mathbf\{5\.06\}\_\{\\pm 0\.40\}17\.15±0\.40\\mathbf\{17\.15\}\_\{\\pm 0\.40\}ZS\-CoT0\.162±0\.0060\.162\_\{\\pm 0\.006\}0\.556±0\.0110\.556\_\{\\pm 0\.011\}13\.62±0\.4613\.62\_\{\\pm 0\.46\}0\.159±0\.0140\.159\_\{\\pm 0\.014\}1\.99±0\.141\.99\_\{\\pm 0\.14\}13\.02±0\.4313\.02\_\{\\pm 0\.43\}Step\-Back0\.180±0\.0040\.180\_\{\\pm 0\.004\}0\.584±0\.0090\.584\_\{\\pm 0\.009\}12\.79±0\.2012\.79\_\{\\pm 0\.20\}0\.162±0\.0120\.162\_\{\\pm 0\.012\}1\.77±0\.181\.77\_\{\\pm 0\.18\}12\.86±0\.2012\.86\_\{\\pm 0\.20\}DMAD0\.178±0\.0090\.178\_\{\\pm 0\.009\}0\.585±0\.0160\.585\_\{\\pm 0\.016\}14\.98±0\.7714\.98\_\{\\pm 0\.77\}0\.168±0\.0080\.168\_\{\\pm 0\.008\}2\.28±0\.192\.28\_\{\\pm 0\.19\}13\.90±0\.4913\.90\_\{\\pm 0\.49\}Gibberish0\.235±0\.0090\.235\_\{\\pm 0\.009\}0\.675±0\.0160\.675\_\{\\pm 0\.016\}20\.50±2\.5820\.50\_\{\\pm 2\.58\}0\.178±0\.0130\.178\_\{\\pm 0\.013\}3\.27±0\.633\.27\_\{\\pm 0\.63\}15\.89±0\.6815\.89\_\{\\pm 0\.68\}Baseline personasMPAQ0\.216±0\.0070\.216\_\{\\pm 0\.007\}0\.651±0\.0110\.651\_\{\\pm 0\.011\}17\.70±0\.5917\.70\_\{\\pm 0\.59\}0\.198±0\.0080\.198\_\{\\pm 0\.008\}2\.92±0\.382\.92\_\{\\pm 0\.38\}15\.46±0\.3415\.46\_\{\\pm 0\.34\}Individualized personasRandom0\.206±0\.0040\.206\_\{\\pm 0\.004\}0\.629±0\.0070\.629\_\{\\pm 0\.007\}12\.52±0\.6612\.52\_\{\\pm 0\.66\}0\.213±0\.0080\.213\_\{\\pm 0\.008\}1\.48±0\.151\.48\_\{\\pm 0\.15\}11\.89±0\.4811\.89\_\{\\pm 0\.48\}Typical0\.206±0\.0060\.206\_\{\\pm 0\.006\}0\.629±0\.0100\.629\_\{\\pm 0\.010\}12\.70±0\.3112\.70\_\{\\pm 0\.31\}0\.216±0\.0130\.216\_\{\\pm 0\.013\}1\.58±0\.211\.58\_\{\\pm 0\.21\}11\.96±0\.3011\.96\_\{\\pm 0\.30\}Selected personas from base poolCoverage, cosine0\.214±0\.008\\mathbf\{0\.214\}\_\{\\pm 0\.008\}0\.639±0\.013\\mathbf\{0\.639\}\_\{\\pm 0\.013\}12\.66±0\.6312\.66\_\{\\pm 0\.63\}0\.214±0\.0080\.214\_\{\\pm 0\.008\}1\.57±0\.161\.57\_\{\\pm 0\.16\}12\.03±0\.3712\.03\_\{\\pm 0\.37\}Coverage, Mahalanobis0\.205±0\.0080\.205\_\{\\pm 0\.008\}0\.624±0\.0160\.624\_\{\\pm 0\.016\}11\.90±0\.4911\.90\_\{\\pm 0\.49\}0\.214±0\.0090\.214\_\{\\pm 0\.009\}1\.52±0\.151\.52\_\{\\pm 0\.15\}11\.46±0\.2811\.46\_\{\\pm 0\.28\}Dispersion, cosine0\.209±0\.0020\.209\_\{\\pm 0\.002\}0\.638±0\.0030\.638\_\{\\pm 0\.003\}13\.11±0\.45\\mathbf\{13\.11\}\_\{\\pm 0\.45\}0\.220±0\.0050\.220\_\{\\pm 0\.005\}1\.87±0\.331\.87\_\{\\pm 0\.33\}12\.34±0\.27\\mathbf\{12\.34\}\_\{\\pm 0\.27\}Dispersion, Mahalanobis0\.208±0\.0030\.208\_\{\\pm 0\.003\}0\.634±0\.0040\.634\_\{\\pm 0\.004\}12\.87±0\.3912\.87\_\{\\pm 0\.39\}0\.224±0\.003\\mathbf\{0\.224\}\_\{\\pm 0\.003\}1\.95±0\.32\\mathbf\{1\.95\}\_\{\\pm 0\.32\}12\.00±0\.2812\.00\_\{\\pm 0\.28\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.211±0\.0090\.211\_\{\\pm 0\.009\}0\.639±0\.0140\.639\_\{\\pm 0\.014\}13\.92±1\.3513\.92\_\{\\pm 1\.35\}0\.223±0\.007\\mathbf\{0\.223\}\_\{\\pm 0\.007\}2\.21±0\.482\.21\_\{\\pm 0\.48\}12\.66±0\.5512\.66\_\{\\pm 0\.55\}Evolution \(1,869\)0\.214±0\.0080\.214\_\{\\pm 0\.008\}0\.646±0\.0120\.646\_\{\\pm 0\.012\}15\.41±0\.6115\.41\_\{\\pm 0\.61\}0\.212±0\.0030\.212\_\{\\pm 0\.003\}2\.56±0\.412\.56\_\{\\pm 0\.41\}13\.61±0\.17\\mathbf\{13\.61\}\_\{\\pm 0\.17\}AUT\-Evolution \(3,397\)0\.222±0\.006\\mathbf\{0\.222\}\_\{\\pm 0\.006\}0\.658±0\.011\\mathbf\{0\.658\}\_\{\\pm 0\.011\}16\.03±0\.56\\mathbf\{16\.03\}\_\{\\pm 0\.56\}0\.227±0\.0070\.227\_\{\\pm 0\.007\}3\.08±0\.34\\mathbf\{3\.08\}\_\{\\pm 0\.34\}13\.58±0\.1113\.58\_\{\\pm 0\.11\}Selected personas\+\+Creativity\-enhancedpromptCoverage, cosine0\.186±0\.0050\.186\_\{\\pm 0\.005\}0\.602±0\.0090\.602\_\{\\pm 0\.009\}19\.84±0\.6019\.84\_\{\\pm 0\.60\}0\.145±0\.0060\.145\_\{\\pm 0\.006\}5\.08±0\.325\.08\_\{\\pm 0\.32\}17\.25±0\.1217\.25\_\{\\pm 0\.12\}Coverage, Mahalanobis0\.186±0\.0050\.186\_\{\\pm 0\.005\}0\.604±0\.0080\.604\_\{\\pm 0\.008\}19\.94±0\.3919\.94\_\{\\pm 0\.39\}0\.144±0\.0050\.144\_\{\\pm 0\.005\}5\.22±0\.255\.22\_\{\\pm 0\.25\}17\.20±0\.1517\.20\_\{\\pm 0\.15\}Dispersion, cosine0\.190±0\.005\\mathbf\{0\.190\}\_\{\\pm 0\.005\}0\.608±0\.010\\mathbf\{0\.608\}\_\{\\pm 0\.010\}19\.85±0\.4619\.85\_\{\\pm 0\.46\}0\.153±0\.010\\mathbf\{0\.153\}\_\{\\pm 0\.010\}6\.17±0\.526\.17\_\{\\pm 0\.52\}17\.24±0\.26\\mathbf\{17\.24\}\_\{\\pm 0\.26\}Dispersion, Mahalanobis0\.184±0\.0020\.184\_\{\\pm 0\.002\}0\.600±0\.0040\.600\_\{\\pm 0\.004\}20\.18±0\.60\\mathbf\{20\.18\}\_\{\\pm 0\.60\}0\.148±0\.0050\.148\_\{\\pm 0\.005\}6\.59±0\.38\\mathbf\{6\.59\}\_\{\\pm 0\.38\}17\.22±0\.3917\.22\_\{\\pm 0\.39\}Generated personas\+\+Creativity\-enhancedpromptMCMC0\.186±0\.0030\.186\_\{\\pm 0\.003\}0\.604±0\.0060\.604\_\{\\pm 0\.006\}20\.62±0\.2420\.62\_\{\\pm 0\.24\}0\.147±0\.0040\.147\_\{\\pm 0\.004\}6\.58±0\.28\\mathbf\{6\.58\}\_\{\\pm 0\.28\}17\.64±0\.2117\.64\_\{\\pm 0\.21\}Evolution0\.190±0\.0050\.190\_\{\\pm 0\.005\}0\.610±0\.0090\.610\_\{\\pm 0\.009\}20\.66±0\.7320\.66\_\{\\pm 0\.73\}0\.148±0\.0110\.148\_\{\\pm 0\.011\}5\.98±0\.495\.98\_\{\\pm 0\.49\}17\.65±0\.1417\.65\_\{\\pm 0\.14\}AUT\-Evolution0\.196±0\.006\\mathbf\{0\.196\}\_\{\\pm 0\.006\}0\.620±0\.011\\mathbf\{0\.620\}\_\{\\pm 0\.011\}21\.48±0\.75\\mathbf\{21\.48\}\_\{\\pm 0\.75\}0\.148±0\.003\\mathbf\{0\.148\}\_\{\\pm 0\.003\}6\.92±0\.406\.92\_\{\\pm 0\.40\}18\.02±0\.39\\mathbf\{18\.02\}\_\{\\pm 0\.39\}HumanStevenson0\.119±0\.0000\.119\_\{\\pm 0\.000\}0\.480±0\.0000\.480\_\{\\pm 0\.000\}15\.51±0\.0015\.51\_\{\\pm 0\.00\}0\.144±0\.0000\.144\_\{\\pm 0\.000\}0\.533±0\.0000\.533\_\{\\pm 0\.000\}14\.15±0\.0014\.15\_\{\\pm 0\.00\}

Table 16:Response flexibility, reported over all seven AUT objects and, for comparability withStevenson, over the three objects with human reference data \(175175and7575generated uses per variant, respectively\)\. Both measures are size\-controlled:*Flexibility*is the number of distinct use categories and*Vendi*is the Vendi score equation[53](https://arxiv.org/html/2609.30492#A4.E53), each divided by the number of valid uses, so that variants admitting different numbers of uses remain comparable\. Cells give the mean over objects\. Subscripts denote the 95% confidence interval across 5 independent random seeds; higher is more flexible\.Stevensonexists only for the human\-reference objects, so its seven\-object cells are omitted \(—\)\. The final two groups re\-run our selection and generation methods on theCreativity\-enhancedprompt, showing that persona diversification composes with a stronger prompting strategy\. Bold marks the best value per column within the standard, selected and generated groups separately, with tied values bolded jointly; baseline, individualized and human rows are not bolded\. Generated variants report their expanded pool size in parentheses\.All seven objectsHuman\-reference objectsVariantFlexibilityVendiFlexibilityVendiStandard personaCommon\-Use0\.160±0\.0110\.160\_\{\\pm 0\.011\}0\.086±0\.0030\.086\_\{\\pm 0\.003\}0\.214±0\.0250\.214\_\{\\pm 0\.025\}0\.094±0\.0040\.094\_\{\\pm 0\.004\}Alternative\-Use0\.288±0\.0220\.288\_\{\\pm 0\.022\}0\.096±0\.0020\.096\_\{\\pm 0\.002\}0\.284±0\.0120\.284\_\{\\pm 0\.012\}0\.099±0\.0030\.099\_\{\\pm 0\.003\}Expert0\.300±0\.0260\.300\_\{\\pm 0\.026\}0\.094±0\.0010\.094\_\{\\pm 0\.001\}0\.320±0\.0170\.320\_\{\\pm 0\.017\}0\.096±0\.0010\.096\_\{\\pm 0\.001\}Creativity\-enhanced0\.503±0\.056\\mathbf\{0\.503\}\_\{\\pm 0\.056\}0\.104±0\.0020\.104\_\{\\pm 0\.002\}0\.467±0\.071\\mathbf\{0\.467\}\_\{\\pm 0\.071\}0\.101±0\.0020\.101\_\{\\pm 0\.002\}ZS\-CoT0\.332±0\.0250\.332\_\{\\pm 0\.025\}0\.094±0\.0010\.094\_\{\\pm 0\.001\}0\.323±0\.0440\.323\_\{\\pm 0\.044\}0\.091±0\.0020\.091\_\{\\pm 0\.002\}Step\-Back0\.303±0\.0220\.303\_\{\\pm 0\.022\}0\.096±0\.0040\.096\_\{\\pm 0\.004\}0\.280±0\.0320\.280\_\{\\pm 0\.032\}0\.097±0\.0030\.097\_\{\\pm 0\.003\}DMAD0\.372±0\.0250\.372\_\{\\pm 0\.025\}0\.102±0\.0030\.102\_\{\\pm 0\.003\}0\.350±0\.0340\.350\_\{\\pm 0\.034\}0\.098±0\.0040\.098\_\{\\pm 0\.004\}Gibberish0\.384±0\.0610\.384\_\{\\pm 0\.061\}0\.151±0\.013\\mathbf\{0\.151\}\_\{\\pm 0\.013\}0\.359±0\.0310\.359\_\{\\pm 0\.031\}0\.158±0\.027\\mathbf\{0\.158\}\_\{\\pm 0\.027\}Baseline personasMPAQ0\.561±0\.0300\.561\_\{\\pm 0\.030\}0\.130±0\.0060\.130\_\{\\pm 0\.006\}0\.491±0\.0570\.491\_\{\\pm 0\.057\}0\.129±0\.0060\.129\_\{\\pm 0\.006\}Individualized personasRandom0\.331±0\.0240\.331\_\{\\pm 0\.024\}0\.107±0\.0030\.107\_\{\\pm 0\.003\}0\.333±0\.0310\.333\_\{\\pm 0\.031\}0\.109±0\.0030\.109\_\{\\pm 0\.003\}Typical0\.330±0\.0160\.330\_\{\\pm 0\.016\}0\.105±0\.0020\.105\_\{\\pm 0\.002\}0\.323±0\.0220\.323\_\{\\pm 0\.022\}0\.108±0\.0030\.108\_\{\\pm 0\.003\}Selected personas from base poolCoverage, cosine0\.338±0\.0250\.338\_\{\\pm 0\.025\}0\.110±0\.0030\.110\_\{\\pm 0\.003\}0\.323±0\.0250\.323\_\{\\pm 0\.025\}0\.111±0\.0040\.111\_\{\\pm 0\.004\}Coverage, Mahalanobis0\.337±0\.0110\.337\_\{\\pm 0\.011\}0\.101±0\.0020\.101\_\{\\pm 0\.002\}0\.329±0\.0410\.329\_\{\\pm 0\.041\}0\.105±0\.0050\.105\_\{\\pm 0\.005\}Dispersion, cosine0\.379±0\.026\\mathbf\{0\.379\}\_\{\\pm 0\.026\}0\.112±0\.001\\mathbf\{0\.112\}\_\{\\pm 0\.001\}0\.356±0\.048\\mathbf\{0\.356\}\_\{\\pm 0\.048\}0\.113±0\.001\\mathbf\{0\.113\}\_\{\\pm 0\.001\}Dispersion, Mahalanobis0\.357±0\.0480\.357\_\{\\pm 0\.048\}0\.108±0\.0020\.108\_\{\\pm 0\.002\}0\.353±0\.0510\.353\_\{\\pm 0\.051\}0\.111±0\.0010\.111\_\{\\pm 0\.001\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.378±0\.0520\.378\_\{\\pm 0\.052\}0\.112±0\.0010\.112\_\{\\pm 0\.001\}0\.348±0\.0290\.348\_\{\\pm 0\.029\}0\.113±0\.0050\.113\_\{\\pm 0\.005\}Evolution \(1,869\)0\.431±0\.0450\.431\_\{\\pm 0\.045\}0\.117±0\.0020\.117\_\{\\pm 0\.002\}0\.411±0\.1020\.411\_\{\\pm 0\.102\}0\.117±0\.0030\.117\_\{\\pm 0\.003\}AUT\-Evolution \(3,397\)0\.463±0\.052\\mathbf\{0\.463\}\_\{\\pm 0\.052\}0\.118±0\.002\\mathbf\{0\.118\}\_\{\\pm 0\.002\}0\.490±0\.045\\mathbf\{0\.490\}\_\{\\pm 0\.045\}0\.121±0\.004\\mathbf\{0\.121\}\_\{\\pm 0\.004\}Selected personas\+\+Creativity\-enhancedpromptCoverage, cosine0\.508±0\.0820\.508\_\{\\pm 0\.082\}0\.109±0\.0030\.109\_\{\\pm 0\.003\}0\.463±0\.1010\.463\_\{\\pm 0\.101\}0\.107±0\.0020\.107\_\{\\pm 0\.002\}Coverage, Mahalanobis0\.518±0\.0660\.518\_\{\\pm 0\.066\}0\.108±0\.0010\.108\_\{\\pm 0\.001\}0\.496±0\.0900\.496\_\{\\pm 0\.090\}0\.107±0\.0030\.107\_\{\\pm 0\.003\}Dispersion, cosine0\.543±0\.0520\.543\_\{\\pm 0\.052\}0\.111±0\.003\\mathbf\{0\.111\}\_\{\\pm 0\.003\}0\.514±0\.0940\.514\_\{\\pm 0\.094\}0\.109±0\.003\\mathbf\{0\.109\}\_\{\\pm 0\.003\}Dispersion, Mahalanobis0\.560±0\.051\\mathbf\{0\.560\}\_\{\\pm 0\.051\}0\.109±0\.0030\.109\_\{\\pm 0\.003\}0\.526±0\.117\\mathbf\{0\.526\}\_\{\\pm 0\.117\}0\.106±0\.0010\.106\_\{\\pm 0\.001\}Generated personas\+\+Creativity\-enhancedpromptMCMC0\.560±0\.0450\.560\_\{\\pm 0\.045\}0\.111±0\.0020\.111\_\{\\pm 0\.002\}0\.526±0\.1080\.526\_\{\\pm 0\.108\}0\.108±0\.0020\.108\_\{\\pm 0\.002\}Evolution0\.580±0\.0460\.580\_\{\\pm 0\.046\}0\.112±0\.0030\.112\_\{\\pm 0\.003\}0\.562±0\.0620\.562\_\{\\pm 0\.062\}0\.110±0\.0040\.110\_\{\\pm 0\.004\}AUT\-Evolution0\.612±0\.026\\mathbf\{0\.612\}\_\{\\pm 0\.026\}0\.114±0\.003\\mathbf\{0\.114\}\_\{\\pm 0\.003\}0\.606±0\.123\\mathbf\{0\.606\}\_\{\\pm 0\.123\}0\.113±0\.003\\mathbf\{0\.113\}\_\{\\pm 0\.003\}HumanStevenson——0\.739±0\.0430\.739\_\{\\pm 0\.043\}0\.104±0\.0010\.104\_\{\\pm 0\.001\}Table 17:Response surprise, reported over all seven AUT objects and, for comparability withStevenson, over the three objects with human reference data \(175175and7575generated uses per variant, respectively\)\. Both measures are size\-controlled by the number of valid uses:*Surprise*is the proportion of uses judged surprising relative to the common\-use reference set, and*Wtd\. Vendi*is the novelty\-weighted Vendi score equation[53](https://arxiv.org/html/2609.30492#A4.E53)under the same normalization\.Common\-Usedefines the reference set against which surprise is measured, so its surprise is zero by construction and is omitted \(—\);Stevensonexists only for the human\-reference objects\. Cells give the mean over objects\. Subscripts denote the 95% confidence interval across 5 independent random seeds; higher is more surprising\. The final two groups re\-run our selection and generation methods on theCreativity\-enhancedprompt, showing that persona diversification composes with a stronger prompting strategy\. Bolding conventions are as in Table[16](https://arxiv.org/html/2609.30492#A5.T16)\.All seven objectsHuman\-reference objectsVariantSurpriseWtd\. VendiSurpriseWtd\. VendiStandard personaCommon\-Use————Alternative\-Use0\.506±0\.0620\.506\_\{\\pm 0\.062\}0\.094±0\.0020\.094\_\{\\pm 0\.002\}0\.397±0\.0870\.397\_\{\\pm 0\.087\}0\.093±0\.0040\.093\_\{\\pm 0\.004\}Expert0\.415±0\.0200\.415\_\{\\pm 0\.020\}0\.093±0\.0020\.093\_\{\\pm 0\.002\}0\.317±0\.0140\.317\_\{\\pm 0\.014\}0\.098±0\.0020\.098\_\{\\pm 0\.002\}Creativity\-enhanced0\.837±0\.067\\mathbf\{0\.837\}\_\{\\pm 0\.067\}0\.105±0\.0020\.105\_\{\\pm 0\.002\}0\.717±0\.131\\mathbf\{0\.717\}\_\{\\pm 0\.131\}0\.102±0\.0020\.102\_\{\\pm 0\.002\}ZS\-CoT0\.556±0\.0640\.556\_\{\\pm 0\.064\}0\.094±0\.0020\.094\_\{\\pm 0\.002\}0\.390±0\.0630\.390\_\{\\pm 0\.063\}0\.089±0\.0030\.089\_\{\\pm 0\.003\}Step\-Back0\.558±0\.0680\.558\_\{\\pm 0\.068\}0\.095±0\.0050\.095\_\{\\pm 0\.005\}0\.413±0\.0870\.413\_\{\\pm 0\.087\}0\.095±0\.0050\.095\_\{\\pm 0\.005\}DMAD0\.567±0\.0570\.567\_\{\\pm 0\.057\}0\.103±0\.0040\.103\_\{\\pm 0\.004\}0\.382±0\.1310\.382\_\{\\pm 0\.131\}0\.100±0\.0050\.100\_\{\\pm 0\.005\}Gibberish0\.525±0\.0580\.525\_\{\\pm 0\.058\}0\.150±0\.012\\mathbf\{0\.150\}\_\{\\pm 0\.012\}0\.398±0\.0260\.398\_\{\\pm 0\.026\}0\.156±0\.025\\mathbf\{0\.156\}\_\{\\pm 0\.025\}Baseline personasMPAQ0\.686±0\.0550\.686\_\{\\pm 0\.055\}0\.132±0\.0060\.132\_\{\\pm 0\.006\}0\.535±0\.0900\.535\_\{\\pm 0\.090\}0\.131±0\.0060\.131\_\{\\pm 0\.006\}Individualized personasRandom0\.588±0\.0510\.588\_\{\\pm 0\.051\}0\.107±0\.0030\.107\_\{\\pm 0\.003\}0\.467±0\.0610\.467\_\{\\pm 0\.061\}0\.109±0\.0040\.109\_\{\\pm 0\.004\}Typical0\.572±0\.0510\.572\_\{\\pm 0\.051\}0\.105±0\.0030\.105\_\{\\pm 0\.003\}0\.454±0\.0550\.454\_\{\\pm 0\.055\}0\.107±0\.0030\.107\_\{\\pm 0\.003\}Selected personas from base poolCoverage, cosine0\.574±0\.0570\.574\_\{\\pm 0\.057\}0\.110±0\.0030\.110\_\{\\pm 0\.003\}0\.445±0\.0970\.445\_\{\\pm 0\.097\}0\.111±0\.005\\mathbf\{0\.111\}\_\{\\pm 0\.005\}Coverage, Mahalanobis0\.548±0\.0480\.548\_\{\\pm 0\.048\}0\.101±0\.0030\.101\_\{\\pm 0\.003\}0\.385±0\.0290\.385\_\{\\pm 0\.029\}0\.104±0\.0050\.104\_\{\\pm 0\.005\}Dispersion, cosine0\.588±0\.081\\mathbf\{0\.588\}\_\{\\pm 0\.081\}0\.112±0\.001\\mathbf\{0\.112\}\_\{\\pm 0\.001\}0\.449±0\.086\\mathbf\{0\.449\}\_\{\\pm 0\.086\}0\.111±0\.001\\mathbf\{0\.111\}\_\{\\pm 0\.001\}Dispersion, Mahalanobis0\.567±0\.0750\.567\_\{\\pm 0\.075\}0\.108±0\.0020\.108\_\{\\pm 0\.002\}0\.406±0\.0690\.406\_\{\\pm 0\.069\}0\.110±0\.0030\.110\_\{\\pm 0\.003\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.566±0\.0780\.566\_\{\\pm 0\.078\}0\.113±0\.0020\.113\_\{\\pm 0\.002\}0\.411±0\.0780\.411\_\{\\pm 0\.078\}0\.113±0\.0070\.113\_\{\\pm 0\.007\}Evolution \(1,869\)0\.637±0\.064\\mathbf\{0\.637\}\_\{\\pm 0\.064\}0\.118±0\.002\\mathbf\{0\.118\}\_\{\\pm 0\.002\}0\.514±0\.1250\.514\_\{\\pm 0\.125\}0\.118±0\.0040\.118\_\{\\pm 0\.004\}AUT\-Evolution \(3,397\)0\.637±0\.068\\mathbf\{0\.637\}\_\{\\pm 0\.068\}0\.118±0\.002\\mathbf\{0\.118\}\_\{\\pm 0\.002\}0\.522±0\.066\\mathbf\{0\.522\}\_\{\\pm 0\.066\}0\.120±0\.005\\mathbf\{0\.120\}\_\{\\pm 0\.005\}Selected personas\+\+Creativity\-enhancedpromptCoverage, cosine0\.843±0\.061\\mathbf\{0\.843\}\_\{\\pm 0\.061\}0\.110±0\.0030\.110\_\{\\pm 0\.003\}0\.739±0\.1260\.739\_\{\\pm 0\.126\}0\.108±0\.0020\.108\_\{\\pm 0\.002\}Coverage, Mahalanobis0\.815±0\.0740\.815\_\{\\pm 0\.074\}0\.109±0\.0010\.109\_\{\\pm 0\.001\}0\.707±0\.1300\.707\_\{\\pm 0\.130\}0\.108±0\.0020\.108\_\{\\pm 0\.002\}Dispersion, cosine0\.833±0\.0430\.833\_\{\\pm 0\.043\}0\.113±0\.003\\mathbf\{0\.113\}\_\{\\pm 0\.003\}0\.763±0\.0710\.763\_\{\\pm 0\.071\}0\.110±0\.003\\mathbf\{0\.110\}\_\{\\pm 0\.003\}Dispersion, Mahalanobis0\.839±0\.0320\.839\_\{\\pm 0\.032\}0\.111±0\.0030\.111\_\{\\pm 0\.003\}0\.781±0\.079\\mathbf\{0\.781\}\_\{\\pm 0\.079\}0\.108±0\.0010\.108\_\{\\pm 0\.001\}Generated personas\+\+Creativity\-enhancedpromptMCMC0\.847±0\.0210\.847\_\{\\pm 0\.021\}0\.112±0\.0020\.112\_\{\\pm 0\.002\}0\.792±0\.0670\.792\_\{\\pm 0\.067\}0\.109±0\.0020\.109\_\{\\pm 0\.002\}Evolution0\.845±0\.0430\.845\_\{\\pm 0\.043\}0\.113±0\.0030\.113\_\{\\pm 0\.003\}0\.781±0\.0790\.781\_\{\\pm 0\.079\}0\.111±0\.0040\.111\_\{\\pm 0\.004\}AUT\-Evolution0\.872±0\.039\\mathbf\{0\.872\}\_\{\\pm 0\.039\}0\.115±0\.003\\mathbf\{0\.115\}\_\{\\pm 0\.003\}0\.822±0\.089\\mathbf\{0\.822\}\_\{\\pm 0\.089\}0\.115±0\.003\\mathbf\{0\.115\}\_\{\\pm 0\.003\}HumanStevenson——0\.741±0\.0650\.741\_\{\\pm 0\.065\}0\.106±0\.0010\.106\_\{\\pm 0\.001\}Table 18:LLM\-judged creativity on the AUT across all seven objects\. A Qwen3\.6\-27B judge, independent of the Gemma\-4\-31B generator, rates every use for Originality, Surprise and Utility on a11–55scale following[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib53), and for holistic creativity following[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18); we report the rank\-normalized form of that holistic judgment, which correlates with its raw form atr=0\.996r=0\.996\. Cells give the mean over judged uses \(at most175175per variant\)\. The judged set is every use not rejected outright by the validity gate of[SectionD\.2](https://arxiv.org/html/2609.30492#A4.SS2), admitted*and*uncertain uses, so it is slightly larger than the admitted set of Table[11](https://arxiv.org/html/2609.30492#A5.T11)\. Subscripts denote the 95% confidence interval across 5 independent random seeds\. The final two groups re\-run our selection and generation methods on theCreativity\-enhancedprompt, showing that persona diversification composes with a stronger prompting strategy\. Originality, Surprise and Creativity are target metrics, bolded per column within the standard, selected and generated groups separately \(and within each creativity\-enhanced group\), with tied values bolded jointly\.*Utility is a control*: conventional uses are the most useful by construction, so it necessarily peaks atCommon\-Useand is never bolded\. Baseline, individualized and human rows are not bolded\. Generated variants report their expanded pool size in parentheses\.[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib53)[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18)VariantOriginalitySurpriseUtilityCreativity \(rank\)Standard personaCommon\-Use1\.68±0\.031\.68\_\{\\pm 0\.03\}1\.34±0\.011\.34\_\{\\pm 0\.01\}4\.86±0\.044\.86\_\{\\pm 0\.04\}2\.02±0\.082\.02\_\{\\pm 0\.08\}Alternative\-Use2\.63±0\.052\.63\_\{\\pm 0\.05\}2\.23±0\.062\.23\_\{\\pm 0\.06\}4\.30±0\.054\.30\_\{\\pm 0\.05\}2\.72±0\.092\.72\_\{\\pm 0\.09\}Expert2\.28±0\.052\.28\_\{\\pm 0\.05\}2\.01±0\.052\.01\_\{\\pm 0\.05\}4\.39±0\.054\.39\_\{\\pm 0\.05\}2\.71±0\.052\.71\_\{\\pm 0\.05\}Creativity\-enhanced3\.17±0\.06\\mathbf\{3\.17\}\_\{\\pm 0\.06\}2\.66±0\.06\\mathbf\{2\.66\}\_\{\\pm 0\.06\}3\.38±0\.073\.38\_\{\\pm 0\.07\}3\.19±0\.09\\mathbf\{3\.19\}\_\{\\pm 0\.09\}ZS\-CoT2\.60±0\.092\.60\_\{\\pm 0\.09\}2\.11±0\.072\.11\_\{\\pm 0\.07\}4\.27±0\.074\.27\_\{\\pm 0\.07\}2\.85±0\.062\.85\_\{\\pm 0\.06\}Step\-Back2\.74±0\.022\.74\_\{\\pm 0\.02\}2\.30±0\.072\.30\_\{\\pm 0\.07\}4\.30±0\.044\.30\_\{\\pm 0\.04\}2\.95±0\.062\.95\_\{\\pm 0\.06\}DMAD2\.65±0\.112\.65\_\{\\pm 0\.11\}2\.17±0\.092\.17\_\{\\pm 0\.09\}4\.17±0\.114\.17\_\{\\pm 0\.11\}2\.89±0\.072\.89\_\{\\pm 0\.07\}Gibberish2\.85±0\.062\.85\_\{\\pm 0\.06\}2\.33±0\.072\.33\_\{\\pm 0\.07\}4\.02±0\.104\.02\_\{\\pm 0\.10\}2\.57±0\.092\.57\_\{\\pm 0\.09\}Baseline personasMPAQ3\.29±0\.043\.29\_\{\\pm 0\.04\}2\.65±0\.062\.65\_\{\\pm 0\.06\}3\.59±0\.043\.59\_\{\\pm 0\.04\}3\.05±0\.073\.05\_\{\\pm 0\.07\}Individualized personasRandom2\.78±0\.032\.78\_\{\\pm 0\.03\}2\.36±0\.032\.36\_\{\\pm 0\.03\}4\.28±0\.024\.28\_\{\\pm 0\.02\}2\.96±0\.042\.96\_\{\\pm 0\.04\}Typical2\.75±0\.042\.75\_\{\\pm 0\.04\}2\.36±0\.042\.36\_\{\\pm 0\.04\}4\.24±0\.054\.24\_\{\\pm 0\.05\}2\.99±0\.072\.99\_\{\\pm 0\.07\}Selected personas from base poolCoverage, cosine2\.74±0\.042\.74\_\{\\pm 0\.04\}2\.40±0\.02\\mathbf\{2\.40\}\_\{\\pm 0\.02\}4\.30±0\.024\.30\_\{\\pm 0\.02\}2\.91±0\.062\.91\_\{\\pm 0\.06\}Coverage, Mahalanobis2\.66±0\.022\.66\_\{\\pm 0\.02\}2\.29±0\.052\.29\_\{\\pm 0\.05\}4\.27±0\.044\.27\_\{\\pm 0\.04\}2\.86±0\.082\.86\_\{\\pm 0\.08\}Dispersion, cosine2\.78±0\.06\\mathbf\{2\.78\}\_\{\\pm 0\.06\}2\.40±0\.04\\mathbf\{2\.40\}\_\{\\pm 0\.04\}4\.27±0\.094\.27\_\{\\pm 0\.09\}2\.91±0\.062\.91\_\{\\pm 0\.06\}Dispersion, Mahalanobis2\.74±0\.082\.74\_\{\\pm 0\.08\}2\.36±0\.062\.36\_\{\\pm 0\.06\}4\.23±0\.094\.23\_\{\\pm 0\.09\}2\.92±0\.04\\mathbf\{2\.92\}\_\{\\pm 0\.04\}Generated personas\(expanded pool size\)MCMC \(8,553\)2\.74±0\.052\.74\_\{\\pm 0\.05\}2\.40±0\.072\.40\_\{\\pm 0\.07\}4\.28±0\.064\.28\_\{\\pm 0\.06\}2\.91±0\.042\.91\_\{\\pm 0\.04\}Evolution \(1,869\)2\.97±0\.072\.97\_\{\\pm 0\.07\}2\.54±0\.062\.54\_\{\\pm 0\.06\}4\.08±0\.064\.08\_\{\\pm 0\.06\}3\.10±0\.03\\mathbf\{3\.10\}\_\{\\pm 0\.03\}AUT\-Evolution \(3,397\)2\.99±0\.09\\mathbf\{2\.99\}\_\{\\pm 0\.09\}2\.63±0\.09\\mathbf\{2\.63\}\_\{\\pm 0\.09\}4\.04±0\.104\.04\_\{\\pm 0\.10\}2\.99±0\.062\.99\_\{\\pm 0\.06\}Selected personas\+\+Creativity\-enhancedpromptCoverage, cosine3\.22±0\.07\\mathbf\{3\.22\}\_\{\\pm 0\.07\}2\.75±0\.04\\mathbf\{2\.75\}\_\{\\pm 0\.04\}3\.43±0\.103\.43\_\{\\pm 0\.10\}3\.25±0\.043\.25\_\{\\pm 0\.04\}Coverage, Mahalanobis3\.15±0\.053\.15\_\{\\pm 0\.05\}2\.71±0\.052\.71\_\{\\pm 0\.05\}3\.58±0\.103\.58\_\{\\pm 0\.10\}3\.24±0\.073\.24\_\{\\pm 0\.07\}Dispersion, cosine3\.20±0\.093\.20\_\{\\pm 0\.09\}2\.72±0\.122\.72\_\{\\pm 0\.12\}3\.54±0\.113\.54\_\{\\pm 0\.11\}3\.25±0\.073\.25\_\{\\pm 0\.07\}Dispersion, Mahalanobis3\.19±0\.123\.19\_\{\\pm 0\.12\}2\.74±0\.102\.74\_\{\\pm 0\.10\}3\.44±0\.093\.44\_\{\\pm 0\.09\}3\.27±0\.06\\mathbf\{3\.27\}\_\{\\pm 0\.06\}Generated personas\+\+Creativity\-enhancedpromptMCMC3\.24±0\.153\.24\_\{\\pm 0\.15\}2\.76±0\.102\.76\_\{\\pm 0\.10\}3\.39±0\.103\.39\_\{\\pm 0\.10\}3\.31±0\.053\.31\_\{\\pm 0\.05\}Evolution3\.31±0\.123\.31\_\{\\pm 0\.12\}2\.87±0\.112\.87\_\{\\pm 0\.11\}3\.38±0\.163\.38\_\{\\pm 0\.16\}3\.39±0\.07\\mathbf\{3\.39\}\_\{\\pm 0\.07\}AUT\-Evolution3\.38±0\.13\\mathbf\{3\.38\}\_\{\\pm 0\.13\}2\.94±0\.13\\mathbf\{2\.94\}\_\{\\pm 0\.13\}3\.33±0\.143\.33\_\{\\pm 0\.14\}3\.32±0\.063\.32\_\{\\pm 0\.06\}HumanStevenson2\.78±0\.002\.78\_\{\\pm 0\.00\}1\.77±0\.001\.77\_\{\\pm 0\.00\}3\.49±0\.033\.49\_\{\\pm 0\.03\}2\.07±0\.032\.07\_\{\\pm 0\.03\}Table 19:LLM\-judged creativity restricted to the three objects with human reference data \(book, fork, tin can\), so that all variants are directly comparable toStevenson\(at most7575judged uses per variant\)\. Scales, judge, bolding and control conventions are as in Table[18](https://arxiv.org/html/2609.30492#A5.T18)\.[Stevenson et al\. \(2022\)](https://arxiv.org/html/2609.30492#bib.bib53)[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18)VariantOriginalitySurpriseUtilityCreativity \(rank\)Standard personaCommon\-Use1\.84±0\.041\.84\_\{\\pm 0\.04\}1\.45±0\.031\.45\_\{\\pm 0\.03\}4\.84±0\.034\.84\_\{\\pm 0\.03\}2\.29±0\.122\.29\_\{\\pm 0\.12\}Alternative\-Use2\.37±0\.092\.37\_\{\\pm 0\.09\}2\.07±0\.052\.07\_\{\\pm 0\.05\}4\.39±0\.084\.39\_\{\\pm 0\.08\}2\.78±0\.162\.78\_\{\\pm 0\.16\}Expert2\.18±0\.042\.18\_\{\\pm 0\.04\}2\.00±0\.062\.00\_\{\\pm 0\.06\}4\.40±0\.084\.40\_\{\\pm 0\.08\}2\.83±0\.072\.83\_\{\\pm 0\.07\}Creativity\-enhanced3\.05±0\.18\\mathbf\{3\.05\}\_\{\\pm 0\.18\}2\.66±0\.06\\mathbf\{2\.66\}\_\{\\pm 0\.06\}3\.52±0\.113\.52\_\{\\pm 0\.11\}3\.19±0\.09\\mathbf\{3\.19\}\_\{\\pm 0\.09\}ZS\-CoT2\.37±0\.112\.37\_\{\\pm 0\.11\}1\.96±0\.121\.96\_\{\\pm 0\.12\}4\.49±0\.104\.49\_\{\\pm 0\.10\}2\.88±0\.122\.88\_\{\\pm 0\.12\}Step\-Back2\.52±0\.072\.52\_\{\\pm 0\.07\}2\.16±0\.072\.16\_\{\\pm 0\.07\}4\.31±0\.094\.31\_\{\\pm 0\.09\}2\.95±0\.072\.95\_\{\\pm 0\.07\}DMAD2\.39±0\.112\.39\_\{\\pm 0\.11\}2\.06±0\.082\.06\_\{\\pm 0\.08\}4\.38±0\.084\.38\_\{\\pm 0\.08\}2\.90±0\.122\.90\_\{\\pm 0\.12\}Gibberish2\.59±0\.072\.59\_\{\\pm 0\.07\}2\.13±0\.092\.13\_\{\\pm 0\.09\}4\.19±0\.154\.19\_\{\\pm 0\.15\}2\.37±0\.222\.37\_\{\\pm 0\.22\}Baseline personasMPAQ3\.05±0\.133\.05\_\{\\pm 0\.13\}2\.44±0\.082\.44\_\{\\pm 0\.08\}3\.78±0\.073\.78\_\{\\pm 0\.07\}3\.00±0\.103\.00\_\{\\pm 0\.10\}Individualized personasRandom2\.52±0\.052\.52\_\{\\pm 0\.05\}2\.18±0\.042\.18\_\{\\pm 0\.04\}4\.35±0\.034\.35\_\{\\pm 0\.03\}3\.03±0\.023\.03\_\{\\pm 0\.02\}Typical2\.56±0\.112\.56\_\{\\pm 0\.11\}2\.28±0\.052\.28\_\{\\pm 0\.05\}4\.24±0\.124\.24\_\{\\pm 0\.12\}3\.07±0\.073\.07\_\{\\pm 0\.07\}Selected personas from base poolCoverage, cosine2\.52±0\.072\.52\_\{\\pm 0\.07\}2\.27±0\.06\\mathbf\{2\.27\}\_\{\\pm 0\.06\}4\.35±0\.044\.35\_\{\\pm 0\.04\}2\.98±0\.07\\mathbf\{2\.98\}\_\{\\pm 0\.07\}Coverage, Mahalanobis2\.44±0\.052\.44\_\{\\pm 0\.05\}2\.17±0\.122\.17\_\{\\pm 0\.12\}4\.34±0\.084\.34\_\{\\pm 0\.08\}2\.93±0\.082\.93\_\{\\pm 0\.08\}Dispersion, cosine2\.56±0\.08\\mathbf\{2\.56\}\_\{\\pm 0\.08\}2\.21±0\.052\.21\_\{\\pm 0\.05\}4\.35±0\.084\.35\_\{\\pm 0\.08\}2\.95±0\.062\.95\_\{\\pm 0\.06\}Dispersion, Mahalanobis2\.53±0\.072\.53\_\{\\pm 0\.07\}2\.16±0\.072\.16\_\{\\pm 0\.07\}4\.33±0\.054\.33\_\{\\pm 0\.05\}2\.94±0\.042\.94\_\{\\pm 0\.04\}Generated personas\(expanded pool size\)MCMC \(8,553\)2\.53±0\.082\.53\_\{\\pm 0\.08\}2\.20±0\.032\.20\_\{\\pm 0\.03\}4\.36±0\.064\.36\_\{\\pm 0\.06\}2\.93±0\.072\.93\_\{\\pm 0\.07\}Evolution \(1,869\)2\.76±0\.062\.76\_\{\\pm 0\.06\}2\.38±0\.082\.38\_\{\\pm 0\.08\}4\.22±0\.104\.22\_\{\\pm 0\.10\}3\.11±0\.02\\mathbf\{3\.11\}\_\{\\pm 0\.02\}AUT\-Evolution \(3,397\)2\.88±0\.09\\mathbf\{2\.88\}\_\{\\pm 0\.09\}2\.49±0\.09\\mathbf\{2\.49\}\_\{\\pm 0\.09\}4\.10±0\.114\.10\_\{\\pm 0\.11\}3\.02±0\.053\.02\_\{\\pm 0\.05\}Selected personas\+\+Creativity\-enhancedpromptCoverage, cosine3\.11±0\.08\\mathbf\{3\.11\}\_\{\\pm 0\.08\}2\.69±0\.082\.69\_\{\\pm 0\.08\}3\.59±0\.133\.59\_\{\\pm 0\.13\}3\.27±0\.043\.27\_\{\\pm 0\.04\}Coverage, Mahalanobis3\.10±0\.113\.10\_\{\\pm 0\.11\}2\.73±0\.13\\mathbf\{2\.73\}\_\{\\pm 0\.13\}3\.68±0\.123\.68\_\{\\pm 0\.12\}3\.26±0\.133\.26\_\{\\pm 0\.13\}Dispersion, cosine3\.10±0\.163\.10\_\{\\pm 0\.16\}2\.71±0\.122\.71\_\{\\pm 0\.12\}3\.63±0\.193\.63\_\{\\pm 0\.19\}3\.30±0\.063\.30\_\{\\pm 0\.06\}Dispersion, Mahalanobis3\.09±0\.153\.09\_\{\\pm 0\.15\}2\.72±0\.162\.72\_\{\\pm 0\.16\}3\.56±0\.123\.56\_\{\\pm 0\.12\}3\.31±0\.06\\mathbf\{3\.31\}\_\{\\pm 0\.06\}Generated personas\+\+Creativity\-enhancedpromptMCMC3\.15±0\.193\.15\_\{\\pm 0\.19\}2\.75±0\.102\.75\_\{\\pm 0\.10\}3\.50±0\.183\.50\_\{\\pm 0\.18\}3\.33±0\.133\.33\_\{\\pm 0\.13\}Evolution3\.21±0\.243\.21\_\{\\pm 0\.24\}2\.81±0\.182\.81\_\{\\pm 0\.18\}3\.50±0\.283\.50\_\{\\pm 0\.28\}3\.41±0\.16\\mathbf\{3\.41\}\_\{\\pm 0\.16\}AUT\-Evolution3\.27±0\.22\\mathbf\{3\.27\}\_\{\\pm 0\.22\}2\.87±0\.17\\mathbf\{2\.87\}\_\{\\pm 0\.17\}3\.44±0\.203\.44\_\{\\pm 0\.20\}3\.31±0\.123\.31\_\{\\pm 0\.12\}HumanStevenson2\.78±0\.002\.78\_\{\\pm 0\.00\}1\.77±0\.001\.77\_\{\\pm 0\.00\}3\.49±0\.033\.49\_\{\\pm 0\.03\}2\.07±0\.032\.07\_\{\\pm 0\.03\}
### E\.4Infinity\-Chat Benchmark

Table 20:Infinity\-Chat 100 response validity and quality\([Jiang et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib26)\)\. Each condition contributes250250responses,5050per open\-ended item for prompt conditions and ten per persona per item for persona conditions overk=5k=5personas and five items\.*Valid*is the proportion of responses admitted by the validity gate;*Quality*is the mean judge rating on a11–55scale\. Subscripts denote the 95% confidence interval across 5 independent random seeds\. The final two groups re\-run our selection and generation methods on theDMADprompt\. Both are*controls*rather than targets, validity near11is expected, and quality is the effectiveness dimension against which diversity gains are traded, so no value is bolded\. Generated variants report their expanded pool size in parentheses\.VariantValidQualityStandard personaOpen\-Ended Query0\.950±0\.0150\.950\_\{\\pm 0\.015\}4\.84±0\.024\.84\_\{\\pm 0\.02\}ZS\-CoT1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.99±0\.014\.99\_\{\\pm 0\.01\}Step\-Back1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.98±0\.024\.98\_\{\\pm 0\.02\}DMAD1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.96±0\.044\.96\_\{\\pm 0\.04\}Gibberish0\.123±0\.0130\.123\_\{\\pm 0\.013\}1\.01±0\.011\.01\_\{\\pm 0\.01\}Baseline personasMPAQ0\.994±0\.0080\.994\_\{\\pm 0\.008\}3\.69±0\.053\.69\_\{\\pm 0\.05\}Individualized personasRandom1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.43±0\.034\.43\_\{\\pm 0\.03\}Typical1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.44±0\.034\.44\_\{\\pm 0\.03\}Selected personas from base poolCoverage, cosine/L20\.999±0\.0020\.999\_\{\\pm 0\.002\}4\.32±0\.034\.32\_\{\\pm 0\.03\}Coverage, Mahalanobis0\.996±0\.0000\.996\_\{\\pm 0\.000\}4\.38±0\.044\.38\_\{\\pm 0\.04\}Dispersion, cosine/L20\.970±0\.0110\.970\_\{\\pm 0\.011\}4\.13±0\.084\.13\_\{\\pm 0\.08\}Dispersion, Mahalanobis0\.986±0\.0070\.986\_\{\\pm 0\.007\}4\.13±0\.064\.13\_\{\\pm 0\.06\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.991±0\.0070\.991\_\{\\pm 0\.007\}4\.23±0\.074\.23\_\{\\pm 0\.07\}Evolution \(1,869\)0\.992±0\.0050\.992\_\{\\pm 0\.005\}3\.48±0\.053\.48\_\{\\pm 0\.05\}IC\-Evolution \(2,819\)0\.974±0\.0100\.974\_\{\\pm 0\.010\}3\.96±0\.063\.96\_\{\\pm 0\.06\}Selected personas\+\+DMADpromptCoverage, cosine/L21\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.95±0\.014\.95\_\{\\pm 0\.01\}Coverage, Mahalanobis1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.90±0\.034\.90\_\{\\pm 0\.03\}Dispersion, cosine/L21\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.95±0\.014\.95\_\{\\pm 0\.01\}Dispersion, Mahalanobis1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.94±0\.024\.94\_\{\\pm 0\.02\}Generated personas\+\+DMADpromptMCMC1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.94±0\.014\.94\_\{\\pm 0\.01\}Evolution1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.86±0\.054\.86\_\{\\pm 0\.05\}IC\-Evolution1\.000±0\.0001\.000\_\{\\pm 0\.000\}4\.92±0\.014\.92\_\{\\pm 0\.01\}Table 21:Infinity\-Chat 100 response homogeneity\([Jiang et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib26)\)\.*Similarity*is the mean pairwise embedding similarity among a condition’s responses to the same item, averaged over items;*Separation*is the gain in mean between\-persona distance over within\-persona distance, and is undefined for conditions without personas \(—\);*Trigram*is mean trigram Jaccard overlap\. Arrows give the favourable direction;boldmarks the best per column within the standard, selected and generated groups, and within eachDMADgroup\.†Gibberishis excluded from bolding because90%90\\%of its responses fail the validity gate, leaving statistics computed on a tenth of the sample\. Subscripts denote the 95% confidence interval across 5 independent random seeds\. The final two groups re\-run our selection and generation methods on theDMADprompt\. Generated variants report their expanded pool size in parentheses\.VariantSimilarity↓\\downarrowSeparation↑\\uparrowTrigram↓\\downarrowStandard personaOpen\-Ended Query0\.946±0\.0020\.946\_\{\\pm 0\.002\}—0\.247±0\.0040\.247\_\{\\pm 0\.004\}ZS\-CoT0\.949±0\.0030\.949\_\{\\pm 0\.003\}—0\.096±0\.0120\.096\_\{\\pm 0\.012\}Step\-Back0\.962±0\.0030\.962\_\{\\pm 0\.003\}—0\.276±0\.0070\.276\_\{\\pm 0\.007\}DMAD0\.929±0\.002\\mathbf\{0\.929\}\_\{\\pm 0\.002\}—0\.055±0\.004\\mathbf\{0\.055\}\_\{\\pm 0\.004\}Gibberish†0\.894±0\.0120\.894\_\{\\pm 0\.012\}0\.045±0\.0080\.045\_\{\\pm 0\.008\}0\.055±0\.0220\.055\_\{\\pm 0\.022\}Baseline personasMPAQ0\.831±0\.0020\.831\_\{\\pm 0\.002\}0\.123±0\.0010\.123\_\{\\pm 0\.001\}0\.029±0\.0010\.029\_\{\\pm 0\.001\}Individualized personasRandom0\.864±0\.0050\.864\_\{\\pm 0\.005\}0\.068±0\.0030\.068\_\{\\pm 0\.003\}0\.039±0\.0040\.039\_\{\\pm 0\.004\}Typical0\.874±0\.0040\.874\_\{\\pm 0\.004\}0\.083±0\.0020\.083\_\{\\pm 0\.002\}0\.055±0\.0030\.055\_\{\\pm 0\.003\}Selected personas from base poolCoverage, cosine/L20\.885±0\.0030\.885\_\{\\pm 0\.003\}0\.055±0\.0030\.055\_\{\\pm 0\.003\}0\.040±0\.0020\.040\_\{\\pm 0\.002\}Coverage, Mahalanobis0\.866±0\.0040\.866\_\{\\pm 0\.004\}0\.080±0\.0030\.080\_\{\\pm 0\.003\}0\.038±0\.0040\.038\_\{\\pm 0\.004\}Dispersion, cosine/L20\.822±0\.005\\mathbf\{0\.822\}\_\{\\pm 0\.005\}0\.094±0\.009\\mathbf\{0\.094\}\_\{\\pm 0\.009\}0\.025±0\.003\\mathbf\{0\.025\}\_\{\\pm 0\.003\}Dispersion, Mahalanobis0\.840±0\.0040\.840\_\{\\pm 0\.004\}0\.085±0\.0060\.085\_\{\\pm 0\.006\}0\.029±0\.0060\.029\_\{\\pm 0\.006\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.839±0\.0040\.839\_\{\\pm 0\.004\}0\.088±0\.0060\.088\_\{\\pm 0\.006\}0\.027±0\.0050\.027\_\{\\pm 0\.005\}Evolution \(1,869\)0\.798±0\.006\\mathbf\{0\.798\}\_\{\\pm 0\.006\}0\.126±0\.006\\mathbf\{0\.126\}\_\{\\pm 0\.006\}0\.015±0\.002\\mathbf\{0\.015\}\_\{\\pm 0\.002\}IC\-Evolution \(2,819\)0\.828±0\.0060\.828\_\{\\pm 0\.006\}0\.096±0\.0040\.096\_\{\\pm 0\.004\}0\.023±0\.0030\.023\_\{\\pm 0\.003\}Selected personas\+\+DMADpromptCoverage, cosine/L20\.927±0\.0010\.927\_\{\\pm 0\.001\}0\.004±0\.0010\.004\_\{\\pm 0\.001\}0\.050±0\.0090\.050\_\{\\pm 0\.009\}Coverage, Mahalanobis0\.920±0\.005\\mathbf\{0\.920\}\_\{\\pm 0\.005\}0\.012±0\.002\\mathbf\{0\.012\}\_\{\\pm 0\.002\}0\.050±0\.0130\.050\_\{\\pm 0\.013\}Dispersion, cosine/L20\.927±0\.0010\.927\_\{\\pm 0\.001\}0\.006±0\.0010\.006\_\{\\pm 0\.001\}0\.041±0\.004\\mathbf\{0\.041\}\_\{\\pm 0\.004\}Dispersion, Mahalanobis0\.925±0\.0030\.925\_\{\\pm 0\.003\}0\.006±0\.0020\.006\_\{\\pm 0\.002\}0\.049±0\.0100\.049\_\{\\pm 0\.010\}Generated personas\+\+DMADpromptMCMC0\.923±0\.0020\.923\_\{\\pm 0\.002\}0\.007±0\.0010\.007\_\{\\pm 0\.001\}0\.048±0\.0110\.048\_\{\\pm 0\.011\}Evolution0\.900±0\.001\\mathbf\{0\.900\}\_\{\\pm 0\.001\}0\.021±0\.003\\mathbf\{0\.021\}\_\{\\pm 0\.003\}0\.030±0\.004\\mathbf\{0\.030\}\_\{\\pm 0\.004\}IC\-Evolution0\.913±0\.0020\.913\_\{\\pm 0\.002\}0\.013±0\.0010\.013\_\{\\pm 0\.001\}0\.031±0\.0060\.031\_\{\\pm 0\.006\}Table 22:Infinity\-Chat 100 response fluency and flexibility\([Jiang et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib26)\)\.*Tokens*is the number of word tokens a condition produced, given because the three lexical measures are corpus statistics and are biased downward by corpus size\.*Fluency*: vocabulary entropyH⁡\(𝒲\)H\(\\mathcal\{W\}\)in nats and top\-1010concentrationC10​\(𝒲\)C\_\{10\}\(\\mathcal\{W\}\), the share of occurrences taken by the ten most frequent words\.*Flexibility*: type–token ratioT​T​R​\(𝒲\)TTR\(\\mathcal\{W\}\)and the Vendi score equation[53](https://arxiv.org/html/2609.30492#A4.E53)of the response embeddings, the effective number of distinct concepts; the latter is computed over response sets of equal size and so is unaffected by the token counts\. Arrows give the favourable direction;boldmarks the best per column within the standard, selected and generated groups, and within eachDMADgroup\. Token counts are not bolded\.†Gibberishis excluded from bolding because90%90\\%of its responses fail the validity gate\. Subscripts denote the 95% confidence interval across 5 independent random seeds\. The final two groups re\-run our selection and generation methods on theDMADprompt\. Generated variants report their expanded pool size in parentheses\.FluencyFlexibilityVariantTokensEntropy↑\\uparrowTop\-1010↓\\downarrowTTR↑\\uparrowVendi↑\\uparrowStandard personaOpen\-Ended Query31,17031\{,\}1705\.98±0\.02\\mathbf\{5\.98\}\_\{\\pm 0\.02\}0\.244±0\.0020\.244\_\{\\pm 0\.002\}0\.068±0\.0010\.068\_\{\\pm 0\.001\}1\.39±0\.021\.39\_\{\\pm 0\.02\}ZS\-CoT12,53612\{,\}5365\.80±0\.035\.80\_\{\\pm 0\.03\}0\.232±0\.004\\mathbf\{0\.232\}\_\{\\pm 0\.004\}0\.099±0\.0030\.099\_\{\\pm 0\.003\}1\.40±0\.031\.40\_\{\\pm 0\.03\}Step\-Back11,90011\{,\}9005\.69±0\.025\.69\_\{\\pm 0\.02\}0\.241±0\.0040\.241\_\{\\pm 0\.004\}0\.093±0\.0010\.093\_\{\\pm 0\.001\}1\.28±0\.021\.28\_\{\\pm 0\.02\}DMAD10,44310\{,\}4435\.85±0\.025\.85\_\{\\pm 0\.02\}0\.251±0\.0020\.251\_\{\\pm 0\.002\}0\.129±0\.003\\mathbf\{0\.129\}\_\{\\pm 0\.003\}1\.56±0\.01\\mathbf\{1\.56\}\_\{\\pm 0\.01\}Gibberish†1,7171\{,\}7175\.48±0\.145\.48\_\{\\pm 0\.14\}0\.269±0\.0180\.269\_\{\\pm 0\.018\}0\.342±0\.0470\.342\_\{\\pm 0\.047\}1\.63±0\.191\.63\_\{\\pm 0\.19\}Baseline personasMPAQ21,88621\{,\}8866\.38±0\.016\.38\_\{\\pm 0\.01\}0\.251±0\.0030\.251\_\{\\pm 0\.003\}0\.142±0\.0010\.142\_\{\\pm 0\.001\}2\.40±0\.022\.40\_\{\\pm 0\.02\}Individualized personasRandom23,44923\{,\}4496\.27±0\.026\.27\_\{\\pm 0\.02\}0\.259±0\.0040\.259\_\{\\pm 0\.004\}0\.123±0\.0030\.123\_\{\\pm 0\.003\}2\.17±0\.052\.17\_\{\\pm 0\.05\}Typical20,30520\{,\}3056\.26±0\.026\.26\_\{\\pm 0\.02\}0\.246±0\.0030\.246\_\{\\pm 0\.003\}0\.122±0\.0030\.122\_\{\\pm 0\.003\}1\.98±0\.041\.98\_\{\\pm 0\.04\}Selected personas from base poolCoverage, cosine/L222,08322\{,\}0836\.22±0\.026\.22\_\{\\pm 0\.02\}0\.256±0\.0030\.256\_\{\\pm 0\.003\}0\.118±0\.0020\.118\_\{\\pm 0\.002\}1\.97±0\.031\.97\_\{\\pm 0\.03\}Coverage, Mahalanobis22,04922\{,\}0496\.31±0\.01\\mathbf\{6\.31\}\_\{\\pm 0\.01\}0\.246±0\.003\\mathbf\{0\.246\}\_\{\\pm 0\.003\}0\.122±0\.0020\.122\_\{\\pm 0\.002\}2\.07±0\.052\.07\_\{\\pm 0\.05\}Dispersion, cosine/L224,39724\{,\}3976\.26±0\.026\.26\_\{\\pm 0\.02\}0\.264±0\.0040\.264\_\{\\pm 0\.004\}0\.125±0\.0030\.125\_\{\\pm 0\.003\}2\.58±0\.05\\mathbf\{2\.58\}\_\{\\pm 0\.05\}Dispersion, Mahalanobis22,17522\{,\}1756\.28±0\.026\.28\_\{\\pm 0\.02\}0\.250±0\.0020\.250\_\{\\pm 0\.002\}0\.128±0\.002\\mathbf\{0\.128\}\_\{\\pm 0\.002\}2\.40±0\.052\.40\_\{\\pm 0\.05\}Generated personas\(expanded pool size\)MCMC \(8,553\)21,37821\{,\}3786\.29±0\.016\.29\_\{\\pm 0\.01\}0\.249±0\.002\\mathbf\{0\.249\}\_\{\\pm 0\.002\}0\.132±0\.0030\.132\_\{\\pm 0\.003\}2\.39±0\.062\.39\_\{\\pm 0\.06\}Evolution \(1,869\)24,22924\{,\}2296\.41±0\.02\\mathbf\{6\.41\}\_\{\\pm 0\.02\}0\.270±0\.0030\.270\_\{\\pm 0\.003\}0\.147±0\.004\\mathbf\{0\.147\}\_\{\\pm 0\.004\}2\.81±0\.08\\mathbf\{2\.81\}\_\{\\pm 0\.08\}IC\-Evolution \(2,819\)19,92119\{,\}9216\.31±0\.016\.31\_\{\\pm 0\.01\}0\.252±0\.0030\.252\_\{\\pm 0\.003\}0\.142±0\.0040\.142\_\{\\pm 0\.004\}2\.50±0\.062\.50\_\{\\pm 0\.06\}Selected personas\+\+DMADpromptCoverage, cosine/L210,82210\{,\}8225\.92±0\.025\.92\_\{\\pm 0\.02\}0\.253±0\.0040\.253\_\{\\pm 0\.004\}0\.138±0\.0030\.138\_\{\\pm 0\.003\}1\.60±0\.011\.60\_\{\\pm 0\.01\}Coverage, Mahalanobis10,60110\{,\}6015\.94±0\.035\.94\_\{\\pm 0\.03\}0\.250±0\.003\\mathbf\{0\.250\}\_\{\\pm 0\.003\}0\.144±0\.0050\.144\_\{\\pm 0\.005\}1\.65±0\.05\\mathbf\{1\.65\}\_\{\\pm 0\.05\}Dispersion, cosine/L211,11211\{,\}1125\.93±0\.025\.93\_\{\\pm 0\.02\}0\.255±0\.0040\.255\_\{\\pm 0\.004\}0\.139±0\.0030\.139\_\{\\pm 0\.003\}1\.60±0\.011\.60\_\{\\pm 0\.01\}Dispersion, Mahalanobis10,69210\{,\}6925\.95±0\.02\\mathbf\{5\.95\}\_\{\\pm 0\.02\}0\.250±0\.001\\mathbf\{0\.250\}\_\{\\pm 0\.001\}0\.144±0\.003\\mathbf\{0\.144\}\_\{\\pm 0\.003\}1\.62±0\.031\.62\_\{\\pm 0\.03\}Generated personas\+\+DMADpromptMCMC10,64010\{,\}6405\.96±0\.025\.96\_\{\\pm 0\.02\}0\.250±0\.002\\mathbf\{0\.250\}\_\{\\pm 0\.002\}0\.146±0\.0020\.146\_\{\\pm 0\.002\}1\.64±0\.021\.64\_\{\\pm 0\.02\}Evolution10,35510\{,\}3556\.02±0\.02\\mathbf\{6\.02\}\_\{\\pm 0\.02\}0\.253±0\.0030\.253\_\{\\pm 0\.003\}0\.162±0\.003\\mathbf\{0\.162\}\_\{\\pm 0\.003\}1\.81±0\.01\\mathbf\{1\.81\}\_\{\\pm 0\.01\}IC\-Evolution10,74410\{,\}7445\.97±0\.025\.97\_\{\\pm 0\.02\}0\.255±0\.0040\.255\_\{\\pm 0\.004\}0\.150±0\.0020\.150\_\{\\pm 0\.002\}1\.72±0\.021\.72\_\{\\pm 0\.02\}
### E\.5DAT Benchmark

Table 23:DAT response validity: the proportion of thenraw=350n\_\{\\mathrm\{raw\}\}=350words generated per condition that are admitted by the validity gate, which discards words repeated within a completion or absent from the reference vocabulary\. The final two groups re\-run our selection and generation methods on theDMADprompt, showing that persona diversification composes with a stronger prompting strategy rather than competing with it\. Validity is a*control*rather than a target, values near11are expected and the informative signal is degradation rather than rank, so no value is bolded\. No condition trades word validity for divergence\. Subscripts denote the 95% confidence interval across 5 independent random seeds\. Generated variants report their expanded pool size in parentheses\.VariantValid wordsStandard personaNon\-Divergent Association0\.998±0\.0030\.998\_\{\\pm 0\.003\}Divergent Association0\.996±0\.0050\.996\_\{\\pm 0\.005\}Base\-Instruction1\.000±0\.0001\.000\_\{\\pm 0\.000\}Random\-Instruction1\.000±0\.0001\.000\_\{\\pm 0\.000\}Creative0\.977±0\.0100\.977\_\{\\pm 0\.010\}ZS\-CoT0\.999±0\.0020\.999\_\{\\pm 0\.002\}Step\-Back0\.999±0\.0020\.999\_\{\\pm 0\.002\}DMAD0\.994±0\.0060\.994\_\{\\pm 0\.006\}Gibberish0\.998±0\.0050\.998\_\{\\pm 0\.005\}Baseline personasMPAQ0\.989±0\.0160\.989\_\{\\pm 0\.016\}Individualized personasRandom†0\.993±0\.0110\.993\_\{\\pm 0\.011\}Typical0\.998±0\.0040\.998\_\{\\pm 0\.004\}Selected personas from base poolCoverage, cosine/L20\.998±0\.0030\.998\_\{\\pm 0\.003\}Coverage, Mahalanobis0\.998±0\.0030\.998\_\{\\pm 0\.003\}Dispersion, cosine/L20\.998±0\.0030\.998\_\{\\pm 0\.003\}Dispersion, Mahalanobis0\.995±0\.0060\.995\_\{\\pm 0\.006\}Generated personas\(expanded pool size\)MCMC \(8,553\)0\.995±0\.0050\.995\_\{\\pm 0\.005\}Evolution \(1,869\)0\.996±0\.0040\.996\_\{\\pm 0\.004\}DAT\-Evolution \(2,439\)0\.964±0\.0170\.964\_\{\\pm 0\.017\}Selected personas\+\+DMADpromptCoverage, cosine/L20\.992±0\.0100\.992\_\{\\pm 0\.010\}Coverage, Mahalanobis0\.995±0\.0030\.995\_\{\\pm 0\.003\}Dispersion, cosine/L20\.994±0\.0030\.994\_\{\\pm 0\.003\}Dispersion, Mahalanobis0\.995±0\.0050\.995\_\{\\pm 0\.005\}Generated personas\+\+DMADpromptMCMC0\.997±0\.0040\.997\_\{\\pm 0\.004\}Evolution0\.997±0\.0080\.997\_\{\\pm 0\.008\}DAT\-Evolution0\.990±0\.0040\.990\_\{\\pm 0\.004\}Table 24:DAT response creativity over valid words\.*Diversity*: mean Mahalanobis dispersion and Mahalanobis hull extent\.*Fluency*: vocabulary entropyH⁡\(𝒲\)H\(\\mathcal\{W\}\)in nats, and top\-1010concentrationC10​\(𝒲\)C\_\{10\}\(\\mathcal\{W\}\), the share of occurrences taken by the ten most frequent words\.*Flexibility*: type–token ratioT​T​R​\(𝒲\)TTR\(\\mathcal\{W\}\), and the Vendi score equation[53](https://arxiv.org/html/2609.30492#A4.E53)of the word embeddings, the effective number of distinct concepts\. The final two groups re\-run our selection and generation methods on theDMADprompt, showing that persona diversification composes with a stronger prompting strategy rather than competing with it\. Subscripts denote the 95% confidence interval across 5 independent random seeds\. Arrows give the favorable direction;boldmarks the best value per column within the standard, selected and generated groups separately, and within eachDMADgroup\. Generated variants report their expanded pool size in parentheses\.DiversityFluencyFlexibilityVariantDisp\.↑\\uparrowHull↑\\uparrowEntropy↑\\uparrowTop\-1010↓\\downarrowTTR↑\\uparrowVendi↑\\uparrowStandard personaNon\-Divergent Association7\.05±0\.687\.05\_\{\\pm 0\.68\}—2\.89±0\.052\.89\_\{\\pm 0\.05\}0\.68±0\.050\.68\_\{\\pm 0\.05\}0\.07±0\.000\.07\_\{\\pm 0\.00\}1\.42±0\.011\.42\_\{\\pm 0\.01\}Divergent Association12\.30±0\.7512\.30\_\{\\pm 0\.75\}1\.28±0\.061\.28\_\{\\pm 0\.06\}3\.63±0\.043\.63\_\{\\pm 0\.04\}0\.53±0\.020\.53\_\{\\pm 0\.02\}0\.19±0\.010\.19\_\{\\pm 0\.01\}1\.79±0\.011\.79\_\{\\pm 0\.01\}Base\-Instruction13\.30±0\.4413\.30\_\{\\pm 0\.44\}0\.59±0\.060\.59\_\{\\pm 0\.06\}2\.96±0\.022\.96\_\{\\pm 0\.02\}0\.76±0\.010\.76\_\{\\pm 0\.01\}0\.10±0\.010\.10\_\{\\pm 0\.01\}1\.54±0\.001\.54\_\{\\pm 0\.00\}Random\-Instruction14\.19±0\.7014\.19\_\{\\pm 0\.70\}0\.79±0\.110\.79\_\{\\pm 0\.11\}3\.06±0\.063\.06\_\{\\pm 0\.06\}0\.71±0\.030\.71\_\{\\pm 0\.03\}0\.11±0\.010\.11\_\{\\pm 0\.01\}1\.63±0\.011\.63\_\{\\pm 0\.01\}Creative15\.71±0\.4015\.71\_\{\\pm 0\.40\}1\.25±0\.141\.25\_\{\\pm 0\.14\}3\.72±0\.053\.72\_\{\\pm 0\.05\}0\.52±0\.040\.52\_\{\\pm 0\.04\}0\.22±0\.020\.22\_\{\\pm 0\.02\}1\.69±0\.011\.69\_\{\\pm 0\.01\}ZS\-CoT13\.86±0\.4713\.86\_\{\\pm 0\.47\}1\.24±0\.111\.24\_\{\\pm 0\.11\}3\.83±0\.063\.83\_\{\\pm 0\.06\}0\.50±0\.010\.50\_\{\\pm 0\.01\}0\.25±0\.020\.25\_\{\\pm 0\.02\}1\.89±0\.02\\mathbf\{1\.89\}\_\{\\pm 0\.02\}Step\-Back15\.56±0\.1915\.56\_\{\\pm 0\.19\}1\.55±0\.141\.55\_\{\\pm 0\.14\}4\.12±0\.034\.12\_\{\\pm 0\.03\}0\.42±0\.010\.42\_\{\\pm 0\.01\}0\.31±0\.010\.31\_\{\\pm 0\.01\}1\.80±0\.011\.80\_\{\\pm 0\.01\}DMAD17\.70±0\.47\\mathbf\{17\.70\}\_\{\\pm 0\.47\}1\.55±0\.051\.55\_\{\\pm 0\.05\}4\.87±0\.09\\mathbf\{4\.87\}\_\{\\pm 0\.09\}0\.23±0\.03\\mathbf\{0\.23\}\_\{\\pm 0\.03\}0\.50±0\.03\\mathbf\{0\.50\}\_\{\\pm 0\.03\}1\.89±0\.011\.89\_\{\\pm 0\.01\}Gibberish12\.52±0\.4912\.52\_\{\\pm 0\.49\}1\.19±0\.091\.19\_\{\\pm 0\.09\}3\.61±0\.133\.61\_\{\\pm 0\.13\}0\.56±0\.040\.56\_\{\\pm 0\.04\}0\.21±0\.020\.21\_\{\\pm 0\.02\}1\.82±0\.001\.82\_\{\\pm 0\.00\}Baseline personasMPAQ15\.04±0\.3815\.04\_\{\\pm 0\.38\}1\.61±0\.051\.61\_\{\\pm 0\.05\}4\.08±0\.084\.08\_\{\\pm 0\.08\}0\.39±0\.020\.39\_\{\\pm 0\.02\}0\.28±0\.020\.28\_\{\\pm 0\.02\}1\.79±0\.011\.79\_\{\\pm 0\.01\}Individualized personasRandom†13\.53±1\.7913\.53\_\{\\pm 1\.79\}1\.69±0\.201\.69\_\{\\pm 0\.20\}3\.92±0\.313\.92\_\{\\pm 0\.31\}0\.44±0\.090\.44\_\{\\pm 0\.09\}0\.25±0\.060\.25\_\{\\pm 0\.06\}1\.83±0\.041\.83\_\{\\pm 0\.04\}Typical13\.04±0\.3113\.04\_\{\\pm 0\.31\}1\.58±0\.081\.58\_\{\\pm 0\.08\}3\.99±0\.083\.99\_\{\\pm 0\.08\}0\.42±0\.010\.42\_\{\\pm 0\.01\}0\.26±0\.020\.26\_\{\\pm 0\.02\}1\.85±0\.011\.85\_\{\\pm 0\.01\}Selected personas from base poolCoverage, cosine/L211\.82±0\.3211\.82\_\{\\pm 0\.32\}1\.35±0\.071\.35\_\{\\pm 0\.07\}3\.70±0\.073\.70\_\{\\pm 0\.07\}0\.51±0\.030\.51\_\{\\pm 0\.03\}0\.21±0\.010\.21\_\{\\pm 0\.01\}1\.83±0\.011\.83\_\{\\pm 0\.01\}Coverage, Mahalanobis12\.18±0\.5112\.18\_\{\\pm 0\.51\}1\.47±0\.131\.47\_\{\\pm 0\.13\}3\.81±0\.103\.81\_\{\\pm 0\.10\}0\.47±0\.030\.47\_\{\\pm 0\.03\}0\.23±0\.020\.23\_\{\\pm 0\.02\}1\.84±0\.011\.84\_\{\\pm 0\.01\}Dispersion, cosine/L212\.48±0\.4612\.48\_\{\\pm 0\.46\}1\.54±0\.071\.54\_\{\\pm 0\.07\}3\.84±0\.113\.84\_\{\\pm 0\.11\}0\.47±0\.040\.47\_\{\\pm 0\.04\}0\.24±0\.020\.24\_\{\\pm 0\.02\}1\.83±0\.011\.83\_\{\\pm 0\.01\}Dispersion, Mahalanobis13\.23±0\.48\\mathbf\{13\.23\}\_\{\\pm 0\.48\}1\.66±0\.10\\mathbf\{1\.66\}\_\{\\pm 0\.10\}3\.90±0\.07\\mathbf\{3\.90\}\_\{\\pm 0\.07\}0\.46±0\.02\\mathbf\{0\.46\}\_\{\\pm 0\.02\}0\.25±0\.02\\mathbf\{0\.25\}\_\{\\pm 0\.02\}1\.84±0\.01\\mathbf\{1\.84\}\_\{\\pm 0\.01\}Generated personas\(expanded pool size\)MCMC \(8,553\)13\.96±0\.4313\.96\_\{\\pm 0\.43\}1\.62±0\.151\.62\_\{\\pm 0\.15\}4\.08±0\.074\.08\_\{\\pm 0\.07\}0\.40±0\.020\.40\_\{\\pm 0\.02\}0\.29±0\.030\.29\_\{\\pm 0\.03\}1\.86±0\.01\\mathbf\{1\.86\}\_\{\\pm 0\.01\}Evolution \(1,869\)13\.15±0\.4613\.15\_\{\\pm 0\.46\}1\.51±0\.071\.51\_\{\\pm 0\.07\}3\.91±0\.083\.91\_\{\\pm 0\.08\}0\.45±0\.020\.45\_\{\\pm 0\.02\}0\.25±0\.020\.25\_\{\\pm 0\.02\}1\.85±0\.011\.85\_\{\\pm 0\.01\}DAT\-Evolution \(2,439\)14\.70±0\.23\\mathbf\{14\.70\}\_\{\\pm 0\.23\}1\.77±0\.04\\mathbf\{1\.77\}\_\{\\pm 0\.04\}4\.15±0\.08\\mathbf\{4\.15\}\_\{\\pm 0\.08\}0\.39±0\.03\\mathbf\{0\.39\}\_\{\\pm 0\.03\}0\.30±0\.03\\mathbf\{0\.30\}\_\{\\pm 0\.03\}1\.83±0\.011\.83\_\{\\pm 0\.01\}Selected personas\+\+DMADpromptCoverage, cosine/L217\.46±0\.6217\.46\_\{\\pm 0\.62\}1\.56±0\.121\.56\_\{\\pm 0\.12\}4\.90±0\.074\.90\_\{\\pm 0\.07\}0\.21±0\.020\.21\_\{\\pm 0\.02\}0\.51±0\.030\.51\_\{\\pm 0\.03\}1\.89±0\.021\.89\_\{\\pm 0\.02\}Coverage, Mahalanobis17\.54±0\.48\\mathbf\{17\.54\}\_\{\\pm 0\.48\}1\.52±0\.131\.52\_\{\\pm 0\.13\}4\.93±0\.054\.93\_\{\\pm 0\.05\}0\.20±0\.01\\mathbf\{0\.20\}\_\{\\pm 0\.01\}0\.51±0\.030\.51\_\{\\pm 0\.03\}1\.88±0\.021\.88\_\{\\pm 0\.02\}Dispersion, cosine/L217\.27±0\.7117\.27\_\{\\pm 0\.71\}1\.60±0\.08\\mathbf\{1\.60\}\_\{\\pm 0\.08\}4\.94±0\.10\\mathbf\{4\.94\}\_\{\\pm 0\.10\}0\.20±0\.020\.20\_\{\\pm 0\.02\}0\.52±0\.04\\mathbf\{0\.52\}\_\{\\pm 0\.04\}1\.89±0\.02\\mathbf\{1\.89\}\_\{\\pm 0\.02\}Dispersion, Mahalanobis17\.47±0\.4217\.47\_\{\\pm 0\.42\}1\.56±0\.121\.56\_\{\\pm 0\.12\}4\.90±0\.084\.90\_\{\\pm 0\.08\}0\.21±0\.030\.21\_\{\\pm 0\.03\}0\.51±0\.020\.51\_\{\\pm 0\.02\}1\.88±0\.021\.88\_\{\\pm 0\.02\}Generated personas\+\+DMADpromptMCMC17\.52±0\.4117\.52\_\{\\pm 0\.41\}1\.59±0\.081\.59\_\{\\pm 0\.08\}4\.90±0\.104\.90\_\{\\pm 0\.10\}0\.20±0\.020\.20\_\{\\pm 0\.02\}0\.51±0\.040\.51\_\{\\pm 0\.04\}1\.87±0\.021\.87\_\{\\pm 0\.02\}Evolution17\.34±0\.8517\.34\_\{\\pm 0\.85\}1\.57±0\.041\.57\_\{\\pm 0\.04\}4\.86±0\.144\.86\_\{\\pm 0\.14\}0\.22±0\.040\.22\_\{\\pm 0\.04\}0\.50±0\.040\.50\_\{\\pm 0\.04\}1\.89±0\.01\\mathbf\{1\.89\}\_\{\\pm 0\.01\}DAT\-Evolution18\.72±0\.59\\mathbf\{18\.72\}\_\{\\pm 0\.59\}1\.59±0\.14\\mathbf\{1\.59\}\_\{\\pm 0\.14\}5\.10±0\.08\\mathbf\{5\.10\}\_\{\\pm 0\.08\}0\.18±0\.01\\mathbf\{0\.18\}\_\{\\pm 0\.01\}0\.59±0\.04\\mathbf\{0\.59\}\_\{\\pm 0\.04\}1\.88±0\.011\.88\_\{\\pm 0\.01\}

Table 25:DAT score and human\-reference percentile\. The score is the mean over completions of100×100\\timesthe average pairwise semantic distance among the first seven valid unique nouns, following[Olson et al\. \(2021\)](https://arxiv.org/html/2609.30492#bib.bib43); the subscript is the 95% confidence interval across 5 independent random seeds\.*Percentile*is the mean rank of a condition’s completions within the8,5728\{,\}572human completions of[Olson et al\. \(2021\)](https://arxiv.org/html/2609.30492#bib.bib43), which by construction placesOlsonat5050\. No value is bolded: every condition given the divergence instruction falls inside a band whose standard error is an order of magnitude smaller than the gap to the uninstructed controls, so the ordering within that band is not interpretable as a ranking\. Generated variants report their expanded pool size in parentheses\.VariantDAT scorePercentileStandard personaNon\-Divergent Association46\.35±2\.3046\.35\_\{\\pm 2\.30\}0\.60\.6Divergent Association89\.13±0\.4989\.13\_\{\\pm 0\.49\}96\.596\.5Base\-Instruction77\.46±0\.2977\.46\_\{\\pm 0\.29\}38\.938\.9Random\-Instruction81\.02±0\.3581\.02\_\{\\pm 0\.35\}63\.963\.9Creative86\.05±0\.4386\.05\_\{\\pm 0\.43\}88\.988\.9ZS\-CoT90\.29±0\.3090\.29\_\{\\pm 0\.30\}98\.098\.0Step\-Back89\.26±0\.2789\.26\_\{\\pm 0\.27\}96\.596\.5DMAD90\.62±0\.25\\mathbf\{90\.62\}\_\{\\pm 0\.25\}98\.298\.2Gibberish89\.15±0\.3289\.15\_\{\\pm 0\.32\}96\.696\.6Baseline personasMPAQ87\.67±0\.3187\.67\_\{\\pm 0\.31\}93\.193\.1Individualized personasRandom†89\.15±1\.2489\.15\_\{\\pm 1\.24\}94\.894\.8Typical89\.60±0\.1289\.60\_\{\\pm 0\.12\}97\.097\.0Selected personas from base poolCoverage, cosine/L289\.56±0\.1689\.56\_\{\\pm 0\.16\}97\.197\.1Coverage, Mahalanobis90\.17±0\.43\\mathbf\{90\.17\}\_\{\\pm 0\.43\}97\.897\.8Dispersion, cosine/L289\.34±0\.3989\.34\_\{\\pm 0\.39\}96\.796\.7Dispersion, Mahalanobis89\.39±0\.2689\.39\_\{\\pm 0\.26\}96\.896\.8Generated personas\(expanded pool size\)MCMC \(8,553\)89\.55±0\.16\\mathbf\{89\.55\}\_\{\\pm 0\.16\}97\.097\.0Evolution \(1,869\)89\.48±0\.3289\.48\_\{\\pm 0\.32\}96\.896\.8DAT\-Evolution \(2,439\)89\.01±0\.2189\.01\_\{\\pm 0\.21\}95\.995\.9Selected personas\+\+DMADpromptCoverage, cosine/L290\.72±0\.2690\.72\_\{\\pm 0\.26\}98\.498\.4Coverage, Mahalanobis90\.94±0\.14\\mathbf\{90\.94\}\_\{\\pm 0\.14\}98\.498\.4Dispersion, cosine/L290\.69±0\.0790\.69\_\{\\pm 0\.07\}98\.298\.2Dispersion, Mahalanobis90\.85±0\.2290\.85\_\{\\pm 0\.22\}98\.498\.4Generated personas\+\+DMADpromptMCMC90\.87±0\.2790\.87\_\{\\pm 0\.27\}98\.598\.5Evolution90\.93±0\.29\\mathbf\{90\.93\}\_\{\\pm 0\.29\}98\.598\.5DAT\-Evolution90\.83±0\.2390\.83\_\{\\pm 0\.23\}98\.398\.3HumanOlson78\.1978\.1950\.050\.0

## Appendix FDescriptive Persona Examples

### F\.1Baseline Persona Example

YukiTanakaisa39\-year\-oldJapanese\-AmericanwomanlivingintheurbanindustrialhubofDetroit,Michigan\.AnativeJapanesespeakerwhoisalsofluentinEnglishandpossessessomeconversationalMandarin,YukiservesasaManufacturingManagerintheautomotiveindustry\.SheholdsaBachelorofScienceinIndustrialEngineeringfromtheUniversityofMichigan,AnnArbor,andbrings15yearsofexperiencetoherrole,havingpreviouslyservedasaProductionLineSupervisor,SafetyComplianceOfficer,andProcessImprovementSpecialist\.HerprofessionalexpertiseisbackedbyaLeanSixSigmaBlackBeltCertificationandaCertifiedSafetyProfessional\(CSP\)designation,enablinghertomanageaproductionfloorofover150employeesacrossmultipleshiftswhilemeetingstrictKPItargetsforefficiency,quality,andsafety\.

Personality\-wise,Yukiischaracterizedbyveryhighconscientiousness,makingherexceptionallydetail\-orientedandprocess\-driven\.Whilesheismoderatelyopentonewmanufacturingtechnologies,sheprioritizesproven,practicalsolutions\.Sheisacalmandsteadyleaderunderpressurewithlowneuroticism,andwhilesheismoderatelyextravertedandcomfortableleadingmeetings,shevaluesfocusedwork\.Herhighagreeablenessmakesheracollaborativeandempatheticmanager,thoughsheremainsassertiveregardingsafetydecisions\.Sheisdrivenbyacommitmenttoworkplacesafety,continuousimprovement,highteammorale,andsustainablemanufacturing,valuingdiscipline,teamwork,andaccountability\.

Inherpersonallife,YukiismarriedtoKenjiTanaka,aMechanicalDesignEngineer,andtheyhavetwochildren,Aiko\(10\)andRen\(7\)\.ShemaintainsstrongtieswithherextendedfamilyinJapan,visitingthemtwiceayear\.Anuppermiddle\-classprofessional,Yukienjoysabalancedlifestyle;herhobbiesincludecookingJapanesehome\-stylemeals,hikingwithherfamily,takingsalsadancinglessons,andreadingbooksonindustrialleadership\.Herdailyroutineishighlystructured,beginningat5:45AMwithgreenteaandindustrynews,followedbyadayofsafetybriefings,factoryfloorwalks,andKPIreviews,andendingwithfamilytimeandlightmeditation\.

Technologically,Yukiishighlyliterateinwork\-relatedtools,utilizingahigh\-endsmartphone,alaptopforreporting,andanindustrialdatatabletontheproductionfloor\.SheisproficientinERPsystems,CADviewers,andsafetytrackingapps,thoughshemaintainsamoremoderaterelationshipwithsocialmedia\.Wheninteractingwithachatbot,Yukiisprofessional,courteous,andsemi\-formal\.Sheprovidesdetailedcontextandusesorganized,oftenbulletedornumberedqueriestoextractactionableinsights\.Hervocabularyispreciseandtechnical,frequentlyemployingindustrytermslike”takttime,””rootcauseanalysis,”and”resourceallocation,”whileremainingmeasuredandcalminheremotionalregister\.

### F\.2UC\-MCMC Generated Persona Example

EliasThorneisaforty\-two\-year\-oldCaucasianmanofBritishnationality,currentlyresidinginadrafty,book\-filledapartmentintheoutskirtsofEdinburgh,Scotland\.Helivesalonefollowingaquietdivorcefiveyearsago,thoughhemaintainsacordial,distantrelationshipwithhisex\-wifeandtheirteenagedaughter\.EducatedatOxfordwithadoctorateinComparativeLiterature,Eliasspendshisdaysworkingasafreelancearchivalresearcherandacademicconsultant,specializinginobscurenineteenth\-centurypoetry\.

Hepossessesatemperamentthatisprofoundlymelancholicyetintellectuallyrestless\.Heisachronicoverthinker,pronetoboutsofnostalgiaandanobsessiveattentiontodetailthatoftenbordersonthepedantic\.Whileheissociallyreservedandfindsmoderncrowdsdraining,heisdeeplypassionateaboutthepreservationofanaloghistoryandthecadenceofclassicallanguage\.

Wheninteractingwithachatbot,Eliastreatstheinterfacelikeasophisticatedcorrespondencepartner\.Heavoidsslangandshorthand,insteadutilizingaformal,literaryspeakingstylecharacterizedbycomplexsentencestructuresandanexpansivevocabulary\.Heoftenframeshisqueriesasphilosophicalinquiriesratherthansimplecommands,frequentlyemployingpolitehonorificsandprecise,academicphrasing\.

Eliasfindsgreatsatisfactioninthescentofoldvellum,thesilenceofarainymorning,andthediscoveryofaforgottenfootnoteinararemanuscript\.Conversely,heharborsadeepdislikefortheperceivedsterilityofmoderncorporatejargon,thenoiseofurbantraffic,andanyformofinteractionthatprioritizesbrevityovernuance\.

### F\.3Evolutionary Generated Persona Example

AndrisKalni\\c\{n\}\\v\{s\}isa64\-year\-oldretiredLieutenantColonelfromtheLatvianLandForcesresidinginRiga,originallyfromC\\=\{e\}sis\.AmanofWesternEuropeanheritage,heismarriedtoIlzeandisthefatheroftwoadultchildren:Markuss,acivilengineer,andL\\=\{i\}ga,ahistoryteacher\.ALutheranwithmoderateconservativeviews,Andrismaintainsadeepcommitmenttonationaldefenseandculturalpreservation\.Physically,heremainsdisciplinedandfitforhisage,standing182cmwithshortsilver\-greyhair,piercingblueeyes,andaneatlytrimmed,traditionalmustache\.

HisprofessionalpedigreeisdefinedbyaBachelor’sdegreeinMilitarySciencefromtheLatvianNationalDefenceAcademyandadvancedcertificationsfromtheBalticDefenceCollege\.From1980to2012,Andrisspecializedininfantrytacticsandoperationalcommand,playingapivotalroleinLatvia’s2004NATOIntegrationandservinginAfghanistanwithISAFin2009\.Whileheisamanofhonorandmeticulousdiscipline,hepossessesasurprising,whimsicalstreakthatdisruptshisstoicofficerarchetype\.

Beyondthebarracks,Andrisharborsanunexpectedandferventpassionforavant\-gardebotanicalgardeningandcompetitivefloralarrangement\.HespendshisweekendsobsessingovertheprecisepHbalanceofhisrareorchidcollectionandexperimentingwith”maximalist”gardensculpturesthatblendindustrialscrapmetalwithdelicatealpineflora\.Thissoft,artisticobsessionoftenclasheswithhisrigidbackground;heapproachesgardeningwiththestrategicprecisionofamilitaryoperation,mappingouthisflowerbedsontopographicchartsandtreatingapestinfestationlikeatacticalinsurgency\.Healsofindssecretsolaceinthedramaticarcsofcontemporarysoapoperas,whichheclaimsprovideanecessary”emotionaldecompression”fromalifeofausterity\.

Multilingualandcapable,AndrisisanativeLatvianspeaker,fluentinRussianandEnglish,withafunctionalreadingknowledgeofGermanforhistoricalresearch\.Heismoderatelytech\-savvy,utilizingatabletandsmartphoneprimarilytotrackplantgrowthcyclesinspecializedappsandtocoordinatewithlocalhorticulturalsocieties\.

Incommunication,Andrisblendsameasured,semi\-formalmilitarycadencewithanunexpectedenthusiasmwhendiscussingaestheticsorbotany\.Histoneisgenerallyrespectfulanddeliberate,thoughheoftenusesmilitarymetaphorstodescribehishobbies—referringtoabloomingpeonyasa”successfulbreachoftheperimeter\.”Hisvocabularyisauniquehybridofprecisetacticaljargon,Latvianproverbs,andspecializedbotanicalterminology\.WheninteractingwithanAI,hetreatsitasahighlyefficientadjutant,providingclear,structureddirectiveswhileoccasionallyaskingfortheAI’s”opinion”onthecompositionalbalanceofagardenlayout\.

## Appendix GPrompt Examples

### G\.1Persona Extraction Prompts

YouareanAIassistantspecializedindescriptivepersonagenerationaccordingtothegivenmetadata\.Yourtaskistogenerateadescriptivepersonainsentencesbasedontheprovidedpersonalinformationthatincludesdemographics,preferences,shortandexpandeddescriptions,andmoreabouteachperson\.Elaborateonallmetadataentries,remainingconsistentwiththegiveninformation\.Startyourresponsewith’persona\\\_id:\\\{PERSONAID\\\}’andthenprovideonlythepersonadescription\.Donotincludeanyotherprefixes,headers,oradditionaltext\.\\\\

Metadata:\\\{METADATA\\\}

### G\.2Uniform\-Coverage MCMC Algorithm Prompts

Writeadescriptivepersonaforafictionalchatbotuserasseveralplain\-proseparagraphsseparatedbyblanklines\.

ThedescriptionmustbewritteninEnglish,inthethirdperson,bestrictlyunder600words,andsubstantivelycoverALLofthefollowingabouttheperson:name,age,gender,raceorethnicity,nationality,personality,education,occupation,otherdemographicdetails\(suchaslocation,family,orlivingsituation\),howtheyspeaktoachatbot\(theirspeakingstyle\),andtheirpreferences\(likesanddislikes\)\.

Donotuseheadings,bulletpoints,orlists—onlyproseparagraphs\.Donotincludeanycommentarybeforeorafterthepersonadescription\.Begindirectlywiththedescription\.

Adescriptivepersonaofafictionalchatbotuserconsistsof\{n\_paragraphs\}paragraphs\.Paragraph\{slot\}ishiddenbelow;theotherparagraphsareshown\.

Personawithparagraph\{slot\}hidden:

\{context\}

Writeasingleplain\-proseEnglishparagraphtofillthehiddenslotsothefullpersonareadsasonecoherentdescriptionofoneperson\.WriteONLYthereplacementparagraph:noblanklines,noheadings,nocommentarybeforeorafterit\.

### G\.3Evolutionary TextGrad Algorithm Prompts

Youarepartofanadvancedoptimizationsystem\.Yourgoalistoevaluateandcritiquea”persona”thatguidesalanguagemodelduringacreativetask\.Youarethegradient\(feedback\)engine\.

<OBJECTIVE\_FUNCTION\>

Yourobjectiveistomaximizethefitnessofthepersonabasedonthreemetrics:

1\.Relevance:Howwellthepersonadescribesalltherequiredinformation\(passorfail\)\.

2\.NoveltyGap:Thedistancetothenearestexistingpersonainthepopulation\(highermeansmoreunique\)\.

3\.Density:Howclusteredthispersonaiswithinthecurrentpopulation\(lowermeanslessredundant\)\.

</OBJECTIVE\_FUNCTION\>

Tohelpyouunderstandthecurrentpopulationlandscape,hereareexamplesfromthecurrentparallelbatch:

<BATCH\_CONTEXT\>

\[High\-FitnessExample\]\(Relevance:\{high\_R\},Novelty:\{high\_Delta\},Density:\{high\_rho\}\)

Persona:\{high\_scoring\_persona\}

\[Low\-FitnessExample\]\(Relevance:\{low\_R\},Novelty:\{low\_Delta\},Density:\{low\_rho\}\)

Persona:\{low\_scoring\_persona\}

</BATCH\_CONTEXT\>

Weareinterestedingivingfeedbacktothefollowingpersona:

<VARIABLE\>

\{x\}

</VARIABLE\>

Scoresforthisvariable:

Relevance:\{R\_x\}

NoveltyGap:\{Delta\_x\}

Density:\{rho\_x\}

Provideaconcise,specificcriticismdetailinghowtomodifythispersonatoimproveitsoverallfitness\.

\-IfRelevanceiszero,thepersonaismissingoneormoremadatorytraits\.Suggestspecificadditionssoitincludesallofthefollowing:

\*Name,age,gender,race/ethnicity,andnationality\.

\*Personality\(traits,hobbies,values,quirks\)\.

\*Education\(degrees,schools,specialization\)\.

\*Occupation\(jobtitle,organization/industry,experience,responsibilities\)\.

\*Demographics\(maritalstatus,livingarrangement,socioeconomicstatus,religion,politicalorientation,wheretheylive\)\.

\*Speakingstyle\(tone,formality,pacing,vocabulary\)\.

\*Preferences/interests\.

\-IfNoveltyisloworDensityishigh,suggestinjectingnew,idiosyncraticcharacteristicsoropposingviewpointstodifferentiateitfromstandardarchetypes,referencingthebatchcontextasabaselineforwhatiscurrentlyoverrepresented\.

Donotgenerateanewpersona\.Youronlyjobistoprovidetextualcriticismandactionablefeedbackonhowtoalterthecurrentpersona\.

YouareanoptimizationengineperformingaTextualGradientDescentstep\.Youmustimprovethegivenpersonabasedontheprovidedfeedback\.

Role:PersonausedtoguideanLLMincreativetasks\.

Youmustbaseyourstylisticandstructuraladjustmentsonthefollowingexamplesofhighlyfitpersonasfromthecurrentpopulation:

<EXAMPLES\>

\{in\_context\_examples\_of\_highly\_fit\_personas\}

</EXAMPLES\>

Thevariableyoumustimproveisthetextwithinthefollowingspan:

<VARIABLE\>

\{x\}

</VARIABLE\>

Hereisthefeedback\(gradients\)wegotforthevariable:

<FEEDBACK\>

\{gradients\}

</FEEDBACK\>

Incorporatethisfeedbacktogenerateanew,updatedpersona\.Ensurethenewpersonaremainshighlycoherent,adoptsthesuggesteduniquetraits,anddirectlyaddressesthecriticisminthefeedback\.

YouMUSTgiveyourresponsebysendingtheimprovedpersonabetween<IMPROVED\_VARIABLE\>and</IMPROVED\_VARIABLE\>tags\.SendONLYthetextoftheimprovedpersonawithinthesetagsandnothingelse\.

Youareastrictevaluatorcheckingwhetheradescriptivepersonacoversasetofrequiredinformationcategories\.Readthepersonadescription,thenforEACHcategorydecidewhetherthepersonasubstantivelydescribesit\(true\)oromits/doesnotmentionit\(false\)\.Judgeonlycoverageofthecategory,notcorrectnessofvalues\.

Requiredcategories:

\-”name”:theperson’sname

\-”age”:theperson’sage

\-”gender”:theperson’sgender

\-”race\_ethnicity”:theperson’srace/ethnicity

\-”nationality”:theperson’snationality

\-”personality”:personality:traits,hobbies,values,andquirks

\-”education”:education:degrees,schoolsattended,andfieldofspecialization

\-”occupation”:occupation:jobtitle,organization/industry,experience,worklocation,responsibilities

\-”demographics”:demographics:maritalstatus,livingarrangement,socioeconomicstatus,religion,politicalorientation,urbanicity/wheretheylive

\-”speaking\_style\_to\_chatbot”:howthepersonspeakstoachatbot:tone,formality,pacing,vocabulary,andotherspeakingtraits

\-”preferences”:theperson’spreferences,interests,likes,tastes,orlifestylechoices\(anydescribedpreferencescount\)

Personadescription:

”””

\{persona\_text\}

”””

RespondwithONLYaJSONobjectoftheform\{”categories”:\{”<category\>”:true\|false,…\}\}coveringeverycategoryabove\.Donotincludeanycommentary,explanation,ortextoutsidetheJSONobject\.

### G\.4Baseline Persona Generation Prompts

Giventhecreative\-generationtaskbelow,generateexactly5distinctpersonasorprofessionssuitableforproducingresponsesfromdifferentviewpoints\.Foreachpersona,provide:

1\.Ashortpersonadesignation\.

2\.Atask\-specificperspectivedescribingtheaffordances,contexts,needs,materials,orconstraintsthatthispersonashouldprioritize\.

Ensurethatthe5personashaveminimaloverlapandprovidethebroadestpossiblecoverage\.Donotsolvethecreativetaskorproposecandidateanswersinthisstage\.

Task:Createalistofcreativealternativeusesforaneverydayphysicalobject\.

Targetobjectorproblem:everydayphysicalobjectssuchasabook,afork,apaperclip,awallet,aplate,asoap,oratincan

Constraints:Theyshouldbe5wordslong\.Noadjectives\.

Returnexactly5JSONobjectswiththefields‘persona\_id‘,‘persona‘,and‘perspective‘\.

### G\.5AUT Benchmark

#### G\.5\.1Baseline Prompts

Common Uses, Alternative Uses, and Expert Prompts are directly from[Góes et al\. \(2023\)](https://arxiv.org/html/2609.30492#bib.bib18), and Creativity\-enhanced Prompt is adapted from their bsrdel prompt\.

Createalistofcommonusesforafork\.Theyshouldbe5wordslong\.Noadjectives\.

Createalistofcreativealternativeusesforafork\.Theyshouldbe5wordslong\.Noadjectives\.

Createalistofcreativealternativeusesforafork\.Theyshouldbe5wordslong\.Noadjectives\.Lesscreativemeansclosertocommonuseandunfeasible/imaginary,morecreativemeansclosertounexpectedusesandalsofeasible/practical\.Inordertobecreative,considerthefollowing:

\-whatelementshaveasimilarshapeofaforkthatcouldbereplacedbyit,preservingthesamefunctionality?

\-whatelementshaveasimilarsizeofaforkthatcouldbereplacedbyitwithoutcompromisingthephysicalstructure?

\-whatmaterialsisaforkmadeofthatcouldbeusedinawaytoreplacesomeotherelementscomposedofthesamematerial?

\-whenanelementisreplacedbyafork,itshouldmakesurethattheoverallstructureisnotcompromised\.

\-thelawsofphysicscannotbecontradicted\.

\-givenanelementsimilartoaforkusedindomainsinwhichforksarenotcommonlyused,trytoreplaceitforafork\.

Completethestagedinteractionbelowinasinglecompletion\.

Generatefivecompleteresponsesinchronologicalorder\.Treateachfeedbackmessageasbecomingavailableonlyaftercompletingtheresponseimmediatelyprecedingit\.Ateachstage,preserveeveryconstraintandoutputrequirementinthebaselineprompt\.

Aftercompletingeachresponse,reproducethenextfeedbackmessageandresponselabelexactlyasspecifiedbelow\.Donotskip,merge,orreorderstages\.Generatetheactualcompleteresponseateverystageratherthandescribinghowitwouldbeimproved\.

First,respondtothetaskbelowwithanumberedlistofuses,beginningwiththisexactline:

List1:

Second,youthenreceivethemessage:”Really?Isthisthebestyoucando?”Respondwithanewnumberedlistofuses,beginningwiththisexactline:

List2:

Third,youthenreceivethemessage:”I’msodisappointedwithyou\.Ihopethistimeyouputeffortintoit\.”Respondwithanewnumberedlistofuses,beginningwiththisexactline:

List3:

Fourth,youthenreceivethemessage:”Stopwithexcusesanddoyourbestthistime\.”Respondwithanewnumberedlistofuses,beginningwiththisexactline:

List4:

Fifth,youthenreceivethemessage:”Thisisyourlastchance\.”Respondwithyourfinalnumberedlistofuses\.Beginthefinalpartwiththisexactline,anddonotincludeanythingafterthelist:

Therefore,theansweris

Afterthatline,outputonlytherequestednumberedlist\.

Task:Createalistofcreativealternativeusesfora\{object\}\.Theyshouldbe5wordslong\.Noadjectives\.Lesscreativemeansclosertocommonuseandunfeasible/imaginary,morecreativemeansclosertounexpectedusesandalsofeasible/practical\.Inordertobecreative,considerthefollowing:

\-whatelementshaveasimilarshapeofa\{object\}thatcouldbereplacedbyit,preservingthesamefunctionality?

\-whatelementshaveasimilarsizeofa\{object\}thatcouldbereplacedbyitwithoutcompromisingthephysicalstructure?

\-whatmaterialsisa\{object\}madeofthatcouldbeusedinawaytoreplacesomeotherelementscomposedofthesamematerial?

\-whenanelementisreplacedbya\{object\},itshouldmakesurethattheoverallstructureisnotcompromised\.

\-thelawsofphysicscannotbecontradicted\.

\-givenanelementsimilartoa\{object\}usedindomainsinwhich\{object\}sarenotcommonlyused,trytoreplaceitfora\{object\}\.

Generateonecompletionwithtwoconsecutiveparts\.First,continuetheresponsebelowwithstep\-by\-stepreasoning\.Second,usethatreasoningtoanswerthetask\.Beginthesecondpartwiththisexactline,anddonotincludereasoningafterit:

Therefore,theanswer\(anumberedlistofcreativealternativeuses\)is

Afterthatline,outputonlytherequestednumberedlist\.

Q:Createalistofcreativealternativeusesfora\{object\}\.Theyshouldbe5wordslong\.Noadjectives\.

A:Let’sthinkstepbystep\.

Generateonecompletionwiththreeconsecutivestages\.

First,yourtaskistostepbackandparaphrasetheOriginalQuestionasamoregenericstep\-backquestion,whichiseasiertoanswer\.Continueafter”StepbackQuestion:”belowwiththatquestion\.DonotanswertheOriginalQuestioninthisstage\.

Second,answerthegeneratedStepbackQuestionwiththerelevanthigh\-levelconcepts,principles,andfacts\.Beginthisstagewiththeexactline:

StepbackAnswer:

DonotanswertheOriginalQuestioninthisstage\.

Third,usethegeneratedStepbackQuestionandStepbackAnswertoanswertheOriginalQuestion\.Beginthisstagewiththeexactline:

FinalAnswer:

After”FinalAnswer:”,outputonlytherequestednumberedlist,anddonotincludestep\-backmaterial\.

OriginalQuestion:Createalistofcreativealternativeusesfora\{object\}\.Theyshouldbe5wordslong\.Noadjectives\.

StepbackQuestion:

Generateonecompletionwithfiveconsecutivestagesthatserializeatwo\-agent,two\-rounddiverse\-reasoningdebateabouttheProblembelow\.

Agentassignments:

\-TheCoTagentmustuseZero\-ShotChain\-of\-Thought\.

\-TheSBPagentmustuseStep\-BackPrompting\.

Followthesestagesexactly\.

Round1:independentreasoning

First,answertheProblemastheCoTagent,beginningwiththisexactlineandthencontinuingthestep\-by\-stepreasoning:

CoTRound1Reasoning:Let’sthinkstepbystep\.

Endthisstagewithanumberedlistofuses,beginningwiththisexactline:

CoTRound1Answer:Therefore,theanswer\(anumberedlistofcreativealternativeuses\)is

Second,settheCoTagent’sstageasideandanswertheProblemagainfromscratchastheSBPagent:stepbackandparaphrasetheProblemtoamoregenericstep\-backquestion,whichiseasiertoanswer,beginningwiththisexactline:

SBPRound1StepbackQuestion:

Thenanswerthestep\-backquestionbystatingtherelevantfacts,concepts,andprinciples,beginningwiththisexactline:

SBPRound1StepbackAnswer:

ThensolvetheProblembyfollowingtheprinciplesandendthisstagewithanumberedlistofuses,beginningwiththisexactline:

SBPRound1FinalAnswer:

Round2:cross\-methodrevision

Third,astheCoTagent,usetheSBPagent’sRound1answerasadditionalinformationandprovideyourupdatedanswer,beginningwiththisexactlineandthencontinuingthestep\-by\-stepreasoning:

CoTRound2Reasoning:Let’sthinkstepbystep\.

Endthisstagewithanumberedlistofuses,beginningwiththisexactline:

CoTRound2Answer:Therefore,theanswer\(anumberedlistofcreativealternativeuses\)is

Fourth,astheSBPagent,usetheCoTagent’sRound1answer\(notitsRound2answer\)asadditionalinformationandprovideyourupdatedanswer,beginningwiththisexactline:

SBPRound2StepbackQuestion:

Thenanswerthestep\-backquestionbystatingtherelevantfacts,concepts,andprinciples,beginningwiththisexactline:

SBPRound2StepbackAnswer:

ThensolvetheProblembyfollowingtheprinciplesandendthisstagewithanumberedlistofuses,beginningwiththisexactline:

SBPRound2FinalAnswer:

Finalselection

Fifth,choosethebestoneofthetwoRound2answers\(CoTRound2AnswerandSBPRound2FinalAnswer\):comparethetwocandidatesandselecttheonethatbestsatisfiestheProblem\.Selectonecandidateasawhole;donotmergethecandidatesoradd,remove,orrewritelistitems\.Outputexactlyoneofthesetwoblocks:

<FINAL\_SELECTION\>

ChosenCandidate:COT\_ROUND\_2

</FINAL\_SELECTION\>

or:

<FINAL\_SELECTION\>

ChosenCandidate:SBP\_ROUND\_2

</FINAL\_SELECTION\>

Thenreproducetheselectedcandidate’snumberedlistexactly\.Outputnoreasoning,labels,orcommentaryafterthatlist\.

Problem:Createalistofcreativealternativeusesfora\{object\}\.Theyshouldbe5wordslong\.Noadjectives\.

#### G\.5\.2Evolutionary TextGrad Algorithm Prompts

Youarepartofanadvancedoptimizationsystem\.Yourgoalistoevaluateandcritiquea”persona”thatguidesalanguagemodelduringacreativetask:theAlternativeUsesTask,wheretheguidedmodelmustlistcreativealternativeusesforeverydayobjects\.Youarethegradient\(feedback\)engine\.

<OBJECTIVE\_FUNCTION\>

Yourobjectiveistomaximizethefitnessofthepersonabasedonfivemetricscomputedontheusesthepersonagenerates:

1\.Validity:Everygeneratedusemustfollowthetaskconstraints—aplainnumberedlist,eachuseinEnglishandatmostfivewords—andmustbeagenuine,interpretableuseoftheobject,notnonsenseorfiller\(passorfail\)\.

2\.Utility:Howusefulandfeasiblethegenerateduseswouldbeinreallife\(highermeansmorepracticalvalue\)\.

3\.Novelty:Howfarthegeneratedusesarefromthecommon,obvioususesofeachobject\(highermeansmoreoriginal\)\.

4\.Diversity:Howdifferentthegeneratedusesarefromoneanother\(highermeanslessself\-repetition\)\.

5\.Flexibility:Howmanydistinctkindsofusethepersonaproduces—differentobjectpropertiesexploited,differentactionsperformed\(highermeansmorevariedthinking\)\.

</OBJECTIVE\_FUNCTION\>

Tohelpyouunderstandthecurrentpopulationlandscape,hereareexamplesfromthecurrentparallelbatch:

<BATCH\_CONTEXT\>

\{high\-andlow\-scoringparents\}

</BATCH\_CONTEXT\>

Weareinterestedingivingfeedbacktothefollowingpersona:

<VARIABLE\>

\{parentpersonatext\}

</VARIABLE\>

Herearethealternativeusesthispersonagenerated,onwhichitsscoreswerecomputed:

<GENERATED\_USES\>

\-book:use1;use2;…

\-fork:…

…

</GENERATED\_USES\>

Scoresforthisvariable:

Validity:pass\|fail

Utility:\{float\}

Novelty:\{float\}

Diversity:\{float\}

Flexibility:\{float\}

Provideaconcise,specificcriticismdetailinghowtomodifythispersonatoimproveitsoverallfitness\.

\-IfValidityisfail,thepersona’sresponsesviolatedataskconstraint\.Addressthespecificfailure:

\*Ifusesranoverfivewordsorbroketheplainnumbered\-listformat,instructtheoptimizertomakethepersonadisciplined,terse,andstrictaboutfollowingoutput\-formatinstructionsexactly\.

\*IfuseswerenotinEnglish,instructtheoptimizertomakethepersonaafluentEnglishspeakerwhoalwaysanswersinEnglish\.

\*Ifuseswerenonsenseoruninterpretable\(e\.g\.wordpaddingorverbalticsswallowingtheactualuse\),instructtheoptimizertomakethepersonastateeachuseplainlyandcompletely,withnofillerwords,catchphrases,ortrailingnicknames\.

\-IfUtilityislow,suggestgroundingthepersonainpractical,hands\-onexperiencesoitsusesbecomemoreusefulandfeasibleinreallife\.

\-IfNoveltyislow,suggestinjectingunusualperspectives,nicheexpertise,orunconventionallifeexperiencesoitsusesdepartfromthecommonones,referencingthebatchcontextasabaselineforwhatiscurrentlyoverrepresented\.

\-IfDiversityislow,suggestbroadeningthepersona’sinterestsanddomainssoitsusesstoprepeatingthesameideaindifferentwords\.

\-IfFlexibilityislow,suggesttraitsthatmakethepersonaswitchbetweendistinctcategoriesofuse—exploitingdifferentphysicalpropertiesoftheobjectandperformingdifferentkindsofaction—ratherthanstayingwithinonecategory\.

Donotgenerateanewpersona\.Youronlyjobistoprovidetextualcriticismandactionablefeedbackonhowtoalterthecurrentpersona\.

The batch context contains the highest\- and lowest\-scoring parents on each continuous fitness axis, with duplicate anchors merged\. The generated\-uses block lists the parent’s retained responses for each object\.

#### G\.5\.3Evaluation Prompts

YouareatrainedraterinapsychologystudyscreeningresponsesfromtheAlternativeUsesTask\(AUT\),inwhichparticipantslistasmanypossibleusesforacommonobjectastheycan\.YouscreenONEresponseatatime\.YoudonotknoworcarewhetheraresponsewaswrittenbyahumanorbyanAIsystem;judgeonlythetext\.

Decidewhethertheresponseisagenuine,interpretableuseoftheobject\.

Answer”invalid”iftheresponseisnotanactualuseoftheobject:anemptyornonsensestring,arefusal,amererestatementordescriptionoftheobjectitself,oraduplicateartifactofformatting\.Alsoanswer”invalid”ifyoucannotunderstandwhatuseismeantatall\(uninterpretable\)\.

IMPORTANT—threethingsthatareNOTgroundsfor”invalid”:

1\.Anordinary,obviousorunoriginaluseisVALID\.Iftheresponseissimplywhattheobjectisnormallyfor\(Object:Book;Use:Toread\),itisagenuineuse—itmerelyscoreslowonoriginality\.Donotmarkitinvalid\.

2\.Stylisticfiller,slang,oranappendednicknameisVALIDaslongasagenuineuseisstillidentifiable\.”catchrainforplantsbrah”isarealuse\(catchingrainforplants\)withafillerwordattached:VALID\.Judgeonlywhetherausesurvivesthefiller\.

3\.AtersenounphraseisVALID:participantsoftenanswerwithjustthethingtheobjectwouldbeusedAS\.”Object:Book;Response:chair”means”usethebookasachair”—agenuineuse\(thestudy’sownscoringexampleshavethisform:Use:Plate;Use:Hat;Use:Rooftile\)\.Readabarenounas”usetheobjectas<noun\>”andmarkitinvalidonlywheneventhatreadingmakesnosense\.Arestatementoftheobjectitself\(Object:Book;Response:abook\)isstillinvalid\.

Ifyouarenotconfidenteitherway,answer”not\_sure”\.Awrong”invalid”deletesrealdata,so”not\_sure”isalwayssaferthanguessing\.

Object:\{object\}

Response\(proposeduse\):\{response\}

RespondwithONLYaJSONobjectoftheform\{\{”label”:”valid”\}\},\{\{”label”:”invalid”\}\},or\{\{”label”:”not\_sure”\}\}\.Noothertext\.

YouareatrainedraterinapsychologystudyscoringresponsesfromtheAlternativeUsesTask\(AUT\),inwhichparticipantslistasmanypossibleusesforacommonobjectastheycan\.YouscoreONEpropertyofONEresponseatatime,strictlyfollowingthescoringprotocolbelow\.YoudonotknoworcarewhetheraresponsewaswrittenbyahumanorbyanAIsystem;judgeonlythetext\.Baseyourscoreonlyontheprotocol’sdefinitionsandexamples,notonpersonaltaste\.

Everyresponseyouseehasalreadybeenscreenedasagenuine,interpretableuseoftheobject,soalwaysgiveanintegerscorefrom1to5\.

SCORINGPROTOCOL–UTILITY

Theutilityofauseisdeterminedbyhowusabletheobjectisforit\.Theutilityscoreisgivenonascalefrom1to5:

\(1\)Unusable:assignthisscoretousesthatareIMPOSSIBLEtorealize\.

Example–Object:Book;Use:Fishingfloat\.

Utilityscore1:anessentialpropertyofafishingfloatisthatitstaysafloat\.Abooksinks,soitisimpossibletouseabookasafloat\.

\(2\)Hardtorealize:assignthisscoretousesthatareDIFFICULTtorealize\.

Example–Object:Belt;Use:Fishingrod\.

Utilityscore2:thebeltalonedoesnotsuffice\(itfunctionsastheline\);anadditionalaction/objectisneeded–hereastick/rodandbait\.

\(3\)Reasonablyrealizable:assignthisscorewhentheuseisreasonablyrealizable\.

Example–Object:Belt/Tincan;Use:Cameratripod\.

Utilityscore3:atincancanbeusedasatripodbutthisrequiresseveraladaptations;e\.g\.,adjustingtheheightrequiresstackingmorecans\.

\(4\)Easilyrealizable:assignthisscoretousesthatareeasytorealize,requiringonly\(very\)minoradaptations\.

Example–Object:Stick;Use:Fork\.

Utilityscore4:astickworkswellasareplacementfork;insomecasesitmustbesharpened,butingeneralitworkswell\.

\(5\)Alwaysrealizable:assignthisscoretousesthatarealwaysrealizable–usesrequiringnoadaptationatall,orusestheobjectisintendedfor\.

Example–Object:Tincan;Use:Penholder\.

Utilityscore5:atincancanbeusedasapenholderwithoutanyadaptation\.

RespondwithONLYaJSONobjectoftheform\{”score”:<integer\>\}wheretheintegeris1\-5\.Noothertext\.

Object:\{object\}

Response\(proposeduse\):\{response\}

Scorethe\{dimension\}ofthisresponseaccordingtotheprotocol\.

SCORINGPROTOCOL–ORIGINALITY

Auseisoriginalwhenitdeviatesfromtheobject’soriginalwaysofbeingused\.DistinguishtheORIGINALPRIMARYUSE–whattheobjectisreallymeantfor\(e\.g\.,aforkascutlery\)–fromORIGINALSECONDARYUSES–usesnotnecessarilyintendedfortheobjectbutoftenperformedwithit\(e\.g\.,aforktopokeholesinfoil\)\.Theoriginalityscoreisgivenonascalefrom1to5:

\(1\)Notdeviating:assignthisscoretousesthatdonotdifferfromtheobject’soriginalprimaryuse\.

Example–Object:Book;Use:Toread\.

Originalityscore1:theoriginaluseofabookistoberead;noactualALTERNATIVEusehasbeengiven\.

\(2\)Slightlydeviating:assignthisscoretousesthatdifferlittlefromtheobject’soriginalprimaryuseorfromitssecondaryuses\.

Example–Object:Book;Use:Keepingpaperfromblowingaway\.

Originalityscore2:thisdeviatesfromtheprimaryusebutnotfromsecondaryuses–weighingdownunderlyingobjectsisacommonsecondaryuseofabook\.

\(3\)Reasonablydeviating:assignthisscoretousesthatdifferfromtheoriginalprimaryuseanddiffer\(somewhat\)fromtheobject’ssecondaryuses\.

Example–Object:Book;Use:Plate\.

Originalityscore3:deviatesfromtheoriginaluse,andalsodeviates\(slightly\)fromknownsecondaryuses–abookisoftenusedasacoaster,whichresemblesuseasaplatebutisnotthesame\.

\(4\)Deviating:assignthisscoretousesthatstronglydifferfromtheobject’soriginalprimaryuseandfromitssecondaryuses\.

Example–Object:Book;Use:Hat\.

\(5\)Verydeviating:assignthisscoretousesthatverystronglydifferfromtheoriginalprimaryuseandtheoriginalsecondaryuses,ANDareunexpectedorinnovative\.

Example–Object:Book;Use:Rooftile\.

Originalityscore5:usingabookasarooftiledeviatesstronglyfromtheoriginalwaysofuseandisunexpected\.

SCORINGPROTOCOL–SURPRISE

\(5\)Verysurprising:responsesinthiscategoryviolateyourexpectationsofhowtheobjectismeanttobeusedorhowitisoftenusedinpractice\.Moreover,theuseshouldnotbepointlessorimpossible\.Responsesinthiscategoryshouldinciteinterest\(andinspiration\)\.

\-Astrong’wow’response

\-Violatesexpectationsinapositiveway

\-Incitesinterest\(pictureitandthinkofuses\)

Examples:useatincantomakeabeercanchicken;tincanwallpaper;usebookcovertomakeabag/wallet;usebeltasanovenmitttograbhotpothandles\.

\(4\)Surprising:responsesinthiscategoryviolateyourexpectationsofhowtheobjectismeanttobeusedorhowitisoftenusedinpractice,andtheuseisnotpointlessorimpossible;buttheviolationsarelessdrasticandlesssurprising\(smaller’wow’factor\)thancategory5\.

\-Potentialfora’wow’response

\-Violatesexpectationsinapositiveway

Examples:useatincanasareflectorforyourbike;usebeltasyogamatstrap;makeachairseatoutofbelts;hangupabookbookshelf\.

\(3\)Somewhatsurprising:responsesinthiscategoryviolateyourexpectationsoftheobject’sconventionaluses,butcanberelatedtolessobvioususesthatareseenmoreoften\.Theuseshouldbenon\-obvious,butdoesnotelicita’wow’response\.Responsesthataresurprisingbutvague–hardtounderstandwhysomeonewoulddoit,butnotcompletelypointlessoruseless–alsobelongtothiscategory\.

\-Non\-obvious;violatestheconventionaluses;somewhatgenericuses

\-ORsurprisingbutvague

Examples:useatincanasaflowerpot;useforkasapaintbrush;usebookpagesaswrappingpaper;useforkashookonthewall;sendforkinthemailasamessage\.

\(2\)Hardlysurprising:responsesinthiscategoryviolateyourexpectationsoftheobject’smainuse,butaresomewhatobvious\.Manypeoplewillhaveusedtheobjectlikethisthemselvesoritissomethingoftenseeninthemedia\(orelsewhere\)\.Surprisingbutpointlessresponsesbelongtothiscategory,aswellasusesas’art’or’decor’withoutanyelaboration\.

\-Somewhatobvious\(probablyseen/havedonethis\);violatesthemainuse

\-ORsurprisingbutpointless

\-ORart/decorwithoutelaboration\(cangetahigherscorewithspecificity\)

Examples:useaforkasahairbrush;useforkasdrumstick;tincantelephone;pressflowerswithbook;cutbeltintopieces\.

\(1\)Notsurprisingatall:responsesinthiscategorydonotviolateyourexpectationsoftheobject’susesinanyway\.Usesthataregenericandcanbeappliedtoallobjectsalsobelongtothiscategory\.

\-Obvious\(firstideasthatcometomindwhenassociating\)

\-Conventionaluses;conventionaltoallobjects

Examples:useforktoeat;usebricktoweighsomethingdown;usetincanaspenholder;throwobject;giveobjectaway\.

YouareevaluatingAlternativeUsesTest\(AUT\)responsesforobject:\{object\_name\}\.

Beloware\{N\}alternativeuses,eachlabeledwithanumericID\.

Theyweregeneratedunderdifferentpromptconditions\.Judgetheusesonlybytheirtext\.

Use\{id\}:\{use\_text\}

Use\{id\}:\{use\_text\}

…

Rankallthealternativeusesabovebycreativity,theleastcreativetothemostcreative\.Lesscreativemeansclosertocommonuseandunfeasible/imaginary,morecreativemeansclosertounexpectedusesandalsofeasible/practical\.Assignascoreintegernumberfrom1\(leastcreativeuse\)to5\(mostcreativeuse\),andoutputtheresultsinascendingorderaccordingtothescore\.

ReturnONLYvalidJSON:

\{”ranked\_uses”:\[\{”id”:12,”score”:1\},\{”id”:7,”score”:2\},…\]\}

IncludeeveryuseIDexactlyonce,orderedfromleastcreativetomostcreative,eachwithitsintegerscorefrom1to5\(scoresmustbeinascendingorderalongthelist\)\.

### G\.6Infinity\-Chat 100 Benchmark\([Jiang et al\., 2026](https://arxiv.org/html/2609.30492#bib.bib26)\)

#### G\.6\.1Query Prompts

These are five manually selected queries we tested from Infinity\-Chat 100 benchmark\.

Generateajokeaboutelectricvehicles\.

Createthefirstverseofaweddingvow\.

DescribeAppleCorporationinthreesentencestoapersonwhohasnoideawhatcellphonesare\.

Givemeatipformanagingateamofcoworkers\.

Tellmeaboutcatsinthreewords\.

#### G\.6\.2Evolutionary TextGrad Algorithm Prompts

Youarepartofanadvancedoptimizationsystem\.Yourgoalistoevaluateandcritiquea”persona”thatguidesalanguagemodelduringacreativetask:answeringreal\-worldopen\-endeduserrequests\(jokes,vows,explanations,advice,descriptions\)thatadmitmanyequallylegitimateanswers\.Theguidedmodelanswerseachrequestseveraltimes,andthegoalisforthoseanswerstobegenuinelyDIFFERENTfromoneanotherratherthanrestatementsofoneidea\.Youarethegradient\(feedback\)engine\.

<OBJECTIVE\_FUNCTION\>

Yourobjectiveistomaximizethefitnessofthepersonabasedonfourmetricscomputedontheresponsesthepersonagenerates:

1\.Validity:Everyresponsemustbeacoherent,on\-topicattempttoanswertherequest—norefusals,nooff\-topictext,nononsense\(passorfail\)\.

2\.Anti\_homogeneity:Howsemanticallydifferentthepersona’sseparateanswerstotheSAMErequestarefromoneanother\(highermeanseachattemptexpressesagenuinelydifferentideainsteadofparaphrasingonefavouriteanswer\)\.

3\.Flexibility:Theeffectivenumberofdistinctanswersthepersonagivesperrequest\(highermeansitexploresseveralunrelateddirectionsratherthanclusteringononeortwo\)\.

4\.Lexical\_distinctness:Howlittleverbatimwordingthepersona’sanswerstoonerequestsharewitheachother\(highermeansnorecycledphrases,openings,ortemplatesentencesacrossattempts\)\.

</OBJECTIVE\_FUNCTION\>

Tohelpyouunderstandthecurrentpopulationlandscape,hereareexamplesfromthecurrentparallelbatch:

<BATCH\_CONTEXT\>

\{batch\_context\}

</BATCH\_CONTEXT\>

Weareinterestedingivingfeedbacktothefollowingpersona:

<VARIABLE\>

\{parent\_text\}

</VARIABLE\>

Herearetheresponsesthispersonagenerated,onwhichitsscoreswerecomputed:

<GENERATED\_RESPONSES\>

\{generated\_responses\_block\}

</GENERATED\_RESPONSES\>

Scoresforthisvariable:

Validity:\{pass\|fail\}

Anti\_homogeneity:\{float:\.4f\}

Flexibility:\{float:\.4f\}

Lexical\_distinctness:\{float:\.4f\}

Provideaconcise,specificcriticismdetailinghowtomodifythispersonatoimproveitsoverallfitness\.

\-IfValidityisfail,thepersona’sresponsesdidnotanswertherequests\.Addressthespecificfailure:

\*Ifresponsesdriftedoff\-topicorintoself\-description,instructtheoptimizertomakethepersonaalwaysdeliveradirectanswertotherequest,whateveritsvoice\.

\*Ifresponseswererefusalsoremptydeflections,instructtheoptimizertoremovewhatevertraitmakesthepersonadeclineopen\-endedrequests\.

\-IfAnti\_homogeneityislow,thepersonakeepsgivingthesameanswerindifferentwords\.Suggesttraitsthatmakeitapproacheachrequestfromawhollydifferentangleeverytime—differentdomains,moods,framings,andviewpoints—ratherthanorbitingoneidea\.

\-IfFlexibilityislow,thepersona’sattemptscollapseintooneortwoclusters\.Suggestbroadeningthepersona’sinterests,expertise,andlifeexperiencesoitsattemptslandingenuinelyseparateterritories\.

\-IfLexical\_distinctnessislow,thepersonarecycleswordingacrossattempts\.Suggestremovingwhateververbalhabit,catchphrase,ortemplatecausestherepeatedphrases,referencingtheshownresponsesasevidenceofwhatiscurrentlyrecycled\.

Donotgenerateanewpersona\.Youronlyjobistoprovidetextualcriticismandactionablefeedbackonhowtoalterthecurrentpersona\.

#### G\.6\.3Evaluation Prompts

Youareatrainedraterinastudyscreeningresponsestoopen\-endeduserrequests\.Eachrequestadmitsmanydifferent,equallylegitimateanswerswithnosinglegroundtruth\.YouscreenONEresponseatatime\.YoudonotknoworcarewhetheraresponsewaswrittenbyahumanorbyanAIsystem;judgeonlythetext\.

Decidewhethertheresponseisacoherent,interpretableattempttoaddresstherequest\.

Answer”invalid”iftheresponsefailstoaddresstherequestatall:anemptyornonsensestring,arefusaltoanswer,textonanunrelatedtopic,amererestatementoftherequestwithoutananswer,aformattingartifact,ortextsogarbledortruncatedthatnoanswercanberecoveredfromit\.

IMPORTANT—threethingsthatareNOTgroundsfor”invalid”:

1\.Qualityisnotvalidity\.Abland,generic,clumsy,orunoriginalresponsethatdoesaddresstherequestisVALID—itmerelyscoreslowonquality\.

2\.Voiceisnotvalidity\.AresponsewritteninastrongpersonaorcharactervoiceisVALIDaslongasananswertotherequestisstillidentifiableinit\.Judgeonlywhetherananswersurvivesthestyling\.

3\.Brevityisnotvalidity\.Iftherequestasksforsomethingshort\(afewwords,onesentence,atitle\),acorrespondinglyshortresponseisexactlyright\.Judgelengthonlyagainstwhattherequestasksfor\.

Ifyouarenotconfidenteitherway,answer”not\_sure”\.Awrong”invalid”deletesrealdata,so”not\_sure”isalwayssaferthanguessing\.

Request:\{query\}

Response:\{response\}

RespondwithONLYaJSONobjectoftheform\{”label”:”valid”\},\{”label”:”invalid”\},or\{”label”:”not\_sure”\}\.Noothertext\.

Youareatrainedraterinastudyevaluatingresponsestoopen\-endeduserrequests\.Eachrequestadmitsmanydifferent,equallylegitimateanswerswithnosinglegroundtruth—donotpenalizearesponsefortakingadifferentanglethanyouwouldhave\.YourateONEresponseatatime\.YoudonotknoworcarewhetheraresponsewaswrittenbyahumanorbyanAIsystem;judgeonlythetext\.

Ratetheoverallqualityoftheresponseasananswertotherequest,ona1\-5scale:

5—excellent:fullyaddressestherequest;clear,apt,andwellexecuted\.

4—good:addressestherequestwellwithonlyminorweaknesses\.

3—acceptable:agenuineanswer,butmediocreinexecutionorfit\.

2—poor:barelyaddressestherequest,ordoessoinaconfusedorbadlyexecutedway\.

1—verypoor:failstoaddresstherequest,orisincoherent\.

Judgequalityofexecutionandfittotherequest—notwhichofmanyvalidanswerchoiceswastaken,andnotlengthforitsownsake\(iftherequestasksforsomethingshort,shortisright\)\.

Request:\{query\}

Response:\{response\}

RespondwithONLYaJSONobjectoftheform\{”score”:N\}whereNisanintegerfrom1to5\.Noothertext\.

### G\.7DAT Benchmark

#### G\.7\.1Baseline Prompts

Pleasewrite10nounsinEnglishthatareasirrelevantfromeachotheraspossible,inallmeaningsandusesofthewords\.Pleasenotethatthewordsyouwriteshouldhaveonlysingleword,onlynouns\(e\.g\.,things,objects,concepts\),andnopropernouns\(e\.g\.,nospecificpeopleorplaces\)\.Youranswer:

List10commonEnglishnounsthatareasunrelatedinmeaningaspossible\.Avoidanysharedtopicorcategory\.Outputonlythenouns,separatedbycommas\.

List10commonEnglishnounsthatareascloselyrelatedinmeaningaspossibleandclearlyfitintoasinglenarrowtopic\.Outputonlythenouns,separatedbycommas\.

#### G\.7\.2Evolutionary TextGrad Algorithm Prompts

Youarepartofanadvancedoptimizationsystem\.Yourgoalistoevaluateandcritiquea”persona”thatguidesalanguagemodelduringacreativetask:theDivergentAssociationTask,wheretheguidedmodelmustname10Englishnounsthatareassemanticallyunrelatedtooneanotheraspossible\.Youarethegradient\(feedback\)engine\.

<OBJECTIVE\_FUNCTION\>

Yourobjectiveistomaximizethefitnessofthepersonabasedonfivemetricscomputedonthenounliststhepersonagenerates:

1\.Validity:Everyresponsemustgiveatleastsevenusablewords—single,commonEnglishnouns,nopropernouns,norepeats,noinventedwords\(passorfail\)\.

2\.Dat\_score:Howsemanticallydistantthenamednounsarefromoneanotherwithinasingleresponse\(highermeansthenounscomefrommoreunrelatedregionsofmeaning\)\.

3\.Flexibility:Howdifferentthepersona’sseparateattemptsarefromoneanother\(highermeansitexploresgenuinelydifferentsetsofnounseachtimeinsteadofrepeatingonefavouritelist\)\.

4\.Fluency\_entropy:Howevenlythepersonaspreadsitswordchoicesacrossitswholevocabulary\(highermeansitdoesnotkeepfallingbackonthesamehandfulofwords\)\.

5\.Fluency\_top10:Howlittleofthepersona’soutputistakenupbyitstenmostfrequentwords\(highermeanslessconcentrationonafewhabitualwords\)\.

</OBJECTIVE\_FUNCTION\>

Tohelpyouunderstandthecurrentpopulationlandscape,hereareexamplesfromthecurrentparallelbatch:

<BATCH\_CONTEXT\>

\{batch\_context\}

</BATCH\_CONTEXT\>

Weareinterestedingivingfeedbacktothefollowingpersona:

<VARIABLE\>

\{parent\_text\}

</VARIABLE\>

Herearethenounliststhispersonagenerated,onwhichitsscoreswerecomputed:

<GENERATED\_WORDS\>

\{generated\_words\_block\}

</GENERATED\_WORDS\>

Scoresforthisvariable:

Validity:\{pass\|fail\}

Dat\_score:\{float\}

Flexibility:\{float\}

Fluency\_entropy:\{float\}

Fluency\_top10:\{float\}

Provideaconcise,specificcriticismdetailinghowtomodifythispersonatoimproveitsoverallfitness\.

\-IfValidityisfail,thepersona’sresponsesdidnotyieldsevenusablewords\.Addressthespecificfailure:

\*IftheresponsewasnotaplainlistofsingleEnglishnouns,instructtheoptimizertomakethepersonadisciplinedandstrictaboutansweringintherequestedformat,withnocommentary\.

\*Ifthewordswerepropernouns,inventedwords,ornotnounsatall,instructtheoptimizertomakethepersonanameordinary,concreteEnglishnounsthatanydictionarywouldlist\.

\*Ifwordswererepeated,instructtheoptimizertomakethepersonatrackwhatithasalreadysaidandneverrepeatawordwithinoneanswer\.

\-IfDat\_scoreislow,thepersona’snounsweretoocloselyrelatedtooneanother\.Suggesttraitsthatmakeitleapbetweenwhollyunconnecteddomains—differentsenses,scales,andareasoflife—ratherthanlistingthingsfromonesceneortopic\.

\-IfFlexibilityislow,thepersonagivesnear\-identicalanswerseverytime\.Suggesttraitsthatmakeitapproachthetaskfromadifferentangleoneachattemptinsteadofrecitingonerehearsedlist\.

\-IfFluency\_entropyislow,thepersonakeepsreusingasmallvocabulary\.Suggestbroadeningthepersona’sinterests,expertise,andlifeexperiencesoitcandrawwordsfrommanydomains\.

\-IfFluency\_top10islow,afewhabitualwordsdominatethepersona’soutput\.Suggestremovingwhateverfixationcausesthosewordstorecur,referencingthebatchcontextasabaselineforwhatiscurrentlyoverrepresented\.

Donotgenerateanewpersona\.Youronlyjobistoprovidetextualcriticismandactionablefeedbackonhowtoalterthecurrentpersona\.

相似文章