Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct
Summary
This arXiv preprint studies the semantic dispersion of sixteen language models forming ensembles, showing that ensemble diversity is small on average and that model identity only partially explains which model is most divergent. The authors propose a per-model dissent contribution metric and find that dispersion is organized by clinical content rather than interpretive openness.
View Cached Full Text
Cached at: 08/04/26, 07:41 AM
# Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct Source: [https://arxiv.org/abs/2608.00285](https://arxiv.org/abs/2608.00285) [View PDF](https://arxiv.org/pdf/2608.00285) > Abstract:Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1\.69 distinct formulations of a psychotherapeutic case, against a single\-model baseline of 1\.43 from one model's own runs\. Ensembles place more than one reading before a decision\-maker on the premise that several models supply several perspectives\. Dispersion over their outputs is measured both as diversity and as uncertainty, and both traditions validate it against a correctness criterion that this task does not admit\. Measuring diversity is a solved problem: the Vendi Score, the exponential of the von Neumann entropy of a similarity matrix, is an effective number of distinct elements\. What a single aggregate does not say is where the diversity comes from\. We define a per\-model dissent contribution, the complement of a model's mean similarity to the other members of its ensemble: a magnitude from the same matrix, not a decomposition of the spectral index, whose maximum identifies the most divergent voice\. Crossing model and case, we test as a preregistered hypothesis whether model identity accounts for a non\-zero share of the variance in dissent, and characterise the structure that test detects\. The panel formulated fifteen stratified vignettes, yielding 7,082 formulations for analysis\. Model identity was a detectable structuring factor of the dissent that remained, but the usual categories recovered it only partly: scale differences pointed in opposite directions across pairs, family grouped models on only five two\-member lines, and the most divergent voice changed with panel composition, so that the surfaced outlier describes the ensemble rather than the model\. Dissent did not track the interpretive openness for which the case bank was stratified; it was organised by clinical content instead, leaving the dispersion an ensemble produces a property to measure rather than assume\. ## Submission history From: Mario Vega\-Barbas \[[view email](https://arxiv.org/show-email/7e96ec84/2608.00285)\] **\[v1\]**Fri, 31 Jul 2026 20:42:53 UTC \(732 KB\)
Similar Articles
Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models
This paper tests whether cognitive diversity (personas, temperature, model identity) drives multi-agent debate gains in small open-weight models, rejecting the hypothesis across 5,500+ budget-matched runs and arguing reported MAD improvements are essentially an ensemble-sampling effect rather than a diversity benefit.
A million people, a million personal AIs, three base models. Is that a diverse deliberation — and how would you measure it?
A critical reflection on whether using only three base models for millions of personal AI agents can produce genuinely diverse deliberation, arguing that correlated errors across models may create false unanimity and seeking operational metrics—drawn from ensemble learning—to measure true human representational diversity.
Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs
This paper investigates whether stochastic sampling (self-consistency) in LLMs can capture cross-question structure similar to diverse ensembles. Using a Marchenko–Pastur test, the authors find that within a single model, stochastic variation yields at most one significant dimension, while an ensemble of 24 models yields four, revealing a dimensionality gap that limits self-consistency as an ensemble substitute.
Do Large Language Models Capture the Diversity in their Training Data?
The paper investigates the conditional diversity gap in large language models by comparing the entropy of generated outputs with training data and proposes an information-theoretic framework to measure and mitigate this gap.
Reach Into The CHOIR: Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles
The paper introduces CHOIR (Collective Hierarchically-Ordered Inquiry Responses), a framework adapting free-list elicitation from cognitive anthropology to probe and measure hidden diversity beneath surface agreement in LLM ensembles, finding that base-model identity is the strongest recoverable signature of distinct model voices.