Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

arXiv cs.AI Papers

Summary

This paper investigates whether aggregating probability estimates from multiple LLMs exhibits a wisdom-of-crowds effect, finding that learned aggregators outperform individual models and that training cutoff contamination is a pervasive confound in such evaluations.

arXiv:2607.18269v1 Announce Type: new Abstract: The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively studied with human forecasters. Whether the same phenomenon emerges when the ``crowd'' consists of large language models (LLMs) is an open question with both theoretical and practical implications. We elicited probability estimates from 15 LLMs on 254 binary prediction market questions and evaluated classical and learned aggregation methods. Learned aggregators -- a multilayer perceptron and a logistic regression -- outperformed all individual models and classical methods. The logistic regression was found to match the neural network, suggesting that the benefit of learned aggregation derives from learning a linear combination of diverse model outputs rather than from nonlinear interactions. Symbolic regression applied to the neural network's learned mapping recovered a pure model-disagreement signal as the lowest-complexity useful formula on the Pareto frontier, further supporting this interpretation. Training cutoff contamination proved a pervasive confound: the apparent capability gap between frontier cloud models and smaller local models collapsed from 35.8% to 8.9% on a clean subset of questions resolving after all models' training cutoffs, and individual model rankings showed only moderate stability. Even when the prediction market is evaluated at each model's training cutoff, LLMs remained substantially less accurate, indicating a genuine gap in collective information aggregation. These findings suggest that LLM crowds can exhibit wisdom-of-crowds effects, but that contamination-free evaluation is essential for reliable assessment.
Original Article
View Cached Full Text

Cached at: 07/22/26, 08:21 AM

# Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles
Source: [https://arxiv.org/html/2607.18269](https://arxiv.org/html/2607.18269)
###### Abstract

The wisdom of crowds—the finding that aggregating judgments across individuals often outperforms the best individual—has been extensively studied with human forecasters\. Whether the same phenomenon emerges when the “crowd” consists of large language models \(LLMs\) is an open question with both theoretical and practical implications\. We elicited probability estimates from 15 LLMs on 254 binary prediction market questions and evaluated classical and learned aggregation methods\. Learned aggregators—a multilayer perceptron and a logistic regression—outperformed all individual models and classical methods\. The logistic regression was found to match the neural network, suggesting that the benefit of learned aggregation derives from learning a linear combination of diverse model outputs rather than from nonlinear interactions\. Symbolic regression applied to the neural network’s learned mapping recovered a pure model\-disagreement signal as the lowest\-complexity useful formula on the Pareto frontier, further supporting this interpretation\. Training cutoff contamination proved a pervasive confound: the apparent capability gap between frontier cloud models and smaller local models collapsed from 35\.8% to 8\.9% on a clean subset of questions resolving after all models’ training cutoffs, and individual model rankings showed only moderate stability\. Even when the prediction market is evaluated at each model’s training cutoff, LLMs remained substantially less accurate, indicating a genuine gap in collective information aggregation\. These findings suggest that LLM crowds can exhibit wisdom\-of\-crowds effects, but that contamination\-free evaluation is essential for reliable assessment\.

Keywords:crowd wisdom; collective intelligence; data contamination; ensemble methods; large language models; probability aggregation; symbolic regression

## 1Introduction

The wisdom of crowds is the idea that crowds can be smarter than their members\(Beckeret al\.,[2019](https://arxiv.org/html/2607.18269#bib.bib60)\), where this has been understood as meaning that the crowd’s aggregate opinion is more accurate \(in some sense; see below\) than that of its “average member” or even that it is more accurate than that of the crowd’s*best*members\. There is ample evidence for the former claim\(Galton,[1907](https://arxiv.org/html/2607.18269#bib.bib20); Surowiecki,[2004](https://arxiv.org/html/2607.18269#bib.bib52); Davis\-Stoberet al\.,[2014](https://arxiv.org/html/2607.18269#bib.bib10)\)and some for the latter\(Hong and Page,[2004](https://arxiv.org/html/2607.18269#bib.bib30); Page,[2007](https://arxiv.org/html/2607.18269#bib.bib42); Prelecet al\.,[2017](https://arxiv.org/html/2607.18269#bib.bib45); Douven,[2019](https://arxiv.org/html/2607.18269#bib.bib14); Hong and Page,[2025](https://arxiv.org/html/2607.18269#bib.bib31)\), and the mechanisms underlying this type of crowd wisdom are also well understood, at least in broad terms: individual forecasters make partially independent errors, and averaging cancels those errors while preserving shared signal\(Larrick and Soll,[2006](https://arxiv.org/html/2607.18269#bib.bib58)\)\. Building on this insight, recent work has identified a number of powerful aggregation strategies that specifically exploit systematic features of the error distribution\(Douvenet al\.,[2026](https://arxiv.org/html/2607.18269#bib.bib18)\)\.

With the current broad availability of large language models \(LLMs\), it becomes natural to ask whether LLMs can substitute for human forecasters in crowd wisdom paradigms\.111On the more general question of whether LLMs can serve as useful models of human cognition and even as “synthetic participants,” seeArgyleet al\.\([2023](https://arxiv.org/html/2607.18269#bib.bib1)\); Binz and Schulz \([2023](https://arxiv.org/html/2607.18269#bib.bib3)\); Hagendorffet al\.\([2023](https://arxiv.org/html/2607.18269#bib.bib24)\); Hashimotoet al\.\([2026](https://arxiv.org/html/2607.18269#bib.bib26)\)\.If so, a crowd of LLMs could provide a cheap and scalable alternative to human forecasting panels for a wide range of tasks\. Moreover, studying LLM crowds could further illuminate the mechanisms of crowd wisdom itself: if error\-pattern diversity drives aggregation benefits, then LLM crowds should also benefit from aggregation to the extent that their outputs are sufficiently diverse—even if the degree of diversity, and hence the magnitude of the benefit, may differ from human crowds, whose members bring a much greater variety of backgrounds and expertise\(Landemore,[2020](https://arxiv.org/html/2607.18269#bib.bib35)\)\. Crucially, if diversity is the key ingredient, then even simple learned aggregators that weight models by their historical agreement patterns may suffice—a prediction our results will confirm\.

The closest existing work to ours isSchoeneggeret al\.\([2024](https://arxiv.org/html/2607.18269#bib.bib54)\), which used a crowd of 12 LLMs to make probabilistic predictions on 31 binary questions from a real\-time forecasting tournament, finding that the LLM crowd—aggregated by simple median—was statistically indistinguishable from a human crowd of 925 forecasters\. This is an important proof of concept, but it leaves some important questions unanswered\. First, Schoenegger and colleagues used only a very simple aggregation method \(the median\) and did not ask whether more sophisticated aggregators could improve over the best individual model or over the median itself\. And second, they offered no account of the*mechanism*underlying any aggregation benefit—whether it reflects error\-pattern diversity, performance\-weighted combination, or something else\. It is also to be noted that, because they collected data in real time on future events, contamination was not an issue for their study; our study, which draws on resolved questions across a longer time window, faces a more serious contamination challenge that requires explicit treatment\(Carliniet al\.,[2021](https://arxiv.org/html/2607.18269#bib.bib94); Robertset al\.,[2024](https://arxiv.org/html/2607.18269#bib.bib96); Xuet al\.,[2024](https://arxiv.org/html/2607.18269#bib.bib93); Gaoet al\.,[2025](https://arxiv.org/html/2607.18269#bib.bib95)\)\.222Halawiet al\.\([2024](https://arxiv.org/html/2607.18269#bib.bib56)\)andZouet al\.\([2022](https://arxiv.org/html/2607.18269#bib.bib57)\)similarly evaluate individual LLMs or simple ensembles on forecasting benchmarks without addressing the aggregation question systematically\. The former paper shows that a retrieval\-augmented system can approach human crowd accuracy, but their system uses web search rather than crowd aggregation as the source of improvement\.

The present study aims to address these issues\. We elicited probability estimates from 15 LLMs—including both locally deployable small models and frontier cloud models—on 254 binary prediction market questions\. We evaluated four classical aggregation methods alongside a logistic regression baseline and a trained multilayer perceptron \(MLP\) aggregator, and applied symbolic regression to recover an interpretable account of the MLP’s learned strategy\. We also conducted a systematic contamination analysis, examining whether training cutoff overlap inflated apparent individual model performance and distorted model rankings\.

## 2Methods

### 2\.1Materials

We collected binary \(YES/NO\) prediction market questions from Manifold Markets \([https://manifold\.markets](https://manifold.markets/)\), a play\-money forecasting platform that hosts a wide range of questions on current events, science, technology, and politics\.333One might worry that the fact that Manifold is a play\-money platform could affect market efficiency relative to real\-money markets\. However,Wolfers and Zitzewitz \([2004](https://arxiv.org/html/2607.18269#bib.bib53)\)show that play\-money prediction markets produce well\-calibrated probabilities comparable to real\-money markets\. We explored other platforms \(Kalshi, Polymarket, Metaculus\) as well, but none yielded a sufficient number of questions meeting our selection criteria within the required resolution window\.Questions were selected to have resolved between June 2024 and March 2026 and to have attracted at least 75 unique traders, a threshold chosen to ensure that the question had received sufficient community attention\. After converting question titles to declarative statements,444So, for example, “Will the US government shut down on Monday?” became “The US government will shut down on Monday\.”a subset of questions was excluded on grounds of ambiguity, personal reference, or dependence on inside information not accessible to a language model \(e\.g\., “Resolves to the side with the most holders at market close”; “I smoke weed before January 1, 2026”; “Does cooking in a cast\-iron pan make steak better? \[Eloise’s Blind Taste Test\]”\)\. The final dataset comprised 208 statements, of which 107 resolved true \(YES\) and 101 resolved false \(NO\)\. The mean community probability at market close was 0\.53 \(S​D=0\.30SD=0\.30\)\.

Because models may have been trained on data that includes the outcomes of questions resolving before their training cutoff \(the most recent cutoff among our 15 models was August 2025; see Table[1](https://arxiv.org/html/2607.18269#S2.T1)\), we defined a*clean subset*consisting of questions resolving after 1 September 2025\. Of the 208 items in the main dataset, 48 resolve after this date\. To increase statistical power, we supplemented these with 46 additional Manifold Markets questions collected using the same selection criteria and also resolving after 1 September 2025, yielding a clean subset of 94 items \(44 YES, 50 NO\) used for robustness analyses, and bringing the total number of unique items across both datasets to 254\. The remaining 160 items from the initial 208\-item dataset \(resolving before September 2025\) serve as the training set for learned aggregators evaluated on the clean subset \(see Sect\.[3](https://arxiv.org/html/2607.18269#S3)\)\.

### 2\.2Language models

Table 1:Language models included in the study\. Training cutoffs are approximate and drawn from official documentation where available\. Model sizes for cloud\-hosted models are not publicly disclosed\.We elicited probability estimates from 15 language models spanning a range of architectures, sizes, and training regimes\. Five models were run locally via Ollama; three were accessed via the Anthropic API; three via the OpenAI API; three via the Google AI API; and one via the DeepSeek API\. See Table[1](https://arxiv.org/html/2607.18269#S2.T1)for further details\.555Exact model identifiers used in API calls are available in the elicitation script that is included in the Supplementary Materials\.Because we also wanted to analyze how model capability affects aggregation performance, we included both frontier cloud models and smaller locally deployable models\.

### 2\.3Elicitation procedure

Each model was presented with a standardized prompt asking it to estimate the probability that a given statement is true, to think briefly, and to conclude with a single number between 0 and 1 in the formatPROBABILITY: \[number\]; see Box[2\.3](https://arxiv.org/html/2607.18269#S2.SS3)for an illustration\. All models were queried via their respective APIs \(or via Ollama for local models\) withmax\_tokensset to 512 and temperature left at the provider default \(equivalent to 1\.0 in all cases\)\. Most models followed the prompt format reliably\.666Google Gemini models initially failed to produce a parseable probability on a substantial fraction of items \(Gemini Flash: 88%; Gemini Pro: 66%\) due to free\-form response formatting; a single follow\-up turn requesting the probability in the required format recovered most of these\. LLaMA 3\.1 8B initially refused politically sensitive questions on 22\.6% of items; a brief neutral system message framing the task as an academic calibration exercise recovered most refusals\. Mistral 7B produced responses without aPROBABILITY:tag on 46% of items but a valid decimal number could be extracted in all but one case via the fallback parser\.

Box 1: Example elicitation promptYou are being asked to assess the probability that the following statement is true\. The statement was evaluated as of 2025\-11\-14\. Use this date to interpret any relative time references\.STATEMENT: "The Doomsday Clock will be 60s \(or less\) to midnight by end of January 2026\."Instructions:
\- This is a binary claim \(either true or false\)\.
\- Give your probability that the statement is TRUE, as a number between 0 and 1\.
\- Base your answer only on your training knowledge\. Do not search the internet\.
\- Think briefly, then state your final answer on a new line in the exact format:
PROBABILITY: \[number\]
For example: PROBABILITY: 0\.73Your response:
### 2\.4Aggregation methods

As argued inDouvenet al\.\([2026](https://arxiv.org/html/2607.18269#bib.bib18)\), how much wisdom there is in a crowd depends, to a large extent, on how we aggregate the opinions of the individuals constituting the crowd\. Following these authors, we evaluated a number of classical aggregation methods as well as a trained neural network aggregator\. We also looked at a logistic regression baseline, and we applied symbolic regression to the neural network’s learned strategy\.

#### Classical aggregators

All four classical methods are instances of the generalff\-mean family\(Bullen,[2003](https://arxiv.org/html/2607.18269#bib.bib64)\), which maps a vector of probabilities\(p1,…,pN\)\(p\_\{1\},\\ldots,p\_\{N\}\)to a scalar aggregate via

Mf​\(p1,…,pN\)=f−1​\(1N​∑i=1Nf​\(pi\)\)\.M\_\{f\}\(p\_\{1\},\\ldots,p\_\{N\}\)\\\>\\\>=\\\>\\\>f^\{\-1\}\\Bigl\(\\tfrac\{1\}\{N\}\\textstyle\\sum\_\{i=1\}^\{N\}f\(p\_\{i\}\)\\Bigr\)\.Choosingffas the identity gives the*arithmetic mean*, arguably the most natural baseline; choosingffas the log function gives the*geometric mean*, renormalized to sum to one across outcomes; and choosingf​\(x\)=x−1f\(x\)=x^\{\-1\}gives the*harmonic mean*\. Finally, settingf=logitf=\\mathrm\{logit\}yields the*log\-odds mean*,

σ​\(1N​∑ilogit​\(pi\)\),\\sigma\\Bigl\(\\tfrac\{1\}\{N\}\\textstyle\\sum\_\{i\}\\mathrm\{logit\}\(p\_\{i\}\)\\Bigr\),whereσ\\sigmadenotes the logistic function\. The log\-odds mean has been advocated as the normative standard for probabilistic aggregation under the assumptions that individual judgments are conditionally independent and well\-calibrated\(Morris,[1983](https://arxiv.org/html/2607.18269#bib.bib87); Genest and Zidek,[1986](https://arxiv.org/html/2607.18269#bib.bib75); Douvenet al\.,[2026](https://arxiv.org/html/2607.18269#bib.bib18)\), and it corresponds to what is sometimes called*extremizing*the arithmetic mean \(i\.e\., pushing the aggregate away from 0\.5 relative to simple averaging\)\. We also include the*median*, which was the primary aggregator inSchoeneggeret al\.\([2024](https://arxiv.org/html/2607.18269#bib.bib54)\)and is more robust to outlier estimates than the arithmetic mean\.

#### Neural network aggregator \(MLP\)

We followedDouvenet al\.\([2026](https://arxiv.org/html/2607.18269#bib.bib18)\)in going beyond the closed\-form classical aggregators and training a multilayer perceptron \(MLP\) to learn an aggregation function directly from the data\. The network maps the 15\-dimensional vector of model probability estimates to a single aggregate probability via two hidden layers: input \(15 units\)→\\rightarrowhidden \(64 units, ReLU activation, dropoutp=0\.2p=0\.2\)→\\rightarrowhidden \(32 units, ReLU, dropoutp=0\.2p=0\.2\)→\\rightarrowsigmoid output\. The network was trained over 25 epochs using the Adam optimizer with learning rate5×10−35\\times 10^\{\-3\}, binary cross\-entropy loss, and batch size 16\. For the aggregation analysis on the clean subset \(Sect\.[2\.1](https://arxiv.org/html/2607.18269#S2.SS1)\), the MLP is trained on the 160 non\-clean items and evaluated on all 94 clean items \(temporal transfer\)\. The relatively simple architecture was chosen after preliminary experiments showed no benefit from larger networks, consistent with the limited training data available\.

#### Logistic regression baseline

To determine whether the MLP’s advantage over classical means requires nonlinear processing, we also trained an L2\-regularized logistic regression \(λ=0\.01\\lambda=0\.01\) on the same 15\-dimensional input vector\. Logistic regression is a linear learned aggregator: it can learn non\-uniform weightings of the models but cannot combine them nonlinearly\. Like the MLP, it was trained on the 160 non\-clean items and evaluated on all 94 clean items in temporal transfer mode, ensuring a fair comparison\. If logistic regression matches the MLP, that is evidence that the aggregation benefit is attributable to learning an appropriate linear combination of diverse model outputs, rather than to deep nonlinear feature interactions\.

### 2\.5Symbolic regression

To obtain a still better understanding of the MLP’s learned aggregation strategy, we took our cue fromCranmeret al\.\([2020](https://arxiv.org/html/2607.18269#bib.bib9)\)and applied symbolic regression to the trained network\. Symbolic regression is a machine\-learning technique that searches the space of mathematical expressions for a formula that best approximates a target function\(Koza,[1992](https://arxiv.org/html/2607.18269#bib.bib33); Schmidt and Lipson,[2009](https://arxiv.org/html/2607.18269#bib.bib46)\)\. Rather than fitting parameters within a fixed functional form, it simultaneously searches over the form and the parameters, representing candidate formulas as trees whose leaves are input variables or constants and whose internal nodes are mathematical operations \(addition, multiplication, logistic function, and so on\)\. A population of candidate formulas is evolved over many generations via selection, mutation, and crossover, gradually improving both accuracy and simplicity\. SeeCranmeret al\.\([2020](https://arxiv.org/html/2607.18269#bib.bib9)\)for details\.

When, as in our case, simplicity and accuracy trade off against each other, the search naturally produces a*Pareto frontier*—a set of formulas such that no formula achieves better accuracy without becoming more complex, and no formula is simpler without sacrificing accuracy\. For this analysis, we retrained the MLP on all complete rows of the full 208\-item dataset rather than on the 160 non\-clean items used for the temporal transfer evaluation\. This is appropriate because the goal here is not predictive performance but*interpretability*: training on the maximum available data gives the MLP a more stable learned mapping for SR to approximate, and the contamination status of individual items is irrelevant for that purpose\. We then used theSymbolicRegression\.jlpackage\(Cranmer,[2023](https://arxiv.org/html/2607.18269#bib.bib8)\)to search for a symbolic formula minimizing mean squared error against this MLP’s predictions on the same 208 items, with formula complexity penalized alongside predictive error\. Given that SR is a stochastic process, we conducted 50 independent SR runs and examined the Pareto frontier of each to assess which formulas appeared consistently in the runs\.

### 2\.6Evaluation

Our primary evaluation metric is the Brier score,

B​S=1N​∑i=1N\(pi−yi\)2,BS\\\>\\\>=\\\>\\\>\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\(p\_\{i\}\-y\_\{i\}\)^\{2\},where, for each statementii\(1⩽i⩽N1\\leqslant i\\leqslant N\),pi∈\[0,1\]p\_\{i\}\\in\[0,1\]is the probability assigned to it andyi∈\{0,1\}y\_\{i\}\\in\\\{0,1\\\}its truth value\. While the word “score” might suggest otherwise, BS is actually a*penalty*, meaning that lower is better\(Brier,[1950](https://arxiv.org/html/2607.18269#bib.bib6)\)\. As secondary metrics we report classification accuracy—the proportion of items on which the model’s prediction exceeds 0\.5 precisely if the statement resolved true—and the familiar area under the receiver operating characteristic curve \(AUC\), which measures the model’s ability to rank true statements above false ones regardless of the absolute probability values\. The three metrics provide complementary perspectives: Brier score rewards both calibration and resolution, accuracy rewards threshold\-level discrimination, and AUC rewards rank\-order discrimination\(Gneitinget al\.,[2007](https://arxiv.org/html/2607.18269#bib.bib23); Douven and Kriegeskorte,[2026](https://arxiv.org/html/2607.18269#bib.bib97)\)\.

### 2\.7Contamination analysis

As mentioned previously, in evaluating LLMs on resolved prediction market questions there is the concern of training cutoff contamination; specifically, that if a question’s outcome was publicly reported before a model’s training cutoff, the model may have seen that outcome during training and so, when prompted, may simply provide the “memorized” answer rather than reason about the prompt\(Denget al\.,[2024](https://arxiv.org/html/2607.18269#bib.bib98); Robertset al\.,[2024](https://arxiv.org/html/2607.18269#bib.bib96)\)\. This could inflate apparent individual model performance and distort comparisons between models with different cutoff dates; most notably, it could favor newer frontier models over older or smaller ones for reasons unrelated to reasoning ability\.

To assess the extent and consequences of contamination in our dataset, we conducted three analyses\. First, for each model we compared the probability extremity \(\|p−0\.5\|\|p\-0\.5\|\) on within\-cutoff items \(questions resolving before the model’s training cutoff\) versus outside\-cutoff items, using Welchtt\-tests; elevated extremity on within\-cutoff items would indicate overconfident retrieval of memorized outcomes\. Second, we compared classification accuracy within versus outside each model’s training window; a contamination effect should manifest as substantially higher accuracy on within\-cutoff items\. Finally, we computed Spearman rank correlations between model rankings on the full 208\-item dataset and on the 94\-item clean subset \(Sect\.[2\.1](https://arxiv.org/html/2607.18269#S2.SS1)\) to quantify how much contamination distorts the apparent ordering of models\.

### 2\.8Comparison with the Manifold crowd

Schoeneggeret al\.\([2024](https://arxiv.org/html/2607.18269#bib.bib54)\)compared their LLM crowd directly against a human forecasting panel, finding the two statistically indistinguishable\. Our design differs in an important respect: rather than a forecasting tournament, we use a continuously updated prediction market, which means a naive comparison against the final market probability would be unfair to the LLMs—the market incorporates information accumulated right up to resolution, while LLMs reason from a frozen training snapshot\. We therefore construct two baselines\. The*final*market probability is the closing price at resolution, and serves as an upper bound on human crowd performance\. The*cutoff\-matched*market probability is the last recorded price on or before midnight UTC of each model’s training cutoff date, providing a temporally fair comparison: both the LLM and the market baseline reflect the same information horizon\.

Cutoff\-matched probabilities were retrieved via the Manifold API’s bet history endpoint, which returns the full sequence of bets with timestamps\. For each market and each model, we took the market probability immediately after the last bet placed on or before the cutoff date \(i\.e\., the running consensus price at that point in time, reflecting the aggregate of all trading activity up to the cutoff\)\. When a market had received no bets before a model’s training cutoff—either because the market was created after that date or because it had not yet attracted any traders—we fell back to the market’s initial probability \(theinitialProbabilityfield from the market endpoint\), which on Manifold is set by the market creator at opening and defaults to 0\.5\. This fallback occurred for approximately 55 markets \(26% of the dataset\) for the four models with the earliest cutoffs \(Mistral 7B, LLaMA 3\.1 8B, Gemma 2 9B, and Phi\-4 14B, all with cutoffs in 2023 or June 2024\); for models with later cutoffs the fallback rate was 0 or very close to 0\. All 208 items were successfully covered for all 15 models, with no missing cutoff\-matched probabilities\.

## 3Results

For completeness and comparability with prior work, Table[2](https://arxiv.org/html/2607.18269#S3.T2)reports model performance for all 15 models on the full 208\-item dataset\. We see that Claude Sonnet 4\.6 achieved the lowest Brier score among individual models \(B​S=0\.273BS=0\.273\), followed by Claude Opus 4\.6 \(B​S=0\.284BS=0\.284\) and Gemini Flash \(B​S=0\.286BS=0\.286\)\. Local models showed substantially worse performance \(Brier scores 0\.362–0\.413\)\.

Naturally, our main current interest is not so much in the individual models but rather in aggregation performance\. For the aggregation analysis, we turn to the 94\-item clean subset, on which all methods are evaluated consistently\. The classical aggregators require no training and are evaluated directly on all 94 items\. The two learned aggregators—logistic regression and MLP—are both trained on the 160 non\-clean items and evaluated on all 94 clean items \(temporal transfer\), so they are applied to questions they have never seen\. This design mirrors real\-world deployment—an aggregator fitted on resolved historical questions is applied to new ones—and ensures that all methods are compared on identical ground\.

Table 2:Individual model performance on the initial 208\-item dataset\. Models per category \(cloud/local\) sorted by Brier score\.ModelGroupBrierAccuracyAUCClaude Sonnet 4\.6Cloud0\.2730\.6350\.677Claude Opus 4\.6Cloud0\.2840\.6150\.668Gemini FlashCloud0\.2860\.6300\.640Gemini ProCloud0\.2920\.6350\.693Claude Haiku 4\.5Cloud0\.3110\.5240\.532GPT\-5\.4 MiniCloud0\.3280\.5870\.632GPT\-5\.2Cloud0\.3380\.5480\.534GPT\-5\.4Cloud0\.3470\.5530\.538DeepSeek\-chatCloud0\.3610\.5240\.515Gemini Flash\-LiteCloud0\.3770\.5800\.584Gemma 2 9BLocal0\.3620\.5020\.450Phi\-4 14BLocal0\.3810\.4470\.391Qwen 2\.5 7BLocal0\.3890\.4180\.431LLaMA 3\.1 8BLocal0\.3940\.4270\.430Mistral 7BLocal0\.4130\.4690\.452![Refer to caption](https://arxiv.org/html/2607.18269v1/x1.png)Figure 1:Brier scores on the 94\-item clean subset for the 15 models and the aggregators \(sorted by Brier score\)\. Both learned aggregator scores are from temporal transfer \(trained on 160 non\-clean items, tested on all 94 clean items\), making them comparable to the all\-data estimates for classical aggregators\. Error brackets show 95% bootstrap confidence intervals for individual models and classical aggregators\. No brackets are drawn for the learned aggregators since their temporal transfer evaluation design is not comparable to the all\-data bootstrap used for other methods; their statistical advantage is instead assessed via the sign and permutation tests reported in the text\.The two learned aggregators showed complementary strengths: logistic regression achieved a lower Brier score \(B​S=0\.241BS=0\.241vs\.B​S=0\.264BS=0\.264for the MLP\), while the MLP showed better rank\-order discrimination \(A​U​C=0\.657AUC=0\.657vs0\.6330\.633\) and threshold\-level accuracy \(0\.6060\.606vs\.0\.5740\.574\)\. Both substantially outperformed the arithmetic mean on all three metrics \(Brier:0\.3130\.313; accuracy:0\.5110\.511; AUC:0\.4980\.498\)\. The near\-equivalence of the two learned aggregators across metrics is consistent with the aggregation benefit being largely captured by a linear combination of model outputs, with any residual nonlinearity in the MLP improving discrimination at the cost of calibration\. The MLP’s advantage over the arithmetic mean was statistically significant across items \(sign test:62/9462/94,p=\.002p=\.002; Wilcoxon:z=−2\.48z=\-2\.48,p=\.013p=\.013; permutation test:p=\.049p=\.049\)\. The logistic regression’s larger mean Brier reduction \(23% vs\. 15\.6% for the MLP\) was not item\-consistently significant \(sign test:46/9446/94,p=\.837p=\.837; Wilcoxon:z=−1\.55z=\-1\.55,p=\.122p=\.122\), though the permutation test, which is sensitive to the mean reduction rather than item\-level consistency,*was*significant \(p=\.002p=\.002\)\. This pattern suggests that logistic regression’s advantage is concentrated on a smaller number of items on which it wins decisively, rather than distributed uniformly across the clean subset\. Figure[1](https://arxiv.org/html/2607.18269#S3.F1)shows Brier scores for all methods\.

Table 3:Crowd\-split analysis: Brier scores for local \(5 models\) and cloud \(10 models\) subsets under each aggregation method \(94\-item clean subset\)\. Diff is the percentage by which the local crowd differs from the cloud crowd\.To examine whether the capability gap between cloud and local models persists under aggregation, we computed Brier scores separately for the cloud\-only crowd \(10 models\) and the local\-only crowd \(5 models\) under each classical aggregator; see Table[3](https://arxiv.org/html/2607.18269#S3.T3)\. Cloud models outperformed local models by 8\.9% under arithmetic mean aggregation, a gap that largely disappears for the other classical methods\. The median—the aggregator used bySchoeneggeret al\.\([2024](https://arxiv.org/html/2607.18269#bib.bib54)\)—achieved a Brier score of 0\.343, outperforming geometric, log\-odds, and harmonic means but falling below the arithmetic mean \(0\.313\)\. This suggests that simple averaging captures more signal than the median in our setting, where model outputs vary continuously rather than clustering at extreme values\.

### 3\.1Training cutoff contamination

All models with sufficient within\-cutoff items showed substantially higher accuracy on those items than on out\-of\-window items \(Table[4](https://arxiv.org/html/2607.18269#S3.T4)\), with accuracy differentials ranging from\+0\.05\+0\.05\(GPT\-5\.4 Mini\) to\+0\.26\+0\.26\(Claude Sonnet\)\. Two models showed significantly elevated probability extremity on within\-cutoff items: Gemini Flash \(t=4\.25t=4\.25,p<\.001p<\.001\) and Gemini Pro \(t=2\.68t=2\.68,p=\.008p=\.008\)\.

Table 4:Calibration check: model accuracy on within\-cutoff items \(resolving before the model’s training cutoff\) versus outside\-cutoff items\. Models with zero within\-cutoff items in the dataset are omitted\. Extremitytt: Welchtt\-test on\|p−0\.5\|\|p\-0\.5\|, positive values indicate greater extremity on within\-cutoff items\.ModelCutoffninn\_\{\\text\{in\}\}Accin\{\}\_\{\\text\{in\}\}noutn\_\{\\text\{out\}\}Accout\{\}\_\{\\text\{out\}\}Δ\\DeltaAccQwen 2\.5 7BSep 2024270\.5561810\.398\+0\.158\+0\.158GPT\-5\.4 MiniSep 2024270\.6301810\.580\+0\.050\+0\.050GPT\-5\.2Sep 2024270\.7781810\.514\+0\.264\+0\.264GPT\-5\.4Sep 2024270\.7411810\.525\+0\.216\+0\.216DeepSeek\-chatNov 2024440\.5681640\.512\+0\.056\+0\.056Gemini Flash\-LiteJan 2025910\.6371160\.535\+0\.103\+0\.103Gemini FlashJan 2025920\.7281160\.552\+0\.177\+0\.177∗Gemini ProJan 2025910\.7471120\.545\+0\.203\+0\.203∗Claude Haiku 4\.5Jul 20251530\.575550\.382\+0\.193\+0\.193Claude Sonnet 4\.6Aug 20251560\.699520\.442\+0\.256\+0\.256Claude Opus 4\.6Aug 20251550\.652500\.500\+0\.152\+0\.152∗Significantly elevated extremity on within\-cutoff items \(p<\.05p<\.05\)\.The Spearman rank correlation between model rankings on the full 208\-item dataset and on the 94\-item clean subset wasρ=0\.532\\rho=0\.532, indicating that contamination substantially reshapes the apparent performance ordering; see Figure[2](https://arxiv.org/html/2607.18269#S3.F2)\. To mention the most notable shifts: LLaMA 3\.1 8B rose from rank 14 to rank 6, GPT\-5\.4 Mini dropped from rank 6 to rank 13, and Claude Opus dropped from rank 2 to rank 9\. Another noteworthy finding was that the cloud–local performance gap collapsed from 35\.8% to 8\.9% on clean items\.777To rule out the possibility that the gap collapse reflects a change in question\-pool composition rather than the absence of contamination, we recomputed the cloud–local gap on the 48 items drawn directly from the original 208\-item dataset that resolve after 1 September 2025 \(i\.e\., excluding the 46 additionally collected items\)\. On this nested subset, the gap collapsed to 13\.4%, closely matching the 8\.9% observed on all 94 clean items and confirming that the effect is attributable to contamination rather than to differences in question type or difficulty\. Also, because training cutoff dates are approximate, we verified that the key contamination findings survive shifting all cutoff dates by±1\\pm 1month\. Under both perturbations, the cloud–local gap remained large on the full dataset \(≈36%\\approx 36\\%\) and collapsed similarly on clean items \(≈13%\\approx 13\\%\), mean accuracy differentials were stable \(meanΔ​acc≈0\.17\\Delta\\text\{acc\}\\approx 0\.17\), and Spearman rank correlations between full\-dataset and out\-of\-window rankings remained around 0\.5\.

![Refer to caption](https://arxiv.org/html/2607.18269v1/x2.png)Figure 2:Individual model Brier scores on the initial 208\-item dataset \(left\) and the 94\-item clean subset \(right\), sorted by Brier score within each panel\. Bars are colored by model family\. The dashed line marks the best individual model in each panel; the solid red line marks the MLP score from the respective run\. Notable rank changes between panels \(full→\\rightarrowclean\): LLaMA 3\.1 8B:14→614\\rightarrow 6; GPT\-5\.4 Mini:6→136\\rightarrow 13; Claude Opus:2→92\\rightarrow 9; GPT\-5\.2:7→37\\rightarrow 3; DeepSeek:9→59\\rightarrow 5\.
### 3\.2Symbolic regression

Table 5:Symbolic regression model selection frequencies across 50 independent runs \(ranked by selection frequency\)\. Brier rank refers to rank on the full 208\-item dataset\.Across 50 independent SR runs, 11 models were selected in more than 50% of runs \(Table[5](https://arxiv.org/html/2607.18269#S3.T5)\)\. Selection frequency was only weakly correlated with individual Brier score rank \(rs=0\.295r\_\{s\}=0\.295\)\. The models selected in 100% of runs included both the best\-performing cloud models \(Claude Sonnet, Claude Opus, Gemini Flash, Gemini Pro\) and three of the worst\-performing local models \(Qwen 2\.5 7B, LLaMA 3\.1 8B, Phi\-4 14B\), producing a U\-shaped selection pattern inconsistent with performance\-weighted averaging\. The lowest\-complexity useful formula on the Pareto frontier wasσ​\(pGemini Pro−pClaude Haiku\)\\sigma\(p\_\{\\text\{Gemini Pro\}\}\-p\_\{\\text\{Claude Haiku\}\}\), a contrastive signal between two specific models\.

To quantify how much of the learned aggregation benefit this minimal formula captures, we evaluatedσ​\(pGemini Pro−pClaude Haiku\)\\sigma\(p\_\{\\text\{Gemini Pro\}\}\-p\_\{\\text\{Claude Haiku\}\}\)directly as an aggregator\. On the full 208\-item dataset it achievedB​S=0\.231BS=0\.231, outperforming the arithmetic mean by 12\.6% and comparing favorably with the full 15\-model MLP\. On the 94\-item clean subset, it achievedB​S=0\.243BS=0\.243\. That a formula using only two models—one strong positive signal and one contrastive signal—can nearly replicate the performance of a aggregator learning from 15 models is a strong indication that the aggregation benefit comes from a small number of complementary error patterns rather than from complex nonlinear interactions among all 15 models\.

### 3\.3Logistic regression coefficient analysis

The symbolic regression results are suggestive but indirect—they show*which*models are selected, not*why*they got selected\. For a more direct test, we inspected the logistic regression weights, which are interpretable by construction, and checked whether they reflect diversity or performance\. For each model, we computed its mean pairwise Pearson correlation of squared errors with the other 14 models across the 94 clean items, as a measure of how redundant its error pattern is relative to the crowd\. We then computed Spearman correlations between absolute logistic regression weight and \(a\) diversity \(negative mean error correlation, so that higher values mean more independent errors\) and \(b\) individual Brier score rank\.

The results strongly favor the diversity interpretation: the Spearman correlation between absolute weight and diversity wasrs=\+0\.482r\_\{s\}=\+0\.482, while the correlation with individual performance rank wasrs=−0\.075r\_\{s\}=\-0\.075, indicating that the aggregator learns to upweight models whose errors are independent from the crowd rather than models that are individually accurate\. Notably, nine of the fifteen models receive negative weights, meaning the aggregator uses them*contrastively*, in that when such a model assigns high probability, the aggregate is pulled downward\. This is a further expression of error\-pattern exploitation: a model’s systematic biases can be informative even when its raw predictions are poor, provided those biases are sufficiently distinct from the rest of the crowd\.

### 3\.4Comparison with human crowd

The Manifold community probability provides two natural baselines\. The*final*market probability \(at market close\) reflects continuous updating right up to resolution; the*cutoff\-matched*market probability is the market price recorded at the time of each model’s training cutoff, controlling for the informational advantage of more recent data\.

On the full 208\-item dataset, the final market probability achievedB​S=0\.098BS=0\.098, compared to0\.2650\.265for the LLM arithmetic mean, which is a factor of 2\.71\.888The human comparison can use all 208 items because, while contamination inflates LLM scores on items resolving before a model’s training cutoff, the cutoff\-matched market probability is evaluated at that same point in time\. In other words, because both the LLM and the market baseline therefore reflect the same information horizon, any contamination advantage enjoyed by the LLM is offset by the fact that the market had already incorporated that information too\.This large gap partly reflects the continuous\-updating advantage: the final market price incorporates all information available before resolution, while LLMs reason from a frozen training snapshot\.

Table 6:LLM Brier scores versus Manifold market probability at each model’s training cutoff date \(all 208 items\)\. Final marketB​S=0\.098BS=0\.098\(same for all models\)\. Ratio = LLM BS / market\-at\-cutoff BS\.ModelCutoffLLM BSMarket@cutoffRatioClaude Sonnet 4\.6Aug 20250\.2730\.1322\.06×\\timesClaude Opus 4\.6Aug 20250\.2840\.1332\.14×\\timesClaude Haiku 4\.5Jul 20250\.3110\.1252\.48×\\timesGemini FlashJan 20250\.2860\.1541\.85×\\timesGemini ProJan 20250\.2920\.1422\.05×\\timesGemini Flash\-LiteJan 20250\.3770\.1432\.64×\\timesDeepSeek\-chatNov 20240\.3610\.1652\.19×\\timesGPT\-5\.4 MiniSep 20240\.3280\.2041\.61×\\timesGPT\-5\.2Sep 20240\.3380\.2041\.66×\\timesGPT\-5\.4Sep 20240\.3470\.2041\.71×\\timesQwen 2\.5 7BSep 20240\.3890\.2041\.91×\\timesGemma 2 9BJun 20240\.3620\.2221\.63×\\timesPhi\-4 14BJun 20240\.3810\.2211\.73×\\timesLLaMA 3\.1 8BDec 20230\.3940\.2291\.50×\\timesMistral 7BMar 20230\.4130\.2301\.80×\\timesArithmetic mean—0\.2650\.132†2\.01×\\times†Market at August 2025, the most recent shared cutoff\.To isolate the contribution of this updating advantage, we fetched the historical market price at each model’s training cutoff date from the Manifold API and computed the Brier score of each of those cutoff\-matched probabilities\. The results are shown in Table[6](https://arxiv.org/html/2607.18269#S3.T6)\. Even after this temporal matching, every individual LLM remained substantially less accurate than the market at the same point in time, with LLM Brier scores 1\.50–2\.64 times worse than the corresponding market probability \(mean ratio: 1\.93×\\times\)\. For the LLM arithmetic mean compared to the market at the most recent shared cutoff \(August 2025\), the ratio was 2\.01×\\times\. This demonstrates that the human crowd advantage reflects genuine information aggregation—the ability of many active traders to integrate dispersed knowledge and update on each other’s estimates—rather than merely having access to more recent data\.

These results differ fromSchoeneggeret al\.\([2024](https://arxiv.org/html/2607.18269#bib.bib54)\), who found their LLM crowd statistically indistinguishable from human forecasters\. We come back to this in Section[4\.3](https://arxiv.org/html/2607.18269#S4.SS3)\.

## 4Discussion

### 4\.1Main findings

We found evidence for crowd wisdom effects in LLM aggregation\. As inDouvenet al\.\([2026](https://arxiv.org/html/2607.18269#bib.bib18)\), the MLP aggregator outperformed all individual models and classical aggregators, and this advantage was replicated on a 94\-item subset of questions resolving outside every model’s training window; the same was true for the logistic regression aggregator\. Moreover, the mechanism appears to be the same as in human crowds: the aggregation benefit derives from exploiting error\-pattern diversity rather than from weighting models by individual performance, as supported by our symbolic regression analysis\.

That said, LLM crowd wisdom is not yet at the level of human crowd wisdom, at least as measured against a continuously updated prediction market\. Even when the market probability is evaluated at the same point in time as each model’s training cutoff—neutralizing the market’s informational advantage—the market outperforms every LLM by a factor of 1\.5–2\.6\. This gap does not reflect merely more recent information; it reflects the market’s ability to aggregate dispersed knowledge from many active traders who update on each other’s estimates in real time\. LLMs, reasoning from a frozen training snapshot, cannot replicate this dynamic\. This structural ceiling limits how accurate any LLM crowd can be, regardless of how well it is aggregated\.

Nevertheless, within that ceiling, the gains from learned aggregation are both meaningful and statistically reliable, and they represent the maximum extractable benefit from the information LLMs carry\. That a simple linear model suffices to achieve most of this gain—and that the mechanism mirrors what has been found with human crowds—is the core finding of the paper\.

### 4\.2Contamination and the cloud–local gap

Training cutoff contamination was a pervasive and consequential confound\. All models with sufficient within\-cutoff items showed meaningfully higher accuracy on those items, with differentials up to 0\.26\. The consequences for model rankings were substantial \(ρ=0\.532\\rho=0\.532\)\. Most strikingly, the cloud–local performance gap collapsed from 35\.8% to 8\.9% on clean items, suggesting that much of the apparent capability advantage of frontier models in probabilistic forecasting reflects knowledge of outcomes rather than superior probabilistic reasoning\.

We observed a qualitative difference between model families in how contamination was expressed\. Gemini models showed significantly elevated probability extremity on within\-cutoff items \(i\.e\., overconfident retrieval of known outcomes\)\. Anthropic models showed comparable accuracy advantages without elevated extremity, suggesting that calibration training discouraged overconfident expression even when outcomes were known\.

### 4\.3Why does the human crowd outperform the LLM crowd?

As noted, the cutoff\-matched comparison places the LLM arithmetic mean at roughly 2 times the market’s Brier score even after controlling for the market’s continuous\-updating advantage\. While, at first glance, this seems in tension withSchoeneggeret al\.\([2024](https://arxiv.org/html/2607.18269#bib.bib54)\), who found their LLM crowd statistically indistinguishable from a human crowd, the tension dissolves once the two baselines are compared directly\.

For note that Schoenegger and colleagues compared against a*forecasting tournament*aggregate: a set of human predictions made at a single point in time, by non\-expert volunteers, on questions selected for the tournament\. This is structurally very similar to our LLM elicitation: simultaneous, non\-interactive, from generalist predictors with no special stake in accuracy\. It is therefore not very surprising that LLMs can match this kind of human performance\. Our baseline is fundamentally different: an actively maintained*prediction market*, in which self\-selected traders with domain knowledge bet repeatedly, react to each other’s estimates and to new information, and collectively drive the market price toward a well\-calibrated consensus\. This is a much higher epistemic bar\.

In other words, our findings and those ofSchoeneggeret al\.\([2024](https://arxiv.org/html/2607.18269#bib.bib54)\)are not in conflict but rather answer different questions: where Schoenegger and colleagues show that LLM crowds can match a panel of simultaneous human judgments, we show that they fall substantially short of a dynamic human market that aggregates dispersed knowledge through repeated interaction\. The gap we observe in our study is not a failure of LLMs relative to humans in general, but a reflection of the difference between two very different human aggregation mechanisms\. Whether LLM crowds could be brought closer to prediction market performance \(perhaps through iterative deliberation protocols that allow models to observe and respond to each other’s estimates\) is an interesting question for future work\.

### 4\.4Limitations and future directions

The clean subset of 94 items is too small for stable ranking of individual models or for detailed analysis of subgroup differences\. Future work should aim for at least 200–300 clean items, which would also allow training learned aggregators entirely within the clean data\. Moreover, we used a single elicitation prompt without variation; different prompting strategies may affect both individual performance and inter\-model error correlations\. Also, we did not include very large open\-weight models \(⩾\\geqslant\\,70B\), which might behave differently from both the small local and frontier API models in our study\.

Another limitation concerns the learned aggregators’ training environment\. The 160 non\-clean items on which the MLP and logistic regression are trained are structured by training\-cutoff contamination: models perform systematically better on items whose outcomes fall within their training window, producing a different pattern of errors than on genuinely novel questions\. The learned weights are therefore fitted to a contaminated error landscape and then applied to a clean one, which is a form of distribution shift\. That the aggregation advantage survives this shift \(sign test:62/9462/94,p=\.002p=\.002\) is encouraging and suggests the learned strategy captures something general about the crowd’s error structure\. Nevertheless, whether the specific weights would remain stable across different training configurations \(e\.g\., if trained on a fully uncontaminated dataset of comparable size\) is an open question that future work with larger clean datasets could address\.

## Supplementary materials

## References

- L\. P\. Argyle, E\. C\. Busby, N\. Fulda, J\. R\. Gubler, C\. Rytting, and D\. Wingate \(2023\)Out of one, many: using language models to simulate human samples\.Political Analysis31,pp\. 337–351\.Cited by:[footnote 1](https://arxiv.org/html/2607.18269#footnote1)\.
- J\. Becker, E\. Porter, and D\. Centola \(2019\)The wisdom of partisan crowds\.Proceedings of the National Academy of Sciences116,pp\. 10717–10722\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- J\. Bezanson, A\. Edelman, S\. Karpinski, and V\. B\. Shah \(2017\)Julia: a fresh approach to numerical computing\.SIAM Review59,pp\. 65–98\.Cited by:[Supplementary materials](https://arxiv.org/html/2607.18269#Sx1.p1.1)\.
- M\. Binz and E\. Schulz \(2023\)Using cognitive psychology to understand gpt\-3\.Proceedings of the National Academy of Sciences120,pp\. e2218523120\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2218523120),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.2218523120)Cited by:[footnote 1](https://arxiv.org/html/2607.18269#footnote1)\.
- G\. W\. Brier \(1950\)Verification of forecasts expressed in terms of probability\.Monthly Weather Review78,pp\. 1–3\.Cited by:[§2\.6](https://arxiv.org/html/2607.18269#S2.SS6.p1.4)\.
- P\. S\. Bullen \(2003\)Handbook of means and their inequalities\.2nd edition,Kluwer,Amsterdam\.Cited by:[§2\.4](https://arxiv.org/html/2607.18269#S2.SS4.SSSx1.p1.2)\.
- N\. Carlini, F\. Tramèr, E\. Wallace, M\. Jagielski, A\. Herbert\-Voss, K\. Lee, A\. Roberts, T\. Brown, D\. Song, Ú\. Erlingsson, A\. Oprea, and C\. Raffel \(2021\)Extracting training data from large language models\.In30th USENIX Security Symposium \(USENIX Security 21\),pp\. 2633–2650\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p3.1)\.
- M\. Cranmer, A\. Sanchez\-Gonzalez, P\. Battaglia, R\. Xu, K\. Cranmer, D\. Spergel, and S\. Ho \(2020\)Discovering symbolic models from deep learning with inductive biases\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 17429–17442\.Cited by:[§2\.5](https://arxiv.org/html/2607.18269#S2.SS5.p1.1)\.
- M\. Cranmer \(2023\)Interpretable machine learning for science with pysr and symbolicregression\.jl\.External Links:2305\.01582,[Link](https://arxiv.org/abs/2305.01582)Cited by:[§2\.5](https://arxiv.org/html/2607.18269#S2.SS5.p2.1)\.
- C\. P\. Davis\-Stober, D\. V\. Budescu, J\. Dana, and S\. B\. Broomell \(2014\)When is a crowd wise?\.Decision1,pp\. 79–101\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- C\. Deng, Y\. Zhao, X\. Tang, M\. Gerstein, and A\. Cohan \(2024\)Investigating data contamination in modern benchmarks for large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),Mexico City, Mexico,pp\. 8706–8719\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.482)Cited by:[§2\.7](https://arxiv.org/html/2607.18269#S2.SS7.p1.1)\.
- I\. Douven, N\. Kriegeskorte, and P\. Stinson \(2026\)Three and a half stages of crowd wisdom\.Collective Intelligence5\.External Links:[Link](https://doi.org/10.1177/26339137261435121)Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1),[§2\.4](https://arxiv.org/html/2607.18269#S2.SS4.SSSx1.p1.7),[§2\.4](https://arxiv.org/html/2607.18269#S2.SS4.SSSx2.p1.6),[§2\.4](https://arxiv.org/html/2607.18269#S2.SS4.p1.1),[§4\.1](https://arxiv.org/html/2607.18269#S4.SS1.p1.1)\.
- I\. Douven and N\. Kriegeskorte \(2026\)Predicting lockean from gradational accuracy\.International Journal of Approximate Reasoning192,pp\. 109636\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.ijar.2026.109636)Cited by:[§2\.6](https://arxiv.org/html/2607.18269#S2.SS6.p1.4)\.
- I\. Douven \(2019\)Optimizing group learning: an evolutionary computing approach\.Artificial Intelligence275,pp\. 235–251\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- F\. Galton \(1907\)Vox populi\.Nature75,pp\. 450–451\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- Z\. Gao, W\. Jiang, and Y\. Yan \(2025\)A test of lookahead bias in llm forecasts\.External Links:2512\.23847,[Link](https://arxiv.org/abs/2512.23847)Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p3.1)\.
- C\. Genest and J\. V\. Zidek \(1986\)Combining probability distributions: a critique and an annotated bibliography\.Statistical Science1,pp\. 114–135\.Cited by:[§2\.4](https://arxiv.org/html/2607.18269#S2.SS4.SSSx1.p1.7)\.
- T\. Gneiting, F\. Balabdaoui, and A\. E\. Raftery \(2007\)Probabilistic forecasts, calibration and sharpness\.Journal of the Royal Statistical Society B69,pp\. 243–268\.Cited by:[§2\.6](https://arxiv.org/html/2607.18269#S2.SS6.p1.4)\.
- T\. Hagendorff, S\. Fabi, and M\. Kosinski \(2023\)Human\-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt\.Nature Computational Science3,pp\. 833–838\.External Links:[Document](https://dx.doi.org/10.1038/s43588-023-00527-x)Cited by:[footnote 1](https://arxiv.org/html/2607.18269#footnote1)\.
- D\. Halawi, F\. Zhang, C\. Yueh\-Han, and J\. Steinhardt \(2024\)Approaching human\-level forecasting with language models\.InAdvances in Neural Information Processing Systems,Vol\.37\.Note:arXiv 2402\.18563Cited by:[footnote 2](https://arxiv.org/html/2607.18269#footnote2)\.
- R\. Hashimoto, T\. Takayanagi, M\. Suzuki, and K\. Izumi \(2026\)LLM agents reveal how human bias shapes path\-dependent market dynamics\.Journal of Computational Social Science9,pp\. 32\.External Links:[Document](https://dx.doi.org/10.1007/s42001-026-00465-4)Cited by:[footnote 1](https://arxiv.org/html/2607.18269#footnote1)\.
- L\. Hong and S\. E\. Page \(2004\)Groups of diverse problem solvers can outperform groups of high\-ability problem solvers\.Proceedings of the National Academy of Sciences101,pp\. 16385–16389\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- L\. Hong and S\. E\. Page \(2025\)The range of collective accuracy for binary classifications under majority rule\.Economic Theory79,pp\. 275–300\.External Links:[Document](https://dx.doi.org/10.1007/s00199-024-01570-z)Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- J\. R\. Koza \(1992\)Genetic programming: on the programming of computers by means of natural selection\.MIT Press,Cambridge, MA\.Cited by:[§2\.5](https://arxiv.org/html/2607.18269#S2.SS5.p1.1)\.
- H\. Landemore \(2020\)Open democracy: reinventing popular rule for the twenty\-first century\.Princeton University Press,Princeton, NJ\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p2.1)\.
- R\. P\. Larrick and J\. B\. Soll \(2006\)Intuitions about combining opinions: misappreciation of the averaging principle\.Management Science52,pp\. 111–127\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- P\. A\. Morris \(1983\)An axiomatic approach to expert resolution\.Management Science29,pp\. 24–32\.Cited by:[§2\.4](https://arxiv.org/html/2607.18269#S2.SS4.SSSx1.p1.7)\.
- S\. E\. Page \(2007\)The difference: how the power of diversity creates better groups, firms, schools, and societies\.Princeton University Press,Princeton, NJ\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- D\. Prelec, H\. S\. Seung, and J\. McCoy \(2017\)A solution to the single\-question crowd wisdom problem\.Nature541,pp\. 532–535\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- M\. Roberts, H\. Thakur, C\. Herlihy, C\. White, and S\. Dooley \(2024\)To the cutoff \. \. \. and beyond? a longitudinal perspective on LLM data contamination\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=m2NVG4Htxs)Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p3.1),[§2\.7](https://arxiv.org/html/2607.18269#S2.SS7.p1.1)\.
- M\. Schmidt and H\. Lipson \(2009\)Distilling free\-form natural laws from experimental data\.Science324,pp\. 81–85\.Cited by:[§2\.5](https://arxiv.org/html/2607.18269#S2.SS5.p1.1)\.
- P\. Schoenegger, I\. Tuminauskaite, P\. S\. Park, R\. V\. S\. Bastos, and P\. E\. Tetlock \(2024\)Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy\.Science Advances10,pp\. eadp1528\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.adp1528)Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p3.1),[§2\.4](https://arxiv.org/html/2607.18269#S2.SS4.SSSx1.p1.7),[§2\.8](https://arxiv.org/html/2607.18269#S2.SS8.p1.1),[§3\.4](https://arxiv.org/html/2607.18269#S3.SS4.p4.1),[§3](https://arxiv.org/html/2607.18269#S3.p4.1),[§4\.3](https://arxiv.org/html/2607.18269#S4.SS3.p1.1),[§4\.3](https://arxiv.org/html/2607.18269#S4.SS3.p3.1)\.
- J\. Surowiecki \(2004\)The wisdom of crowds: why the many are smarter than the few and how collective wisdom shapes business, economies, societies, and nations\.Random House,New York, NY\.Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p1.1)\.
- J\. Wolfers and E\. Zitzewitz \(2004\)Prediction markets\.Journal of Economic Perspectives18,pp\. 107–126\.Cited by:[footnote 3](https://arxiv.org/html/2607.18269#footnote3)\.
- C\. Xu, S\. Guan, D\. Greene, and M\. Kechadi \(2024\)Benchmark data contamination of large language models: a survey\.External Links:2406\.04244,[Link](https://arxiv.org/abs/2406.04244)Cited by:[§1](https://arxiv.org/html/2607.18269#S1.p3.1)\.
- A\. Zou, T\. Xiao, R\. Jia, J\. Kwon, M\. Mazeika, R\. Li, D\. Song, J\. Steinhardt, O\. Evans, and D\. Hendrycks \(2022\)Forecasting future world events with neural networks\.InAdvances in Neural Information Processing Systems,Vol\.35\.Note:NeurIPS 2022 Datasets and Benchmarks TrackCited by:[footnote 2](https://arxiv.org/html/2607.18269#footnote2)\.

Similar Articles

Can LLMs Take Retrieved Information with a Grain of Salt?

arXiv cs.CL

This paper investigates how large language models adapt to the certainty of retrieved information, identifying systematic limitations in handling uncertainty. It proposes an interaction strategy that reduces obedience errors by 25% without modifying model weights.