Tag
This paper audits five diversity measures for LLM ensembles, finding that their associations with majority-vote gain are heavily entangled with model capability and are unstable after controlling for capability. The only robust signal is a modest residual pairwise co-failure association.
This paper investigates whether aggregating probability estimates from multiple LLMs exhibits a wisdom-of-crowds effect, finding that learned aggregators outperform individual models and that training cutoff contamination is a pervasive confound in such evaluations.