llm-ensembles

Tag

Cards List
#llm-ensembles

Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles

arXiv cs.CL · 2d ago Cached

This paper audits five diversity measures for LLM ensembles, finding that their associations with majority-vote gain are heavily entangled with model capability and are unstable after controlling for capability. The only robust signal is a modest residual pairwise co-failure association.

0 favorites 0 likes
#llm-ensembles

Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

arXiv cs.AI · 4d ago Cached

This paper investigates whether aggregating probability estimates from multiple LLMs exhibits a wisdom-of-crowds effect, finding that learned aggregators outperform individual models and that training cutoff contamination is a pervasive confound in such evaluations.

0 favorites 0 likes
← Back to home

Submit Feedback