Tag
The paper introduces diversity profiles as a curve-valued method to evaluate diversity in AI-generated content, addressing the limitations of ambiguous scalar metrics by providing a more transparent and resolution-aware framework.
This paper audits five diversity measures for LLM ensembles, finding that their associations with majority-vote gain are heavily entangled with model capability and are unstable after controlling for capability. The only robust signal is a modest residual pairwise co-failure association.