Tag
This paper audits five diversity measures for LLM ensembles, finding that their associations with majority-vote gain are heavily entangled with model capability and are unstable after controlling for capability. The only robust signal is a modest residual pairwise co-failure association.
This paper proposes a multi-factor scoring system for evaluating LLM responses, integrating accuracy, conciseness, factual consistency, readability, and coherence. Applied to the TruthfulQA dataset, it reveals strengths and limitations of mainstream models, offering a transparent evaluation framework.