Tag
This paper introduces generative-process diversity as a measure for language models using Normalized Compression Distance, demonstrating that it predicts correlated failure across benchmarks better than semantic similarity, with implications for safety in multi-model systems.
This paper identifies a fundamental constraint on multi-model LLM systems: accuracy is capped by the rate at which all models fail on the same query. Across 67 frontier models, the all-wrong rate is significantly underestimated by common metrics, limiting gains from voting, routing, and ensemble strategies.