Tag
This paper proposes a human-grounded framework to measure distributional breadth (cultural reach) of LLM-generated content in open-ended tasks, introducing metrics LLM Coverage and In-Boundary Rate. Experiments show current LLMs produce plausible but narrow content concentrated near the center of human response space.
A critical reflection on whether using only three base models for millions of personal AI agents can produce genuinely diverse deliberation, arguing that correlated errors across models may create false unanimity and seeking operational metrics—drawn from ensemble learning—to measure true human representational diversity.
This paper introduces a domain-agnostic multi-layered framework for unsupervised extraction of perspectives to evaluate pluralism in LLM-generated text, finding that rare perspectives are disproportionately underrepresented.