Tag
This paper introduces a two-level meta-rubric framework for evaluating factual completeness in open-ended generation, instantiated as the GAMUT benchmark. It features 1,813 questions across 10 domains and finds the benchmark challenging and discriminative, with top models scoring 58.7%.