Tag
This paper introduces a factuality-specific annotation policy using failure-space analysis to improve human-anchored factuality evaluation for LLMs under limited annotation budgets, achieving significant efficiency gains on benchmark systems.
This paper introduces TriQua, a framework for LLM factuality evaluation that adaptively represents facts as triples or hyperrelational facts with contextual qualifiers, along with TriQuaScore for fine-grained factuality scoring. It demonstrates strong alignment with human annotations and improved evidence-based verification over existing methods.