Tag
This paper introduces a benchmark for evaluating LLM judges' ability to detect omissions in AI-generated clinical notes, finding that standard judges struggle with omissions but restructuring the task into per-fact verification improves detection.