Tag
This paper introduces a benchmark for evaluating LLM judges' ability to detect omissions in AI-generated clinical notes, finding that standard judges struggle with omissions but restructuring the task into per-fact verification improves detection.
This paper introduces EHR-ReasonCon, a reasoning-intensive benchmark for consistency verification between clinical notes and structured tables in electronic health records, and EHR-Inspector, an LLM-based framework that achieves state-of-the-art performance in detecting discrepancies.