Tag
This paper introduces a benchmark for evaluating LLM judges' ability to detect omissions in AI-generated clinical notes, finding that standard judges struggle with omissions but restructuring the task into per-fact verification improves detection.
A look at the growing use of ambient AI scribes in doctor visits, the questions patients should ask, and the legal landscape around consent and HIPAA compliance.