I benchmarked 8 LLMs for medical scribing. Hallucinations were rare; omissions need attention.
Summary
A benchmark of 8 LLMs for medical scribing found hallucinations rare but omissions a concern.
Similar Articles
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
This paper investigates whether standard benchmarks underestimate LLM performance by re-evaluating hallucination detection datasets using an LLM-first, human-adjudicated assessment method. The study finds that incorporating LLM reasoning into the adjudication process improves agreement and suggests that model-assisted re-evaluation yields more reliable benchmarks for ambiguity-prone tasks.
One of the most interesting benchmarks and its implications for alignment
The article discusses how public benchmarking of hallucination in LLMs has driven rapid improvements and explores implications for alignment, emphasizing the need for benchmarks on transparency and honesty.
Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention
This survey paper presents a lifecycle-based framework for understanding hallucinations in LLMs, covering causes, detection, mitigation, and prevention across data, training, and inference stages.
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
This paper presents the first rigorous study of how LLM watermarking schemes affect medical performance, evaluating five watermarks across multiple LLMs and VLMs on clinical reasoning tasks. The authors find that watermarks can cause degradation in medical text quality, including hallucinations and lexical corruption, which are masked by general-domain benchmarks.
Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing
This paper proposes a method to detect hallucinations in LLMs by analyzing topological signatures in attention graphs, showing improvements over existing baselines across multiple benchmarks.