I benchmarked 8 LLMs for medical scribing. Hallucinations were rare; omissions need attention.

Reddit r/LocalLLaMA News

Summary

A benchmark of 8 LLMs for medical scribing found hallucinations rare but omissions a concern.

No content available
Original Article

Similar Articles

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

arXiv cs.CL

This paper investigates whether standard benchmarks underestimate LLM performance by re-evaluating hallucination detection datasets using an LLM-first, human-adjudicated assessment method. The study finds that incorporating LLM reasoning into the adjudication process improves agreement and suggests that model-assisted re-evaluation yields more reliable benchmarks for ambiguity-prone tasks.

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

arXiv cs.AI

This paper presents the first rigorous study of how LLM watermarking schemes affect medical performance, evaluating five watermarks across multiple LLMs and VLMs on clinical reasoning tasks. The authors find that watermarks can cause degradation in medical text quality, including hallucinations and lexical corruption, which are masked by general-domain benchmarks.