Tag
CRAFT is a new LLM framework for iterative refinement of temporal reasoning over clinical narratives, introducing a verifier-based feedback mechanism and the MedTempo benchmark for vaccine adverse-event reports.
This paper presents the first rigorous study of how LLM watermarking schemes affect medical performance, evaluating five watermarks across multiple LLMs and VLMs on clinical reasoning tasks. The authors find that watermarks can cause degradation in medical text quality, including hallucinations and lexical corruption, which are masked by general-domain benchmarks.