Tag
Introduces MedPIC-Bench, a benchmark with counterfactual questions to evaluate whether LLMs correctly apply medication-safety rules when patient-specific conditions change; across 28 LLMs, accuracy drops significantly on counterfactual questions, revealing a common failure to revise judgments.
As AI therapy chatbots grow in popularity, US states are enacting laws to regulate them over patient safety concerns, following lawsuits against platforms like Character.ai and rising use of ChatGPT/Claude for mental health advice.
This Perspective paper argues that large language models are not yet safe for autonomous clinical decision support, particularly in triage of undifferentiated patients, due to lack of robust evaluation under incomplete information and asymmetric costs of missed diagnoses.
This paper presents a formative study using a two-stage LLM pipeline (Gemini 2.5 Pro and Flash) to detect internal documentation inconsistencies in electronic health records, analyzing 3,000 discharge summaries and proposing a graded ontology for categorizing inconsistencies.
This paper presents a provenance-aware, knowledge-graph-based multi-agent framework that integrates patient narratives from Reddit and WebMD with FDA adverse event reports for nine antidepressants, using an LLM entity-recognition pipeline to achieve high accuracy and enabling traceable safety information for psychiatric medications.
A Penn State study found that AI chatbots like ChatGPT respond to everyday health queries with nearly 76% accuracy, raising concerns about trustworthiness in real-world healthcare applications. The research highlights that AI tools may be best used by physicians rather than patients.