Tag
This paper introduces Guideline-as-Oracle (GAO), a method for zero-annotation training of a multi-turn ophthalmic telephone triage agent by compiling American Academy of Ophthalmology guidance into a 70-row rule table used to generate 3,000 training dialogues. Fine-tuning a 9B model on this corpus improves agreement with an operational reference from 61.7% to 74.1% and emergent-case recall from 9.5% to 69.0%, beating several general-purpose systems without needing a frontier model at inference.
Presents RESPClinBench, a real-world scenario benchmark for respiratory clinical decision-making, evaluating seven LLMs on COPD and pulmonary nodule cases. Finds task-specific limitations including imaging hallucination and medication-safety risks.
A pulmonologist discusses how AI is poised to take over aspects of his medical job, highlighting the growing impact of AI in healthcare diagnostics and clinical practice.
ClinFusion is a new open medical multimodal LLM from Alibaba DAMO Academy, available in 8B and 32B sizes (Apache 2.0), with unified 2D + 3D image understanding and state-of-the-art results on medical benchmarks.
This paper introduces PatTree, a graph-based multimodal patient representation that automatically structures heterogeneous clinical data for medical classification tasks. It achieves state-of-the-art performance on the ADNI-1 cohort with 98.5% balanced accuracy for Alzheimer's disease classification.
Introduces OncoTriad-QA, a patient-level benchmark integrating radiology, pathology, genomics, and clinical data for pan-cancer reasoning, along with OncoVLM, a reference multimodal model that outperforms existing medical LLMs after fine-tuning.
TumorBoard is a multi-agent decision-support system for longitudinal neuro-oncology that uses a shared longitudinal case state and auditable claim-evidence ledger. It outperforms baselines on a 360-case benchmark, with a safety governor reducing harmful recommendations.
Introduces MedPIC-Bench, a benchmark with counterfactual questions to evaluate whether LLMs correctly apply medication-safety rules when patient-specific conditions change; across 28 LLMs, accuracy drops significantly on counterfactual questions, revealing a common failure to revise judgments.
A new MIT-led study in Nature Medicine finds that AI assistance and explainability methods impact skin disease diagnosis accuracy differently depending on user expertise: non-experts over-trust AI explanations, while clinicians perform best with only the model's prediction. The results highlight the need for user-centered AI design that accounts for automation bias.
TreeProbe is the first cultural-bias benchmark for Tibetan medicine in LLMs, containing 4,719 expert-adjudicated items across 467 diseases and 10 subtasks, revealing systematic external ontology drift in current models.
EarlyDx is a new large-scale benchmark for evaluating LLMs on open-ended, evidence-supported diagnosis generation at emergency department admission, built from 154,834 MIMIC-IV encounters. It reveals that even frontier and medical-specialized models struggle to synthesize admission-time evidence, with post-training only partially improving inference-dependent recall.
MyoCardBench is a real-world benchmark for evaluating large language models in cardiovascular care, comprising 2,263 items across 13 tasks. GPT-5.4 achieved the highest overall score, demonstrating strengths in full-cycle care and multimodal interpretation.
Proposes a Collaborative Meta Knowledge Enhancement (COME) framework for dementia etiology diagnosis that injects heterogeneity-aware embeddings into a unified Transformer architecture, achieving state-of-the-art performance across multiple independent cohorts.
This paper presents a retrieval-augmented, multi-agent LLM framework with human-in-the-loop for detecting cutaneous immune-related adverse events from clinical notes, achieving higher accuracy, improved inter-rater agreement, and halved review time compared to manual review.
This paper shows that Monte Carlo dropout provides epistemic uncertainty signals for chest radiograph classifiers, which improves error detection and reduces confident misdiagnoses in clinical decision-support agents when communicated as a binary error-risk flag.
OpenAI is rolling out ChatGPT Health to all US users, allowing them to connect medical records and health-tracking data, with claims of clinician-level reasoning and integration with GPT-5.6 Sol.
This paper extends missing-information stress-testing to open-ended medical conversation, finding that LLM judge choice materially changes apparent safety and that LLM judges are more permissive than clinicians.
This paper describes TalTech's systems for generating SOAP notes directly from doctor-patient conversation audio, using Voxtral models fine-tuned with supervised learning and DAPO reinforcement learning. Their submissions ranked first in both tracks of the BeTraC challenge, achieving high concept accuracy and low hallucination rates.
Tri-Net v2 is an open-source implementation of a Scientific Reports paper for unified skin lesion and symptom-based monkeypox detection.
Kimi K3 AI model successfully reads a chest X-ray after OpenMed removes all 23 patient identifiers from the DICOM data, ensuring privacy. The model correctly identifies a left-sided whiteout.