Tag
This paper presents ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools for thyroid ultrasound, storing outputs as auditable evidence records. Developed on a large multicentre dataset, it achieves strong results in nodule segmentation, benign-malignant classification, and report generation.
RadFusion is a framework that adds threshold controllability to radiology report generation by fusing a multi-label classifier with a VQA-based generator and an LLM rewrite step, enabling sensitivity-specificity trade-offs and ROC-based validation. Experiments on MIMIC-CXR show improved diagnostic accuracy and clinically adaptable report behavior.
Introduces CliniCARE-Bench, a deployment-oriented benchmark for evaluating AI agents on clinical audit tasks over longitudinal EHR data, with 25 clinician-validated scenarios and 750 patient cases. It assesses verdict accuracy, evidence grounding, policy adherence, and calibrated abstention, finding that raw accuracy overstates investigation quality.
This paper presents a model-agnostic framework for per-modality failure analysis in multimodal clinical AI, distinguishing loud vs silent failures when a modality is dropped. Validated on planted ground truth and applied to EchoJEPA and HuBERT-ECG embeddings for LVEF prediction, it shows that dropping echo nearly doubles error.
Introduces a new problem domain for fine-grained analysis of children's gait behaviors from standard RGB video, along with a new dataset of over 1,100 high-frame-rate sequences and a unified framework, demonstrating that current SOTA methods and MLLMs fail on this clinical task.
MyoCardBench is a real-world benchmark for evaluating large language models in cardiovascular care, comprising 2,263 items across 13 tasks. GPT-5.4 achieved the highest overall score, demonstrating strengths in full-cycle care and multimodal interpretation.
ClinFusion is a vision-centric multimodal large language model for holistic medical understanding that unifies 2D and 3D medical image analysis using a cascaded vision encoder. It achieves state-of-the-art results on 20 out of 24 benchmarks and outperforms proprietary models like GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.
This paper extends missing-information stress-testing to open-ended medical conversation, finding that LLM judge choice materially changes apparent safety and that LLM judges are more permissive than clinicians.
LLM4EHR proposes a clinical foundation model that temporally aligns Electronic Health Record time series with medical event sequences using a domain-adapted large language model and a regularized contrastive objective, improving downstream prediction tasks.
GraphDx is a cost-aware, knowledge-enhanced multi-agent framework for sequential diagnosis that uses LLM-constructed medical knowledge graphs and three collaborative agents to improve diagnostic success rates and reduce test costs.
Introduces Safe-Psych, a sequential benchmark for evaluating how large language models handle diagnostic uncertainty in psychiatry, revealing that even strong models often fail to abstain or seek clarification when clinical evidence is incomplete.
SAGEAgent is an LLM-based clinical agent that sequentially decides which diagnostic modalities to acquire for cancer patients to balance predictive accuracy with clinical invasiveness, reducing acquisition burden by 55% while maintaining competitive survival prediction performance.
This survey examines recent progress in medical LLMs, presenting a dual-view approach that connects clinical practice with computational methods, and introduces a benchmark dataset for evaluating medical reasoning capabilities across 18 state-of-the-art models.
This paper introduces the Nimblemind Multi-Agent System (nMAS) for extracting evidence of H. pylori infection from gastric biopsy reports, achieving 98.61% accuracy across 216 feature-case decisions and demonstrating substantial time savings over manual review.
This paper proposes a reinforcement learning framework for evidence-seeking diagnostic reasoning using LLMs. The RL-trained 7B model outperforms larger models in multilingual clinical consultation tasks, showing that specialized RL can distill high-level clinical reasoning.
This paper examines the use of reinforcement learning from world feedback for clinical protocol-execution tasks in FHIR environments, identifies structural barriers like high silent-finish ceilings and zero-gradient tasks, and introduces MedAgentBench-v3 with a lower ceiling. It shows that pure RL underperforms rule-based SFT due to these barriers, and proposes a combined SFT+RL approach.
This paper introduces Manana, a non-parametric prompt-learning framework that teaches LLMs to recommend anti-seizure medications and defer uncertain cases in underrepresented epilepsy care settings, improving accuracy on Ugandan cohorts and enabling selective prediction with high precision.
This paper presents a blinded evaluation of clinical AI tools using real point-of-care queries from physicians, comparing specialized and general-purpose models across five dimensions. The specialized tool (OpenEvidence) outperformed general-purpose models on all axes, and the authors release the Real-POCQi benchmark.
This paper proposes the Clinical Harness, a runtime governance architecture for registering, orchestrating, guarding, and monitoring AI-enabled clinical capabilities, using osteoporosis as a demonstration case.
Introduces PhysAssistBench, a benchmark for evaluating LLMs in interactive doctor-patient-EHR assistance. Experiments show current models are unreliable in this setting, highlighting the need for coordinated capabilities.