Tag
Introduces CANOE, a multi-agent neuro-symbolic framework for open-ended care plan coordination that uses argumentative computation and human-in-the-loop contestation to improve transparency, safety, and clinical correctness.
This paper presents an evaluation of multi-turn multimodal diagnostic reasoning using challenging real-world clinical cases, aiming to assess AI models' ability to handle complex medical scenarios.
This paper presents PATHFinder Agent, an end-to-end conversational AI system that generates personalized prenatal care plans following ACOG's PATH guidelines, integrating patient intake, dynamic dialogue, plan synthesis, and clinician oversight. Evaluation of frontier LLMs, including GPT-5.2, shows promising but incomplete performance, highlighting gaps in antenatal testing recommendations.
MissHyper is a new hypergraph forecasting model that restores clinical synchronicity by aggregating co-timestamp records before message passing, achieving consistent gains on PhysioNet 2012, MIMIC-III, and MIMIC-IV benchmarks.
Presents a lightweight knowledge-injection framework for zero-shot ICU delirium prediction that augments structured EHR data summaries with external clinical knowledge at inference time, improving AUROC by up to 8.57 percentage points on LLaMA models without fine-tuning.
This paper develops a multi-dimensional framework to evaluate discrimination, calibration, interpretability, and algorithmic fairness for machine learning-based type 2 diabetes risk prediction models, revealing significant performance degradation under real-world distribution shifts and biases by age and obesity.
CardioMeta is a calibrated multi-task framework for jointly predicting diabetes, hypertension, and cardiovascular disease across NHANES and MIMIC-IV data, emphasizing leakage control, calibration, and transparent reliability.
A former Mayo Clinic research director alleges the hospital ignored high error rates in its AI tools (MAYA) and retaliated against her for blowing the whistle, raising serious concerns about AI deployment in healthcare.
Pythia is a multi-agent system that autonomously writes and optimizes extraction prompts for clinical concepts without manual prompt engineering or fine-tuning, using a locally hosted open-weights model. It achieves mean sensitivity of 0.76 and specificity of 0.95 on clinical symptom detection, outperforming lexicon-based methods on specificity.
This study proposes a feature-guided zero-shot framework using LLMs for early chronic kidney disease screening, achieving consistent improvements with minimal community-accessible features across heterogeneous datasets.
Proposes a graph-constrained traversal policy that reformulates ICD-10-CM code prediction as a finite-horizon decision process over a pruned code hierarchy, outperforming flat baselines on MIMIC-IV discharge summaries.
This paper proposes an explicit multimodal routing framework for clinical prediction using EHR data, enabling interpretable, robust, and auditable reasoning across structured variables, clinical notes, and chest X-rays via discrete unimodal, bimodal, and trimodal routes with inference-time route masking for missing modality simulation.
The author shares personal views from participating in training a vertical medical model at a top domestic hospital, highlighting core challenges such as medical data not leaving the hospital, high cost of on-premises deployment, and weak willingness to pay, and suggests partnering with hardware vendors. They also note that general medical models (like Baichuan) already perform well.
This survey examines recent progress in medical LLMs, presenting a dual-view approach that connects clinical practice with computational methods, and introduces a benchmark dataset for evaluating medical reasoning capabilities across 18 state-of-the-art models.
This paper proposes FedDualAtt, a personalized federated learning approach for ECG classification that splits transformer attention heads into globally aggregated and locally private branches to handle data heterogeneity across clinical sites. Experiments on the FedCVD benchmark show improved performance over existing methods.
This paper presents CISM, a channel-independent spectrogram framework that treats missingness as a predictive signal for clinical multivariate time series prediction. Experiments on MIMIC-IV show it outperforms baselines for in-hospital mortality prediction.
This paper presents an Attention-Based Multiple Instance Learning (ABMIL) framework that leverages patient-level labels from cancer registries to train deep learning classifiers for tumor group classification without requiring per-report annotations, achieving a macro F1 of 0.83 on tasks at the BC Cancer Registry.
Fertility companies are using AI to improve IVF success rates by better predicting pregnancy outcomes and screening embryos, though experts raise ethical and privacy concerns. A 2023 review found AI models can more accurately predict successful pregnancy than embryologists.
This paper introduces HealthAgentBench, a suite of 54 realistic healthcare tasks for evaluating frontier AI agents. It finds that even the best agent (Codex GPT-5.5) achieves only ~42% success, highlighting substantial room for improvement.
A benchmark of 8 LLMs for medical scribing found hallucinations rare but omissions a concern.