Tag
Researchers propose a calibrated hybrid ensemble combining five deep learning and three classical ML models via a logistic regression stacking meta-learner for EEG-based epileptic seizure forecasting. Evaluated on CHB-MIT with Leave-One-Patient-Out cross-validation, the model reaches 74.2% seizure-level sensitivity at 1.24 false alarms per hour with ~16.9 minutes of warning time.
SMILESGNN introduces a multimodal architecture combining SMILES transformers and graph neural networks with cross-attention for interpretable drug toxicity prediction, achieving competitive performance on benchmarks like ClinTox and Tox21 with minimal parameters.
The paper introduces the Modality Discrepancy Transformer (MDT), a novel multimodal fusion framework that enhances cross-modal discrepancy modeling for recognizing ambivalence and hesitancy in clinical videos, outperforming baselines on the BAH dataset.
SHIFT-M3 is a lightweight pre-fusion screen that measures alignment consistency between LLM-generated and clinical report summaries of ECG records, achieving high accuracy in detecting data integrity issues.
This study examines the action-level reliability of clinical LLM agents by rerunning tasks with identical inputs and comparing orders, finding significant divergence that benchmarks may miss and proposing enhanced evaluation methods.
This paper proposes a composite objective for cone beam CT report generation that prioritizes factual entailment over lexical overlap, demonstrating that optimizing for lexical metrics harms factual accuracy. It releases a dataset and code, and presents a system that generates constrained clinical reports under polarity, laterality, and tooth level consistency constraints.
This paper proposes Fed-Equilibrium, a federated learning framework that balances robustness and fairness in clinical networks using topological Pareto control to ensure minority nodes achieve convergence comparable to dominant hubs.
This paper introduces LogiMed-RoB, a benchmark for evaluating large language models' hierarchical logical consistency in medical risk-of-bias assessment, revealing that high atomic consistency can conceal critical reasoning flaws in clinical deployment.
This paper proposes a Program-Solve interface where clinical language models generate Python code for a deterministic executor to perform math calculations, evaluating on MedCalc-Bench and finding improved accuracy for larger models like Qwen2.5-32B compared to direct arithmetic and hand-written libraries.
This paper uses mechanistic interpretability on Gemma-3-27B-PT to extract and align symptom vectors for depression with clinician judgments, demonstrating potential for interpretable clinical assessment tools.
This paper proposes AI Morbidity and Mortality (AI M&M), a blameless framework for case-based review of clinical AI failures, aiming to convert individual errors into actionable institutional learning.
This paper introduces an episode-level evaluation protocol for healthcare NLP agents to better assess performance in clinical workflows beyond static benchmarks.
The paper evaluates a configurable multi-agent system (nMAS) for extracting structured oncology data from fragmented clinical documents, achieving high performance compared to a baseline model.
This review formalizes Explainable AI (XAI) methods in computational pathology by introducing definitions, a taxonomy, and task-driven recommendations to address clinical adoption challenges.
The paper audits large language models on their refusal and fabrication behavior in clinical pain speech transcripts, finding that authority-framed prompts lead to confident fabrication in models like Gemini 2.5 Flash and Llama 3.1 8B, while cooperative prompting shows robust abstention.
The article introduces EEG-to-Report, a browser-based framework for annotating clinical EEG data and extracting features to create AI-ready datasets, with an auto-report module combining convolutional networks and large language models for generating clinical narratives.
The paper proposes a multimodal prompt-learning framework to handle missing modalities in electronic health records for robust clinical prediction in intensive care units, introducing four prompt types to capture dependencies and interactions.
This paper introduces CAIR, a two-stage framework for imputing physiological time-series data under realistic missingness, outperforming existing methods by incorporating gap mechanisms and curriculum-aware training.
This dissertation addresses security and efficiency in AI by analyzing backdoor attacks in language and vision-language models, proposing detection frameworks and novel attack methods, and introducing efficient multimodal models for clinical applications.
This paper investigates optimal-transport explanations for clinical data, showing that while heatmaps can localize synthetic lesions, they fail to localize real disease, highlighting a synthetic-to-real gap in explainable AI for healthcare.