clinical-ai

Tag

Cards List
#clinical-ai

Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

arXiv cs.AI ↗ · 2026-08-14 Cached

This paper presents ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools for thyroid ultrasound, storing outputs as auditable evidence records. Developed on a large multicentre dataset, it achieves strong results in nodule segmentation, benign-malignant classification, and report generation.

0 favorites 0 likes
#clinical-ai

RadFusion: Towards Threshold-Controllable Radiology Report Generation

arXiv cs.AI ↗ · 2026-08-12 Cached

RadFusion is a framework that adds threshold controllability to radiology report generation by fusing a multi-label classifier with a VQA-based generator and an LLM rewrite step, enabling sensitivity-specificity trade-offs and ROC-based validation. Experiments on MIMIC-CXR show improved diagnostic accuracy and clinically adaptable report behavior.

0 favorites 0 likes
#clinical-ai

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

arXiv cs.AI ↗ · 2026-08-11 Cached

Introduces CliniCARE-Bench, a deployment-oriented benchmark for evaluating AI agents on clinical audit tasks over longitudinal EHR data, with 25 clinician-validated scenarios and 750 patient cases. It assesses verdict accuracy, evidence grounding, policy adherence, and calibrated abstention, finding that raw accuracy overstates investigation quality.

0 favorites 0 likes
#clinical-ai

Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI

Hugging Face Daily Papers ↗ · 2026-08-02 Cached

This paper presents a model-agnostic framework for per-modality failure analysis in multimodal clinical AI, distinguishing loud vs silent failures when a modality is dropped. Validated on planted ground truth and applied to EchoJEPA and HuBERT-ECG embeddings for LVEF prediction, it shows that dropping echo nearly doubles error.

0 favorites 0 likes
#clinical-ai

Decoding Children's Gait Behavior

Hugging Face Daily Papers ↗ · 2026-08-01 Cached

Introduces a new problem domain for fine-grained analysis of children's gait behaviors from standard RGB video, along with a new dataset of over 1,100 high-frame-rate sequences and a unified framework, demonstrating that current SOTA methods and MLLMs fail on this clinical task.

0 favorites 0 likes
#clinical-ai

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

arXiv cs.CL ↗ · 2026-07-29 Cached

MyoCardBench is a real-world benchmark for evaluating large language models in cardiovascular care, comprising 2,263 items across 13 tasks. GPT-5.4 achieved the highest overall score, demonstrating strengths in full-cycle care and multimodal interpretation.

0 favorites 0 likes
#clinical-ai

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Hugging Face Daily Papers ↗ · 2026-07-27 Cached

ClinFusion is a vision-centric multimodal large language model for holistic medical understanding that unifies 2D and 3D medical image analysis using a cascaded vision encoder. It achieves state-of-the-art results on 20 out of 24 benchmarks and outperforms proprietary models like GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.

0 favorites 0 likes
#clinical-ai

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

arXiv cs.AI ↗ · 2026-07-22 Cached

This paper extends missing-information stress-testing to open-ended medical conversation, finding that LLM judge choice materially changes apparent safety and that LLM judges are more permissive than clinicians.

0 favorites 0 likes
#clinical-ai

LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models

arXiv cs.LG ↗ · 2026-07-20 Cached

LLM4EHR proposes a clinical foundation model that temporally aligns Electronic Health Record time series with medical event sequences using a domain-adapted large language model and a regularized contrastive objective, improving downstream prediction tasks.

0 favorites 0 likes
#clinical-ai

GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis

arXiv cs.AI ↗ · 2026-07-20 Cached

GraphDx is a cost-aware, knowledge-enhanced multi-agent framework for sequential diagnosis that uses LLM-constructed medical knowledge graphs and three collaborative agents to improve diagnostic success rates and reduce test costs.

0 favorites 0 likes
#clinical-ai

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

arXiv cs.AI ↗ · 2026-07-16 Cached

Introduces Safe-Psych, a sequential benchmark for evaluating how large language models handle diagnostic uncertainty in psychiatry, revealing that even strong models often fail to abstain or seek clarification when clinical evidence is incomplete.

0 favorites 0 likes
#clinical-ai

SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction

arXiv cs.AI ↗ · 2026-07-13 Cached

SAGEAgent is an LLM-based clinical agent that sequentially decides which diagnostic modalities to acquire for cancer patients to balance predictive accuracy with clinical invasiveness, reducing acquisition burden by 55% while maintaining competitive survival prediction performance.

0 favorites 0 likes
#clinical-ai

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

arXiv cs.AI ↗ · 2026-07-10 Cached

This survey examines recent progress in medical LLMs, presenting a dual-view approach that connects clinical practice with computational methods, and introduces a benchmark dataset for evaluating medical reasoning capabilities across 18 state-of-the-art models.

0 favorites 0 likes
#clinical-ai

Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports

arXiv cs.AI ↗ · 2026-07-08 Cached

This paper introduces the Nimblemind Multi-Agent System (nMAS) for extracting evidence of H. pylori infection from gastric biopsy reports, achieving 98.61% accuracy across 216 feature-case decisions and demonstrating substantial time savings over manual review.

0 favorites 0 likes
#clinical-ai

Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models

arXiv cs.AI ↗ · 2026-07-07 Cached

This paper proposes a reinforcement learning framework for evidence-seeking diagnostic reasoning using LLMs. The RL-trained 7B model outperforms larger models in multilingual clinical consultation tasks, showing that specialized RL can distill high-level clinical reasoning.

0 favorites 0 likes
#clinical-ai

World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments

arXiv cs.AI ↗ · 2026-07-03 Cached

This paper examines the use of reinforcement learning from world feedback for clinical protocol-execution tasks in FHIR environments, identifies structural barriers like high silent-finish ceilings and zero-gradient tasks, and introduces MedAgentBench-v3 with a lower ceiling. It shows that pure RL underperforms rule-based SFT due to these barriers, and proposes a combined SFT+RL approach.

0 favorites 0 likes
#clinical-ai

Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care

arXiv cs.LG ↗ · 2026-07-01 Cached

This paper introduces Manana, a non-parametric prompt-learning framework that teaches LLMs to recommend anti-seizure medications and defer uncertain cases in underrepresented epilepsy care settings, improving accuracy on Ugandan cohorts and enabling selective prediction with high precision.

0 favorites 0 likes
#clinical-ai

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

arXiv cs.AI ↗ · 2026-06-30 Cached

This paper presents a blinded evaluation of clinical AI tools using real point-of-care queries from physicians, comparing specialized and general-purpose models across five dimensions. The specialized tool (OpenEvidence) outperformed general-purpose models on all axes, and the authors release the Real-POCQi benchmark.

0 favorites 0 likes
#clinical-ai

Clinical Harness for Governable Medical AI Skill Ecosystems

arXiv cs.AI ↗ · 2026-06-26 Cached

This paper proposes the Clinical Harness, a runtime governance architecture for registering, orchestrating, guarding, and monitoring AI-enabled clinical capabilities, using osteoporosis as a demonstration case.

0 favorites 0 likes
#clinical-ai

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

arXiv cs.CL ↗ · 2026-06-18 Cached

Introduces PhysAssistBench, a benchmark for evaluating LLMs in interactive doctor-patient-EHR assistance. Experiments show current models are unreliable in this setting, highlighting the need for coordinated capabilities.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback