medical-ai

Tag

Cards List
#medical-ai

Conformal Risk Prediction for Non-Alcoholic Fatty Liver Disease Using Gradient Boosting with Distribution-Free Coverages

arXiv cs.LG · 2026-06-10 Cached

This paper presents LiverRisk, a machine learning framework for NAFLD risk prediction that combines gradient-boosted decision trees with conformal prediction to provide calibrated, distribution-free coverage guarantees on individual risk estimates, achieving high AUROC on internal and external cohorts.

0 favorites 0 likes
#medical-ai

A Multi-modal Agentic Co-pilot for Evidence Grounded Computational Pathology

arXiv cs.AI · 2026-06-09 Cached

PathPocket is a multimodal AI agentic co-pilot for evidence-grounded pathology, utilizing a comprehensive evidence corpus and hypergraph to outperform existing state-of-the-art methods on over 200,000 real-world cases.

0 favorites 0 likes
#medical-ai

Automatic Extraction of Structured Information from Brain MRI Reports Using an Open-Weight Large Language Model

arXiv cs.AI · 2026-06-09 Cached

This paper evaluates the open-weight LLM LLaMA 3.1 for automatic extraction of structured data from Dutch brain MRI reports, achieving high performance on visual rating scores and accurate detection of findings, with few-shot prompting improving extraction of numerical variables.

0 favorites 0 likes
#medical-ai

When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG

arXiv cs.CL · 2026-06-04 Cached

A large-scale study across 5 models (7B–72B), 10 biomedical QA datasets, 4 retrieval methods, and 4 corpora finds that RAG yields only small and inconsistent gains (1–2 points) over no-retrieval baselines in biomedical question answering. The study concludes that the main bottleneck is not retrieval quality but models' limited ability to effectively use retrieved evidence.

0 favorites 0 likes
#medical-ai

An A.I. Aggregator?

Reddit r/AI_Agents · 2026-06-03

A user shares their experience using ChatGPT for complex medical caregiving and proposes the idea of aggregating multiple AI models to improve reliability by seeking consensus among different LLMs.

0 favorites 0 likes
#medical-ai

Cross-Modal Contrastive Learning of ECG and Angiography Representations for Severe Stenosis Classification

arXiv cs.LG · 2026-06-03 Cached

This paper introduces StenCE, a pretraining framework that uses cross-modal contrastive learning between ECG and X-ray angiography representations to detect severe coronary stenosis from ECGs, achieving high performance and enabling early diagnosis even in asymptomatic patients.

0 favorites 0 likes
#medical-ai

Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases

Hugging Face Daily Papers · 2026-06-03

Researchers introduce MedSP1000, a 1,638-case interactive benchmark derived from standardized patient scenarios to evaluate LLMs as dynamic clinical agents across multi-turn encounters. Results show even the best model (GPT-5.5) completes only 60.4% of expert rubric items, suggesting current LLMs are not yet reliable enough for clinical practice.

0 favorites 0 likes
#medical-ai

Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages

arXiv cs.CL · 2026-06-01 Cached

This paper investigates whether compact, task-specific bi-encoders fine-tuned on synthetic data from large language models can outperform general-purpose embeddings for clinical code retrieval in non-English languages, achieving state-of-the-art results on Spanish benchmarks CodiESP and DISTEMIST.

0 favorites 0 likes
#medical-ai

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs

arXiv cs.AI · 2026-06-01 Cached

EHRBench is an automated and reliable benchmark for evaluating LLMs on clinical decision-making tasks using real-world electronic health records, covering nearly 1M QA items across diagnosis, treatment, and prognosis tasks.

0 favorites 0 likes
#medical-ai

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

Hugging Face Daily Papers · 2026-06-01 Cached

AutoMedBench is a workflow-aware benchmark for autonomous medical-AI research, evaluating agents across five stages on diverse medical imaging tasks. Stage-level scoring reveals validation as the weakest stage, highlighting the need for reliable verification in agentic workflows.

0 favorites 0 likes
#medical-ai

Most of reddit badmouths AI, but my experience in medicine:

Reddit r/singularity · 2026-05-29

A medical professional shares their positive experience using ChatGPT to assist in diagnostic pathology, demonstrating the AI's ability to provide accurate and detailed analysis comparable to a dermatopathologist.

0 favorites 0 likes
#medical-ai

Reliable Extraction of Clinical Follow-Up Instructions: A Hybrid Neural-Symbolic Pipeline

arXiv cs.CL · 2026-05-27 Cached

This paper presents a hybrid neural-symbolic pipeline for extracting follow-up instructions from clinical notes, using BioBERT and deterministic date arithmetic. It achieves high performance (Pair F1 ~0.99) compared to generative baselines.

0 favorites 0 likes
#medical-ai

Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents

arXiv cs.AI · 2026-05-27 Cached

This paper addresses the problem of tool failures in medical AI agents by proposing a GRPO-based reinforcement learning framework that leverages instance-level selection, disagreement-aware synergy learning, and entropy-guided sampling to correct erroneous tool consensus and improve reliability across seven medical benchmarks.

0 favorites 0 likes
#medical-ai

MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning

arXiv cs.AI · 2026-05-27 Cached

MedGuideX transforms clinical practice guidelines into executable decision logic to generate factual and counterfactual QA data for training medical LLMs, achieving a 10.28% relative improvement in average accuracy across clinical reasoning benchmarks.

0 favorites 0 likes
#medical-ai

HRVConformer: Neonatal Hypoxic-Ischemic Encephalopathy Classification from the Heart Rate signals

arXiv cs.LG · 2026-05-27 Cached

This paper introduces HRVConformer, a hybrid Convolution-Transformer architecture for classifying neonatal hypoxic-ischemic encephalopathy directly from raw heart rate signals, achieving an AUC of 83.23% and outperforming baseline models like ResNet50 and Transformer.

0 favorites 0 likes
#medical-ai

RAG4Outcome: A Retrieval-Augmented Multimodal Framework for Prognostic Prediction in Chronic Osteomyelitis

arXiv cs.AI · 2026-05-25 Cached

Proposes RAG4Outcome, a retrieval-augmented generation framework integrating multimodal clinical data (PET-CT reports, surgical records, follow-up notes) to improve prognostic prediction in chronic osteomyelitis, enhancing interpretability and clinical reliability.

0 favorites 0 likes
#medical-ai

MedExpMem: Adapting Experience Memory for Differential Diagnosis

arXiv cs.LG · 2026-05-25 Cached

Proposes MedExpMem, an experience memory framework that enables medical vision-language models to accumulate and retrieve discriminative diagnostic experience from past cases, improving differential diagnosis accuracy by up to 7.0% on a radiology benchmark.

0 favorites 0 likes
#medical-ai

@itsolelehmann: marc andreessen just went on Rogan and casually dropped a TON of AI alpha full pod is 3 hours and 20 minutes, but i pul…

X AI KOLs Following · 2026-05-22 Cached

Marc Andreessen on the Joe Rogan podcast shares 17 provocative takes on AI, including claims that AGI has arrived, top models surpass human experts, and AI is transforming medicine, therapy, coding, and science.

0 favorites 0 likes
#medical-ai

Prompting language influences diagnostic reasoning and accuracy of large language models

arXiv cs.CL · 2026-05-20 Cached

This study evaluates how prompting language (English vs. French) affects diagnostic reasoning and accuracy across five LLMs using 180 clinical vignettes, finding that most models perform significantly better in English, with o3 being the only exception.

0 favorites 0 likes
#medical-ai

Forecasting Medium-Horizon Alzheimer's Disease Progression: Residual Gap-Aware Transformers for 24-Month CDR-SB Change from ADNI Clinical and Biomarker Histories

arXiv cs.LG · 2026-05-19 Cached

This paper proposes a residual gap-aware transformer that combines a mixed-effects statistical reference with transformer-based residual learning to forecast 24-month CDR-SB change from ADNI clinical and biomarker histories, achieving reduced MSE and improved correlation over baselines.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback