medical-ai

Tag

Cards List
#medical-ai

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

arXiv cs.CL · 3d ago Cached

This paper introduces Guideline-as-Oracle (GAO), a method for zero-annotation training of a multi-turn ophthalmic telephone triage agent by compiling American Academy of Ophthalmology guidance into a 70-row rule table used to generate 3,000 training dialogues. Fine-tuning a 9B model on this corpus improves agreement with an operational reference from 61.7% to 74.1% and emergent-case recall from 9.5% to 69.0%, beating several general-purpose systems without needing a frontier model at inference.

0 favorites 0 likes
#medical-ai

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

arXiv cs.CL · 3d ago Cached

Presents RESPClinBench, a real-world scenario benchmark for respiratory clinical decision-making, evaluating seven LLMs on COPD and pulmonary nodule cases. Finds task-specific limitations including imaging hallucination and medication-safety risks.

0 favorites 0 likes
#medical-ai

Pulmonologist illustrates why how AI is about to take over his job

Reddit r/singularity · 3d ago

A pulmonologist discusses how AI is poised to take over aspects of his medical job, highlighting the growing impact of AI in healthcare diagnostics and clinical practice.

0 favorites 0 likes
#medical-ai

@AdinaYakup: ClinFusion new open medical multimodal LLM from Alibaba DAMO Academy - 8B/32B (Apache2.0) - Unified 2D + 3D image under…

X AI KOLs Following · 4d ago Cached

ClinFusion is a new open medical multimodal LLM from Alibaba DAMO Academy, available in 8B and 32B sizes (Apache 2.0), with unified 2D + 3D image understanding and state-of-the-art results on medical benchmarks.

0 favorites 0 likes
#medical-ai

PatTree: a novel approach for automated creation of multimodal, graph-based patient representations for medical classification tasks

arXiv cs.LG · 4d ago Cached

This paper introduces PatTree, a graph-based multimodal patient representation that automatically structures heterogeneous clinical data for medical classification tasks. It achieves state-of-the-art performance on the ADNI-1 cohort with 98.5% balanced accuracy for Alzheimer's disease classification.

0 favorites 0 likes
#medical-ai

OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

arXiv cs.CL · 4d ago Cached

Introduces OncoTriad-QA, a patient-level benchmark integrating radiology, pathology, genomics, and clinical data for pan-cancer reasoning, along with OncoVLM, a reference multimodal model that outperforms existing medical LLMs after fine-tuning.

0 favorites 0 likes
#medical-ai

TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology

arXiv cs.AI · 4d ago Cached

TumorBoard is a multi-agent decision-support system for longitudinal neuro-oncology that uses a shared longitudinal case state and auditable claim-evidence ledger. It outperforms baselines on a 360-case benchmark, with a safety governor reducing harmful recommendations.

0 favorites 0 likes
#medical-ai

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

arXiv cs.AI · 4d ago Cached

Introduces MedPIC-Bench, a benchmark with counterfactual questions to evaluate whether LLMs correctly apply medication-safety rules when patient-specific conditions change; across 28 LLMs, accuracy drops significantly on counterfactual questions, revealing a common failure to revise judgments.

0 favorites 0 likes
#medical-ai

The benefits of medical AI assistance vary based on user expertise

MIT News — Artificial Intelligence · 5d ago Cached

A new MIT-led study in Nature Medicine finds that AI assistance and explainability methods impact skin disease diagnosis accuracy differently depending on user expertise: non-experts over-trust AI explanations, while clinicians perform best with only the model's prediction. The results highlight the need for user-centered AI design that accounts for automation bias.

0 favorites 0 likes
#medical-ai

TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

arXiv cs.CL · 5d ago Cached

TreeProbe is the first cultural-bias benchmark for Tibetan medicine in LLMs, containing 4,719 expert-adjudicated items across 467 diseases and 10 subtasks, revealing systematic external ontology drift in current models.

0 favorites 0 likes
#medical-ai

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

arXiv cs.AI · 6d ago Cached

EarlyDx is a new large-scale benchmark for evaluating LLMs on open-ended, evidence-supported diagnosis generation at emergency department admission, built from 154,834 MIMIC-IV encounters. It reveals that even frontier and medical-specialized models struggle to synthesize admission-time evidence, with post-training only partially improving inference-dependent recall.

0 favorites 0 likes
#medical-ai

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

arXiv cs.CL · 2026-07-29 Cached

MyoCardBench is a real-world benchmark for evaluating large language models in cardiovascular care, comprising 2,263 items across 13 tasks. GPT-5.4 achieved the highest overall score, demonstrating strengths in full-cycle care and multimodal interpretation.

0 favorites 0 likes
#medical-ai

Dementia Etiology Diagnosis via Collaborative Meta Knowledge Enhancement

arXiv cs.LG · 2026-07-28 Cached

Proposes a Collaborative Meta Knowledge Enhancement (COME) framework for dementia etiology diagnosis that injects heterogeneity-aware embeddings into a unified Transformer architecture, achieving state-of-the-art performance across multiple independent cohorts.

0 favorites 0 likes
#medical-ai

Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events

arXiv cs.CL · 2026-07-24 Cached

This paper presents a retrieval-augmented, multi-agent LLM framework with human-in-the-loop for detecting cutaneous immune-related adverse events from clinical notes, achieving higher accuracy, improved inter-rater agreement, and halved review time compared to manual review.

0 favorites 0 likes
#medical-ai

Bayesian uncertainty estimation improves clinical decision making in medical AI agents

arXiv cs.LG · 2026-07-24 Cached

This paper shows that Monte Carlo dropout provides epistemic uncertainty signals for chest radiograph classifiers, which improves error detection and reduces confident misdiagnoses in clinical decision-support agents when communicated as a binary error-risk flag.

0 favorites 0 likes
#medical-ai

OpenAI is making big claims as it rolls out ChatGPT Health to everyone

The Verge · 2026-07-23 Cached

OpenAI is rolling out ChatGPT Health to all US users, allowing them to connect medical records and health-tracking data, with claims of clinician-level reasoning and integration with GPT-5.6 Sol.

0 favorites 0 likes
#medical-ai

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

arXiv cs.AI · 2026-07-22 Cached

This paper extends missing-information stress-testing to open-ended medical conversation, finding that LLM judge choice materially changes apparent safety and that LLM judges are more permissive than clinicians.

0 favorites 0 likes
#medical-ai

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

arXiv cs.CL · 2026-07-21 Cached

This paper describes TalTech's systems for generating SOAP notes directly from doctor-patient conversation audio, using Voxtral models fine-tuned with supervised learning and DAPO reinforcement learning. Their submissions ranked first in both tracks of the BeTraC challenge, achieving high concept accuracy and low hallucination rates.

0 favorites 0 likes
#medical-ai

Tri-Net v2: Open-source implementation of our Scientific Reports paper on unified skin lesion and symptom-based monkeypox detection [R]

Reddit r/MachineLearning · 2026-07-21

Tri-Net v2 is an open-source implementation of a Scientific Reports paper for unified skin lesion and symptom-based monkeypox detection.

0 favorites 0 likes
#medical-ai

@MaziyarPanahi: I finally got Kimi K3 to read a chest X-ray, and it never saw who the patient is OpenMed stripped 23 identifiers off th…

X AI KOLs Timeline · 2026-07-20 Cached

Kimi K3 AI model successfully reads a chest X-ray after OpenMed removes all 23 patient identifiers from the DICOM data, ensuring privacy. The model correctly identifies a left-sided whiteout.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback