medical-ai

Tag

Cards List
#medical-ai

TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

arXiv cs.CL · 2026-08-04 Cached

TreeProbe is the first cultural-bias benchmark for Tibetan medicine in LLMs, containing 4,719 expert-adjudicated items across 467 diseases and 10 subtasks, revealing systematic external ontology drift in current models.

0 favorites 0 likes
#medical-ai

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

arXiv cs.AI · 2026-08-03 Cached

EarlyDx is a new large-scale benchmark for evaluating LLMs on open-ended, evidence-supported diagnosis generation at emergency department admission, built from 154,834 MIMIC-IV encounters. It reveals that even frontier and medical-specialized models struggle to synthesize admission-time evidence, with post-training only partially improving inference-dependent recall.

0 favorites 0 likes
#medical-ai

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

arXiv cs.CL · 2026-07-29 Cached

MyoCardBench is a real-world benchmark for evaluating large language models in cardiovascular care, comprising 2,263 items across 13 tasks. GPT-5.4 achieved the highest overall score, demonstrating strengths in full-cycle care and multimodal interpretation.

0 favorites 0 likes
#medical-ai

Dementia Etiology Diagnosis via Collaborative Meta Knowledge Enhancement

arXiv cs.LG · 2026-07-28 Cached

Proposes a Collaborative Meta Knowledge Enhancement (COME) framework for dementia etiology diagnosis that injects heterogeneity-aware embeddings into a unified Transformer architecture, achieving state-of-the-art performance across multiple independent cohorts.

0 favorites 0 likes
#medical-ai

Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events

arXiv cs.CL · 2026-07-24 Cached

This paper presents a retrieval-augmented, multi-agent LLM framework with human-in-the-loop for detecting cutaneous immune-related adverse events from clinical notes, achieving higher accuracy, improved inter-rater agreement, and halved review time compared to manual review.

0 favorites 0 likes
#medical-ai

Bayesian uncertainty estimation improves clinical decision making in medical AI agents

arXiv cs.LG · 2026-07-24 Cached

This paper shows that Monte Carlo dropout provides epistemic uncertainty signals for chest radiograph classifiers, which improves error detection and reduces confident misdiagnoses in clinical decision-support agents when communicated as a binary error-risk flag.

0 favorites 0 likes
#medical-ai

OpenAI is making big claims as it rolls out ChatGPT Health to everyone

The Verge · 2026-07-23 Cached

OpenAI is rolling out ChatGPT Health to all US users, allowing them to connect medical records and health-tracking data, with claims of clinician-level reasoning and integration with GPT-5.6 Sol.

0 favorites 0 likes
#medical-ai

Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety

arXiv cs.AI · 2026-07-22 Cached

This paper extends missing-information stress-testing to open-ended medical conversation, finding that LLM judge choice materially changes apparent safety and that LLM judges are more permissive than clinicians.

0 favorites 0 likes
#medical-ai

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

arXiv cs.CL · 2026-07-21 Cached

This paper describes TalTech's systems for generating SOAP notes directly from doctor-patient conversation audio, using Voxtral models fine-tuned with supervised learning and DAPO reinforcement learning. Their submissions ranked first in both tracks of the BeTraC challenge, achieving high concept accuracy and low hallucination rates.

0 favorites 0 likes
#medical-ai

Tri-Net v2: Open-source implementation of our Scientific Reports paper on unified skin lesion and symptom-based monkeypox detection [R]

Reddit r/MachineLearning · 2026-07-21

Tri-Net v2 is an open-source implementation of a Scientific Reports paper for unified skin lesion and symptom-based monkeypox detection.

0 favorites 0 likes
#medical-ai

@MaziyarPanahi: I finally got Kimi K3 to read a chest X-ray, and it never saw who the patient is OpenMed stripped 23 identifiers off th…

X AI KOLs Timeline · 2026-07-20 Cached

Kimi K3 AI model successfully reads a chest X-ray after OpenMed removes all 23 patient identifiers from the DICOM data, ensuring privacy. The model correctly identifies a left-sided whiteout.

0 favorites 0 likes
#medical-ai

@MaziyarPanahi: I played Thinking Machines' new Inkling a 2-minute doctor's visit, on my Mac Studio The patient came in about her knee.…

X AI KOLs Timeline · 2026-07-16 Cached

Thinking Machines' Inkling, a 975B parameter model running locally on a Mac Studio via llama.cpp, listened to a doctor's visit audio and accurately diagnosed heart failure from subtle cues in small talk, demonstrating advanced medical reasoning without leaving the machine.

0 favorites 0 likes
#medical-ai

@MaziyarPanahi: I finally got GLM-5.2 to work an entire 3-year patient chart that only Bonsai 27B was allowed to read 292 encounters li…

X AI KOLs Timeline · 2026-07-15 Cached

@MaziyarPanahi runs GLM-5.2 and Bonsai 27B locally on a Mac Studio using llama.cpp to process a 3-year patient chart, catching a dangerous drug interaction that was previously flagged but overlooked. The models operate entirely on-device under Apache-2.0, with Bonsai answering queries in ~2s and @PrismML claiming a 1-bit build fits an iPhone.

0 favorites 0 likes
#medical-ai

Agentic systems for breast cancer treatment recommendations

arXiv cs.CL · 2026-07-15 Cached

Evaluates agentic LLM systems for generating breast cancer treatment recommendations using 72 clinical cases, finding that the best system (Claude Opus 4.8 with D&C+SA pipeline) achieved a global score of 0.594 but remains insufficient for unsupervised clinical use due to persistent errors.

0 favorites 0 likes
#medical-ai

Information-seeking failures of large language models in agentic clinical reasoning

arXiv cs.AI · 2026-07-14 Cached

The paper develops an agentic evaluation framework for clinical reasoning in hematologic oncology, finding that LLMs primarily fail due to systematic information-seeking deficits rather than insufficient knowledge, with error patterns resembling cognitive biases in novice clinicians.

0 favorites 0 likes
#medical-ai

From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation

arXiv cs.AI · 2026-07-14 Cached

This paper proposes a framework that uses the Toulmin model of argumentation to structure ML-based retinal diagnosis from OCT images, integrating biomarker extraction, medical LLM reasoning (MedGemma), and similarity measures (MedSigLip) for interpretable and evidence-based diagnostic assistance.

0 favorites 0 likes
#medical-ai

@seclink: I recall verifying earlier that Ant Ling seems to give away a certain amount of tokens every day (maybe 1 million?), and it's especially good at the healthcare industry. Interested friends can give it a try, it's from a major company, reliable. https://developer.ant-ling.com/zh-CN/docs/gett…

X AI KOLs Following · 2026-07-11 Cached

Recommend Ant Ling large model API, which gives away 1 million tokens daily, excels in healthcare, supports OpenAI SDK compatible integration, and provides quick start documentation.

0 favorites 0 likes
#medical-ai

Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance

arXiv cs.AI · 2026-07-10 Cached

This paper presents HCC-STAR, a clinically aligned large language model for risk stratification and treatment guidance in hepatocellular carcinoma, aiming to improve precision therapy by leveraging electronic medical records.

0 favorites 0 likes
#medical-ai

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

arXiv cs.AI · 2026-07-10 Cached

MentalHospital is a virtual environment designed to evaluate AI agents and human experts on psychiatric clinical encounters, covering interviewing, examination, diagnosis, and treatment planning. Experiments compare human experts, trainees, and various LLMs on objective and subjective metrics.

0 favorites 0 likes
#medical-ai

Reasoning-Medical0.1-27B (Qwen3.5-27B medical finetune, claims to surpass MedGemma)

Reddit r/LocalLLaMA · 2026-07-09 Cached

EpistemeAI released Reasoning-Medical0.1-27B, a fine-tuned version of Qwen3.5-27B for medical reasoning, claiming to surpass MedGemma on several medical benchmarks by incorporating chain-of-thought reasoning on a curated dataset of 100,000 records.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback