medical-llm

Tag

Cards List
#medical-llm

Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination

arXiv cs.CL · 2026-07-02 Cached

This paper investigates whether hallucination in medical LLMs can be detected and controlled at the neuron level. The authors find that while hallucination signals are detectable across many neurons (AUROC 0.77-0.86), they are not easily corrected by steering those same neurons.

0 favorites 0 likes
#medical-llm

Primary ICD Category Prediction using LLM-based Probing

arXiv cs.AI · 2026-06-30 Cached

This paper presents a method that uses frozen medical large language model (LLM) representations as a shared embedding space to predict primary ICD diagnosis categories from both structured and unstructured electronic health record data, achieving improved accuracy over baseline methods on MIMIC-IV and showing transferability to MIMIC-III.

0 favorites 0 likes
#medical-llm

Could it be that there aren’t really any medical LLM APIs available right now? [D]

Reddit r/MachineLearning · 2026-06-24

The author notes a surprising lack of publicly available APIs for medical-oriented LLMs, despite models like MedGemma and BioMistral existing on Hugging Face, and seeks information on any available options.

0 favorites 0 likes
#medical-llm

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

arXiv cs.CL · 2026-06-18 Cached

Introduces PhysAssistBench, a benchmark for evaluating LLMs in interactive doctor-patient-EHR assistance. Experiments show current models are unreliable in this setting, highlighting the need for coordinated capabilities.

0 favorites 0 likes
#medical-llm

Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy

arXiv cs.AI · 2026-06-09 Cached

This paper introduces AI-MASLD, a stress-audit framework for medical LLMs that reveals how benchmark accuracy can hide serious safety failures, and demonstrates that open-weight models can match or exceed proprietary ones on safety dimensions.

0 favorites 0 likes
#medical-llm

HiMed: Incentivizing Hindi Reasoning in Medical LLMs

arXiv cs.CL · 2026-05-26 Cached

Introduces HiMed, a Hindi reasoning medical corpus and benchmark suite, and HiMed-8B, a Hindi-form medical reasoning LLM using decaying scaffolding reward, demonstrating improved Hindi medical reasoning and reduced English–Hindi accuracy gap.

0 favorites 0 likes
#medical-llm

When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering

arXiv cs.CL · 2026-05-22 Cached

Introduces OGCaReBench, a free-form retrieval benchmark for evaluating LLMs on clinical questions that require reasoning beyond standard guidelines. Experiments show that even the best model achieves only 56% accuracy, but retrieval augmentation boosts performance to 82%.

0 favorites 0 likes
#medical-llm

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models

arXiv cs.CL · 2026-05-21 Cached

This paper presents a large-scale assessment of medical LLMs, including custom MedGPTs and open-source models, finding 25-30% exhibit low factual accuracy and 33.6-54.3% violate operational thresholds, highlighting systemic safety risks.

0 favorites 0 likes
#medical-llm

How NOT to fine-tune your medical LLM; a look into Mark Kaplan's healtthruth.ai - "override and reframe foundational training"

Reddit r/ArtificialInteligence · 2026-05-15

This article critiques Mark Kaplan's approach to fine-tuning medical LLMs via his platform healtthruth.ai, highlighting pitfalls in overriding foundational training for healthcare AI.

0 favorites 0 likes
← Back to home

Submit Feedback