clinical

Tag

Cards List
#clinical

PatientAct: Theory-Grounded Mental Health Client Simulation

arXiv cs.CL · 2026-08-14 Cached

PatientAct is a theory-grounded framework for LLM-based simulated mental health clients, integrating clinical case formulation and dynamic memory with trust thresholds to produce more realistic resistance and behavior.

0 favorites 0 likes
#clinical

@JakobWasserthal: TotalSegmentator now has an MCP server. Connect it to e.g. Codex and ask questions like: “Does this CT show signs of he…

X AI KOLs Following · 2026-08-07 Cached

TotalSegmentator now has an MCP server, enabling AI agents like Codex to run it and answer clinical questions about CT scans, e.g., detecting hepatosplenomegaly or NAFLD/NASH.

0 favorites 0 likes
#clinical

VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]

Reddit r/MachineLearning · 2026-08-01

This paper highlights that VLMs for chest x-ray report generation can score well on benchmarks while erasing clinically meaningful terms and introducing biased language, and proposes a framework to measure these failures.

0 favorites 0 likes
#clinical

MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models

arXiv cs.AI · 2026-07-28 Cached

MedLoCoMo is a new benchmark for evaluating LLMs on long-context, multi-session medical dialogue reasoning, constructed from MIMIC-IV data. It tests single-admission, cross-admission, and adversarial unanswerable questions, revealing that cross-admission reasoning remains challenging even for models with long context windows.

0 favorites 0 likes
#clinical

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

arXiv cs.AI · 2026-07-13 Cached

MedRealMM is a new multimodal benchmark for Chinese online medical consultation, built from real-world patient-doctor interactions, evaluating LLMs on next-response generation with clinical rubrics.

0 favorites 0 likes
#clinical

OpenMed 1.8: Apache-2.0 clinical de-identification that runs fully local, now on Android, iOS, and in the browser. 400+ open issues if you want in on 1.9

Reddit r/LocalLLaMA · 2026-07-09

OpenMed 1.8 is an Apache-2.0 toolkit for clinical de-identification that runs entirely locally, with new support for Android, iOS, and browser platforms, and invites community contributions for version 1.9.

0 favorites 0 likes
#clinical

Mental Health Disorder Detection Beyond Social Media: A Systematic Review of Available Datasets

arXiv cs.CL · 2026-07-07 Cached

A systematic review of non-social media free-text datasets for mental health disorder detection, identifying biases and gaps in current resources.

0 favorites 0 likes
#clinical

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

arXiv cs.AI · 2026-07-07 Cached

The paper introduces MedCalc-Pro, a new benchmark for evaluating LLMs in complex medical calculations involving single, multi, and nested calculator settings, along with an agent framework that improves performance through multi-tool selection and structured validation.

0 favorites 0 likes
#clinical

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

arXiv cs.CL · 2026-06-26 Cached

This paper audits multilingual clinical ASR systems on psychiatric interviews in Indian languages and proposes SamaVaani, a unified debiasing technique to improve performance and fairness across demographic groups.

0 favorites 0 likes
#clinical

Expert-Level Crisis Detection in Mental Health Conversations

arXiv cs.CL · 2026-06-10 Cached

Introduces CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in mental health conversations, along with an Alert–Confirm evaluation protocol and a synthetic training corpus plus a 32B model that outperforms existing open-source and proprietary models.

0 favorites 0 likes
#clinical

Meddies PII: An Open Multilingual De-identification Model for Clinical Text

Reddit r/LocalLLaMA · 2026-06-08

Meddies PII is an open multilingual model and dataset for clinical text de-identification, designed to remove patient identifiers while preserving clinical facts. It uses synthetic data generated with dynamic prompting to handle diverse real-world formats.

0 favorites 0 likes
#clinical

MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents

arXiv cs.AI · 2026-06-03 Cached

MedCUA-Bench is a new benchmark for evaluating computer-use agents on clinical software tasks, covering 18 scenarios across 10 medical domains with safety dimensions. Results show that current agents perform poorly, especially on real OpenEMR, highlighting a significant gap in reliability.

0 favorites 0 likes
#clinical

AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis

arXiv cs.LG · 2026-06-01 Cached

AMNESIA is the first large-scale open-source benchmark for medical unlearning, comprising 70,560 QA pairs from 8,820 patient notes across 11 diseases, designed to evaluate forgetting of both factual and reasoning knowledge in LLMs.

0 favorites 0 likes
#clinical

On the Role of Inductive Bias in Time-Series Pretraining: A Case Study in Learning Generalizable Representations for Clinical Time Series

arXiv cs.LG · 2026-05-27 Cached

This paper investigates the role of inductive bias in time-series pretraining for clinical data, proposing PathoFM, an encoder-centric transformer pretrained on multivariate gait windows. The study compares different pretraining objectives and finds that dynamics-centric mixtures yield the most balanced transfer across classification and regression tasks.

0 favorites 0 likes
#clinical

When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure

arXiv cs.AI · 2026-05-26 Cached

This paper investigates how large language models maintain correct beliefs under adversarial pressure in clinical settings, proposing R-FT fine-tuning to improve epistemic resilience while balancing corrigibility, and demonstrating significant robustness gains on medical benchmarks.

0 favorites 0 likes
#clinical

AnchorDiff: Topology-Aware Masked Diffusion with Confidence-based Rewriting for Radiology Report Generation

arXiv cs.AI · 2026-05-19 Cached

AnchorDiff proposes a topology-aware masked diffusion framework for radiology report generation, integrating RadGraph-derived clinical anchors and confidence-based rewriting to achieve state-of-the-art results on MIMIC-CXR and MIMIC-RG4 benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback