Tag
PatientAct is a theory-grounded framework for LLM-based simulated mental health clients, integrating clinical case formulation and dynamic memory with trust thresholds to produce more realistic resistance and behavior.
TotalSegmentator now has an MCP server, enabling AI agents like Codex to run it and answer clinical questions about CT scans, e.g., detecting hepatosplenomegaly or NAFLD/NASH.
This paper highlights that VLMs for chest x-ray report generation can score well on benchmarks while erasing clinically meaningful terms and introducing biased language, and proposes a framework to measure these failures.
MedLoCoMo is a new benchmark for evaluating LLMs on long-context, multi-session medical dialogue reasoning, constructed from MIMIC-IV data. It tests single-admission, cross-admission, and adversarial unanswerable questions, revealing that cross-admission reasoning remains challenging even for models with long context windows.
MedRealMM is a new multimodal benchmark for Chinese online medical consultation, built from real-world patient-doctor interactions, evaluating LLMs on next-response generation with clinical rubrics.
OpenMed 1.8 is an Apache-2.0 toolkit for clinical de-identification that runs entirely locally, with new support for Android, iOS, and browser platforms, and invites community contributions for version 1.9.
A systematic review of non-social media free-text datasets for mental health disorder detection, identifying biases and gaps in current resources.
The paper introduces MedCalc-Pro, a new benchmark for evaluating LLMs in complex medical calculations involving single, multi, and nested calculator settings, along with an agent framework that improves performance through multi-tool selection and structured validation.
This paper audits multilingual clinical ASR systems on psychiatric interviews in Indian languages and proposes SamaVaani, a unified debiasing technique to improve performance and fairness across demographic groups.
Introduces CRADLE-Dialogue, a clinician-annotated benchmark for turn-level crisis detection in mental health conversations, along with an Alert–Confirm evaluation protocol and a synthetic training corpus plus a 32B model that outperforms existing open-source and proprietary models.
Meddies PII is an open multilingual model and dataset for clinical text de-identification, designed to remove patient identifiers while preserving clinical facts. It uses synthetic data generated with dynamic prompting to handle diverse real-world formats.
MedCUA-Bench is a new benchmark for evaluating computer-use agents on clinical software tasks, covering 18 scenarios across 10 medical domains with safety dimensions. Results show that current agents perform poorly, especially on real OpenEMR, highlighting a significant gap in reliability.
AMNESIA is the first large-scale open-source benchmark for medical unlearning, comprising 70,560 QA pairs from 8,820 patient notes across 11 diseases, designed to evaluate forgetting of both factual and reasoning knowledge in LLMs.
This paper investigates the role of inductive bias in time-series pretraining for clinical data, proposing PathoFM, an encoder-centric transformer pretrained on multivariate gait windows. The study compares different pretraining objectives and finds that dynamics-centric mixtures yield the most balanced transfer across classification and regression tasks.
This paper investigates how large language models maintain correct beliefs under adversarial pressure in clinical settings, proposing R-FT fine-tuning to improve epistemic resilience while balancing corrigibility, and demonstrating significant robustness gains on medical benchmarks.
AnchorDiff proposes a topology-aware masked diffusion framework for radiology report generation, integrating RadGraph-derived clinical anchors and confidence-based rewriting to achieve state-of-the-art results on MIMIC-CXR and MIMIC-RG4 benchmarks.