A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks
Summary
This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across five datasets, finding that standard evaluation protocols may overestimate model performance and that leaderboard rankings lack stability.
View Cached Full Text
Cached at: 05/26/26, 09:00 AM
# A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks Source: [https://arxiv.org/abs/2605.23977](https://arxiv.org/abs/2605.23977) [View PDF](https://arxiv.org/pdf/2605.23977) > Abstract:This paper audits benchmark evaluation in clinical\-interview depression detection through four complementary probes across DAIC/E\-DAIC, CMDC, ANDROIDS, MODMA, and PDCH\. First, we re\-evaluate E\-DAIC under strict subject\-disjoint leave\-one\-subject\-out cross\-validation\. A lightweight hybrid text\-plus\-LLM\-score model reaches macro\-F1 = 0\.723 \- the highest reported under this protocol, to our knowledge \- providing a conservative out\-of\-fold reference point that does not depend on the privileged official holdout\. Second, we test whether the E\-DAIC official split supports fine\-grained leaderboard rankings by sweeping 96 model configurations across modality bundles, pooling strategies, and learners\. Development\-side cross\-validation and official\-test rankings align only moderately: the best cross\-validation configuration ranks twentieth on the official test, the official\-test winner ranks forty\-first by cross\-validation, top\-3 overlap is zero, and the apparent winner is rank\-1 in only 32\.3% of subject bootstraps\. Third, we externally validate strong public CMDC and ANDROIDS baselines that achieve near\-ceiling in\-domain performance\. Zero\-shot transfer to external corpora is substantially weaker\. Finally, we stress\-test E\-DAIC text and audio models using paired symptom\-dense versus symptom\-light interview slices defined by an SRDS\-based annotator\. Text scores rise sharply on symptom\-dense slices, whereas audio scores remain nearly flat; the text\-minus\-audio gap is positive across all five seeds\. ## Submission history From: Takehiro Ishikawa \[[view email](https://arxiv.org/show-email/0f13edff/2605.23977)\] **\[v1\]**Wed, 13 May 2026 17:32:41 UTC \(347 KB\)
Similar Articles
MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models
MedBench v5 is a dynamic, process-oriented benchmark for clinical multimodal models that integrates hallucination detection and stress testing, moving beyond static QA to evaluate reasoning and stability under information-flow stressors.
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
Introduces LingxiDiagBench, a large-scale multi-agent benchmark for evaluating LLMs on Chinese psychiatric consultation and diagnosis. Key findings show high accuracy on binary classification but poor performance on multi-way differential diagnosis, highlighting a decoupling between conversational quality and diagnostic accuracy.
Fine-tuning LLMs for Passive Depression Severity Estimation from AI Mental Health Dialogue
This paper presents a method for fine-tuning LLMs to predict PHQ-9 depression severity scores directly from transcripts of conversations with an AI mental health application, achieving strong correlation with clinical thresholds using a augmented dataset of 6,283 users.
ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV
This paper introduces ClinicalBench and the EpiKG system, evaluating assertion-aware retrieval for clinical question answering on MIMIC-IV data across multiple LLMs. It demonstrates that handling negation and temporality in retrieval significantly improves performance over standard baselines.
When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening
This paper introduces a SCID-anchored benchmark of 555 interviews to evaluate five LLMs for psychiatric screening, finding that while models show potential, they tend to discount symptom evidence in the presence of preserved functioning or protective context, requiring careful validation.