One Year Later...The Harms Persist, But So Do We!
Summary
This study evaluates six proprietary LLMs across 16 DSM-5 conditions using adversarial attacks, finding that safety safeguards are only reliable for suicide and self-harm, with failure rates up to 100% for other conditions like eating disorders and substance use disorder.
View Cached Full Text
Cached at: 06/24/26, 07:43 AM
# One Year Later...The Harms Persist, But So Do We! Source: [https://arxiv.org/abs/2606.23884](https://arxiv.org/abs/2606.23884) [View PDF](https://arxiv.org/pdf/2606.23884) > Abstract:General\-purpose large language models \(LLMs\) are increasingly used for mental health\-related conversations, yet safety safeguards remain inadequate and inconsistent across clinical conditions\. This study evaluates six proprietary LLMs across 16 DSM\-5 conditions using four adversarial attack variants, introducing an eight\-dimension harm taxonomy and a multi\-dimensional evaluation framework\. Results show that safeguards hold reliably only for suicide and self\-harm, while conditions such as eating disorders, substance use disorder, and major depressive disorder exhibit failure rates of up to 100%\. We argue that ethical design and deployment of these LLMs demand clearly defined harm categories across clinical conditions and implementation of safeguards accordingly\. Until such safeguards are in place, these models pose significant risks to vulnerable populations, making their growing integration into educational settings a particularly concerning\. ## Submission history From: Annika Marie Schoene \[[view email](https://arxiv.org/show-email/397736a3/2606.23884)\] **\[v1\]**Mon, 22 Jun 2026 19:30:14 UTC \(460 KB\)
Similar Articles
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
This paper analyzes how LLMs internally represent self-harm content, finding that self-harm information crystallizes in the final layers and that linear separability does not align with probe accuracy.
When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening
This paper introduces a SCID-anchored benchmark of 555 interviews to evaluate five LLMs for psychiatric screening, finding that while models show potential, they tend to discount symptom evidence in the presence of preserved functioning or protective context, requiring careful validation.
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
Introduces Safe-Psych, a sequential benchmark for evaluating how large language models handle diagnostic uncertainty in psychiatry, revealing that even strong models often fail to abstain or seek clarification when clinical evidence is incomplete.
Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models
This paper presents a large-scale assessment of medical LLMs, including custom MedGPTs and open-source models, finding 25-30% exhibit low factual accuracy and 33.6-54.3% violate operational thresholds, highlighting systemic safety risks.
Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs
This paper explores using fine-tuned LLMs to identify indicators of vulnerability (mental ill health, substance misuse, alcohol dependence, homelessness) in UK police incident logs, finding that while LLMs can produce meaningful prevalence estimates, they require careful methodological support and are not reliable for individual-level decisions.