HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

arXiv cs.CL Papers

Summary

HarmProfile is a benchmark dataset for analyzing harmful outputs from frontier LLMs, covering over 80,000 artifacts across 23 models and 15 harm categories to define model risk profiles.

arXiv:2608.14577v1 Announce Type: new Abstract: Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large-scale, high-quality collections of frontier-LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content-centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful-output distribution as a model-level risk profile. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface. Our source code is available at https://github.com/fresh-ma/HarmProfile .
Original Article
View Cached Full Text

Cached at: 08/18/26, 09:42 AM

# HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Source: [https://arxiv.org/abs/2608.14577](https://arxiv.org/abs/2608.14577)
[View PDF](https://arxiv.org/pdf/2608.14577)

> Abstract:Frontier large language models \(LLMs\) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis\. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large\-scale, high\-quality collections of frontier\-LLM misbehavior are difficult to obtain\. To address this gap, we introduce HarmProfile, a content\-centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful\-output distribution as a model\-level risk profile\. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures\. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories\. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface\. Our source code is available at[this https URL](https://github.com/fresh-ma/HarmProfile)\.

## Submission history

From: Zhouyuan Ma \[[view email](https://arxiv.org/show-email/ffc148b6/2608.14577)\] **\[v1\]**Thu, 11 Jun 2026 09:39:32 UTC \(12,311 KB\)

Similar Articles

The Role of Fine-grained Harm Signals in LLM Safety

arXiv cs.CL

This paper explores the role of fine-grained category-specific harmfulness representations in LLMs for safety. It shows that category residuals, orthogonal to general harm, vary in encoding harmfulness and can induce refusal, with implications for understanding LLM safety mechanisms.