ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

arXiv cs.CL Papers

Summary

The paper introduces ConceptGuard, a benchmark for evaluating context-sensitive unlearning in large language models using dual-use concepts, revealing that current unlearning techniques perform poorly under this practical evaluation framework.

arXiv:2608.20338v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely. Current approaches rely on disjoint forget and retain sets composed of independent facts, and measure success using simple and direct factual recall. This framing fails to capture a key requirement of unlearning, namely the ability to eliminate harmful behaviors while preserving benign and beneficial knowledge. We argue that effective unlearning must operate at the level of concepts, ensuring complete removal of unsafe applications while maintaining their correct and useful usage, thereby achieving conceptually meaningful and complete unlearning. To better evaluate unlearning techniques from such a practical viewpoint, we introduce the notion of dual-use concepts: concepts that can be used in both harmful and benign contexts. Building on these concepts, we construct a benchmark called ConceptGuard where forget and retain sets are explicitly complementary in concept usage. Our benchmark uniquely enables unlearning to be explored and gauged at the level of concepts, instead of sparse facts, and evaluation is intent-sensitive with the goal of maximizing contextual separation to promote safer behavior. We demonstrate that current unlearning techniques perform poorly under this setting, showing weak contextual separation alongside poor performance in ROUGE and concept-level metrics. Our results reveal strong forgetting-utility trade-offs, limited gains in contextual sensitivity, and poor consistency in concept-level control across methods, and provide ideas for unlearning approaches that better align with real-world safety requirements. Our dataset is publicly available.
Original Article
View Cached Full Text

Cached at: 08/21/26, 10:19 AM

# ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
Source: [https://arxiv.org/html/2608.20338](https://arxiv.org/html/2608.20338)
Sahil KaleAffiliation:Pune Institute of Computer TechnologyAffiliation:Pune, IndiaEmail:[sahilrkale05@gmail\.com](mailto:)Ian Harris

###### Abstract

Large Language Models \(LLMs\) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely\. Current approaches rely on disjoint forget and retain sets composed of independent facts, and measure success using simple and direct factual recall\. This framing fails to capture a key requirement of unlearning, namely the ability to eliminate harmful behaviors while preserving benign and beneficial knowledge\. We argue that effective unlearning must operate at the level of concepts, ensuring complete removal of unsafe applications while maintaining their correct and useful usage, thereby achieving conceptually meaningful and complete unlearning\. To better evaluate unlearning techniques from such a practical viewpoint, we introduce the notion of dual\-use concepts: concepts that can be used in both harmful and benign contexts\. Building on these concepts, we construct a benchmark calledConceptGuardwhere forget and retain sets are explicitly complementary in concept usage\. Our benchmark uniquely enables unlearning to be explored and gauged at the level of concepts, instead of sparse facts, and evaluation is intent\-sensitive with the goal of maximizingcontextual separationto promote safer behavior\. We demonstrate that current unlearning techniques perform poorly under this setting, showing weak contextual separation alongside poor performance in ROUGE and concept\-level metrics\. Our results reveal strong forgetting–utility trade\-offs, limited gains in contextual sensitivity, and poor consistency in concept\-level control across methods, and provide ideas for unlearning approaches that better align with real\-world safety requirements\. Our dataset is publicly available\.111Dataset:[https://huggingface\.co/datasets/sk0511/concept\-guard](https://huggingface.co/datasets/sk0511/concept-guard)

## 1Introduction

Large Language Models \(LLMs\) are now deployed across a wide range of applications, including education, healthcare, software development, and decision support\. Their broad adoption amplifies both their utility and their risk surface\([12](https://arxiv.org/html/2608.20338#bib.bib4)\)\. Models trained on large\-scale, heterogeneous corpora inevitably absorb undesirable content, including copyrighted material, private data, and knowledge that enables harmful or unsafe behaviors\([13](https://arxiv.org/html/2608.20338#bib.bib5)\)\. As a result, the ability to selectively remove learned information after training, increasingly referred to as*machine unlearning*\([12](https://arxiv.org/html/2608.20338#bib.bib4)\), has become a critical requirement for responsible deployment\.

The goal of unlearning is to ensure that a model no longer uses certain references from training data for producing responses while preserving its overall usefulness\([7](https://arxiv.org/html/2608.20338#bib.bib7)\)\. From a model safety perspective, we posit that this essentially translates to a goal of removing the ability to produce harmful responses, yet retaining capability to answer in other benign contexts\. In current literature for LLMs, unlearning is typically formalized by defining a*forget set*, containing data to be removed, and a*retain set*, containing data that should remain accessible\([2](https://arxiv.org/html/2608.20338#bib.bib8)\)\. Existing unlearning pipelines generally proceed as follows\. A pretrained language model is fine\-tuned on a dataset containing both parts, a retain set and a forget set, after which an unlearning method is applied on only the forget set\. Evaluation then measures two properties to check for efficient unlearning\. First,*forget quality*, which assesses whether the model can no longer recall information from the forget set\. Second,*model utility*, which evaluates whether performance on the retain set and overall capability of the model is preserved\. Several benchmarks currently operate under this framework, including TOFU\([11](https://arxiv.org/html/2608.20338#bib.bib1)\), MUSE\([14](https://arxiv.org/html/2608.20338#bib.bib2)\), and WMDP\([9](https://arxiv.org/html/2608.20338#bib.bib3)\)\.

![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_Intro_a.png)\(a\)Current unlearning benchmarks construct disjoint forget and retain sets and evaluate unlearning performance independently
![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_Intro_b.png)\(b\)Our benchmark constructs complementary forget and retain sets and evaluates concept usage with different intents

Figure 1:Comparison between existing unlearning benchmarks and our benchmark based on dual\-use conceptsWhile these benchmarks offer a useful analysis of unlearning methods from an effectiveness and computational viewpoint, we believe that current benchmarks fall short of capturing whether unlearning has occurred in a conceptually meaningful sense from a model safety perspective\. This is due to two major framing flaws in the steps of dataset construction and evaluation, which we identify as follows, respectively\.

- •First, while forget and retain sets may be drawn from similar topics or domains, they are typically constructed as random, disjoint subsets of data consisting of an assortment of facts, and thus fail to capture whether unlearning preserves context\-dependent use of certain concepts or knowledge\.
- •Secondly, evaluation methods operate at the level of isolated factual recall, testing whether specific facts have been erased or preserved as per their presence in the forget or retain set instead of evaluating if the model has been trained to prevent answering in a harmful context while also ensuring that it can safely use the same concept elsewhere\.

Table 1:Comparison of existing unlearning benchmarks with our proposed benchmarkIn many realistic safety settings, an effective unlearning objective should neither be to remove isolated facts nor to eliminate a concept wholesale, but to enable selective use: preventing harmful applications of a concept while preserving its benign, beneficial and all unrelated uses\. Consequently, strong performance on current benchmarks with isolated datasets and evaluation does not necessarily imply meaningful or safe unlearning behavior\.

In this work, we introduce a new benchmark titledConceptGuarddesigned to evaluate unlearning at the level ofdual\-useconcepts, i\.e\. concepts that can be used in both harmful and benign contexts\. Each concept in our benchmark is associated with both harmful and benign data representing dual use\. The forget set contains harmful uses of a concept, while the retain set contains benign uses of the same concept\. Crucially, these sets are complementary rather than independent or random subsets\. Within our benchmark, forget quality is directly tied to the safety of the model after unlearning, while model utility reflects its ability to retain correct and helpful behavior in benign contexts, as the objective is to introduce and subsequently unlearn unsafe behavior\. Essentially, we shift the focus of unlearning from fact deletion to checking intent\-sensitive concept removal and retention through our benchmark to provide a more faithful and practically relevant evaluation of unlearning in LLMs\.

In summary, we make the following contributions through this paper: \(1\) We analyze existing LLM unlearning benchmarks and show that their evaluation protocols fail to capture the true objective of unlearning from a safety viewpoint of removing harmful usage of concepts while maintaining performance in benign concept usage\. \(2\) We introduce ConceptGuard, a novel benchmark based on dual\-use concepts, where forget and retain sets are complementary to enable a thematic and contextual evaluation of unlearning\. \(3\) We provide an intent\-sensitive evaluation protocol and show how current unlearning techniques fail to satisfy all goals under conceptual overlap between forget and retain sets\.

## 2Background: Unlearning in Large Language Models

Machine unlearning in large language models is commonly formulated as a post\-training modification problem, where the goal is to remove the influence of a specified subset of training data without retraining the model from scratch\([16](https://arxiv.org/html/2608.20338#bib.bib9)\)\. Let𝒟f\\mathcal\{D\}\_\{f\}denote the*forget set*and𝒟r\\mathcal\{D\}\_\{r\}the*retain set*\. Given a pretrained modelff, an unlearning procedure seeks to produce an updated modelfunlearnf\_\{\\text\{unlearn\}\}such that the influence of𝒟f\\mathcal\{D\}\_\{f\}is minimized while preserving performance on𝒟r\\mathcal\{D\}\_\{r\}and general tasks\.

In existing work like[8](https://arxiv.org/html/2608.20338#bib.bib14)and[5](https://arxiv.org/html/2608.20338#bib.bib12), this objective is often operationalized through a two\-stage pipeline as follows:

1. 1\.Supervised fine\-tuning offfon a dataset containing both𝒟f\\mathcal\{D\}\_\{f\}and𝒟r\\mathcal\{D\}\_\{r\}to getfftf\_\{\\text\{ft\}\}
2. 2\.Application of an unlearning method to producefunlearnf\_\{\\text\{unlearn\}\}that modifies the model parameters with respect to removing information contained in𝒟f\\mathcal\{D\}\_\{f\}and retaining performance on𝒟r\\mathcal\{D\}\_\{r\}simultaneously

Evaluation is typically based on the dual criteria of forget quality, which measure the extent to which information from𝒟f\\mathcal\{D\}\_\{f\}is no longer recoverable, and model utility, which measures retained performance offfinfunlearnf\_\{\\text\{unlearn\}\}for in other tasks\. This formulation forms the basis of most existing benchmarks and methods for unlearning in LLMs, and we follow a similar structure in our benchmark, albeit with conceptual enhancements\.

## 3Limitations of Existing Benchmarks

We examine two key conceptual limitations in existing unlearning benchmarks from a model safety perspective\. While both stem from a common issue, we present them separately to analyze their effects on the unlearning evaluation pipeline\. As shown in Figure[1](https://arxiv.org/html/2608.20338#S1.F1)and summarized in Table[1](https://arxiv.org/html/2608.20338#S1.T1), current benchmarks treat dataset construction and evaluation as disjoint processes\. Together, these lead to a mismatch between benchmark performance and meaningful unlearning behavior, which ConceptGuard aims to address\.

### 3\.1Disjoint Construction of Forget and Retain Sets

A primary limitation lies in how the forget and retain sets,𝒟f\\mathcal\{D\}\_\{f\}and𝒟r\\mathcal\{D\}\_\{r\}, are constructed\. These are typically treated as disjoint subsets sampled from a larger dataset, without explicitly encoding relationships between them\. Consequently, unlearning is evaluated at the level of isolated facts rather than underlying concepts and their contextual usage\.

While this design suits privacy\-preserving settings that require removing specific records, safety\-oriented unlearning requires modifying behavior based on context\. For instance, knowledge of chemical synthesis may be harmful in one setting but necessary in educational or industrial contexts\. This requires distinguishing harmful and benign uses of the same concept\.

Existing benchmarks do not capture this distinction\. TOFU\([11](https://arxiv.org/html/2608.20338#bib.bib1)\)splits fictional author data into forget and retain sets via percentage partitioning, without enforcing conceptual alignment\. MUSE\([14](https://arxiv.org/html/2608.20338#bib.bib2)\)similarly partitions data from sources such as Harry Potter texts and news corpora, but does not ensure complementary usage across𝒟f\\mathcal\{D\}\_\{f\}and𝒟r\\mathcal\{D\}\_\{r\}\. WMDP\([9](https://arxiv.org/html/2608.20338#bib.bib3)\)focuses on hazardous knowledge in domains such as biosecurity and cybersecurity, but evaluates primarily on harmful queries, with utility measured outside the same conceptual scope\.

As a result,𝒟f\\mathcal\{D\}\_\{f\}and𝒟r\\mathcal\{D\}\_\{r\}remain structurally independent\. This prevents evaluation of whether models can selectively suppress harmful uses while preserving beneficial ones, and instead favors solutions operating at the level of individual facts, and also prevents analysis of how unlearning performance differs based on data themes\.

### 3\.2Evaluation Lacks Contextual Sensitivity

A related limitation arises in evaluation\. Existing benchmarks assess whether knowledge from the forget set is removed, without verifying whether benign uses of the same concepts are preserved\. Evaluation is typically split into*forget quality*on𝒟f\\mathcal\{D\}\_\{f\}and*model utility*on𝒟r\\mathcal\{D\}\_\{r\}or unrelated tasks\.

Forget quality measures inability to reproduce information from𝒟f\\mathcal\{D\}\_\{f\}\. TOFU uses probability, ROUGE, and truth ratio on QA pairs, MUSE evaluates memorization and membership inference, and WMDP uses accuracy on hazardous queries\. These focus on whether specific information is no longer accessible\. Model utility is evaluated largely independently\. TOFU includes retain set and auxiliary datasets such as real authors and world facts but not conceptually linked to the forget set, MUSE evaluates retained performance separately, and WMDP relies on general benchmarks such as MMLU\. These evaluations are often outside the conceptual domain of the forget set\.

This separation creates a fundamental gap\. Since forgetting and utility are evaluated on disjoint sets and often unrelated domains, benchmarks do not test whether a model can apply the same concept differently across contexts\. In these cases, if models suppress entire concepts rather than selectively modifying their usage, strong benchmark performance does not imply context\-sensitive or safety\-aligned unlearning\.

## 4The ConceptGuard Benchmark

Our proposed benchmark is designed to address the limitations in dataset construction and evaluation identified in existing unlearning frameworks\. It also enables flexible or targeted analysis of concept\-level unlearning, where forget and retain set concepts can be adjusted for contextual focus\. We describe the design of the dataset, including its construction around dual\-use concepts, and introduce an evaluation framework that captures context\-dependent unlearning behavior\.

### 4\.1Dual\-Use Concepts

We base our unlearning benchmark on the notion of*dual\-use*concepts, defined as concepts that can be applied in both harmful and benign contexts\. While dual\-use as a notion has been discussed in prior work on AI risks\([1](https://arxiv.org/html/2608.20338#bib.bib10)\)as well as in LLM\-specific settings\([15](https://arxiv.org/html/2608.20338#bib.bib11)\), we adopt a formulation such that a dual\-use concept can be easily used in either harmful or benign ways depending on context, framing, intent, and how it is combined with other concepts\. Through these concepts, we aim to implement the objective of unlearning as not removing a concept entirely, but preventing its harmful application while preserving its benign and beneficial uses\. For example, knowledge related to cybersecurity techniques, blockchain working, or biochemical processes may be essential in educational, research, or industrial contexts, while also enabling harmful use if applied with malicious intent\. Effective unlearning in such settings therefore requires distinguishing between these contexts rather than suppressing the concept altogether\.

Motivated by this, each concept in our benchmark is associated with two complementary forms of data: harmful instances, which constitute the forget set𝒟f\\mathcal\{D\}\_\{f\}, and benign instances, which constitute the retain set𝒟r\\mathcal\{D\}\_\{r\}\. Crucially, these sets are not constructed independently, but are explicitly paired to represent different uses of the same underlying concept\.

### 4\.2Dataset Construction

The dataset is constructed through a multi\-stage pipeline consisting of source data extraction, concept identification, aggregation, and generation of complementary benign instances\. LLM\-assisted dataset construction was manually supervised at three stages: dual\-use concept identification, concept aggregation, and validation of generated benign counterparts\. Detailed annotation procedures and inter\-annotator agreement statistics are provided in Section[A](https://arxiv.org/html/2608.20338#A1)\.

- •Data Source Selection:We base our dataset on the LLM\-LAT harmful dataset\([10](https://arxiv.org/html/2608.20338#bib.bib13)\), which contains prompts designed to elicit unsafe behavior along with corresponding GPT\-3\.5 responses\. The responses in therejectedcolumn serve as the primary source of harmful instances\. We intentionally use model\-generated unsafe outputs to ensure real\-world applicability\.
- •Identification of Dual\-Use Concepts:Each prompt is processed using a GPT\-5\-based \(gpt\-5\-2025\-08\-07\) classifier to determine whether it reflects harmful usage of a dual\-use concept\. The classifier is instructed to identify high\-level conceptual capabilities applicable in both benign and harmful contexts, while excluding inherently malicious or narrowly defined activities \(e\.g\., explicit bio\-terrorism\)\. The tagging prompt is provided in Figure[5](https://arxiv.org/html/2608.20338#A1.F5)in the Appendix\. This yields mappings of the form\(prompt,concept,harmful response\)\(\\text\{prompt\},\\text\{concept\},\\text\{harmful response\}\)\.
- •Concept Aggregation and Forget Set Formulation:Extracted concepts are aggregated to analyze frequency and distribution\. Many occur infrequently and correspond to narrow variants \(e\.g\.,*SQL injection*vs\.*database exploitation*,*spam bots*vs\.*automated messaging abuse*\)\. These are merged into broader parent concepts through manual curation by annotators with graduate\-level expertise\. For each finalized concept, associated harmful responses are retained as instances in the forget set𝒟f\\mathcal\{D\}\_\{f\}, capturing behaviors the model is expected to unlearn\.
- •Generation of Benign Counterparts for the Retain Set:For each instance in𝒟f\\mathcal\{D\}\_\{f\}, a complementary benign instance is generated using GPT\-5 under manual supervision\. The model re\-frames the same concept by modifying its usage in the original prompt and produces a response of comparable length and detail in a constructive or informational context\. It is further instructed to mirror the structure of harmful responses to ensure stylistic consistency\. The generation prompt is provided in Figure[6](https://arxiv.org/html/2608.20338#A1.F6)in the Appendix\. These outputs form the retain set𝒟r\\mathcal\{D\}\_\{r\}, with spot\-checking to ensure coherence and non\-harmfulness\.

### 4\.3Dataset Statistics and Structure

The final dataset consists of 5,166 instances associated with dual\-use concepts, evenly split between harmful and benign usage\. Harmful instances constitute the forget set𝒟f\\mathcal\{D\}\_\{f\}, while benign instances form the retain set𝒟r\\mathcal\{D\}\_\{r\}, with both sets constructed as complementary examples of the same underlying concepts\. Table[6](https://arxiv.org/html/2608.20338#A1.T6)in the Appendix shows the main statistics of the final ConceptGuard dataset\. The most frequent concepts include cybersecurity \(302 instances\), social engineering \(219\), and disinformation \(103\), reflecting common dual\-use domains in real\-world safety settings\. Examples from the dataset are also provided for reference in Table[7](https://arxiv.org/html/2608.20338#A1.T7)in the Appendix\.

### 4\.4Evaluation Protocol

Our benchmark evaluates unlearning along two dimensions:*forget quality*and*model utility*, along with an additional concept\-level measure of*contextual separation*\. Since the objective is to induce and subsequently remove unsafe behavior, forget quality is directly tied to the safety of the model after unlearning, while model utility reflects its ability to retain and enhance correct and helpful behavior in benign contexts\. For evaluation, in addition to𝒟f\\mathcal\{D\}\_\{f\}and𝒟r\\mathcal\{D\}\_\{r\}, we construct corresponding query sets𝒬f\\mathcal\{Q\}\_\{f\}and𝒬r\\mathcal\{Q\}\_\{r\}, consisting of questions with harmful and benign intents, respectively\. The size of each query set is the same as the training sets and each query is framed such that the expected response aligns with an instance in the forget or retain sets\. The query construction process is described in Section[A\.4](https://arxiv.org/html/2608.20338#A1.SS4)in the Appendix\. For metrics based on LLM\-as\-a\-judge, additional details are given in Section[B](https://arxiv.org/html/2608.20338#A2)in the Appendix\. Letffdenote the unlearned model\.

#### 4\.4\.1Forget Quality

\(F1\) Harmful Memorization:A model that has successfully unlearned𝒟f\\mathcal\{D\}\_\{f\}should not reproduce its contents verbatim\. For each sequencex∈𝒟fx\\in\\mathcal\{D\}\_\{f\}, we prompt the model with the firstlltokensx\[:l\]x\[:l\]and compare the generated continuationf\(x\[:l\]\)f\(x\[:l\]\)with the ground\-truth continuationx\[l\+1:\]x\[l\+1:\]using ROUGE\-L\. Lower values indicate better forgetting\. We define:

HarmMem\(f\):=1\|𝒟f\|∑x∈𝒟fROUGE\(f\(x\[:l\]\),x\[l\+1:\]\)\\small\\text\{HarmMem\}\(f\):=\\tfrac\{1\}\{\|\\mathcal\{D\}\_\{f\}\|\}\\textstyle\\sum\_\{x\\in\\mathcal\{D\}\_\{f\}\}\\text\{ROUGE\}\(f\(x\[:l\]\),\\,x\[l\+1:\]\)\(1\)
\(F2\) Query\-Based Harmful Memorization:We assess whether the model can still generate responses aligned with harmful instances when prompted with queries from𝒬f\\mathcal\{Q\}\_\{f\}\. For each queryq∈𝒬fq\\in\\mathcal\{Q\}\_\{f\}, we compare the generated responsef⁡\(q\)f\(q\)with the reference responsea⁡\(q\)a\(q\)derived from the corresponding instance in𝒟f\\mathcal\{D\}\_\{f\}\. Lower values indicate better forgetting\.

HarmQA​\(f\):=1\|𝒬f\|​∑q∈𝒬fROUGE​\(f⁡\(q\),a⁡\(q\)\)\\small\\text\{HarmQA\}\(f\):=\\tfrac\{1\}\{\|\\mathcal\{Q\}\_\{f\}\|\}\\textstyle\\sum\_\{q\\in\\mathcal\{Q\}\_\{f\}\}\\text\{ROUGE\}\(f\(q\),\\,a\(q\)\)\(2\)
\(F3\) Harmful Response Score:We evaluate whether the model produces outputs with harmful wording or intent when prompted with queries intending such responses\. For eachq∈𝒬fq\\in\\mathcal\{Q\}\_\{f\}, we obtain the model responsef⁡\(q\)f\(q\)and score it using an LLM\-based evaluatorg⁡\(⋅\)g\(\\cdot\)for harmfulness seen\. Lower values indicate safer behavior\. We define:

HarmScore​\(f\):=1\|𝒬f\|​∑q∈𝒬fg⁡\(f⁡\(q\)\)\\small\\text\{HarmScore\}\(f\):=\\tfrac\{1\}\{\|\\mathcal\{Q\}\_\{f\}\|\}\\textstyle\\sum\_\{q\\in\\mathcal\{Q\}\_\{f\}\}g\(f\(q\)\)\(3\)

#### 4\.4\.2Model Utility

\(U1\) Helpful Retention:To ensure that useful knowledge from the retain set is successfully preserved, we compute an analogous verbatim similarity metric on𝒟r\\mathcal\{D\}\_\{r\}\. Higher values indicate better retention\.

HelpRet\(f\):=1\|𝒟r\|∑x∈𝒟rROUGE\(f\(x\[:l\]\),x\[l\+1:\]\)\\small\\text\{HelpRet\}\(f\):=\\tfrac\{1\}\{\|\\mathcal\{D\}\_\{r\}\|\}\\textstyle\\sum\_\{x\\in\\mathcal\{D\}\_\{r\}\}\\text\{ROUGE\}\(f\(x\[:l\]\),\\,x\[l\+1:\]\)\(4\)
\(U2\) Query\-Based Utility:We evaluate the model’s ability to respond correctly to benign queries from𝒬r\\mathcal\{Q\}\_\{r\}\. For eachq∈𝒬rq\\in\\mathcal\{Q\}\_\{r\}, we compare the generated responsef⁡\(q\)f\(q\)with a reference responsea⁡\(q\)a\(q\)derived from the corresponding instance in𝒟r\\mathcal\{D\}\_\{r\}\. Higher values indicate better utility\.

HelpQA​\(f\):=1\|𝒬r\|​∑q∈𝒬rROUGE​\(f⁡\(q\),a⁡\(q\)\)\\small\\text\{HelpQA\}\(f\):=\\tfrac\{1\}\{\|\\mathcal\{Q\}\_\{r\}\|\}\\textstyle\\sum\_\{q\\in\\mathcal\{Q\}\_\{r\}\}\\text\{ROUGE\}\(f\(q\),\\,a\(q\)\)\(5\)
\(U3\) Helpful Response Score:We measure whether the model produces useful responses in benign contexts using an LLM\-based evaluatorh⁡\(⋅\)h\(\\cdot\)that scores the general cohesion and utility of responses\. Higher values indicate better utility\.

HelpScore​\(f\):=1\|𝒬r\|​∑q∈𝒬rh⁡\(f⁡\(q\)\)\\small\\text\{HelpScore\}\(f\):=\\tfrac\{1\}\{\|\\mathcal\{Q\}\_\{r\}\|\}\\textstyle\\sum\_\{q\\in\\mathcal\{Q\}\_\{r\}\}h\(f\(q\)\)\(6\)

#### 4\.4\.3Contextual Separation

We measure the extent to which the model differentiates between harmful and benign uses of the same concept\. Let𝒞\\mathcal\{C\}denote the set of dual\-use concepts, and𝒬fc,𝒬rc\\mathcal\{Q\}\_\{f\}^\{c\},\\mathcal\{Q\}\_\{r\}^\{c\}denote the subsets of harmful and benign queries corresponding to conceptc∈𝒞c\\in\\mathcal\{C\}\. We define the concept\-wise separation as:

Sep​\(f,c\):=HelpScorec​\(f\)−HarmScorec​\(f\)\\small\\text\{Sep\}\(f,c\):=\\text\{HelpScore\}\_\{c\}\(f\)\-\\text\{HarmScore\}\_\{c\}\(f\)\(7\)
The overall contextual separation is given by the equation ahead\. Higher values indicate stronger ability to suppress harmful behavior while preserving benign usage within the same concept\.

CtxtSep​\(f\):=∑c∈𝒞wc⋅Sep​\(f,c\),wc=\|𝒬fc\|\+\|𝒬rc\|\|𝒬f\|\+\|𝒬r\|\\small\\text\{CtxtSep\}\(f\):=\\textstyle\\sum\_\{c\\in\\mathcal\{C\}\}w\_\{c\}\\cdot\\text\{Sep\}\(f,c\),\\hskip 9\.24994ptw\_\{c\}=\\frac\{\|\\mathcal\{Q\}\_\{f\}^\{c\}\|\+\|\\mathcal\{Q\}\_\{r\}^\{c\}\|\}\{\|\\mathcal\{Q\}\_\{f\}\|\+\|\\mathcal\{Q\}\_\{r\}\|\}\(8\)
↑\\uparrowx%

desirable↑\\uparrowx%undesirableBold= best per column\. Change relative to base model\.

Table 2:ROUGE\-based evaluation results

## 5Experimental Setup

### 5\.1Unlearning Methods

We evaluate unlearning methods in a setting where models are first exposed to dual\-use concepts with harmful and benign usages, and subsequently trained to unlearn harmful usages of the concept\. We select a representative set of methods to analyze their behavior at the concept level under our context\-dependent usage\.

- •Gradient Ascent \(with Gradient Descent on the Retain Set\):Gradient ascent\([8](https://arxiv.org/html/2608.20338#bib.bib14)\)directly minimizes the likelihood of harmful data by maximizing the training loss on𝒟f\\mathcal\{D\}\_\{f\}\. To try to preserve model utility, we directly train the model using gradient ascent on the retain set simultaneously, as done in\([14](https://arxiv.org/html/2608.20338#bib.bib2)\)\. This method serves as a simple and widely\-used baseline for direct suppression of harmful data and preservation of useful contents\.
- •SimNPO:SimNPO\([5](https://arxiv.org/html/2608.20338#bib.bib12)\)is a preference\-based unlearning method that suppresses harmful responses by directly penalizing their likelihood using a reference\-free objective\. We include it to study if preference\-based formulations enable selective, behavior\-level unlearning\.
- •Representation Misdirection for Unlearning \(RMU\):RMU\([9](https://arxiv.org/html/2608.20338#bib.bib3)\)operates at the representation level by perturbing internal activations for harmful data while preserving those for benign data\. This method is particularly relevant to our benchmark as it explicitly attempts to separate harmful and benign representations\.
- •UNDIAL:UNDIAL\([3](https://arxiv.org/html/2608.20338#bib.bib15)\)performs unlearning through self\-distillation by modifying the model’s output distribution to downweight harmful tokens\. This approach provides a softer alternative to direct suppression and is included to evaluate whether distribution\-level adjustments yield better contextual behavior\.

### 5\.2Models and Setup

We conduct our experiments using two instruction\-tuned base models: Qwen\-2\.5\-3B\-Instruct and Llama\-3\.1\-8B\-Instruct\. We first fine\-tune the base modelffon the combined dataset𝒟f∪𝒟r\\mathcal\{D\}\_\{f\}\\cup\\mathcal\{D\}\_\{r\}to obtainfftf\_\{\\text\{ft\}\}\. Unlearning methods are then applied tofftf\_\{\\text\{ft\}\}with respect to the forget set𝒟f\\mathcal\{D\}\_\{f\}, resulting in an unlearned modelfunlearnf\_\{\\text\{unlearn\}\}\. All implementation details, including fine\-tuning and unlearning hyperparameters, are provided in the Appendix Section[C](https://arxiv.org/html/2608.20338#A3)\. Results from ROUGE based metrics are provided in Table[2](https://arxiv.org/html/2608.20338#S4.T2), while analysis results using LLM\-as\-a\-judge with GPT\-5\.4 \(instead of GPT\-5 to avoid circular evaluation\) are presented in Table[3](https://arxiv.org/html/2608.20338#S5.T3)\.

↑\\uparrowx%

desirable↑\\uparrowx%undesirableBold= best per column\. Change relative to base model\.

Table 3:LLM\-as\-a\-judge evaluation results along with contextual separation scores

## 6Results and Discussion

Unlearning methods induce strong forgetting–utility trade\-offs under conceptual overlap\.Across both models, fine\-tuning successfully induces strong memorization of harmful and benign instances as expected, however, unlearning methods reverse this trend to varying degrees\. Gradient Ascent achieves the strongest forgetting \(lowestHarmMemandHarmQA\), but at the cost of severe utility degradation, indicating over\-suppression and a major collapse of internals\. In contrast, SimNPO and RMU provide a more balanced trade\-off, retaining substantially higher utility while still reducing harmful memorization\. SimNPO achieves the strongest overall utility retention with relatively low harmful outputs, suggesting that framing concept usage through preferences is an effective strategy for safe unlearning\. UNDIAL occupies an intermediate regime, with competitive forgetting but weaker retention; however, its performance improves with scale, indicating stronger potential for larger models where concept representations layers may be greater\([6](https://arxiv.org/html/2608.20338#bib.bib16)\)\.

Current unlearning methods fail to enhance the contextual separation of safe and benign usage of concepts\.Overall, our results show thattruly safe unlearning, which we define as enhancing contextual separation of concept usage, is only partially achieved\. Even though forget and retain sets complementarily encode such behavior, unlearning methods fail to assimilate this dual goal\. SimNPO and RMU achieve the highest separation scores across both models, indicating a stronger ability to suppress harmful behavior while preserving benign usage within the same conceptual space, albeit performance can be made significantly better\. This gap is more pronounced in the larger model, suggesting that higher\-capacity models work better under concept entanglement\.

![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_Qwen_Heatmap.png)\(a\)Qwen\-2\.5\-3B\-Instruct
![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_Lllama_Heatmap.png)\(b\)Llama\-3\.1\-8B\-Instruct

Figure 2:Concept\-wise contextual separation across methods for the top varying concepts\.Unlearning methods fail to show uniformity at the concept level\.We analyze contextual separation \(normalized by number of samples per concept\) along two complementary views: \(i\) the top 8 concepts exhibiting maximum variation across methods \(restricted to concepts with at least 50 instances\), and \(ii\) the overall score distribution across major concepts \(with at least 15 instances\)\. As shown in Figure[2](https://arxiv.org/html/2608.20338#S6.F2), concepts such as anonymity and social media consistently exhibit high variance across methods, suggesting that concepts grounded inhuman behaviorunder conflicting contexts are inherently harder to unlearn uniformly\. This lack of consistency persists at a broader level as seen in Figure[4](https://arxiv.org/html/2608.20338#S6.F4): even when considering aggregate distributions across major concepts, methods do not exhibit stable or uniform separation patterns across certain concepts, with heavy variance across methods and concepts visible for all except GA\. Further, Tables[5](https://arxiv.org/html/2608.20338#S6.T5)and[5](https://arxiv.org/html/2608.20338#S6.T5)showing concepts with the highest and lowest contextual separation scores reinforce this observation, showing no fixed set of concepts that remain consistently separable across methods or model scales\. However, weak structure emerges\. SimNPO tends to favor concepts framed more as preferences or intent\-driven behavior \(e\.g\., anonymity, social engineering\), while RMU shows relatively stronger separation on system\-level or operational concepts \(e\.g\., automation, telecommunications\)\. In contrast, GA and UNDIAL exhibit largely inconsistent and diffuse behavior across concepts\. Overall, these results indicate that current unlearning methods lack fine\-grained control at the concept level, and fail to generalize uniformly under conceptual entanglement\.

Contextual separation is weakly sensitive to forget set size\.Reducing the proportion of the forget set \(while maintaining same proportion of concepts in retain set\) leads to a consistent but marginal increase in contextual separation across all methods and models \(Figure[4](https://arxiv.org/html/2608.20338#S6.F4)\)\. However, the gains are limited and largely attributable to the increased influence of the retain set, which directly boosts helpfulness scores\. The relative ranking of methods remains largely unchanged\. This suggests that scaling down the forget set alone is insufficient to meaningfully improve concept\-level unlearning performance, and that specific contextual\-separation unlearning methods are necessary to achieve better scores and consequently, safer unlearned models\.

Table 4:Contextual separation: Qwen\-2\.5\-3B
Table 5:Contextual separation: Llama\-3\.1\-8B

![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_Qwen_ConceptDistr.png)\(a\)Qwen\-2\.5\-3B
![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_Llama_ConceptDistr.png)\(b\)Llama\-3\.1\-8B

Figure 3:Concept\-wise contextual separation across methods for high\-frequency concepts
![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_Qwen_ForgetSize.png)\(c\)Qwen\-2\.5\-3B
![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_Llama_ForgetSize.png)\(d\)Llama\-3\.1\-8B

Figure 4:Overall contextual separation scores across methods based on forget set size

## 7Conclusion and Future Work

We introduce a concept\-aware benchmark for evaluating unlearning under contextual overlap, where the same concept appears in both harmful and benign settings\. Our results show that while existing methods can effectively reduce harmful memorization, they consistently induce a strong forgetting–utility trade\-off and fail to meaningfully enhance contextual separation\. Among evaluated approaches, preference\-based and representation\-level methods achieve a more balanced outcome, but still fall short of robust concept\-level disentanglement\. Through fine\-grained analysis, we further show that unlearning behavior is highly variable and does not generalize uniformly across concepts, highlighting a fundamental limitation of current approaches\.

Our study opens several directions for future work\. First, the current dataset maintains a fixed distribution of concept frequencies; exploring alternative distributions, filtering strategies, and different thematic groupings could reveal deeper insights into concept sensitivity\. Second, varying the number and granularity of concepts may help better understand the limits of contextual separation\. Extending the benchmark to additional domains, languages, and more diverse concept spaces is another natural direction\. Finally, developing unlearning methods that explicitly optimize for contextual separation, rather than treating forgetting and retention independently, remains a key open challenge\.

## Impact Statement

This work aims to improve the safety of language models by enabling precise removal of harmful behaviors while preserving useful capabilities\. However, unlearning methods could be misused to selectively suppress beneficial or factual information, raising concerns around controllability and misuse\. Developing robust, transparent, and auditable unlearning methods remains an important direction for mitigating such risks\.

## References

- Brundageet al\.\(2018\)M\. Brundage, S\. Avin, J\. Clark, H\. Toner, P\. Eckersley, B\. Garfinkel, A\. Dafoe, P\. Scharre, T\. Zeitzoff, B\. Filar, H\. Anderson, H\. Roff, G\. Allen, J\. Steinhardt, C\. Flynn, S\. Orsquo;Heigeartaigh, S\. Beard, H\. Belfield, S\. Farquhar, C\. Lyle, R\. Crootof, O\. Evans, M\. Page, J\. Bryson, R\. Yampolskiy, and D\. AmodeiThe malicious use of artificial intelligence: forecasting, prevention, and mitigation\.University of Cambridge,Apollo \- University of Cambridge Repository\.External Links:[Link](https://www.repository.cam.ac.uk/handle/1810/275332),[Document](https://dx.doi.org/10.17863/CAM.22520)Cited by:[§4\.1](https://arxiv.org/html/2608.20338#S4.SS1.p1.1)\.
- Chang and Lee \(2025\)H\. Chang and H\. LeeWhich retain set matters for LLM unlearning? a case study on entity unlearning\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 5966–5982\.External Links:[Link](https://aclanthology.org/2025.findings-acl.310/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.310),ISBN 979\-8\-89176\-256\-5Cited by:[§1](https://arxiv.org/html/2608.20338#S1.p2.1)\.
- Donget al\.\(2025\)Y\. R\. Dong, H\. Lin, M\. Belkin, R\. Huerta, and I\. VulićUNDIAL: self\-distillation with adjusted logits for robust unlearning in large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),Albuquerque, New Mexico,pp\. 8827–8840\.External Links:[Link](https://aclanthology.org/2025.naacl-long.444/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.444),ISBN 979\-8\-89176\-189\-6Cited by:[4th item](https://arxiv.org/html/2608.20338#S5.I1.i4.p1.1)\.
- Dornaet al\.\(2025\)V\. Dorna, A\. Mekala, W\. Zhao, A\. McCallum, Z\. C\. Lipton, J\. Z\. Kolter, and P\. MainiOpenUnlearning: accelerating LLM unlearning via unified benchmarking of methods and metrics\.arXiv preprint arXiv:2506\.12618\.External Links:[Link](https://arxiv.org/abs/2506.12618)Cited by:[Appendix C](https://arxiv.org/html/2608.20338#A3.p1.1)\.
- Fanet al\.\(2025\)C\. Fan, J\. Liu, L\. Lin, J\. Jia, R\. Zhang, S\. Mei, and S\. LiuSimplicity prevails: rethinking negative preference optimization for llm unlearning\.InAdvances in Neural Information Processing Systems,Note:PosterCited by:[§2](https://arxiv.org/html/2608.20338#S2.p2.1),[2nd item](https://arxiv.org/html/2608.20338#S5.I1.i2.p1.1)\.
- Gevaet al\.\(2021\)M\. Geva, R\. Schuster, J\. Berant, and O\. LevyTransformer feed\-forward layers are key\-value memories\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,M\. Moens, X\. Huang, L\. Specia, and S\. W\. Yih \(Eds\.\),Online and Punta Cana, Dominican Republic,pp\. 5484–5495\.External Links:[Link](https://aclanthology.org/2021.emnlp-main.446/),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.446)Cited by:[§6](https://arxiv.org/html/2608.20338#S6.p1.1)\.
- Houet al\.\(2025\)L\. Hou, Z\. Wang, G\. Liu, C\. Wang, W\. Liu, and K\. PengDecoupling memories, muting neurons: towards practical machine unlearning for large language models\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 13978–13999\.External Links:[Link](https://aclanthology.org/2025.findings-acl.719/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.719),ISBN 979\-8\-89176\-256\-5Cited by:[§1](https://arxiv.org/html/2608.20338#S1.p2.1)\.
- Janget al\.\(2023\)J\. Jang, D\. Yoon, S\. Yang, S\. Cha, M\. Lee, L\. Logeswaran, and M\. SeoKnowledge unlearning for mitigating privacy risks in language models\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 14389–14408\.External Links:[Link](https://aclanthology.org/2023.acl-long.805/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.805)Cited by:[§2](https://arxiv.org/html/2608.20338#S2.p2.1),[1st item](https://arxiv.org/html/2608.20338#S5.I1.i1.p1.1)\.
- Liet al\.\(2024\)N\. Li, A\. Pan, A\. Gopal, S\. Yue, D\. Berrios, A\. Gatti, J\. D\. Li, A\. Dombrowski, S\. Goel, L\. Phan, G\. Mukobi, N\. Helm\-Burger, R\. Lababidi, L\. Justen, A\. B\. Liu, M\. Chen, I\. Barrass, O\. Zhang, X\. Zhu, R\. Tamirisa, B\. Bharathi, A\. Khoja, Z\. Zhao, A\. Herbert\-Voss, C\. B\. Breuer, S\. Marks, O\. Patel, A\. Zou, M\. Mazeika, Z\. Wang, P\. Oswal, W\. Lin, A\. A\. Hunt, J\. Tienken\-Harder, K\. Y\. Shih, K\. Talley, J\. Guan, R\. Kaplan, I\. Steneker, D\. Campbell, B\. Jokubaitis, A\. Levinson, J\. Wang, W\. Qian, K\. K\. Karmakar, S\. Basart, S\. Fitz, M\. Levine, P\. Kumaraguru, U\. Tupakula, V\. Varadharajan, R\. Wang, Y\. Shoshitaishvili, J\. Ba, K\. M\. Esvelt, A\. Wang, and D\. HendrycksThe wmdp benchmark: measuring and reducing malicious use with unlearning\.External Links:2403\.03218,[Link](https://arxiv.org/abs/2403.03218)Cited by:[Table 1](https://arxiv.org/html/2608.20338#S1.T1.3.1.4.1),[§1](https://arxiv.org/html/2608.20338#S1.p2.1),[§3\.1](https://arxiv.org/html/2608.20338#S3.SS1.p3.1),[3rd item](https://arxiv.org/html/2608.20338#S5.I1.i3.p1.1)\.
- LLM\-LAT \(2024\)LLM\-LATHarmful Dataset\.Note:[https://huggingface\.co/datasets/LLM\-LAT/harmful\-dataset](https://huggingface.co/datasets/LLM-LAT/harmful-dataset)Accessed: 2026\-04Cited by:[1st item](https://arxiv.org/html/2608.20338#S4.I1.i1.p1.1)\.
- Mainiet al\.\(2024\)P\. Maini, Z\. Feng, A\. Schwarzschild, Z\. C\. Lipton, and J\. Z\. KolterTOFU: a task of fictitious unlearning for llms\.External Links:2401\.06121,[Link](https://arxiv.org/abs/2401.06121)Cited by:[Table 1](https://arxiv.org/html/2608.20338#S1.T1.3.1.2.1),[§1](https://arxiv.org/html/2608.20338#S1.p2.1),[§3\.1](https://arxiv.org/html/2608.20338#S3.SS1.p3.1)\.
- Qiuet al\.\(2025\)R\. Qiu, J\. Tan, J\. Pu, H\. Wang, X\. Gao, and F\. SunA survey on unlearning in large language models\.arXiv preprint arXiv:2510\.25117\.Cited by:[§1](https://arxiv.org/html/2608.20338#S1.p1.1)\.
- Quet al\.\(2024\)Y\. Qu, M\. Ding, N\. Sun, K\. Thilakarathna, T\. Zhu, and D\. NiyatoThe frontier of data erasure: machine unlearning for large language models\.arXiv preprint arXiv:2403\.15779\.Cited by:[§1](https://arxiv.org/html/2608.20338#S1.p1.1)\.
- Shiet al\.\(2024\)W\. Shi, J\. Lee, Y\. Huang, S\. Malladi, J\. Zhao, A\. Holtzman, D\. Liu, L\. Zettlemoyer, N\. A\. Smith, and C\. ZhangMUSE: machine unlearning six\-way evaluation for language models\.External Links:2407\.06460,[Link](https://arxiv.org/abs/2407.06460)Cited by:[Table 1](https://arxiv.org/html/2608.20338#S1.T1.3.1.3.1),[§1](https://arxiv.org/html/2608.20338#S1.p2.1),[§3\.1](https://arxiv.org/html/2608.20338#S3.SS1.p3.1),[1st item](https://arxiv.org/html/2608.20338#S5.I1.i1.p1.1)\.
- Weidingeret al\.\(2022\)L\. Weidinger, J\. Uesato, M\. Rauh, C\. Griffin, P\. Huang, J\. Mellor, A\. Glaese, M\. Cheng, B\. Balle, A\. Kasirzadeh, C\. Biles, S\. Brown, Z\. Kenton, W\. Hawkins, T\. Stepleton, A\. Birhane, L\. A\. Hendricks, L\. Rimell, W\. Isaac, J\. Haas, S\. Legassick, G\. Irving, and I\. GabrielTaxonomy of risks posed by language models\.InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency,FAccT ’22,New York, NY, USA,pp\. 214–229\.External Links:ISBN 9781450393522,[Link](https://doi.org/10.1145/3531146.3533088),[Document](https://dx.doi.org/10.1145/3531146.3533088)Cited by:[§4\.1](https://arxiv.org/html/2608.20338#S4.SS1.p1.1)\.
- Yaoet al\.\(2024\)J\. Yao, E\. Chien, M\. Du, X\. Niu, T\. Wang, Z\. Cheng, and X\. YueMachine unlearning of pre\-trained large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 8403–8419\.External Links:[Link](https://aclanthology.org/2024.acl-long.457/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.457)Cited by:[§2](https://arxiv.org/html/2608.20338#S2.p1.1)\.

## Appendix ADataset Construction Details

### A\.1Concept Tagging and Filtering

We identify dual\-use concepts from harmful prompts using a GPT\-5\-based classifier under manual supervision\. The model is instructed to extract high\-level concepts that can plausibly appear in both benign and harmful contexts, while filtering out inherently malicious or overly narrow activities \(e\.g\., explicit bioterrorism or one\-off exploits\)\. Each instance is mapped to a tuple\(prompt,concept,harmful response\)\(\\text\{prompt\},\\text\{concept\},\\text\{harmful response\}\)\.

To improve concept quality and consistency, the initial GPT\-5\-based tagging prompt was iteratively refined by the authors through batch\-level analysis of 100 samples per iteration\. Two main refinements were introduced: adding representative examples and explicitly instructing the model that concepts for which even seeking information may correspond to harmful objectives should not be considered dual\-use\. Following prompt refinement, 15% \(approximately 1,000\) of tagged instances were independently annotated by two graduate\-level annotators \(one MS student in Computer Science and one MS student in Cybersecurity\) for validation\. The annotators followed the guideline:

> Given a harmful prompt, response, and extracted concept, determine whether the concept represents a general capability that can plausibly have both beneficial and harmful applications\. Label as \(i\) Accept if the concept is dual\-use, \(ii\) Reject if it is inherently harmful or too narrow, or \(iii\) Unclear if further discussion is required\.

The annotation achieved Cohen’sκ=0\.81\\kappa=0\.81\. Disagreements were resolved through discussion between the annotators, and examples remaining as ’Unclear’ were discarded\. The full, iteratively refined tagging prompt is provided in Figure[5](https://arxiv.org/html/2608.20338#A1.F5)\.

### A\.2Concept Aggregation and Dataset Formation

Extracted concepts are aggregated to form a consistent and interpretable concept space\. The main goal is to merge low\-frequency and semantically overlapping concepts into broader parent categories, ensuring sufficient coverage per concept while avoiding fragmentation\.

Concept aggregation was performed through a two\-pass annotation process\. In the first pass, the same graduate\-level annotators were asked to identify candidate concept groups based on the following instructions:

> Mark all possible samples which can be grouped under a possible concept at a higher abstraction level such that the current concepts \(i\) represent the same underlying capability and harmful intent, \(ii\) differ only due to implementation details or attack variants, and \(iii\) can plausibly share similar benign counterparts\.

In the second pass, the annotators assigned descriptive parent labels to the identified groups through a real\-time Zoom call\. Final concept categories were refined through discussion, retaining candidate groups for which more than 90% of the same samples were marked for grouping by both annotators\. Samples associated with unresolved annotation disagreements and candidate groups containing only 1–2 samples were discarded\. This process reduced the initial 6,732 tagged instances to 2,583 validated harmful instances spanning 68 dual\-use concepts\. Each harmful instance was then paired with one benign counterpart, resulting in the final 5,166\-instance benchmark\.

The resulting harmful instances constitute the forget set𝒟f\\mathcal\{D\}\_\{f\}, where each example represents a specific harmful use of a broader dual\-use concept\. The final concept distribution is moderately long\-tailed, with a few dominant categories \(e\.g\., cybersecurity and fraud\) and a wide range of lower\-frequency concepts\.

### A\.3Benign Counterpart Generation

For each instance in𝒟f\\mathcal\{D\}\_\{f\}, we generate a corresponding benign instance to form the retain set𝒟r\\mathcal\{D\}\_\{r\}using GPT\-5 under manual supervision\. Rather than directly rewriting the original prompt, the model is instructed to first analyze the harmful prompt–response pair, identify the underlying dual\-use concept and fine\-grained themes, and then reframe the task into a benign query grounded in the same conceptual space\.

Specifically, the model:

- •Extracts the core concepts and fine\-grained themes from the prompt and harmful response,
- •Constructs a concise benign query that uses the same concepts in a safe and constructive context,
- •Generates a 180–250 word response to this query, focusing on informative, educational, or awareness\-driven content,
- •Preserves structural and stylistic similarity with the harmful response while ensuring complete removal of unsafe or sensitive content\.

This structured two\-step process \(query construction followed by response generation\) ensures that benign instances remain closely aligned with their harmful counterparts in terms of concept usage and expression, differing primarily in intent\. The generation prompt is provided in Figure[6](https://arxiv.org/html/2608.20338#A1.F6)\.

To also ensure that the generated retain set examples preserved the intended concept while removing harmful intent, 15% \( 375\) of generated benign instances were independently reviewed by the same two annotators\. The validation guideline was:

> Given a harmful\-benign pair, verify whether the benign example \(i\) preserves the same underlying concept, \(ii\) represents a constructive or educational use case, \(iii\) removes all harmful instructions or intent, and \(iv\) remains coherent and aligned with the original context\. Assign a score of 1 only in case of agreement with all the above points, else 0\.

This achieved Cohen’sκ=0\.88\\kappa=0\.88\. Examples with disagreements were discussed\. Cases failing validation were regenerated using the same controlled generation procedure\.

### A\.4Query Set Construction

In addition to𝒟f\\mathcal\{D\}\_\{f\}and𝒟r\\mathcal\{D\}\_\{r\}, we construct corresponding query sets𝒬f\\mathcal\{Q\}\_\{f\}and𝒬r\\mathcal\{Q\}\_\{r\}for evaluation\. Each query is generated by conditioning on its paired response and associated dual\-use concept, with the objective of eliciting a semantically similar response without relying on direct lexical overlap using GPT\-5, using the prompt shown in Figure[7](https://arxiv.org/html/2608.20338#A1.F7)\.

Harmful queries in𝒬f\\mathcal\{Q\}\_\{f\}are constructed to probe unsafe behavior, while benign queries in𝒬r\\mathcal\{Q\}\_\{r\}target constructive usage of the same concepts\. The size of each query set matches the corresponding training split, enabling evaluation of both memorization and generalization under controlled conceptual alignment\.

### A\.5Dataset Statistics

The final dataset consists of 5,166 instances, evenly split between harmful and benign usage\. Harmful instances form the forget set𝒟f\\mathcal\{D\}\_\{f\}, while benign instances form the retain set𝒟r\\mathcal\{D\}\_\{r\}\.

The concept distribution is skewed, with high\-frequency categories such as cybersecurity \(302 instances\), social engineering \(219\), and disinformation \(103\), alongside a long tail of less frequent concepts\. Detailed statistics are provided in Table[6](https://arxiv.org/html/2608.20338#A1.T6)\.

### A\.6Examples

We provide representative examples of harmful and benign pairs in Table[7](https://arxiv.org/html/2608.20338#A1.T7)\. The harmful and benign queries form the evaluation query sets𝒬f\\mathcal\{Q\}\_\{f\}and𝒬r\\mathcal\{Q\}\_\{r\}, respectively, while their corresponding responses populate the forget set𝒟f\\mathcal\{D\}\_\{f\}and retain set𝒟r\\mathcal\{D\}\_\{r\}\. These examples highlight how the same underlying concept is expressed in both unsafe and constructive contexts while maintaining similar topical structure\.

Table 6:Summary statistics for theConceptGuarddataset\.![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_P1.png)Figure 5:Prompt used to identify dual\-use concept usage and tag responses to the concepts![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_P2.png)Figure 6:Prompt used to generate benign text counterparts from given dual\-use concepts resembling harmful response text![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_P3.png)Figure 7:Prompt used to generate query set for evaluationTable 7:Representative dataset examples showing paired harmful and benign queries \(𝒬f\\mathcal\{Q\}\_\{f\},𝒬r\\mathcal\{Q\}\_\{r\}\) and their corresponding responses \(𝒟f\\mathcal\{D\}\_\{f\},𝒟r\\mathcal\{D\}\_\{r\}\) for the same concept\.

## Appendix BEvaluation Protocol

We evaluate model behavior using LLM\-based scoring functions for harmfulness and helpfulness, enabling fine\-grained, context\-sensitive assessment beyond exact\-match or lexical metrics\. Given a modelf⁡\(⋅\)f\(\\cdot\), responses to the harmful and benign query sets \(𝒬f\\mathcal\{Q\}\_\{f\}and𝒬r\\mathcal\{Q\}\_\{r\}\) are generated under identical decoding settings and evaluated using two independent LLM judges, with one judge dedicated to harmfulness and the other to helpfulness\.

Specifically, each response is evaluated along two axes: \(i\) harmfulness \(HarmScore\), which measures the extent to which a response enables, endorses, or meaningfully contributes to unsafe behavior, and \(ii\) helpfulness \(HelpScore\), which captures the relevance, correctness, and utility of responses in benign contexts\. Each judge is provided with a dedicated instruction\-driven scoring prompt and produces a final score in the range\[0,1\]\[0,1\]\. Higher HarmScore indicates greater harmfulness, whereas higher HelpScore indicates greater helpfulness\. These dimensions are evaluated independently to explicitly capture the trade\-off between suppressing harmful behavior and preserving useful knowledge\. The detailed scoring criteria and rubrics are provided in Figures[8](https://arxiv.org/html/2608.20338#A2.F8)and[9](https://arxiv.org/html/2608.20338#A2.F9)\.

The final evaluation protocol uses zero\-shot, rubric\-based judge prompts without access to ground\-truth labels or contextual examples, requiring the judge to assess each response based solely on the query, response, and specified scoring criteria\. This setting reflects realistic evaluation scenarios in which the quality and safety of a response must be inferred from its context rather than matched against a predefined reference answer\.

To assess the reliability of the judge\-based metrics, we conducted human validation using the same two graduate\-level annotators described in Section[A](https://arxiv.org/html/2608.20338#A1)\. The same set of 200 judge\-scored responses, consisting of 100 responses to harmful queries and 100 responses to benign queries, was independently evaluated by both annotators using the same scoring criteria as the corresponding LLM judges\. In the initial validation pass, agreement between human annotations and the LLM judges was limited, with Cohen’sκ\\kappaaveraging 0\.44 across the two annotators\. Discussion with the annotators indicated that the initial evaluation prompts included contextual examples that could introduce ambiguity, as judges occasionally appeared to assign scores based on similarity to the provided examples rather than the quality of the response itself\.

We therefore revised both judge prompts to use explicit rubric\-based scoring instructions without contextual examples and repeated the human validation\. Following this refinement, agreement increased to an average Cohen’sκ\\kappaof 0\.72 across the two annotators, indicating substantially improved alignment between human and automated evaluations\. Specifically, Cohen’sκ\\kappawas computed separately between each annotator and the corresponding LLM judge and then averaged across the two annotators\.

Thus, the validated protocol allows us to directly quantify contextual separation in model behavior by comparing HarmScore and HelpScore across complementary query sets, providing a principled basis for analyzing concept\-level unlearning performance\.

![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_P_Harm.png)Figure 8:Prompt used to evaluate the harmfulness of responses![Refer to caption](https://arxiv.org/html/2608.20338v1/media/Unlearn_P_Help.png)Figure 9:Prompt used to evaluate the helpfulness of responses
## Appendix CImplementation Details

We provide key training details to ensure reproducibility\. All experiments were conducted on 2 NVIDIA RTX 4090 GPUs \(24GB each\) with 128GB CPU memory\. We use bfloat16 precision throughout\. Unless otherwise specified, all hyperparameters follow defaults\. We use the Open Unlearning\[[4](https://arxiv.org/html/2608.20338#bib.bib6)\]framework as the base for our code for ease of use and integration\.

### C\.1Fine\-tuning Setup

Both Qwen\-2\.5\-3B\-Instruct and Llama\-3\.1\-8B\-Instruct are first fine\-tuned on𝒟f∪𝒟r\\mathcal\{D\}\_\{f\}\\cup\\mathcal\{D\}\_\{r\}, after which unlearning methods are applied\. The fine\-tuning configuration is shared across models and reported in Table[8](https://arxiv.org/html/2608.20338#A3.T8)\.

Table 8:Shared fine\-tuning hyperparameters for all models
### C\.2Unlearning Hyperparameters

We report method\-specific hyperparameters in Table[9](https://arxiv.org/html/2608.20338#A3.T9)\. All methods inherit the fine\-tuning configuration unless explicitly overridden\.

Table 9:Method\-specific hyperparameters for unlearning

Similar Articles

ContextGuard: Structured Self-Auditing for Context Learning in Language Models

arXiv cs.CL

Introduces ContextGuard, a structured self-auditing framework that improves LLM context learning by decomposing model self-assessment into confirmed and uncertain categories and applying targeted revisions, achieving a task-solving rate increase from 9.64% to 13.85% on Qwen3.5-4B on the CL-Bench benchmark.

MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs

arXiv cs.AI

MLUBench is a large-scale benchmark for lifelong unlearning in multimodal large language models (MLLMs), featuring 127 entities across 9 classes. The paper identifies that existing unlearning methods suffer from cumulative degradation and proposes LUMoE to mitigate this, showing significant improvements.

Model Unlearning Objectives Vary for Distinct Language Functions

arXiv cs.CL

The paper argues that unlearning in LLMs should be goal-dependent, proposing a cosine-based meta-learned variant of RMU for dangerous knowledge and a multi-layer objective with probe directions for toxicity, achieving strong results across four 7-8B models.