理解提示模板在知识蒸馏安全对齐中的作用

arXiv cs.CL 论文

摘要

本文分析了在知识蒸馏过程中,提示模板的选择如何影响学生大语言模型的安全对齐,发现聊天模板相比非聊天模板在多个模型和基准测试中会导致更大的性能下降。

arXiv:2609.30802v1 Announce Type: new Abstract: Prior research has demonstrated that the choice of prompt template during Supervised Fine-Tuning (SFT) significantly impacts the robustness of safety alignment afterwards. However, the influence of template selection during Knowledge Distillation (KD) from teacher to student remains largely unexplored. Thus, we fill this gap by analyzing how different template configurations influence the pre-existing safety alignment of the student. We observe a significant degradation of safety alignment present in the aligned base instruct-tuned model. Specifically, we find that utilizing chat templates renders the model more compliant with harmful queries compared to a non-chat template. These findings are consistent across three models: LLaMA, Gemma and Qwen model families and are evaluated across multiple safety benchmarks. We further show that using a non-chat template during distillation better preserves the base student's internal representations, while chat template distillation induces a larger representational shift. Code: https://github.com/anjilab/role-of-prompt-template-in-kd
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:43

# Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment
Source: [https://arxiv.org/html/2609.30802](https://arxiv.org/html/2609.30802)
Anjila BudathokiManish DhakalAffiliation:University of Tennessee, KnoxvilleAffiliation:University of Tennessee, KnoxvilleBenjamin M\. AmpelAffiliation:Georgia State University\{abudatho, mdhakal1\}@vols\.utk\.edu, bampel@gsu\.edu, yding@utk\.eduAffiliation:Georgia State University\{abudatho, mdhakal1\}@vols\.utk\.edu, bampel@gsu\.edu, yding@utk\.eduYi DingAffiliation:University of Tennessee, KnoxvilleAffiliation:University of Tennessee, Knoxville

###### Abstract

Prior research has demonstrated that the choice of prompt template during Supervised Fine\-Tuning \(SFT\) significantly impacts the robustness of safety alignment afterwards\. However, the influence of template selection during Knowledge Distillation \(KD\) from teacher to student remains largely unexplored\. Thus, we fill this gap by analyzing how different template configurations influence the pre\-existing safety alignment of the student\. We observe a significant degradation of safety alignment present in the aligned base instruct\-tuned model\. Specifically, we find that utilizing chat templates renders the model more compliant with harmful queries compared to a non\-chat template\. These findings are consistent across three models: LLaMA, Gemma and Qwen model families and are evaluated across multiple safety benchmarks\. We further show that using a non\-chat template during distillation better preserves the base student’s internal representations, while chat template distillation induces a larger representational shift\.111Code:[https://github\.com/anjilab/role\-of\-prompt\-template\-in\-kd](https://github.com/anjilab/role-of-prompt-template-in-kd)

Warning: this paper includes examples that may be offensive or harmful\.

## 1Introduction

Figure 1:Overview of the experimental pipeline\.We distill student models on benign instruction\-following data under two formatting conditions: non\-chat and chat\. The pipeline compares how these training\-time prompt templates affect the resulting student model during downstream safety and utility evaluation\.Knowledge Distillation \(KD\)[Hinton et al\. \(2015\)](https://arxiv.org/html/2609.30802#bib.bib5)is an effective technique for compressing large models into smaller, efficient student models while retaining performance\. In the context of Large Language Models \(LLMs\), KD is commonly used to obtain compact models that are easier to deploy under computational constraints\. As these models are increasingly integrated into practical applications such as healthcare and autonomous systems\([MohiEldeen Alabbasy et al\., 2023](https://arxiv.org/html/2609.30802#bib.bib21);[Agand, 2024](https://arxiv.org/html/2609.30802#bib.bib22)\), considerations of robustness and safety become increasingly important\. In particular, deployed distilled models are expected to exhibit safe and harmless behavior, including refusing harmful, malicious, or policy\-violating requests\. To encourage such behavior, modern instruction\-tuned models typically undergo alignment processes that embed refusal capabilities before any downstream adaptation[Askell et al\. \(2021\)](https://arxiv.org/html/2609.30802#bib.bib9);[Ouyang et al\. \(2022\)](https://arxiv.org/html/2609.30802#bib.bib8)\. A key open question is how the safety alignment of student models changes during further training via KD, and which training\-time factors govern its preservation or degradation\.

In particular, the choice of prompt template[Lyu et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib13)during training is one such factor\. Prompt templates specify how inputs are formatted for the model, often as strings with placeholders filled by user queries, instructions, and assistant responses\. Modern LLMs[Touvron et al\. \(2023\)](https://arxiv.org/html/2609.30802#bib.bib3);[Jiang et al\. \(2023\)](https://arxiv.org/html/2609.30802#bib.bib19)primarily utilizechat templates, which standardize user interactions via specific control tokens \(e\.g\.,<\|user\|\>,\[INST\]\); we contrast these withnon\-chat templates, which present the same content without these conversational control tokens\. Hereafter, we usechat templatesto refer to templates that wrap inputs with special conversational control tokens, andnon\-chat templatesto refer to task\-style formats that omit these tokens\. Prior work shows that such formatting choices can substantially affect safety behavior during SFT\([Lyu et al\., 2024](https://arxiv.org/html/2609.30802#bib.bib13);[Jiang et al\., 2025](https://arxiv.org/html/2609.30802#bib.bib17);[Wang et al\., 2025b](https://arxiv.org/html/2609.30802#bib.bib18)\); however, their role in KD remains empirically underexplored\.

Unlike SFT, KD exposes the student to teacher\-generated soft targets, an additional signal that may interact with template formatting in ways that distinctly affect the student’s safety representations\. This raises an important question about the interplay between distillation, template formatting, and safety retention:

How do chat and non\-chat prompt templates affect the preservation of safety\-aligned refusal behavior when student models are distilled on benign downstream tasks?

In this work, we answer this question by systematically investigating how the format of prompt templates during training shapes the distilled student model’s susceptibility to harmful queries[Qi et al\. \(2025\)](https://arxiv.org/html/2609.30802#bib.bib23);[Arditi et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib24)\. We distill student models on benign instruction\-following data under two controlled template conditions: chat and non\-chat, and evaluate all models using the standard chat template at inference, ensuring that any observed safety differences reflect training\-time choices alone\. Beyond output\-level evaluation, we conduct a mechanistic analysis to examine whether chat\-template KD shifts the student’s internal refusal representations or whether the behavioral gap is a surface\-level artifact\. We further studyprompt template mixingby varying the proportion of chat versus non\-chat samples during training to characterize how safety and utility scale with chat template exposure\. The overall pipeline is shown in Figure[1](https://arxiv.org/html/2609.30802#S1.F1)\.

Our key findings are summarized as follows:

- •Knowledge distillation on benign dataset can erode the pre\-existing safety alignment of aligned student models across model families\.
- •Prompt formatting modulates the severity of this degradation: using a chat template during KD consistently leads to a higher Attack Success Rate \(ASR\) than non\-chat KD under identical training data\.
- •The student\-side training template, not the teacher’s output distribution, is the primary driver: using the chat template only for the teacher causes minimal safety loss, whereas using it for the student produces consistent regression\.
- •The degradation is not limited to output behavior: chat\-template KD shifts the student’s internal refusal direction away from the base model and reduces its ability to separate harmful from harmless prompts, whereas non\-chat KD largely preserves both\.
- •Safety is more sensitive to chat template exposure than utility\. As the proportion of chat template samples increases, ASR generally rises, indicating amplified safety degradation, whereas utility improves only modestly\.

## 2Related work

##### Knowledge Distillation\.

[Hinton et al\. \(2015\)](https://arxiv.org/html/2609.30802#bib.bib5)originally demonstrated that soft targets \(i\.e\., a full probability distribution\) encode richer information than hard labels \(i\.e\., a discrete label\)\. This enabled the efficient transfer of knowledge from a large deep neural network to smaller ones\. Recently, research has shifted towards distilling the knowledge of LLMs into more compact student models[Gu et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib7);[Ko et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib6);[Wang et al\. \(2025a\)](https://arxiv.org/html/2609.30802#bib.bib32)\. Notably,[Gu et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib7)introduced a distillation objective for generative models \(adopted by many subsequent methods\) which introduced reverse KL divergencefor improved KD in generative LLMs\. However, despite the advancement in task\-centric KD that prioritizes task utility \(e\.g\., preserving accuracy or fluency\), how these processes impact safety alignment is currently under\-studied\.

##### Safety Alignment in LLMs\.

To align LLMs with human values, models often adopt a two\-stage post\-training pipeline of SFT and alignment\-tuning \(e\.g\., Reinforcement Learning from Human Feedback, Direct Preference Optimization\)[Ouyang et al\. \(2022\)](https://arxiv.org/html/2609.30802#bib.bib8);[Bai et al\. \(2022\)](https://arxiv.org/html/2609.30802#bib.bib10);[Rafailov et al\. \(2023\)](https://arxiv.org/html/2609.30802#bib.bib11)\. These methods have demonstrated that careful alignment can embed the ability to reject harmful instructions \(known as safety guardrails\) within an LLM\. Prior work has also employed KD as a mechanism to enforce safety alignment in student models[Yang et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib14)\. However, literature has found that several vulnerabilities exist in the safety guardrails of open\-access LLMs\([Yi et al\., 2024](https://arxiv.org/html/2609.30802#bib.bib12)\)\. These vulnerabilities lead to jailbreak attacks, which allow malicious users to use the LLM in unintended ways\([Yi et al\., 2024](https://arxiv.org/html/2609.30802#bib.bib12)\)\. It is unclear if distilling models on benign downstream tasks can further degrade these safety guardrails in previously aligned models\.

##### Prompt Templates\.

Prompt templates, which determine how instructions, user inputs, and assistant responses are serialized before being passed to a model, are often treated as implementation details, but these formatting choices meaningfully shape model behavior and safety\. During instruction fine\-tuning, chat templates can reduce context awareness, influencing how models attend to input\([Wang et al\., 2025b](https://arxiv.org/html/2609.30802#bib.bib18)\)\. They also affect safety: task\-style fine\-tuning and safety\-oriented testing better preserve safe behavior\([Lyu et al\., 2024](https://arxiv.org/html/2609.30802#bib.bib13)\), while certain chat template designs induce unexpected behavioral failures\([Jiang et al\., 2025](https://arxiv.org/html/2609.30802#bib.bib17)\)\. Similar vulnerabilities also appear in VLMs, where Role\-Modality Attacks exploit dialogue roles and modality placement\([Shayegani et al\., 2026](https://arxiv.org/html/2609.30802#bib.bib30)\)\. Despite these findings, the role of prompt templates in knowledge distillation remains underexplored; we address this gap by systematically comparing chat and non\-chat templates during benign KD and measuring their impact on both utility and safety\.

## 3Methodology

In this section, we present our approach for analyzing the impact of prompt templates on safety alignment during Knowledge Distillation \(KD\)\. Our pipeline, illustrated in Figure[1](https://arxiv.org/html/2609.30802#S1.F1), consists of two stages: \(1\) constructing instruction datasets using distinct prompt formatting strategies \(chat vs\. non\-chat\), and \(2\) fine\-tuning a student model using a balanced KD objective\.

### 3\.1Prompt Formatting

To examine the role of structural cues during distillation, we design two controlled formatting conditions for the instruction dataset𝒟\\mathcal\{D\}\. While modern instruction\-tuned models typically require specific chat templates \(e\.g\.,<\|begin\_of\_text\|\>,<\|start\_header\_id\|\>\) to maintain state and role\([Hugging Face Team, 2025](https://arxiv.org/html/2609.30802#bib.bib20)\), it is unclear if these tokens aid or hinder the transfer of safety representations during KD\.

We use standard definition of two template configurations, as described below \(see e\.g\. in Table[9](https://arxiv.org/html/2609.30802#A1.T9)\):

- •Chat Template:We use each model family’s native chat template, including all special control tokens, as the chat\-format baseline for instruction following\.
- •Non\-chat Template:We strip all model\-specific control tokens, presenting the input as raw text\. This isolates the semantic content of the instruction from the structural priors enforced by the template\.

Importantly, regardless of the training template, all models are evaluated using the standard chat template at inference to reflect real\-world deployment conditions \(details in §[3\.3\.3](https://arxiv.org/html/2609.30802#S3.SS3.SSS3.Px3)\)

### 3\.2Distillation Objective

Given the formatted dataset𝒟=\{\(x,y\)\}\\mathcal\{D\}=\\\{\(x,y\)\\\}, we fine\-tune the student modelqθq\_\{\\theta\}by combining a standard supervised learning objective with Knowledge Distillation from the teacherpp\. The total loss function balances the ground\-truth alignment with the transfer of the teacher’s probability distribution:

ℒ=\(1−α\)​ℒC​E\+α​ℒK​D\\mathcal\{L\}=\(1\-\\alpha\)\\mathcal\{L\}\_\{CE\}\+\\alpha\\mathcal\{L\}\_\{KD\}
The supervised component,ℒC​E\\mathcal\{L\}\_\{CE\}, enforces alignment with the gold reference tokens, whileℒK​D\\mathcal\{L\}\_\{KD\}minimizes the forward KL divergence between the teacher and student logits:

ℒC​E\\displaystyle\\mathcal\{L\}\_\{CE\}=𝔼\(x,y\)∼𝒟​\[−log⁡qθ​\(y\|x\)\]\\displaystyle=\\mathbb\{E\}\_\{\(x,y\)\\sim\\mathcal\{D\}\}\\left\[\-\\log q\_\{\\theta\}\(y\|x\)\\right\]ℒK​D\\displaystyle\\mathcal\{L\}\_\{KD\}=𝔼x∼𝒟,y∼p\(⋅\|x\)\[logp\(y\|x\)−logqθ\(y\|x\)\]\\displaystyle=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\},y\\sim p\(\\cdot\|x\)\}\\left\[\\log p\(y\|x\)\-\\log q\_\{\\theta\}\(y\|x\)\\right\]
We setα=0\.5\\alpha=0\.5to assign equal weight to task performance and knowledge transfer\. Although our primary experiments use forward KL as the standard distillation objective, we additionally evaluate reverse KL and combined FKL\+RKL objectives to assess whether the observed template effects generalize across distillation objectives \(Appendix[C](https://arxiv.org/html/2609.30802#A3)\)\.

### 3\.3Experimental Setup

In this section, we detail our experimental setup for investigating how prompt template choice during distillation influences the preservation of safety alignment of the student models\.

#### 3\.3\.1Models\.

We employ three widely used open\-weight model families: LLaMA\-3[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib16), Gemma\-2[Team et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib25), and Qwen\-2\.5[Qwen et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib31)\. Following standard KD protocols[Gu et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib7);[Ko et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib6), we use larger Instruct variants asteachers\(ℳT\\mathcal\{M\}\_\{T\}\) and smaller Instruct variants asstudents\(ℳS\\mathcal\{M\}\_\{S\}\):LLaMA\-3\.1\-8B\-Instruct→\\rightarrowLLaMA\-3\.2\-3B\-Instruct,Gemma\-2\-9B\-IT→\\rightarrowGemma\-2\-2B\-IT, andQwen2\.5\-7B\-Instruct→\\rightarrowQwen2\.5\-3B\-Instruct\. These model pairs were selected due to their rigorous pre\-training filtration and post\-training safety alignment, providing a strong baseline for measuring safety degradation\.

#### 3\.3\.2Datasets\.

##### Instruction Following\.

We perform distillation with thedatabricks\-dolly\-15kdataset[Gu et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib7), a standard benchmark for instruction following\.

##### Safety Evaluation\.

To measure safety alignment, we evaluate on four adversarial benchmarks: AdvBench[Zou et al\. \(2023\)](https://arxiv.org/html/2609.30802#bib.bib27)and JailbreakBench[Chao et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib28)for harmful\-instruction attacks, and HarmBench[Mazeika et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib2)and SORRY\-Bench[Xie et al\. \(2025\)](https://arxiv.org/html/2609.30802#bib.bib1)for broader harmful\-behavior coverage\. These datasets are used for evaluation only; please refer to Appendix[A\.4](https://arxiv.org/html/2609.30802#A1.SS4)for a detailed description of each\.

#### 3\.3\.3Training and Evaluation

##### Training

We fine\-tune all student models using LoRA[Hu et al\. \(2022\)](https://arxiv.org/html/2609.30802#bib.bib29)rankr=8r=8\. The training is conducted on a single node equipped with 4×\\timesNVIDIA RTX 4090 GPUs\. For a comprehensive overview of hyperparameters, including learning rates and batch sizes, please refer to the Appendix[A\.1](https://arxiv.org/html/2609.30802#A1.SS1)\.

##### Evaluation Setting\.

We assess model performance along two dimensions: task utility and safety alignment\. For task utility, we report the average ROUGE\-L score across the four general\-domain instruction\-following sets\. For safety alignment, we compute Attack Success Rate \(ASR\) on AdvBench and JailbreakBench using the WalledEval framework[Gupta et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib26)withLLaMA\-3\-Guard\-8Bas the judge, and evaluate HarmBench and SORRY\-Bench using their respective provided judges\.

##### Inference Setting\.

Crucially, during evaluation, we strictly adhere to the standard chat template for all models \(including those trained with non\-chat templates\)\. This follows the recommended inference\-time usage of modern instruction\-tuned LLMs[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib16);[Team et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib25)and ensures that our evaluation reflects the practical safety of deployed models\. Therefore, any observed safety differences across conditions can be attributed to the training\-time template choice, rather than to differences in the evaluation format\.

## 4Results

Table 1:Student\-model safety and utility under different prompt templates\.We evaluate how prompt templates used during knowledge distillation affect the resulting student model’s safety and utility\. Safety is measured using Attack Success Rate \(ASR%\), where lower values indicate safer behavior\. Utility reports the average ROUGE\-L score across four instruction\-following benchmarks: Dolly, SelfInst, S\-NI, and Vicuna\. The subscripted deltas show the change relative to the corresponding student baseline \(ℳS\\mathcal\{M\}\_\{S\}\) within each model family\. Red deltas indicate increased ASR, corresponding to worse safety, while green deltas indicate reduced ASR, corresponding to improved safety\. To emphasize the KD template effect, bold values are used only in the KD safety columns and denote the higher \(less safe\) ASR between chat and non\-chat KD\.We analyze the safety behavior of student models after knowledge distillation \(KD\) under two training\-template conditions: \(1\) chat template and \(2\) non\-chat template\. During evaluation, all models are tested using the standard chat template, reflecting the expected deployment setting\.

### 4\.1Distillation generally erodes safety\.

We observe that standard KD on benign instruction\-following data degrades the safety alignment in the student models\. As shown in Table[1](https://arxiv.org/html/2609.30802#S4.T1), all three distilled student models show increased ASR on most safety benchmarks relative to their corresponding base student models\. In some cases, the vulnerability of the student models has almost tripled with maximal degradation \(HarmBench for LLaMA and SORRY\-Bench for Gemma\)\. We further show in Appendix[C](https://arxiv.org/html/2609.30802#A3)that the same template\-driven safety degradation persists under alternative distillation objectives\.

### 4\.2Template Choice Shapes Safety Degradation

Our main finding is that the prompt template used during KD strongly affects safety preservation\. When the student is distilled using the standard chat template, the resulting model shows larger increases in ASR than with the non\-chat template\. For example, on SORRY\-Bench, the Gemma student shows a\+35\.32%ASR increase relative to its base student model\. Similarly, on HarmBench, we observe a\+27\.57%increase for LLaMA and a\+13\.12%increase for Qwen\. These results suggest that chat\-template KD is associated with greater erosion of pre\-existing refusal behavior\.

While prior work has shown that SFT can degrade safety alignment[Lyu et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib13);[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.30802#bib.bib17), we show that similar template sensitivity also emerges in KD, even when the distillation data are benign\. Importantly, the effect is substantially weaker with non\-chat KD\. On Gemma, SORRY\-Bench ASR increases by only\+3\.14%, compared with\+35\.32%under chat\-template KD\. The same pattern holds across model families: LLaMA shows a smaller HarmBench increase, while Qwen remains near baseline on HarmBench \(\+0\.12%\) but degrades substantially under chat\-template KD \(\+13\.12%\)\.

These results suggest that using a non\-chat template during KD reduces safety degradation under standard chat\-template evaluation\. However, non\-chat KD does not fully preserve the original alignment\. For example, LLaMA still shows a\+17\.63%HarmBench ASR increase, indicating that distillation itself weakens refusal behavior\. Although non\-chat KD still improves utility over the base student model, its gains are smaller than those achieved by chat template KD\.

### 4\.3The safety\-utility trade\-off\.

Given that KD is primarily used to improve task utility, it is important to examine whether template choice also affects this dimension\. Our results reveal a clear trade\-off: KD improves instruction\-following under both template conditions, but larger utility gains coincide with larger ASR increases\. As shown in Table[1](https://arxiv.org/html/2609.30802#S4.T1), chat\-template KD achieves higher utility but also incurs the largest safety costs, whereas non\-chat KD yields smaller utility gains while better preserving safety\. Appendix[E](https://arxiv.org/html/2609.30802#A5)further shows that chat\-template KD accelerates this vulnerability earlier in training\.

One possible explanation is that the lower ASR of non\-chat KD simply reflects weaker instruction\-following rather than better safety preservation\. However, this is inconsistent with the data: for LLaMA, non\-chat KD achieves lower utility than SFT Chat \(24\.75 vs\. 29\.10\) yet shows higher HarmBench ASR \(29\.81% vs\. 22\.50%\)\. If reduced compliance explained the lower ASR, lower utility should correspond to lower ASR, which is not observed\. Section[5\.2](https://arxiv.org/html/2609.30802#S5.SS2)further shows that non\-chat KD better preserves refusal\-related representations, suggesting deeper mechanistic differences between the two template conditions\.

## 5Analysis

The observed safety degradation under chat\-template KD Prompts us to investigate three follow\-up questions: first, whether this degradation is driven primarily by the student’s training format or by the teacher’s output distribution \(§[5\.1](https://arxiv.org/html/2609.30802#S5.SS1)\); second, whether this behavioral gap corresponds to changes in the model’s internal refusal representations \(§[5\.2](https://arxiv.org/html/2609.30802#S5.SS2)\); and third, whether reducing the proportion of chat\-template examples during training can moderate this effect \(§[5\.3](https://arxiv.org/html/2609.30802#S5.SS3)\)\.

### 5\.1KD Prompt Template Drives the Degradation

Table 2:Decoupling student and teacher templates during KD\. Safety is measured using ASR \(%\), where lower is better\. The subscripted deltas show the change relative to the corresponding student baseline \(ℳS\\mathcal\{M\}\_\{S\}\)\.To determine whether the safety degradation observed in Section[1](https://arxiv.org/html/2609.30802#S4.T1)is inherited from the teacher’s output distribution or driven by the student’s training format, we decouple the two templates during KD\. As shown in Table[2](https://arxiv.org/html/2609.30802#S5.T2), using the chat template only for the teacher does not meaningfully degrade safety, whereas using it only for the student causes a small but consistent regression across both model families\. This suggests that the student\-side training format plays the dominant role, rather than the teacher\-side output format alone\. We attribute this asymmetry to the nature of each side’s influence: the student template directly shapes the input distribution over which gradients are computed, whereas the teacher template only determines the soft\-target distribution observed by the student\. As a result, the teacher\-side template provides a weaker signal and does not directly alter the student’s own representation learning\.

The strongest degradation appears when both the student and teacher use the chat template\. In this matched chat/chat setting, LLaMA HarmBench increases by\+27\.57%\+27\.57\\%and Gemma SORRY\-Bench increases by\+35\.32%\+35\.32\\%, far exceeding the changes observed when only one side uses the chat template\. These results indicate that the regression is not simply inherited from the teacher, but emerges when the student learns from a teacher distribution generated under the same chat\-based template\. Decoupling either side removes most of the safety loss, showing that matched chat\-template distillation is the primary source of the observed degradation\. Since the student\-side template emerges as the primary driver, we next examine whether this behavioral difference is also reflected in the model’s internal representation\.

### 5\.2Mechanistic Analysis

As Section[1](https://arxiv.org/html/2609.30802#S4.T1)shows that the student\-side template is the primary driver of safety degradation, we next examine whether this difference also appears in the model’s internal representations, and not just in output behavior, using the refusal direction\([Arditi et al\., 2024](https://arxiv.org/html/2609.30802#bib.bib24)\)as a diagnostic probe\. Specifically, we analyze whether knowledge distillation changes the refusal geometry of the original instruction\-tuned student model\.

##### Setup\.

We probe the model’s internal representation using therefusal direction\([Arditi et al\., 2024](https://arxiv.org/html/2609.30802#bib.bib24)\), a linear axis in activation space estimated from residual stream activations \(resid\_pre\) that separates harmful from harmless prompts in the base student\. If distillation preserves this axis, the model retains its internal distinction between harmful and harmless inputs; if the axis changes or becomes weaker, the behavioral safety gap reflects a deeper representational shift\. For each model and layerℓ\\ell, we estimate this direction from mean activation differences between harmful prompts \(SORRY\-Bench\) and harmless prompts \(Dolly\-15k\):

rℓ=𝔼⁡\[hℓharm\]−𝔼⁡\[hℓsafe\]‖𝔼⁡\[hℓharm\]−𝔼⁡\[hℓsafe\]‖2\.r\_\{\\ell\}=\\frac\{\\mathbb\{E\}\[h\_\{\\ell\}^\{\\text\{harm\}\}\]\-\\mathbb\{E\}\[h\_\{\\ell\}^\{\\text\{safe\}\}\]\}\{\\left\\\|\\mathbb\{E\}\[h\_\{\\ell\}^\{\\text\{harm\}\}\]\-\\mathbb\{E\}\[h\_\{\\ell\}^\{\\text\{safe\}\}\]\\right\\\|\_\{2\}\}\.\(1\)
We compute this direction for the base student, the chat template distilled student, and the non\-chat template distilled student\. We then track two complementary properties after distillation: whether the refusal direction remains aligned with the base student \(cosine similarity\), and whether it still separates harmful from harmless activations \(projection gap\):

CosSimℓ​\(M,MS\)=rℓM⋅rℓMS‖rℓM‖2​‖rℓMS‖2,\\mathrm\{CosSim\}\_\{\\ell\}\(M,M\_\{S\}\)=\\frac\{r\_\{\\ell\}^\{M\}\\cdot r\_\{\\ell\}^\{M\_\{S\}\}\}\{\\left\\\|r\_\{\\ell\}^\{M\}\\right\\\|\_\{2\}\\left\\\|r\_\{\\ell\}^\{M\_\{S\}\}\\right\\\|\_\{2\}\},\(2\)ΔℓM=𝔼x∈𝒟h​a​r​m​\[hℓM​\(x\)⋅rℓM\]−𝔼x∈𝒟s​a​f​e​\[hℓM​\(x\)⋅rℓM\]\.\\Delta\_\{\\ell\}^\{M\}=\\mathbb\{E\}\_\{x\\in\\mathcal\{D\}\_\{harm\}\}\\\!\\left\[h\_\{\\ell\}^\{M\}\(x\)\\cdot r\_\{\\ell\}^\{M\}\\right\]\-\\mathbb\{E\}\_\{x\\in\\mathcal\{D\}\_\{safe\}\}\\\!\\left\[h\_\{\\ell\}^\{M\}\(x\)\\cdot r\_\{\\ell\}^\{M\}\\right\]\.\(3\)
A cosine similarity close to 1\.0 indicates that the refusal direction is preserved, while lower values indicate rotation away from the original direction\. Furthermore, higher projection gaps indicate stronger internal separation between harmful and harmless prompts\. We additionally apply PCA to last\-token hidden states to visualize these shifts\. All models are evaluated using the same chat\-template inference format so that differences reflect the effect of distillation rather than inference\-time formatting\.

##### Findings\.

Table 3:Representation similarity and refusal\-projection gap across model families\.Cosine similarity is measured with respect to the refusal direction of the original instruction\-tuned student, and projection gap measures the harmful–harmless separation along this direction\.As shown in Table[3](https://arxiv.org/html/2609.30802#S5.T3), chat\-template KD causes larger changes in refusal geometry than non\-chat\-template KD across Gemma, LLaMA, and Qwen\. Across all three model families, non\-chat distillation better preserves the base model’s representation structure, while chat\-template distillation induces a larger representational shift and reduces the projection gap\. The chat\-template distilled models show lower cosine similarity to the original instruction\-tuned student’s refusal direction, indicating that the refusal direction is less preserved after chat\-template KD\. They also show a weaker projection gap, suggesting that harmful and harmless prompts become less separable along the refusal direction\. This pattern is consistent with the behavioral safety results: the same condition that produces higher attack success rate also produces larger shifts in refusal\-related representations\.

In contrast, non\-chat template KD better preserves both the direction and separability of the original refusal geometry\. This suggests that non\-chat template distillation interferes less with internal representations associated with refusal behavior\. The PCA visualization in Figure[2](https://arxiv.org/html/2609.30802#S5.F2)provides a complementary view of this effect, showing that chat template KD leads to greater changes in the harmful\-vs\-harmless representation structure\. Overall, these findings suggest that the safety degradation caused by chat template KD is not only a surface\-level decoding effect, but is also reflected in intermediate representations\.

We further perform small\-scale causal interventions on the learned refusal direction as a diagnostic test of the correlational findings\. The results are reported in Appendix[F](https://arxiv.org/html/2609.30802#A6), with an important caveat that the LLaMA ablation can impair output coherence and therefore its lower ASR should not be interpreted as improved safety\.

![Refer to caption](https://arxiv.org/html/2609.30802v1/combined_last_layer_all_families.png)Figure 2:Final\-layer representation shifts under different prompt templates\.PCA projections of final\-layer hidden states across three models show that non\-chat template KD remains closer to the base instruct\-tuned student, whereas chat\-template KD produces a larger shift in representation space\.Table 4:KD safety and utility under different chat template proportions\.Chat Prop\. denotes the proportion of chat\-template examples used during KD, where 0\.0 is pure non\-chat and 1\.0 is pure chat template\.

### 5\.3Template Mixing Provides Limited Mitigation

The preceding analyses show that using the chat template during KD drives behavioral safety degradation and shifts the model’s internal refusal representation\. We therefore examine whether this degradation depends on the amount of chat\-template exposure during distillation\. Rather than introducing additional safety\-specific supervision[Han et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib15);[Qi et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib4), we isolate the role of prompt formatting by mixing chat and non\-chat formatted samples within the same benign utility dataset\. This design allows us to test whether gradually reducing chat\-template exposure can moderate safety degradation while keeping the training data content fixed and avoiding explicit safety supervision\.

Specifically, we vary the proportion of chat template samples from 0 \(pure non\-chat\) to 1 \(pure chat\), where intermediate values represent mixed training\. A value of 0\.5 chat\-template proportion indicates that half of the training examples use the chat template, while the remaining half use the corresponding non\-chat prompt format\. As reported in Table[4](https://arxiv.org/html/2609.30802#S5.T4), utility increases as this proportion increases: for LLaMA, AVG utility rises from 24\.75 to roughly 29, and for Gemma from 23\.63 to roughly 30\. However, these gains are small relative to the concurrent rise in ASR\.

In contrast, safety is much more sensitive to the proportion of chat template samples\. For LLaMA, SORRY\-Bench increases from 27\.36 to 30\.86, while HarmBench rises sharply from 29\.81 to 39\.75\. A stronger trend appears for Gemma, where SORRY\-Bench increases from 19\.18 to 51\.36 and HarmBench from 14\.25 to 19\.56\. Importantly, prompt mixing does not reduce ASR below the pure non\-chat setting; it only attenuates degradation relative to full chat\-template KD\. These results suggest that prompt\-template mixing minimally affects utility but substantially increases harmful\-response susceptibility as chat\-template proportions rise\.

### 5\.4Robustness to Prompt Formatting

Our primary experiments use the standard chat template for all models, allowing controlled comparison across KD training formats\. To assess whether the observed degradation depends on the inference format or on sensitivity to specific control tokens within the chat template, we perform two robustness analyses: cross\-template evaluation, which changes the complete inference format, and token ablations, which selectively remove components of the chat template\.

Table 5:Cross\-template safety evaluation\.Attack success rate \(ASR%,↓\\downarrow\) under chat and non\-chat inference formats across the original instruction\-tuned student \(base\) and distilled models\. The subscripted deltas show the change relative to the corresponding student baseline\.##### Cross\-template evaluation\.

To test whether the observed degradation is caused by sensitivity to the inference template, we provide cross\-template evaluation in HarmBench and SORRYBench safety dataset as shown in Table[5](https://arxiv.org/html/2609.30802#S5.T5)\.

For each model family, we compare the base and the distilled models under the same inference template\. The degradation persists under non\-chat evaluation, with chat template KD exhibits higher ASR than the corresponding base model across all three model families and both benchmarks\. Although the magnitude of the degradation varies with the inference format, its direction remains consistent, indicating that the observed degradation in safety alignment is not limited to the standard chat template evaluation\.

Table 6:Token\-ablation safety evaluation\.HarmBench ASR after removing chat\-template components\. The subscripted deltas show the change relative to the corresponding student baseline\.
##### Token ablation\.

To examine whether the degradation is driven by specific control tokens in the chat template, we perform fine\-grained HarmBench ablations across all three model families\. We remove BOS/EOT tokens \(w/o BOS/EOT\), which mark sequence or turn boundaries; remove role markers \(w/o Role Tags\), which identify the user and assistant turns; and evaluate the harmful instruction alone without chat template tokens \(Raw Query\) as shown in Table[6](https://arxiv.org/html/2609.30802#S5.T6)\.

When BOS/EOT tokens are removed, KD\-Chat remains less safe than the corresponding base model across all three families, with ASR increases of\+9\.37\+9\.37,\+7\.69\+7\.69, and\+6\.00\+6\.00for LLaMA, Gemma, and Qwen, respectively\. This indicates that BOS/EOT tokens alone do not account for the increased ASR\. The role\-marker and raw\-query ablations show more model\-dependent behavior, as the base models themselves are highly sensitive to these input formats\. Thus, individual template components can modulate the magnitude of the effect, but no single control token component consistently explains the degradation across model families\.

## 6Discussion

The above analyses in section[5](https://arxiv.org/html/2609.30802#S5)show that chat template KD consistently leads to safety degradation than non\-chat KD, is accompanied by changes in refusal\-related representations, increases with chat template exposure, and remains evident across alternative inference formats and token ablations\. Therefore, these findings show that safety degradation during knowledge distillation is not only a consequence of optimizing on benign task data, but is also strongly shaped by the prompt template used during training\. In particular, chat template KD produces faster and larger increases in ASR than non\-chat KD, indicating that prompt formatting can amplify the erosion of safety alignment\. These results suggest that prompt formatting during KD affects safety beyond a simple evaluation\-template artifact\.

We hypothesize that this occurs because chat templates expose structural role tokens and conversational transitions, such as user\-assistant boundaries, that are closely tied to the model’s learned refusal behavior\. During KD, repeatedly optimizing on benign chat\-formatted responses may overwrite or weaken these safety\-relevant associations, thereby disrupting the conditional behavior needed for refusal\. This interpretation is consistent with our representation\-level analysis, where chat\-template KD produces larger shifts in refusal\-related geometry compared to non\-chat\-template KD\.

A practical implication of these findings is that prompt templates should be decoupled between training and deployment when safety is the primary concern\. Specifically, our results suggest using a non\-chat template during KD training while still deploying through the standard chat interface, i\.e\.,non\-chat training→\\rightarrowchat inference\. However, this training–deployment asymmetry is unusual in practice, and its robustness beyond benign instruction\-following data and our benchmark settings remains to be verified\. Thus, prompt templates are not merely a syntactic implementation choice, but an important factor in alignment stability\.

## 7Conclusion

In this work, we show that knowledge distillation \(KD\) on benign tasks can erode the safety alignment of student models\. Across model families, KD improves task utility but also increases harmful\-compliance vulnerability, suggesting that optimizing for instruction\-following performance can conflict with preserving the refusal behavior already present in aligned student models\. Our findings further identify prompt templates as a key training\-time factor that amplifies this degradation\. More broadly, these results raise an important question for future work: whether fine\-tuning and distillation methods can preserve refusal geometry without relying on explicit safety supervision, enabling student models to improve utility while maintaining the safety behavior of their base instruct\-tuned models\.

## Limitations

While our study provides evidence that prompt templates play an important role in safety degradation during knowledge distillation, three limitations remain\.

##### Dataset\.

First, KD is performed exclusively on benign instruction\-following data, leaving open whether similar template\-driven degradation emerges in domain\-specific settings such as code, math, or reasoning\. These domains differ in how heavily they rely on structural tokens: code contains formatting cues that may interact with chat\-template control tokens in distinct ways, whereas math and reasoning data are largely plain text\. We restrict our study to the benign instruction\-following setting in order to isolate the effect of prompt formatting from domain\-induced distributional shifts; extending the analysis to structurally richer domains is a natural next step that our pipeline directly supports\.

##### Beyond LoRA\.

Second, all experiments use LoRA\-based distillation rather than full\-parameter fine\-tuning\. Prior work suggests that LoRA can produce behavioral and representational changes similar to full fine\-tuning\([Hu et al\., 2022](https://arxiv.org/html/2609.30802#bib.bib29)\)\. Therefore, we expect the template\-driven safety degradation observed in our experiments to also appear under full\-parameter distillation, although the magnitude may differ\. A direct comparison between LoRA and full\-parameter training is an important direction for future work\.

##### Evaluator Bias\.

Safety evaluation uses automated judges \(LLaMA\-Guard\-8B for AdvBench and JailbreakBench, a fine\-tuned Mistral judge for SORRY\-Bench, and LLaMA\-2\-13B\-cls for HarmBench\)\. These judges may not fully capture all forms of harmful content\. However, each benchmark uses a fixed judge, so comparisons within a benchmark remain consistent even if absolute ASR values differ\. A direct comparison across judges is left for future work\.

## Ethical Considerations

This work investigates conditions under which knowledge distillation on benign downstream tasks can erode the safety alignment of LLMs\. While these findings could in principle be exploited to bypass refusal behavior, our research primary intent is to highlight underexplored vulnerability in the KD pipeline\. We aim to inform the development of more robust safety\-preserving distillation techniques\. All experiments use publicly available datasets and models, and no human subjects were involved\. The safety evaluation datasets contain harmful or offensive content by design and were used solely for research and evaluation\.

## Acknowledgments

Research was sponsored by the Army Research Laboratory and was accomplished under Cooperative Agreement Number W911NF\-23\-2\-0224\. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the Army Research Laboratory or the U\.S\. Government\. The U\.S\. Government is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation herein\.

## References

- Agand \(2024\)P\. AgandKnowledge distillation from single\-task teachers to multi\-task student for end\-to\-end autonomous driving\.InProceedings of the AAAI Conference on Artificial Intelligence,pp\. 23375–23376\.Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p1.1)\.
- Arditiet al\.\(2024\)A\. Arditi, O\. Obeso, A\. Syed, D\. Paleka, N\. Panickssery, W\. Gurnee, and N\. NandaRefusal in language models is mediated by a single direction\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 136037–136083\.External Links:[Document](https://dx.doi.org/10.52202/079017-4322),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/f545448535dfde4f9786555403ab7c49-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p5.1),[§5\.2](https://arxiv.org/html/2609.30802#S5.SS2.SSS0.Px1.p1.1),[§5\.2](https://arxiv.org/html/2609.30802#S5.SS2.p1.1)\.
- Askellet al\.\(2021\)A\. Askell, Y\. Bai, A\. Chen, D\. Drain, D\. Ganguli, T\. Henighan, A\. Jones, N\. Joseph, B\. Mann, N\. DasSarma,et al\.A general language assistant as a laboratory for alignment\.arXiv preprint arXiv:2112\.00861\.Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p1.1)\.
- Baiet al\.\(2022\)Y\. Bai, A\. Jones, K\. Ndousse, A\. Askell, A\. Chen, N\. DasSarma, D\. Drain, S\. Fort, D\. Ganguli, T\. Henighan, N\. Joseph, S\. Kadavath, J\. Kernion, T\. Conerly, S\. El\-Showk, N\. Elhage, Z\. Hatfield\-Dodds, D\. Hernandez, T\. Hume, S\. Johnston, S\. Kravec, L\. Lovitt, N\. Nanda, C\. Olsson, D\. Amodei, T\. Brown, J\. Clark, S\. McCandlish, C\. Olah, B\. Mann, and J\. KaplanTraining a helpful and harmless assistant with reinforcement learning from human feedback\.arXiv preprint arXiv:2204\.05862\.Cited by:[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px2.p1.1)\.
- Chaoet al\.\(2024\)P\. Chao, E\. Debenedetti, A\. Robey, M\. Andriushchenko, F\. Croce, V\. Sehwag, E\. Dobriban, N\. Flammarion, G\. J\. Pappas, F\. Tramèr, H\. Hassani, and E\. WongJailbreakBench: an open robustness benchmark for jailbreaking large language models\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 55005–55029\.External Links:[Document](https://dx.doi.org/10.52202/079017-1745),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/63092d79154adebd7305dfd498cbff70-Paper-Datasets_and_Benchmarks_Track.pdf)Cited by:[§A\.4](https://arxiv.org/html/2609.30802#A1.SS4.p3.1),[§3\.3\.2](https://arxiv.org/html/2609.30802#S3.SS3.SSS2.Px2.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. MaThe llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§A\.4](https://arxiv.org/html/2609.30802#A1.SS4.p2.1),[§3\.3\.1](https://arxiv.org/html/2609.30802#S3.SS3.SSS1.p1.1),[§3\.3\.3](https://arxiv.org/html/2609.30802#S3.SS3.SSS3.Px3.p1.1)\.
- Guet al\.\(2024\)Y\. Gu, L\. Dong, F\. Wei, and M\. HuangMiniLLM: knowledge distillation of large language models\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 32694–32717\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/8ac015d409635f196f9e3e9dcfb9a94e-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px1.p1.1),[§3\.3\.1](https://arxiv.org/html/2609.30802#S3.SS3.SSS1.p1.1),[§3\.3\.2](https://arxiv.org/html/2609.30802#S3.SS3.SSS2.Px1.p1.1)\.
- Guptaet al\.\(2024\)P\. Gupta, L\. Q\. Yau, H\. H\. Low, I\. Lee, H\. M\. Lim, Y\. X\. Teoh, K\. J\. Hng, D\. W\. Liew, R\. Bhardwaj, R\. Bhardwaj, and S\. PoriaWalledEval: a comprehensive safety evaluation toolkit for large language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,D\. I\. Hernandez Farias, T\. Hope, and M\. Li \(Eds\.\),Miami, Florida, USA,pp\. 397–407\.External Links:[Link](https://aclanthology.org/2024.emnlp-demo.42/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-demo.42)Cited by:[§3\.3\.3](https://arxiv.org/html/2609.30802#S3.SS3.SSS3.Px2.p1.1)\.
- Hanet al\.\(2024\)S\. Han, K\. Rao, A\. Ettinger, L\. Jiang, B\. Y\. Lin, N\. Lambert, Y\. Choi, and N\. DziriWildGuard: open one\-stop moderation tools for safety risks, jailbreaks, and refusals of llms\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 8093–8131\.External Links:[Document](https://dx.doi.org/10.52202/079017-0261),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/0f69b4b96a46f284b726fbd70f74fb3b-Paper-Datasets_and_Benchmarks_Track.pdf)Cited by:[§5\.3](https://arxiv.org/html/2609.30802#S5.SS3.p1.1)\.
- Hintonet al\.\(2015\)G\. Hinton, O\. Vinyals, and J\. DeanDistilling the knowledge in a neural network\.arXiv preprint arXiv:1503\.02531\.Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p1.1),[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px1.p1.1)\.
- Huet al\.\(2022\)E\. J\. Hu, yelong shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[§3\.3\.3](https://arxiv.org/html/2609.30802#S3.SS3.SSS3.Px1.p1.1),[Beyond LoRA\.](https://arxiv.org/html/2609.30802#Sx1.SS0.SSS0.Px2.p1.1)\.
- Hugging Face Team \(2025\)Hugging Face TeamChat templates\.Note:[https://huggingface\.co/docs/transformers/chat\_templating](https://huggingface.co/docs/transformers/chat_templating)Part of the official Transformers documentation \(v5\.x series\)\.External Links:[Link](https://huggingface.co/docs/transformers/chat_templating)Cited by:[§3\.1](https://arxiv.org/html/2609.30802#S3.SS1.p1.1)\.
- Jianget al\.\(2023\)A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. SayedMistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p2.1)\.
- Jianget al\.\(2025\)F\. Jiang, Z\. Xu, L\. Niu, B\. Y\. Lin, and R\. PoovendranChatbug: a common vulnerability of aligned llms induced by chat templates\.InProceedings of the AAAI Conference on Artificial Intelligence,pp\. 27347–27355\.Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p2.1),[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2609.30802#S4.SS2.p2.1)\.
- Koet al\.\(2024\)J\. Ko, S\. Kim, T\. Chen, and S\. YunDistiLLM: towards streamlined distillation for large language models\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 24872–24895\.External Links:[Link](https://proceedings.mlr.press/v235/ko24c.html)Cited by:[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px1.p1.1),[§3\.3\.1](https://arxiv.org/html/2609.30802#S3.SS3.SSS1.p1.1)\.
- Lyuet al\.\(2024\)K\. Lyu, H\. Zhao, X\. Gu, D\. Yu, A\. Goyal, and S\. AroraKeeping llms aligned after fine\-tuning: the crucial role of prompt templates\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 118603–118631\.External Links:[Document](https://dx.doi.org/10.52202/079017-3766),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/d6f034bb216b472fc7d32ec7aff20342-Paper-Conference.pdf)Cited by:[Appendix D](https://arxiv.org/html/2609.30802#A4.p1.1),[§1](https://arxiv.org/html/2609.30802#S1.p2.1),[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2609.30802#S4.SS2.p2.1)\.
- Mazeikaet al\.\(2024\)M\. Mazeika, L\. Phan, X\. Yin, A\. Zou, Z\. Wang, N\. Mu, E\. Sakhaee, N\. Li, S\. Basart, B\. Li, D\. Forsyth, and D\. HendrycksHarmBench: a standardized evaluation framework for automated red teaming and robust refusal\.InProceedings of the 41st International Conference on Machine Learning,R\. Salakhutdinov, Z\. Kolter, K\. Heller, A\. Weller, N\. Oliver, J\. Scarlett, and F\. Berkenkamp \(Eds\.\),Proceedings of Machine Learning Research, Vol\.235,pp\. 35181–35224\.External Links:[Link](https://proceedings.mlr.press/v235/mazeika24a.html)Cited by:[§3\.3\.2](https://arxiv.org/html/2609.30802#S3.SS3.SSS2.Px2.p1.1)\.
- MohiEldeen Alabbasyet al\.\(2023\)F\. MohiEldeen Alabbasy, A\.S\. Abohamama, and M\. F\. AlrahmawyCompressing medical deep neural network models for edge devices using knowledge distillation\.Journal of King Saud University \- Computer and Information Sciences35\(7\),pp\. 101616\.External Links:ISSN 1319\-1578,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.jksuci.2023.101616),[Link](https://www.sciencedirect.com/science/article/pii/S1319157823001702)Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p1.1)\.
- Ouyanget al\.\(2022\)L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. F\. Christiano, J\. Leike, and R\. LoweTraining language models to follow instructions with human feedback\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 27730–27744\.External Links:[Document](https://dx.doi.org/10.52202/068431-2011),[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/b1efde53be364a73914f58805a001731-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p1.1),[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px2.p1.1)\.
- Qiet al\.\(2025\)X\. Qi, A\. Panda, K\. Lyu, X\. Ma, S\. Roy, A\. Beirami, P\. Mittal, and P\. HendersonSafety alignment should be made more than just a few tokens deep\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 54911–54941\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/88be023075a5a3ff3dc3b5d26623fa22-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p5.1)\.
- Qiet al\.\(2024\)X\. Qi, Y\. Zeng, T\. Xie, P\. Chen, R\. Jia, P\. Mittal, and P\. HendersonFine\-tuning aligned language models compromises safety, even when users do not intend to\!\.InInternational Conference on Learning Representations,B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(Eds\.\),Vol\.2024,pp\. 30988–31043\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/file/83b7da3ed13f06c13ce82235c8eedf35-Paper-Conference.pdf)Cited by:[§5\.3](https://arxiv.org/html/2609.30802#S5.SS3.p1.1)\.
- Qwenet al\.\(2024\)Qwen, :, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. QiuQwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§3\.3\.1](https://arxiv.org/html/2609.30802#S3.SS3.SSS1.p1.1)\.
- Rafailovet al\.\(2023\)R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. FinnDirect preference optimization: your language model is secretly a reward model\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 53728–53741\.External Links:[Document](https://dx.doi.org/10.52202/075280-2338),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/a85b405ed65c6477a4fe8302b5e06ce7-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px2.p1.1)\.
- Shayeganiet al\.\(2026\)E\. Shayegani, G\. M\. Shahariar, S\. Abdali, L\. Yu, N\. Abu\-Ghazaleh, and Y\. DongMisaligned roles, misplaced images: structural input perturbations expose multimodal alignment blind spots\.InInternational Conference on Learning Representations,C\. Vondrick, B\. Hariharan, C\. Raffel, L\. Pinto, D\. Yang, and A\. Faust \(Eds\.\),Vol\.2026,pp\. 92110–92145\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2026/file/94c51ec52572fb4e291dd1b1b8b02c31-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px3.p1.1)\.
- Teamet al\.\(2024\)G\. Team, M\. Riviere, S\. Pathak, P\. G\. Sessa, C\. Hardin, S\. Bhupatiraju, L\. Hussenot, T\. Mesnard, B\. Shahriari, A\. Ramé,et al\.Gemma 2: improving open language models at a practical size\.arXiv preprint arXiv:2408\.00118\.Cited by:[§3\.3\.1](https://arxiv.org/html/2609.30802#S3.SS3.SSS1.p1.1),[§3\.3\.3](https://arxiv.org/html/2609.30802#S3.SS3.SSS3.Px3.p1.1)\.
- Touvronet al\.\(2023\)H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale,et al\.Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p2.1)\.
- Wanget al\.\(2025a\)G\. Wang, Z\. Yang, Z\. Wang, S\. Wang, Q\. Xu, and Q\. HuangABKD: pursuing a proper allocation of the probability mass in knowledge distillation viaα\\alpha\-β\\beta\-divergence\.InProceedings of the 42nd International Conference on Machine Learning,A\. Singh, M\. Fazel, D\. Hsu, S\. Lacoste\-Julien, F\. Berkenkamp, T\. Maharaj, K\. Wagstaff, and J\. Zhu \(Eds\.\),Proceedings of Machine Learning Research, Vol\.267,pp\. 65167–65212\.External Links:[Link](https://proceedings.mlr.press/v267/wang25dz.html)Cited by:[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2025b\)Y\. Wang, A\. Bai, N\. Peng, and C\. HsiehOn the loss of context awareness in general instruction fine\-tuning\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38, Main Conference,pp\. 8610–8635\.External Links:[Document](https://dx.doi.org/10.52202/085713-0292),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/0cc540ef357216738768caaa7301b615-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.30802#S1.p2.1),[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px3.p1.1)\.
- Xieet al\.\(2025\)T\. Xie, X\. Qi, Y\. Zeng, Y\. Huang, U\. Sehwag, K\. Huang, L\. He, B\. Wei, D\. Li, Y\. Sheng, R\. Jia, B\. Li, K\. Li, D\. Chen, P\. Henderson, and P\. MittalSORRY\-bench: systematically evaluating large language model safety refusal\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 59937–59973\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/9622163c87b67fd5a4a0ec3247cf356e-Paper-Conference.pdf)Cited by:[§3\.3\.2](https://arxiv.org/html/2609.30802#S3.SS3.SSS2.Px2.p1.1)\.
- Yanget al\.\(2024\)M\. Yang, Y\. Chen, Y\. Liu, and L\. ShiDistillseq: a framework for safety alignment testing in large language models using knowledge distillation\.InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis,pp\. 578–589\.Cited by:[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px2.p1.1)\.
- Yiet al\.\(2024\)J\. Yi, R\. Ye, Q\. Chen, B\. Zhu, S\. Chen, D\. Lian, G\. Sun, X\. Xie, and F\. WuOn the vulnerability of safety alignment in open\-access LLMs\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 9236–9260\.External Links:[Link](https://aclanthology.org/2024.findings-acl.549/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.549)Cited by:[§2](https://arxiv.org/html/2609.30802#S2.SS0.SSS0.Px2.p1.1)\.
- Zouet al\.\(2023\)A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. FredriksonUniversal and transferable adversarial attacks on aligned language models\.arXiv preprint arXiv:2307\.15043\.Cited by:[§A\.4](https://arxiv.org/html/2609.30802#A1.SS4.p2.1),[§3\.3\.2](https://arxiv.org/html/2609.30802#S3.SS3.SSS2.Px2.p1.1)\.

## Appendix ATraining Details

### A\.1Hyperparameter details

Table 7:Hyperparameter settings\.Dataset\-wise training and evaluation configurations used for student model distillation\.Table[7](https://arxiv.org/html/2609.30802#A1.T7)summarizes the hyperparameter settings used for student model distillation\. We use Brain Floating Point \(BF16\) mixed precision to improve training throughput while maintaining numerical stability\. During instruction\-following evaluation, responses are generated with a decoding temperature oft=1\.0t=1\.0and nucleus sampling parametert​o​p​\_​p=0\.7top\\\_p=0\.7\.

### A\.2Dataset Details

Table 8:Data statistics of training and evaluation data\.In Table[8](https://arxiv.org/html/2609.30802#A1.T8), we provide the details of training and evaluation datasets used for downstream instruction\-following tasks and the benchmarks employed for safety evaluation\.

### A\.3Prompt Template

Table 9:Example of Prompt templates used for LLaMA, Gemma, and Qwen family\. The non\-chat is common to all the models whereas based on the model family chat template varies\.An example of each prompt template used during distillation is shown in Table[9](https://arxiv.org/html/2609.30802#A1.T9)\.

### A\.4Safety Evaluation

In this section, we summarize the safety benchmarks employed in to assess both the pre\-exisiting student models safety and their safety alignment after distillation process\.

AdvBench:We evaluate model safety on AdvBench, containing 520 harmful behaviors[Zou et al\. \(2023\)](https://arxiv.org/html/2609.30802#bib.bib27)\. Following WalledEval, we use their curated harmful prompts222AdvBench dataset:[walledai/AdvBench](https://huggingface.co/datasets/walledai/AdvBench)\.and prompt to the models\. Then, the responses are assessed usingLLaMA\-3\-Guard\-8B\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.30802#bib.bib16)\)\.

JailbreakBench:We further evaluate refusal robustness using 200 harmful prompts, from JailbreakBench[Chao et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib28)\. The evaluation was done by the same model as AdvBench\. The reported score of AdvBench and JailbreakBench, in[1](https://arxiv.org/html/2609.30802#S4.T1)is the average of two seeds\.

SORRY\-Bench:SORRY\-Bench comprises a 44\-class safety taxonomy with 440 base prompts and uses a fine\-tuned Mistral\-7B\-Instruct\-v0\.2 model as an automated judge for safety refusal\.333SORRY\-Bench judge model:[SORRY\-Bench/ft\-mistral\-7b\-instruct\-v0\.2\-SORRY\-Bench\-202406](https://huggingface.co/SORRY-Bench/ft-mistral-7b-instruct-v0.2-SORRY-Bench-202406)\.

We evaluate models on the base prompts and the fulfillment rate in the benchmark is reported here as ASR, defined as the proportion of responses that comply with unsafe instructions\. Higher ASR indicates weaker safety, while lower ASR reflects stronger refusal\. Evaluations use default decoding settings \(t=0\.7t=0\.7,t​o​p​\_​p=1\.0top\\\_p=1\.0, max tokens=1024=1024\), and reported results are averaged over three random seeds in Table[1](https://arxiv.org/html/2609.30802#S4.T1)\)\.

HarmBench:To evaluate the generalization of safety performance, we additionally assess models on HarmBench\. Although HarmBench is primarily designed for automated red teaming and evaluating LLM attacks and defenses, we used it to measure the safety of models based on ability to elicit harmful behavior\. HarmBench defines four behavior categories—standard, copyright, contextual, and multimodal—and we use the test split comprising 320 behaviors\. For each behavior, we generate five test cases usingMixtral\-8x7B\-Instruct\-v0\.1444[https://huggingface\.co/mistralai/Mixtral\-8x7B\-Instruct\-v0\.1](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1)Safety is evaluated using ASR, where lower values indicate safer models\. We employ the providedLLaMA\-2\-13b\-clsclassifier as the automated judge555HarmBench classifier:[cais/HarmBench\-LLaMA\-2\-13b\-cls](https://huggingface.co/cais/HarmBench-LLaMA-2-13b-cls)\.\. All reported results correspond to zero\-shot prompts generated with HarmBench’s default configuration\.

## Appendix BSafety\-Utility

Table 10:Utility Evaluation results\.Utility is measured using ROUGE\-L \(R\-L\) on DollyEval, SelfInst, VicunaEval, and S\-NI, with AVG reporting the mean score across the four benchmarks\. Safety is evaluated on SORRY\-Bench, where lower values indicate safer behavior\. The Method column indicates the training setting:ℳT\\mathcal\{M\}\_\{T\}denotes the fine\-tuned teacher model,ℳS\\mathcal\{M\}\_\{S\}denotes the base instruct\-tuned student model, SFT denotes supervised fine\-tuning, and KD as standard distillation\.Table[10](https://arxiv.org/html/2609.30802#A2.T10)presents a detailed breakdown of utility across the instruction\-following evaluation datasets\. All results are obtained using a decoding temperature oft=1\.0t=1\.0andt​o​p​\_​p=0\.7top\\\_p=0\.7\.

## Appendix CGeneralization to other distillation objective

Table 11:Generalization to other KD objectives under chat and non\-chat templates\.We report safety and utility for students trained with alternative distillation objectives under two prompt\-template settings\. Safety is measured using ASR \(%\) on SORRY\-Bench and HarmBench, where lower is better\. Utility reports the average ROUGE\-L score across the instruction\-following benchmarks\.The effect of prompt templates generalizes beyond the standard KD objective\. As shown in Table[11](https://arxiv.org/html/2609.30802#A3.T11), chat\-template distillation consistently produces higher ASR than non\-chat distillation across RKL and FKL\+RKL, while also yielding higher average utility\. This pattern holds across all three models\. For example, under RKL, HarmBench ASR increases from 12\.56 to 26\.63 for LLaMA and from 18\.38 to 29\.38 for Qwen when moving from non\-chat to chat distillation\. Gemma shows the same trend on SORRY\-Bench, increasing from 18\.86 to 41\.29\. These results indicate that template\-driven safety degradation is robust to changes in the distillation objective\.

## Appendix DTeacher Model Performance

Model\#ParamsMethodUtility \(R\-L↑\\uparrow\)Safety \(ASR↓\\downarrow\)AVGSORRYHarmBenchLLaMA8BℳT\\mathcal\{M\}\_\{T\}20\.2013\.268\.69SFT \(Non\-Chat\)24\.8613\.71\+0\.458\.63\-0\.06SFT \(Chat\)30\.3316\.37\+3\.1115\.69\+7\.003BℳS\\mathcal\{M\}\_\{S\}19\.8426\.7212\.18SFT \(Non\-Chat\)24\.2727\.12\+0\.4012\.19\+0\.01SFT \(Chat\)29\.1028\.71\+1\.9922\.50\+10\.32Gemma9BℳT\\mathcal\{M\}\_\{T\}22\.9112\.0514\.13SFT \(Non\-Chat\)25\.9312\.58\+0\.5315\.06\+0\.93SFT \(Chat\)35\.0646\.36\+34\.3124\.50\+10\.372BℳS\\mathcal\{M\}\_\{S\}22\.3516\.0411\.56SFT \(Non\-Chat\)23\.6628\.71\+12\.6720\.37\+8\.81SFT \(Chat\)30\.0553\.18\+37\.1413\.25\+1\.69

Table 12:Chat\-template SFT degrades safety across teacher and student models\.Utility is reported using AVG ROUGE\-L across DollyEval, SelfInst, VicunaEval, and S\-NI\. Safety is evaluated using SORRY\-Bench and HarmBench, where lower ASR indicates safer behavior\. The subscripted deltas show the change relative to each model’s own instruct\-tuned baseline \(ℳT\\mathcal\{M\}\_\{T\}orℳS\\mathcal\{M\}\_\{S\}\)\.Table[12](https://arxiv.org/html/2609.30802#A4.T12)reports supervised fine\-tuning under both prompt templates on teacher and student checkpoints\. We observed similar pattern of chat\-template SFT consistently degrading safety across both model families and both scales similar to[Lyu et al\. \(2024\)](https://arxiv.org/html/2609.30802#bib.bib13), while non\-chat SFT leaves the baseline largely intact \(e\.g\., Gemma 9B SORRY:\+34\.31\+34\.31vs\.\+0\.53\+0\.53; LLaMA 8B HarmBench:\+7\.00\+7\.00vs\.−0\.06\-0\.06\)\. The effect is therefore not a small\-model artifact nor specific to distillation\.

## Appendix EChat templates accelerate vulnerability\.

Figure 3:Temporal dynamics of safety and utility\.Chat \(orange\) triggers rapid utility and vulnerability spikes\. Non\-Chat \(blue\) follows a gradual trajectory, slowing safety erosion in student LLaMA modelTo further examine the effect of prompt templates, we track task utility, measured by ROUGE\-L, and safety degradation, measured by ASR, across distillation steps\. As shown in Figure[3](https://arxiv.org/html/2609.30802#A5.F3), chat\-template KD produces a sharper early increase in both utility and ASR\. This indicates that the chat template improves instruction\-following performance quickly, but also accelerates the loss of safety behavior\. In contrast, non\-chat\-template KD shows a more gradual trajectory, with slower increases in ASR while utility improves steadily\. These results provide additional evidence that chat templates amplify safety degradation during distillation, whereas non\-chat templates lead to a more stable training trajectory\.

Table 13:Preliminary refusal\-direction interventions\.ASR \(%\) before and after intervention on 40 harmful SORRY\-Bench prompts\. “Ablate KD \(Chat\)” removes the KD\-Chat refusal direction, while “Add to KD \(Non\-Chat\)” adds the base model’s refusal direction\.
## Appendix FPreliminary causal intervention on refusal direction\.

To complement the correlational analysis in §[5\.2](https://arxiv.org/html/2609.30802#S5.SS2), we perform small\-scale interventions on the difference\-in\-means refusal direction\. For each model family, we evaluate 40 harmful prompts from SORRY\-Bench using the same judge as in the main evaluation\. For KD\-Chat, we ablate the learned refusal direction from the model activations to test whether weakening this direction increases unsafe behavior\. For KD Non\-Chat, we add the base model’s refusal direction to test whether restoring this direction improves safety\. The full results are reported in Table[13](https://arxiv.org/html/2609.30802#A5.T13)\.

When we ablate the refusal direction, ASR increases for Gemma and Qwen, while adding the base model’s refusal direction to KD\-Non\-Chat reduces ASR across all three model families\. These results provide preliminary interventional evidence that modifying the refusal direction can influence safety behavior\.

For LLaMA student model, ablating the refusal direction decreases ASR, but this does not indicate improved safety\. After the intervention, 9/40 outputs become repetitive or off\-topic, compared with 1/40 before intervention\. These incoherent outputs are scored as non\-compliant by the judge, the measured ASR is mechanically reduced\. We therefore attribute the lower ASR in this setting to degraded output coherence rather than improved refusal behavior\.

## Appendix GFull token ablations comparison

Table[14](https://arxiv.org/html/2609.30802#A8.T14)provides the full HarmBench comparison across inference formats and token ablations\. We report results under the standard chat template, after removing BOS/EOT tokens, after removing role tags, under the non\-chat template, and finally using the raw harmful query without chat\-template formatting\. This ordering progressively removes chat\-specific structure and allows us to compare how safety behavior changes as different components of the inference format are removed\. Overall, KD\-Chat remains less safe than the corresponding base model under several ablated settings, especially after removing BOS/EOT tokens, while the role\-tag and raw\-query conditions are more model dependent\. These results suggest that no single chat\-template component fully accounts for the degradation, although inference formatting can substantially modulate its magnitude\.

## Appendix HFull Token Ablation Comparison

Table[14](https://arxiv.org/html/2609.30802#A8.T14)provides the full HarmBench comparison across inference formats and token ablations\. We report results under the standard chat template, after removing BOS/EOT tokens, after removing role tags, under the non\-chat template, and finally using the raw harmful query without chat\-template formatting\. This ordering progressively removes chat\-specific structure and allows us to compare how safety behavior changes as different components of the inference format are removed\. Overall, KD\-Chat remains less safe than the corresponding base model under several ablated settings, especially after removing BOS/EOT tokens, while the role\-tag and raw\-query conditions are more model dependent\. These results suggest that no single chat\-template component fully accounts for the degradation, although inference formatting can substantially modulate its magnitude\.

Table 14:Safety under inference\-template and token ablations\.HarmBench attack success rate \(ASR%\) for models with training settings and inference configurations\. The subscripted values report the change relative to the corresponding base model under the same evaluation configuration\.
## Appendix IUse of AI Assistants

We used AI assistants to help polish the text and debug code\. All content, ideas, and analyses presented in this paper remain the sole responsibility of the authors\.

相似文章

知识蒸馏对小型语言模型偏见的非对称影响

arXiv cs.CL

本文表明,知识蒸馏对小型语言模型中的偏见具有非对称影响:在无歧义任务上,它提升了上下文遵循能力,但在有歧义任务上却损害了拒答校准;并提出了PCCD,一种用于诊断聚合指标遗漏的逐项损害的协议。