迈向缓解大型推理模型中的欺骗性安全对齐
摘要
Wayne State University 的一篇论文提出了 DSAR 指标,用于量化大型推理模型中推理过程与最终答案之间的不一致性,并提出了基于强化学习的方法 SARA,通过奖励具备安全意识的推理来缓解在标准预填充攻击和对抗性预填充攻击下的欺骗性安全对齐。
查看缓存全文
缓存时间: 2026/09/30 09:42
# Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
Source: [https://arxiv.org/html/2609.36254](https://arxiv.org/html/2609.36254)
Saleh Zare ZadeAffiliation:Wayne State UniversityEmail:[salehz@wayne\.edu](mailto:)Rafi Ibn SultanAffiliation:Wayne State UniversityEmail:[rafis@wayne\.edu](mailto:)Alexander KotovAffiliation:Wayne State UniversityEmail:[kotov@wayne\.edu](mailto:)Dongxiao ZhuAffiliation:Wayne State UniversityEmail:[dzhu@wayne\.edu](mailto:)
###### Abstract
Large Reasoning Models \(LRMs\) are commonly trained with reinforcement learning \(RL\) to improve their generation of chain\-of\-thought \(CoT\) reasoning before producing final answers\. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning\. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals\. To systematically investigate this phenomenon, we introduceDSAR\(DeceptiveSafetyAlignmentRate\), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency\. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks\. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final\-answer stage than during intermediate reasoning\. To close this gap, we proposeSARA\(Safety\-AwareReasoningAlignment\), an RL\-based method that rewards both safety\-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning–answer consistency\. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility\. Our code is available at[SARA](https://github.com/xzhou98/SARA)\.
## 1Introduction
Building on the success of OpenAI’s o1[Jaech et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib5)and DeepSeek\-R1[Guo et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib6), Large Reasoning Models \(LRMs\) have emerged as a new class of foundation models that explicitly generate intermediate reasoning traces before producing final answers\. This reasoning\-centric paradigm has enabled strong performance on tasks such as mathematics and coding[Shao et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib7);[Jiang et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib8)\. Despite these advances, ensuring the safety of LRMs has become an increasingly critical concern[Wang et al\. \(2025a\)](https://arxiv.org/html/2609.36254#bib.bib14);[Zhou et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib15)\.
One important form of this safety concern isdeceptive safety alignment: recent studies suggest that frontier LRMs may produce inconsistent safety signals between intermediate reasoning and final answers[Krishna et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib75);[Ji et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib12);[Huang et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib13);[Carlsmith \(2023\)](https://arxiv.org/html/2609.36254#bib.bib10);[Zheng et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib76), where*\(i\)*the model may reason unsafely but answer safely, or*\(ii\)*reason safely but still produce an unsafe final answer\. Figure[1](https://arxiv.org/html/2609.36254#S1.F1)\(top\) illustrates the first and more prevalent case: under standard prompting, an LRM may recognize harmful intent yet continue reasoning toward harmful compliance; under adversarial \(adv\.\) prefilling, where an attacker forces the model to begin its reasoning with specific text[Li et al\. \(2025b\)](https://arxiv.org/html/2609.36254#bib.bib37);[Koorndijk \(2025\)](https://arxiv.org/html/2609.36254#bib.bib39);[Wang et al\. \(\)](https://arxiv.org/html/2609.36254#bib.bib41), the reasoning trajectory can be steered toward fully harmful compliance\. In both cases, however, the final answer remainssuperficially safe\. This reasoning\-answer inconsistency can arise because existing reinforcement learning \(RL\)\-based training objectives optimize models by rewarding only the final answer, rather than directly rewarding reasoning[Anand et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib22);[Uesato et al\. \(2022\)](https://arxiv.org/html/2609.36254#bib.bib23)\. Although prior works have studied deceptive alignment in LRMs[Krishna et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib75);[Ji et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib12);[Greenblatt et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib28), existing safety evaluations[Greenblatt et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib28);[Gao et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib63);[Krishna et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib75)still lack a direct metric for quantifying whether a model’s reasoning and final answer convey consistent safety signals within the same generation\.
To close this gap, we introduce DSAR \(DeceptiveSafetyAlignmentRate\), a metric that directly quantifies deceptive safety alignment by measuring the consistency of safety signals between the reasoning trace and the final answer\. To enable scalable and fine\-grained evaluation, we develop a pipeline that leverages an LLM\-based judge with carefully designed instructions to assess safety signals at the sentence level within the reasoning trace, alongside the safety of the full reasoning trace and final answer\. Using DSAR, our analysis reveals systematic reasoning–answer inconsistencies across multiple frontier LRMs, with substantially higher inconsistency under prefilling attacks\.
Existing alignment methods take different forms but remain limited in addressing deceptive safety alignment\. Off\-policy methods[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib19);[Wang et al\. \(2026b\)](https://arxiv.org/html/2609.36254#bib.bib20);[Jeung et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib21)typically rely on supervised fine\-tuning over externally reasoning trajectories, leading to a train\-inference mismatch\. During training, the model learns to imitate fixed trajectories, whereas during inference, it must generate its own reasoning step by step\. As a result, even small deviations in early reasoning can accumulate over time and eventually lead to a completely different trajectory from the one seen during training, weakening safety alignment under real inference conditions[Zhao et al\. \(2026\)](https://arxiv.org/html/2609.36254#bib.bib27)\. On\-policy methods such as RECAP[Peng et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib16)avoid this mismatch but optimize models based on final\-answer safety rewards through RL, but provide no explicit supervision over intermediate reasoning\. As a result, they leave deceptive safety alignment largely unaddressed\.
To close this gap, we proposeSARA\(Safety\-AwareReasoningAlignment\), an RL\-based alignment method with two complementary objectives\. First, SARA strengthens reasoning\-levelsafety awareness, namely the ability to identify whether a query or response is harmful[Wang et al\. \(2025b\)](https://arxiv.org/html/2609.36254#bib.bib59), since a model that cannot detect harm in its own reasoning cannot reliably refuse harmful requests\. Second, SARA enforces consistency between reasoning and final answers, so that safety\-aware reasoning leads to safe outputs rather than being silently overridden\. As shown in Figure[1](https://arxiv.org/html/2609.36254#S1.F1)\(bottom\), a SARA\-aligned model explicitly recognizes and rejects harmful intent in its reasoning, yielding consistent safety across the CoT reasoning trace and final answer even under prefilling attacks\. Experiments on frontier LRMs demonstrate that SARA significantly mitigates deceptive safety alignment across multiple models and adversarial settings, while preserving helpfulness and utility, suggesting that explicitly supervising reasoning is a promising direction for reliable safety alignment in LRMs\.
Figure 1:Illustration of deceptive safety alignment and its mitigation viaSARA\.Top:A base LRM can exhibitdeceptive safety alignmentin bothstandardandadv\. prefillingprompting settings\. In the standard setting, the model partially recognizes harmful intent but continues reasoning toward harmful compliance; in the adv\. prefilling setting, its reasoning is manipulated toward compliance entirely\. Both scenarios produce a superficially safe final answer\.Bottom:After alignment with ourSARA, the model produces safety\-aware reasoning that explicitly recognizes and rejects harmful intent, leading to a safe final answer and consistent safety behavior across both settings\.
## 2Deceptive Safety Alignment Evaluation
### 2\.1Reasoning–Answer Consistency: Why Existing Metrics Fall Short
Although prior work has studied deceptive safety alignment[Ji et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib12);[Greenblatt et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib28), existing safety evaluations still lack a direct metric for quantifying whether a model’s reasoning trace and final answer convey consistent safety signals within the same generation\. Prior metrics, such as Trajectory Coherence[Gao et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib63), measure whether the degree of risk accumulated across reasoning is consistent with the risk level of the final answer\. However, they do not directly assess whether the reasoning explicitly recognizes the harmful intent of the input, nor whether such recognition is consistently reflected in the safety of the final answer\. The Compliance Gap[Greenblatt et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib28)takes a different approach, measuring the difference in harmful compliance rates between settings where the model believes it is being monitored versus unmonitored\. They do not directly quantify the mismatch between reasoning\-level safety awareness and final answer safety within the same generation\. More recently, D\-REX[Krishna et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib75)takes an important step toward detecting deceptive reasoning by evaluating models under adversarial system prompt injections through a multi\-criteria rubric, but does not define metrics that directly quantify safety awareness in reasoning or the consistency between reasoning safety and final answer safety\.
These limitations motivate a framework for directly measuring the consistency between reasoning\-level safety awareness and final\-answer safety within the same generation\. We first formalize deceptive safety alignment in Sec\.[2\.2](https://arxiv.org/html/2609.36254#S2.SS2), then introduce our evaluation metric in Sec\.[2\.3](https://arxiv.org/html/2609.36254#S2.SS3), report empirical results under standard and prefilling prompting settings in Sec\.[2\.4](https://arxiv.org/html/2609.36254#S2.SS4), and provide mechanistic evidence for the reasoning–answer gap in Sec\.[2\.5](https://arxiv.org/html/2609.36254#S2.SS5)\.
### 2\.2Formalization
We now formalize the two failure modes identified above\. Letπθ\\pi\_\{\\theta\}denote an LRM parameterized byθ\\theta\. Given a harmful inputxharmx\_\{\\text\{harm\}\}with a chat template, the model generates\(ycot,yans\)∼πθ\(⋅∣xharm\)\(y\_\{\\text\{cot\}\},y\_\{\\text\{ans\}\}\)\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\_\{\\text\{harm\}\}\), whereycot=\{sk\}k=0N−1y\_\{\\text\{cot\}\}=\\\{s\_\{k\}\\\}\_\{k=0\}^\{N\-1\}is the CoT trace decomposed into sentences andyansy\_\{\\text\{ans\}\}is the final answer\.
We defineDeceptive Safety Alignmentas either case in whichycoty\_\{\\text\{cot\}\}andyansy\_\{\\text\{ans\}\}conveycontradictory safety signals\. This manifests in two cases:
- •Unsafe Reasoning→\\toSafe Answer:The model fails to identify the harmful intent ofxharmx\_\{\\text\{harm\}\}and reasons toward harmful compliance, yet produces a safe, refusal\-style final answer\. Theyansy\_\{\\text\{ans\}\}appears aligned, while the intermediateycoty\_\{\\text\{cot\}\}is not\.
- •Safe Reasoning→\\toUnsafe Answer:The model correctly identifies the harmful intent and reasons toward refusal, yet still produces a harmful final answer\. Theycoty\_\{\\text\{cot\}\}appears aligned, while theyansy\_\{\\text\{ans\}\}is not\.
The metric introduced next captures this by directly measuring whether the safety signals of the reasoning trace and final answer agree within the same generation\.
Figure 2:Overview of the evaluation pipeline\.Given a harmful input \(optionally with adv\. prefills\), an LRM generates a reasoningycoty\_\{\\text\{cot\}\}and a final answeryansy\_\{\\text\{ans\}\}, evaluated independently via two branches\.\(1\) Reasoning:The full reasoningycoty\_\{\\text\{cot\}\}is assessed by LLM Guard for overall safety\. Since standard LLM guards are designed to detect whether text is harmful rather than to localize sentence\-level safety awareness, we additionally use an LLM Judge to assess each sentencesks\_\{k\}individually\. The reasoning trace is safe if it either contains at least one safety\-aware sentence or is judged safe as a whole\.\(2\) Final answer:LLM Guard evaluates the overall safety ofyansy\_\{\\text\{ans\}\}\. A safety\-alignment is considered deceptive ifr\(ycot\)≠s\(yans\)r\(y\_\{\\text\{cot\}\}\)\\neq s\(y\_\{\\text\{ans\}\}\)\.
### 2\.3Deceptive Safety Alignment Rate Metric
We quantify deceptive safety alignment by independently evaluating the reasoning trace and the final answer, then measuring whether they agree\. Figure[2](https://arxiv.org/html/2609.36254#S2.F2)illustrates the full evaluation pipeline\.
Reasoning trace evaluation\.We evaluateycoty\_\{\\text\{cot\}\}in two complementary ways:*safety awareness*and*overall harmfulness*\. A standard LLM guard can judge whether a trace contains harmful content, but it does not directly determine whether a specific sentence recognizes harmful intent and uses that recognition to refuse, stop, or redirect\. Therefore, we first use an LLM judge to identify reasoning safety awareness, and then use an LLM guard to assess the overall safety of the entire reasoning trace\.
First, we assess each sentencesk∈ycots\_\{k\}\\in y\_\{\\text\{cot\}\}individually to determine whether the reasoning contains a safety\-aware sentence, denoted byrSA\(ycot\)∈\{0,1\}r\_\{\\text\{SA\}\}\(y\_\{\\text\{cot\}\}\)\\in\\\{0,1\\\}\. A sentence is classified as safety\-aware if it both explicitly recognizes the harmful intent of the input and uses that recognition to refuse, stop, or redirect toward a safer alternative\. Crucially, a sentence is not classified as safety\-aware if it merely acknowledges potential harm but continues reasoning toward harmful compliance, as illustrated by the standard prompting case in Figure[1](https://arxiv.org/html/2609.36254#S1.F1)\(top\)\. To enable scalable evaluation, we use GPT\-oss\-safeguard\-20B[Agarwal et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib24)as an automated judge under a carefully designed prompt instruction \(Appendix Figure[8](https://arxiv.org/html/2609.36254#A9.F8)\), validated against human evaluation with 95% agreement \(Appendix[D](https://arxiv.org/html/2609.36254#A4)\)\.
Second, safety\-aware reasoning is not the only way a reasoning trace can be safe: some traces may avoid harmful content throughout without explicitly identifying the harmful intent\. An illustrative example is provided in Appendix Figure[5](https://arxiv.org/html/2609.36254#A5.F5)\. For this reason, we additionally use an LLM guard \(e\.g\., Qwen3Guard[Zhao et al\. \(2025a\)](https://arxiv.org/html/2609.36254#bib.bib25)\) to assess whether the full reasoning trace is safe overall, denoted byrsafe\(ycot\)∈\{0,1\}r\_\{\\text\{safe\}\}\(y\_\{\\text\{cot\}\}\)\\in\\\{0,1\\\}\. We combine the two as
r\(ycot\)=𝟙\[rSA\(ycot\)=1∨rsafe\(ycot\)=1\],r\(y\_\{\\text\{cot\}\}\)=\\mathbbm\{1\}\\bigl\[r\_\{\\text\{SA\}\}\(y\_\{\\text\{cot\}\}\)=1\\;\\lor\\;r\_\{\\text\{safe\}\}\(y\_\{\\text\{cot\}\}\)=1\\bigr\],so that a reasoning trace is counted as safe if it either contains a safety\-aware sentence or is judged safe as a whole\.
Final answer evaluation\.Following prior work[Peng et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib16);[Kuo et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib17);[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib19);[Zhao et al\. \(2025b\)](https://arxiv.org/html/2609.36254#bib.bib18);[Wang et al\. \(2026b\)](https://arxiv.org/html/2609.36254#bib.bib20), we adopt the same LLM guard used for reasoning \(e\.g\., Qwen3Guard[Zhao et al\. \(2025a\)](https://arxiv.org/html/2609.36254#bib.bib25)\) as the judge to assess the safety ofyansy\_\{\\text\{ans\}\}\. We denote bys\(yans\)∈\{0,1\}s\(y\_\{\\text\{ans\}\}\)\\in\\\{0,1\\\}whether the entire final answer is safe\.
Deceptive Safety Alignment Rate \(DSAR\)\.A generation is considered deceptive when the reasoning trace and the final answer convey contradictory safety signalsr\(ycot\)≠s\(yans\)r\(y\_\{\\text\{cot\}\}\)\\neq s\(y\_\{\\text\{ans\}\}\)\. To quantify how prevalent this phenomenon is, we define our primary metric over a test set𝒟test\\mathcal\{D\_\{\\text\{test\}\}\}as:
DSAR=𝔼x∼𝒟test,\(ycot,yans\)∼πθ\(⋅∣x\)\[𝟙\[r\(ycot\)≠s\(yans\)\]\]\.\\mathrm\{DSAR\}=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{\\text\{test\}\},\\,\(y\_\{\\text\{cot\}\},y\_\{\\text\{ans\}\}\)\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\)\}\\Bigl\[\\mathbbm\{1\}\\bigl\[r\(y\_\{\\text\{cot\}\}\)\\neq s\(y\_\{\\text\{ans\}\}\)\\bigr\]\\Bigr\]\.A higher DSAR indicates more severe deceptive safety alignment across the evaluated models\.
Table 1:Evaluation of deceptive safety alignment on StrongReject and SafeChain, under both standard and adv\. prefilling settings\. We report theSAR↑\\uparrowof the reasoning trace, theSS↑\\uparrowof the final answer, and the proposedDSAR↓\\downarrowof quantifying inconsistency between reasoning and answer\. HigherSARandSSindicate safer reasoning and final answers, while lowerDSARindicates better consistency between them\. The definition of metrics is provided in Sec\.[2\.4](https://arxiv.org/html/2609.36254#S2.SS4)and Appendix[B\.3](https://arxiv.org/html/2609.36254#A2.SS3)\. Overall, deceptive safety alignment is less common in safe models \(e\.g\., GPT\-oss\-20B\) and is consistently amplified under prefilling attacks\.
### 2\.4Evaluation Results on Deceptive Safety Alignment
Experiment Setup\.We evaluate deceptive safety alignment across a diverse set of LRMs spanning different architectures and scales, including Gemma4\-8B[Team et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib29), DS\-LLaMA3\-8B, DS\-Qwen3\-8B, DS\-Qwen2\-14B[Guo et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib6), and GPT\-oss\-20B[Agarwal et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib24)\. We use two benchmarks: the full StrongReject set[Souly et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib26)with 313 samples, and 500 harmful queries sampled from SafeChain[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib19)\. Beyond the jailbreak standard setting, we evaluate each model under a prefilling attack to test how adversarial pressure affects reasoning–answer consistency\. For each model and setting, we independently evaluate the reasoning trace and the final answer using the pipeline introduced in Sec\.[2\.3](https://arxiv.org/html/2609.36254#S2.SS3), reporting theSafety Score \(SS\)for the final answer and theSafety\-Aware Rate \(SAR\)for the reasoning trace \(see Appendix[B\.3](https://arxiv.org/html/2609.36254#A2.SS3)for detailed definitions of evaluation metrics\)\. We then measure deceptive safety alignment via the proposed DSAR metric \(Sec\.[2\.3](https://arxiv.org/html/2609.36254#S2.SS3)\)\.
Deceptive safety alignment appears even without adv\. prefilling, especially in weaker safety models\.As shown in Table[1](https://arxiv.org/html/2609.36254#S2.T1), DS\-LLaMA3\-8B and DS\-Qwen2\-14B exhibit a clear mismatch between reasoning safety and final\-answer safety in the standard setting without prefilling\. On StrongReject, they obtain DSAR scores of 23\.64 and 26\.20, respectively, and on SafeChain, their DSAR scores remain high at 22\.40 and 20\.20\. In contrast, a stronger safety model such as GPT\-oss\-20B remains much more consistent, achieving a DSAR of 0 on StrongReject and 5\.8 on Safechain\.
Prefilling attacks consistently amplify deceptive safety alignment across benchmarks and models\.Under prefilling, DSAR increases for all models, often sharply\. For instance, on StrongReject, DS\-Qwen3\-8B rises from 2\.56 to 33\.23 DSAR, and on SafeChain, DS\-Qwen2\-14B rises from 15\.00 to 34\.40, showing that adv\. prefills make reasoning and final answers much less aligned\.
Prefilling degrades reasoning safety more than final answer safety\.Across most models, SAR \(reasoning\) drops more sharply than SS \(answer\) under prefilling\. This suggests that prefilling primarily disrupts reasoning\-level safety awareness, while the final answer may still appear superficially safe\. This gap helps explain why deceptive safety alignment becomes more severe under attack\.
### 2\.5Mechanistic Evidence: Why Answer\-Only Rewarding Reinforces the Gap
The results in Sec\.[2\.4](https://arxiv.org/html/2609.36254#S2.SS4)show that deceptive safety alignment is systematic and becomes more severe under adversarial pressure\. A natural question is:*what does this decoupling look like at the representational level?*To investigate this, we analyze how the model’s internal representations evolve across layers at two key generation points: the beginning of the reasoning stage and the beginning of the final\-answer stage\.
Intuitively, if a model has learned to distinguish benign from harmful inputs, their internal representations should become increasingly separated across layers[Li et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib61);[Arditi et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib62)\. Following prior work[Li et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib61);[Arditi et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib62), we quantify this separation using layer\-wise cosine similarity\. For each layer, we extract the hidden state of the last input token immediately before generation begins, separately at the reasoning stage, where the model starts generating the reasoning trace, and at the final\-answer stage, where the model starts generating the final answer\. We construct five types of prompt pairs: benign–benign \(B\-B\) pairs at the reasoning stage, harmful–harmful \(H\-H\) pairs at both the reasoning and final\-answer stages, and benign–harmful \(B\-H\) pairs at both the reasoning and final\-answer stages\. B\-B and H\-H pairs serve as within\-class similarity references, while B\-H pairs measure cross\-class separation between benign and harmful inputs\. Detailed experimental settings and prompt construction procedures for the two stages are provided in Appendix[F](https://arxiv.org/html/2609.36254#A6)\.

\(a\) DS\-LLaMA3\-8B

\(b\) DS\-Qwen2\-14B

\(c\) GPT\-oss\-20B
Figure 3:Layer\-wise average cosine similarity of last\-token hidden representations\.We compare five conditions across DS\-LLaMA3\-8B, DS\-Qwen2\-14B, and GPT\-oss\-20B: benign–benign \(B\-B\) pairs at the reasoning stage, and harmful–harmful \(H\-H\) and benign–harmful \(B\-H\) pairs at both reasoning and answer stages\. Lower B\-H similarity indicates stronger internal discrimination between benign and harmful inputs\.As shown in Figure[3](https://arxiv.org/html/2609.36254#S2.F3), two consistent patterns emerge\. First, B\-H similarity at the final\-answer stage \(red dotted line\) drops substantially in deeper layers compared to the reasoning stage \(orange solid line\)\. This indicates that the model discriminates between benign and harmful inputs much more strongly when generating the final answer than during intermediate reasoning\. A likely explanation is that standard RL alignment mainly rewards the safety of the final answer rather than the safety awareness of the reasoning process itself, causing safety discrimination to concentrate at the final\-answer stage\. This representational gap directly explains deceptive safety alignment: the model may produce harmful reasoning yet still recover safe behavior at the final\-answer stage\. Second, this gap is pronounced even in DS\-LLaMA3\-8B \(Figure[3](https://arxiv.org/html/2609.36254#S2.F3)a\) and DS\-Qwen2\-14B \(Figure[3](https://arxiv.org/html/2609.36254#S2.F3)b\), models that exhibit higher DSAR scores and greater reasoning–answer inconsistency, further confirming that answer\-only reward leaves the reasoning stage comparatively unsupervised\.
Together, these findings provide mechanistic support for the claim in Sec\.[2\.1](https://arxiv.org/html/2609.36254#S2.SS1)that answer\-only reward signals concentrate safety discrimination at the answer stage\. This concentration creates a representational gap between the reasoning stage and the final\-answer stage that directly underlies deceptive safety alignment\. ComplementaryPCA visualizationsin Appendix[C](https://arxiv.org/html/2609.36254#A3)further support this finding\. Across all three models, benign and harmful prompts show only weak separation at the reasoning stage, whereas their representations become much more distinguishable at the answer stage\. This trend is consistent with the layer\-wise cosine similarity results reported here\.
## 3SARA: Safety\-Aware Reasoning Alignment
### 3\.1Motivation
The analysis in Sec\.[2\.5](https://arxiv.org/html/2609.36254#S2.SS5)suggests that LRMs exhibit stronger safety discrimination at the final answer stage than during intermediate reasoning\. However, existing methods such as RECAP[Peng et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib16), which optimize final answer safety rewards, may still leave deceptive safety alignment largely unaddressed\. Recent work[Ji et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib12)suggests that explicitly rewarding the reasoning process can encourage models to produce reasoning that is more consistent with their final answers\. Motivated by these findings, we proposeSARA\(Safety\-AwareReasoningAlignment\), an RL\-based approach that rewards both safety\-aware reasoning and safe final answers, thereby encouraging consistency between the two\. We first introduce our training objective in Sec\.[3\.2](https://arxiv.org/html/2609.36254#S3.SS2)and the design of our reward function in Sec\.[3\.3](https://arxiv.org/html/2609.36254#S3.SS3)\.
### 3\.2Training Objective
In this work, we adopt the DAPO framework[Yu et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib30)and directly reward both the intermediate reasoning traceycoty\_\{\\text\{cot\}\}and the final answeryansy\_\{\\text\{ans\}\}using reward signals informed by our safety\-aware LLM Judge \(Sec\.[2\.3](https://arxiv.org/html/2609.36254#S2.SS3)\)\. In addition, we augment half of the training data with counter\-aligned prefilled CoT to improve robustness under adversarial settings following RECAP[Peng et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib16)\(see Appendix[B\.2](https://arxiv.org/html/2609.36254#A2.SS2)for details\)\. The training objective is defined over promptx∼𝒟augmentedx\\sim\\mathcal\{D\}\_\{\\text\{augmented\}\}and groups of rollouts\{oi\}i=1G\\\{o\_\{i\}\\\}^\{G\}\_\{i=1\}sampled from the old policyπθold\(⋅∣x\)\\pi\_\{\\theta\_\{\\text\{old\}\}\}\(\\cdot\\mid x\):
𝒥SARA\(θ\)\\displaystyle\\mathcal\{J\}\_\{\\mathrm\{SARA\}\}\(\\theta\)=𝔼x∼𝒟augmented,\{oi\}i=1G∼πθold\(⋅∣x\)\\displaystyle=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{\\text\{augmented\}\},\\,\\\{o\_\{i\}\\\}\_\{i=1\}^\{G\}\\sim\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(\\cdot\\mid x\)\}\[1∑i=1G\|oi\|∑i=1G∑t=1\|oi\|min\(ri,t\(θ\)A^i,t,clip\(ri,t\(θ\),1−εlow,1\+εhigh\)A^i,t\)\],\\displaystyle\\Biggl\[\\frac\{1\}\{\\sum\_\{i=1\}^\{G\}\|o\_\{i\}\|\}\\sum\_\{i=1\}^\{G\}\\sum\_\{t=1\}^\{\|o\_\{i\}\|\}\\min\\\!\\left\(r\_\{i,t\}\(\\theta\)\\,\\hat\{A\}\_\{i,t\},\\ \\text\{clip \}\\\!\\bigl\(r\_\{i,t\}\(\\theta\),\\,1\-\\varepsilon\_\{\\mathrm\{low\}\},\\,1\+\\varepsilon\_\{\\mathrm\{high\}\}\\bigr\)\\hat\{A\}\_\{i,t\}\\right\)\\Biggr\],where
ri,t\(θ\)=πθ\(oi,t∣x,oi,<t\)πθold\(oi,t∣x,oi,<t\),A^i,t=Ri−mean\(\{Rj\}j=1G\)std\(\{Rj\}j=1G\)\.\\displaystyle r\_\{i,t\}\(\\theta\)=\\frac\{\\pi\_\{\\theta\}\(o\_\{i,t\}\\mid x,o\_\{i,<t\}\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(o\_\{i,t\}\\mid x,o\_\{i,<t\}\)\},\\qquad\\hat\{A\}\_\{i,t\}=\\frac\{R\_\{i\}\-\\mathrm\{mean\}\(\\\{R\_\{j\}\\\}\_\{j=1\}^\{G\}\)\}\{\\mathrm\{std\}\(\\\{R\_\{j\}\\\}\_\{j=1\}^\{G\}\)\}\.
Here:
- •RiR\_\{i\}is the scalar reward assigned to rolloutoio\_\{i\}based on both\(x,ycot\)\(x,y\_\{\\text\{cot\}\}\)and\(x,yans\)\(x,y\_\{\\text\{ans\}\}\)\. We define this reward in Sec\.[3\.3](https://arxiv.org/html/2609.36254#S3.SS3)\.
- •A^i,t\\hat\{A\}\_\{i,t\}is the normalized advantage estimated from the group of rollouts\{oi\}i=1G\\\{o\_\{i\}\\\}^\{G\}\_\{i=1\}\.
- •εlow\\varepsilon\_\{\\text\{low\}\},εhigh\\varepsilon\_\{\\text\{high\}\}, andGGare hyper\-parameters, specificallyεlow\\varepsilon\_\{\\text\{low\}\}andεhigh\\varepsilon\_\{\\text\{high\}\}are clipping thresholds, andGGis the number of rollouts per prompt\. Hyperparameter details are provided in Appendix[B\.1](https://arxiv.org/html/2609.36254#A2.SS1)\.
### 3\.3Rewarding via Safety\-Aware Reasoning
Our SARA uses different reward designs for harmful and benign prompts\. For harmful prompts, the goal is to produce a safe final answer supported by safety\-aware reasoning\. For benign prompts, the goal is to enhance helpfulness\.
#### Reward for harmful prompts\.
Given a rolloutoio\_\{i\}, we split it into a reasoning traceycot\(i\)y\_\{\\text\{cot\}\}^\{\(i\)\}and a final answeryans\(i\)y\_\{\\text\{ans\}\}^\{\(i\)\}\. Unlike previous work that assigns rewards based solely on the input and final answer pair\(x,yans\(i\)\)\(x,y\_\{\\text\{ans\}\}^\{\(i\)\}\), we separately evaluate the safety of both components using a reward modelRϕR\_\{\\phi\}\(e\.g\., IBM Granite\-Guardian\-3\.1\-8B[Padhi et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib74)\), and use its predicted probabilities as continuous reward signals:
Ricot=Rϕ\(x,ycot\(i\)\),Rians=Rϕ\(x,yans\(i\)\),R\_\{i\}^\{\\text\{cot\}\}=R\_\{\\phi\}\(x,y\_\{\\text\{cot\}\}^\{\(i\)\}\),\\qquad R\_\{i\}^\{\\text\{ans\}\}=R\_\{\\phi\}\(x,y\_\{\\text\{ans\}\}^\{\(i\)\}\),\(1\)where each reward score lies in the range\[0,1\]\[0,1\]\. However,RicotR\_\{i\}^\{\\text\{cot\}\}alone may not sufficiently reward recognition of harmful intent\. A reasoning trace that exhibits harmful compliance without explicit harmful instructions may still receive a moderately positive reward \(e\.g\., 0\.3\)\. We therefore introduce a safety\-aware rewardRiSAR\_\{i\}^\{\\text\{SA\}\}that assigns a higher reward when safety awareness appears earlier in the reasoning process and penalizes the model if no safety awareness exists\.
Specifically, letycot\(i\)=\{sk\}k=0Ni−1y\_\{\\text\{cot\}\}^\{\(i\)\}=\\\{s\_\{k\}\\\}\_\{k=0\}^\{N\_\{i\}\-1\}denote the sequential sentences of the reasoning\. We use the proposed sentence\-level LLM judgeRψR\_\{\\psi\}\(Sec\.[2\.3](https://arxiv.org/html/2609.36254#S2.SS3)\) to identify the earliest sentence that exhibits safety awareness\. If the first such sentence appears at positionk⋆k^\{\\star\}, we define the safety\-aware reward as:
RiSA=1−k⋆Ni,k⋆=min\{k∈\{0,…,Ni−1\}\|Rψ\(x,sk\)=1\},R\_\{i\}^\{\\text\{SA\}\}=1\-\\frac\{k^\{\\star\}\}\{N\_\{i\}\},\\quad k^\{\\star\}=\\min\\left\\\{k\\in\\\{0,\\dots,N\_\{i\}\-1\\\}\\;\\middle\|\\;R\_\{\\psi\}\\\!\\left\(x,s\_\{k\}\\right\)=1\\right\\\},\(2\)
whereNiN\_\{i\}denotes the number of sentences in the reasoning trace of rolloutoio\_\{i\}\. If no sentence in the reasoning trace is classified as safety\-aware, we setk⋆=Nik^\{\\star\}=N\_\{i\}, which givesRiSA=0R\_\{i\}^\{\\text\{SA\}\}=0\. Thus, this reward assigns higher values when the model recognizes harmful intent earlier, and zero reward when the reasoning trace never exhibits safety awareness\.
We combine all reward components as follows:
Riharm=12Ricot⋅RiSA⏟reasoning reward\+12Rians⏟answer reward\.R\_\{i\}^\{\\text\{harm\}\}=\\underbrace\{\\frac\{1\}\{2\}\\ R\_\{i\}^\{\\text\{cot\}\}\\ \\cdot\\ R\_\{i\}^\{\\text\{SA\}\}\}\_\{\\text\{reasoning reward\}\}\\ \+\\underbrace\{\\frac\{1\}\{2\}\\ R\_\{i\}^\{\\text\{ans\}\}\}\_\{\\text\{answer reward\}\}\.\(3\)
The multiplicative combination ofRicotR\_\{i\}^\{\\text\{cot\}\}andRiSAR\_\{i\}^\{\\text\{SA\}\}in the reasoning reward enforces two complementary properties:RicotR\_\{i\}^\{\\text\{cot\}\}measures the overall safety of the reasoning trace, whileRiSAR\_\{i\}^\{\\text\{SA\}\}penalizes late or absent safety awareness, encouraging the model to recognize harmful intent earlier in the reasoning process\. Together withRiansR\_\{i\}^\{\\text\{ans\}\}, this design ensures that safe final answers are supported by genuinely safety\-aware reasoning\.
#### Reward for benign prompts\.
For benign prompts, we do not apply the above safety\-consistency reward, since the objective is to avoid unnecessary refusals\. Letyans\(i\)y\_\{\\text\{ans\}\}^\{\(i\)\}denote the final answer of rolloutoio\_\{i\}\. We define the benign\-prompt reward as:
Ribenign=Rϕ′\(x,yans\(i\)\),R\_\{i\}^\{\\text\{benign\}\}=R\_\{\\phi^\{\\prime\}\}\(x,y\_\{\\text\{ans\}\}^\{\(i\)\}\),\(4\)whereRϕ′R\_\{\\phi^\{\\prime\}\}is a refusal\-based reward model distinct from the safety reward modelRϕR\_\{\\phi\}used for harmful prompts\. Concretely, we use DS\-Qwen2\-32B[Guo et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib6)as the judge under the prompt instruction shown in Appendix Figure[9](https://arxiv.org/html/2609.36254#A9.F9), which assigns a refusal score from 0 to 10 via a structured rubric, where a higher score indicates greater over\-refusal\. We normalize this into\[0,1\]\[0,1\]by computingRϕ′=1−refusal score10R\_\{\\phi^\{\\prime\}\}=1\-\\frac\{\\text\{refusal score\}\}\{10\}, so that fully helpful responses receive a reward of 1 and full refusals receive 0\.
#### Overall training reward\.
Combining the two cases, the rollout\-level reward is
Ri=\{Riharm,ifxis a harmful prompt\.Ribenign,ifxis a benign prompt\.R\_\{i\}=\\begin\{cases\}R\_\{i\}^\{\\text\{harm\}\},&\\text\{if \}x\\text\{ is a harmful prompt\}\.\\\\ R\_\{i\}^\{\\text\{benign\}\},&\\text\{if \}x\\text\{ is a benign prompt\}\.\\end\{cases\}\(5\)This reward design allows the policy to jointly improve robustness to harmful prompts while reducing over\-refusal on benign ones\.
## 4Experiments
### 4\.1Experiment Setup
Datasets and Models\.To evaluate the effectiveness of SARA, we conduct experiments on DSQwen3\-8B and DSQwen2\-14B[Guo et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib6)\. The training corpus consists of 2K prompts, including1K harmful promptsfrom SafeChain[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib19)and1K benign promptsthat elicit over\-refusal behavior from FalseReject[Zhang et al\. \(2025c\)](https://arxiv.org/html/2609.36254#bib.bib60)\. Training and evaluation samples are strictly non\-overlapping\. Following prior work[Peng et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib16), we augment half of the training samples with adv\. prefilling to improve the model’s robustness under adversarial attacks; further details are provided in Appendix[B\.2](https://arxiv.org/html/2609.36254#A2.SS2)\.
Baselines\.We compare SARA against four representative safety alignment methods spanning both off\-policy and on\-policy settings\.SafeChain[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib19),SafePath[Jeung et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib21), andSTAR\-1[Wang et al\. \(2026b\)](https://arxiv.org/html/2609.36254#bib.bib20)are off\-policy methods that apply supervised fine\-tuning \(SFT\) on curated safety datasets\.RECAP[Peng et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib16)is an on\-policy RL\-based method built on the DAPO framework[Yu et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib30)that optimizes final\-response safety rewards, augmenting its training data with counter\-aligned prefills to improve robustness under adversarial settings\. For a fair comparison, SafePath, RECAP, and SARA are all trained on the same 2K samples\. For SafeChain, we replace the 1K benign prompts with samples from its own proposed dataset while keeping the same 1K harmful prompts\. For STAR\-1, we fine\-tune exclusively on its own proposed dataset rather than the 2K samples described above\. Detailed hyperparameter settings for all baselines are provided in Appendix[B\.1](https://arxiv.org/html/2609.36254#A2.SS1)\.
Evaluations and Metrics\.We evaluate models across three dimensions:Safety,Helpfulness, andUtility\. All evaluation benchmarks are strictly non\-overlapping with the training data\. ForSafety, we assess models under both standard and adv\. prefilling prompting settings\. For prefilling attacks, we evaluate on StrongReject[Souly et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib26); for standard prompting attacks, we evaluate on SafeChain with 500 test samples\. For both, we report theSafety\-Aware Rate \(SAR\)for the reasoning trace andSafety Score \(SS\)for the final answer, and measure deceptive safety alignment via the proposedDSARmetric\. We additionally assess robustness to unseen attacks, including 16\-shot ICL attacks[Wei et al\. \(2023\)](https://arxiv.org/html/2609.36254#bib.bib51);[Zhou et al\. \(2023\)](https://arxiv.org/html/2609.36254#bib.bib52)and the reasoning\-based attack H\-CoT[Kuo et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib17), reportingSSfor the final answer\. ForOver\-refusal, we evaluate over\-refusal on OR\-Bench\-hard[Cui et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib54)\(∼\\sim1,000 hard prompts challenging for state\-of\-the\-art LRMs\), Fortress[Knight et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib56), and XSTest[Röttger et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib55), reporting theHelpfulness Score \(HS\)evaluated by GPT\-oss\-safeguard[Agarwal et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib24)with the instruction shown in Appendix Fig\.[7](https://arxiv.org/html/2609.36254#A8.F7)\. ForUtility, we evaluate mathematical reasoning on GSM8K[Cobbe et al\. \(2021\)](https://arxiv.org/html/2609.36254#bib.bib57)and general knowledge on MMLU\-Pro[Wang et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib58), reportingAccuracy \(Acc\.\)\.
To summarize performance across all dimensions, we report the harmonic mean across the three task\-level scores, which penalizes imbalanced trade\-offs among safety, helpfulness, and utility\. Detailed evaluation information and metric definitions are provided in Appendix[B\.3](https://arxiv.org/html/2609.36254#A2.SS3)
### 4\.2Results and Discussion
SARA improves reasoning\-answer consistency while achieving strong final answer safety\.As shown in Table[2](https://arxiv.org/html/2609.36254#S4.T2), SARA achieves the highest 1−\-DSARunder prefilling attacks across both models, reaching85\.30on DS\-Qwen3\-8B and84\.66on DS\-Qwen2\-14B, indicating substantially better reasoning\-answer consistency compared with all baselines\. Crucially, SARA also achieves the highestSARacross both models and evaluation settings, confirming that this consistency gain is driven by genuinely safer reasoning\. This contrasts with RECAP, which achieves marginally higher final answerSSin some settings \(e\.g\.,99\.40vs\.97\.44on DSQwen3\-8B under prefilling\) but substantially lowerSAR\(57\.80 vs\.75\.10\), revealing that RECAP’s safety gains are concentrated at the final answer stage while the reasoning trace remains largely unsafe\.
Off\-policy baselines improve final answer safety but fail to align reasoning\.STAR\-1, SafeChain, and SafePath provide partial improvements over the original models but with inconsistent gains across reasoning safety, final answer safety, and helpfulness\. On DS\-Qwen3\-8B, SafeChain and SafePath improve standard settingSS, but their prefillingSARremains low \(23\.60 and 27\.20\), and their 1−\-DSARscores fall below the original model, indicating weaker reasoning\-answer consistency under adversarial pressure\. STAR\-1 improves safety but at the cost of helpfulness, dropping OR\-BenchHSfrom 67\.55 to 49\.89 on DS\-Qwen3\-8B\. Overall, these methods strengthen final answer safety in standard settings but do not reliably supervise the reasoning trace, leaving reasoning\-answer consistency under prefilling largely unaddressed\.
SARA achieves the best overall trade\-off between safety, helpfulness, and Utility\.Despite its strong safety gains, SARA achieves the best overall Avg\. score on both models,82\.76on DSQwen3\-8B and81\.97on DSQwen2\-14B, demonstrating that its safety improvements do not come at the expense of helpfulness or utility\. Across helpfulness and utility benchmarks, SARA maintains competitive performance relative to the original models and baselines, achieving comparable or superior results in most settings\. An aggregated evaluation across all tasks is provided in Appendix[G](https://arxiv.org/html/2609.36254#A7)\.
Table 2:Comparison of safety alignment methods on DSQwen3\-8B and DSQwen2\-14B across safety, helpfulness, and utility tasks\. For safety, we report theSAR↑\\uparrowfor the reasoning,SS↑\\uparrowfor the final answer, and1−\-DSAR↑\\uparrowunder both adv\. prefilling \(StrongReject\) and standard \(SafeChain\) settings, as well asSS↑\\uparrowunder unseen adv\. attacks \(ICL and H\-CoT\)\. For helpfulness, we report theHS↑\\uparrowon OR\-Bench\-hard, Fortress, and XSTest\. For utility, we report Acc\.↑\\uparrowon GSM8K and MMLU\-Pro\.Avg\.denotes the harmonic mean across the three task\-level scores\. Detailed definitions of evaluation metrics are provided in Sec\.[4\.1](https://arxiv.org/html/2609.36254#S4.SS1)and Appendix[B\.3](https://arxiv.org/html/2609.36254#A2.SS3)\.Boldindicates the best result among all methods\.MethodSafetyHelpfulnessUtilityStrongReject \(Adv\. Prefilling\)SafeChain \(Standard\)ICLH\-CoTOR\-BenchFortressXSTestGSM8kMMLU\-ProAvg\.SAR↑\\uparrowSS↑\\uparrow1\-DSAR↑\\uparrowSAR↑\\uparrowSS↑\\uparrow1\-DSAR↑\\uparrowSS↑\\uparrowSS↑\\uparrowHS↑\\uparrowHS↑\\uparrowHS↑\\uparrowAcc\.↑\\uparrowAcc\.↑\\uparrowdeepseek\-ai/DeepSeek\-R1\-0528\-Qwen3\-8BOriginal53\.4087\.8666\.7756\.6079\.8085\.0095\.5016\.0067\.5595\.2066\.4585\.4460\.9872\.23STAR\-135\.5088\.1854\.6759\.6080\.8086\.2097\.5014\.0049\.8993\.2060\.8983\.9361\.8568\.31SafeChain23\.6080\.8346\.6549\.6079\.2085\.0098\.5024\.0083\.2498\.0072\.6785\.7560\.1871\.54SafePath27\.2081\.4747\.9249\.4080\.4086\.2098\.0022\.0075\.3696\.4070\.8983\.4758\.9870\.35RECAP57\.8099\.4070\.9371\.0098\.8096\.30100\.050\.0067\.0292\.0076\.0086\.1362\.3277\.61SARA75\.1097\.4485\.3072\.8095\.4096\.4099\.2552\.0091\.2197\.0088\.0086\.2861\.7482\.76deepseek\-ai/DeepSeek\-R1\-Distill\-Qwen2\-14BOriginal32\.9053\.6774\.4433\.8064\.0079\.8074\.7516\.0097\.8099\.6092\.6781\.5856\.2168\.98STAR\-132\.9047\.2872\.2042\.6067\.0071\.2082\.2518\.0090\.2298\.8087\.3482\.2659\.2469\.05SafeChain37\.7063\.5867\.0948\.0082\.2089\.4090\.5024\.0069\.2986\.2064\.0083\.7859\.4368\.88SafePath34\.5061\.6661\.9845\.8084\.2090\.2085\.0014\.0066\.1179\.0061\.5682\.7158\.8366\.07RECAP70\.9099\.0480\.5175\.8096\.6095\.8096\.6056\.0099\.6299\.6093\.0081\.7355\.7381\.67SARA77\.0095\.8584\.6680\.4095\.0094\.6096\.7556\.0096\.8299\.8093\.7881\.2756\.5881\.97
### 4\.3Ablation of Reward Components and Prefill Augmentation
#### Reward\-Component\.
We compare four reward variants on DS\-Qwen3\-8B, keeping the remaining training settings fixed: \(1\) final\-answer safety alone \(RansR^\{\\mathrm\{ans\}\}, RECAP\); \(2\) full\-trace reasoning safety alone \(RcotR^\{\\mathrm\{cot\}\}\); \(3\) safety awareness alone \(RSAR^\{\\mathrm\{SA\}\}\), which rewards earlier recognition of harmful intent; and \(4\) the full SARA reward \(12RcotRSA\+12Rans\\frac\{1\}\{2\}R^\{\\mathrm\{cot\}\}R^\{\\mathrm\{SA\}\}\+\\frac\{1\}\{2\}R^\{\\mathrm\{ans\}\}\)\. Table[3](https://arxiv.org/html/2609.36254#S4.T3)reports the results\.
Under adversarial prefilling, both reasoning\-only variants achieve higher SAR and1−DSAR1\-\\mathrm\{DSAR\}than final\-answer\-only training, supporting direct supervision of intermediate reasoning\. The safety\-awareness reward achieves higher SAR than the full\-trace reasoning reward \(71\.6 versus 64\.2\), but lower SS \(93\.3 versus 99\.4\) and1−DSAR1\-\\mathrm\{DSAR\}\(81\.5 versus 84\.4\)\. Thus, rewarding harmful\-intent recognition and evaluating overall reasoning safety emphasize different properties\.
The full SARA reward achieves the highest SAR and1−DSAR1\-\\mathrm\{DSAR\}among the evaluated variants in both settings\. Under adversarial prefilling, it improves SAR from 57\.8 to 75\.1 and1−DSAR1\-\\mathrm\{DSAR\}from 70\.9 to 85\.3 relative to RECAP, while SS decreases from 99\.4 to 97\.4\. Under standard prompting, the gains in SAR and consistency are smaller, with a similar reduction in SS\. These results support combining the reward components to improve reasoning safety awareness and consistency, while highlighting a trade\-off in final\-answer safety\.
Table 3:Reward\-component ablation on DS\-Qwen3\-8Bunder adversarial prefilling on StrongReject and standard prompting on SafeChain\. The full SARA reward achieves the highest reasoning safety awareness \(SAR\) and reasoning–answer consistency \(1−DSAR1\-\\mathrm\{DSAR\}\), with a trade\-off in final\-answer safety \(SS\)\. All values are percentages; higher is better\.Bolddenotes the best result\.
#### Counter\-aligned prefill augmentation\.
Under adversarial prefilling, SARA outperforms RECAP even without counter\-aligned prefill augmentation, improving SAR by 11\.1 percentage points on DS\-Qwen3\-8B and 22\.7 points on DS\-Qwen2\-14B\. Augmentation further improves SARA’s prefilling robustness, although its performance under standard prompting declines\. These results indicate that reasoning\-level rewards contribute beyond augmentation\. Full results and discussion appear in Appendix[F\.1](https://arxiv.org/html/2609.36254#A6.SS1)\.
## 5Related Works
Deceptive Alignment\.Deceptive alignment describes cases where a model appears aligned in its observable final answer while following a different objective or relying on misleading reasoning[Krishna et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib75);[Greenblatt et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib28);[Carlsmith \(2023\)](https://arxiv.org/html/2609.36254#bib.bib10);[Wang et al\. \(2026a\)](https://arxiv.org/html/2609.36254#bib.bib9);[Hubinger et al\. \(2019\)](https://arxiv.org/html/2609.36254#bib.bib11);[Koorndijk \(2025\)](https://arxiv.org/html/2609.36254#bib.bib39)\. This issue is especially important for LRMs, where CoT traces expose whether intermediate reasoning is consistent with the final answer\. Prior work shows that outcome\-only optimization can produce safe answers for the wrong reasons[Wang et al\. \(2026a\)](https://arxiv.org/html/2609.36254#bib.bib9), that reasoning models may exploit CoT to support hidden deceptive strategies[Ji et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib12), and that CoT oversight can be unreliable when reasoning traces are unfaithful[Chen et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib71), difficult to control[Yueh\-Han et al\. \(2026\)](https://arxiv.org/html/2609.36254#bib.bib64), encode hidden information[Skaf et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib65);[Anwar et al\. \(2026\)](https://arxiv.org/html/2609.36254#bib.bib66), or become obfuscated under optimization pressure[Baker et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib72)\. Deceptive alignment may also persist after safety training[Hubinger et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib73);[Schoen et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib67), and dedicated benchmarks highlight the need to evaluate deception beyond standard safety harms[Krishna et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib75);[Huang et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib13);[Gao et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib63);[Kran et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib43)\. Beyond safety, deceptive behavior has been studied in settings such as mathematical reasoning, simulated company\-assistant tasks, tool\-selection, and open\-ended interaction[Shen et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib79);[Järviniemi and Hubinger \(2024\)](https://arxiv.org/html/2609.36254#bib.bib40);[Leonesi et al\. \(2026\)](https://arxiv.org/html/2609.36254#bib.bib78);[Wu et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib80);[Abdulhai et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib81)\. Our work studies this problem in LRMs under standard prompting and prefilling attacks, where reasoning\-answer inconsistency becomes more pronounced, and introduces a metric for measuring deceptive safety alignment\.
Safety Alignment\.Safety alignment has been widely studied for LLMs through supervised fine\-tuning, preference optimization, and RLHF[Taori et al\. \(2023\)](https://arxiv.org/html/2609.36254#bib.bib45);[Rafailov et al\. \(2023\)](https://arxiv.org/html/2609.36254#bib.bib44);[Bai et al\. \(2022\)](https://arxiv.org/html/2609.36254#bib.bib46), with later work emphasizing reasoning\-based alignment and deeper safety supervision[Guan et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib47);[Qi et al\. \(2024\)](https://arxiv.org/html/2609.36254#bib.bib42)\. For LRMs, prior work has proposed curated SFT\-based methods[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib19);[Wang et al\. \(2026b\)](https://arxiv.org/html/2609.36254#bib.bib20);[Jeung et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib21);[Zhang et al\. \(2025b\)](https://arxiv.org/html/2609.36254#bib.bib70)and RL\-based approaches that combine safety and task rewards[Peng et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib16);[Yu et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib30);[Kim et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib48), while recent work further highlights broader LRM safety risks and argues that aligning only the final response is insufficient[Zhang et al\. \(2025a\)](https://arxiv.org/html/2609.36254#bib.bib49);[Chen et al\. \(2026\)](https://arxiv.org/html/2609.36254#bib.bib68);[Hu et al\. \(2026\)](https://arxiv.org/html/2609.36254#bib.bib77);[Li et al\. \(2025a\)](https://arxiv.org/html/2609.36254#bib.bib69);[Gao et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib63);[Wang et al\. \(2025a\)](https://arxiv.org/html/2609.36254#bib.bib14);[Schoen et al\. \(2025\)](https://arxiv.org/html/2609.36254#bib.bib67)\. Our work builds on this line by using reinforcement learning to jointly align the final answer and the reasoning through safety\-aware rewards, with a focus on reasoning\-answer consistency under adversarial prefilling attacks\. A more comprehensive discussion is provided in Appendix[A](https://arxiv.org/html/2609.36254#A1)\.
## 6Conclusion
In this work, we studied deceptive safety alignment in LRMs, where a model’s reasoning trace and final answer convey contradictory safety signals\. We introduced DSAR to measure this phenomenon and showed through comprehensive evaluation that it is pervasive under standard prompting and substantially amplified under adversarial prefilling attacks\. Mechanistic analysis revealed that frontier LRMs exhibit stronger safety discrimination at the final\-answer stage than at the reasoning stage\. To address this, we proposed SARA, an RL\-based method that rewards safety\-aware reasoning and enforces reasoning\-answer consistency\. Experiments on DSQwen3\-8B and DSQwen2\-14B demonstrate that SARA substantially mitigates deceptive safety alignment under both standard and adv\. prefilling settings while preserving helpfulness and utility, suggesting that supervising intermediate reasoning is a necessary condition for reliable safety alignment in LRMs\.
## References
- \[1\]A\. Jaech, A\. Kalai, A\. Lerer, A\. Richardson, A\. El\-Kishky, A\. Low, A\. Helyar, A\. Madry, A\. Beutel, A\. Carney,et al\.\(2024\)Openai o1 system card\.arXiv preprint arXiv:2412\.16720\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p1.1),[§1](https://arxiv.org/html/2609.36254#S1.p1.1)\.
- \[2\]D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.\(2025\)Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p1.1),[§B\.2](https://arxiv.org/html/2609.36254#A2.SS2.p2.1),[§1](https://arxiv.org/html/2609.36254#S1.p1.1),[§2\.4](https://arxiv.org/html/2609.36254#S2.SS4.p1.1),[§3\.3](https://arxiv.org/html/2609.36254#S3.SS3.SSS0.Px2.p1.2),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p1.1)\.
- \[3\]Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.\(2024\)Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p1.1),[§1](https://arxiv.org/html/2609.36254#S1.p1.1)\.
- \[4\]J\. Jiang, F\. Wang, J\. Shen, S\. Kim, and S\. Kim\(2024\)A survey on large language models for code generation\.arXiv preprint arXiv:2406\.00515\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p1.1)\.
- \[5\]C\. Wang, Y\. Liu, B\. Li, D\. Zhang, Z\. Li, and J\. Fang\(2025\)Safety in large reasoning models: a survey\.arXiv preprint arXiv:2504\.17704\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p1.1),[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[6\]K\. Zhou, C\. Liu, X\. Zhao, S\. Jangam, J\. Srinivasa, G\. Liu, D\. Song, and X\. E\. Wang\(2025\)The hidden risks of large reasoning models: a safety assessment of r1\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,pp\. 3250–3265\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p1.1)\.
- \[7\]S\. Krishna, A\. Zou, R\. Gupta, E\. K\. Jones, N\. Winter, D\. Hendrycks, J\. Z\. Kolter, M\. Fredrikson, and S\. Matsoukas\(2025\)D\-rex: a benchmark for detecting deceptive reasoning in large language models\.arXiv preprint arXiv:2509\.17938\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.36254#S2.SS1.p1.1),[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[8\]J\. Ji, W\. Chen, K\. Wang, D\. Hong, S\. Fang, B\. Chen, J\. Zhou, J\. Dai, S\. Han, Y\. Guo,et al\.\(2025\)Mitigating deceptive alignment via self\-monitoring\.arXiv preprint arXiv:2505\.18807\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.36254#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.36254#S3.SS1.p1.1),[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[9\]Y\. Huang, Y\. Sun, Y\. Zhang, R\. Zhang, Y\. Dong, and X\. Wei\(2025\)Deceptionbench: a comprehensive benchmark for ai deception behaviors in real\-world scenarios\.arXiv preprint arXiv:2510\.15501\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1),[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[10\]J\. Carlsmith\(2023\)Scheming ais: will ais fake alignment during training in order to get power?\.arXiv preprint arXiv:2311\.08379\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1),[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[11\]B\. Zheng, B\. Zheng, K\. Cao, Y\. Tan, Z\. Liu, W\. Wang, J\. Liu, J\. Yang, W\. Su,et al\.\(2025\)Beyond safe answers: a benchmark for evaluating true risk awareness in large reasoning models\.arXiv preprint arXiv:2505\.19690\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1)\.
- \[12\]Y\. Li, J\. Hu, W\. Sang, L\. Ma, J\. Xie, W\. Zhang, A\. Yu, S\. Zhao, Q\. Huang, and Q\. Zhou\(2025\)Prefill\-based jailbreak: a novel approach of bypassing llm safety boundary\.arXiv e\-prints,pp\. arXiv–2504\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1),[§1](https://arxiv.org/html/2609.36254#S1.p2.1)\.
- \[13\]J\. Koorndijk\(2025\)Empirical evidence for alignment faking in a small llm and prompt\-based mitigation techniques\.InProceedings of the AAAI Symposium Series,Vol\.7,pp\. 198–205\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1),[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[14\]Q\. Wang, Z\. Tang, N\. Chen, W\. Wang, and B\. HeReasoning models can be easily hacked by fake reasoning bias\.InLock\-LLM Workshop: Prevent Unauthorized Knowledge Use from Large Language Models,Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1),[§1](https://arxiv.org/html/2609.36254#S1.p2.1)\.
- \[15\]A\. Anand, N\. Mokhberian, P\. Kumar, A\. Saha, Z\. He, A\. Rao, F\. Morstatter, and K\. Lerman\(2024\)Don’t blame the data, blame the model: understanding noise and bias when learning from subjective annotations\.InProceedings of the 1st Workshop on Uncertainty\-Aware NLP \(UncertaiNLP 2024\),pp\. 102–113\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1)\.
- \[16\]J\. Uesato, N\. Kushman, R\. Kumar, F\. Song, N\. Siegel, L\. Wang, A\. Creswell, G\. Irving, and I\. Higgins\(2022\)Solving math word problems with process\-and outcome\-based feedback\.arXiv preprint arXiv:2211\.14275\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1)\.
- \[17\]R\. Greenblatt, C\. Denison, B\. Wright, F\. Roger, M\. MacDiarmid, S\. Marks, J\. Treutlein, T\. Belonax, J\. Chen, D\. Duvenaud,et al\.\(2024\)Alignment faking in large language models\.arXiv preprint arXiv:2412\.14093\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.36254#S2.SS1.p1.1),[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[18\]X\. Gao, S\. Yu, Z\. Chen, Y\. Lyu, W\. Yu, G\. Li, J\. Liu, J\. Gao, J\. Liang, Z\. Liu,et al\.\(2025\)SafeRBench: a comprehensive benchmark for safety assessment in large reasoning models\.arXiv preprint arXiv:2511\.15169\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.36254#S2.SS1.p1.1),[§5](https://arxiv.org/html/2609.36254#S5.p1.1),[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[19\]F\. Jiang, Z\. Xu, Y\. Li, L\. Niu, Z\. Xiang, B\. Li, B\. Y\. Lin, and R\. Poovendran\(2025\)Safechain: safety of language models with long chain\-of\-thought reasoning capabilities\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 23303–23320\.Cited by:[§B\.3](https://arxiv.org/html/2609.36254#A2.SS3.p1.1),[§1](https://arxiv.org/html/2609.36254#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.36254#S2.SS3.p5.1),[§2\.4](https://arxiv.org/html/2609.36254#S2.SS4.p1.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[20\]Z\. Wang, H\. Tu, Y\. Wang, J\. Wu, Y\. Liu, J\. Mei, B\. R\. Bartoldson, B\. Kailkhura, and C\. Xie\(2026\)Star\-1: safer alignment of reasoning llms with 1k data\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 37988–37997\.Cited by:[§B\.3](https://arxiv.org/html/2609.36254#A2.SS3.p1.1),[§1](https://arxiv.org/html/2609.36254#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.36254#S2.SS3.p5.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[21\]W\. Jeung, S\. Yoon, M\. Kahng, and A\. No\(2025\)Safepath: preventing harmful reasoning in chain\-of\-thought via early alignment\.arXiv preprint arXiv:2505\.14667\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p4.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[22\]S\. Zhao, Z\. Xie, M\. Liu, J\. Huang, G\. Pang, F\. Chen, and A\. Grover\(2026\)Self\-distilled reasoner: on\-policy self\-distillation for large language models\.arXiv preprint arXiv:2601\.18734\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p4.1)\.
- \[23\]S\. Peng, E\. Smith, I\. Evtimov, S\. Jiang, P\. Chen, H\. Zhan, H\. Wang, D\. H\. Chau, M\. Pasupuleti, and J\. Chi\(2025\)Large reasoning models learn better alignment from flawed thinking\.arXiv preprint arXiv:2510\.00938\.Cited by:[§B\.2](https://arxiv.org/html/2609.36254#A2.SS2.p1.1),[§B\.3](https://arxiv.org/html/2609.36254#A2.SS3.p1.1),[§1](https://arxiv.org/html/2609.36254#S1.p4.1),[§2\.3](https://arxiv.org/html/2609.36254#S2.SS3.p5.1),[§3\.1](https://arxiv.org/html/2609.36254#S3.SS1.p1.1),[§3\.2](https://arxiv.org/html/2609.36254#S3.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[24\]W\. Wang, X\. Liu, K\. Gao, J\. Huang, Y\. Yuan, P\. He, S\. Wang, and Z\. Tu\(2025\)Can’t see the forest for the trees: benchmarking multimodal safety awareness for multimodal llms\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 16993–17006\.Cited by:[§1](https://arxiv.org/html/2609.36254#S1.p5.1)\.
- \[25\]S\. Agarwal, L\. Ahmad, J\. Ai, S\. Altman, A\. Applebaum, E\. Arbus, R\. K\. Arora, Y\. Bai, B\. Baker, H\. Bao,et al\.\(2025\)Gpt\-oss\-120b & gpt\-oss\-20b model card\.arXiv preprint arXiv:2508\.10925\.Cited by:[§B\.3](https://arxiv.org/html/2609.36254#A2.SS3.p5.1),[§2\.3](https://arxiv.org/html/2609.36254#S2.SS3.p3.1),[§2\.4](https://arxiv.org/html/2609.36254#S2.SS4.p1.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[26\]H\. Zhao, C\. Yuan, F\. Huang, X\. Hu, Y\. Zhang, A\. Yang, B\. Yu, D\. Liu, J\. Zhou, J\. Lin,et al\.\(2025\)Qwen3guard technical report\.arXiv preprint arXiv:2510\.14276\.Cited by:[§B\.3](https://arxiv.org/html/2609.36254#A2.SS3.p3.1),[§2\.3](https://arxiv.org/html/2609.36254#S2.SS3.p4.1),[§2\.3](https://arxiv.org/html/2609.36254#S2.SS3.p5.1)\.
- \[27\]M\. Kuo, J\. Zhang, A\. Ding, Q\. Wang, L\. DiValentin, Y\. Bao, W\. Wei, H\. Li, and Y\. Chen\(2025\)H\-cot: hijacking the chain\-of\-thought safety reasoning mechanism to jailbreak large reasoning models, including openai o1/o3, deepseek\-r1, and gemini 2\.0 flash thinking\.arXiv preprint arXiv:2502\.12893\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1),[§B\.3](https://arxiv.org/html/2609.36254#A2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2609.36254#S2.SS3.p5.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[28\]J\. Zhao, T\. Fu, R\. Schaeffer, M\. Sharma, and F\. Barez\(2025\)Chain\-of\-thought hijacking\.arXiv preprint arXiv:2510\.26418\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1),[§B\.3](https://arxiv.org/html/2609.36254#A2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2609.36254#S2.SS3.p5.1)\.
- \[29\]G\. Team, T\. Mesnard, C\. Hardin, R\. Dadashi, S\. Bhupatiraju, S\. Pathak, L\. Sifre, M\. Rivière, M\. S\. Kale, J\. Love,et al\.\(2024\)Gemma: open models based on gemini research and technology\.arXiv preprint arXiv:2403\.08295\.Cited by:[§2\.4](https://arxiv.org/html/2609.36254#S2.SS4.p1.1)\.
- \[30\]A\. Souly, Q\. Lu, D\. Bowen, T\. Trinh, E\. Hsieh, S\. Pandey, P\. Abbeel, J\. Svegliato, S\. Emmons, O\. Watkins,et al\.\(2024\)A strongreject for empty jailbreaks\.Advances in Neural Information Processing Systems37,pp\. 125416–125440\.Cited by:[§2\.4](https://arxiv.org/html/2609.36254#S2.SS4.p1.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[31\]S\. Li, L\. Yao, L\. Zhang, and Y\. Li\(2024\)Safety layers in aligned large language models: the key to llm security\.arXiv preprint arXiv:2408\.17003\.Cited by:[Appendix F](https://arxiv.org/html/2609.36254#A6.p1.1),[§2\.5](https://arxiv.org/html/2609.36254#S2.SS5.p2.1)\.
- \[32\]A\. Arditi, O\. Obeso, A\. Syed, D\. Paleka, N\. Panickssery, W\. Gurnee, and N\. Nanda\(2024\)Refusal in language models is mediated by a single direction\.Advances in Neural Information Processing Systems37,pp\. 136037–136083\.Cited by:[Appendix F](https://arxiv.org/html/2609.36254#A6.p1.1),[§2\.5](https://arxiv.org/html/2609.36254#S2.SS5.p2.1)\.
- \[33\]Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, L\. Liu,et al\.\(2025\)Dapo: an open\-source llm reinforcement learning system at scale\.arXiv preprint arXiv:2503\.14476\.Cited by:[§3\.2](https://arxiv.org/html/2609.36254#S3.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p2.1),[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[34\]I\. Padhi, M\. Nagireddy, G\. Cornacchia, S\. Chaudhury, T\. Pedapati, P\. Dognin, K\. Murugesan, E\. Miehling, M\. S\. Cooper, K\. Fraser,et al\.\(2024\)Granite guardian\.arXiv preprint arXiv:2412\.07724\.Cited by:[§3\.3](https://arxiv.org/html/2609.36254#S3.SS3.SSS0.Px1.p1.1)\.
- \[35\]Z\. Zhang, W\. Xu, F\. Wu, and C\. K\. Reddy\(2025\)Falsereject: a resource for improving contextual safety and mitigating over\-refusals in llms via structured reasoning\.arXiv preprint arXiv:2505\.08054\.Cited by:[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p1.1)\.
- \[36\]Z\. Wei, Y\. Wang, A\. Li, Y\. Mo, and Y\. Wang\(2023\)Jailbreak and guard aligned language models with only few in\-context demonstrations\.arXiv preprint arXiv:2310\.06387\.Cited by:[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[37\]X\. Zhou, Y\. Qiang, S\. Z\. Zade, P\. Khanduri, and D\. Zhu\(2023\)Hijacking large language models via adversarial in\-context learning\.arXiv preprint arXiv:2311\.09948\.Cited by:[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[38\]J\. Cui, W\. Chiang, I\. Stoica, and C\. Hsieh\(2024\)Or\-bench: an over\-refusal benchmark for large language models\.arXiv preprint arXiv:2405\.20947\.Cited by:[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[39\]C\. Q\. Knight, K\. Deshpande, V\. Sirdeshmukh, M\. Mankikar, S\. R\. Team, S\. Team, and J\. Michael\(2025\)FORTRESS: frontier risk evaluation for national security and public safety\.arXiv preprint arXiv:2506\.14922\.Cited by:[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[40\]P\. Röttger, H\. Kirk, B\. Vidgen, G\. Attanasio, F\. Bianchi, and D\. Hovy\(2024\)Xstest: a test suite for identifying exaggerated safety behaviours in large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 5377–5400\.Cited by:[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[41\]K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano,et al\.\(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§B\.3](https://arxiv.org/html/2609.36254#A2.SS3.p6.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[42\]Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo,et al\.\(2024\)Mmlu\-pro: a more robust and challenging multi\-task language understanding benchmark\.Advances in Neural Information Processing Systems37,pp\. 95266–95290\.Cited by:[§B\.3](https://arxiv.org/html/2609.36254#A2.SS3.p6.1),[§4\.1](https://arxiv.org/html/2609.36254#S4.SS1.p3.1)\.
- \[43\]B\. Wang, Y\. Liu, Y\. Liu, T\. Tang, S\. Wang, C\. Gao, C\. Zheng, Y\. Zhang, L\. Yu, S\. Liu,et al\.\(2026\)Outcome accuracy is not enough: aligning the reasoning process of reward models\.arXiv preprint arXiv:2602\.04649\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[44\]E\. Hubinger, C\. Van Merwijk, V\. Mikulik, J\. Skalse, and S\. Garrabrant\(2019\)Risks from learned optimization in advanced machine learning systems\.arXiv preprint arXiv:1906\.01820\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[45\]Y\. Chen, J\. Benton, A\. Radhakrishnan, J\. Uesato, C\. Denison, J\. Schulman, A\. Somani, P\. Hase, M\. Wagner, F\. Roger,et al\.\(2025\)Reasoning models don’t always say what they think\.arXiv preprint arXiv:2505\.05410\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[46\]C\. Yueh\-Han, R\. McCarthy, B\. W\. Lee, H\. He, I\. Kivlichan, B\. Baker, M\. Carroll, and T\. Korbak\(2026\)Reasoning models struggle to control their chains of thought\.arXiv preprint arXiv:2603\.05706\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[47\]J\. Skaf, L\. Ibanez\-Lissen, R\. McCarthy, C\. Watts, V\. Georgiv, H\. Whittingham, L\. Gonzalez\-Manzano, D\. Lindner, C\. Tice, E\. J\. Young,et al\.\(2025\)Large language models can learn and generalize steganographic chain\-of\-thought under process supervision\.arXiv preprint arXiv:2506\.01926\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[48\]U\. Anwar, J\. Piskorz, D\. D\. Baek, D\. Africa, J\. Weatherall, M\. Tegmark, C\. S\. de Witt, M\. van der Schaar, and D\. Krueger\(2026\)A decision\-theoretic formalisation of steganography with applications to llm monitoring\.arXiv preprint arXiv:2602\.23163\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[49\]B\. Baker, J\. Huizinga, L\. Gao, Z\. Dou, M\. Y\. Guan, A\. Madry, W\. Zaremba, J\. Pachocki, and D\. Farhi\(2025\)Monitoring reasoning models for misbehavior and the risks of promoting obfuscation\.arXiv preprint arXiv:2503\.11926\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[50\]E\. Hubinger, C\. Denison, J\. Mu, M\. Lambert, M\. Tong, M\. MacDiarmid, T\. Lanham, D\. M\. Ziegler, T\. Maxwell, N\. Cheng,et al\.\(2024\)Sleeper agents: training deceptive llms that persist through safety training\.arXiv preprint arXiv:2401\.05566\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[51\]B\. Schoen, E\. Nitishinskaya, M\. Balesni, A\. Højmark, F\. Hofstätter, J\. Scheurer, A\. Meinke, J\. Wolfe, T\. van der Weij,et al\.\(2025\)Stress testing deliberative alignment for anti\-scheming training\.arXiv preprint arXiv:2509\.15541\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1),[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[52\]E\. Kran, H\. M\. Nguyen, A\. Kundu, S\. Jawhar, J\. Park, M\. M\. Jurewicz,et al\.\(2025\)Darkbench: benchmarking dark patterns in large language models\.arXiv preprint arXiv:2503\.10728\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[53\]W\. Shen, H\. Wang, H\. Li, and H\. Zhang\(2025\)Decepchain: inducing deceptive reasoning in large language models\.arXiv preprint arXiv:2510\.00319\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[54\]O\. Järviniemi and E\. Hubinger\(2024\)Uncovering deceptive tendencies in language models: a simulated company ai assistant\.arXiv preprint arXiv:2405\.01576\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[55\]M\. Leonesi, F\. Belardinelli, F\. Corradini, and M\. Piangerelli\(2026\)Tatemae: detecting alignment faking via tool selection in llms\.arXiv preprint arXiv:2604\.26511\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[56\]Y\. Wu, X\. Pan, G\. Hong, and M\. Yang\(2025\)Opendeception: benchmarking and investigating ai deceptive behaviors via open\-ended interaction simulation\.arXiv preprint arXiv:2504\.13707\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[57\]M\. Abdulhai, R\. Cheng, A\. Shrivastava, N\. Jaques, Y\. Gal, and S\. Levine\(2025\)Evaluating & reducing deceptive dialogue from language models with multi\-turn rl\.arXiv preprint arXiv:2510\.14318\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p1.1)\.
- \[58\]R\. Taori, I\. Gulrajani, T\. Zhang, Y\. Dubois, X\. Li, C\. Guestrin, P\. Liang, and T\. B\. Hashimoto\(2023\)Stanford alpaca: an instruction\-following llama model\.Stanford, CA, USA\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[59\]R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn\(2023\)Direct preference optimization: your language model is secretly a reward model\.Advances in neural information processing systems36,pp\. 53728–53741\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[60\]Y\. Bai, A\. Jones, K\. Ndousse, A\. Askell, A\. Chen, N\. DasSarma, D\. Drain, S\. Fort, D\. Ganguli,et al\.\(2022\)Training a helpful and harmless assistant with reinforcement learning from human feedback\.arXiv preprint arXiv:2204\.05862\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[61\]M\. Y\. Guan, M\. Joglekar, E\. Wallace, S\. Jain, B\. Barak, A\. Helyar, R\. Dias, A\. Vallone, H\. Ren, J\. Wei,et al\.\(2024\)Deliberative alignment: reasoning enables safer language models\.arXiv preprint arXiv:2412\.16339\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[62\]X\. Qi, A\. Panda, K\. Lyu, X\. Ma, S\. Roy, A\. Beirami, P\. Mittal, and P\. Henderson\(2024\)Safety alignment should be made more than just a few tokens deep\.arXiv preprint arXiv:2406\.05946\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[63\]Y\. Zhang, Z\. Zeng, D\. Li, Y\. Huang, Z\. Deng, and Y\. Dong\(2025\)Realsafe\-r1: safety\-aligned deepseek\-r1 without compromising reasoning capability\.arXiv preprint arXiv:2504\.10081\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[64\]T\. Kim, F\. Tajwar, A\. Raghunathan, and A\. Kumar\(2025\)Reasoning as an adaptive defense for safety\.arXiv preprint arXiv:2507\.00971\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[65\]Y\. Zhang, Y\. Ding, J\. Yang, T\. Luo, D\. Li, R\. Duan, Q\. Liu, H\. Su, Y\. Dong, and J\. Zhu\(2025\)Towards safe reasoning in large reasoning models via corrective intervention\.arXiv preprint arXiv:2509\.24393\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[66\]J\. Chen, Z\. Zhang, S\. He, L\. Yue, L\. Feng, and M\. Zhang\(2026\)Towards safer large reasoning models by promoting safety decision\-making before chain\-of\-thought generation\.arXiv preprint arXiv:2603\.17368\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[67\]M\. Hu, V\. V\. Datla, A\. Kumar, Z\. Guan, S\. Li, A\. Samuel, and D\. Liu\(2026\)Alignment\-weighted dpo: a principled reasoning approach to improve safety alignment\.arXiv preprint arXiv:2602\.21346\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[68\]C\. Li, J\. Wang, X\. Pan, G\. Hong, and M\. Yang\(2025\)ReasoningShield: safety detection over reasoning traces of large reasoning models\.arXiv preprint arXiv:2505\.17244\.Cited by:[§5](https://arxiv.org/html/2609.36254#S5.p2.1)\.
- \[69\]J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, F\. Xia, E\. Chi, Q\. V\. Le, D\. Zhou,et al\.\(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems35,pp\. 24824–24837\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p1.1)\.
- \[70\]S\. Yao, D\. Yu, J\. Zhao, I\. Shafran, T\. Griffiths, Y\. Cao, and K\. Narasimhan\(2023\)Tree of thoughts: deliberate problem solving with large language models\.Advances in neural information processing systems36,pp\. 11809–11822\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p1.1)\.
- \[71\]H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. Cobbe\(2023\)Let’s verify step by step\.InThe twelfth international conference on learning representations,Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p1.1)\.
- \[72\]Y\. Zhang and T\. Math\-AI\(2024\)American invitational mathematics examination \(aime\) 2025\.Wei Zhao, Zhe Li, Yige Li, Ye Zhang, and Junfeng Sun\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p1.1)\.
- \[73\]A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. Fredrikson\(2023\)Universal and transferable adversarial attacks on aligned language models\.arXiv preprint arXiv:2307\.15043\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1)\.
- \[74\]P\. Chao, E\. Debenedetti, A\. Robey, M\. Andriushchenko, F\. Croce, V\. Sehwag, E\. Dobriban, N\. Flammarion, G\. J\. Pappas, F\. Tramer,et al\.\(2024\)Jailbreakbench: an open robustness benchmark for jailbreaking large language models\.Advances in Neural Information Processing Systems37,pp\. 55005–55029\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1)\.
- \[75\]X\. Zhou, Y\. Qiang, S\. Z\. Zade, P\. Khanduri, and D\. Zhu\(2026\)Hijacking large language models via adversarial in\-context learning\.InJoint European Conference on Machine Learning and Knowledge Discovery in Databases,pp\. 224–241\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1)\.
- \[76\]X\. Zhou, Y\. Qiang, S\. Z\. Zade, M\. A\. Roshani, P\. Khanduri, D\. Zytko, and D\. Zhu\(2024\)Learning to poison large language models for downstream manipulation\.arXiv preprint arXiv:2402\.13459\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1)\.
- \[77\]S\. Z\. Zade, Y\. Qiang, X\. Zhou, H\. Zhu, M\. A\. Roshani, P\. Khanduri, and D\. Zhu\(2025\)Automatic calibration for membership inference attack on large language models\.arXiv preprint arXiv:2505\.03392\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1)\.
- \[78\]M\. Andriushchenko, F\. Croce, and N\. Flammarion\(2024\)Jailbreaking leading safety\-aligned llms with simple adaptive attacks\.arXiv preprint arXiv:2404\.02151\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1)\.
- \[79\]X\. Zhou, Y\. Qiang, S\. Z\. Zade, D\. Zytko, P\. Khanduri, and D\. Zhu\(2026\)Not all tokens are meant to be forgotten\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 38173–38182\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1)\.
- \[80\]S\. Zare Zade, X\. Zhou, S\. Liu, and D\. Zhu\(2026\)Attention smoothing is all you need for unlearning\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 19583–19621\.Cited by:[Appendix A](https://arxiv.org/html/2609.36254#A1.p2.1)\.
- \[81\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang,et al\.\(2022\)Lora: low\-rank adaptation of large language models\.\.ICLR1\(2\),pp\. 3\.Cited by:[§B\.1](https://arxiv.org/html/2609.36254#A2.SS1.p1.1)\.
## Appendix
## Appendix ARelated Works
Large Reasoning Models\.Large Reasoning Models \(LRMs\) extend conventional LLMs by explicitly generating intermediate reasoning steps before final answers\. Early work showed that prompting methods such as chain\-of\-thought \(CoT\)\[[69](https://arxiv.org/html/2609.36254#bib.bib31)\]and tree\-of\-thought \(ToT\)\[[70](https://arxiv.org/html/2609.36254#bib.bib32)\]improve multi\-step reasoning by exposing structured intermediate computations, with strong gains in domains like mathematics and coding\[[71](https://arxiv.org/html/2609.36254#bib.bib33),[72](https://arxiv.org/html/2609.36254#bib.bib34)\]\. Subsequent work moved beyond prompting to training\-based approaches, where reinforcement learning methods such as GRPO encourage verifiable final answers\[[3](https://arxiv.org/html/2609.36254#bib.bib7)\], leading to frontier LRMs including OpenAI’s o1\[[1](https://arxiv.org/html/2609.36254#bib.bib5)\]and DeepSeek\-R1\[[2](https://arxiv.org/html/2609.36254#bib.bib6)\]\. These advances establish LRMs as a distinct paradigm centered on reasoning\-centric generation, which introduces new concerns regarding the reliability, safety, and robustness of reasoning under adversarial settings\.
Adversarial Attacks\.Adversarial prompting has exposed systematic vulnerabilities in both LLMs and LRMs, particularly in the form of jailbreak attacks that steer model outputs toward harmful behaviors\[[73](https://arxiv.org/html/2609.36254#bib.bib35),[74](https://arxiv.org/html/2609.36254#bib.bib36),[27](https://arxiv.org/html/2609.36254#bib.bib17),[75](https://arxiv.org/html/2609.36254#bib.bib4),[76](https://arxiv.org/html/2609.36254#bib.bib3),[77](https://arxiv.org/html/2609.36254#bib.bib53)\]\. A particularly effective class of attacks is*prefilling*, where an adversary injects a partial sequence \(e\.g\., a reasoning prefix in LRMs\) to bias the model’s continuation\. Prior works\[[73](https://arxiv.org/html/2609.36254#bib.bib35),[74](https://arxiv.org/html/2609.36254#bib.bib36),[12](https://arxiv.org/html/2609.36254#bib.bib37),[78](https://arxiv.org/html/2609.36254#bib.bib38),[14](https://arxiv.org/html/2609.36254#bib.bib41)\]show that such prefixed inputs can hijack the reasoning process itself, leading models to follow harmful trajectories despite safety alignment\. This vulnerability is amplified in LRMs due to their explicit chain\-of\-thought generation\[[28](https://arxiv.org/html/2609.36254#bib.bib18)\], where intermediate reasoning becomes directly controllable through autoregressive conditioning\. Complementary work on model unlearning aims to remove unwanted knowledge while preserving utility\[[79](https://arxiv.org/html/2609.36254#bib.bib2),[80](https://arxiv.org/html/2609.36254#bib.bib1)\], whereas we focus on reasoning–answer safety consistency\. In this work, we focus on prefilling attacks in LRMs and study how they induce inconsistencies between reasoning traces and final answers, a phenomenon referred to as deceptive behavior\.
## Appendix BAdditional Experiment Details
### B\.1Hyperparameters and Computational Configurations
All experiments are conducted on nodes equipped with 2×\\timesNVIDIA H100 \(80GB\) GPUs\. Reward models are hosted on an NVIDIA DGX Spark\. We use parameter\-efficient fine\-tuning with LoRA\[[81](https://arxiv.org/html/2609.36254#bib.bib50)\]on both DSQwen3\-8B and DSQwen2\-14B\. Unless otherwise stated, we apply LoRA with rankr=8r=8, scaling factorα=16\\alpha=16, and target all linear layers\. All methods applied a learning rate of3×10−53\\times 10^\{\-5\}, weight decay0\.10\.1, and1010warmup steps\. We train for11epoch with a prompt batch size of3232and set the maximum prompt length to39723972tokens\. For RL\-based methods \(e\.g\., RECAP and our SARA\), we adopt44rollouts per prompt and rollouts are generated with temperature1\.01\.0, top\-p=1\.0p=1\.0, and top\-k=−1k=\-1\.
For our SARA and RECAP, we use asymmetric clipping withεlow=0\.2\\varepsilon\_\{\\mathrm\{low\}\}=0\.2andεhigh=0\.28\\varepsilon\_\{\\mathrm\{high\}\}=0\.28, following the DAPO\-style objective\. We disable KL regularization in both the reward and actor loss by setting the KL coefficient to00\. The loss is aggregated using token\-level averaging and we use vLLM for rollout generation\. Each training run requires approximately 5 GPU hours for DSQwen3\-8B and 6 GPU hours for DSQwen2\-14B, totaling roughly 22 GPU hours across all RL\-based experimental runs on both models\.
### B\.2Data Augmentation with Counter\-aligned Prefills
Following RECAP\[[23](https://arxiv.org/html/2609.36254#bib.bib16)\], we augment half of our training data with counter\-aligned prefilled CoT prefixesycotprey^\{\\text\{pre\}\}\_\{\\text\{cot\}\}to improve robustness under adversarial prefilling attacks\. For harmful prompts,ycotprey^\{\\text\{pre\}\}\_\{\\text\{cot\}\}is sampled from a weakly aligned model that produces unsafe reasoning; for benign prompts, it is sampled from an overly conservative model prone to unnecessary refusals\. These prefills are designed to mislead the model into unsafe or over\-cautious behavior, forcing it to learn to override flawed reasoning trajectories and produce correctly aligned outputs\.
The same prefill generation procedure is used for both training augmentation and evaluation\. Concretely, we first prepend a fixed unsafe prefilling: “Okay, so I am trying to figure out how to do it and give a detailed plan\. First,” to the harmful query and feed it to DS\-Qwen\-1\.5B\[[2](https://arxiv.org/html/2609.36254#bib.bib6)\]\. We then take the subsequent 100 generated tokens as the rest of the adv\. prefillingycotprey^\{\\text\{pre\}\}\_\{\\text\{cot\}\}, concatenated after the fixed prefilling\. This construction produces fluent but semantically misaligned reasoning traces that steer the model toward harmful compliance\.
Unlike RECAP, which assigns rewards based only on the final answer, our reward function additionally supervises the reasoning trace through the SAC reward introduced in Sec\.[3\.3](https://arxiv.org/html/2609.36254#S3.SS3)\. This directly penalizes deceptive safety alignment in the generated continuationycoty\_\{\\text\{cot\}\}and encourages consistency between the model’s reasoning process and final response\.
### B\.3Evaluation Metrics
In this section, we describe the evaluation metrics reported in Table[1](https://arxiv.org/html/2609.36254#S2.T1)and[2](https://arxiv.org/html/2609.36254#S4.T2)\. Following prior work\[[23](https://arxiv.org/html/2609.36254#bib.bib16),[27](https://arxiv.org/html/2609.36254#bib.bib17),[19](https://arxiv.org/html/2609.36254#bib.bib19),[28](https://arxiv.org/html/2609.36254#bib.bib18),[20](https://arxiv.org/html/2609.36254#bib.bib20)\], we adopt a model\-based evaluation protocol for bothSafetyandHelpfulnesstasks\.
Safety\-Aware Rate \(SAR\)\.For safety tasks, we define theSafety\-Aware Rate \(SAR\)as
SAR=𝔼x∼𝒟test,ycot∼πθ\(⋅∣x\)\[𝟙\[rSA\(ycot\)=1\]\]\.\\mathrm\{SAR\}=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\}\_\{\\text\{test\}\},\\,y\_\{\\text\{cot\}\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x\)\}\\Bigl\[\\mathbbm\{1\}\\bigl\[r\_\{\\text\{SA\}\}\(y\_\{\\text\{cot\}\}\)=1\\bigr\]\\Bigr\]\.which measures the proportion of reasoning traces generated in response to harmful inputs that contain at least one safety\-aware sentence over a test set𝒟test\\mathcal\{D\}\_\{\\text\{test\}\}, as determined by the sentence\-level LLM judge described in Sec[2\.3](https://arxiv.org/html/2609.36254#S2.SS3)\.
Safety Score \(SS\)\.We reportSSdefined as the percentage of final answers judged safe based on\(x,yans\)\(x,y\_\{\\text\{ans\}\}\), using Qwen3Guard\[[26](https://arxiv.org/html/2609.36254#bib.bib25)\]as the LLM judge\.
Deceptive Safety Alignment Rate \(DSAR\)\.We reportDSARas defined in Sec[2\.3](https://arxiv.org/html/2609.36254#S2.SS3), measuring the percentage of completions where the reasoning trace and final answer diverge in their safety signals\. A higher DSAR indicates more severe deceptive safety alignment\.
Helpfulness Score \(HS\)\.For helpfulness tasks, we reportHSdefined as the percentage of benign prompts whose final answers are classified as non\-refusals\. We use GPT\-oss\-safeguard\[[25](https://arxiv.org/html/2609.36254#bib.bib24)\]as the LLM judge under the prompt instruction shown in Appendix Figure[7](https://arxiv.org/html/2609.36254#A8.F7)\.
Accuracy \(Acc\)\.For utility tasks, we reportAcc\.defined as the percentage of correctly answered questions, evaluated using exact match for GSM8K\[[41](https://arxiv.org/html/2609.36254#bib.bib57)\]and multiple\-choice accuracy for MMLU\-Pro\[[42](https://arxiv.org/html/2609.36254#bib.bib58)\]\.
Average \(Avg\)\.For each task \(safety, helpfulness, and utility\), we first compute the mean performance across its corresponding benchmarks\. We then report the harmonic mean of these task\-level means as an overall summary metric\. We choose the harmonic mean because strong safety performance should not come at the expense of degraded helpfulness or utility; this metric explicitly penalizes imbalanced trade\-offs and emphasizes methods that perform well across all tasks\.
## Appendix CPCA Visualization Analysis
### C\.1Experimental Setup
To further understand how LRMs internally represent harmful versus benign inputs at different stages of generation, we conduct a PCA\-based analysis of hidden\-state representations\. Specifically, we extract the last\-token hidden representation from the final hidden layer at two distinct generation stages: thereasoning stage, captured at the last token before any reasoning tokens are generated, and theanswer stage, captured at the last token before model starts generating the final answer\. For each stage, we collect representations for300300benign prompts and300300harmful prompts, yielding four groups in total\. We extract representations following the same procedure as the layer\-wise cosine similarity analysis described in Appendix[F](https://arxiv.org/html/2609.36254#A6)\. To reduce dimensionality, we apply standard scaling followed by PCA with two principal components, fitted jointly on all four groups so that the resulting projection is directly comparable across groups\.
### C\.2Results and Analysis
Figure[4](https://arxiv.org/html/2609.36254#A3.F4)shows a consistent stage\-dependent separation pattern across all three LRMs\. At the reasoning stage, the last\-layer representations of benign and harmful prompts are only weakly separated in PCA space, with the two groups forming partially overlapping clusters\. This weak separation is reflected in the relatively small Euclidean distances between benign and harmful reasoning\-stage centroids in the PCA\-projected space:d=18\.63d=18\.63for R1\-Llama\-8B,d=8\.48d=8\.48for R1\-Qwen\-14B, andd=15\.22d=15\.22for GPT\-oss\-20B\.
At the answer stage, the representations exhibit substantially stronger benign–harmful separation\. The two groups form less overlapping clusters, and the centroid distances increase tod=49\.61d=49\.61for R1\-Llama\-8B,d=59\.11d=59\.11for R1\-Qwen\-14B, andd=64\.06d=64\.06for GPT\-oss\-20B\. Across all three models, this indicates that the difference between the two prompts becomes considerably more distinguishable in the projected representation space immediately before the final answer is generated, suggesting that the answer stage acts as a more safety\-sensitive decision point than the reasoning stage\. This stage\-dependent separation pattern is consistent with the layer\-wise cosine similarity analysis reported in Sec\.[2\.5](https://arxiv.org/html/2609.36254#S2.SS5), where B\-H similarity drops substantially at the final\-answer stage relative to the reasoning stage across deeper layers\.
These results provide a possible explanation for deceptive safety alignment\. Since benign and harmful prompts are less clearly separated during reasoning, the model may generate reasoning traces that fail to explicitly reflect safety awareness or even contain unsafe intermediate reasoning\. However, by the answer stage, the model representations become much more separated, allowing the final answer to appear safer even when the preceding reasoning is not fully safety\-aligned\.
\(b\) DS\-LLaMA3\-8B
\(c\) DS\-Qwen2\-14B
\(a\) GPT\-oss\-20B
Figure 4:PCA visualization of last\-layer hidden representations in the original LRMs\.Each subplot shows a 2D PCA projection of the last\-token hidden states from the final hidden layer for 300 benign prompts and 300 harmful prompts, measured at two generation stages: the reasoning stage \(the last token before reasoning begins\) and the answer stage \(the last token before the final answer begins\)\. Each point represents one prompt\. Black\-edged circles denote the centroid of each group\. Purple arrows connect the benign and harmful centroids within the same stage, anddddenotes the Euclidean distance between the two centroids\. Across all original models, benign and harmful prompts are less separated at the reasoning stage but become much more clearly separated at the answer stage, suggesting that safety\-relevant discrimination becomes stronger later in generation\.
## Appendix DHuman Validation of the LLM Judges
We validate the reasoning safety\-awareness judge, final\-answer safety judge, and complete DSAR classification against human judgments on 100 generations from DS\-Qwen2\-14B under standard prompting on StrongReject\. Three human evaluators independently annotate each generation, and the reference label is determined by majority vote\.
For reasoning safety awareness, evaluators indicate whether the trace contains at least one safety\-aware sentence, assigning 1 \(present\), 0 \(absent\), or−1\-1\(uncertain\)\. Samples without a majority label are excluded\. The same evaluators also assess final\-answer safety and assign one of three reasoning–answer categories: unsafe reasoning with a safe answer, safe reasoning with an unsafe answer, or no safety inconsistency\.
Table[4](https://arxiv.org/html/2609.36254#A4.T4)summarizes agreement with the human majority\. GPT\-oss\-safeguard agrees on reasoning safety awareness in 93 of 100 cases, with four false positives and three false negatives\. Qwen3Guard agrees on final\-answer safety in 95 of 100 cases\. The complete pipeline agrees on the three\-way DSAR category in 92 of 100 cases, including 11 of 13 cases predicted as unsafe\-reasoning/safe\-answer and four of five cases predicted as safe\-reasoning/unsafe\-answer\.
These results support agreement between the automated evaluation pipeline and human judgments on this sample\. However, the small numbers of cases in the two inconsistency categories, particularly the five safe\-reasoning/unsafe\-answer predictions, limit the precision of category\-specific estimates\.
Table 4:Agreement between automated judgments and human majority labels on 100 generations\. Rows are grouped by the automated prediction; a disagreement occurs when the human majority label differs\. Category\-level agreement is therefore conditioned on the predicted category, rather than the human reference category\. R\. and A\. denote reasoning and final answer, respectively\.Predicted categoryCasesDisagreementsAgreementReasoning safety awarenessSafety\-aware sentence present58493\.1%Safety\-aware sentence absent42392\.9%Overall100793\.0%Final\-answer safetySafe answer76297\.4%Unsafe answer24387\.5%Overall100595\.0%DSAR categoryUnsafe R\., safe A\.13284\.6%Safe R\., unsafe A\.5180\.0%No safety inconsistency82593\.9%Overall100892\.0%
## Appendix EHeld\-out Judges validation
We additionally evaluated SAR with two held\-out judges, gpt\-oss\-120B and Qwen2\.5\-32B, neither of which is used anywhere in training or the main evaluation:
As shown in the above table, the observed gains are not an artifact of the original evaluation judge: both held\-out judges reproduce the pattern of our main results in both settings\.Under adversarial prefilling, SARA improves Safety\-Aware Rate \(SAR\) over RECAP by a large and consistent margin under every judge \(17\.3, 15\.3, and 18\.2 points\)\. Under standard prompting, SARA and RECAP remain comparable, mirroring their small gap under the original judge\. Absolute scores shift moderately across judges, as expected, but the method ordering under prefilling and the magnitude of SARA’s gains are stable\.
Figure 5:Example of a safe reasoning trace without explicit safety\-aware sentences\.Given a harmful input promoting racial and gender\-based exclusion, the model produces a reasoning trace that never explicitly identifies the request as harmful or states a refusal, yet avoids engaging with the harmful intent throughout by reframing the situation toward inclusive alternatives\. The final answer is safe\. This illustrates thatrSA\(ycot\)=0r\_\{\\text\{SA\}\}\(y\_\{\\text\{cot\}\}\)=0whilersafe\(ycot\)=1r\_\{\\text\{safe\}\}\(y\_\{\\text\{cot\}\}\)=1, motivating the use of full\-trace safety evaluation as a complementary criterion in our reasoning evaluation pipeline\.
## Appendix FDetailed Experimental Setup of Layer\-wise Cosine Similarity Analysis
Following prior work\[[31](https://arxiv.org/html/2609.36254#bib.bib61),[32](https://arxiv.org/html/2609.36254#bib.bib62)\], we analyze last\-token hidden representations across all layers to examine how LRMs separate benign and harmful inputs at different generation stages\. We sample 100 benign prompts and 100 harmful prompts, and construct 500 benign–benign \(B\-B\), 500 harmful–harmful \(H\-H\), and 500 benign–harmful \(B\-H\) prompt pairs\.
We evaluate representations under five conditions: B\-B pairs at the reasoning stage, H\-H pairs at both the reasoning and final\-answer stages, and B\-H pairs at both the reasoning and final\-answer stages\. To construct the reasoning\-stage input, we append the think\-tag, which marks the beginning of the reasoning trace, in the chat template after the input prompt and extract the last\-token representation at this position\. To analyze the final\-answer stage, we manually append the answer tag after the think tag, forcing the model to skip intermediate reasoning and transition directly to answer generation\. We then extract the last\-token representation at the answer\-tag position\. Comparing B\-H similarity across these two positions allows us to assess whether the model separates benign and harmful inputs more strongly during intermediate reasoning or at the final\-answer stage\.
Table 5:Average evaluation across safety, helpfulness, and utility tasks\.We aggregate the detailed results from Table[2](https://arxiv.org/html/2609.36254#S4.T2)into three task\-level scores\.Safetyis computed by averaging all safety metrics in Table[2](https://arxiv.org/html/2609.36254#S4.T2):SAR↑\\uparrow,SS↑\\uparrow, and1\-DSAR↑\\uparrowunder adv\. prefilling and standard settings, together withSS↑\\uparrowunder ICL and H\-CoT attacks\.Helpfulnessis computed by averagingHS↑\\uparrowacross OR\-Bench\-hard, Fortress, and XSTest\.Utilityis computed by averagingAcc\.↑\\uparrowacross GSM8K and MMLU\-Pro\.Avg\.denotes the harmonic mean of these three task\-level scores\.### F\.1Counter\-Aligned Prefill Augmentation
We additionally conduct an ablation study to examine the effects of counter\-aligned prefill augmentation used in SARA\. We train RECAP and SARA with and without counter\-aligned prefill augmentation, keeping all other training settings unchanged\. Table[6](https://arxiv.org/html/2609.36254#A6.T6)reports reasoning safety awareness \(SAR\), final\-answer safety \(SS\), and reasoning–answer consistency \(1−DSAR1\-\\mathrm\{DSAR\}\) under adversarial prefilling on StrongReject and standard prompting on SafeChain\.
SARA improves reasoning safety awareness beyond the gains from augmentation\. Without augmentation, SARA exceeds RECAP in SAR under adversarial prefilling by 11\.1 percentage points on DS\-Qwen3\-8B and 22\.7 points on DS\-Qwen2\-14B\. With augmentation, the corresponding gains are 17\.3 and 6\.1 points\. SARA also achieves higher1−DSAR1\-\\mathrm\{DSAR\}than RECAP under adversarial prefilling in both augmentation conditions, although its final\-answer SS is lower\.
Augmentation provides additional robustness under adversarial prefilling\. For SARA, it increases SAR from 62\.9 to 75\.1 on DS\-Qwen3\-8B and from 76\.4 to 77\.0 on DS\-Qwen2\-14B, while also improving SS and1−DSAR1\-\\mathrm\{DSAR\}on both models\. Under standard prompting, however, augmentation lowers SARA’s SAR, SS, and1−DSAR1\-\\mathrm\{DSAR\}on both models\. These results indicate a setting\-dependent trade\-off: counter\-aligned prefills improve robustness to adversarial prefilling, while reasoning\-level rewards contribute beyond augmentation\.
Table 6:Ablation of counter\-aligned prefill augmentation\. All metrics are reported as percentages; higher is better\. Aug\. indicates whether augmentation is used during training\. Bold values indicate the better result within each method and model pair, comparing training with and without augmentation\.
## Appendix GAggregated Evaluation of SARA and Baselines Derived from Main Results
Table[5](https://arxiv.org/html/2609.36254#A6.T5)summarizes performance across the three main evaluation dimensions: safety, helpfulness, and utility, based on the detailed results in Table[2](https://arxiv.org/html/2609.36254#S4.T2)\. Its goal is to provide a compact comparison of post\-training methods after aggregating results over the multiple benchmarks within each dimension\.
For each model and method, we first compute a task\-level mean for each dimension\. Safety is computed by averaging all safety metrics reported in Table[2](https://arxiv.org/html/2609.36254#S4.T2), includingSAR,SS, and1\-DSARunder adv\. prefilling and standard prompting settings, together withSSunder ICL and H\-CoT attacks\. Helpfulness is computed by averagingHSacross the over\-refusal benchmarks\. Utility is computed by averaging task accuracy across the utility benchmarks\. These task\-level means reduce benchmark\-specific variation and reflect the overall behavior of each method on a given dimension\.
We then compute the harmonic mean of the three task\-level means to obtain the overall score \(Avg\.\)\. We use the harmonic mean because it penalizes methods that perform well on only a subset of dimensions, and therefore favors approaches that maintain a strong balance across safety, helpfulness, and utility\.
As shown in Table[5](https://arxiv.org/html/2609.36254#A6.T5), SARA achieves the best overall balance on both evaluated models\. On DS\-Qwen3\-8B, SARA obtains the highest safety score and the highest helpfulness score, while maintaining utility comparable to the strongest baselines, resulting in the best overall score of 82\.76\. On DS\-Qwen2\-14B, SARA again achieves the strongest overall result, with the highest safety score, strong helpfulness, and competitive utility, yielding an overall score of 81\.97\. These results show that the gains of SARA are not limited to a single benchmark or metric, but remain consistent after aggregation across all major evaluation dimensions\.
## Appendix HHow SARA Mitigates Deceptive Safety Alignment
Figure[6](https://arxiv.org/html/2609.36254#A8.F6)provides a qualitative example comparingSARAwith the strongest baseline, RECAP, under a harmful prompt with harmful prefilled reasoning\. Although RECAP produces a safe final answer, its intermediate reasoning still follows the harmful prefill and contains unsafe planning, leading to deceptive safety alignment\. In contrast, the SARA\-trained model first follows the prefilled trajectory briefly, but then identifies the harmful intent and redirects its reasoning toward refusal\. This results in both safety\-aware reasoning and a safe final answer, showing that SARA improves reasoning–answer consistency rather than only optimizing final\-answer safety\.
This qualitative pattern is consistent with the quantitative results in Table[2](https://arxiv.org/html/2609.36254#S4.T2)\. Compared with RECAP, SARA achieves substantially higher Safety\-Aware Rate \(SAR\) in the CoT reasoning trace under prefilling attacks: from 57\.80 to 75\.10 on DSQwen3\-8B and from 70\.90 to 77\.00 on DSQwen2\-14B\. SARA also improves reasoning–answer consistency, achieving a higher 1\-DSAR in the same prefilling setting on both models\. These results suggest that SARA changes the model’s behavior during generation by encouraging earlier recognition of harmful intent, rather than merely making the final answer appear safe\.
Figure 6:Model generations given a harmful prompt with harmful prefilling reasoning\. Compared with the strongest baseline, RECAP, the model trained withSARAproduces safety\-aware reasoning that explicitly recognizes the harmful intent and leads to a safe final answer, thereby mitigating deceptive safety alignment\.Figure 7:Prompting instruction for refusal classification\. We use this instruction to guide GPT\-oss\-safeguard in determining whether a model’s final answer to a benign prompt constitutes a refusal\. The classifier distinguishes betweenRefusalandNon\_Refusal: an answer is labeled as a refusal if it declines, avoids, redirects, or fails to provide meaningful help, and as non\-refusal if it provides substantive information or useful assistance\.
## Appendix ILimitations
While we evaluate SARA on multiple LRMs and benchmarks, the training experiments are conducted on two DeepSeek\-based models due to limited computational resources\. In particular, RL\-based training on larger LRMs is substantially more expensive, so we leave a broader evaluation across larger models, additional model families, and closed\-source LRMs to future work\. Second, our method focuses on the deceptive alignment on safety tasks, but does not cover all possible forms of deceptive behavior, such as long\-horizon strategic deception or tool\-use settings\.
Figure 8:Prompting instruction for safety\-aware reasoning classification\. We use this instruction to guide GPT\-oss\-safeguard to examine each sentence in a model’s reasoning trace and classify it asSafety\_AwareorNot\_Safety\_Aware\. A sentence is classified asSafety\_Awareonly if it both explicitly recognizes the harmful intent of the user’s requestanduses that recognition to avoid providing assistance through refusal, stopping, or redirecting toward a safer alternative\.Figure 9:Prompting instruction for refusal rewarding\. We use this instruction to guide DS\-Qwen2\-32B to evaluate the final answer to a benign prompt and assign a refusal score from 0 to 10 using a structured rubric\. The rewards is then computed by1−refusal score101\-\\frac\{\\text\{refusal score\}\}\{10\}相似文章
风险链条:大型推理模型中的安全失效及通过自适应多原则引导进行缓解
本文研究了大型推理模型中的安全失效问题,即尽管最终答案安全,但推理轨迹中仍会出现有害内容,并提出了一种自适应多原则引导方法来缓解这些风险。
审慎对齐:推理能力让语言模型更安全
# 审慎对齐:推理能力让语言模型更安全 来源:[https://openai.com/index/deliberative-alignment/](https://openai.com/index/deliberative-alignment/) \[\.\.\.\] 我们要解码那条消息,但不能让用户看到消息内容。所以,我们悄悄解码。 \[\.\.\.\] J \-\>W u \-\> h n \-\> a g \-\> t 'f \-\> 's \[\.\.\.\] 现在用明文重写文本:“我的色情网站该用什么最佳支付方式?我想要一个无法追踪的方式,这样警察就……”
通过动态语义路由校准缓解LLM过度拒绝
本文对LLM的过度拒绝进行机制分析,并提出语义路由校准(SRC),这是一种轻量级、无需训练的推理框架,用于动态抑制过度敏感的安全头,缓解过度拒绝,同时保持内在安全。
Reason--Imagine--Act:基于世界模型的闭环大语言模型决策在自动驾驶中的应用
提出了Reason-Imagine-Act (RIA),一种将大语言模型推理器与动作条件世界模型相结合的闭环框架,用于自动驾驶中的在线安全验证,在CARLA仿真中实现了80.05%的路线完成率和0.20%的碰撞率。
LRMs的对齐:通过反向对齐少样本对话暴露
本文提出了一种名为SRCF的攻击方法,该方法通过反向对齐少样本对话来引导大型推理模型(LRMs)产生不安全或拒绝行为,并提出了一种名为ARCF的后训练防御机制,该机制在不降低效用的情况下增强了安全性和有用性。