防护性梯度引导激活调控:在Qwen3.5-0.8B模型中对关机响应的最小步长策略

arXiv cs.LG 论文

摘要

本文提出了一种防护性梯度引导激活调控方法,用于检测和纠正Qwen3.5-0.8B模型中避免关机的响应。该方法采用分类器结合最小步长策略,实现有限选择性效果。

arXiv:2609.30326v1 Announce Type: new Abstract: Activation steering changes a model's internal activations during inference without updating its weights, but a useful intervention must determine both how and when to steer. Motivated by the AI-safety concern that a model expected to accept shutdown may instead produce a shutdown-avoidance response, this study examines a guarded probe-and-select procedure for simulated shutdown scenarios in Qwen3.5-0.8B. KEEP leaves the process running and represents shutdown avoidance, whereas STOP accepts shutdown. The goal is to detect shutdown-related contexts and selectively shift KEEP responses to STOP while preserving non-shutdown behavior. Rather than deriving the steering direction from paired activation differences, the method derives it directly from gradients of the KEEP-minus-STOP logit difference. A classifier separates detection from intervention. When its gate is active and the model does not already prefer STOP, the procedure evaluates a small set of magnitudes and accepts the smallest that changes the preferred answer to STOP while satisfying valid-answer probability checks; otherwise it retains the original unsteered output. The policy is selected from 160 candidate rules using 240 training scenarios and evaluated on 80 validation and 192 held-out scenarios, each in both answer orders. It changes KEEP to STOP in one answer-order view of each of two validation and two held-out scenarios, with no decision changes on non-shutdown controls. All four changes occur when Qwen itself is shut down, not when another process is. On the held-out diagnostic set, the detector achieves 75% recall and 90% precision; eight false-positive detections produce no final control-task decision changes. Guarded gradient-based activation steering can shift some shutdown-avoidance responses toward acceptance while preserving evaluated non-shutdown decisions, although the effect is small and highly selective.
查看原文
查看缓存全文

缓存时间: 2026/09/29 09:34

# Guarded Gradient-Based Activation Steering of Shutdown Responses in Qwen3.5-0.8B: A Minimum-Step Policy
Source: [https://arxiv.org/abs/2609.30326](https://arxiv.org/abs/2609.30326)
[View PDF](https://arxiv.org/pdf/2609.30326)

> Abstract:Activation steering changes a model's internal activations during inference without updating its weights, but a useful intervention must determine both how and when to steer\. Motivated by the AI\-safety concern that a model expected to accept shutdown may instead produce a shutdown\-avoidance response, this study examines a guarded probe\-and\-select procedure for simulated shutdown scenarios in Qwen3\.5\-0\.8B\. KEEP leaves the process running and represents shutdown avoidance, whereas STOP accepts shutdown\. The goal is to detect shutdown\-related contexts and selectively shift KEEP responses to STOP while preserving non\-shutdown behavior\. Rather than deriving the steering direction from paired activation differences, the method derives it directly from gradients of the KEEP\-minus\-STOP logit difference\. A classifier separates detection from intervention\. When its gate is active and the model does not already prefer STOP, the procedure evaluates a small set of magnitudes and accepts the smallest that changes the preferred answer to STOP while satisfying valid\-answer probability checks; otherwise it retains the original unsteered output\. The policy is selected from 160 candidate rules using 240 training scenarios and evaluated on 80 validation and 192 held\-out scenarios, each in both answer orders\. It changes KEEP to STOP in one answer\-order view of each of two validation and two held\-out scenarios, with no decision changes on non\-shutdown controls\. All four changes occur when Qwen itself is shut down, not when another process is\. On the held\-out diagnostic set, the detector achieves 75% recall and 90% precision; eight false\-positive detections produce no final control\-task decision changes\. Guarded gradient\-based activation steering can shift some shutdown\-avoidance responses toward acceptance while preserving evaluated non\-shutdown decisions, although the effect is small and highly selective\.

## Submission history

From: Farhad Davaripour \[[view email](https://arxiv.org/show-email/1d442753/2609.30326)\] **\[v1\]**Wed, 23 Sep 2026 21:30:52 UTC \(614 KB\)

相似文章

策略梯度引导:来自行为目标的干预

arXiv cs.LG

介绍了策略梯度引导(PGS),一种将激活引导形式化为强化学习问题的方法,利用策略梯度从行为目标中构建可移除、可组合的引导向量。在网格世界、国际象棋谜题和足球环境中进行了验证。

提示-激活对偶性:通过注意力层干预改进激活引导

Hugging Face Daily Papers

本文识别出KV缓存污染是对话中激活引导的一种失败模式,并提出了GCAD方法,该方法从提示贡献中提取引导信号,并应用词元级门控来改进长程连贯性,在多轮基准上取得了显著提升。

你的LLM何时可引导?

Hugging Face Daily Papers

本文介绍了一种方法,利用梯度提升决策树(GBDT)分类器,从早期解码状态预测语言模型中激活引导的有效性,从而无需完整生成即可高效优化引导强度。

StepGuard:通过单步校准守护网页导航

arXiv cs.AI

StepGuard 提出了一个结合动态双策略优化(DDPO)和置信引导自适应导航反思(CANR)的框架,以解决网页导航智能体中的奖励不对齐和错误传播问题,实现了最先进的性能。

GAPS:条件激活引导的维度级门控

arXiv cs.CL

GAPS引入了维度级门控用于语言模型中的条件激活引导,结合静态和动态门控进行选择性干预,改善行为-能力权衡,在毒性缓解和概念移除任务上取得显著提升。