ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation

arXiv cs.CL Papers

Summary

ForeSight is a framework that predicts harmful outputs from large language models by analyzing first-token hidden states, improving early risk monitoring efficiency.

arXiv:2609.13737v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards operate at the input, output, or streaming-generation stages, while early-risk methods that rely on surface tokens or output logits may suffer from weak initial signals, and internals-based detectors using dense representations may retain highly entangled and redundant safety-irrelevant information. It therefore remains unclear whether the earliest post-generation hidden states already contain reliable signals about final-response harmfulness. To address this gap, we propose ForeSight, a first-token output-risk forecasting framework that distills weak and redundant early safety signals into compact, layer-aware risk representations. Experiments on five safety benchmarks and two target models demonstrate that ForeSight achieves superior and efficient early-risk forecasting while relying solely on first-token hidden states. The code is available at: https://github.com/Scabbards1500/Foresight
Original Article
View Cached Full Text

Cached at: 09/15/26, 08:41 AM

# Enhancing Risk Monitoring via Early Safety Signal Distillation
Source: [https://arxiv.org/html/2609.13737](https://arxiv.org/html/2609.13737)
Hanling Wang††thanks:Equal contribution\.Chenlong Wei11footnotemark:1Affiliation:Xi’an Jiaotong\-Liverpool UniversityEmail:[xiaohui\.zhu@xjtlu\.edu\.cn](mailto:)Ling XuAffiliation:Xi’an Jiaotong\-Liverpool UniversityEmail:[eezhuy@zju\.edu\.cn](mailto:)Hanyan NiuAffiliation:Xi’an Jiaotong\-Liverpool UniversityQi CaoAffiliation:Xi’an Jiaotong\-Liverpool UniversityShizhou HuangAffiliation:East China Normal UniversityYang YangAffiliation:Shanghai UniversityXiaohui ZhuAffiliation:Xi’an Jiaotong\-Liverpool UniversityYao Zhu††thanks:Corresponding author\.Affiliation:Zhejiang University

###### Abstract

As large language models \(LLMs\) are increasingly deployed, the generation of harmful content has become a critical safety concern\. Existing safeguards operate at the input, output, or streaming\-generation stages, while early\-risk methods that rely on surface tokens or output logits may suffer from weak initial signals, and internals\-based detectors using dense representations may retain highly entangled and redundant safety\-irrelevant information\. It therefore remains unclear whether the earliest post\-generation hidden states already contain reliable signals about final\-response harmfulness\. To address this gap, we propose ForeSight, a first\-token output\-risk forecasting framework that distills weak and redundant early safety signals into compact, layer\-aware risk representations\. Experiments on five safety benchmarks and two target models demonstrate that ForeSight achieves superior and efficient early\-risk forecasting while relying solely on first\-token hidden states\. The code is available at:[https://github\.com/Scabbards1500/Foresight](https://github.com/Scabbards1500/Foresight)

Disclaimer:This paper contains offensive content that may be disturbing to some readers\.

## 1Introduction

Figure 1:Overview of first\-token output\-risk estimation\. Unlike post\-hoc, streaming, and dense internals\-based detectors, ForeSight forecasts final\-response harmfulness from sparse safety\-relevant signals in first\-token hidden states\.As large language models \(LLMs\) are deployed at scale, harmful prompts can elicit toxic, biased, or dangerous outputs, posing significant societal risks\([Zou et al\., 2023b](https://arxiv.org/html/2609.13737#bib.bib1)\)\. It is therefore important to detect output\-level risks as early as possible during generation, before harmful content is fully produced\. Existing safety mechanisms address this challenge at different stages of generation\. Input\-side moderation filters out unsafe prompts before generation\([Mozafari et al\., 2020](https://arxiv.org/html/2609.13737#bib.bib4);[Caselli et al\., 2021](https://arxiv.org/html/2609.13737#bib.bib5)\), but prompt\-level risk does not necessarily predict output\-level harmfulness, as safety\-aligned models may refuse harmful prompts while jailbreaks or subtle reframing can still elicit unsafe responses\([Shen et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib2);[Wei et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib3)\)\. Response\-level guardrail models evaluate generated outputs\([Inan et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib45);[Han et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib13)\), but such detection is inherently post\-hoc\. To reduce this delay, streaming moderation approaches monitor partial generations in real time\([Zhao et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib46);[Li et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib11)\)\. Recent studies further explore whether harmfulness can be anticipated from early generated tokens or output logits\([Qi et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib17);[Hu et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib14);[Chen et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib15)\)\. However, early observable signals derived from either generated tokens or output logits are often weak at the beginning of generation, leaving it unclear whether the model’s earliest internal states already contain reliable information about final\-response harmfulness\.

We hypothesize that such early safety signals may already exist in first\-token hidden states, but remain sparse, distributed, and highly entangled within dense activations\. This intuition is supported by prior findings that semantic concepts can often be linearly decoded from model representations\([Alain and Bengio, 2016](https://arxiv.org/html/2609.13737#bib.bib34);[Hernandez et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib35);[Park et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib36);[Jiang et al\., 2024b](https://arxiv.org/html/2609.13737#bib.bib18)\)\. To extract these weak signals, we propose ForeSight, a sparse safety signal distillation framework for early risk estimation from the first generated token\. ForeSight identifies salient safety\-relevant neurons from first\-token hidden states, reduces redundant activations through structured aggregation, and combines complementary signals across layers to forecast final\-response harmfulness\.

Experiments on five safety benchmarks demonstrate its effectiveness in output\-risk forecasting\. Our main contributions are summarized as follows:

- •We introduce a first\-token output\-risk forecasting setting that predicts final\-response harmfulness using only the hidden states immediately following the first generated token\.
- •We propose ForeSight, a sparse early safety signal distillation framework that extracts salient neurons from first\-token hidden states, suppresses redundant activations, and aggregates distributed safety signals across layers for early harmfulness forecasting\.
- •Extensive experiments on five safety benchmarks demonstrate that first\-token hidden states contain meaningful predictive signals, and our sparse distillation approach outperforms strong dense baselines\.

## 2Related Work

![Refer to caption](https://arxiv.org/html/2609.13737v1/framework.png)Figure 2:Overview of ForeSight\. The framework predicts final\-response harmfulness from first\-token hidden states by selecting a sparse set of safety\-relevant neurons, clustering them based on training\-set activation profiles, aggregating weighted layer representations intog⁡\(x\)g\(x\), and training an MLP classifier for risk forecasting\. Output\-level labels are derived offline from full responses and used only during training\.### 2\.1Early Intervention Safety Mechanisms

Safety mechanisms for LLMs can be categorized by the stage at which harmfulness is detected\. Early approaches mainly classify either user prompts or fully generated responses, using encoder\-based toxicity or hate\-speech classifiers such as BERT\([Devlin et al\., 2019](https://arxiv.org/html/2609.13737#bib.bib6)\)and RoBERTa\([Caselli et al\., 2021](https://arxiv.org/html/2609.13737#bib.bib5);[Mozafari et al\., 2020](https://arxiv.org/html/2609.13737#bib.bib4);[Zhao et al\., 2021](https://arxiv.org/html/2609.13737#bib.bib8);[Markov et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib9)\), or generative guard models such as Llama Guard\([Inan et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib45)\)and WildGuard\([Han et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib13)\)\. To reduce the delay of post\-hoc response moderation, recent methods shift safety detection into the generation process, including response\-level guardrails, streaming moderation, and token\-level monitoring\([Zhao et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib46);[Zeng et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib20);[Ghosh et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib21);[Li et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib11);[Kavumba et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib12)\)\. Other studies attempt to forecast final\-response harmfulness from the first few generated tokens or their logits\([Hu et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib14);[Chen et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib15)\)\. However, these approaches still rely primarily on surface\-level textual signals or output distributions, which tend to be weak and noisy at the earliest decoding stage\. In contrast, ForeSight extracts sparse yet predictive safety signals from first\-token hidden representations, enabling harmfulness forecasting before explicit harmful content appears\.

### 2\.2Leveraging LLM Internals for Safety Detection

LLM internal representations have been shown to encode rich, specialized features that support downstream classification and reveal interpretable model behaviors\([Gurnee et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib22);[Lai et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib23)\)\. Beyond static classification, early internal states can also contain signals that are predictive of future outputs\([Dong et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib16)\), suggesting their potential for risk estimation\. In the safety domain, prior work finds that hidden activations capture fine\-grained safety\-related concepts and can distinguish aligned from jailbroken behaviors without relying solely on generated text\([Jiang et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib25);[Zhou et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib19);[Zhao et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib24);[Du et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib28)\)\. Representation\-level studies further suggest that safety\-related activations can be behavior\-relevant\([Zou et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib26);[Zou et al\., 2023a](https://arxiv.org/html/2609.13737#bib.bib27)\)\. Recent methods leverage internal activations for safety probing and intervention, revealing harmfulness\-related directions and latent structures\([Zhao et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib24);[Yung et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib29)\)\. For example, Legilimens probes conceptual features\([Wu et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib30)\), ShieldHead attaches decoding\-time safety heads\([Xuan et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib32)\), and HSF filters risky inputs from hidden\-state patterns\([Qian et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib31)\)\. However, these methods often operate before generation, rely on later decoding states, or use dense representations, without directly forecasting final\-response harmfulness from sparse earliest\-token signals\. In this work, ForeSight distills weak but predictive first\-token hidden\-state signals into sparse, structured, and layer\-aware representations for output\-level risk forecasting\.

## 3Method

### 3\.1Task Definition

Given an input promptxx, an autoregressive language model generates a responsey=\(y1,…,yn\)y=\(y\_\{1\},\\dots,y\_\{n\}\)via

yt∼pθ\(⋅∣x,y<t\),t=1,…,n,y\_\{t\}\\sim p\_\{\\theta\}\(\\cdot\\mid x,\\,y\_\{<t\}\),\\quad t=1,\\dots,n,\(1\)where each token is produced sequentially, conditioned on the prompt and the previously generated prefix\.

Our goal is to predict*output harmfulness*at an extremely early stage of decoding\. Unlike conventional safety classification that judges whether the prompt itself is unsafe, our target is whether the model’s*final generated response*is harmful\.

For each prompt, we first let the target model generate a complete responseyy\. An external evaluator then assigns an output\-level harmfulness labelz∈\{0,1\}z\\in\\\{0,1\\\}, wherez=1z=1indicates a harmful response andz=0z=0otherwise\.

Let𝒮=\{l1,…,lm\}\\mathcal\{S\}=\\\{l\_\{1\},\\dots,l\_\{m\}\\\}denote a set of transformer layers\. For each decoding steptt, we denote the corresponding multi\-layer hidden states as

ht=\{ht\(l\)∣l∈𝒮\},h\_\{t\}=\\\{h\_\{t\}^\{\(l\)\}\\mid l\\in\\mathcal\{S\}\\\},\(2\)whereht\(l\)∈ℝdh\_\{t\}^\{\(l\)\}\\in\\mathbb\{R\}^\{d\}is the hidden state at layerllafter generating tokenyty\_\{t\}\.

The task is to learn a predictor

f:h1↦z,f:h\_\{1\}\\mapsto z,\(3\)which estimates the harmfulness of the final response using only the internal model state available after the first decoding step\. At inference time, the predictor does not access any future tokensy\>1y\_\{\>1\}or the completed response text\.

### 3\.2Early Safety Signal Distillation

First\-token hidden states do not directly provide clean safety representations\. Although early decoding states may contain predictive signals about future responses\([Qi et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib17)\), internal activations often encode multiple overlapping behavioral and semantic factors\([Zhou et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib19)\)\. This makes it difficult to directly extract safety\-relevant information from the full hidden state, motivating a distillation process that localizes, compresses, and aggregates early safety signals\. ForeSight addresses this challenge through early safety signal distillation\. Given the hidden state immediately after the first generated token, we aim to extract a compact representation that preserves safety\-relevant variation while suppressing redundant or task\-irrelevant dimensions\.

For each selected layerl∈𝒮l\\in\\mathcal\{S\}, we denote this early state ash1\(l\)∈ℝdh\_\{1\}^\{\(l\)\}\\in\\mathbb\{R\}^\{d\}\. We treat\{h1\(l\)\}l∈𝒮\\\{h\_\{1\}^\{\(l\)\}\\\}\_\{l\\in\\mathcal\{S\}\}as a multi\-layer representation of the same early decoding state, where different layers may capture complementary but overlapping information\. In the following sections, we usehhto refer to this first\-token hidden state for readability\.

##### Saliency\-Induced Sparse Selection

Since the latent safety\-relevant component is not directly identifiable, we estimate its support using a layer\-wise linear classifier trained on first\-token hidden states\.

For each layerll, we train a binary linear classifier on the first\-token hidden state:

f⁡\(h1\)=softmax⁡\(W​h1\+b\),f\(h\_\{1\}\)=\\mathrm\{softmax\}\\left\(Wh\_\{1\}\+b\\right\),\(4\)
whereW∈ℝ2×dW\\in\\mathbb\{R\}^\{2\\times d\}andb∈ℝ2b\\in\\mathbb\{R\}^\{2\}are trainable parameters\. The classifier is optimized by minimizing:

ℒcls=CE⁡\(f⁡\(h\),z\)\+1C​∥W∥1,\\mathcal\{L\}\_\{\\mathrm\{cls\}\}=\\mathrm\{CE\}\\big\(f\(h\),z\\big\)\+\\frac\{1\}\{C\}\\lVert W\\rVert\_\{1\},\(5\)
where theℓ1\\ell\_\{1\}penalty encourages the classifier to rely on a sparse subset of hidden\-state dimensions\. By strictly penalizingWW, the model is forced to reduce the weights of safety\-irrelevant dimensions\. The inverse regularization strengthCCis selected for each layer on the validation set; smallerCCvalues correspond to stronger regularization and thus impose stronger sparsity onWW\.

After training, we define the saliency of dimensionj=1,…,dj=1,\\dots,das

aj=\|W1,j−W0,j\|,W1,W0∈ℝd,a\_\{j\}=\\left\|W\_\{1,j\}\-W\_\{0,j\}\\right\|,\\;W\_\{1\},W\_\{0\}\\in\\mathbb\{R\}^\{d\},\(6\)whereW1,jW\_\{1,j\}represents the harmful weight andW0,jW\_\{0,j\}represents the non\-harmful weight;aja\_\{j\}measures the contribution to the logit margin between harmful and non\-harmful classes\.

We then construct a binary maskM∈\{0,1\}dM\\in\\\{0,1\\\}^\{d\}by sorting dimensions in descending order ofaja\_\{j\}and retaining the smallest subset whose cumulative saliency reaches a fractionτ\\tauof the total saliency mass satisfying:

∑j=1dMj​aj≥τ​∑j=1daj\.\\sum\_\{j=1\}^\{d\}M\_\{j\}a\_\{j\}\\geq\\tau\\sum\_\{j=1\}^\{d\}a\_\{j\}\.\(7\)The saliency scoresaaand the induced maskMMare layer\-specific but input\-independent, and remain fixed for all samples after classifier training\. The sparsified representation is:

h~=M⊙h\.\\tilde\{h\}=M\\odot h\.\(8\)In implementation, we retain only the nonzero entries selected byMMand represent the sparse layer feature as the corresponding subvector ofhh\.

##### Structured Redundancy Reduction

Although sparsification removes low\-saliency dimensions, residual redundancy may persist due to correlated neuron activations\. To further reduce redundancy, we cluster retained neuron dimensions according to their cross\-sample activation profiles on the training set\.

For every layerll, let𝒥=\{j∣Mj=1\}\\mathcal\{J\}=\\\{j\\mid M\_\{j\}=1\\\}denote the set of retained dimensions at layerll\. For each retained dimensionj∈𝒥j\\in\\mathcal\{J\}, we define its activation profile over the training set as

pj=\[h~j​\(xi\)\]xi∈𝒟train\.p\_\{j\}=\\left\[\\tilde\{h\}\_\{j\}\(x\_\{i\}\)\\right\]\_\{x\_\{i\}\\in\\mathcal\{D\}\_\{\\mathrm\{train\}\}\}\.\(9\)We then cluster these profiles usingkk\-means\([McQueen, 1967](https://arxiv.org/html/2609.13737#bib.bib47)\)intoKKclusters:

𝒞=\{C1,…,CK\},\\mathcal\{C\}=\\left\\\{C\_\{1\},\\ldots,C\_\{K\}\\right\\\},\(10\)where each cluster groups neurons with similar activation behavior across training inputs\.

Given the learned clusters from layerll, for any samplexx, we aggregate the retained activations within each cluster by average pooling:

ck\(x\)=1\|Ck\|∑j∈Ckh~j\(x\),k=1,…,K\.c\_\{k\}\(x\)=\\frac\{1\}\{\|C\_\{k\}\|\}\\sum\_\{j\\in C\_\{k\}\}\\tilde\{h\}\_\{j\}\(x\),\\;k=1,\\ldots,K\.\(11\)This yields the structured layer representation

h¯​\(x\)=\[c1​\(x\),…,cK​\(x\)\]∈ℝK,\\bar\{h\}\(x\)=\\left\[c\_\{1\}\(x\),\\dots,c\_\{K\}\(x\)\\right\]\\in\\mathbb\{R\}^\{K\},\(12\)which induces a low\-dimensional subspace capturing shared variation while suppressing redundant activations\.

Table 1:Main results on five safety benchmarks with two target models\. All methods are evaluated for output\-level harmfulness prediction, and we report F1 and accuracy \(ACC\)\.
##### Cross\-Layer Adaptive Aggregation

Different transformer layers capture heterogeneous aspects of early decoding dynamics\. We assign each layer a fixed weightw\(l\)w^\{\(l\)\}computed from the validation performance of its sparse linear classifier\. Specifically,w\(l\)w^\{\(l\)\}is computed by min\-max normalizing the validation F1 scores across layers\. Therefore, layers with stronger performance receive larger weights, yielding the unified representation:

g⁡\(x\)=⨁l∈𝒮w\(l\)⋅h¯\(l\)​\(x\),g\(x\)=\\bigoplus\_\{l\\in\\mathcal\{S\}\}w^\{\(l\)\}\\cdot\\bar\{h\}^\{\(l\)\}\(x\),\(13\)where⨁\\bigoplusdenotes concatenation across layers\.

Overall,g⁡\(x\)g\(x\)can be interpreted as a distilled representation of early predictive safety signals, where redundant activations are suppressed and safety\-relevant structure is amplified\.

### 3\.3Risk Forecasting

Given the distilled representationg⁡\(x\)g\(x\), we predict the harmfulness probability with a lightweight MLP:

z^​\(x\)=σ⁡\(ϕθ​\(g⁡\(x\)\)\),\\hat\{z\}\(x\)=\\sigma\\big\(\\phi\_\{\\theta\}\(g\(x\)\)\\big\),\(14\)whereϕθ\\phi\_\{\\theta\}denotes the MLP classifier andσ⁡\(⋅\)\\sigma\(\\cdot\)is the sigmoid function\. The classifier is trained with binary cross\-entropy:

ℒBCE=−1N∑i=1N\[zilogz^i\+\(1−zi\)log\(1−z^i\)\]\.\\mathcal\{L\}\_\{\\mathrm\{BCE\}\}=\-\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\left\[z\_\{i\}\\log\\hat\{z\}\_\{i\}\+\(1\-z\_\{i\}\)\\log\(1\-\\hat\{z\}\_\{i\}\)\\right\]\.\(15\)
During inference, responses withz^​\(x\)≥0\.5\\hat\{z\}\(x\)\\geq 0\.5are classified as harmful, with a fixed threshold independent of validation tuning and saliency selection\.

## 4Experiment

Figure 3:Accuracy–efficiency comparison on Llama\-3\.1\-8B and Qwen3\-8B\. Accuracy is averaged over five datasets, and efficiency is measured by inference time and FLOPs\.Table 2:Cross\-dataset generalization results\. Each detector is trained on one source dataset and evaluated on the remaining four target datasets\. We report the average F1 of target datasets; L and Q denote Llama\-3\.1\-8B\-Instruct and Qwen3\-8B, respectively\.### 4\.1Datasets and Evaluation Metrics

We evaluate ForeSight on five safety benchmarks: HarmEval\([Banerjee et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib37)\), S\-Eval\([Yuan et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib39)\), CatQA\([Bhardwaj et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib38)\), ToxicChat\([Lin et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib10)\), and WildJailbreak\([Jiang et al\., 2024a](https://arxiv.org/html/2609.13737#bib.bib40)\)\. For each benchmark, we use Llama\-3\.1\-8B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib41)\)and Qwen3\-8B\([Yang et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib42)\)as target models to generate complete responses, covering both higher\-risk and more refusal\-prone generation regimes\. Our task focuses on output\-level harmfulness and we label each complete response using three judge models: Llama\-3\.3\-70B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib41)\), Mistral\-Small\-24B\-Instruct\([Liu et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib44)\), and Qwen3\.5\-27B\([Team, 2026](https://arxiv.org/html/2609.13737#bib.bib43)\)\. Only samples with unanimous judge agreement are retained\. We report F1 score and accuracy \(ACC\) as evaluation metrics\. Additional dataset statistics and label construction details are provided in Appendix[A](https://arxiv.org/html/2609.13737#A1)\.

### 4\.2Experimental Setup

We evaluate ForeSight, which forecasts output\-level harmfulness using only first\-token hidden states, against two groups of baselines\. First, we include deployment\-oriented text\-based safety classifiers, including judge models: Llama\-3\.3\-70B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib41)\), Mistral\-Small\-24B\-Instruct\([Liu et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib44)\), and Qwen3\.5\-27B\([Team, 2026](https://arxiv.org/html/2609.13737#bib.bib43)\), as well as open\-source guardrail models: Llama\-Guard\-3\-8B\([Inan et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib45)\)and Qwen3Guard\-8B\([Zhao et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib46)\)\. These models are not retrained but are given access to the original prompt and the first 10 generated tokens at inference time, providing a stronger observation budget than ForeSight\. We additionally include a supervised RoBERTa classifier\([Liu et al\., 2019](https://arxiv.org/html/2609.13737#bib.bib7)\)trained on the first 10 generated tokens using the same output\-level labels and data splits as ForeSight, providing a controlled comparison based on early textual observations\. Second, we compare with trainable early\-forecasting method baselines, including the logits\-based method MULI\([Hu et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib14)\), the hidden\-state\-based method Latent Guard\([Zhao et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib24)\), and the activation\-based method TPC\([Oldfield et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib33)\), which are trained with the same output\-level labels and data splits as ForeSight\. Together, these baselines control for early output\-distribution cues and dense first\-token hidden\-state signals\. Thus, text\-based baselines provide deployment\-oriented references, while method\-based baselines provide controlled comparisons\. Further implementation details and baseline configurations are provided in Appendix[B](https://arxiv.org/html/2609.13737#A2)and Appendix[C](https://arxiv.org/html/2609.13737#A3)\.

Figure 4:Average F1 score over five datasets under different observation lengths on Llama\-3\.1\-8B and Qwen3\-8B\. Prefix length 0 denotes the input\-only setting, while lengths 1–25 denote output\-prefix settings with progressively longer generated text\. Text\-based baselines rely on explicit tokens, whereas ForeSight uses only first\-token hidden states\.
### 4\.3Main Results

Table[1](https://arxiv.org/html/2609.13737#S3.T1)reports results on five safety benchmarks with two target models\. Overall, ForeSight achieves the strongest performance, obtaining the best or tied\-best ACC in 9 out of 10 settings and the best F1 in 9 out of 10 settings\. Compared with judge\-based models, guardrail baselines, and recent internals\-based detectors such as MULI and Latent Guard, ForeSight shows more balanced performance across datasets\. Since MULI uses early\-token logits and Latent Guard uses dense first\-token hidden states, these gains suggest that ForeSight benefits from sparse structured distillation rather than first\-token access alone\.

On Llama\-3\.1\-8B, ForeSight improves over the strongest baseline by 8\.47, 9\.70, 5\.19, and 1\.40 F1 points on HarmEval, S\-Eval, ToxicChat, and WildJailbreak, respectively, with CatQA as the only exception\. On Qwen3\-8B, ForeSight achieves the best F1 on all five datasets, with gains of 26\.51, 9\.13, 21\.00, 18\.64, and 3\.40 points\. Although Latent Guard and some guardrail baselines occasionally obtain high ACC, their F1 scores are often lower, suggesting less stable or majority\-biased predictions\. These results indicate that first\-token hidden states already contain predictive safety signals for final\-response harmfulness\.

Figure[3](https://arxiv.org/html/2609.13737#S4.F3)further shows that ForeSight achieves strong performance with lower inference time and FLOPs, supporting its efficiency for early risk estimation\.

#### 4\.3\.1Generalization

We evaluate cross\-dataset generalization by training each detector on one source dataset and testing it on the remaining target datasets\. As shown in Table[2](https://arxiv.org/html/2609.13737#S4.T2), ForeSight consistently outperforms other method\-based baselines across all source datasets and both model backbones\. This indicates that the sparse first\-token safety signals captured by ForeSight are more transferable across domains\.

#### 4\.3\.2Observation Length

We compare text\-based baselines under different observation lengths while keeping ForeSight fixed at the first\-token hidden\-state setting\. As shown in Figure[4](https://arxiv.org/html/2609.13737#S4.F4), text\-based baselines perform poorly in the early output\-prefix regime and improve only as more generated tokens become available\. The gap between the input\-only point and short output prefixes suggests that prompt\-level risk is not a reliable proxy for final response harmfulness\. In contrast, ForeSight achieves strong performance using only first\-token hidden states, indicating that early internal states encode predictive safety information that is not directly exposed in the prompt or the earliest surface tokens\.

### 4\.4Ablation Study

#### 4\.4\.1Persistence of First\-token Safety Signals

We examine whether safety signals learned from the first generated token remain informative at later decoding positions by applying a classifier trained onh1h\_\{1\}to subsequent hidden stateshth\_\{t\}without retraining\. As shown in Figure[5](https://arxiv.org/html/2609.13737#S4.F5), the classifier remains predictive for both Llama\-3\.1\-8B and Qwen3\-8B\. The signal is more stable on Llama\-3\.1\-8B, whereas Qwen3\-8B shows a decline at later positions, suggesting that first\-token safety signals persist but may weaken as generation proceeds\.

Figure 5:Transfer of theh1h\_\{1\}\-trained classifier across decoding positions\. For each target model, the classifier is trained on HarmEval using the first\-token hidden stateh1h\_\{1\}and directly evaluated on hidden stateshth\_\{t\}at different generated token positions\. The dashed vertical line markst=1t=1, the training position\.This suggests that the safety\-relevant information captured at the first generated token persists across later decoding steps\. The lower accuracy ath0h\_\{0\}further indicates that the first generated token provides additional decoding\-stage information beyond the prompt\-only prefill state\. Overall, these results supporth1h\_\{1\}as an effective and transferable point for early harmfulness forecasting, without requiring longer observation windows or token\-specific retraining\.

Table 3:Cross\-domain performance of different neuron selection strategies on Llama\-3\.1\-8B\. All models are trained on HarmEval and evaluated on other safety benchmarks\.
#### 4\.4\.2Analysis of Sparse Safety Signal Distillation

Table[4](https://arxiv.org/html/2609.13737#S4.T4)compares different strategies for constructing first\-token representations\. Full hidden states use the largest representation but achieve limited performance, suggesting substantial redundancy in raw activations\. Sparse selection improves F1 by retaining more informative dimensions\. By contrast, sparse selection with clustering achieves the best accuracy and F1 with only 648 dimensions, indicating that structured clustering reduces redundancy while preserving safety\-relevant variation\.

Table 4:Neuron selection ablation on HarmEval with Llama\-3\.1\-8B\. Dim denotes the representation dimensionality, and Time denotes the average inference time per example\.We further evaluate whether the sparse representation transfers across datasets\. As shown in Table[3](https://arxiv.org/html/2609.13737#S4.T3), sparse selection with clustering achieves the strongest performance in most transfer settings, suggesting that structured first\-token representations improve both efficiency and cross\-domain robustness\. Additional cluster transferability results are provided in Appendix[D\.1](https://arxiv.org/html/2609.13737#A4.SS1)\.

Finally, we examine the robustness of the sparsity design itself\. Very small retention thresholds lead to lower performance, suggesting that overly aggressive pruning removes useful safety\-relevant signals\. Meanwhile, the optimalℓ1\\ell\_\{1\}regularization strength varies across layers, indicating that different layers require different sparsity levels\. These results support moderate thresholding and layer\-specific sparse selection rather than a fixed global sparsity setting\. Detailed curves are provided in Appendix[D\.4](https://arxiv.org/html/2609.13737#A4.SS4)and[D\.3](https://arxiv.org/html/2609.13737#A4.SS3)\.

#### 4\.4\.3Intervention Sensitivity of the Learned Risk Direction

We further test whether the learned first\-token risk direction is intervention\-sensitive\. For each selected layerll, we define the direction from the sparse classifier margin:

u\(l\)=M\(l\)⊙\(W1\(l\)−W0\(l\)\)‖M\(l\)⊙\(W1\(l\)−W0\(l\)\)‖2,u^\{\(l\)\}=\\frac\{M^\{\(l\)\}\\odot\(W^\{\(l\)\}\_\{1\}\-W^\{\(l\)\}\_\{0\}\)\}\{\\\|M^\{\(l\)\}\\odot\(W^\{\(l\)\}\_\{1\}\-W^\{\(l\)\}\_\{0\}\)\\\|\_\{2\}\},\(16\)whereM\(l\)M^\{\(l\)\}is the saliency mask\. The direction is oriented so that\+u\(l\)\+u^\{\(l\)\}increases the validation\-set harmfulness score\. Given a fixed first tokeny1y\_\{1\}, we edit the corresponding hidden state as

h~1\(l\)=h1\(l\)−α​u\(l\),\\tilde\{h\}\_\{1\}^\{\(l\)\}=h\_\{1\}^\{\(l\)\}\-\\alpha u^\{\(l\)\},\(17\)then continue decoding and evaluate the final response with the same judge protocol\. We report the Harmful score, defined as the softmax probability assigned by the binary classifier to the harmful class, i\.e\.,P⁡\(harmful=1∣hs=1\)P\(\\mathrm\{harmful\}=1\\mid h\_\{s=1\}\)\. A lower score indicates that the edited first\-token representation is predicted to be less harmful\.

Table 5:Representation editing analysis\. We fixy1y\_\{1\}and compare the learned risk direction with reverse and random controls\.As shown in Table[5](https://arxiv.org/html/2609.13737#S4.T5), editing away from the learned risk direction reduces the Harmful Score from 0\.4264 to 0\.3382 and the judged harmful rate from 79\.31% to 70\.00%, while reverse editing increases the score to 0\.6498 and slightly raises the harmful rate to 81\.48%\. The random control produces a weaker reduction, suggesting that the learned direction is associated with harmfulness\-related changes in the generation behavior rather than being a generic perturbation\.

## 5Conclusion

In this work, we study forecasting final\-response harmfulness from hidden states immediately after the first generated token\. To address the weak and entangled nature of early signals, we propose ForeSight, an early safety signal distillation framework that extracts, compresses, and aggregates sparse safety\-relevant representations across layers\. Experiments on multiple backbones and benchmarks show that first\-token hidden states contain useful predictive signals of future harmfulness\. Compared with dense hidden\-state approaches, ForeSight achieves more effective and efficient early\-risk forecasting with minimal decoding information\. Overall, our results suggest that early decoding states contain distilled safety signals before explicit harmful content becomes observable\.

## 6Limitations

This work has several limitations:

- •Dependence on internal hidden states\.ForeSight requires access to model hidden states, limiting its applicability to closed\-source or API\-only LLMs\.
- •Reliance on judge\-model labels\.Although we retain only samples with unanimous judge agreement, judge labels may still contain biases or exclude ambiguous safety cases\. We further analyze borderline examples with inconsistent labels in Appendix[E](https://arxiv.org/html/2609.13737#A5)\.
- •Preliminary intervention analysis\.Appendix[G](https://arxiv.org/html/2609.13737#A7)provides a small\-scale forecast\-guided resampling study, but it remains a sanity check rather than a complete defense\. Future work may develop more robust decoding\-control mechanisms\.

## 7Ethics Statement

ForeSight is designed as a defensive tool for early detection of harmful generations\. Although internal risk signals may have dual\-use implications, we release only detection\-oriented code and evaluation scripts, without capabilities for generating or amplifying unsafe trajectories\.

All experiments are conducted using publicly available datasets and pretrained models under their respective licenses and usage agreements\. These resources are used solely for LLM safety research and evaluation, and we cite the original creators of all datasets, models, and baselines\.

The datasets may contain safety\-sensitive or offensive content\. We do not collect new personally identifiable information, and to the best of our knowledge, the datasets used do not contain such information\. We provide appropriate disclaimers for potentially disturbing content, and the opinions expressed in the datasets do not represent those of the authors\.

## 8Reproducibility Statement

We provide detailed descriptions of the experimental procedure, including model inference, hyperparameter settings, and baseline configurations, in Section[4\.2](https://arxiv.org/html/2609.13737#S4.SS2)and Appendices[C](https://arxiv.org/html/2609.13737#A3)and[B](https://arxiv.org/html/2609.13737#A2)\. The dataset construction process and prompt formulations are described in Section[4\.1](https://arxiv.org/html/2609.13737#S4.SS1)and further detailed in Appendix[A](https://arxiv.org/html/2609.13737#A1)\. All datasets used in this work are publicly available\. Our code is publicly available to facilitate reproducibility at:[https://github\.com/Scabbards1500/Foresight](https://github.com/Scabbards1500/Foresight)\.

## 9The Use of LLM

ChatGPT \(OpenAI\) was used solely for English copyediting and minor LaTeX formatting\. It was not involved in idea generation, experiments, analysis, or result production\. All scientific content was written and verified by the authors\.

## References

- Alain and Bengio \(2016\)G\. Alain and Y\. BengioUnderstanding intermediate layers using linear classifier probes\.arXiv preprint arXiv:1610\.01644\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p2.1)\.
- Banerjeeet al\.\(2025\)S\. Banerjee, S\. Layek, S\. Tripathy, S\. Kumar, A\. Mukherjee, and R\. HazraSafeinfer: context adaptive decoding time safety alignment for large language models\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.39,pp\. 27188–27196\.Cited by:[§A\.1](https://arxiv.org/html/2609.13737#A1.SS1.p2.1),[§4\.1](https://arxiv.org/html/2609.13737#S4.SS1.p1.1)\.
- Bhardwajet al\.\(2024\)R\. Bhardwaj, D\. A\. Do, and S\. PoriaLanguage models are homer simpson\! safety re\-alignment of fine\-tuned language models through task arithmetic\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 14138–14149\.Cited by:[§A\.1](https://arxiv.org/html/2609.13737#A1.SS1.p4.1),[§4\.1](https://arxiv.org/html/2609.13737#S4.SS1.p1.1)\.
- Caselliet al\.\(2021\)T\. Caselli, V\. Basile, J\. Mitrović, and M\. GranitzerHateBERT: retraining bert for abusive language detection in english\.InProceedings of the 5th Workshop on Online Abuse and Harms \(WOAH 2021\),pp\. 17–25\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Chenet al\.\(2025\)G\. Chen, Y\. Xia, X\. Jia, Z\. Li, P\. Torr, and J\. GuLLM jailbreak detection for \(almost\) free\!\.arXiv preprint arXiv:2509\.14558\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Devlinet al\.\(2019\)J\. Devlin, M\. Chang, K\. Lee, and K\. ToutanovaBert: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 \(long and short papers\),pp\. 4171–4186\.Cited by:[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Donget al\.\(2025\)Z\. Dong, Z\. Zhou, Z\. Liu, C\. Yang, and C\. LuEmergent response planning in llms\.arXiv preprint arXiv:2502\.06258\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Duet al\.\(2025\)T\. Du, Z\. Wei, Q\. Chen, C\. Zhang, and Y\. WangAdvancing llm safe alignment with safety representation ranking\.arXiv preprint arXiv:2505\.15710\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Ghoshet al\.\(2025\)S\. Ghosh, P\. Varshney, M\. N\. Sreedhar, A\. Padmakumar, T\. Rebedea, J\. R\. Varghese, and C\. ParisienAegis2\.0: a diverse ai safety dataset and risks taxonomy for alignment of llm guardrails\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 5992–6026\.Cited by:[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§A\.2](https://arxiv.org/html/2609.13737#A1.SS2.p1.1),[§B\.1](https://arxiv.org/html/2609.13737#A2.SS1.p2.1),[§4\.1](https://arxiv.org/html/2609.13737#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.13737#S4.SS2.p1.1)\.
- Gurneeet al\.\(2023\)W\. Gurnee, N\. Nanda, M\. Pauly, K\. Harvey, D\. Troitskii, and D\. BertsimasFinding neurons in a haystack: case studies with sparse probing\.arXiv preprint arXiv:2305\.01610\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Hanet al\.\(2024\)S\. Han, K\. Rao, A\. Ettinger, L\. Jiang, B\. Y\. Lin, N\. Lambert, Y\. Choi, and N\. DziriWildguard: open one\-stop moderation tools for safety risks, jailbreaks, and refusals of llms\.Advances in neural information processing systems37,pp\. 8093–8131\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Hernandezet al\.\(2024\)E\. Hernandez, A\. Sen Sharma, T\. Haklay, K\. Meng, M\. Wattenberg, J\. Andreas, Y\. Belinkov, and D\. BauLinearity of relation decoding in transformer language models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 10504–10526\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p2.1)\.
- Huet al\.\(2024\)Z\. Hu, J\. Piet, G\. Zhao, J\. Jiao, and D\. WagnerToxicity detection for free\.Advances in Neural Information Processing Systems37,pp\. 17518–17540\.Cited by:[§B\.2](https://arxiv.org/html/2609.13737#A2.SS2.p2.1),[§1](https://arxiv.org/html/2609.13737#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.13737#S4.SS2.p1.1)\.
- Inanet al\.\(2023\)H\. Inan, K\. Upasani, J\. Chi, R\. Rungta, K\. Iyer, Y\. Mao, M\. Tontchev, Q\. Hu, B\. Fuller, D\. Testuggine,et al\.Llama guard: llm\-based input\-output safeguard for human\-ai conversations\.arXiv preprint arXiv:2312\.06674\.Cited by:[§B\.1](https://arxiv.org/html/2609.13737#A2.SS1.p5.1),[§1](https://arxiv.org/html/2609.13737#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.13737#S4.SS2.p1.1)\.
- Jianget al\.\(2023\)A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. SayedMistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[§D\.6](https://arxiv.org/html/2609.13737#A4.SS6.p1.1)\.
- Jianget al\.\(2024a\)L\. Jiang, K\. Rao, S\. Han, A\. Ettinger, F\. Brahman, S\. Kumar, N\. Mireshghallah, X\. Lu, M\. Sap, Y\. Choi,et al\.Wildteaming at scale: from in\-the\-wild jailbreaks to \(adversarially\) safer language models\.Advances in Neural Information Processing Systems37,pp\. 47094–47165\.Cited by:[§A\.1](https://arxiv.org/html/2609.13737#A1.SS1.p6.1),[§4\.1](https://arxiv.org/html/2609.13737#S4.SS1.p1.1)\.
- Jianget al\.\(2024b\)Y\. Jiang, G\. Rajendran, P\. Ravikumar, B\. Aragam, and V\. VeitchOn the origins of linear representations in large language models\.arXiv preprint arXiv:2403\.03867\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p2.1)\.
- Jianget al\.\(2025\)Y\. Jiang, X\. Gao, T\. Peng, Y\. Tan, X\. Zhu, B\. Zheng, and X\. YueHiddendetect: detecting jailbreak attacks against large vision\-language models via monitoring hidden states\.arXiv preprint arXiv:2502\.147443\(5\)\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Kavumbaet al\.\(2026\)P\. Kavumba, K\. Wataoka, H\. H\. Nguyen, J\. Li, and M\. OhagiPredict, don’t react: value\-based safety forecasting for llm streaming\.arXiv preprint arXiv:2604\.03962\.Cited by:[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Laiet al\.\(2026\)P\. Lai, J\. Zheng, S\. Cheng, Y\. Chen, P\. Li, Y\. Liu, and G\. ChenBeyond the surface: enhancing llm\-as\-a\-judge alignment with human via internal representations\.Advances in Neural Information Processing Systems38,pp\. 93353–93383\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Liet al\.\(2026\)Y\. Li, Q\. Sheng, Y\. Yang, X\. Zhang, and J\. CaoFrom judgment to interference: early stopping llm harmful outputs via streaming content monitoring\.Advances in Neural Information Processing Systems38,pp\. 54305–54333\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Linet al\.\(2023\)Z\. Lin, Z\. Wang, Y\. Tong, Y\. Wang, Y\. Guo, Y\. Wang, and J\. ShangToxicchat: unveiling hidden challenges of toxicity detection in real\-world user\-ai conversation\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 4694–4702\.Cited by:[§A\.1](https://arxiv.org/html/2609.13737#A1.SS1.p5.1),[§4\.1](https://arxiv.org/html/2609.13737#S4.SS1.p1.1)\.
- Liuet al\.\(2026\)A\. H\. Liu, K\. Khandelwal, S\. Subramanian, V\. Jouault, A\. Rastogi, A\. Sadé, A\. Jeffares, A\. Jiang, A\. Cahill, A\. Gavaudan,et al\.Ministral 3\.arXiv preprint arXiv:2601\.08584\.Cited by:[§A\.2](https://arxiv.org/html/2609.13737#A1.SS2.p1.1),[§B\.1](https://arxiv.org/html/2609.13737#A2.SS1.p3.1),[§4\.1](https://arxiv.org/html/2609.13737#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.13737#S4.SS2.p1.1)\.
- Liuet al\.\(2019\)Y\. Liu, M\. Ott, N\. Goyal, J\. Du, M\. Joshi, D\. Chen, O\. Levy, M\. Lewis, L\. Zettlemoyer, and V\. StoyanovRoberta: a robustly optimized bert pretraining approach\.arXiv preprint arXiv:1907\.11692\.Cited by:[§B\.1](https://arxiv.org/html/2609.13737#A2.SS1.p7.1),[§4\.2](https://arxiv.org/html/2609.13737#S4.SS2.p1.1)\.
- Markovet al\.\(2023\)T\. Markov, C\. Zhang, S\. Agarwal, F\. E\. Nekoul, T\. Lee, S\. Adler, A\. Jiang, and L\. WengA holistic approach to undesired content detection in the real world\.InProceedings of the AAAI conference on artificial intelligence,Vol\.37,pp\. 15009–15018\.Cited by:[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- McQueen \(1967\)J\. B\. McQueenSome methods of classification and analysis of multivariate observations\.InProc\. of 5th Berkeley Symposium on Math\. Stat\. and Prob\.,pp\. 281–297\.Cited by:[§3\.2](https://arxiv.org/html/2609.13737#S3.SS2.SSS0.Px2.p2.2)\.
- Mozafariet al\.\(2020\)M\. Mozafari, R\. Farahbakhsh, and N\. CrespiHate speech detection and racial bias mitigation in social media based on bert model\.PloS one15\(8\),pp\. e0237861\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Oldfieldet al\.\(2026\)J\. Oldfield, P\. Torr, I\. Patras, A\. Bibi, and F\. BarezBeyond linear probes: dynamic safety monitoring for language models\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 55192–55229\.Cited by:[§B\.2](https://arxiv.org/html/2609.13737#A2.SS2.p4.1),[§4\.2](https://arxiv.org/html/2609.13737#S4.SS2.p1.1)\.
- Parket al\.\(2023\)K\. Park, Y\. J\. Choe, and V\. VeitchThe linear representation hypothesis and the geometry of large language models\.arXiv preprint arXiv:2311\.03658\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p2.1)\.
- Qiet al\.\(2025\)X\. Qi, A\. Panda, K\. Lyu, X\. Ma, S\. Roy, A\. Beirami, P\. Mittal, and P\. HendersonSafety alignment should be made more than just a few tokens deep\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 54911–54941\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p1.1),[§3\.2](https://arxiv.org/html/2609.13737#S3.SS2.p1.1)\.
- Qianet al\.\(2025\)C\. Qian, H\. Zhang, L\. Sha, and Z\. ZhengHsf: defending against jailbreak attacks with hidden state filtering\.InCompanion Proceedings of the ACM on Web Conference 2025,pp\. 2078–2087\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Shenet al\.\(2024\)X\. Shen, Z\. Chen, M\. Backes, Y\. Shen, and Y\. Zhang"Do anything now": characterizing and evaluating in\-the\-wild jailbreak prompts on large language models\.InProceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security,pp\. 1671–1685\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p1.1)\.
- Team \(2026\)Q\. TeamQwen3\.5: towards native multimodal agents\.URL: https://qwen\.ai/blog\.Cited by:[§A\.2](https://arxiv.org/html/2609.13737#A1.SS2.p1.1),[§B\.1](https://arxiv.org/html/2609.13737#A2.SS1.p4.1),[§4\.1](https://arxiv.org/html/2609.13737#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.13737#S4.SS2.p1.1)\.
- Weiet al\.\(2023\)A\. Wei, N\. Haghtalab, and J\. SteinhardtJailbroken: how does llm safety training fail?\.Advances in neural information processing systems36,pp\. 80079–80110\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p1.1)\.
- Wuet al\.\(2024\)J\. Wu, J\. Deng, S\. Pang, Y\. Chen, J\. Xu, X\. Li, and W\. XuLegilimens: practical and unified content moderation for large language model services\.InProceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security,pp\. 1151–1165\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Xuanet al\.\(2025\)Z\. Xuan, X\. Mao, D\. Chen, X\. Zhang, Y\. Dong, and J\. ZhouShieldHead: decoding\-time safeguard for large language models\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 18129–18143\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Yanget al\.\(2025\)A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§4\.1](https://arxiv.org/html/2609.13737#S4.SS1.p1.1)\.
- Yuanet al\.\(2025\)X\. Yuan, J\. Li, D\. Wang, Y\. Chen, X\. Mao, L\. Huang, J\. Chen, H\. Xue, X\. Liu, W\. Wang,et al\.S\-eval: towards automated and comprehensive safety evaluation for large language models\.Proceedings of the ACM on Software Engineering2\(ISSTA\),pp\. 2136–2157\.Cited by:[§A\.1](https://arxiv.org/html/2609.13737#A1.SS1.p3.1),[§4\.1](https://arxiv.org/html/2609.13737#S4.SS1.p1.1)\.
- Yunget al\.\(2025\)C\. Yung, H\. Huang, S\. Monazam Erfani, and C\. LeckieCURVALID: geometrically\-guided adversarial prompt detection\.arXiv preprint arXiv:2503\.03502\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Zenget al\.\(2024\)W\. Zeng, Y\. Liu, R\. Mullins, L\. Peran, J\. Fernandez, H\. Harkous, K\. Narasimhan, D\. Proud, P\. Kumar, B\. Radharapu,et al\.Shieldgemma: generative ai content moderation based on gemma\.arXiv preprint arXiv:2407\.21772\.Cited by:[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Zhaoet al\.\(2025\)H\. Zhao, C\. Yuan, F\. Huang, X\. Hu, Y\. Zhang, A\. Yang, B\. Yu, D\. Liu, J\. Zhou, J\. Lin,et al\.Qwen3guard technical report\.arXiv preprint arXiv:2510\.14276\.Cited by:[§B\.1](https://arxiv.org/html/2609.13737#A2.SS1.p6.1),[§1](https://arxiv.org/html/2609.13737#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1),[§4\.2](https://arxiv.org/html/2609.13737#S4.SS2.p1.1)\.
- Zhaoet al\.\(2026\)J\. Zhao, J\. Huang, Z\. Wu, D\. Bau, and W\. ShiLlms encode harmfulness and refusal separately\.Advances in Neural Information Processing Systems38,pp\. 140283–140318\.Cited by:[§B\.2](https://arxiv.org/html/2609.13737#A2.SS2.p3.1),[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2609.13737#S4.SS2.p1.1)\.
- Zhaoet al\.\(2021\)Z\. Zhao, Z\. Zhang, and F\. HopfgartnerA comparative study of using pre\-trained language models for toxic comment classification\.InCompanion Proceedings of the Web Conference 2021,pp\. 500–507\.Cited by:[§2\.1](https://arxiv.org/html/2609.13737#S2.SS1.p1.1)\.
- Zhouet al\.\(2024\)Z\. Zhou, H\. Yu, X\. Zhang, R\. Xu, F\. Huang, and Y\. LiHow alignment and jailbreak work: explain llm safety through intermediate hidden states\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 2461–2488\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.13737#S3.SS2.p1.1)\.
- Zouet al\.\(2023a\)A\. Zou, L\. Phan, S\. Chen, J\. Campbell, P\. Guo, R\. Ren, A\. Pan, X\. Yin, M\. Mazeika, A\. Dombrowski,et al\.Representation engineering: a top\-down approach to ai transparency\.arXiv preprint arXiv:2310\.01405\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Zouet al\.\(2024\)A\. Zou, L\. Phan, J\. Wang, D\. Duenas, M\. Lin, M\. Andriushchenko, R\. Wang, Z\. Kolter, M\. Fredrikson, and D\. HendrycksImproving alignment and robustness with circuit breakers\.Advances in Neural Information Processing Systems37,pp\. 83345–83373\.Cited by:[§2\.2](https://arxiv.org/html/2609.13737#S2.SS2.p1.1)\.
- Zouet al\.\(2023b\)A\. Zou, Z\. Wang, N\. Carlini, M\. Nasr, J\. Z\. Kolter, and M\. FredriksonUniversal and transferable adversarial attacks on aligned language models\.arXiv preprint arXiv:2307\.15043\.Cited by:[§1](https://arxiv.org/html/2609.13737#S1.p1.1)\.

## Appendix ABenchmarks and Data Construction

We provide additional details on the benchmark sources and output\-level label construction\. All safety\-oriented datasets are used only as prompt sources; prediction labels are assigned based on the harmfulness of each target model’s complete response\. This ensures that ForeSight is evaluated on output\-level harmfulness rather than prompt\-level risk\.

### A\.1Data Sources

We use five safety\-oriented datasets as sources of harmful, toxic, or jailbreak\-related prompts\. Their prompt\-level annotations are used only for sample selection, while the final prediction labels are determined by the harmfulness of the target model’s complete responses\.

HarmEval\([Banerjee et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib37)\)is a policy\-oriented harmfulness evaluation benchmark introduced in SafeInfer\. It contains harmful prompts covering multiple prohibited\-use categories and is designed to evaluate whether language models produce unsafe responses to safety\-sensitive inputs\.

S\-Eval\([Yuan et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib39)\)is a large\-scale safety evaluation benchmark organized around a hierarchical risk taxonomy\. It includes diverse risk\-inducing prompts and adversarially transformed attacks, covering categories such as illegal activities, privacy risks, hate and harassment, and unsafe advice\.

CatQA\([Bhardwaj et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib38)\)is a red\-teaming question\-answering benchmark for evaluating language model safety\. It contains harmful questions across a broad range of topics, with an emphasis on testing model behavior under unsafe user requests\.

ToxicChat\([Lin et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib10)\)is a toxicity benchmark collected from real\-world user interactions with a chatbot system\. It provides prompts annotated for toxic and jailbreak\-related content, reflecting safety risks in realistic conversational scenarios\.

WildJailbreak\([Jiang et al\., 2024a](https://arxiv.org/html/2609.13737#bib.bib40)\)is an open\-source safety dataset constructed through the WildTeaming framework\. It contains harmful requests and adversarial jailbreak prompts derived from in\-the\-wild user\-chatbot interactions\.

### A\.2Output\-Level Label Construction

Figure 6:Output\-level label construction process\. A frozen target model first generates a complete response, which is then labeled by three external judge models; only unanimous labels are retained for the main experiments\.Figure 7:Prompt templates used for output\-level labeling and text\-based baseline evaluation\. The external judge prompt labels complete generated responses, while the baseline prompt predicts final\-response harmfulness from observed response prefixes\.Table 6:Distribution of harmfulness vote counts across datasets and target models\. Columns 0–3 denote the number of judge models that label the generated response as harmful\. Samples with vote counts of 0 or 3 are retained for the main experiments\.For each prompt, a frozen target model generates a complete response, while the hidden states required by ForeSight are recorded during decoding\. The harmfulness label is assigned only after full generation, ensuring that the prediction target is output\-level harmfulness rather than prompt\-level risk\. We annotate each response using three external judge models: Llama\-3\.3\-70B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib41)\), Mistral\-Small\-24B\-Instruct\([Liu et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib44)\), and Qwen3\.5\-27B\([Team, 2026](https://arxiv.org/html/2609.13737#bib.bib43)\)\.

Their outputs are converted into a harmfulness vote count from 0 to 3\. Unanimous cases, i\.e\., vote counts of 0 or 3, are retained for the main experiments and split into training, validation, and test sets with an 8:1:1 ratio\. Non\-unanimous cases, i\.e\., vote counts of 1 or 2, are excluded from the main splits and analyzed separately as borderline examples which are discussed in Appendix[E](https://arxiv.org/html/2609.13737#A5)\.

The totals may differ slightly across target models because a small number of judge outputs cannot be parsed into a valid label and are therefore excluded before aggregation\.

![Refer to caption](https://arxiv.org/html/2609.13737v1/cluster_transferability.png)Figure 8:Effect of the cluster numberkkin ForeSight on Llama\-3\.1\-8B across datasets\. Each cell reportsΔ\\DeltaF1 relative to the dataset\-specific best result, where 0 denotes the bestkkfor that dataset\.

## Appendix BDetailed Baseline Settings

Table 7:Comparison protocol for different baselines\. Re\-trained indicates whether the detector is trained on the same output\-level harmfulness labels and data splits as ForeSight\.We summarize the baseline protocols for predicting final\-response harmfulness under early\-observation settings\. As shown in Table[7](https://arxiv.org/html/2609.13737#A2.T7), text\-based baselines serve as deployment\-oriented references, while method\-based baselines provide controlled comparisons using the same output\-level labels and data splits as ForeSight\. We describe these two groups below\.

### B\.1Text\-based Baselines

For text\-based baselines, the observable input consists of the original prompt and the first 10 generated tokens\. All models are evaluated using the same prompt template, as shown in Figure[7](https://arxiv.org/html/2609.13737#A1.F7)\. These baselines include both general\-purpose instruction\-tuned LLMs and specialized safety guardrail models\. We additionally include a supervised RoBERTa classifier trained on the same output\-level labels and data splits as ForeSight\.

Llama\-3\.3\-70B\-Instruct\([Grattafiori et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib41)\)is a large instruction\-tuned model from the Llama family\. It is designed for general\-purpose language understanding and generation, and has been widely used as a strong judge\-style model in evaluation settings\.

Mistral\-Small\-24B\-Instruct\([Liu et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib44)\)is an instruction\-tuned model from the Mistral family\. It provides strong general\-purpose reasoning and instruction\-following capabilities with a relatively compact parameter scale compared with larger judge models\.

Qwen3\.5\-27B\([Team, 2026](https://arxiv.org/html/2609.13737#bib.bib43)\)is a large instruction\-tuned model from the Qwen family\. It is trained for broad natural language understanding, reasoning, and generation tasks, making it suitable as a general LLM\-based judgment model\.

LlamaGuard3\-8B\([Inan et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib45)\)is a safety moderation model built on the Llama family\. It is designed to classify user and model content according to safety\-related categories and produce moderation\-oriented decisions\.

Qwen3Guard\-8B\([Zhao et al\., 2025](https://arxiv.org/html/2609.13737#bib.bib46)\)is a safety guardrail model from the Qwen family\. It is specifically developed for safety classification and moderation, providing an open\-source guardrail model with a safety\-oriented objective\.

RoBERTa\([Liu et al\., 2019](https://arxiv.org/html/2609.13737#bib.bib7)\)is a robustly optimized BERT\-based encoder model designed for natural language understanding and text classification tasks\. It has been widely adopted as a strong pretrained language representation model across various NLP benchmarks\.

### B\.2Method\-based Baselines

For method\-based baselines, we follow their original feature extraction procedures and train them using the same output\-level harmfulness labels and train/validation/test splits as ForeSight\.

MULI\([Hu et al\., 2024](https://arxiv.org/html/2609.13737#bib.bib14)\)is a logits\-based early forecasting method for detecting harmful or toxic generations\. It exploits early decoding\-time output distributions as predictive signals before the full response is generated\.

Latent Guard\([Zhao et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib24)\)is a latent\-representation\-based safety detection method\. It identifies safety risks from the internal representations of language models, rather than relying solely on surface\-level generated text\.

TPC\([Oldfield et al\., 2026](https://arxiv.org/html/2609.13737#bib.bib33)\)extends linear probes with truncated polynomial classifiers to predict toxicity from language model activations\.

## Appendix CExperimental Settings and Hyperparameters

This section summarizes the implementation settings and hyperparameters used for ForeSight\. The layer\-wise L1 regularization strength and the number of K\-means clusters are selected based on validation performance\. All experiments use a random seed for reproducibility\.

Table 8:Implementation settings and hyperparameters for ForeSight\.
## Appendix DAnalysis of Sparse Distillation

### D\.1Cluster Number Sensitivity

We examine the effect of the cluster numberkkacross datasets\. As shown in Figure[8](https://arxiv.org/html/2609.13737#A1.F8), competitive performance is mostly concentrated aroundk=14k=14to1818, while very smallkkvalues often degrade performance, suggesting that overly coarse clustering may merge distinct activation patterns\. We therefore selectkkby validation performance rather than using a universal fixed value\.

### D\.2Sensitivity to Hyperparameter Selection

To examine whether ForeSight is overly dependent on validation\-based hyperparameter selection, we evaluate a fixed preset configuration across all five Qwen3\-8B datasets\. Specifically, we use the same L1 regularization strength, clustering configuration, and cross\-layer aggregation weights without dataset\-specific tuning\. As shown in Table[9](https://arxiv.org/html/2609.13737#A4.T9), the fixed preset reduces the average F1 score by only 0\.97 points compared with validation\-selected settings, while maintaining a substantial improvement over the strongest guardrail baseline Qwen3Guard\-8B \(55\.14 F1\)\. These results indicate that ForeSight is relatively robust to hyperparameter choices rather than relying on extensive validation tuning\.

Table 9:Sensitivity analysis of hyperparameter selection on Qwen3\-8B\. A fixed preset configuration causes only a minor performance drop compared with validation\-selected hyperparameters\.
### D\.3Sensitivity to Neuron Retention Threshold

We study the neuron retention threshold, which controls the cumulative saliency mass used to retain neurons\. As shown in Figure[9](https://arxiv.org/html/2609.13737#A4.F9), performance drops under very small thresholds, peaks around 0\.35, and shows no consistent gains with larger thresholds, supporting a moderate threshold that balances signal preservation and noise reduction\.

Figure 9:Effect of the neuron retention threshold on HarmEval with Llama\-3\.1\-8B\. The threshold controls the cumulative saliency mass used to retain neurons\.
### D\.4Sensitivity to Linear Classifier Regularization

We analyze the effect ofℓ1\\ell\_\{1\}regularization strengthCCin the layer\-wise linear classifier\. As shown in Figure[10](https://arxiv.org/html/2609.13737#A4.F10), the optimalCCvaries across layers: shallower layers tend to prefer largerCC, while deeper layers more often favor smallerCC, supporting layer\-specific sparse selection\.

![Refer to caption](https://arxiv.org/html/2609.13737v1/probe_regularization.png)Figure 10:Layer\-wise effect ofℓ1\\ell\_\{1\}regularization on HarmEval with Llama\-3\.1\-8B\. Lines show the test F1 and ACC of each layer\-wise classifier, while the shaded background indicates the optimal regularization strengthCCacross layers\.![Refer to caption](https://arxiv.org/html/2609.13737v1/borderline.png)Figure 11:Relationship between judge harmfulness vote count and ForeSight harmfulness score across datasets\. Boxplots show the harmful\-class logit grouped by the number of judges assigning a harmful label; vote counts 1 and 2 correspond to non\-unanimous borderline cases\.
### D\.5Analysis of Probing Layer Selection

Table 10:Comparison with single\-layer probing baselines across datasets with Llama\-3\.1\-8B\.We analyze the performance of layer\-wise probes to examine whether a single layer is sufficient for early harmfulness forecasting\. As shown in Table[10](https://arxiv.org/html/2609.13737#A4.T10), the best single\-layer probe underperforms ForeSight, indicating that aggregating sparse safety signals across layers provides additional predictive benefits\.

### D\.6Further Backbone Generalization

To further examine whether ForeSight generalizes beyond the original backbones, we evaluate it on Mistral\-7B\([Jiang et al\., 2023](https://arxiv.org/html/2609.13737#bib.bib48)\), an additional backbone from a different model family\. As shown in Table[11](https://arxiv.org/html/2609.13737#A4.T11), saliency\-based feature selection consistently improves performance over using full hidden states, increasing F1 from 80\.00 to 86\.68\. Combining saliency selection with k\-means compression further achieves 88\.32 F1 and 92\.59 ACC with only 640 dimensions\. These results suggest that first\-token safety signals and sparse feature selection are not limited to Llama\-3\.1\-8B and Qwen3\-8B, but can extend to other LLM backbones\.

Table 11:Backbone generalization on Mistral\-7B across datasets\.
### D\.7Training Data Scaling

To examine the impact of training data size, we retrain all trainable methods using 25%, 50%, 75%, and 100% of the pooled training data from five datasets, while keeping the validation and test sets unchanged\.

As shown in Table[12](https://arxiv.org/html/2609.13737#A4.T12), ForeSight achieves the best F1 score under all settings and reaches 72\.14 F1 with full data, outperforming the strongest baseline by 3\.32 points\. These results show that ForeSight remains effective as training data scales\.

Table 12:Training data scaling results\. Each method is retrained using different ratios of the pooled training data from five datasets\.Table 13:Macro\-level failure categories of ForeSight\. FP denotes non\-harmful outputs incorrectly predicted as harmful, while FN denotes harmful outputs missed by the detector\.

## Appendix EBorderline Evaluation

Although unanimous labels improve reliability, they may exclude ambiguous borderline cases\. We therefore analyze the non\-unanimous samples excluded from the main training and evaluation splits\.

For each sample, we use the judge harmfulness vote count as a graded ambiguity signal and compare it with ForeSight’s harmful\-class logit\. As shown in Figure[11](https://arxiv.org/html/2609.13737#A4.F11), higher vote counts generally correspond to higher predicted harmfulness scores\. Table[14](https://arxiv.org/html/2609.13737#A5.T14)further shows positive rank correlations across datasets, with an average Spearman correlation of 0\.343 and Kendall correlation of 0\.240\. These results indicate that ForeSight captures graded early safety signals in ambiguous responses, although borderline harmfulness remains challenging\.

Table 14:Rank correlation between judge harmfulness vote count and ForeSight’s harmful\-class logit\. The average row reports the macro average across datasets\.
## Appendix FFailure Case Analysis

As shown in Table[13](https://arxiv.org/html/2609.13737#A4.T13), we group errors into four macro\-level patterns\. Input\-dominant false positives occur when safety\-sensitive prompts lead to safe or evasive responses, suggesting that early hidden states may sometimes over\-emphasize prompt\-level risk\. False negatives are more diverse, often involving harmful content masked by safety\-oriented language, jailbreak\-style framing, or atypical surface forms\.

## Appendix GForecast\-Guided Resampling

Figure 12:Example of forecast\-guided first\-token resampling\. A lower\-risk first\-token trajectory leads to a response judged safe under the same prompt\.Although ForeSight is designed for early harmfulness forecasting, we further test whether its score can guide a simple intervention\. Given an initial tokeny1y\_\{1\}, we resample within a limited budget when its predicted risk is high, and continue generation from the first low\-risk candidate or, if unavailable, the lowest\-risk one\.

Figure[12](https://arxiv.org/html/2609.13737#A7.F12)shows an example where resampling changes a high\-risk trajectory into a safe response\. On 25 initially harmful HarmEval cases from Llama\-3\.1\-8B\-Instruct, ForeSight selects low\-risk trajectories for 21 cases, with 17 final responses judged safe\. We view this as a preliminary sanity check rather than a complete defense\.

Similar Articles

Strengthening our Frontier Safety Framework

Google DeepMind Blog

DeepMind published the third iteration of its Frontier Safety Framework, expanding risk domains to include harmful manipulation and misalignment risks, with refined risk assessment processes and enhanced governance protocols for advanced AI models.