Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training

arXiv cs.CL Papers

Summary

This paper proposes LayerTracer, an interpretable framework for layer allocation in continued pre-training, demonstrating that freezing deep layers while training shallow ones outperforms full-parameter fine-tuning. It offers a low-cost, actionable strategy for resource-constrained teams optimizing Large Language Models.

arXiv:2605.11416v1 Announce Type: new Abstract: Selective layer-wise updates are essential for low-cost continued pre-training of Large Language Models (LLMs), yet determining which layers to freeze or train remains an empirical black-box problem due to the lack of interpretable guidance. To address this issue, we propose LayerTracer, an architecture-agnostic diagnostic framework that reveals the evolution patterns of layer-wise representations and stability by locating task execution positions and quantifying layer sensitivity. Analysis results reveal that deep layers act as critical regions for task execution and maintain high stability against disruptive updates. Guided by this finding, we conduct three controlled continued pre-training trials to compare diverse freeze-train strategies, demonstrating that training shallow layers while freezing deep layers consistently outperforms full-parameter fine-tuning and the opposite allocation on both C-Eval and CMMLU benchmarks. We further present a hybrid model case study, which validates that placing high-quality pre-trained modules in deep layers effectively preserves inherent knowledge of the model. This work delivers a low-cost and interpretable solution for resource-constrained teams, offering actionable guidance for layer-wise parameter allocation in continued pre-training and hybrid model construction.
Original Article
View Cached Full Text

Cached at: 05/13/26, 06:13 AM

# Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training
Source: [https://arxiv.org/html/2605.11416](https://arxiv.org/html/2605.11416)
Yu\-Hang Wu1,2, Qin\-Yuan Liu1, Qiu\-Yang Zhao1, Bo Jiang1, Jiang\-Feng Yang1, Qing\-Wei Cong1 1Nanhu Research Institute of China Electronic Science and Technology 2School of Electronic and Electrical Engineering, Shanghai University of Engineering Science

###### Abstract

Selective layer\-wise updates are essential for low\-cost continued pre\-training of Large Language Models \(LLMs\), yet determining which layers to freeze or train remains an empirical black\-box problem due to the lack of interpretable guidance\. To address this issue, we propose LayerTracer, an architecture\-agnostic diagnostic framework that reveals the evolution patterns of layer\-wise representations and stability by locating task execution positions and quantifying layer sensitivity\. Analysis results reveal that deep layers act as critical regions for task execution and maintain high stability against disruptive updates\. Guided by this finding, we conduct three controlled continued pre\-training trials to compare diverse freeze\-train strategies, demonstrating that training shallow layers while freezing deep layers consistently outperforms full\-parameter fine\-tuning and the opposite allocation on both C\-Eval and CMMLU benchmarks\. We further present a hybrid model case study, which validates that placing high\-quality pre\-trained modules in deep layers effectively preserves inherent knowledge of the model\. This work delivers a low\-cost and interpretable solution for resource\-constrained teams, offering actionable guidance for layer\-wise parameter allocation in continued pre\-training and hybrid model construction\.

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre\-Training

Yu\-Hang Wu1,2, Qin\-Yuan Liu1, Qiu\-Yang Zhao1, Bo Jiang1, Jiang\-Feng Yang1, Qing\-Wei Cong1††thanks:Corresponding author\.1Nanhu Research Institute of China Electronic Science and Technology2School of Electronic and Electrical Engineering, Shanghai University of Engineering Science

![Refer to caption](https://arxiv.org/html/2605.11416v1/x1.png)Figure 1:Comparison of three layer\-wise pre\-training strategies on the CEval datasetHuanget al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib57)\)using the Qwen3\-710M architecture\. \(1\) \(Freeze\-Deep/Train\-Shallow\) achieves the best performance, followed by \(2\) \(Full Training\), while \(3\) \(Freeze\-Shallow/Train\-Deep\) yields the worst results\.## 1Introduction

With the rapid development of Large Language Models \(LLMs\), leading general\-purpose models such as GPTOpenAI \([2023](https://arxiv.org/html/2605.11416#bib.bib5)\), LLaMATouvronet al\.\([2023b](https://arxiv.org/html/2605.11416#bib.bib7),[a](https://arxiv.org/html/2605.11416#bib.bib6)\); Dubeyet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib8)\), and QwenTeam \([2024b](https://arxiv.org/html/2605.11416#bib.bib1),[a](https://arxiv.org/html/2605.11416#bib.bib2)\); Yanget al\.\([2025a](https://arxiv.org/html/2605.11416#bib.bib3)\)have achieved significant breakthroughs in semantic understanding, logical reasoning, and generation, thereby accelerating industrial adoption and engineering deploymentZhaoet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib17)\); Liuet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib40)\); Abo El\-Enenet al\.\([2025](https://arxiv.org/html/2605.11416#bib.bib18)\); Wuet al\.\([2025](https://arxiv.org/html/2605.11416#bib.bib19)\)\. However, the superior performance of these models relies heavily on massive high\-quality pretraining corpora and large\-scale computing clusters, creating substantial technical barriersBaiet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib37)\); Touvronet al\.\([2023b](https://arxiv.org/html/2605.11416#bib.bib7)\)\. For small and medium\-sized teams with limited computational power and data reserves, pretraining general\-purpose models from scratch is largely infeasible\. Consequently, a mainstream paradigm for vertical domain modeling has emerged\. This approach leverages the stacked decoder layers of open\-source pretrained models by freezing a portion of the decoder layers, while only training the remaining native layers as well as newly added custom layers\. Combined with high\-quality domain\-low\-budget continued pre\-training data for domain adaptation, this strategy enables the rapid deployment of commercially viable vertical domain models at minimal computational and data costsBaoet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib45)\); Chenet al\.\([2023b](https://arxiv.org/html/2605.11416#bib.bib46)\); Roziereet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib47)\); Chenet al\.\([2023a](https://arxiv.org/html/2605.11416#bib.bib48)\); Labraket al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib49)\)\.

Although this approach has established a standardized pipeline, it suffers from a critical limitation: the lack of interpretable guidance for layer\-wise freeze–train allocation during continued pre\-training\. Thus, small and medium\-sized teams are forced to rely on empirical heuristics when determining which layers to freeze or train, leaving the allocation process largely opaque\. This reliance on trial\-and\-error often leads to undesirable overwriting of pre\-trained knowledge and performance fluctuations, while substantially increasing computational costs\. As illustrated in Figure[1](https://arxiv.org/html/2605.11416#S0.F1), the specific arrangement of frozen and trainable layers critically dictates convergence stability and final performance\. Therefore, uncovering functional differentiation patterns across layers and deriving actionable freeze\-train rules emerges as a critical imperative\. Specifically, we aim to address the following two research questions:

- •RQ1: What are the functional differentiation patterns across layers in pretrained models?
- •RQ2: How can these patterns inform practical layer\-wise freeze/train allocation strategies for continued pre\-training?

![Refer to caption](https://arxiv.org/html/2605.11416v1/x2.png)Figure 2:The architectures of Qwen3 model and Qwen3\.5 model\. Qwen3 adopts a single architecture, while Qwen3\.5 is a hybrid architecture with a 3:1 ratio of Full Attention to Linear Attention\.Nevertheless, existing research offers limited guidance for layer\-wise parameter allocation in continued pre\-training\. Although prior layer probing studiesBelinkov \([2022b](https://arxiv.org/html/2605.11416#bib.bib25)\); Gurneeet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib27)\); Juet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib29)\); Xiaoet al\.\([2025](https://arxiv.org/html/2605.11416#bib.bib30)\); Eisenstadtet al\.\([2025](https://arxiv.org/html/2605.11416#bib.bib32)\); Zhanget al\.\([2026](https://arxiv.org/html/2605.11416#bib.bib31)\)and parameter\-efficient tuning methodsBaoet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib45)\); Liuet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib40)\)have yielded valuable insights into internal model dynamics and fine\-tuning strategies, they mainly focus on interpreting model behaviors or task\-specific fine\-tuning, rather than producing actionable rules for layer freeze\-train decisions\. More importantly, few existing efforts develop quantitative paradigms to characterize intrinsic layer\-wise properties, resulting in a lack of interpretable principles to address the core research questionsRQ1andRQ2\.

To address these issues, we propose LayerTracer, a diagnostic framework for layer\-wise freeze/train allocation in continued pre\-training\. It quantifies two core metrics:Task Particle \(TP\)identifies the layers where task probability undergoes meaningful relative shifts, marking where task evidence actively consolidates\.Layer\-wise Sensitivity \(LS\)measures the relative change in Jensen\-Shannon divergenceLin \([1991](https://arxiv.org/html/2605.11416#bib.bib20)\)across consecutive layers under controlled perturbation, capturing zones Sensitive to disruptive updates\. Across the Qwen3 series, we observe a consistent finding: shallow layers exhibit higher sensitivity to perturbations, while deep layers consolidate task evidence and stabilize execution\. Based on this finding, we derive a practical allocation rule: freeze deep pre\-trained layers and train shallow ones\. We validate this rule through three controlled continued pre\-training experiments, complemented by a hybrid architecture case study to simulate its practical value in resource\-constrained industrial scenarios\. Results show that Our train\-shallow/freeze\-deep strategy achieves a notable relative improvement over the reverse allocation, with an average gain of 15\.72% across C\-EvalHuanget al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib57)\)and CMMLULiet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib58)\)\.

## 2Related Work

### 2\.1Evolution of LLM Architectures

The TransformerVaswaniet al\.\([2017](https://arxiv.org/html/2605.11416#bib.bib9)\)remains the dominant backbone across modern foundation modelsOpenAI \([2023](https://arxiv.org/html/2605.11416#bib.bib5)\); Touvronet al\.\([2023a](https://arxiv.org/html/2605.11416#bib.bib6),[b](https://arxiv.org/html/2605.11416#bib.bib7)\); Dubeyet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib8)\); Yanget al\.\([2025a](https://arxiv.org/html/2605.11416#bib.bib3)\)\. To overcome the scalability bottleneck of self\-attention, efficient alternatives such as State Space ModelsGu and Dao \([2024](https://arxiv.org/html/2605.11416#bib.bib10)\); Dao and Gu \([2024](https://arxiv.org/html/2605.11416#bib.bib11)\), Linear AttentionAhnet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib12)\), and GatedDeltaNetYanget al\.\([2025b](https://arxiv.org/html/2605.11416#bib.bib15)\)have been proposed, giving rise to hybrid systems like JambaLieberet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib14)\)and Qwen3\.5Team \([2025](https://arxiv.org/html/2605.11416#bib.bib4)\)that interleave full\-attention and efficient layers as illustrated in Figure[2](https://arxiv.org/html/2605.11416#S1.F2)\. Such designs enable low\-cost domain adaptation by freezing pre\-trained foundation layers while tuning lightweight modulesBaoet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib45)\); Chenet al\.\([2023b](https://arxiv.org/html/2605.11416#bib.bib46)\); Roziereet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib47)\); Chenet al\.\([2023a](https://arxiv.org/html/2605.11416#bib.bib48)\); Labraket al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib49)\)\. Nevertheless, layer\-wise freeze–train decisions remain purely heuristic\. Without quantitative and interpretable guidance, improper allocation easily breaks knowledge coherence and degrades robustness, especially in fragile hybrid architectures\.

### 2\.2Layer\-Wise Representation Analysis

Layer\-wise interpretability has evolved from early linguistic probingPimentelet al\.\([2020](https://arxiv.org/html/2605.11416#bib.bib24)\); Belinkov \([2022a](https://arxiv.org/html/2605.11416#bib.bib23)\); Youssefet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib26)\)to advanced mechanistic analysis that locates task\-critical components and knowledge boundariesGurneeet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib27)\); Menget al\.\([2022](https://arxiv.org/html/2605.11416#bib.bib36)\); Gevaet al\.\([2021](https://arxiv.org/html/2605.11416#bib.bib35)\); Juet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib29)\); Xiaoet al\.\([2025](https://arxiv.org/html/2605.11416#bib.bib30)\)\. Closely related are logit/tuned lens and activation patching and causal tracing techniquesMenget al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib41)\); Hernandezet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib42)\); Liuet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib43)\), which project intermediate states through the LM head or intervene on activations to reveal internal dynamics\. However, these works focus on post\-hoc interpretation rather than actionable guidance\. They lack quantitative metrics for layer robustness and cannot be directly generalized to hybrid models\. No prior framework unifies task localization and sensitivity measurement for continued pre\-training\.

### 2\.3Parameter\-Efficient Adaptation and Allocation

Parameter\-efficient tuning methods such as LoRAHuet al\.\([2021](https://arxiv.org/html/2605.11416#bib.bib50)\), AdaLoRAZhanget al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib52)\), and DCFTZhanget al\.\([2025](https://arxiv.org/html/2605.11416#bib.bib51)\)optimize layer\-wise allocation for fine\-tuning, with extensions to instruction tuning and alignment via RLHFSchulmanet al\.\([2017](https://arxiv.org/html/2605.11416#bib.bib53)\); Rafailovet al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib54)\); Guoet al\.\([2025](https://arxiv.org/html/2605.11416#bib.bib55)\)\. Yet these approaches rely on task\-specific gradients or labeled data, making them unsuitable for unsupervised continued pre\-training on raw text\. Existing strategies also ignore intrinsic layer stability and fail to generalize to hybrid architectures\. This creates a critical need for a gradient\-free, interpretable diagnostic to guide principled layer allocation\.

![Refer to caption](https://arxiv.org/html/2605.11416v1/x3.png)Figure 3:Overview of the LayerTracer framework\. \(a\) Baseline projection: hidden states at each layer are projected via the shared LM head to obtain the target token probabilityPt​\(l\)P\_\{t\}\(l\), witht∗t^\{\*\}selected from the final distributionPP\. \(b\) Task Particle: we compute the relative probability shiftRatio​\(l\)\\mathrm\{Ratio\}\(l\)across consecutive layers\. Layers satisfyingRatio​\(l\)\>0\\mathrm\{Ratio\}\(l\)\>0form a execution interval, marking depths where task evidence actively consolidates\. \(c\) Layer\-wise Sensitivity: we apply context\-targeted masking perturbations at layerlland compute the relative fluctuationΔ​JS​\(l\)\\Delta\\mathrm\{JS\}\(l\)of JS divergence across adjacent layers, identifying zones highly sensitive to parameter updates and information flow disruptions\.

## 3Method

To answerRQ1andRQ2raised in this paper, this section proposes a hierarchical layer analysis framework named LayerTracer\. The overview of the LayerTracer framework is illustrated in Figure[3](https://arxiv.org/html/2605.11416#S2.F3)\.

### 3\.1Preliminary

We first define unified notations for layer\-wise analysis\. The input follows a composite task structuret=s1⊕s2t=s\_\{1\}\\bm\{\\oplus\}s\_\{2\}, wheres1s\_\{1\}denotes context examples ands2s\_\{2\}denotes the query input\. Given such structured input, the model outputs a series of hidden states\{h1,h2,…,hN\}\\\{h\_\{1\},h\_\{2\},\\dots,h\_\{N\}\\\}, wherehlh\_\{l\}represents the hidden state at thell\-th layer\. All layer\-wise hidden states are projected to the vocabulary space via the shared final\-layer LM head, ensuring consistent distribution calibration and controlled variables across layers\.

The final output probability distributionPPis obtained by projecting the last\-layer hidden statehNh\_\{N\}through the shared LM head\. We define the target tokent∗t^\{\*\}as the token with the maximum probability inPP\. This token is determined by the model’s own final prediction rather than external labels, and is used to trace how the model’s preferred answer emerges across layers\. LetPt​\(l\)P\_\{t\}\(l\)denote the probability oft∗t^\{\*\}derived from the hidden statehlh\_\{l\}projected by the shared LM head\. For brevity, we refer toPt​\(l\)P\_\{t\}\(l\)as task probability in subsequent analysis, and the collection of layers with meaningfulPt​\(l\)P\_\{t\}\(l\)shifts as task evidence consolidation zones\. For a masking perturbation applied at layerll, we mask only the contexts1s\_\{1\}while keeping the querys2s\_\{2\}unchanged and useQ​\(l\)Q\(l\)to denote the final output distribution after perturbation\. The Jensen\-Shannon divergence between two distributionsPPandQQis denoted asJS​\(P∥Q\)\\mathrm\{JS\}\(P\\parallel Q\)\.

### 3\.2Task Particle

TheTask Particle \(TP\)is designed to locate the critical layer where the model initiates task execution\. Following the notation in section[3\.1](https://arxiv.org/html/2605.11416#S3.SS1), for each layerll, we compute the probability of the target tokent∗t^\{\*\}and measure its relative increase ratio between consecutive layers:

Ratio​\(l\)=\|Pt​\(l\)−Pt​\(l−1\)Pt​\(l\)\+ϵ\|,\\mathrm\{Ratio\}\(l\)=\\left\|\\frac\{P\_\{t\}\(l\)\-P\_\{t\}\(l\-1\)\}\{P\_\{t\}\(l\)\+\\epsilon\}\\right\|,\(1\)whereϵ=10−6\\epsilon=10^\{\-6\}ensures numerical stability, andPt​\(l−1\)P\_\{t\}\(l\-1\)denotes the task probability from the preceding layer\. The Task Particle is formally defined as the set of layers whereRatio​\(l\)\>0\\mathrm\{Ratio\}\(l\)\>0, forming a continuous task\-execution interval\. Unlike single\-point localization methods, this interval\-based view captures all depths where task evidence undergoes meaningful relative shifts, encompassing both task probability promotion and suppression\. This continuous zone marks the region where the model actively consolidates and executes task\-relevant representations\.

![Refer to caption](https://arxiv.org/html/2605.11416v1/x4.png)\(\(a\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x5.png)\(\(b\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x6.png)\(\(c\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x7.png)\(\(d\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x8.png)\(\(e\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x9.png)\(\(f\)\)

Figure 4:Distribution details in the AntSynNET dataset, where 500 samples are evenly divided into 10 groups of 50 samples, denoted as Task 1–Task 10 for visualization\. \(a\)–\(c\) present the variation ofRatio\\mathrm\{Ratio\}across different layers for Qwen3\-0\.6B\-Base to Qwen3\-14B\-Base\. A deeper green color indicates a largerRatio\\mathrm\{Ratio\}, meaning that the target token probability changes more significantly at this layer, marking a key position for task execution\. \(d\)–\(f\) present the variation ofΔ​JS\\Delta\\mathrm\{JS\}across different layers, which measures the relative fluctuation of JS divergence between adjacent layers\. A deeper color indicates a larger abrupt change in layer sensitivity, reflecting higher vulnerability and lower robustness\.
### 3\.3Layer\-wise Sensitivity

To evaluate layer\-wise stability under perturbations, we propose the Layer\-wise Sensitivity \(LS\) metric\. Different from layer\-level evaluation that directly adopts raw JS divergence, LS focuses on the relative variation trend across consecutive layers, thereby capturing abrupt changes in the stability of context information flow\.

Specifically, letℐc\\mathcal\{I\}\_\{c\}andℐq\\mathcal\{I\}\_\{q\}denote the token indices of the context exampless1s\_\{1\}and the querys2s\_\{2\}, respectively\. For a perturbation applied at layerll, we intervene on the residual hidden states after thell\-th Transformer block by masking only the context\-token positions:

h~l,i=\{𝟎,i∈ℐc,hl,i,i∈ℐq\.\\tilde\{h\}\_\{l,i\}=\\begin\{cases\}\\mathbf\{0\},&i\\in\\mathcal\{I\}\_\{c\},\\\\ h\_\{l,i\},&i\\in\\mathcal\{I\}\_\{q\}\.\\end\{cases\}\(2\)
The perturbed hidden statesH~l\\tilde\{H\}\_\{l\}are then fed into the remaining layersl\+1,…,Nl\+1,\\dots,Nto obtain the final output distributionQ​\(l\)Q\(l\)\.

Given the original intact output distributionPPand the perturbed distributionQ​\(l\)Q\(l\)at layerll, we define the JS divergence at layerllas:

JS​\(l\)=JS​\(P∥Q​\(l\)\)\.\\mathrm\{JS\}\(l\)=\\mathrm\{JS\}\\big\(P\\parallel Q\(l\)\\big\)\.\(3\)
The standard Jensen\-Shannon divergence is formulated as:

JS​\(P∥Q​\(l\)\)\\displaystyle\\mathrm\{JS\}\\big\(P\\parallel Q\(l\)\\big\)=12​KL​\(P∥M​\(l\)\)\\displaystyle=\\frac\{1\}\{2\}\\mathrm\{KL\}\\big\(P\\parallel M\(l\)\\big\)\(4\)\+12​KL​\(Q​\(l\)∥M​\(l\)\),\\displaystyle\+\\frac\{1\}\{2\}\\mathrm\{KL\}\\big\(Q\(l\)\\parallel M\(l\)\\big\),whereM​\(l\)=12​\(P\+Q​\(l\)\)M\(l\)=\\frac\{1\}\{2\}\\big\(P\+Q\(l\)\\big\)represents the average mixed distribution, andKL\(⋅∥⋅\)\\mathrm\{KL\}\(\\cdot\\parallel\\cdot\)denotes the Kullback\-Leibler divergence:

KL​\(X∥Y\)=∑iX​\(i\)​log⁡X​\(i\)Y​\(i\)\.\\mathrm\{KL\}\(X\\parallel Y\)=\\sum\_\{i\}X\(i\)\\log\\frac\{X\(i\)\}\{Y\(i\)\}\.\(5\)
Based on the layer\-level JS divergence sequence, we further compute the relative fluctuation between adjacent layers:

Δ​JS​\(l\)=\|JS​\(l\)−JS​\(l−1\)JS​\(l−1\)\+ϵ\|,\\Delta\\mathrm\{JS\}\(l\)=\\left\|\\frac\{\\mathrm\{JS\}\(l\)\-\\mathrm\{JS\}\(l\-1\)\}\{\\mathrm\{JS\}\(l\-1\)\+\\epsilon\}\\right\|,\(6\)A largeΔ​JS​\(l\)\\Delta\\mathrm\{JS\}\(l\)value indicates an abrupt shift in layer sensitivity, meaning the stability of information flow changes significantly at this layer\. This metric effectively identifies critical layers that are sensitive to parameter updates\.

## 4Experiment

To empirically validate the diagnostic utility of LayerTracer and derive actionable freeze\-train rules, we conduct comprehensive layer\-wise analysis across Qwen3 variants\. Figure[4](https://arxiv.org/html/2605.11416#S3.F4)illustrates the distribution of Task Particle and Layer\-wise Sensitivity\.

### 4\.1Experimental Settings

Dataset\.We use a composite dataset constructed upon AntSynNETNguyenet al\.\([2017](https://arxiv.org/html/2605.11416#bib.bib16)\), following four standard categories for layer\-wise reasoning analysis: Knowledge–Algorithmic, Extractive–Knowledge, Extractive–Algorithmic, and Knowledge–Translation\. We reorganize each sample into a structured formt=s1⊕s2t=s\_\{1\}\\oplus s\_\{2\}\. This structural separation enables𝒔𝟏\\bm\{s\_\{1\}\}\-targeted masking perturbation, which helps distinguish the roles of shallow and deep layers by isolating context influence during layer\-wise sensitivity evaluation\. For clarity, dataset details are presented in Appendix[B](https://arxiv.org/html/2605.11416#A2)\.

Experimental Details\.All experiments are conducted on a server equipped with 8 NVIDIA H200 141G GPUs\.This study uses three Qwen3 variants of 0\.6B, 8B, and 14B parameters, as the primary subjects for layer\-wise analysis\. The experiment aims to characterize the distribution trends of Task Particles and Layer\-wise Sensitivity across varying model capacities\. For each layer, the hidden state is projected via the language modeling head to obtain the full vocabulary distribution, from which we retain only the top\-10 candidate tokens with the highest probabilities\. We verify in Appendix[C\.1](https://arxiv.org/html/2605.11416#A3.SS1)that results remain consistent when extending to top\-50, confirming the robustness of our layer\-wise patterns\.

### 4\.2Experimental Results

#### 4\.2\.1Deep Layers as Execution Zones

![[Uncaptioned image]](https://arxiv.org/html/2605.11416v1/x10.png)

Figure[4\(a\)](https://arxiv.org/html/2605.11416#S3.F4.sf1)–[4\(c\)](https://arxiv.org/html/2605.11416#S3.F4.sf3)visualize the layer\-wise distribution ofRatio​\(l\)\\mathrm\{Ratio\}\(l\), the relative change of task probability defined in Equation[1](https://arxiv.org/html/2605.11416#S3.E1)\. Following our definition in Section[3\.2](https://arxiv.org/html/2605.11416#S3.SS2), layers withRatio​\(l\)\>0\\mathrm\{Ratio\}\(l\)\>0constitute the Task Particle interval, marking where task evidence undergoes meaningful shifts\. We observe that highRatio\\mathrm\{Ratio\}values \(darker green\) concentrate in deep layers across all model scales, forming a stable task\-execution zone\. This indicates that deep layers serve as the primary region where the model executes task\-specific reasoning, while shallow layers show minimal task\-related activity\.

![Refer to caption](https://arxiv.org/html/2605.11416v1/x11.png)\(\(a\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x12.png)\(\(b\)\)

Figure 5:Layer\-wise analysis results on Nemotron\-12B\-Base, a Mamba\-based non\-Transformer architecture\. Consistent with observations on the Qwen3 series, the Task Particle concentrates in deep layers while shallow layers exhibit higher instability, verifying the universality of our findings across diverse model architectures\.
#### 4\.2\.2Shallow Layers as Sensitive Zones

![[Uncaptioned image]](https://arxiv.org/html/2605.11416v1/x13.png)

Figure[4\(d\)](https://arxiv.org/html/2605.11416#S3.F4.sf4)–[4\(f\)](https://arxiv.org/html/2605.11416#S3.F4.sf6)illustrateΔ​JS​\(l\)\\Delta\\mathrm\{JS\}\(l\), which quantifies the relative change in layer\-wise sensitivity under masking perturbations\. Based on Equation[6](https://arxiv.org/html/2605.11416#S3.E6), higherΔ​JS\\Delta\\mathrm\{JS\}values \(darker red\) indicate layers where sensitivity to context perturbations changes abruptly\. The results reveal that shallow layers exhibit significantly largerΔ​JS\\Delta\\mathrm\{JS\}fluctuations compared to deep layers, which maintain stable sensitivity profiles\. This indicates that shallow layers remain highly responsive to parameter updates, making them natural candidates for adaptation, while deep layers have consolidated stable knowledge representations\.

ModelsCEvalSTEMSocial ScienceHumanitiesOtherAverageQwen3\-710M\-Base \(Trainable\)26\.60 / 26\.5527\.19 / 27\.2126\.18 / 25\.8326\.53 / 25\.6726\.62 / 26\.32Qwen3\-710M\-Base \(Frozen/Trainable\)24\.86/24\.9725\.33/25\.6025\.51/25\.7125\.85/25\.9225\.39/25\.55Qwen3\-710M\-Base \(Trainable/Frozen\)27\.76/27\.6532\.07/32\.1728\.94/29\.3928\.82/28\.9929\.40/29\.55Table 1:Performance evaluation on the C\-Eval benchmarkHuanget al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib57)\)across three training strategies for Qwen3\-710M\-Base\. Each cell reports Generate / Logit scores\. The first row denotes the baseline with full\-parameter pre\-tuning\. The second row represents the strategy of freezing shallow layers and training deep layers\. The third row illustrates our strategy of training shallow layers and freezing deep layers\.Red valuesindicates performance better than the full\-parameter baseline, whilegreen valuesindicates performance worse than the baseline\.Bold valuesdenote the best performance across all settings\. Standard deviation details over five repeated runs for all experiments are provided in Appendix[C\.3](https://arxiv.org/html/2605.11416#A3.SS3)\.ModelsCMMLUSTEMSocial ScienceHumanitiesOtherAverageQwen3\-710M\-Base \(Trainable\)26\.60 / 26\.5527\.19 / 27\.2126\.18 / 25\.8326\.53 / 25\.6726\.62 / 26\.32Qwen3\-710M\-Base \(Frozen/Trainable\)24\.28/24\.1424\.98/26\.0125\.14/25\.5026\.30/26\.8025\.61/25\.11Qwen3\-710M\-Base \(Trainable/Frozen\)26\.69/26\.7230\.62/30\.6527\.93/28\.0027\.76/28\.3228\.25/28\.42Table 2:Performance evaluation on the CMMLU benchmarkLiet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib58)\)across the same three pre\-training strategies\. Settings and notations follow Table[1](https://arxiv.org/html/2605.11416#S4.T1)\.
#### 4\.2\.3Core Principle: Matching Zones with Strategies

![[Uncaptioned image]](https://arxiv.org/html/2605.11416v1/x14.png)

The above results reveal a consistent shallow\-sensitive and deep\-stable hierarchy, suggesting a simple layer\-wise allocation principle for continued pre\-training\. We instantiate this principle with a model\-midpoint split: the shallow half is treated as sensitive zones, where highΔ​JS\\Delta\\mathrm\{JS\}indicates greater sensitivity to perturbations and thus higher plasticity for absorbing domain\-specific changes\. In contrast, the deep half is treated as execution zones, where highRatio\\mathrm\{Ratio\}values indicate active task\-evidence consolidation and stabilized task execution\. Since these deep layers carry task\-execution functions and inherit high\-quality parameters from strong pre\-trained models, they should be preserved rather than frequently updated\. This midpoint split provides a symmetric parameter partition and a clean controlled setting for validation, instead of exhaustively searching for the optimal boundary\.

As illustrated in Figure[6](https://arxiv.org/html/2605.11416#S4.F6), we therefore match different strategies to different layer zones: shallow sensitive zones are trained to absorb domain\-specific patterns, while deep execution zones are frozen to preserve strong pre\-trained parameters and avoid disrupting consolidated task evidence\. Details as elaborated in[C\.2](https://arxiv.org/html/2605.11416#A3.SS2)\.

![Refer to caption](https://arxiv.org/html/2605.11416v1/x15.png)Figure 6:Illustration of the proposed zone\-strategy matching principle\. Shallow layers are treated as sensitive zones, where active training supports domain\-specific updates, while deep layers are treated as execution zones, where freezing preserves pre\-trained task evidence\.

## 5Why is LayerTracer Effective?

Guided by the hierarchical patterns revealed by LayerTracer in Section[4](https://arxiv.org/html/2605.11416#S4), we further evaluate whether the derived strategy leads to measurable gains in practical continued pre\-training settings\. While Section[4](https://arxiv.org/html/2605.11416#S4)focuses on representation dynamics and intrinsic layer properties, this section examines whether the LayerTracer\-guided allocation strategy improves downstream task performance\.

![[Uncaptioned image]](https://arxiv.org/html/2605.11416v1/x16.png)

### 5\.1Pre\-Training and Evaluation Setup

We conduct continued pre\-training on Qwen3\-710M\-Base using the high\-quality Chinese corpus CCI3\.0\-HQWanget al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib59)\)with 262B tokens, where all training samples are in the form of natural text streams rather than question–answer pairs, and compare three strategies: full\-parameter pre\-training as the baseline, freezing the shallow module while training the deep module, and training the shallow module while freezing the deep module\. Evaluation is performed on two standard Chinese benchmarks, C\-EvalHuanget al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib57)\)and CMMLULiet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib58)\), using two mainstream protocols widely adopted for pre\-trained model assessment: logit\-based evaluation infers answers directly from model logits to reflect inherent knowledge fitting, while generation\-based evaluation judges whether the correct answer is included within a 16\-token autoregressive output withdo\_sample=Falseto characterize generation quality, and consistency between these two metrics indicates reliable model behavior\. Detailed settings are summarized in Appendix[A](https://arxiv.org/html/2605.11416#A1)\.

DatasetTokensQwen Private Corpus36TCCI3\.0\-HQ262BTable 3:Statistics of pre\-training corpora\. The Qwen private corpus is a large\-scale high\-quality proprietary dataset built by the Qwen team, while CCI3\.0\-HQWanget al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib59)\)is a public high\-quality Chinese corpus with a smaller scale\.
### 5\.2Results and Analysis

As shown in Table[1](https://arxiv.org/html/2605.11416#S4.T1)and Table[2](https://arxiv.org/html/2605.11416#S4.T2), the train\-shallow/freeze\-deep strategy consistently achieves the best performance across both benchmarks and both evaluation protocols\. On C\-Eval, it improves the average score from 26\.62% / 26\.32% under full\-parameter continued pre\-training to 29\.40% / 29\.55%, yielding absolute gains of 2\.78% in Generate evaluation and 3\.23% in Logit evaluation\. In contrast, the reverse strategy, which freezes shallow layers and trains deep layers, drops to 25\.39% / 25\.55%, falling below the full\-parameter baseline\.

A similar trend is observed on CMMLU\. The train\-shallow/freeze\-deep strategy reaches 28\.25% / 28\.42% on average, outperforming the full\-parameter baseline by 1\.63% and 2\.10% under Generate and Logit evaluation, respectively\. By comparison, the freeze\-shallow/train\-deep strategy obtains only 25\.61% / 25\.11%, again showing clear degradation\. These results indicate that updating deep layers can disrupt pre\-trained task evidence, while preserving deep execution layers and adapting shallow layers leads to more reliable continued pre\-training\.

## 6A Case Study: Industrial Implications

![[Uncaptioned image]](https://arxiv.org/html/2605.11416v1/x17.png)

STEMSocial ScienceHumanitiesOtherAverageQwen3\-Nemotron\-Base10\.73 / 22\.749\.64 / 25\.7310\.54 / 20\.3511\.23 / 21\.649\.37 / 22\.62Nemotron\-Qwen3\-Base19\.40/28\.4018\.77/33\.3019\.37/25\.7819\.19/27\.4819\.19/28\.73Table 4:Performance evaluation on the C\-Eval benchmark for two hybrid architectures\. Nemotron\-Qwen3\-Base places pre\-trained Qwen3 decoder layers in the deep half, while Qwen3\-Nemotron\-Base places them in the shallow half\. Bold values denote the best performance\.ModelsCMMLUSTEMSocial ScienceHumanitiesOtherAverageQwen3\-Nemotron\-Base9\.72 / 22\.7410\.64 / 25\.7310\.07 / 22\.359\.23 / 20\.249\.87 / 22\.65Nemotron\-Qwen3\-Base16\.79/25\.8720\.70/28\.9421\.20/26\.9518\.00/27\.4319\.17/27\.30Table 5:Performance evaluation on the CMMLU benchmark for the same two hybrid strategies\. Settings and notations follow Table[4](https://arxiv.org/html/2605.11416#S6.T4)\.![Refer to caption](https://arxiv.org/html/2605.11416v1/x18.png)Figure 7:The architectures of Nemotron\-Qwen model and Qwen\-Nemotron model\. They are hybrid architectures with a 1:1 ratio of Full Attention to Linear Attention\.This section examines the industrial value of LayerTracer for resource\-constrained LLM development\. As hybrid architectures become more common, a key challenge is how to place heterogeneous modules while preserving the generalization ability of foundation models\. Guided by the LayerTracer principle of preserving deep execution zones while adapting shallow sensitive zones, we provide a practical case study for hybrid model construction\. Figure[7](https://arxiv.org/html/2605.11416#S6.F7)illustrates the hybrid architecture used in this study\.

To illustrate its industrial value, we construct two hybrid models, Nemotron\-Qwen3\-Base and Qwen3\-Nemotron\-Base, both with 705M parameters\. In Nemotron\-Qwen3\-Base, Mamba2\-based Nemotron layers are placed in the first half of the network and actively trained, while Qwen3 pre\-trained decoder layers are placed in the second half and frozen to preserve high\-quality pre\-trained parameters\. In Qwen3\-Nemotron\-Base, the allocation is reversed, placing Qwen3 layers in the first half and Nemotron layers in the second half\. We adopt the same pre\-training data, benchmarks, and evaluation metrics as Section[5](https://arxiv.org/html/2605.11416#S5)\.

Experimental results are reported in Table[4](https://arxiv.org/html/2605.11416#S6.T4)and Table[5](https://arxiv.org/html/2605.11416#S6.T5)\. On C\-Eval, placing Qwen3 pre\-trained layers in the deep half improves the average score from 9\.37% / 22\.62% to 19\.19% / 28\.73%, yielding absolute gains of 9\.82 and 6\.11 percentage points under Generate and Logit evaluation, respectively\. On CMMLU, the same strategy improves the average score from 9\.87% / 22\.65% to 19\.17% / 27\.30%, with gains of 9\.30% and 4\.65%\. The improvements are consistent across STEM, Social Science, Humanities, and Other categories\.

## 7Conclusion

This paper proposes LayerTracer, an interpretable diagnostic framework for layer\-wise freeze\-train allocation in continued pre\-training\. By combining Task Particle and Layer\-wise Sensitivity, LayerTracer reveals a consistent hierarchy: shallow layers are more sensitive and suitable for adaptation, while deep layers form stable execution zones where task evidence is consolidated\. Based on this finding, we derive a train\-shallow/freeze\-deep strategy\. Experiments show that this strategy outperforms both full\-parameter continued pre\-training and the reverse allocation, while the hybrid architecture case study further demonstrates the value of preserving high\-quality pre\-trained parameters in deep layers\. Overall, LayerTracer provides practical and low\-cost guidance for efficient model adaptation and hybrid model construction\.

## Limitations

LayerTracer is designed as an interpretable diagnostic framework for continued pre\-training, and its main limitation is computational scale rather than methodological scope\. Due to limited resources, we validate it on mainstream model sizes instead of extending the analysis to much larger parameter scales\. Even so, our experiments already cover comparatively large continued pre\-training settings, and the results are fully reproducible, stable, and consistent across benchmarks\. We believe the framework provides practical and effective guidance for layer\-wise allocation, while larger\-scale validation remains an important direction for future work\.

## Ethical Statement

This work focuses on fundamental interpretability and efficient adaptation techniques for large language models, with no involvement of unethical data collection, harmful content generation, or biased application scenarios\. All experiments comply with standard AI research ethics and model usage norms\. The proposed LayerTracer framework aims to improve training efficiency and stability while reducing computational costs, which supports green and responsible development of large models\.

## References

- M\. Abo El\-Enen, S\. Saad, and T\. Nazmy \(2025\)A survey on retrieval\-augmentation generation \(rag\) models for healthcare applications\.Neural Computing and Applications37\(33\),pp\. 28191–28267\.External Links:[Document](https://dx.doi.org/10.1007/s00521-025-11234-5)Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1)\.
- K\. Ahn, X\. Cheng, M\. Song, C\. Yun, A\. Jadbabaie, and S\. Sra \(2024\)Linear attention is \(maybe\) all you need \(to understand transformer optimization\)\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- J\. Bai, S\. Bai, Y\. Chu, Z\. Cui, K\. Dang, X\. Deng, Y\. Fan, W\. Ge, Y\. Han, F\. Huang,et al\.\(2023\)Qwen technical report\.arXiv preprint arXiv:2309\.16609\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1)\.
- Z\. Bao, W\. Chen, S\. Xiao, K\. Ren, J\. Wu, C\. Zhong, J\. Peng, X\. Huang, and Z\. Wei \(2023\)Disc\-medllm: bridging general large language models and real\-world medical consultation\.arXiv preprint arXiv:2308\.14346\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§1](https://arxiv.org/html/2605.11416#S1.p4.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- Y\. Belinkov \(2022a\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.Cited by:[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- Y\. Belinkov \(2022b\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p4.1)\.
- J\. Chen, X\. Wang, K\. Ji, A\. Gao, F\. Jiang, S\. Chen, H\. Zhang, D\. Song, W\. Xie, C\. Kong,et al\.\(2023a\)Huatuogpt\-ii, one\-stage training for medical adaption of llms\.arXiv preprint arXiv:2311\.09774\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- Z\. Chen, A\. H\. Cano, A\. Romanou, A\. Bonnet, K\. Matoba, F\. Salvi, M\. Pagliardini, S\. Fan, A\. Köpf, A\. Mohtashami,et al\.\(2023b\)Meditron\-70b: scaling medical pretraining for large language models\.arXiv preprint arXiv:2311\.16079\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- T\. Dao and A\. Gu \(2024\)Transformers are ssms: generalized models and efficient algorithms through structured state space duality\.arXiv preprint arXiv:2405\.21060\.Cited by:[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- A\. Dubey, A\. Jauhri, A\. Pandey,et al\.\(2024\)The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- R\. Eisenstadt, I\. Zimerman, and L\. Wolf \(2025\)Overclocking llm reasoning: monitoring and controlling thinking path lengths in llms\.External Links:2506\.07240Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p4.1)\.
- M\. Geva, R\. Schuster, J\. Berant, and O\. Levy \(2021\)Transformer feed\-forward layers are key\-value memories\.InProceedings of the Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 5484–5495\.Cited by:[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- A\. Gu and T\. Dao \(2024\)Mamba: linear\-time sequence modeling with selective state spaces\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.\(2025\)DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.Cited by:[§2\.3](https://arxiv.org/html/2605.11416#S2.SS3.p1.1)\.
- W\. Gurnee, N\. Nanda, M\. Pauly, K\. Harvey, D\. Troitskii, and D\. Bertsimas \(2023\)Finding neurons in a haystack: case studies with sparse probing\.InAdvances in Neural Information Processing Systems,Vol\.36\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p4.1),[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- E\. Hernandez, A\. S\. Sharma, T\. Haklay, K\. Meng, M\. Wattenberg, J\. Andreas, Y\. Belinkov, and D\. Bau \(2024\)Linearity of relation decoding in transformer language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen \(2021\)LoRA: low\-rank adaptation of large language models\.External Links:2106\.09685Cited by:[§2\.3](https://arxiv.org/html/2605.11416#S2.SS3.p1.1)\.
- Y\. Huang, Y\. Bai, Z\. Zhu, J\. Zhang, J\. Zhang, T\. Su, J\. Liu, C\. Lv, Y\. Zhang, J\. Lei, Y\. Fu, M\. Sun, and J\. He \(2023\)C\-eval: a multi\-level multi\-discipline chinese evaluation suite for foundation models\.InAdvances in Neural Information Processing Systems,Cited by:[§A\.4](https://arxiv.org/html/2605.11416#A1.SS4.p1.2),[Figure 1](https://arxiv.org/html/2605.11416#S0.F1),[§1](https://arxiv.org/html/2605.11416#S1.p5.1),[Table 1](https://arxiv.org/html/2605.11416#S4.T1),[§5\.1](https://arxiv.org/html/2605.11416#S5.SS1.p1.1)\.
- T\. Ju, W\. Sun, W\. Du, X\. Yuan, Z\. Ren, and G\. Liu \(2024\)How large language models encode context knowledge? a layer\-wise probing study\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p4.1),[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- Y\. Labrak, A\. Bazoge, E\. Morin, P\. Gourraud, M\. Rouvier, and R\. Dufour \(2024\)Biomistral: a collection of open\-source pretrained large language models for medical domains\.InFindings of the association for computational linguistics: acl 2024,Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- H\. Li, Y\. Zhang, F\. Koto, Y\. Yang, H\. Zhao, Y\. Gong, N\. Duan, and T\. Baldwin \(2024\)CMMLU: measuring massive multitask language understanding in Chinese\.InFindings of the Association for Computational Linguistics,Cited by:[§A\.4](https://arxiv.org/html/2605.11416#A1.SS4.p1.2),[§1](https://arxiv.org/html/2605.11416#S1.p5.1),[Table 2](https://arxiv.org/html/2605.11416#S4.T2),[§5\.1](https://arxiv.org/html/2605.11416#S5.SS1.p1.1)\.
- A\. Lieber, O\. Sharir, B\. Lenz, and Y\. Shoham \(2024\)Jamba: a hybrid transformer\-mamba language model\.arXiv preprint arXiv:2403\.19887\.Cited by:[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- J\. Lin \(1991\)Divergence measures based on the shannon entropy\.IEEE Transactions on Information Theory37\(1\),pp\. 145–151\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p5.1)\.
- F\. Liu, P\. Shi, X\. Liu, Y\. Zhang, and G\. Neubig \(2024\)What do probes actually probe? on the role of surface statistics in linguistic probing\.InProceedings of the Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 12345–12359\.Cited by:[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- T\. Liu, Y\. Li, Q\. Xie, X\. Wang, and H\. Li \(2023\)A survey on the robustness of large language models\.InFindings of the Conference on Empirical Methods in Natural Language Processing \(EMNLP Findings\),pp\. 14521–14538\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§1](https://arxiv.org/html/2605.11416#S1.p4.1)\.
- K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov \(2022\)Locating and editing factual associations in gpt\.InAdvances in Neural Information Processing Systems \(NeurIPS\),pp\. 17359–17372\.Cited by:[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- K\. Meng, D\. Bau, A\. Andonian, and Y\. Belinkov \(2023\)Mass editing memory in a transformer\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- K\. A\. Nguyen, S\. Schulte im Walde, and N\. T\. Vu \(2017\)Distinguishing antonyms and synonyms in a pattern\-based neural network\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers,Cited by:[§B\.1](https://arxiv.org/html/2605.11416#A2.SS1.p1.4),[§4\.1](https://arxiv.org/html/2605.11416#S4.SS1.p1.2)\.
- OpenAI \(2023\)GPT\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- T\. Pimentel, J\. Valvoda, R\. H\. Maudslay, R\. Zmigrod, A\. Williams, and R\. Cotterell \(2020\)Information\-theoretic probing for linguistic structure\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,pp\. 4609–4622\.Cited by:[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.Advances in neural information processing systems36,pp\. 53728–53741\.Cited by:[§2\.3](https://arxiv.org/html/2605.11416#S2.SS3.p1.1)\.
- B\. Roziere, J\. Gehring, F\. Gloeckle, S\. Sootla, I\. Gat, X\. E\. Tan, Y\. Adi, J\. Liu, R\. Sauvestre, T\. Remez,et al\.\(2023\)Code llama: open foundation models for code\.arXiv preprint arXiv:2308\.12950\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov \(2017\)Proximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[§2\.3](https://arxiv.org/html/2605.11416#S2.SS3.p1.1)\.
- Q\. Team \(2024a\)Qwen2 technical report\.arXiv preprint arXiv:2407\.10671\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1)\.
- Q\. Team \(2024b\)Qwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1)\.
- Q\. Team \(2025\)Qwen3\.5 official repository\.Cited by:[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- H\. Touvron, T\. Lavril, G\. Izacard,et al\.\(2023a\)LLaMA: open and efficient foundation language models\.arXiv preprint arXiv:2302\.13971\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- H\. Touvron, L\. Martin, K\. Stone,et al\.\(2023b\)LLaMA 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- A\. Vaswani, N\. Shazeer, N\. Parmar,et al\.\(2017\)Attention is all you need\.InAdvances in Neural Information Processing Systems \(NeurIPS\),pp\. 5998–6008\.Cited by:[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- L\. Wang, B\. Zhang, C\. Wu, H\. Zhao, X\. Shi, S\. Gu, J\. Li, Q\. Ma, T\. Pan, and G\. Liu \(2024\)CCI3\.0\-hq: a large\-scale chinese dataset of high quality designed for pre\-training large language models\.External Links:2410\.18505Cited by:[§5\.1](https://arxiv.org/html/2605.11416#S5.SS1.p1.1),[Table 3](https://arxiv.org/html/2605.11416#S5.T3)\.
- Y\. Wu, Y\. Xiong, H\. Zhang, J\. Zhang, and Z\. Zhou \(2025\)Sugar\-coated poison: benign generation unlocks llm jailbreaking\.arXiv preprint arXiv:2504\.05652\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1)\.
- C\. Xiao, H\. P\. Chan, H\. Zhang, M\. Aljunied, L\. Bing, N\. A\. Moubayed, and Y\. Rong \(2025\)Analyzing llms’ knowledge boundary cognition across languages through the lens of internal representation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics,Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p4.1),[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- A\. Yang, A\. Li, B\. Yang,et al\.\(2025a\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1),[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- S\. Yang, J\. Kautz, and A\. Hatamizadeh \(2025b\)Gated delta networks: improving mamba2 with delta rule\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.1](https://arxiv.org/html/2605.11416#S2.SS1.p1.1)\.
- P\. Youssef, O\. A\. Koras, M\. Li, J\. Schlötterer, and C\. Seifert \(2023\)Give me the facts\! a survey on factual knowledge probing in pre\-trained language models\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 15588–15605\.Cited by:[§2\.2](https://arxiv.org/html/2605.11416#S2.SS2.p1.1)\.
- H\. Zhang, Z\. Zhang, and M\. Wang \(2026\)Locate, steer, and improve: a practical survey of actionable mechanistic interpretability in large language models\.External Links:2601\.14004Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p4.1)\.
- J\. Zhang, Y\. Xiong, C\. Xia, D\. Zhu, and X\. Qiu \(2025\)Parameter\-efficient fine\-tuning of large language models via deconvolution in subspace\.InProceedings of the 31st International Conference on Computational Linguistics,Cited by:[§2\.3](https://arxiv.org/html/2605.11416#S2.SS3.p1.1)\.
- Q\. Zhang, M\. Chen, A\. Bukharin, P\. He, Y\. Cheng, W\. Chen, and T\. Zhao \(2023\)Adaptive budget allocation for parameter\-efficient fine\-tuning\.InThe Eleventh International Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2605.11416#S2.SS3.p1.1)\.
- W\. X\. Zhao, K\. Zhou, J\. Li,et al\.\(2023\)A survey of large language models\.arXiv preprint arXiv:2303\.18223\.Cited by:[§1](https://arxiv.org/html/2605.11416#S1.p1.1)\.

## Appendix ATraining and Evaluation Details

To ensure rigorous and reproducible comparisons, all continued pre\-training experiments were conducted under a unified configuration\. This appendix details the hardware setup, hyperparameter settings, and training efficiency metrics across all evaluated layer\-wise allocation strategies\.

### A\.1Hardware and Implementation\.

All experiments were implemented using the native Accelerate training framework and executed on a single server equipped with 8×\\timesNVIDIA H200 141G\. To strictly isolate the impact of layer\-wise parameter allocation, we maintained identical data pipelines, hardware environments, and training schedules across all trials\. Wall\-clock training time was measured from the first optimizer step to the final checkpoint save, explicitly excluding data loading and model initialization overhead\.

### A\.2Hyperparameter Configuration

We adopted a consistent hyperparameter configuration for all continued pre\-training runs, as summarized in Table[6](https://arxiv.org/html/2605.11416#A1.T6)\. Models were trained on the CCI3\.0\-HQ corpus \(~262B tokens\) for a single epoch with a sequence length of 4096 and BF16 precision\. Optimization was performed using AdamW \(β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.95\\beta\_\{2\}=0\.95\) with a learning rate of3×10−53\\times 10^\{\-5\}, a linear warmup ratio of 0\.1, and a weight decay of 0\.01\. The effective global batch size of 64 was achieved using a per\-GPU micro\-batch size of 8 with a single gradient accumulation step\.

### A\.3Training Efficiency and Strategy Mapping\.

Table[7](https://arxiv.org/html/2605.11416#A1.T7)reports the computational efficiency across all evaluated strategies\. Strategies \(1\)–\(3\) correspond to the vanilla Qwen3\-710M\-Base variants reported in Table[1](https://arxiv.org/html/2605.11416#S4.T1)of the main text: \(1\) full\-parameter training, \(2\) Train\-Shallow/Freeze\-Deep, and \(3\) Freeze\-Shallow/Train\-Deep\. Strategies \(4\) and \(5\) correspond to the hybrid architectures evaluated in Table[4](https://arxiv.org/html/2605.11416#S6.T4): \(4\) Nemotron\-Qwen3\-Base \(training shallow NemotronDecoderLayer while freezing deep QwenDecoderLayer\) and \(5\) Qwen3\-Nemotron\-Base \(training shallow QwenDecoderLayer while freezing deep NemotronDecoderLayer\)\. As demonstrated, partial\-parameter strategies reduce the trainable parameter count by approximately 50% relative to full\-parameter training, yielding consistent reductions in training time while preserving downstream performance\. Figure[8](https://arxiv.org/html/2605.11416#A1.F8)illustrates the training loss curves of the three models investigated in Section[4](https://arxiv.org/html/2605.11416#S4)\.

![Refer to caption](https://arxiv.org/html/2605.11416v1/x19.png)\(\(a\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x20.png)\(\(b\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x21.png)\(\(c\)\)

Figure 8:Training loss curves across three layer\-wise allocation strategies on the CCI3\.0\-HQ corpus\. \(a\) Train\-Shallow/Freeze\-Deep: shallow layers are updated while deep layers are frozen, showing stable convergence and the lowest final loss\. \(b\) Freeze\-Shallow/Train\-Deep: deep layers are updated while shallow layers are frozen, exhibiting slower convergence and higher final loss due to disruption of task\-critical representations\. \(c\) Full\-parameter Training: all layers are updated, serving as the baseline with intermediate convergence behavior\. The consistent pattern validates that preserving deep execution zones while adapting shallow sensitive zones yields more efficient and stable continued pre\-training\.
### A\.4Evaluation Details

To ensure fair comparison with existing baselines and official leaderboards, all evaluations are conducted in a zero\-shot setting without any task\-specific demonstrations or fine\-tuning\. The evaluation prompts strictly follow the official templates provided by C\-EvalHuanget al\.\([2023](https://arxiv.org/html/2605.11416#bib.bib57)\)and CMMLULiet al\.\([2024](https://arxiv.org/html/2605.11416#bib.bib58)\), preserving the original question formatting, option ordering, and answer extraction rules\. For generation\-based evaluation, we adopt the official decoding configuration with temperature set to0\.950\.95, top\-p sampling disabled, and a maximum output length of1616tokens to align with the benchmark protocols\. For logit\-based evaluation, we directly extract the probabilities of candidate answer tokens from the model’s vocabulary logits without any decoding,\. All evaluations are performed on the validation sets of both benchmarks, and results are reported as accuracy percentages across four domain categories \(STEM, Social Science, Humanities, and Other\) as well as the overall average\.

HyperparameterValueSequence Length4096PrecisionBF16Optimizer \(AdamW\)β1=0\.9\\beta\_\{1\}=0\.9,β2=0\.95\\beta\_\{2\}=0\.95Learning Rate3×10−53\\times 10^\{\-5\}Warmup Ratio0\.1Weight Decay0\.01Per\-GPU Micro Batch Size8Gradient Accumulation Steps1Global Batch Size64Training Epochs1Training Tokens~262BTable 6:Unified hyperparameter configuration for all continued pre\-training experiments\.StrategyTrainable ParametersTraining Time\(1\) Full\-Parameter0\.71B~141h\(2\) Train\-Shallow / Freeze\-Deep0\.35B~112h\(3\) Freeze\-Shallow / Train\-Deep0\.35B~107h\(4\)–\(5\) Hybrid Architectures0\.35B~115hTable 7:Training efficiency comparison across layer\-wise allocation strategies\. Strategies \(1\)\-\(3\) correspond to vanilla Qwen3\-710M\-Base variants, while \(4\)\-\(5\) correspond to the hybrid architectures \(Nemotron\-Qwen3\-705M\-Base and Qwen3\-Nemotron\-705M\-Base\)\.IDPrompt1Example:good\-\>Bad, no\-Yes; Query:bad\-\>2Example:hot\-\>Cold, big\-Small; Query:cold\-\>3Example:fast\-\>Slow, light\-Heavy; Query:slow\-\>4Example:love\-\>Hate, start\-End; Query:hate\-\>5Example:day\-\>Night, up\-Down; Query:night\-\>6Example:rich\-\>Poor, win\-Lose; Query:poor\-\>7Example:early\-\>Late, high\-Low; Query:late\-\>8Example:strong\-\>Weak, loud\-Quiet; Query:weak\-\>9Example:hard\-\>Soft, dark\-Bright; Query:soft\-\>10Example:full\-\>Empty, open\-Close; Query:empty\-\>Table 8:Representative structured prompts from the AntSynNET\-based dataset\. These examples illustrate the prompt format rather than the full set of visualization groups\.

## Appendix BExperimental Dataset

### B\.1Dataset Construction

The original AntSynNET dataset\(Nguyenet al\.,[2017](https://arxiv.org/html/2605.11416#bib.bib16)\)provides word pairs for lexical semantic relation tasks, where each instance is represented as a tuple⟨w1,w2⟩\\langle w\_\{1\},w\_\{2\}\\rangleindicating an antonym or synonym relation, such as\(good, bad\)or\(hot, cold\)\. To support our analysis in Section[4](https://arxiv.org/html/2605.11416#S4), we reformulate each instance into a structured promptt=s1⊕s2t=s\_\{1\}\\oplus s\_\{2\}, wheres1s\_\{1\}contains contextual demonstration pairs ands2s\_\{2\}contains the query word to be predicted\. The structured prompt follows the template:

Example: \[pair\_1\], \[pair\_2\];Query: \[target\_word\]\-\>
Here,s1s\_\{1\}spans the contextual examples before the query, whiles2s\_\{2\}corresponds to the query segment\. Specifically, we perturb token positions belonging tos1s\_\{1\}while keeping the query segments2s\_\{2\}unchanged\. In this way, the task query itself remains intact, and the resulting distributional change mainly reflects how each layer responds to contextual information disruption\. Thus,s1s\_\{1\}\-targeted masking serves as a unified and controlled disturbance for measuring layer\-wise sensitivity\.

### B\.2Why Controlled Lexical Prompts

LayerTracer tracks how the probability of a target token evolves across layers by projecting intermediate hidden states through the LM head\. Therefore, the diagnostic setting requires the first predicted token to be meaningful and directly related to the answer\. Complex QA benchmarks, such as math or medical QA, are less suitable for this purpose because the model may need to generate a reasoning process or follow a specific answer format\. In contrast, the structured AntSynNET prompts provide short and controlled one\-step lexical prediction tasks, making it easier to isolate how contextual demonstrations affect target\-token emergence across layers\.

Our analysis uses 500 structured samples in total\. For visualization, these samples are evenly divided into 10 groups, with 50 samples per group, denoted as Task 1–Task 10\. Layer\-wise metrics are computed for each sample and then averaged within each group to produce the heatmap rows\. Table[8](https://arxiv.org/html/2605.11416#A1.T8)provides 10 representative prompts to illustrate the dataset format\.

![Refer to caption](https://arxiv.org/html/2605.11416v1/x22.png)\(\(a\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x23.png)\(\(b\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x24.png)\(\(c\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x25.png)\(\(d\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x26.png)\(\(e\)\)
![Refer to caption](https://arxiv.org/html/2605.11416v1/x27.png)\(\(f\)\)

Figure 9:Robustness analysis under Top\-50 probability support on the AntSynNET dataset, where 500 samples are evenly divided into 10 groups of 50 samples for visualization\. Here, Top\-50 refers to the probability support used for JS\-based layer\-wise sensitivity computation, which is distinct from the top\-10 candidate\-token display in decoding\. The single target\-token probability used inRatio\\mathrm\{Ratio\}remains unchanged\. \(a\)–\(c\) show the variation ofRatio\\mathrm\{Ratio\}across layers for Qwen3\-0\.6B\-Base to Qwen3\-14B\-Base\. A deeper green color indicates a largerRatio\\mathrm\{Ratio\}, meaning that the target token probability changes more significantly at this layer\. \(d\)–\(f\) show the variation ofΔ​JS\\Delta\\mathrm\{JS\}across layers under the same Top\-50 setting\. A deeper color indicates a larger abrupt change in layer sensitivity, reflecting higher vulnerability and lower robustness\.

## Appendix CAdditional Experiments

### C\.1Robustness to Top\-KKProbability Support

To examine whether the Layer\-wise Sensitivity results are affected by the choice of probability support, we further repeat the analysis using a larger Top\-KKprobability support withK=50K=50, as shown in Figure[9](https://arxiv.org/html/2605.11416#A2.F9)\. For each layer, the JS divergence is computed over the aligned Top\-50 probability distribution rather than a smaller truncated support\. This Top\-50 setting refers to the probability support used for JS\-based sensitivity computation, which is distinct from the top\-10 candidate\-token selection used for target\-token analysis and visualization\.

The results show that enlarging the probability support to Top\-50 preserves the same qualitative trend observed in the main analysis: shallow layers still exhibit stronger sensitivity fluctuations, while deep layers remain comparatively stable and act as execution\-dominant zones\. Consequently, it further justifies our practical allocation strategy in Section[4\.2\.3](https://arxiv.org/html/2605.11416#S4.SS2.SSS3), where we adopt a symmetric midpoint split to train the shallow sensitive zones and freeze the deep execution zones\.

This consistency indicates that the LayerTracer diagnosis is robust to the choice of Top\-KKprobability support, further supporting the conclusion that shallow layers are more suitable for adaptation whereas deep layers should be preserved during continued pre\-training\.

### C\.2Boundary Selection

An exhaustive training sweep over all possible split boundaries can identify a model\-specific optimum, but this is not the goal of our study\. Our objective is to derive a general allocation rule from layer\-wise diagnostics to addressRQ2\. Figure[10](https://arxiv.org/html/2605.11416#A3.F10)shows the normalized TP and LS profiles of Qwen3\-710M\-Base\. LS is concentrated in shallow and middle layers, indicating that these layers are sensitive to perturbation and suitable for adaptation\. In contrast, TP becomes more prominent in later layers, suggesting that deep layers mainly serve as task\-evidence consolidation and execution zones\.

Based on this observation, we use the 50% midpoint as a simple and model\-agnostic split in the main continued pre\-training experiments\. This setting keeps a substantial sensitive shallow layer trainable while preserving the deep execution layer from disruptive updates\. It also provides a symmetric comparison with matched trainable\-parameter budgets\. Thus, the midpoint is not treated as an optimal boundary, but as a controlled test of the Train\-Shallow/Freeze\-Deep principle\.

To examine whether this allocation rule depends on the midpoint, we further conduct a split\-boundary diagnostic using LayerTracer\. For a model withNNlayers and a split ratiorr, we set the split layer to the nearest feasible layerb=round⁡\(r​N\)b=\\operatorname\{round\}\(rN\)\. Layers11tobbform the shallow layer, and layersb\+1b\+1toNNform the deep layers\. A desirable Train\-Shallow/Freeze\-Deep split should place high\-LS layers in the trainable layer and high\-TP layers in the frozen layer\. We therefore define the following boundary alignment score:

𝒮​\(b\)\\displaystyle\\mathcal\{S\}\(b\)=LS^¯1:b\+TP^¯b\+1:N\\displaystyle=\\overline\{\\widehat\{\\mathrm\{LS\}\}\}\_\{1:b\}\+\\overline\{\\widehat\{\\mathrm\{TP\}\}\}\_\{b\+1:N\}\(7\)−TP^¯1:b−LS^¯b\+1:N,\\displaystyle\\quad\-\\overline\{\\widehat\{\\mathrm\{TP\}\}\}\_\{1:b\}\-\\overline\{\\widehat\{\\mathrm\{LS\}\}\}\_\{b\+1:N\},whereTP^\\widehat\{\\mathrm\{TP\}\}andLS^\\widehat\{\\mathrm\{LS\}\}denote normalized TP and LS values\. A positive score indicates that the split is better aligned with the Train\-Shallow/Freeze\-Deep rule than with the reverse allocation\.

As shown in Table[9](https://arxiv.org/html/2605.11416#A3.T9),𝒮​\(b\)\\mathcal\{S\}\(b\)remains positive under 33%, 50%, and 66% shallow\-layer splits\. This result indicates that the allocation direction is not tied to the exact midpoint\. The performance deltas in Figure[10](https://arxiv.org/html/2605.11416#A3.F10)further support this interpretation: freezing the shallow sensitive layer and training only deep layers leads to performance drops, whereas training shallow layers while freezing deep layers yields consistent gains over full\-parameter training\.

Trainable Shallow RatioSplit Layer𝒮​\(b\)\\mathcal\{S\}\(b\)33%90\.00450%140\.13866%190\.392Table 9:Split\-boundary diagnostic based on normalized TP and LS\. Positive scores indicate that sensitive layers are assigned to the trainable layer and task\-execution layers are assigned to the frozen layer, supporting the Train\-Shallow/Freeze\-Deep rule across multiple split boundaries\.BenchmarkGenerate Std\.Logit Std\.C\-Eval0\.000\.00CMMLU0\.000\.00Table 10:Standard deviation over five repeated evaluations for all models used in Section[5](https://arxiv.org/html/2605.11416#S5)and Section[6](https://arxiv.org/html/2605.11416#S6)\. All values are zero under both generation\-based and logit\-based protocols, since generation uses deterministic decoding and logit evaluation directly computes scores from fixed candidate\-token probabilities\.![Refer to caption](https://arxiv.org/html/2605.11416v1/x28.png)Figure 10:Layer\-wise TP and LS profiles of Qwen3\-710M\-Base\. The light green region denotes the shallow layer used in the midpoint split, where LS is relatively stronger and layers are more sensitive to perturbation\. The light red region denotes the deep layers, where TP becomes more prominent and task evidence is mainly consolidated\. The dashed line marks the 50% split used in the main experiments\. The inset reports CEval and CMMLU changes relative to full\-parameter training: freezing shallow layers hurts performance, while freezing deep layers and training shallow layers brings consistent gains\.
### C\.3Standard Deviation under Deterministic Evaluation

As shown in Table[10](https://arxiv.org/html/2605.11416#A3.T10), all models used in Section[5](https://arxiv.org/html/2605.11416#S5)and Section[6](https://arxiv.org/html/2605.11416#S6)achieve zero standard deviation over five repeated evaluations on both C\-Eval and CMMLU under the two evaluation protocols\. For the generation\-based evaluation, we use greedy decoding withdo\_sample=False, so the generated answer is deterministic given the same model and input\. For the logit\-based evaluation, the score is computed directly from the fixed probabilities of the candidate answer tokens, without any decoding randomness\. Therefore, the repeated evaluations produce identical results\. This confirms the reproducibility of our reported scores under the fixed evaluation setting\.

Similar Articles

Attribution-Guided Continual Learning for Large Language Models

arXiv cs.LG

This paper proposes an attribution-guided continual fine-tuning framework for large language models that estimates task-specific parameter importance in Transformer layers and modulates gradients accordingly, mitigating catastrophic forgetting while maintaining performance on new tasks.

Uncovering the Latent Potential of Deep Intermediate Representations

arXiv cs.LG

This paper introduces LOES (Layer-wise Optimal Embedding Selection) and GeoReg (Geometric Regularization Loss), methods that select and fuse task-relevant intermediate layers from deep models to improve transfer learning performance, demonstrating consistent gains across architectures and modalities.