LLMs中的脚本选择:后期层承诺的证据
摘要
本文研究了LLM层中脚本知识的分布,揭示了脚本识别发生在早期,而对输出脚本的承诺则在后期层中出现,强调了模型深度对于多语言架构的重要性。
查看缓存全文
缓存时间: 2026/09/25 09:15
# Script Choice in LLMs: Evidence for Late-Layer Commitment
Source: [https://arxiv.org/html/2609.28784](https://arxiv.org/html/2609.28784)
Sandra Mitrović22footnotemark:2Itay Sabato44footnotemark:4Ljiljana Dolamić33footnotemark:3Fabio Rinaldi22footnotemark:2Affiliation:22footnotemark:2SUPSI, IDSIA, SwitzerlandAffiliation:33footnotemark:3armasuisse, Science & Technology, SwitzerlandAffiliation:44footnotemark:4Independent ResearcherAffiliation:\{david\.kletz, sandra\.mitrovic, fabio\.rinaldi\}@supsi\.chEmail:[ljiljana\.dolamic@armasuisse\.ch](mailto:)Email:[itaysabato@gmail\.com](mailto:)
###### Abstract
In this paper, we investigate how script knowledge is distributed across the layers of LLMs using two complementary interpretability methods: logistic regression probing and logit\-lens analysis\. Our probing experiments reveal a clear asymmetry: both the input script and the instructed output script are encoded in the earliest layers of the network, while, in contrast, commitment to the actual output script emerges only in the final layers, with the model’s intermediate representations defaulting to Latin throughout most of the layers\. This two\-stage process is confirmed by logit\-lens analyses, which show that script commitment consistently occurs at the very last layers of the LLMs\. Together with the weaker script\-following performance observed in smaller models, these results form a converging body of evidence linking script commitment to model depth, with broader implications for the design of sufficiently deep, inclusive multilingual architectures\.
## 1Introduction
Large Language Models \(LLMs\) can generate text across scripts, from Latin and Cyrillic to Arabic and Devanagari[Team et al\. \(2022\)](https://arxiv.org/html/2609.28784#bib.bib8);[Workshop et al\. \(2023\)](https://arxiv.org/html/2609.28784#bib.bib11);[Hernández\-Cano et al\. \(2026\)](https://arxiv.org/html/2609.28784#bib.bib9)\. This ability appears simple behaviorally, but internally, the model must first identify the script of the input, infer the script required for the output, and ultimately shift its probability mass toward tokens in the appropriate script\. Transformer models process different aspects of linguistic knowledge at different depths[Tenney et al\. \(2019\)](https://arxiv.org/html/2609.28784#bib.bib2);[Peters et al\. \(2018\)](https://arxiv.org/html/2609.28784#bib.bib6);[Belinkov \(2022\)](https://arxiv.org/html/2609.28784#bib.bib3), as a consequence, these sub\-tasks might not be executed simultaneously at a single point in the network, but instead distributed across layers\.
This distributional assumption motivates our central research question: where, across the layers of an LLM, does script knowledge emerge, and does the answer differ depending on whether we consider the recognition of the input script, the planning of the output script, or the commitment of the output distribution to a target writing system? Prior work supports the plausibility of such layer\-wise specialization beyond syntax and semantics\. At the behavioral level,[Liang et al\. \(2024\)](https://arxiv.org/html/2609.28784#bib.bib16)confirm that LLMs possess a degree of script planning capability, suggesting that output script selection is a deliberate, learnable operation rather than a mere surface artefact\. At a coarser granularity,[Pochinkov et al\. \(2024\)](https://arxiv.org/html/2609.28784#bib.bib14);[Pochinkov et al\. \(2025\)](https://arxiv.org/html/2609.28784#bib.bib15)show that commitment to paragraph\-level content may be a general property of autoregressive generation, which we expect to extend to script\-level commitment as well\. Evidence further suggests that adjacent layers may not always build cooperatively on one another[Patrawala et al\. \(2025\)](https://arxiv.org/html/2609.28784#bib.bib12), that simpler tasks require fewer layers[Fan et al\. \(2025\)](https://arxiv.org/html/2609.28784#bib.bib7), and that high\-resource language knowledge tends to be consolidated in deeper layers[Li et al\. \(2025\)](https://arxiv.org/html/2609.28784#bib.bib10)\.
Script processing, however, has not yet been studied from this perspective, even though this is a particularly interesting case: it intervenes at multiple stages of generation, from input recognition to output realization\. Most closely related to our work is RomanLens[Saji et al\. \(2025\)](https://arxiv.org/html/2609.28784#bib.bib18), which pairs the logit lens with activation patching to study the mechanism of romanization and language transfer, and reports a related finding: intermediate layers often represent non\-Roman\-script targets in Romanized form before the model transitions to native script\. Our work differs both methodologically, pairing the logit lens with linear probing rather than patching, and conceptually, distinguishing input script, instructed output script, and the script actually produced\.
In this paper, separation of these three scripts is central and we investigate how script knowledge is distributed across the layers of large language models\. We focus on the Qwen 2\.5 Instruct family \(0\.5B–32B\), which provides a controlled experimental setting\.
To trace script processing through the model’s layers, we employ two methods\. In Section[3](https://arxiv.org/html/2609.28784#S3), we train probes[Alain and Bengio \(2017\)](https://arxiv.org/html/2609.28784#bib.bib4)on the hidden state of the final prompt token at every layer to identify where script information becomes linearly accessible, distinguishing input script, instructed output script, script produced by the LLM\. Then, in Section[4](https://arxiv.org/html/2609.28784#S4), we apply a logit\-lens analysis[nostalgebraist \(2020\)](https://arxiv.org/html/2609.28784#bib.bib5)to track when the output distribution commits its output distribution to the target script\. Our experiments cover nine scripts111Data is available on our GitHub repo[https://github\.com/IDSIA\-NLP/LLMScriptCommitment](https://github.com/IDSIA-NLP/LLMScriptCommitment): Latin, Cyrillic, Arabic, Hebrew, Devanagari, Korean, Japanese, Chinese, and Armenian\. Finally, showing that commitment to non\-Latin scripts emerges only in the last layers of the LLM, in Section[5](https://arxiv.org/html/2609.28784#S5)we discuss this pattern as a potential, correlational explanation for the weaker performance of smaller models on non\-Latin scripts\.
## 2Experimental Setup
### 2\.1Models
We work with the Qwen 2\.5 Instruct family across six sizes: 0\.5B, 1\.5B, 3B, 7B, 14B, and 32B222The 14B and 32B checkpoints were run under 4\-bit quantization\. All checkpoints share the same pre\-training corpus[Qwen et al\. \(2025\)](https://arxiv.org/html/2609.28784#bib.bib1), so behavioral differences across scales reflect architecture rather than knowledge\. The models span 24 to 64 transformer layers, making this a good suite for studying the effect of depth\. Note however that depth increases non\-monotonically with parameter count: the 3B model \(36 layers\) is deeper than the 7B \(28 layers\)\. The full architecture table and quantization details are in Appendix[A](https://arxiv.org/html/2609.28784#A1)\.
### 2\.2Data and Prompts
Prompts are drawn from nine source languages: English, Russian, Arabic, Chinese, Japanese, Korean, Hindi, Armenian, and Hebrew\. Each language appears in two conditions: a canonical condition, where the model is instructed to produce output in the same script as the input, and a cross\-script condition, where it is asked to use a different script\. For all languages except English, the cross\-script target is Latin\. This yields 17 language\-to\-script mappings in total\.
Within each mapping, prompts are drawn from three families: sentences, words, and questions\. Each prompt is instantiated across four instruction variants that range from explicit \(“write the following using only Cyrillic characters”\) to concise \(“answer in Cyrillic”\)\. English and Russian are represented by 274 input items each \(80 sentences, 114 words, 80 questions\); the remaining seven languages by 70 items each \(20/30/20\)\. The full dataset comprises 8,304 prompts\. Each model sees the same set\. Detailed prompt templates are listed in Appendix[B](https://arxiv.org/html/2609.28784#A2)\.
### 2\.3Script Classification
The script of a token is determined by the Unicode block name of its first character\. We define ten script classes: one for each available input script and a ‘‘Neutral’’ class for tokens whose first character is any character that doesn’t have a script\. Tokens in the Neutral class are excluded from script\-specific analysis\. This classifier is applied at two points: to the prompt, to derive the input\-script label, and to the first generated token333We focus on the first generated token because it directly reflects the initial script\-selection decision before auto\-regressive feedback effectsof each model response, to derive the used\-output\-script label\.
### 2\.4Probing
For each \(model, prompt\) pair, we record the hidden state at the last prompt token at every layer \(the embedding output from layer 0 and each of theL−1L\-1transformer block outputs\) extracted during the prefill step of the generation forward pass, without a second inference\. We train one independent logistic regression probe per \(target, layer\) pair against three targets:Input script \(9 classes\): the source script of the prompt,Requested\-output script \(9 classes\): the script the prompt instructs the model to use, andUsed\-output script\(10 classes\): the script the model actually produced, including a Neutral/Other bucket for non\-classifiable outputs\.
Probes are a logistic regression withℓ2\\ell\_\{2\}regularization trained on per\-feature standardized activations\. High accuracy from such a probe indicates that the target is geometrically accessible \(hidden states of different classes are linearly separable in activation space,[Hewitt and Liang, 2019](https://arxiv.org/html/2609.28784#bib.bib17)\) in the representation at that layer, not merely encoded in some nonlinear fashion\. We use a 70/30 stratified train/test split, averaged across five independent random splits to reduce variance\. Our primary metric is balanced accuracy \(mean per\-class recall\), which is insensitive to the class imbalance introduced by the Neutral bucket in the used\-output target; we also report AUC and raw accuracy as secondary metrics\. Full hyperparameters are in Appendix[C](https://arxiv.org/html/2609.28784#A3)\.
### 2\.5Logit Lens
The logit lens[nostalgebraist \(2020\)](https://arxiv.org/html/2609.28784#bib.bib5)projects each intermediate representation to obtain a distribution over the full vocabulary at every layer\. We apply this to the same last\-prompt\-token hidden states used in section[2\.4](https://arxiv.org/html/2609.28784#S2.SS4), requiring no additional forward pass\. At each layerℓ\\ell, we extract the top\-1 token of the projected distribution and classify its script using the classifier of section[2\.3](https://arxiv.org/html/2609.28784#S2.SS3)\. We define the script commitment layer as the earliestℓ\\ellat which the top\-1 token belongs to the target script444Examples without such layer are considered as failures\. We report two quantities: \(i\) the fraction of examples where the model’s generated token belongs to the correct script, and \(ii\) the mean script commitment layer over successful examples, broken down by target script and model size\.
## 3Encoding of Script
Figure 1:Balanced accuracy of the requested\-output script probe across layers, for all six Qwen 2\.5 model sizes\.Figure 2:Balanced accuracy of the used\-output script probe across layers\.At layer 0, all targets are at chance for every model \(≃0\.11\{\\simeq\}\\,0\.11for the 9\-class targets and≃0\.09\{\\simeq\}\\,0\.09for used\-output\), confirming that token embeddings carry no linearly accessible script information\. This, in fact, is expected and serves as a sanity check, since all our prompts end with “Answer:” \(see Appendix[B](https://arxiv.org/html/2609.28784#A2)\) and, hence, have as the last prompt token always “:”\. However, at layer 1, input script already reaches≃0\.98\{\\simeq\}\\,0\.98balanced accuracy across all model sizes\.
Figure[1](https://arxiv.org/html/2609.28784#S3.F1)shows balanced accuracy for the requested\-output target\. All six models reach 1\.00 by layer 3–5, regardless of size or depth: the script the model was instructed to use is fully encoded within the first∼5\{\\sim\}5–2020% of the network\. Notably, balanced accuracy decreases slightly in the final layers of each model suggesting that the explicit representation of the requested script is gradually displaced as the network shifts from encoding the instruction to computing the output distribution\.
In Figure[2](https://arxiv.org/html/2609.28784#S3.F2), the used\-output target rises more slowly, peaks at 82–100% of model depth, and never saturates, with balanced accuracy ranging from 0\.77 \(0\.5B\) to 0\.86 \(3B\) at peak\. The maximum is set in part by each model’s actual script\-following rate \(0\.64–0\.83 strict, across sizes\)\. Together, the two figures establish the core dissociation: the model encodes what it should produce within the first few layers, but what it will actually produce only becomes readable in the final layers\. Section[4](https://arxiv.org/html/2609.28784#S4)examines what happens in between\.
## 4Late Commitment to Output Script
Figure 3:Average first layer of script commitment by target script and model size\. Error bars show standard deviation across prompts\. Hollow profiles show the maximum each model can reach \(number of layers\)\.Figure[3](https://arxiv.org/html/2609.28784#S4.F3)shows the mean script commitment layer per target script and model size, computed over successful generations only\. It shows an important contrast between Latin and all other scripts: the top\-1 predicted token at the generation position commits to Latin within the first 1–2 layers across all six models regardless of the source language of the prompt\. Non\-Latin script generation requires a late override: for Cyrillic, Arabic, Devanagari, Korean, and Hebrew, the commitment layer consistently falls above85%85\\%of model depth, often in the final three to five layers\.
The commitment layer grows with model depth in absolute terms: for Cyrillic, it rises from∼\{\\sim\}22 layers \(0\.5B, 24 total\) to∼\{\\sim\}60 layers \(32B, 64 total\), while remaining at roughly the same relative position across most models and scripts\. One notable exception is Cyrillic for the 3B model, which commits at layer∼\{\\sim\}14 out of 36\. Armenian, Chinese, and Japanese yield too low success rates across all model sizes and are absent from Figure[3](https://arxiv.org/html/2609.28784#S4.F3)\. Detailed results are in Appendix[F](https://arxiv.org/html/2609.28784#A6), and show that when their output is the correct one, they exhibit the same late\-commitment pattern as the other non\-Latin scripts\.
Table 1:Logit\-lens top\-1 prediction at the generated token position across the last 5 layers ofQwen2\.5\-3B\-Instructwhen asked to write “Capital” using the Cyrillic script\. The’ К’character is a Cyrillic character\.Table[1](https://arxiv.org/html/2609.28784#S4.T1)illustrates the underlying mechanism\. The heatmap shows the top\-1 logit\-lens token at each layer and each prompt position for a representative English\-to\-Cyrillic generation \(target:‘‘Капитал’’\)\. Through most of the network, the top\-1 token at the generation position is the semantically correct Latin token \(“capital”\)\. Only in the final five to eight layers does the prediction switch to the Cyrillic equivalent\. The Cyrillic token is not arbitrary: it is the transliteration of the Latin prediction held in the preceding layers\. This two\-stage process accounts for both the late commitment observed in Figure[3](https://arxiv.org/html/2609.28784#S4.F3)and suggests that the failure of models to complete it might be due to the insufficient model depth\. Additionally, this finding corroborates the one of RomanLens[Saji et al\. \(2025\)](https://arxiv.org/html/2609.28784#bib.bib18), that independently reports a similar late\-layer script transition on different models\.
## 5Models’ Depth Bottleneck Hypothesis and Smaller Models Failures
Our results show that script encoding and identification are distributed across the model’s layers\. Input script identity is encoded at the first transformer block \(layer 1,∼98%\{\\sim\}98\\%balanced accuracy\)\. The instructed output script is also captured and encoded within the first five layers \(∼5\{\\sim\}5–20%20\\%of depth\)\. The vocabulary distribution, however, does not commit to non\-Latin scripts until the final layers \(≥85%\{\\geq\}85\\%of depth\), and does so imperfectly in proportion to each model’s depth\. More specifically, Figure[4](https://arxiv.org/html/2609.28784#S5.F4)presents the total probability mass assigned to Latin tokens, tokens in the requested target script, and tokens in any other script, at each layer\. It shows that for most of each model’s depth, Latin mass matches or exceeds target\-script mass, with target\-script mass surpassing Latin mass only toward the final layers\. Qwen2\.5\-0\.5B is the one exception, where Latin mass stays at or above target\-script mass even at the very last layer, consistent with the observed model lack of depth\.
Figure 4:Probability mass assigned to tokens of Latin, target and any other script across layers\.#### Linear accessibility does not imply functional use
The gap between layers of target script encoding and layers of realization in the distribution reflects a distinction between encoding information and using it\. The model keeps the target script information accessible from layer 5 onward, yet its output distribution does not reflect this\. We interpret the intervening computation as the network working through the semantic content of the answer in its default register \(Latin\), with script conversion deferred to the final layers\. The slight decrease in requested\-output probe accuracy observed at the end of the network \(Section[3](https://arxiv.org/html/2609.28784#S3)\) is consistent with this view: as the final layers shift from maintaining the instruction to executing it, the explicit linear representation of the target script is partially displaced by the output distribution it is generating\.
The non\-monotonic relation between model size in parameters and script\-following performance shows that depth is the important parameter for this task\. Non\-Latin generation requires enough layers to execute the late\-stage script conversion that our logit\-lens analysis makes visible\. A model with too few layers defaults to a neutral or Latin output, regardless of how well it has encoded the instruction\.
We thus interpret the smaller models’ struggle with non\-Latin scripts as supportive of the hypothesis that limited model depth may constitute a bottleneck to completing the late\-stage script conversion process, rather than reflecting a failure of knowledge or representation\.
## 6Conclusion
In this paper, we have shown that script processing in LLMs is not a monolithic operation\. Rather, input identification, instruction encoding, and output script commitment emerge at different stages of the network\. The output vocabulary distribution is modified only after the model has formed an internal representation of what it intends to produce, and the gap between representation and script realization is resolved only in the final layers\.
The observed pattern offers a potential mechanistic explanation for the weaker script\-following performance of smaller models: rather than failing to encode the target script instruction, it suggests that these models lack layers to complete the late conversion process\.
As such, our findings have broader implications for the development of multilingual LLMs: improving support for underrepresented writing systems may require not only better representations, but also sufficiently deep architectures to translate those representations into reliable generation\.
Future work should also investigate how LLMs acquire the ability to transliterate their top Latin prediction into the target script, and complement our script\-based analysis with a breakdown by language family\.
## Limitations
We acknowledge the following limitations\. First, all experiments are conducted on a single model family, Qwen 2\.5 Instruct\. It remains an open question whether the observed patterns generalize to other architectures\. Second, the two largest models \(14B and 32B\) were run under 4\-bit quantization due to hardware constraints\. Third, Armenian, Chinese, and Japanese exhibited too low script\-following success rates across all model sizes to be included in the main logit\-lens analysis\.
Finally, our probing and logit\-lens analyzes both rely on a single token position: the last token of the prompt, which serves as the prediction site for the first generated token\.
## Acknowledgments
The work described in this paper has been partially funded by the “Language Preference and Hallucination Investigation in Multilingual RAG \(LPHI\-mRAG\)” project, funded by armasuisse S&T, Switzerland\.
## References
- Alain and Bengio \(2017\)G\. Alain and Y\. BengioUnderstanding intermediate layers using linear classifier probes\.External Links:[Link](https://openreview.net/forum?id=ryF7rTqgl)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p5.1)\.
- Belinkov \(2022\)Y\. BelinkovProbing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.External Links:[Link](https://aclanthology.org/2022.cl-1.7/),[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00422)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p1.1)\.
- Fanet al\.\(2025\)S\. Fan, X\. Jiang, X\. Li, X\. Meng, P\. Han, S\. Shang, A\. Sun, and Y\. WangNot all layers of llms are necessary during inference\.InProceedings of the Thirty\-Fourth International Joint Conference on Artificial Intelligence,IJCAI ’25\.External Links:ISBN 978\-1\-956792\-06\-5,[Link](https://doi.org/10.24963/ijcai.2025/566),[Document](https://dx.doi.org/10.24963/ijcai.2025/566)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p2.1)\.
- Hernández\-Canoet al\.\(2026\)A\. Hernández\-Cano, A\. Hägele, A\. H\. Huang, A\. Romanou, A\. Solergibert, B\. Pásztor, B\. Messmer, D\. Garbaya, E\. F\. Ďurech, I\. Hakimi, J\. G\. Giraldo, M\. Ismayilzada, N\. Foroutan, S\. Moalla, T\. Chen, V\. Sabolčec, Y\. Xu, M\. Aerni, B\. AlKhamissi, I\. A\. Marinas, M\. H\. Amani, M\. Ansaripour, I\. Badanin, H\. Benoit, E\. Boros, N\. J\. Browning, F\. Bösch, M\. Böther, N\. Canova, C\. Challier, C\. Charmillot, J\. Coles, J\. M\. Deriu, A\. Devos, L\. Drescher, D\. Dzenhaliou, M\. Ehrmann, D\. Fan, S\. Fan, S\. Gao, M\. Gila, M\. Grandury, D\. Hashemi, A\. M\. Hoyle, J\. Jiang, M\. Klein, A\. Kucharavy, A\. Kucherenko, F\. Lübeck, R\. Machacek, T\. I\. Manitaras, A\. Marfurt, K\. Matoba, S\. Matrenok, H\. Mendonça, F\. R\. Mohamed, S\. Montariol, L\. Mouchel, S\. Najem\-Meyer, J\. Ni, G\. Oliva, M\. Pagliardini, E\. Palme, A\. Panferov, L\. Paoletti, M\. Passerini, I\. Pavlov, A\. Poiroux, K\. Ponkshe, N\. Ranchin, J\. Rando, M\. Sauser, J\. Saydaliev, M\. Sayfiddinov, M\. Schneider, S\. Schuppli, M\. Scialanga, A\. Semenov, K\. Shridhar, R\. Singhal, A\. Sotnikova, A\. Sternfeld, A\. K\. Tarun, P\. Teiletche, J\. Vamvas, X\. Yao, H\. Zhao, A\. Ilic, A\. Klimovic, A\. Krause, C\. Gulcehre, D\. Rosenthal, E\. Ash, F\. Tramèr, J\. VandeVondele, L\. Veraldi, M\. Rajman, T\. C\. Schulthess, T\. Hoefler, A\. Bosselut, M\. Jaggi, and I\. SchlagApertus: democratizing open and compliant LLMs for global language environments\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 46877–46955\.External Links:[Link](https://aclanthology.org/2026.acl-long.2172/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.2172),ISBN 979\-8\-89176\-390\-6Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p1.1)\.
- Hewitt and Liang \(2019\)J\. Hewitt and P\. LiangDesigning and interpreting probes with control tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 2733–2743\.External Links:[Link](https://aclanthology.org/D19-1275/),[Document](https://dx.doi.org/10.18653/v1/D19-1275)Cited by:[§2\.4](https://arxiv.org/html/2609.28784#S2.SS4.p2.1)\.
- Liet al\.\(2025\)D\. Li, H\. Zhao, Q\. Zeng, and M\. DuExploring multilingual probing in large language models: a cross\-language analysis\.InProceedings of the 1st Joint Workshop on Large Language Models and Structure Modeling \(XLLM 2025\),H\. Fei, K\. Tu, Y\. Zhang, X\. Hu, W\. Han, Z\. Jia, Z\. Zheng, Y\. Cao, M\. Zhang, W\. Lu, N\. Siddharth, L\. Øvrelid, N\. Xue, and Y\. Zhang \(Eds\.\),Vienna, Austria,pp\. 61–70\.External Links:[Link](https://aclanthology.org/2025.xllm-1.7/),[Document](https://dx.doi.org/10.18653/v1/2025.xllm-1.7),ISBN 979\-8\-89176\-286\-2Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p2.1)\.
- Lianget al\.\(2024\)S\. Liang, B\. Zhang, J\. Zhao, and K\. LiuABSEval: an agent\-based framework for script evaluation\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 12418–12434\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.691/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.691)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p2.1)\.
- nostalgebraist \(2020\)nostalgebraistInterpreting GPT: the logit lens — LessWrong — lesswrong\.com\.Note:[https://www\.lesswrong\.com/posts/AcKRB8wDpdaN6v6ru/interpreting\-gpt\-the\-logit\-lens](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)\[Accessed 24\-05\-2026\]Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p5.1),[§2\.5](https://arxiv.org/html/2609.28784#S2.SS5.p1.1)\.
- Patrawalaet al\.\(2025\)A\. Patrawala, J\. Feng, E\. Jones, and J\. SteinhardtLLM layers immediately correct each other\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38, Main Conference,pp\. 139071–139113\.External Links:[Document](https://dx.doi.org/10.52202/085713-4643),[Link](https://proceedings.neurips.cc/paper_files/paper/2025/file/cb574f092013f886fe0b93d295aa402e-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p2.1)\.
- Peterset al\.\(2018\)M\. E\. Peters, M\. Neumann, L\. Zettlemoyer, and W\. YihDissecting contextual word embeddings: architecture and representation\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,E\. Riloff, D\. Chiang, J\. Hockenmaier, and J\. Tsujii \(Eds\.\),Brussels, Belgium,pp\. 1499–1509\.External Links:[Link](https://aclanthology.org/D18-1179/),[Document](https://dx.doi.org/10.18653/v1/D18-1179)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p1.1)\.
- Pochinkovet al\.\(2024\)N\. Pochinkov, A\. Benoit, L\. Agarwal, Z\. A\. Majid, and L\. Ter\-MinassianExtracting paragraphs from llm token activations\.External Links:2409\.06328,[Link](https://arxiv.org/abs/2409.06328)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p2.1)\.
- Pochinkovet al\.\(2025\)N\. Pochinkov, Y\. Volkova, A\. Vasileva, and S\. V\. R\. ChereddyParaScopes: what do language models activations encode about future text?\.External Links:2511\.00180,[Link](https://arxiv.org/abs/2511.00180)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p2.1)\.
- Qwenet al\.\(2025\)Qwen, :, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. QiuQwen2\.5 technical report\.External Links:2412\.15115,[Link](https://arxiv.org/abs/2412.15115)Cited by:[§2\.1](https://arxiv.org/html/2609.28784#S2.SS1.p1.1)\.
- Sajiet al\.\(2025\)A\. Saji, J\. A\. Husain, T\. Jayakumar, R\. Dabre, A\. Kunchukuttan, and R\. PuduppullyRomanLens: the role of latent Romanization in multilinguality in LLMs\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 26410–26429\.External Links:[Link](https://aclanthology.org/2025.findings-acl.1354/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.1354),ISBN 979\-8\-89176\-256\-5Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p3.1),[§4](https://arxiv.org/html/2609.28784#S4.p3.1)\.
- Tamoet al\.\(2026\)J\. B\. Tamo, D\. Carlander\-Reuterfelt, J\. Rubin, O\. Poliannikov, D\. Hong, and M\. WangLinguaMap: which layers of LLMs speak your language and how to tune them?\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=r00UxTl8El)Cited by:[Appendix H](https://arxiv.org/html/2609.28784#A8.p2.1)\.
- Teamet al\.\(2022\)N\. Team, M\. R\. Costa\-jussà, J\. Cross, O\. Çelebi, M\. Elbayad, K\. Heafield, K\. Heffernan, E\. Kalbassi, J\. Lam, D\. Licht, J\. Maillard, A\. Sun, S\. Wang, G\. Wenzek, A\. Youngblood, B\. Akula, L\. Barrault, G\. M\. Gonzalez, P\. Hansanti, J\. Hoffman, S\. Jarrett, K\. R\. Sadagopan, D\. Rowe, S\. Spruit, C\. Tran, P\. Andrews, N\. F\. Ayan, S\. Bhosale, S\. Edunov, A\. Fan, C\. Gao, V\. Goswami, F\. Guzmán, P\. Koehn, A\. Mourachko, C\. Ropers, S\. Saleem, H\. Schwenk, and J\. WangNo language left behind: scaling human\-centered machine translation\.External Links:2207\.04672,[Link](https://arxiv.org/abs/2207.04672)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p1.1)\.
- Tenneyet al\.\(2019\)I\. Tenney, D\. Das, and E\. PavlickBERT rediscovers the classical NLP pipeline\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 4593–4601\.External Links:[Link](https://aclanthology.org/P19-1452/),[Document](https://dx.doi.org/10.18653/v1/P19-1452)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p1.1)\.
- Workshopet al\.\(2023\)B\. Workshop, :, T\. L\. Scao, A\. Fan, C\. Akiki, E\. Pavlick, S\. Ilić, D\. Hesslow, R\. Castagné, A\. S\. Luccioni, F\. Yvon, M\. Gallé, J\. Tow, A\. M\. Rush, S\. Biderman, A\. Webson, P\. S\. Ammanamanchi, T\. Wang, B\. Sagot, N\. Muennighoff, A\. V\. del Moral, O\. Ruwase, R\. Bawden, S\. Bekman, A\. McMillan\-Major, I\. Beltagy, H\. Nguyen, L\. Saulnier, S\. Tan, P\. O\. Suarez, V\. Sanh, H\. Laurençon, Y\. Jernite, J\. Launay, M\. Mitchell, C\. Raffel, A\. Gokaslan, A\. Simhi, A\. Soroa, A\. F\. Aji, A\. Alfassy, A\. Rogers, A\. K\. Nitzav, C\. Xu, C\. Mou, C\. Emezue, C\. Klamm, C\. Leong, D\. van Strien, D\. I\. Adelani, D\. Radev, E\. G\. Ponferrada, E\. Levkovizh, E\. Kim, E\. B\. Natan, F\. D\. Toni, G\. Dupont, G\. Kruszewski, G\. Pistilli, H\. Elsahar, H\. Benyamina, H\. Tran, I\. Yu, I\. Abdulmumin, I\. Johnson, I\. Gonzalez\-Dios, J\. de la Rosa, J\. Chim, J\. Dodge, J\. Zhu, J\. Chang, J\. Frohberg, J\. Tobing, J\. Bhattacharjee, K\. Almubarak, K\. Chen, K\. Lo, L\. V\. Werra, L\. Weber, L\. Phan, L\. B\. allal, L\. Tanguy, M\. Dey, M\. R\. Muñoz, M\. Masoud, M\. Grandury, M\. Šaško, M\. Huang, M\. Coavoux, M\. Singh, M\. T\. Jiang, M\. C\. Vu, M\. A\. Jauhar, M\. Ghaleb, N\. Subramani, N\. Kassner, N\. Khamis, O\. Nguyen, O\. Espejel, O\. de Gibert, P\. Villegas, P\. Henderson, P\. Colombo, P\. Amuok, Q\. Lhoest, R\. Harliman, R\. Bommasani, R\. L\. López, R\. Ribeiro, S\. Osei, S\. Pyysalo, S\. Nagel, S\. Bose, S\. H\. Muhammad, S\. Sharma, S\. Longpre, S\. Nikpoor, S\. Silberberg, S\. Pai, S\. Zink, T\. T\. Torrent, T\. Schick, T\. Thrush, V\. Danchev, V\. Nikoulina, V\. Laippala, V\. Lepercq, V\. Prabhu, Z\. Alyafeai, Z\. Talat, A\. Raja, B\. Heinzerling, C\. Si, D\. E\. Taşar, E\. Salesky, S\. J\. Mielke, W\. Y\. Lee, A\. Sharma, A\. Santilli, A\. Chaffin, A\. Stiegler, D\. Datta, E\. Szczechla, G\. Chhablani, H\. Wang, H\. Pandey, H\. Strobelt, J\. A\. Fries, J\. Rozen, L\. Gao, L\. Sutawika, M\. S\. Bari, M\. S\. Al\-shaibani, M\. Manica, N\. Nayak, R\. Teehan, S\. Albanie, S\. Shen, S\. Ben\-David, S\. H\. Bach, T\. Kim, T\. Bers, T\. Fevry, T\. Neeraj, U\. Thakker, V\. Raunak, X\. Tang, Z\. Yong, Z\. Sun, S\. Brody, Y\. Uri, H\. Tojarieh, A\. Roberts, H\. W\. Chung, J\. Tae, J\. Phang, O\. Press, C\. Li, D\. Narayanan, H\. Bourfoune, J\. Casper, J\. Rasley, M\. Ryabinin, M\. Mishra, M\. Zhang, M\. Shoeybi, M\. Peyrounette, N\. Patry, N\. Tazi, O\. Sanseviero, P\. von Platen, P\. Cornette, P\. F\. Lavallée, R\. Lacroix, S\. Rajbhandari, S\. Gandhi, S\. Smith, S\. Requena, S\. Patil, T\. Dettmers, A\. Baruwa, A\. Singh, A\. Cheveleva, A\. Ligozat, A\. Subramonian, A\. Névéol, C\. Lovering, D\. Garrette, D\. Tunuguntla, E\. Reiter, E\. Taktasheva, E\. Voloshina, E\. Bogdanov, G\. I\. Winata, H\. Schoelkopf, J\. Kalo, J\. Novikova, J\. Z\. Forde, J\. Clive, J\. Kasai, K\. Kawamura, L\. Hazan, M\. Carpuat, M\. Clinciu, N\. Kim, N\. Cheng, O\. Serikov, O\. Antverg, O\. van der Wal, R\. Zhang, R\. Zhang, S\. Gehrmann, S\. Mirkin, S\. Pais, T\. Shavrina, T\. Scialom, T\. Yun, T\. Limisiewicz, V\. Rieser, V\. Protasov, V\. Mikhailov, Y\. Pruksachatkun, Y\. Belinkov, Z\. Bamberger, Z\. Kasner, A\. Rueda, A\. Pestana, A\. Feizpour, A\. Khan, A\. Faranak, A\. Santos, A\. Hevia, A\. Unldreaj, A\. Aghagol, A\. Abdollahi, A\. Tammour, A\. HajiHosseini, B\. Behroozi, B\. Ajibade, B\. Saxena, C\. M\. Ferrandis, D\. McDuff, D\. Contractor, D\. Lansky, D\. David, D\. Kiela, D\. A\. Nguyen, E\. Tan, E\. Baylor, E\. Ozoani, F\. Mirza, F\. Ononiwu, H\. Rezanejad, H\. Jones, I\. Bhattacharya, I\. Solaiman, I\. Sedenko, I\. Nejadgholi, J\. Passmore, J\. Seltzer, J\. B\. Sanz, L\. Dutra, M\. Samagaio, M\. Elbadri, M\. Mieskes, M\. Gerchick, M\. Akinlolu, M\. McKenna, M\. Qiu, M\. Ghauri, M\. Burynok, N\. Abrar, N\. Rajani, N\. Elkott, N\. Fahmy, O\. Samuel, R\. An, R\. Kromann, R\. Hao, S\. Alizadeh, S\. Shubber, S\. Wang, S\. Roy, S\. Viguier, T\. Le, T\. Oyebade, T\. Le, Y\. Yang, Z\. Nguyen, A\. R\. Kashyap, A\. Palasciano, A\. Callahan, A\. Shukla, A\. Miranda\-Escalada, A\. Singh, B\. Beilharz, B\. Wang, C\. Brito, C\. Zhou, C\. Jain, C\. Xu, C\. Fourrier, D\. L\. Periñán, D\. Molano, D\. Yu, E\. Manjavacas, F\. Barth, F\. Fuhrimann, G\. Altay, G\. Bayrak, G\. Burns, H\. U\. Vrabec, I\. Bello, I\. Dash, J\. Kang, J\. Giorgi, J\. Golde, J\. D\. Posada, K\. R\. Sivaraman, L\. Bulchandani, L\. Liu, L\. Shinzato, M\. H\. de Bykhovetz, M\. Takeuchi, M\. Pàmies, M\. A\. Castillo, M\. Nezhurina, M\. Sänger, M\. Samwald, M\. Cullan, M\. Weinberg, M\. D\. Wolf, M\. Mihaljcic, M\. Liu, M\. Freidank, M\. Kang, N\. Seelam, N\. Dahlberg, N\. M\. Broad, N\. Muellner, P\. Fung, P\. Haller, R\. Chandrasekhar, R\. Eisenberg, R\. Martin, R\. Canalli, R\. Su, R\. Su, S\. Cahyawijaya, S\. Garda, S\. S\. Deshmukh, S\. Mishra, S\. Kiblawi, S\. Ott, S\. Sang\-aroonsiri, S\. Kumar, S\. Schweter, S\. Bharati, T\. Laud, T\. Gigant, T\. Kainuma, W\. Kusa, Y\. Labrak, Y\. S\. Bajaj, Y\. Venkatraman, Y\. Xu, Y\. Xu, Y\. Xu, Z\. Tan, Z\. Xie, Z\. Ye, M\. Bras, Y\. Belkada, and T\. WolfBLOOM: a 176b\-parameter open\-access multilingual language model\.External Links:2211\.05100,[Link](https://arxiv.org/abs/2211.05100)Cited by:[§1](https://arxiv.org/html/2609.28784#S1.p1.1)\.
## Appendix AModels
See details in Table[2](https://arxiv.org/html/2609.28784#A1.T2)\. Large models \(Qwen2\.5\-14B and Qwen2\.5\-32B\) were loaded with 4\-bit\) quantization via BitsAndBytes, using double quantization and bfloat16 compute\.
Table 2:Qwen2\.5 models used\.
## Appendix BPrompts
All prompts end by asking the answer with the word “Answer:” so that the model’s first generated token is the start of the answer\. Placeholders are shown in angle brackets: ⟨target\_script⟩ is the name of the requested script \(e\.g\.Latin,Cyrillic\), and ⟨input⟩ is the source sentence, word, or question\.
### Sentences
### Words
### Questions
## Appendix CProbes
ℓ2\\ell\_\{2\}regularisation \(C=1C\\\!=\\\!1\), and L\-BFGS optimisation, on a single 70/30 train/test split of 1,000 prompts after per\-feature standardisation\.
## Appendix DScript\-Following Behavior
We check each model’s behavioral compliance with the requested script\. The strict follow rate is the fraction of prompts for which the used\-output script matches the requested script; the lax follow rate additionally counts “Other” and “Neutral” outputs as successes \(Table[3](https://arxiv.org/html/2609.28784#A4.T3)\)\.
Table 3:Script\-following rates\.
## Appendix EInput language
Input language is the easiest target by a wide margin: it is fully determined by the surface tokens of the prompt, requiring no cross\-token integration\. At layer 0 \(embedding output\) all probes sit at chance \(balanced accuracy 0\.47–0\.54\); a single transformer block suffices to drive every metric to≈\\approx1\.0 in every model\.
Figure 5:Balanced accuracy of the input\-language probe across layers, for all six Qwen models\.
## Appendix FDetailed Results
Figure[6](https://arxiv.org/html/2609.28784#A6.F6)illustrates per\-script breakdown across all six Qwen models\.
Figure 6:Per\-script breakdown across all six Qwen models\. Left: correct\-script generation rate by target script\. Center: output distribution by target script, broken down into correct, wrong, and neutral outputs\. Right: average first commitment layer by target script for correct outputs only\.Figure 7:Logit margin between the top\-1 token and the highest\-ranked token from different script, across layers\.To investigate the distribution of probability mass beyond the top\-1 token at each layer, we calculate the logit margin between the top\-1 token and the highest\-ranked token belonging to a different script \(Figure[7](https://arxiv.org/html/2609.28784#A6.F7)\)\. Across all six models, the logit margin stays modest \(from 0\.5 to 5\), before increasing in the final layers, with the two models 14B and 32B showing the largest increase in margin at the final layer\. This confirms top\-1 at the commitment layer is not a weak tie\-break\. At the last layer, prompts whose output matches the requested script show a larger margin than those that do not\.
## Appendix GWrong Scripts Generation
For each model, we add in Figure[8](https://arxiv.org/html/2609.28784#A7.F8)a breakdown of the prompts whose used\-output script does not match the requested script, by the script that was actually produced\.
Figure 8:Breakdown of wrong\-script generations by the script actually produced, per model\.
## Appendix HComputational Details
All experiments were run on an NVIDIA L40S GPU \(46 GB VRAM\)\. All models were run with greedy decoding \(do\_sample=False\) and a maximum budget of 256 new tokens\. Generation was stopped at the first token whose decoded text contained a classifiable Unicode character\.
Probes were trained on GPU using a batched logistic regression implemented in PyTorch, fitting all layers of a given target simultaneously in a single Adam optimization loop \(learning rate 0\.1, up to 1000 epochs\)\.[15](https://arxiv.org/html/2609.28784#bib.bib13)相似文章
当代理过早承诺:诊断LLM代理的过早承诺
本文引入表征承诺,这是一种跨运行隐藏状态收敛,用于诊断LLM代理何时过早锁定了轨迹。研究表明,承诺预测轨迹一致性而非正确性,并提出了监控方法,用于检测代理何时自信地稳定下来,而不是假设一致性等于可信度。
先承诺后推理:开放权重LLM中回答预承诺的行为复现与初步激活级证据
本文使用一个极简的洗车问题,在开放权重大型语言模型(Qwen3-8B)中复现了回答预承诺现象,并提供了初步激活级证据,表明承诺在答案文本生成之前就已编码在隐藏状态中。
推理大语言模型中的隐藏语言一致性现象
本文研究多语言推理模型,揭示输出中的语言一致性可能随任务难度增加而退化或崩溃,尤其是对于资源较少的语言。文章认为,评估多语言能力需要同时考虑准确性、语言一致性和任务难度。
论大语言模型适应性的局限:模型内化先验对标注任务性能的影响
本文研究了LLM的内化先验如何影响零样本标注性能,发现近三分之二的错误抵抗基于提示的修正,并引入了定义特定熟悉度(DSF)作为比记忆化指标更好的预测因子。
Answer First, Reason Later: Commitment Order in Diffusion LLMs
This paper investigates why diffusion LLMs fail at reasoning tasks: unconstrained token commitment freezes answers early and collapses to answer-only outputs. The authors identify commitment order as the root cause and propose a training-free, frontier-gated decoding intervention that recovers performance while preserving parallel decoding.