Late Transformer Layers Recode Syntax Canonically: Evidence from Greek Scrambling and Cross-Layer Generalisation
Summary
This paper investigates how syntactic information is processed in later transformer layers using Greek language data, revealing a representational shift towards canonical word order through cross-layer generalization analysis.
View Cached Full Text
Cached at: 09/02/26, 05:51 AM
# Late Transformer Layers Recode Syntax Canonically: Evidence from Greek Scrambling and Cross-Layer Generalisation
Source: [https://arxiv.org/html/2609.00416](https://arxiv.org/html/2609.00416)
Christos Nikolaos ZacharopoulosRevekka KyriakoglouIndependent ResearcherUniversité Paris 8 Vincennes–Saint\-DenisParis, Francechristonik@gmail\.comrevekka\.kyriakoglou@univ\-paris8\.frChara TsoukalaThéo DesbordesInstitute for LanguageDept\. of Basic Neurosciencesand Speech ProcessingFaculty of Medicine, University of GenevaAthena Research Center, Athens, GreeceGeneva, Switzerlandchara\.tsoukala@athenarc\.grtheo\.desbordes@unige\.ch
###### Abstract
Probing studies have established that syntactic information is decodable in early and middle transformer layers, but what happens to that information in later layers remains poorly understood\. We apply a cross\-layer generalisation analysis to three Greek\-tuned large language models evaluated on tightly controlled minimal pairs: object\-relative constructions in Modern Greek, where canonical \(Subject\-Verb\-Object; SVO\) and non\-canonical \(Verb\-Subject\-Object; VSO\) orders differ only in within\-clause word order, while preserving propositional meaning\. When a probe trained on late layers \(20–31\) is tested on each early layer individually, it produces below\-chance transfer \(cluster\-corrected,p<0\.01p<0\.01\), classifying 99\.3% of non\-canonical sentences as canonical\. Probe coefficients reverse sign around layer 22, indicating a directional recoding toward the canonical form rather than simple information loss\. These findings characterise a representational format change in late transformer layers that goes beyond the well\-established decline in syntactic decodability, and they generate a directly testable prediction for human EEG and MEG decoding studies using the same stimuli\. Code and stimuli are publicly available on[OSF](https://osf.io/5d3w8/overview?view_only=d42af279745543808cc377b9f96cb1af)
## 1Introduction
Consider two sentences in Modern Greek that carry the same meaning:
\(1\)\\acctonos\\acctonosς\\acctonos\.\(canonical SVO\) def\.f\.nomAnna\.nomsee\.pst\.3sgdef\.m\.accdog\.acc ‘Anna saw the dog\.’
\(2\)\\acctonos\\acctonosς\\acctonos\.\(non\-canonical VSO\) see\.pst\.3sgdef\.f\.nomAnna\.nomdef\.m\.accdog\.acc ‘Anna saw the dog\.’ lit\. ‘Saw Anna the dog’
Because case endings mark grammatical roles, the verb may precede the subject without altering the proposition\([Georgiafentis et al\., 2025](https://arxiv.org/html/2609.00416#bib.bib1);[Katsika and Allen, 2013](https://arxiv.org/html/2609.00416#bib.bib2)\)\. A competent reader resolves both orders to the same meaning\. The question we address is not whether this resolution occurs, but what representational changes enable it inside a transformer language model\.
Probing studies since BERT have established that syntactic information peaks in middle transformer layers and declines at the output\.[Tenney et al\. \(2019a\)](https://arxiv.org/html/2609.00416#bib.bib8)showed that classical NLP tasks are resolved in a layer\-wise pipeline;[Tenney et al\. \(2019b\)](https://arxiv.org/html/2609.00416#bib.bib9)and[Hewitt and Manning \(2019\)](https://arxiv.org/html/2609.00416#bib.bib5)confirmed that structural properties are most decodable in middle layers;[Coenen et al\. \(2019\)](https://arxiv.org/html/2609.00416#bib.bib10)demonstrated geometric clustering of syntactic relations in BERT’s representation space\([Belinkov et al\., 2020](https://arxiv.org/html/2609.00416#bib.bib11), see also\)\. These studies characterisewheresyntactic information resides, but notwhat happens to itafter it peaks\. Does the code simply weaken in later layers, or does the representation change in a specific direction? Standard per\-layer probing cannot distinguish these possibilities, because it treats each layer independently\.
We address this gap with a cross\-layer generalisation analysis borrowed from temporal decoding in cognitive neuroscience\([King and Dehaene, 2014](https://arxiv.org/html/2609.00416#bib.bib12);[Desbordes et al\., 2026](https://arxiv.org/html/2609.00416#bib.bib7)\)\. A linear probe trained on representations at one layer is tested on those at another; transfer failure between layers indicates a change in representational format, and the direction of failure characterises the nature of that change\. This design has, to our knowledge, not previously been applied in the probing literature on word\-order representation in transformer LLMs\.
Characterising these representational dynamics also matters for brain–model alignment research, where layer depth correlates with cortical processing stages\([Caucheteux et al\., 2021](https://arxiv.org/html/2609.00416#bib.bib13);[Schrimpf et al\., 2021](https://arxiv.org/html/2609.00416#bib.bib6);[Goldstein et al\., 2025](https://arxiv.org/html/2609.00416#bib.bib4)\)but detailed comparisons reveal divergent mechanisms in specific constructions\([Zacharopoulos et al\., 2023](https://arxiv.org/html/2609.00416#bib.bib3);[Zacharopoulos et al\., 2026](https://arxiv.org/html/2609.00416#bib.bib15)\)\.
Greek object\-relative constructions provide a uniquely controlled test case: case morphology marks grammatical roles independently of word order, so the same words can appear in SVO or VSO order within a relative clause while preserving propositional content, as in examples \(1\)–\(2\)\. Syntactically rigid languages do not permit this degree of experimental control\.
We apply this analysis across all layers of three Greek\-tuned transformer models\. A probe trained on layers 20–31 classifies 99\.3% of non\-canonical sentences as canonical when applied to early layers, producing below\-chance transfer across a contiguous cluster\. Probe coefficients reverse sign around layer 22, accounting for the directional inversion\. This asymmetry generates a directly testable prediction for human neural responses to the same stimuli\.
## 2Materials & Methods
### 2\.1Models
The primary model was[Llama\-Krikri\-8B\-Base](https://huggingface.co/ilsp/Llama-Krikri-8B-Base)\([Roussis et al\., 2025](https://arxiv.org/html/2609.00416#bib.bib14)\), a Greek–English LLM based on the Llama\-3\.1\-8B architecture with state\-of\-the\-art performance on Greek benchmarks\. Two comparison variants were evaluated:[Llama\-Krikri\-8B\-Instruct](https://huggingface.co/ilsp/Llama-Krikri-8B-Instruct)\(instruction\-tuned\) and[Plutus\-8B\-Instruct](https://huggingface.co/TheFinAI/plutus-8B-instruct)\(Low\-Rank Adaptation on a Greek financial corpus\)\. Results for these variants are reported in Appendix[A](https://arxiv.org/html/2609.00416#A1)\.
### 2\.2Experimental Design and Stimuli
Stimuli were object\-relative sentences in Modern Greek\. The SVO/VSO distinction refers to word orderinside the embedded relative clause, not the matrix sentence; the matrix verb is always sentence\-final in both conditions\. The SVO/VSO contrast from examples \(1\)–\(2\) applies inside the embedded relative clause: \[Det N2Vtrans\] vs\. \[VtransDet N2\]; the matrix subject and sentence\-final verb remain fixed\. Complete glossed minimal\-pair stimuli appear in Appendix[B\.4](https://arxiv.org/html/2609.00416#A2.SS4)\.
Stimuli were systematically generated from a predefined lexicon of Greek nouns, determiners, and verbs \(25 masculine and 25 feminine human noun stems; see Appendix[B](https://arxiv.org/html/2609.00416#A2)\)\. Template generation was required because naturally occurring SVO/VSO pairs inevitably differ in lexical content, whereas templates allow all lexical material to be held constant within each minimal pair\.
Stem pairs for N1and N2were always distinct\. The design yielded24=322^\{4\}=32unique fully counterbalanced linguistic conditions\. The main experiment used 128 sentences; a sensitivity analysis on an expanded 1024\-sentence set is reported in Appendix[A](https://arxiv.org/html/2609.00416#A1)\. All sentences were processed with the model’s native subword tokenizer \(see Appendix[B\.1](https://arxiv.org/html/2609.00416#A2.SS1)\)\.
### 2\.3Probing and Layer\-wise Analysis
Word order information was framed as a binary classification problem \(canonical SVO vs\. non\-canonical VSO\) over hidden\-state summaries\. For each sentence and layer, the hidden\-state matrix \(tokens×\\timeshidden dimensions\) for the post\-clause region was extracted\. Four distributional statistics were computed over the token dimension: mean, variance, skewness, and kurtosis\. These four scalars constitute the per\-layer feature vector for a given sentence\. This summary was chosen over raw hidden states because it provides a compact characterisation of the sequence\-level representational geometry and enables the coefficient\-sign analysis in §[3](https://arxiv.org/html/2609.00416#S3)\. Features were normalised with a robust estimator \(median and inter\-quartile range\)\. The classifier was L2\-regularised logistic regression withC=1\.0C=1\.0\(scikit\-learn default\); no hyperparameter search was conducted, as the choice was fixed a priori\.
The clause boundary is the complementizerπου\(onset of the relative clause\); the pre\-clause region contains all tokens up to and including it, and the post\-clause region, all tokens following it\.
Layer\-wise discriminative performance and cross\-layer generalisation were evaluated with ROC–AUC under stratified 10\-fold cross\-validation\. To identify contiguous ranges of layers with above\- or below\-chance performance, a nonparametric, cluster\-based one\-sample permutation test was applied \(two\-sided; 1000 permutations; cluster\-forming threshold from thettdistribution;α=0\.01\\alpha=0\.01\) over AUC−0\.5\-0\.5\.
### 2\.4Cross\-layer Generalisation Analysis
A single logistic\-regression probe was trained on features pooled from layers 20–31 and evaluated on each test layer 0–19\. This GAT\-style design tests whether the late\-layer representational format transfers to early layers, rather than maximising per\-layer accuracy\.
#### Reproducibility\.
## 3Results
\(a\)Layer\-wise classification \(canonical vs\. non\-canonical\)\. Pre\-clause \(dashed\) and post\-clause \(solid\) ROC–AUC; shaded = SEM; bars =p<0\.01p<0\.01clusters\.
\(b\)Cross\-layer generalisation matrix\. Each cell = ROC–AUC when a probe trained on one layer is tested on another\. Contours markp<0\.01p<0\.01clusters\.
\(c\)Probe coefficient sign score per layer\. Positive = net positive coefficient mass; negative = net negative\. Grey curve = mean normalised coefficient\.
\(d\)Late\-to\-early generalisation\. AUC for a probe trained on layers 20–31, tested on each early layer\. Shaded = significantly below chance \(p<0\.01p<0\.01\)\.
Figure 1:Full results for Llama\-Krikri\-8B\-Base \(128\-sentence set\)\. \(a\) Layer\-wise AUC; \(b\) cross\-layer GAT matrix; \(c\) coefficient sign dynamics; \(d\) late\-to\-early transfer\. The sign reversal around layer 22 in \(c\) is consistent with the directional below\-chance transfer in \(d\)\.### 3\.1Sentence\-type information emerges only after the clause boundary
Post\-clause activations yielded above\-chance decoding performance from early layers onward, peaking in the middle layers \(Figure[1\(a\)](https://arxiv.org/html/2609.00416#S3.F1.sf1)\)\. Pre\-clause activations remained at chance level at all layers, as expected: sentence\-type information cannot be inferred before the structure\-defining constituents of the relative clause appear\. Performance declined towards chance in the final layers\. The expanded 1024\-sentence set and both comparison model variants replicated the same qualitative layer\-wise profile \(Appendix[A](https://arxiv.org/html/2609.00416#A1)\)\.
### 3\.2Middle layers support broad cross\-layer generalisation
Cross\-layer generalisation analysis revealed a contiguous middle\-layer region \(approximately layers 5–19\) in which classifiers trained on one layer generalised above chance to a wide range of test layers \(Figure[1\(b\)](https://arxiv.org/html/2609.00416#S3.F1.sf2)\)\. This indicates a shared representational format for sentence type across this portion of the network\. The embeddings from late layers showed reduced cross\-layer generalisation, with AUC values approaching chance \(0\.50\) by the final layers\.
### 3\.3Late\-layer decision boundaries do not transfer to earlier layers
A probe trained only on late\-layer features \(layers 20–31\) produced below\-chance AUC when tested on earlier layers \(Figure[1\(d\)](https://arxiv.org/html/2609.00416#S3.F1.sf4)\), with significant effects across a contiguous cluster of early layers \(p<0\.01p<0\.01\)\. This result indicates that late\-layer representations encode sentence type in a different linear format from early\-layer representations\. On the significant early\-layer cluster, the probe classified 99\.3% of true non\-canonical sentences as canonical and 98\.0% of true canonical sentences as canonical, yielding an overall accuracy of 49\.3%\. The below\-chance transfer therefore reflects a systematic directional bias: the late\-layer decision boundary, when applied to early\-layer representations, treats virtually all items as canonical regardless of their actual word order\. The same qualitative late\-to\-early below\-chance pattern was observed in the 1024\-sentence set and in both comparison models \(Appendix[B](https://arxiv.org/html/2609.00416#A2)\)\.
### 3\.4Probe coefficients reverse sign in later layers
The directional nature of the transfer failure can be traced to the probe weights themselves\. We define a per\-layersign scoreas the proportion of positive probe coefficients minus the proportion of negative ones; a score near\+1\+1indicates that all coefficients are positive, a score near−1\-1that all are negative\. Around layer 22, the sign score crosses zero and the dominant probe coefficients flip sign \(Figure[1\(c\)](https://arxiv.org/html/2609.00416#S3.F1.sf3)\): the same distributional features that predicted non\-canonical class membership in early\-to\-middle layers predict canonical class membership in late layers\. This sign reversal is consistent with a change in the direction of the linearly decodable contrast across layers, and it explains why the late\-layer boundary produces inverted, rather than merely weak, transfer to early\-layer representations\.
## 4Discussion
The central finding is that late transformer layers do not simply lose syntactic sensitivity; they recode it in a specific direction\. Prior probing work established that syntactic information peaks in middle layers and declines thereafter\([Tenney et al\., 2019a](https://arxiv.org/html/2609.00416#bib.bib8);[Tenney et al\., 2019b](https://arxiv.org/html/2609.00416#bib.bib9);[Hewitt and Manning, 2019](https://arxiv.org/html/2609.00416#bib.bib5);[Coenen et al\., 2019](https://arxiv.org/html/2609.00416#bib.bib10)\)\. Those studies reportwhereinformation resides at each layer; the present cross\-layer analysis revealshowthe representational format changes across layers\. The decline is a directional shift toward canonical\-form representations, not a uniform weakening\. This converges with evidence that late layers support syntax\-invariant operations even in the semantic domain\([Zacharopoulos and Kyriakoglou, 2025](https://arxiv.org/html/2609.00416#bib.bib16)\)\.
Three alternative accounts deserve consideration\. Simple information loss predicts near\-chance transfer in both directions; it does not predict a 99\.3% directional bias toward one class\. Feature compression could produce biased transfer, but the coordinated sign reversal across all four distributional features around layer 22 is more consistent with a systematic representational change than with incidental compression\. Probe mismatch predicts noisy, variable errors; instead, misclassification is nearly deterministic\. We therefore interpret the evidence as most consistent with directional recoding, while acknowledging that causal intervention \(e\.g\., ablation of late\-layer representations\) would be needed to establish this conclusively\.
Because SVO order and canonical status are co\-extensive in this design, the recoding could reflect a specifically syntactic shift or a more general compression toward the high\-frequency form \(see Limitations\)\.
Since transformers allow direct inspection of internal states, the cross\-layer analysis generates a testable prediction for human neural data: the GAT\-style design could be replicated with EEG or MEG recordings from participants exposed to the same Greek stimuli, testing whether temporal generalisation over neural responses shows a corresponding directional asymmetry\. Cross\-linguistic extension and constructions where canonicality and surface word order dissociate remain important directions for future work\.
## 5Conclusion
Controlled scrambling in Greek allowed us to isolate syntactic structure from semantic content and track how word\-order representations evolve across transformer layers\. Syntactic information peaks in a mid\-layer zone of stable representations and is recoded in late layers into a format directionally aligned with the canonical word order \(99\.3% non\-canonical→\\tocanonical misclassification; sign reversal at layer 22\)\. These findings characterise a directional recoding process whose signature is directly testable in human EEG and MEG data\.
## 6Limitations
Four constraints qualify our interpretation\. First, linear probes detect only linearly decodable information; nonlinear representations may persist in late layers\. Second, SVO order and canonical status are co\-extensive in this design, so the probe may track either dimension\. Third, residual lexical regularities in template\-generated stimuli might contribute weakly to classification, although the balanced design minimises this; Jabberwocky variants could isolate structural form more fully\. Fourth, the study examines a single language and construction type; cross\-linguistic and cross\-constructional generalisation remains open\.
## 7Ethical Considerations
This study uses synthetic stimuli and pretrained language models\. No human participants were recruited and no personal data were collected\. We identify no direct participant\-related ethical risks\.
## References
- Y\. Belinkov, S\. Gehrmann, and E\. PavlickInterpretability and Analysis in Neural NLP\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts,A\. Savary and Y\. Zhang \(Eds\.\),Online,pp\. 1–5\.External Links:[Link](https://aclanthology.org/2020.acl-tutorials.1/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-tutorials.1)Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p5.1)\.
- Caucheteuxet al\.\(2021\)C\. Caucheteux, A\. Gramfort, and J\. KingLong\-range and hierarchical language predictions in brains and algorithms\.arXiv preprint arXiv:2111\.14232\.Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p7.1)\.
- Coenenet al\.\(2019\)A\. Coenen, E\. Reif, B\. Kim, A\. Pearce, F\. Viégas, and M\. WattenbergVisualizing and measuring the geometry of BERT\.Advances in Neural Information Processing Systems32\.Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p5.1),[§4](https://arxiv.org/html/2609.00416#S4.p1.1)\.
- Desbordeset al\.\(2026\)T\. Desbordes, I\. Olasagasti, N\. Piron, S\. Schwartz, and N\. KazaninaTemporal evolution of neural codes: The added value of a geometric approach to linear coefficients\.NeuroImage327,pp\. 121737\.External Links:ISSN 1053\-8119,[Link](https://www.sciencedirect.com/science/article/pii/S1053811926000558),[Document](https://dx.doi.org/10.1016/j.neuroimage.2026.121737)Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p6.1)\.
- Georgiafentiset al\.\(2025\)M\. Georgiafentis, S\. Skopeteas, and A\. TsokoglouInformation structure in greek: interface and comparative studies\.Journal of Greek Linguistics25\(1\),pp\. 3–9\.Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p4.1)\.
- Goldsteinet al\.\(2025\)A\. Goldstein, E\. Ham, M\. Schain, S\. A\. Nastase, B\. Aubrey, Z\. Zada, A\. Grinstein\-Dabush, H\. Gazula, A\. Feder, W\. Doyle, S\. Devore, P\. Dugan, D\. Friedman, M\. Brenner, A\. Hassidim, Y\. Matias, O\. Devinsky, N\. Siegelman, A\. Flinker, O\. Levy, R\. Reichart, and U\. HassonTemporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models\.Nature Communications16\(1\),pp\. 10529\(en\)\.External Links:ISSN 2041\-1723,[Link](https://www.nature.com/articles/s41467-025-65518-0),[Document](https://dx.doi.org/10.1038/s41467-025-65518-0)Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p7.1)\.
- Hewitt and Manning \(2019\)J\. Hewitt and C\. D\. ManningA Structural Probe for Finding Syntax in Word Representations\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota,pp\. 4129–4138\.External Links:[Link](https://aclanthology.org/N19-1419/),[Document](https://dx.doi.org/10.18653/v1/N19-1419)Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p5.1),[§4](https://arxiv.org/html/2609.00416#S4.p1.1)\.
- Katsika and Allen \(2013\)K\. Katsika and S\. AllenProcessing subject and object relative clauses in a flexible word order language: evidence from greek\.In2013 Conference on Architectures and Mechanisms in Language Processing, Marseille, France,Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p4.1)\.
- King and Dehaene \(2014\)J\. King and S\. DehaeneCharacterizing the dynamics of mental representations: the temporal generalization method\.Trends in Cognitive Sciences18\(4\),pp\. 203–210\(English\)\.External Links:ISSN 1364\-6613, 1879\-307X,[Link](https://www.cell.com/trends/cognitive-sciences/abstract/S1364-6613(14)00019-9),[Document](https://dx.doi.org/10.1016/j.tics.2014.01.002)Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p6.1)\.
- Roussiset al\.\(2025\)D\. Roussis, L\. Voukoutis, G\. Paraskevopoulos, S\. Sofianopoulos, P\. Prokopidis, V\. Papavasileiou, A\. Katsamanis, S\. Piperidis, and V\. KatsourosKrikri: advancing open large language models for greek\.Findings of the Association for Computational Linguistics: EMNLP 2025,pp\. 5012–5033\.Cited by:[§2\.1](https://arxiv.org/html/2609.00416#S2.SS1.p1.1)\.
- Schrimpfet al\.\(2021\)M\. Schrimpf, I\. A\. Blank, G\. Tuckute, C\. Kauf, E\. A\. Hosseini, N\. Kanwisher, J\. B\. Tenenbaum, and E\. FedorenkoThe neural architecture of language: Integrative modeling converges on predictive processing\.Proceedings of the National Academy of Sciences118\(45\),pp\. e2105646118\.External Links:[Link](https://www.pnas.org/doi/full/10.1073/pnas.2105646118),[Document](https://dx.doi.org/10.1073/pnas.2105646118)Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p7.1)\.
- Tenneyet al\.\(2019a\)I\. Tenney, D\. Das, and E\. PavlickBERT Rediscovers the Classical NLP Pipeline\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,pp\. 4593–4601\.Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p5.1),[§4](https://arxiv.org/html/2609.00416#S4.p1.1)\.
- Tenneyet al\.\(2019b\)I\. Tenney, P\. Xia, B\. Chen, A\. Wang, A\. Poliak, R\. T\. McCoy, N\. Kim, B\. Van Durme, S\. R\. Bowman, D\. Das, and E\. PavlickWhat do you learn from context? Probing for sentence structure in contextualized word representations\.InProceedings of the 7th International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p5.1),[§4](https://arxiv.org/html/2609.00416#S4.p1.1)\.
- Zacharopouloset al\.\(2023\)C\. Zacharopoulos, T\. Desbordes, and M\. Sablé\-MeyerAssessing the influence of attractor\-verb distance on grammatical agreement in humans and language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 16081–16090\.Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p7.1)\.
- Zacharopouloset al\.\(2026\)C\. Zacharopoulos, S\. Dehaene, and Y\. LakretzDisentangling hierarchical and sequential computations during sentence processing\.Cortex\.External Links:ISSN 0010\-9452,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.cortex.2026.02.004),[Link](https://www.sciencedirect.com/science/article/pii/S0010945226000456)Cited by:[§1](https://arxiv.org/html/2609.00416#S1.p7.1)\.
- Zacharopoulos and Kyriakoglou \(2025\)C\. Zacharopoulos and R\. KyriakoglouIn machina N400: pinpointing where a causal language model detects semantic violations\.InArtificial Intelligence and Cognitive Science,Communications in Computer and Information Science, Vol\.2950\.Cited by:[§4](https://arxiv.org/html/2609.00416#S4.p1.1)\.
## Appendix ACross\-model GAT matrices
This appendix reports the generalization\-across\-layers \(GAT\) matrices for the two comparison models emphasized in the manuscript, both evaluated on the expanded10241024\-sentence stimulus set rather than the main128128\-sentence set\. In both cases, we observe the same qualitative structure described in the main text: a broad mid\-layer regime of above\-chance generalization together with a late\-to\-early transfer regime that falls below chance\. The precise extent of the clusters varies by model, but the overall representational organization is preserved, indicating that the main findings replicate across comparison systems even under the larger stimulus regime\.
\(a\)KriKri\-Instruct
\(b\)Plutus
Figure 2:Generalization\-across\-layers matrices for KriKri\-Instruct and Plutus, both evaluated on the expanded10241024\-sentence stimulus set\. In both models, the dominant pattern is preserved: a broad zone of above\-chance generalization in middle layers and a later regime showing below\-chance transfer to earlier layers\.To complement the GAT matrices, Figure[3](https://arxiv.org/html/2609.00416#A1.F3)displays the corresponding coefficient plots for the same large\-set runs\. These show the same broad sign\-reversal profile across models: coefficients are predominantly positive in the middle\-to\-late layers where decoding is strongest, and they flip sign in the late regime that drives the below\-chance transfer pattern\.
\(a\)KriKri\-Instruct coefficients
\(b\)Plutus coefficients
Figure 3:Coefficient plots for the two comparison models on the expanded10241024\-sentence stimulus set\. Both models show the same broad sign\-reversal pattern that accompanies the late\-to\-early below\-chance generalization regime under the large stimulus set\.
## Appendix BLexicon and expanded stimulus set
### B\.1Tokenization & Forward pass
All sentences were fed to the model using its native subword tokenizer, which extends Llama 3\.1 with Greek\-specific units\. Inputs were the raw Greek strings from the controlled stimulus set, with sentence\-final punctuation preserved and no prompt or few\-shot context added\. We executed inference\-only forward passes with the pretrained weights, disabling caching and dropout, and extracted the complete set of hidden states at every layer and token position\.
The clause\-initial complementizerserved as a fixed temporal anchor\. For each sentence, we partitioned the token sequence into a “before” region \(up to and including the complementizer\) and an “after” region\. Token–word alignment exploited this anchor and the deterministic SVO/VSO templates: multi\-token words \(e\.g\., morphologically complex nouns or verbs\) were treated as spans, and their word\-level representations were computed as the mean of constituent token vectors\. This procedure allowed us to consistently index the determiner and head noun of the matrix subject \(DetN1Det\\penalty\\ N1,N1N1\), the determiner and head noun inside the relative clause \(DetN2Det\\penalty\\ N2,N2N2\), and the two verbs \(V1 transitive inside the clause, V2 sentence\-final\)\. These layer\-by\-position activations constitute the sole inputs to our probing analyses and all subsequent figures\.
### B\.2Lexicon
The experimental materials were generated from a fixed lexicon of Greek nouns, verbs, and determiners\. The noun inventory consisted of human\-denoting stems with masculine and feminine forms in both singular and plural\. The verbal inventory included distinct transitive and intransitive paradigms for singular and plural forms\. Tables[1](https://arxiv.org/html/2609.00416#A2.T1)and[2](https://arxiv.org/html/2609.00416#A2.T2)report the complete lexicon used by the stimulus generator\.
Table 1:Noun lexicon used to generate the experimental stimuli\.Table 2:Verb and determiner lexicon used to generate the experimental stimuli\.
### B\.3Expanded stimulus set
In addition to the main 128\-sentence experiment, we generated an expanded stimulus set containing10241024sentences for the sensitivity analyses reported in the appendix\. The larger set preserves the same design logic as the main experiment: it contains the same 32 fully counterbalanced linguistic conditions, enforces distinct lexical stems forN1N1andN2N2, and varies only word order while keeping sentence meaning fixed within each minimal\-pair contrast\.
Table 3:Summary of the two generated stimulus sets\.The expanded set was used to test whether the decoding and generalization effects reported in the main text remain stable when the number of lexicalized sentence instances per condition is substantially increased\. The qualitative pattern is preserved under this larger sampling regime, indicating that the reported effects are not an artifact of the smaller 128\-item stimulus set\. The full sentence lists are available in the[OSF project](https://osf.io/5d3w8/overview?view_only=d42af279745543808cc377b9f96cb1af)undercode/stimuli\(greek\_sentences\_128\.csvandgreek\_sentences\_1024\.csv\); see[https://osf\.io/5d3w8/overview?view\_only=d42af279745543808cc377b9f96cb1af](https://osf.io/5d3w8/overview?view_only=d42af279745543808cc377b9f96cb1af)\.
### B\.4Example generated sentences
Below we provide three representative minimal pairs from the 128\-sentence set, illustrating how the template produces matched SVO and VSO stimuli\. In each pair, only the word order inside the bracketed relative clause differs; all lexical items, morphological forms, and propositional content are identical\.
Pair 1 \(masculine singular, transitive =ςµ\\acctonos\):
SVO:\\acctonosς \[ µ\\acctonosςµ\\acctonos\]\\acctonos\. def\.m\.nomteacher\.nomcomp\[def\.f\.nomstudent\.nomlike\.3sg\] leave\.3sg ‘The teacher that the student likes is leaving\.’
VSO:\\acctonosς \[ςµ\\acctonosµ\\acctonos\]\\acctonos\. def\.m\.nomteacher\.nomcomp\[like\.3sgdef\.f\.nomstudent\.nom\] leave\.3sg ‘The teacher that the student likes is leaving\.’
Pair 2 \(feminine singular, transitive =µ\\acctonos\):
SVO:ς\\acctonosµ \[\\acctonosµ\\acctonos\]\\acctonos\. def\.f\.nomnurse\.nomcomp\[def\.m\.nomresearcher\.nomadmire\.3sg\] laugh\.3sg ‘The nurse that the researcher admires is laughing\.’
VSO:ς\\acctonosµ \[µ\\acctonos\\acctonos\]\\acctonos\. def\.f\.nomnurse\.nomcomp\[admire\.3sgdef\.m\.nomresearcher\.nom\] laugh\.3sg ‘The nurse that the researcher admires is laughing\.’
Pair 3 \(masculine plural, transitive =µς\\acctonos\):
SVO:\\acctonos\[ ςς\\acctonosµς\\acctonos\]\\acctonos\. def\.pl\.nomathlete\.nom\.plcomp\[def\.pl\.nomdesigner\.nom\.plhate\.3pl\] cry\.3pl ‘The athletes that the designers hate are crying\.’
VSO:\\acctonos\[µς\\acctonosςς\\acctonos\]\\acctonos\. def\.pl\.nomathlete\.nom\.plcomp\[hate\.3pldef\.pl\.nomdesigner\.nom\.pl\] cry\.3pl ‘The athletes that the designers hate are crying\.’Similar Articles
The Changing Geometry of Grammar: Dimensionality and Neighborhood Reorganization across Transformer Layers
This paper investigates how the grammatical role of tokens shapes the geometry of transformer representations across layers, finding distinct evolution patterns in encoder versus decoder models.
Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words
This paper explores how transformer models handle functional words like pronouns and adverbs, analyzing their embeddings and generalization capabilities. It finds that training on mixed lexicalized and functional sentences helps models learn shared syntactic-semantic structures.
When transformers learn "impossible" languages, what do they learn?
This paper investigates how transformer language models learn 'impossible' languages with unnatural properties, finding that while grammatical sensitivity degrades gradually, generative production shows pronounced failures, suggesting a linking hypothesis for non-attestation.
Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models
This paper systematically compares word order preferences in decoder-only language models across artificial and natural languages, finding that preferences are data-driven and can reduce linguistic diversity with the widespread adoption of LLMs.
Probing Character-level Transformers for the Spanish L-shaped Morphome
This paper probes character-level transformers to investigate whether they encode the Spanish L-shaped morphome, an irregular morphological pattern, as an abstract class or just surface alternations. The authors find that the encoding is item-specific and localized, but does not generalize like human learners.