Probing Character-level Transformers for the Spanish L-shaped Morphome

arXiv cs.CL 论文

摘要

This paper probes character-level transformers to investigate whether they encode the Spanish L-shaped morphome, an irregular morphological pattern, as an abstract class or just surface alternations. The authors find that the encoding is item-specific and localized, but does not generalize like human learners.

arXiv:2608.03452v1 Announce Type: new Abstract: When a transformer learns an irregular morphological pattern, what has it learned? Our test case is the Spanish \emph{L-shaped morphome}, a complex irregular pattern in which the verb's stem alternates in exactly the first-person singular indicative and all subjunctive forms, and whose membership no phonological, semantic, or syntactic feature predicts. Prior studies have shown that character-level transformers can reproduce this pattern, but that evidence describes what models produce, not what they represent. Probing five architectures, twelve trained models each, under lemma-disjoint cross-validation with controls and surface baselines, we show that the models encode the L-shaped class itself, not just its visible alternations. It is decodable above every surface baseline, survives instances in which every form shows the same stem, and probes trained on alternating instances still classify non-alternating ones. The encoding is localized where the stem choice is made, at the stem-final consonant position of the middle decoder, before the alternant is read. And it is item-specific: which verbs a model learned matters far more than which architecture it is. The models store the morphome as an item-specific lexical abstraction, sufficient to reproduce the pattern but not to generalize it as humans do.
查看原文
查看缓存全文

缓存时间: 2026/08/05 07:45

# Probing Character-level Transformers for the Spanish L-shaped Morphome
Source: [https://arxiv.org/html/2608.03452](https://arxiv.org/html/2608.03452)
Akhilesh Kakolu Ramarao1, Kevin Tang1,2, Wiebke Petersen3, Dinah Baer\-Henney4 1Department of English Language and Linguistics, Heinrich Heine University Düsseldorf 2Department of Linguistics, College of Liberal Arts and Sciences, University of Florida 3Institute of Linguistics and Information Science, Heinrich Heine University Düsseldorf 4Institut für Germanistik, Philologische Fakultät, Ruhr\-Universität Bochum \{akhilesh\.kakolu\.ramarao, kevin\.tang,wiebke\.petersen\}@uni\-duesseldorf\.de, dinah\.baer\-henney@rub\.de

###### Abstract

When a transformer learns an irregular morphological pattern, what has it learned? Our test case is the Spanish*L\-shaped morphome*, a complex irregular pattern in which the verb’s stem alternates in exactly the first\-person singular indicative and all subjunctive forms, and whose membership no phonological, semantic, or syntactic feature predicts\. Prior studies have shown that character\-level transformers can reproduce this pattern, but that evidence describes what models produce, not what they represent\. Probing five architectures, twelve trained models each, under lemma\-disjoint cross\-validation with controls and surface baselines, we show that the models encode the L\-shaped class itself, not just its visible alternations\. It is decodable above every surface baseline, survives instances in which every form shows the same stem, and probes trained on alternating instances still classify non\-alternating ones\. The encoding is localized where the stem choice is made, at the stem\-final consonant position of the middle decoder, before the alternant is read\. And it is item\-specific: which verbs a model learned matters far more than which architecture it is\. The models store the morphome as an item\-specific lexical abstraction, sufficient to reproduce the pattern but not to generalize it as humans do\.

Probing Character\-level Transformers for the Spanish L\-shaped Morphome

Akhilesh Kakolu Ramarao1, Kevin Tang1,2, Wiebke Petersen3, Dinah Baer\-Henney41Department of English Language and Linguistics, Heinrich Heine University Düsseldorf2Department of Linguistics, College of Liberal Arts and Sciences, University of Florida3Institute of Linguistics and Information Science, Heinrich Heine University Düsseldorf4Institut für Germanistik, Philologische Fakultät, Ruhr\-Universität Bochum\{akhilesh\.kakolu\.ramarao, kevin\.tang,wiebke\.petersen\}@uni\-duesseldorf\.de, dinah\.baer\-henney@rub\.de

## 1Introduction

Morphomic patterns are among the most puzzling phenomena in inflectional morphology: systematic distributions of stem alternants over paradigm cells that no phonological, semantic, or syntactic property unifies\(Spencer and Aronoff,[1994](https://arxiv.org/html/2608.03452#bib.bib6); Maiden,[2018](https://arxiv.org/html/2608.03452#bib.bib8)\)\. They raise basic questions about morphological patterns\. Can such a complex pattern be learned from exposure to inflected forms alone? Is it stored as a property of individual forms, or of an abstract class of lexemes? And when a learner reproduces the pattern, does it thereby represent it, or only its visible alternations? The Spanish*L\-shaped morphome*is a well\-studied case: in verbs such as*salir*‘to leave’, the first\-person singular indicative \(*salgo*\) shares its stem with every subjunctive form \(*salga*,*salgas*, …\), while the remaining indicative forms use the regular stem \(*sales*,*sale*, …\)\. No phonological, semantic, or syntactic feature unifies exactly these cells, which is what makes it a morphomic pattern\. To choose the right stem, one must therefore know two things: whether the verb belongs to the arbitrary L\-shaped class, and which paradigm cell is being inflected\. Whether human speakers actually represent such a class is contested\(Nevinset al\.,[2015](https://arxiv.org/html/2608.03452#bib.bib7); Cappellaroet al\.,[2024](https://arxiv.org/html/2608.03452#bib.bib41)\)\.

Such questions are difficult to settle from human data alone because what a speaker has internalized can only be inferred from behavior\. Computational modeling offers a complementary approach where we build a learner whose input is fully known, and examine what it acquires\. For morphological inflection, the standard learner is a neural sequence\-to\-sequence model trained to map forms and morphosyntactic tags to inflected forms\. It can be trained on exactly the verbs we choose, and it can be examined both behaviorally, through the forms it produces, and representationally, through its internal states\. Character\-level transformers are the dominant approach to morphological inflection\(Wuet al\.,[2021](https://arxiv.org/html/2608.03452#bib.bib26); Kakolu Ramaraoet al\.,[2025](https://arxiv.org/html/2608.03452#bib.bib22)\), and the SIGMORPHON shared tasks have benchmarked them on complex morphological patterns across typologically diverse languages\(Cotterellet al\.,[2017](https://arxiv.org/html/2608.03452#bib.bib42),[2018](https://arxiv.org/html/2608.03452#bib.bib43); Kodner and Khalifa,[2022](https://arxiv.org/html/2608.03452#bib.bib44)\)\. For the L\-shape specifically, a recent line of work has established three behavioral facts\. Character\-level transformers reproduce the stem alternations, and its performance is highly dependent on the frequency of L\-shaped verbs in the training\(Kakolu Ramaraoet al\.,[2025](https://arxiv.org/html/2608.03452#bib.bib22)\)\. These models and human speakers also make opposite errors, the models apply the alternation to verbs where it does not belong, whereas speakers apply it less often than the pattern would license\(Kakolu Ramaraoet al\.,[2026b](https://arxiv.org/html/2608.03452#bib.bib23)\)\. And among five architectures varying in positional encoding and tag representation, position\-invariant tag encoding enables acquisition of the L\-shaped paradigm even when L\-shaped verbs are scarce, though no architecture generalizes like humans\(Kakolu Ramaraoet al\.,[2026a](https://arxiv.org/html/2608.03452#bib.bib24)\)\. However, these findings are based on model outputs\. A model may produce the correct forms because it has formed an internal category of L\-shaped verbs, or merely tracks the surface phonotactics of L\-shaped stems\.

Making that distinction requires examining the model’s internal representations\. The standard method is probing which involves training small diagnostic classifiers to predict a property from the model’s hidden states\(Conneauet al\.,[2018](https://arxiv.org/html/2608.03452#bib.bib3); Hupkes and Zuidema,[2018](https://arxiv.org/html/2608.03452#bib.bib36); Liuet al\.,[2019](https://arxiv.org/html/2608.03452#bib.bib51)\)\. We extract hidden states from every encoder and decoder layer of the five architectures ofKakolu Ramaraoet al\.\([2026a](https://arxiv.org/html/2608.03452#bib.bib24)\)\. Probing has three known pifalls, a classifier can memorize lemmas, succeed through its own capacity, or recover only what the surface forms already predict\(Hewitt and Liang,[2019](https://arxiv.org/html/2608.03452#bib.bib2); Ravichanderet al\.,[2021](https://arxiv.org/html/2608.03452#bib.bib48)\)\. We therefore evaluate only on lemmas a classifier never saw, consider only accuracy above a shuffled\-label control, and compare all results against surface\-form baselines\. Probing can also do more than detect the class\. At specific character positions, it can show whether the L\-shaped class is spread across the entire form, or concentrated where the stem choice is made, and whether it is available before the alternant is produced\. It can also show whether the encoding is organized as paradigm’s L\-shape\. And across models, it can show whether the class is encoded as a generalization, or is dependent on which particular verbs are chosen\. Therefore, we address three research questions:

- •RQ1Do the models encode which verbs are L\-shaped, beyond what the surface forms themselves predicts?
- •RQ2Where is the L\-shaped membership encoded: in which layers, at what positions within the forms, and does the encoding reflect the L\-shaped organization of the paradigm?
- •RQ3What determines how strongly a model encodes L\-shaped membership: its architecture, or the particular verbs it learned?

## 2Background

### 2\.1The L\-Shaped Morphome

Each Spanish verb has a twelve\-cell present\-tense paradigm, with three persons \(1, 2, 3\), two numbers \(singular, plural\), and two moods \(indicative, subjunctive\)\. A verb is*L\-shaped*when its1sg\.indstem is identical to all six subjunctive stems and distinct from the stems of the other five indicative cells; verbs without this configuration are*NL\-shaped*\(regular verbs\)\. Table[1](https://arxiv.org/html/2608.03452#S2.T1)shows the distribution of stem alternants for*salir*: the cells sharing the alternant*salg\-*trace an upside\-down “L” through the paradigm\. No natural class covers exactly these seven cells:1sg\.indgroups with the subjunctive against its own mood, which is precisely what makes the pattern morphomic, and comparable stem distributions recur across Romance\(Maiden,[2018](https://arxiv.org/html/2608.03452#bib.bib8),[2021](https://arxiv.org/html/2608.03452#bib.bib9)\)\.

Table 1:Present\-tense forms of*salir*‘to leave’, segmented into stem and ending\. The cells built on the alternant*salg\-*\(bold\) form the L\-shaped shape of the paradigm\.
### 2\.2Probing Neural Representations

Diagnostic classifiers are lightweight models trained to predict a linguistic property from frozen hidden states\(Conneauet al\.,[2018](https://arxiv.org/html/2608.03452#bib.bib3); Hupkes and Zuidema,[2018](https://arxiv.org/html/2608.03452#bib.bib36); Liuet al\.,[2019](https://arxiv.org/html/2608.03452#bib.bib51)\)\. Layer\-wise probing of this kind has mapped where linguistic information resides in transformers\(Tenneyet al\.,[2019](https://arxiv.org/html/2608.03452#bib.bib14); Dalviet al\.,[2019](https://arxiv.org/html/2608.03452#bib.bib12)\)and, more recently, in models of morphology\(Astrach and Pinter,[2025](https://arxiv.org/html/2608.03452#bib.bib52)\)\.

Probe accuracy on its own is difficult to interpret as a probe can succeed by memorizing lexical identity, by exploiting class imbalance, or by reading information off the surface string rather than out of the representation\(Belinkov,[2022](https://arxiv.org/html/2608.03452#bib.bib32)\)\. Section[4\.1](https://arxiv.org/html/2608.03452#S4.SS1)addresses each of these three shortcomings: lemma\-disjoint folds and structure\-preserving controls rule out lexical memorization, balanced accuracy removes the majority\-class guessing, and surface baselines measure what the string alone predicts\. How a hidden state is read out matters as well\. Mean\-pooling a layer and reading a single position can expose different information\(Ácset al\.,[2021](https://arxiv.org/html/2608.03452#bib.bib55); Liao and Shi,[2026](https://arxiv.org/html/2608.03452#bib.bib53)\)and our position\-targeted analysis addresses this concern\.

## 3Experimental setup

### 3\.1Data

We use the Spanish verbal paradigms in Seseo IPA transcription released byKakolu Ramaraoet al\.\([2025](https://arxiv.org/html/2608.03452#bib.bib22)\), and probe the publicly available models ofKakolu Ramaraoet al\.\([2026a](https://arxiv.org/html/2608.03452#bib.bib24)\)trained in their10%L\-90%NLcondition, in which L\-shaped verbs are as scarce as they are in the Spanish lexicon\. In that setup, 333 lemmas are sampled and partitioned into training \(233 lemmas\), development \(34\), and test \(66\) sets with no lemma overlap, so the models must generalize to unseen lemmas\.

The probing corpus is the entire test set of 43,560 instances\. An instance is a triple of three inflected forms of one lemma: two source forms and one target form, each from a different cell of the twelve\-cell present\-tense paradigm\. The twelve cells yield 66 undordered source pairs, and each pair combines with any of the 10 remaining cells as target, so every test lemma contributes66×10=66066\\times 10=660instances and the 66 test lemmas together give 43,560\. Only seven of the 66 test lemmas are L\-shaped, giving7×660=4,6207\\times 660=4,620instances, which several analyses below subdivide further\.111Each lemma split has its own seven L\-shaped test lemmas, see Table[3](https://arxiv.org/html/2608.03452#A1.T3)in Appendix[A](https://arxiv.org/html/2608.03452#A1)\.

##### Task\.

The models are trained on two\-source morphological re\-inflection\(Kannet al\.,[2017](https://arxiv.org/html/2608.03452#bib.bib54)\), framed as character\-level sequence\-to\-sequence transduction\(Wuet al\.,[2021](https://arxiv.org/html/2608.03452#bib.bib26)\): given two source form\-tag pairs from a verb’s paradigm and a target feature bundle, produce the target form\. It also mirrors the wug\-test paradigm of the human experiments, in which participants see two forms of a novel verb and produce a third\(Nevinset al\.,[2015](https://arxiv.org/html/2608.03452#bib.bib7)\)\.

##### Input format\.

Each source sequence concatenates two word forms and a target morphosyntactic tag, separated by a delimiter\. For example,

\\tipaencoding

s " a l g o<V;IND;PRS;1;SG\>\# \\tipaencodings " a l g a<V;SBJV;PRS;1;SG\>\#<V;SBJV;PRS;3;PL\>

and the target form is\\tipaencodings " a l g a n \(*salgan*\)\. How different architectures rewrite the tag content is explained in Section[3\.2](https://arxiv.org/html/2608.03452#S3.SS2)and Figure[1](https://arxiv.org/html/2608.03452#S3.F1)\.

### 3\.2Model Architectures

The five models share an encoder\-decoder transformer backbone\(Vaswaniet al\.,[2017](https://arxiv.org/html/2608.03452#bib.bib10)\)with 4 encoder and 4 decoder layers, 4 attention heads, and embedding dimensiond=256d=256, and differ in how the encoder treats morphosyntactic tags: whether tags receive sequential positional encoding or a fixed position, and whether tag content is atomic or decomposed into features\(Kakolu Ramaraoet al\.,[2026a](https://arxiv.org/html/2608.03452#bib.bib24)\)\. Figure[1](https://arxiv.org/html/2608.03452#S3.F1)shows the five variants on the same input: the two dimensions are visible as the positional index assigned to tag content \(sequential vs\. a fixed positional encoding of0\) and as the form in which tag content enters the embedding \(an atomic vocabulary token, decomposed tokens, or a structured feature vector\)\.

Vanillas0a1l2\\tipaencodingg3o4⟨\\langleV;SBJV;PRS;1;PL⟩\\rangle5atomic tag tokenC\-Seps0a1l2\\tipaencodingg3o4V5SBJV6PRS718PL9decomposed tag tokensF\-Invs0a1l2\\tipaencodingg3o4⟨\\langleV;SBJV;PRS;1;PL⟩\\rangle0atomic tag tokenF\-1Hs0a1l2\\tipaencodingg3o401100010one\-hot over \{IND,SBJV\|\|1,2,3\|\|SG,PL\}F\-Geos0a1l2\\tipaencodingg3o411111100one\-hot over \[±\\pmptcp±\\pmauth±\\pmpl±\\pmind\]Figure 1:The five architectures on a common input \(*salgo*\+ the tagv;sbjv;prs;1;pl; the real inputs contain two form–tag pairs and a target tag\)\. Blue boxes are character tokens, orange boxes tag content; the number under each token is its positional index\. The two sequential architectures \(top\) assign tags sequential positions and differ only in whether the tag is one token or five; the three position\-invariant architectures give all tag content the fixed position0and differ in tag content: an atomic token \(F\-Inv\), a one\-hot vector over feature categories \(F\-1H; bits shown forsbjv;1;pl\), or a Harley–Ritter\(Harley and Ritter,[2002](https://arxiv.org/html/2608.03452#bib.bib5)\)feature vector \(F\-Geo;\[\+participant,\+author,\+plural,−indicative\]\[\+\\text\{participant\},\+\\text\{author\},\+\\text\{plural\},\-\\text\{indicative\}\]for the same cell\)\.##### Sequential positional encoding\.

Vanillaconcatenates form characters and feature tags into one flat sequence, with each tag a single token\.Character\-separateddiffers only in decomposing the tag content into individual tag tokens\.

##### Position\-invariant tags\.

The other three architectures place all tag content at a fixed positional index of0and only the characters get sequential positions\. They differ in what the tag itself looks like\.Feature\-invariantkeeps the tag as one atomic token\.Feature\-onehotsplits the tag into its component features \(mood, person, number\) and turns them into a binary one\-hot vector\.Feature\-geometricalso uses a feature vector, but are categorized as aHarley and Ritter \([2002](https://arxiv.org/html/2608.03452#bib.bib5)\)feature \(±\\pmparticipant,±\\pmauthor,±\\pmplural,±\\pmindicative\)\.

##### Training and checkpoints\.

All models are taken from the studies ofKakolu Ramaraoet al\.\([2026a](https://arxiv.org/html/2608.03452#bib.bib24)\): for each architecture, three lemma splits crossed with four training data subsamples yield 12 trained models\.

### 3\.3Representation probing

All probing analyses start from the same extraction procedure\. We run each model under teacher forcing, capture the hidden states of all eight layers \(4 encoder, 4 decoder\), and mean\-pool each layer’s states into one fixed\-length vector per instance per layer\. Teacher forcing gives every model the identical input strings, so a probe cell means the same instance in every architecture, and it keeps the probe labels and the positional indices used in positional readouts of Experiment 2 \(Section[4\.2](https://arxiv.org/html/2608.03452#S4.SS2)\) aligned with the gold forms\. By default we pool over the*content*\(character\) positions only, excluding the morphological\-tag, “\#”\-separator, andbos/eospositions\.

## 4Experiments

The three experiments addresses the three research questions\. Experiment 1 \(Section[4\.1](https://arxiv.org/html/2608.03452#S4.SS1)\) asks whether the models encode the L\-shaped class; Experiment 2 \(Section[4\.2](https://arxiv.org/html/2608.03452#S4.SS2)\) asks where the L\-shaped is encoded; Experiment 3 \(Section[4\.4](https://arxiv.org/html/2608.03452#S4.SS4)\) asks what determines how strongly a given model encodes it\.

### 4\.1Experiment 1: Do the models encode the L\-shaped class?

This first experiment proceeds in two stages, matching the two sides of RQ1: stage one establishes that L\-shapedness is decodable from the representations beyond what the surface forms predict; stage two asks what that decodable information is, an abstract property of the lexeme or an evidence of the visible alternation\.

#### 4\.1\.1Setup

##### Probed properties\.

Three properties are probed per \(architecture, model, layer\):

- •l\-shaped: whether the lemma belongs to an L\-shaped paradigm, the morphome membership itself \(2 classes; 4,620 L\-shaped vs\. 38,940 NL\-shaped\)
- •stem\-final match: whether the stem\-final consonant is shared across the instance’s three paradigm forms, i\.e\. whether the L\-shaped alternation surfaces in the instance \(2 classes; 38,580 shared vs\. 4,980 differing\)
- •conjugation: verb class from the infinitive ending,*\-ar*,*\-er*, or*\-ir*\(3 classes; 35,640 vs\. 4,620 vs\. 3,300\)

l\-shapedis the target of the study as it is the morphome membership itself\.stem\-final matchis its counterpart as it indicates whether the alternation is visible\.conjugationis a second arbitrary lexeme\-level classification and it shows whether inflection\-class information is encoded at all, and it is correlated with L\-shapedness \(no L\-shaped verb is*\-ar*\)\.

##### Probes and cross\-validation\.

We use logistic regression models as linear probes \(ℓ2\\ell\_\{2\}regularization,C=1\.0C=1\.0, L\-BFGS, max 500 iterations\), implemented in scikit\-learn\(Pedregosaet al\.,[2011](https://arxiv.org/html/2608.03452#bib.bib20)\)\.222A small MLP probe \(one hidden layer of 10 units\) exceeds the linear probe by at most0\.050\.05balanced accuracy, so whatever the pooled representations encode is already linearly accessible, and we report linear probes everywhere\.

Cross\-validation uses 5\-fold*lemma\-disjoint*splits \(StratifiedGroupKFoldgrouped by lemma\)\. We use balanced accuracy, the mean of per\-class recall as the metric throughout\.

Probe results depend on the probe’s hyperparameters\(Hewitt and Liang,[2019](https://arxiv.org/html/2608.03452#bib.bib2); Voita and Titov,[2020](https://arxiv.org/html/2608.03452#bib.bib33)\), so we checked the regularization strength directly\. For all 60 models we reran thel\-shapedprobe at thedec2, the layer where decodability of the class most often peaks over all eight layers \(Table[2](https://arxiv.org/html/2608.03452#S4.T2)\), with regularization strengthC∈\{0\.01,0\.1,1,10\}C\\in\\\{0\.01,0\.1,1,10\\\}\. Balanced accuracy changes by at most0\.0070\.007betweenC=1C=1andC=10C=10and by at most0\.0440\.044betweenC=0\.1C=0\.1andC=1C=1\. The probe results are thus insensitive to the regularization choice, and we reportC=1C=1throughout\.

##### Control tasks\.

Every probe is paired with a control, a second probe trained on shuffled labels, whose accuracy shows what a probe can achieve when there is nothing real to find\(Hewitt and Liang,[2019](https://arxiv.org/html/2608.03452#bib.bib2)\)\. For the lemma\-level properties, the control shuffles which label each lemma carries: every lemma keeps one consistent \(wrong\) label, and the overall proportion of labels across lemmas is preserved, so the control preserves exactly the structure a lemma\-disjoint probe could exploit\.stem\-final matchneeds a different control, its label differs from instance to instance of the same lemma, so there is no single lemma label to shuffle, instead, labels are shuffled across individual instances, again preserving their overall proportion\. Control accuracy is averaged over 5 permutations, and we report*selectivity*which is the real probe’s balanced accuracy minus the control’s\.

##### Surface\-form baselines\.

A probe on hidden states is informative only relative to what the surface string already predicts, and in many instances the alternation itself is visible in the string\. The baselines are therefore matched to the probes in everything but the input\. Same labels, same lemma\-disjoint folds as the probes, and only the input differs, the surface string in place of the hidden state\.

##### n\-gram classifier \(surface\)\.

ACountVectorizerover phoneme n\-grams up to ordern∈\{1,2,3\}n\\in\\\{1,2,3\\\}feeds the identical logistic regression used for the representation probes, asking how well the same classifier can predict the label from surface co\-occurrences alone\.

##### n\-gram LM\.

One n\-gram language model is fit per label class \(L or NL\), on that class’s training folds; a held\-out instance is assigned to the class whose model finds its string more probable\.

#### 4\.1\.2Membership is decodable beyond the surface forms

Addressing RQ1 requires, first, that anything be decodable beyond the input string at all\. The setup below operationalizes each:*decodability*is how far a linear probe’s balanced accuracy on held\-out lemmas exceeds its control;*beyond the surface string*is how far it exceeds the surface baselines, which are given the same labels and folds but only the string\. This first experiment applies these measures to the three properties at every layer of every model\.

All three properties are linearly decodable above chance and above control from every architecture\. Figure[2](https://arxiv.org/html/2608.03452#S4.F2)shows the distribution over the twelve trained models of linear\-probe balanced accuracy at each architecture’s best layer; Appendix[B](https://arxiv.org/html/2608.03452#A2)Figure[3](https://arxiv.org/html/2608.03452#A2.F3)shows the same spread separately at each of the eight layers, and Table[2](https://arxiv.org/html/2608.03452#S4.T2)reports the best\-layer means\. Three patterns, one per probed property, are consistent across architectures\.

Table 2:Best\-layer linear\-probe results per architecture and property: balanced accuracy \(mean±\\pmstd over 12 models\) and selectivity over the control\. Chance is0\.500\.50for the binary properties and0\.330\.33forconjugation\. Model abbreviations:Van=Vanilla,C\-Sep=Character\-separated,F\-Inv=Feature\-invariant,F\-1H=Feature\-onehot,F\-Geo=Feature\-geometric\.![Refer to caption](https://arxiv.org/html/2608.03452v1/x1.png)Figure 2:Distribution over the 12 trained models of linear\-probe balanced accuracy at each architecture’s best layer, for the three probed properties\. Dots are individual models\. Dashed lines mark the strongest surface baseline; dotted lines mark chance\.
#### 4\.1\.3Setup: isolating the lexical class

Stage one \(Section[4\.1\.2](https://arxiv.org/html/2608.03452#S4.SS1.SSS2)\) just established that something about the L\-shaped class is decodable\. Stage two, this setup and the results that follow, asks what that something is\. There are two candidates: an abstract property of the lexeme, or the stem alternation visible in the input string itself\. On the full corpus the two are indistinguishable, an L\-shaped verb alternates by definition, but a given instance displays the alternation only when its three forms are taken from both sides of the paradigm’s L\-shape \(Table[1](https://arxiv.org/html/2608.03452#S2.T1)\), and most do\. So telling the two apart requires cases in which they predict different outcomes\. Each of the three manipulations below builds such a case by taking away one kind of surface information\.

##### Removing the visible alternation\.

L\-shapedness is closely tied to the visible stem\-final alternation\. An instance like \(*salgo*,*sales*→\\rightarrow*salga*\) shows both stem alternants, so anything that can compare the two stem\-final consonants can classify it, and success there does not indicate any abstraction\. The informative instances are the ones where no alternation is visible: a triple taken entirely from inside the L\-shaped pattern \(*salgo*,*salga*→\\rightarrow*salgas*\) or entirely from outside it \(*sales*,*sale*→\\rightarrow*salen*\) shows one stem throughout, exactly like the instances of a regular verb\. We call such triples*no\-alternation*instances, and the rest*alternation\-visible*instances\.

We therefore restrict the probing corpus to the instances whose stem\-final consonant is shared across all three paradigm forms and ask whether a probe can still classify L\-shaped vs\. NL\-shaped, under the same lemma\-disjoint folds, balanced accuracy, and controls, with the surface baselines recomputed on the same subset\.

##### Training on one subset, testing on the other\.

If the alternation\-visible and no\-alternation instances share one morphome representation, a probe trained on one subset should transfer to the other\. Within the same five lemma\-disjoint folds as the main probes, the L\-vs\-NL probe is therefore fit on the train\-lemma instances of onestem\-final matchsubset and evaluated on the test\-lemma instances of the other, in both directions\.

##### Removing the conjugation cue\.

No L\-shaped verb belongs to the*\-ar*conjugation, so conjugation alone predicts NL\-shapedness for the majority of verbs, and a probe that partly reads the model’s conjugation information scores above chance on L/NL without any morphome information\. We take this cue away by restricting the L/NL probe to*\-er*/*\-ir*instances, where conjugation carries no information about the class, under the same folds\.

#### 4\.1\.4The L\-shaped class survives without its surface evidence

All three manipulations point to the same conclusion: what the probes read is a property of the lexeme, not the visible alternation \(Appendix[D](https://arxiv.org/html/2608.03452#A4)Table[5](https://arxiv.org/html/2608.03452#A4.T5)\)\. The alternation\-visible subset sets the baseline for comparison\. Here every instance shows the stem change, so L\-shapedness can be read off the string directly\. The surface baselines are strong \(0\.780\.78at best\) and probes reach0\.730\.73–0\.800\.80\.

In the no\-alternation subset, the surface baseline lose the evidence they rely on and fall from0\.780\.78to at most0\.600\.60\. The probes still classify\. Every architecture remains above the best subset baseline \(0\.620\.62–0\.680\.68, selectivity0\.100\.10–0\.190\.19\), with the position\-invariant architectures at the top \(0\.660\.66–0\.680\.68\)\. Thereby, it is not a description of a visible alternation, because there is none and it is information the model itself associates with the lexeme\.

Probes trained only on alternation\-visible instances classify no\-alternation instances of held\-out lemmas above the permutation control in every architecture \(best layer0\.570\.57–0\.650\.65balanced accuracy, selectivity\+0\.08\+0\.08–\+0\.13\+0\.13\), peaking in the middle decoder\. The probe has never seen the test lemma and has never seen an instance without a visible alternation, yet the direction it learned still separates L from NL\. The reverse direction transfers as well, and more strongly \(0\.670\.67–0\.720\.72, selectivity\+0\.09\+0\.09–\+0\.21\+0\.21\)\.

Restricted to*\-er*/*\-ir*instances, where conjugation carries no information about L\-shapedness, the L/NL probe remains above its shuffled\-label control in every architecture, and decodability is now strongest atenc3\(0\.650\.65–0\.720\.72\) and weakens through the decoder\. In the full corpus, part of what the decoder appears to encode about the class is in fact conjugation information\. And the class signal that remains once conjugation is removed is strongest in the encoder\. We return to this encoder\-decoder split in the Discussion \(Section[5](https://arxiv.org/html/2608.03452#S5)\), after Experiment 2 \(Section[4\.2](https://arxiv.org/html/2608.03452#S4.SS2)\) has located the signal within the input\.

### 4\.2Experiment 2: Where is the class encoded?

RQ2 asks where in the model the class is encoded, at which layers the class is encoded, at which positions within the form, and whether the encoding carries the paradigm structure that defines the L\-shape\.

##### Reading at the alternation positions\.

Mean pooling shows whether a property is present somewhere in the averaged representation, but not at which position in the word it is encoded\. The stem\-final consonant is the segment that alternates in L\-shaped verbs \(*salg\-*vs\.*sal\-*\)\. We therefore extract the hidden state at exactly that position of each form and probe it as before\.

##### Reading before the alternant\.

A decoder\-side readout under teacher\-forcing has a catch: the probed state has already read the alternant\. To separate the two we add a pre\-alternant readout for decoder layers, the state at the content position immediately before the stem\-final consonant\. Under teacher forcing the state at characterkkpredicts characterk\+1k\{\+\}1, so this state is about to produce the alternant but has not yet seen it\.

##### Testing the paradigm configuration\.

Decodability at the alternation positions shows where the class is encoded, but not whether the encoding has the shape of the L\-shaped pattern, with1sg\.indgrouped with the subjunctive\. We therefore train a linear mood classifier \(IND vs\. SBJV\) on all target cells*except*1sg\.ind, so that it learns the indicative–subjunctive boundary from other cells \(*sales*vs\.*salga*,*cantas*vs\.*cantes*\), and then apply it to held\-out\-lemma1sg\.indinstances\. If the model organizes this paradigm,*salgo*should fall on the subjunctive side of the boundary while*canto*falls on the indicative side\.

### 4\.3Decodability concentrates where the alternant is chosen

The morphome signal is not spread evenly over the word but it is concentrated at the segment that alternates\. The stem\-final\-position readout \(Appendix[E](https://arxiv.org/html/2608.03452#A5)Table[6](https://arxiv.org/html/2608.03452#A5.T6)\) shows that at the alternation position itself the property becomes clearly decodable \(0\.680\.68–0\.760\.76, against0\.590\.59–0\.610\.61under mean pooling\)\.l\-shapedreaches0\.760\.76–0\.840\.84at the stem\-final position, far above its mean\-pooled values, with across\-model standard deviations roughly halved \(e\.g\.±\.06\\pm\.06forFeature\-invariant, vs\.±\.11\\pm\.11pooled\)\. When restricted to no\-alternation instances, L\-shapedness remains decodable at the very segment that would alternate, at0\.730\.73–0\.810\.81\.

The pre\-alternant readout shows that this decoder signal is predictive\. At the state that is about to produce the stem\-final consonant but has not yet seen it,l\-shapedremains decodable at0\.780\.78–0\.800\.80, within0\.000\.00–0\.050\.05of the stem\-final\-position readout above, and holds at0\.730\.73–0\.770\.77in the no\-alternation subset\.

#### 4\.3\.1The decoder carries the L\-shaped configuration

The encoding also has the shape of the L\-pattern \(Appendix[C](https://arxiv.org/html/2608.03452#A3)Table[4](https://arxiv.org/html/2608.03452#A3.T4)\)\. Atdec2, the mood probe places held\-out1sg\.indinstances of L\-shaped verbs on the subjunctive side far more often than those of NL\-shaped verbs \(meanP​\(sbjv\)P\(\\textsc\{sbjv\}\)differences of0\.110\.11to0\.250\.25forVanillaand the three position\-invariant architectures\) and atenc3the contrast is absent\. The defining cell of the L\-pattern thus groups with the subjunctive, for exactly the verbs it should, and does so in the decoder\.Character\-separatedis the exception, with almost no contrast and this mirrors its behavior, where it also separates the paradigm cells least well\(Kakolu Ramaraoet al\.,[2026a](https://arxiv.org/html/2608.03452#bib.bib24)\)\.

### 4\.4Experiment 3: What determines encoding strength?

In this experiment, we address RQ3 which asks whether encoding strength is governed by the architecture or by the verbs the model learned from\.

Within an architecture, models differ only in their lemma split \(three levels\) and their training subsample \(four levels\), so the across\-model variance can be split between these two factors, reported asη2\\eta^\{2\}, the share of variance lying between the levels of a factor\. All five architectures share the same 12 split×\\timessubsample conditions, so the two architecture groups can be compared pairwise within matched conditions, holding the training data fixed\.

#### 4\.4\.1The lemma split dominates

Which verbs a model is trained on matters more than which architecture it is\. Decodability varies widely from model to model: the 12\-model standard deviations in Table[2](https://arxiv.org/html/2608.03452#S4.T2)are±0\.09\\pm 0\.09–0\.120\.12forl\-shaped, roughly twice the entire spread between the best and worst architecture\. Models of the same architecture fall into a weakly decodable and a strongly decodable group at the same layer, ranging from0\.510\.51to0\.910\.91, and Figure[3](https://arxiv.org/html/2608.03452#A2.F3)shows the split holding across layers\. The lemma split fixes which seven L\-shaped verbs are held out for testing; the training subsample fixes which25%25\\%of each verb’s training triples the model sees\. A model that had learned a general rule for the class should decode it whatever verbs it is tested on, so its decodability would follow the random subsample, not the split\. A model that had instead stored the class verb by verb should decode it only for the verbs it happened to learn, so its decodability would follow the split\. Atdec2the lemma split accounts for4949–98%98\\%of the across\-model variance inl\-shapeddecodability per architecture \(η2=0\.71\\eta^\{2\}=0\.71pooled, after removing architecture means\), and the training subsample explains at most23%23\\%\(pooled0\.030\.03\)\.

#### 4\.4\.2Architecture matters only without surface evidence

Architecture still plays a small role\. Averaging the two groups within each matched condition, the position\-invariant architectures exceed the sequential ones in 10 of 12 conditions in the no\-alternation subset \(mean difference\+0\.04\+0\.04; pairedt​\(11\)=2\.69t\(11\)=2\.69,p=\.021p=\.021; Wilcoxonp=\.034p=\.034\)\. On the full corpus the same comparison is not reliable \(\+0\.02\+0\.02,p=\.24p=\.24\)\. The architectural advantage is therefore specific to the instances where surface evidence is absent\.

## 5Discussion and Conclusion

We probed five character\-level inflection transformer architectures, twelve independently trained models each, for the Spanish L\-shaped morphome under lemma\-disjoint folds, shuffled\-label controls, and surface\-form baselines\. We find that the models encode the L\-shaped class \(RQ1\), that the encoding is concentrated at the stem\-final consonant position of the middle decoder, present before the alternant is read \(RQ2\), and that its strength is governed by the lexical sample a model learned from far more than by its architecture \(RQ3\)\.

What the models encode is an association between the lexeme’s class and the segment that alternates, not the alternation itself\. The weakstem\-final matchresults point the same way: the pooled readout only faintly exposes whether an alternation is visible at all \(0\.590\.59–0\.610\.61\), well below the L\-shape decodability observed \(0\.700\.70–0\.760\.76\)\. Thereby, the probes classify the class better than they can see its surface evidence, so they must be reading something else\.

The pre\-alternant readout shows that the class information is present in the state that has not yet read the alternant, and placed where it can inform the stem choice\. The clustering of1sg\.indwith the subjunctive is a decoder phenomenon, while the conjugation\-independent lexical signal is strongest in the encoder\. This shows that the encoder carries the lexeme’s class and the decoder converts it, together with the mood of the target cell, into the stem choice\.

Architecture matters only where surface evidence is absent: the position\-invariant models’ advantage is reliable in that subset alone, so the choice identified behaviorally as an inductive bias\(Kakolu Ramaraoet al\.,[2026a](https://arxiv.org/html/2608.03452#bib.bib24)\)shows up internally as a stronger item\-specific encoding of the class\. Thereby, what positional\-invariant architectures add is not knowledge independent of the verbs learned, but a stronger encoding of the stored class exactly where no alternation is visible\. Dual\-route accounts store irregular morphology item by item and compute regular morphology by rule\(Prasada and Pinker,[1993](https://arxiv.org/html/2608.03452#bib.bib45)\)\. Neural learners challenge that division as a single network handles regulars and irregulars alike\(Kirov and Cotterell,[2018](https://arxiv.org/html/2608.03452#bib.bib37)\)\. Our single\-route models nonetheless form item\-specific knowledge of the class, so a single mechanism does not rule out storage\-like knowledge\.

## Limitations

The probing corpus contains only seven L\-shaped test lemmas per split \(21 in total\)\. The alternation types are also unevenly spread, they all share\\tipaencoding/s/ outside the L\-shape region \(Appendix[A](https://arxiv.org/html/2608.03452#A1)Table[3](https://arxiv.org/html/2608.03452#A1.T3)\), so the L\-shapedness is partly predictable from the stem\-final segment alone, and decodability there may partly reflect phonology rather than the class\. Whether the decoder causally relies on the decodable morphome signal requires interventional methods\(Ravfogelet al\.,[2020](https://arxiv.org/html/2608.03452#bib.bib35); Elazaret al\.,[2021](https://arxiv.org/html/2608.03452#bib.bib34)\)\. All decoder representations are extracted under teacher forcing, so they reflect the processing of a correct continuation rather than of each model’s own production\.

## Ethics Statement

This work analyzes the internal representations of publicly released models trained on openly available Spanish verbal paradigm data\(Kakolu Ramaraoet al\.,[2026a](https://arxiv.org/html/2608.03452#bib.bib24)\)\. No new data were collected and no human participants were involved, and the human evidence discussed comes from prior studies\. We foresee no ethical risks arising from this analysis\.

## References

- Subword pooling makes a difference\.InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume,P\. Merlo, J\. Tiedemann, and R\. Tsarfaty \(Eds\.\),Online,pp\. 2284–2295\.External Links:[Link](https://aclanthology.org/2021.eacl-main.194/),[Document](https://dx.doi.org/10.18653/v1/2021.eacl-main.194)Cited by:[§2\.2](https://arxiv.org/html/2608.03452#S2.SS2.p2.1)\.
- G\. Astrach and Y\. Pinter \(2025\)Probing subphonemes in morphology models\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 12954–12961\.External Links:[Link](https://aclanthology.org/2025.findings-acl.672/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.672),ISBN 979\-8\-89176\-256\-5Cited by:[§2\.2](https://arxiv.org/html/2608.03452#S2.SS2.p1.1)\.
- Y\. Belinkov \(2022\)Probing classifiers: promises, shortcomings, and advances\.Computational Linguistics48\(1\),pp\. 207–219\.External Links:ISSN 0891\-2017,[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00422),[Link](https://doi.org/10.1162/coli_a_00422),https://direct\.mit\.edu/coli/article\-pdf/48/1/207/2006605/coli\_a\_00422\.pdfCited by:[§2\.2](https://arxiv.org/html/2608.03452#S2.SS2.p2.1)\.
- C\. Cappellaro, N\. Dumrukcic, I\. Fritz, F\. Franzon, and M\. Maiden \(2024\)The cognitive reality of morphomes\. evidence from Italian\.Morphology34\(1\),pp\. 33–71\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1007/s11525-023-09419-2)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p1.1)\.
- A\. Conneau, G\. Kruszewski, G\. Lample, L\. Barrault, and M\. Baroni \(2018\)What you can cram into a single $&\!\#\* vector: probing sentence embeddings for linguistic properties\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),I\. Gurevych and Y\. Miyao \(Eds\.\),Melbourne, Australia,pp\. 2126–2136\.External Links:[Link](https://aclanthology.org/P18-1198/),[Document](https://dx.doi.org/10.18653/v1/P18-1198)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.03452#S2.SS2.p1.1)\.
- R\. Cotterell, C\. Kirov, J\. Sylak\-Glassman, G\. Walther, E\. Vylomova, A\. D\. McCarthy, K\. Kann, S\. J\. Mielke, G\. Nicolai, M\. Silfverberg, D\. Yarowsky, J\. Eisner, and M\. Hulden \(2018\)The CoNLL–SIGMORPHON 2018 shared task: universal morphological reinflection\.InProceedings of the CoNLL–SIGMORPHON 2018 Shared Task: Universal Morphological Reinflection,Brussels,pp\. 1–27\.External Links:[Document](https://dx.doi.org/10.18653/v1/K18-3001),[Link](https://aclanthology.org/K18-3001)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p2.1)\.
- R\. Cotterell, C\. Kirov, J\. Sylak\-Glassman, G\. Walther, E\. Vylomova, P\. Xia, M\. Faruqui, S\. Kübler, D\. Yarowsky, J\. Eisner, and M\. Hulden \(2017\)CoNLL\-SIGMORPHON 2017 shared task: universal morphological reinflection in 52 languages\.InProceedings of the CoNLL SIGMORPHON 2017 Shared Task: Universal Morphological Reinflection,Vancouver,pp\. 1–30\.External Links:[Document](https://dx.doi.org/10.18653/v1/K17-2001),[Link](https://aclanthology.org/K17-2001)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p2.1)\.
- F\. Dalvi, A\. Nortonsmith, A\. Bau, Y\. Belinkov, H\. Sajjad, N\. Durrani, and J\. Glass \(2019\)One size does not fit all: comparing NMT representations of different granularities\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),pp\. 1–10\.External Links:[Document](https://dx.doi.org/10.18653/v1/N19-1080)Cited by:[§2\.2](https://arxiv.org/html/2608.03452#S2.SS2.p1.1)\.
- Y\. Elazar, S\. Ravfogel, A\. Jacovi, and Y\. Goldberg \(2021\)Amnesic probing: behavioral explanation with amnesic counterfactuals\.Transactions of the Association for Computational Linguistics9,pp\. 160–175\.Cited by:[Limitations](https://arxiv.org/html/2608.03452#Sx1.p1.1)\.
- H\. Harley and E\. Ritter \(2002\)Person and number in pronouns: a feature\-geometric analysis\.Language78\(3\),pp\. 482–526\.External Links:ISSN 1535\-0665,[Link](http://dx.doi.org/10.1353/lan.2002.0158),[Document](https://dx.doi.org/10.1353/lan.2002.0158)Cited by:[Figure 1](https://arxiv.org/html/2608.03452#S3.F1),[§3\.2](https://arxiv.org/html/2608.03452#S3.SS2.SSS0.Px2.p1.5)\.
- J\. Hewitt and P\. Liang \(2019\)Designing and interpreting probes with control tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 2733–2743\.External Links:[Link](https://aclanthology.org/D19-1275/),[Document](https://dx.doi.org/10.18653/v1/D19-1275)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p3.1),[§4\.1\.1](https://arxiv.org/html/2608.03452#S4.SS1.SSS1.Px2.p3.8),[§4\.1\.1](https://arxiv.org/html/2608.03452#S4.SS1.SSS1.Px3.p1.1)\.
- D\. Hupkes and W\. Zuidema \(2018\)Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure \(extended abstract\)\.InProceedings of the Twenty\-Seventh International Joint Conference on Artificial Intelligence, IJCAI\-18,pp\. 5617–5621\.External Links:[Document](https://dx.doi.org/10.24963/ijcai.2018/796),[Link](https://doi.org/10.24963/ijcai.2018/796)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.03452#S2.SS2.p1.1)\.
- A\. Kakolu Ramarao, K\. Tang, and D\. Baer\-Henney \(2025\)Frequency matters: modeling irregular morphological patterns in Spanish with transformers\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 4474–4489\.External Links:[Link](https://aclanthology.org/2025.findings-acl.230/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.230),ISBN 979\-8\-89176\-256\-5Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p2.1),[§3\.1](https://arxiv.org/html/2608.03452#S3.SS1.p1.1)\.
- A\. Kakolu Ramarao, K\. Tang, and D\. Baer\-Henney \(2026a\)Character\-aware transformers learn an irregular morphological pattern yet none generalize like humans\.InProceedings of the 15th Workshop on Cognitive Modeling and Computational Linguistics,B\. Oh, T\. Kuribayashi, G\. Rambelli, E\. Takmaz, P\. Wicke, J\. Li, and R\. Yoshida \(Eds\.\),Palma, Mallorca, Spain,pp\. 74–85\.External Links:[Document](https://dx.doi.org/10.63317/3ovkb8stpc2h)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p2.1),[§1](https://arxiv.org/html/2608.03452#S1.p3.1),[§3\.1](https://arxiv.org/html/2608.03452#S3.SS1.p1.1),[§3\.2](https://arxiv.org/html/2608.03452#S3.SS2.SSS0.Px3.p1.1),[§3\.2](https://arxiv.org/html/2608.03452#S3.SS2.p1.2),[§4\.3\.1](https://arxiv.org/html/2608.03452#S4.SS3.SSS1.p1.3),[§5](https://arxiv.org/html/2608.03452#S5.p4.1),[Ethics Statement](https://arxiv.org/html/2608.03452#Sx2.p1.1)\.
- A\. Kakolu Ramarao, K\. Tang, and D\. Baer\-Henney \(2026b\)Transformers over\-extend what humans underlearn: the case of the Spanish L\-shaped morphome\.Note:arXiv:2507\.21556External Links:[Link](https://arxiv.org/abs/2507.21556),[Document](https://dx.doi.org/10.48550/arXiv.2507.21556)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p2.1)\.
- K\. Kann, R\. Cotterell, and H\. Schütze \(2017\)Neural multi\-source morphological reinflection\.InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017, Valencia, Spain, April 3\-7, 2017, Volume 1: Long Papers,M\. Lapata, P\. Blunsom, and A\. Koller \(Eds\.\),pp\. 514–524\.External Links:[Document](https://dx.doi.org/10.18653/v1/e17-1049),[Link](https://doi.org/10.18653/v1/e17-1049)Cited by:[§3\.1](https://arxiv.org/html/2608.03452#S3.SS1.SSS0.Px1.p1.1)\.
- C\. Kirov and R\. Cotterell \(2018\)Recurrent neural networks in linguistic theory: revisiting pinker and prince \(1988\) and the past tense debate\.Transactions of the Association for Computational Linguistics6,pp\. 651–665\.External Links:[Link](https://aclanthology.org/Q18-1045/),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00247)Cited by:[§5](https://arxiv.org/html/2608.03452#S5.p4.1)\.
- J\. Kodner and S\. Khalifa \(2022\)SIGMORPHON–UniMorph 2022 shared task 0: modeling inflection in language acquisition\.InProceedings of the 19th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology,Seattle, Washington,pp\. 157–175\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.sigmorphon-1.18),[Link](https://aclanthology.org/2022.sigmorphon-1.18)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p2.1)\.
- D\. Liao and F\. Shi \(2026\)How tokenization limits phonological knowledge representation in language models and how to improve them\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 13921–13938\.External Links:[Link](https://aclanthology.org/2026.acl-long.634/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.634),ISBN 979\-8\-89176\-390\-6Cited by:[§2\.2](https://arxiv.org/html/2608.03452#S2.SS2.p2.1)\.
- N\. F\. Liu, M\. Gardner, Y\. Belinkov, M\. E\. Peters, and N\. A\. Smith \(2019\)Linguistic knowledge and transferability of contextual representations\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota,pp\. 1073–1094\.External Links:[Link](https://aclanthology.org/N19-1112/),[Document](https://dx.doi.org/10.18653/v1/N19-1112)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p3.1),[§2\.2](https://arxiv.org/html/2608.03452#S2.SS2.p1.1)\.
- M\. Maiden \(2018\)The romance verb: morphomic structure and diachrony\.Oxford University Press\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1093/oso/9780199660216.001.0001)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.03452#S2.SS1.p1.1)\.
- M\. Maiden \(2021\)The morphome\.Annual Review of Linguistics7,pp\. 89–108\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1146/annurev-linguistics-040220-042614)Cited by:[§2\.1](https://arxiv.org/html/2608.03452#S2.SS1.p1.1)\.
- A\. Nevins, C\. Rodrigues, and K\. Tang \(2015\)The rise and fall of the L\-shaped morphome: diachronic and experimental studies\.Probus: International Journal of Latin and Romance Linguistics27\(1\),pp\. 101–155\.External Links:[Document](https://dx.doi.org/10.1515/probus-2015-0002)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p1.1),[§3\.1](https://arxiv.org/html/2608.03452#S3.SS1.SSS0.Px1.p1.1)\.
- F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg, J\. Vanderplas, A\. Passos, D\. Cournapeau, M\. Brucher, M\. Perrot, and É\. Duchesnay \(2011\)Scikit\-learn: machine learning in python\.J\. Mach\. Learn\. Res\.12\(null\),pp\. 2825–2830\.External Links:ISSN 1532\-4435Cited by:[§4\.1\.1](https://arxiv.org/html/2608.03452#S4.SS1.SSS1.Px2.p1.2)\.
- S\. Prasada and S\. Pinker \(1993\)Generalisation of regular and irregular morphological patterns\.Language and Cognitive Processes \- LANG COGNITIVE PROCESS8,pp\. 1–56\.External Links:[Document](https://dx.doi.org/10.1080/01690969308406948)Cited by:[§5](https://arxiv.org/html/2608.03452#S5.p4.1)\.
- S\. Ravfogel, Y\. Elazar, H\. Gonen, M\. Twiton, and Y\. Goldberg \(2020\)Null it out: guarding protected attributes by iterative nullspace projection\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 7237–7256\.External Links:[Link](https://aclanthology.org/2020.acl-main.647/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.647)Cited by:[Limitations](https://arxiv.org/html/2608.03452#Sx1.p1.1)\.
- A\. Ravichander, Y\. Belinkov, and E\. Hovy \(2021\)Probing the probing paradigm: does probing accuracy entail task relevance?\.InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume,P\. Merlo, J\. Tiedemann, and R\. Tsarfaty \(Eds\.\),Online,pp\. 3363–3377\.External Links:[Link](https://aclanthology.org/2021.eacl-main.295/),[Document](https://dx.doi.org/10.18653/v1/2021.eacl-main.295)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p3.1)\.
- A\. Spencer and M\. Aronoff \(1994\)Morphology by itself: stems and inflectional classes\.Language70,pp\. 811\.External Links:[Document](https://dx.doi.org/10.2307/416331)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p1.1)\.
- I\. Tenney, D\. Das, and E\. Pavlick \(2019\)BERT rediscovers the classical NLP pipeline\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 4593–4601\.External Links:[Link](https://aclanthology.org/P19-1452/),[Document](https://dx.doi.org/10.18653/v1/P19-1452)Cited by:[§2\.2](https://arxiv.org/html/2608.03452#S2.SS2.p1.1)\.
- A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin \(2017\)Attention is all you need\.CoRRabs/1706\.03762\.External Links:[Link](http://arxiv.org/abs/1706.03762),1706\.03762Cited by:[§3\.2](https://arxiv.org/html/2608.03452#S3.SS2.p1.2)\.
- E\. Voita and I\. Titov \(2020\)Information\-theoretic probing with minimum description length\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 183–196\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.14/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.14)Cited by:[§4\.1\.1](https://arxiv.org/html/2608.03452#S4.SS1.SSS1.Px2.p3.8)\.
- S\. Wu, R\. Cotterell, and M\. Hulden \(2021\)Applying the transformer to character\-level transduction\.InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume,P\. Merlo, J\. Tiedemann, and R\. Tsarfaty \(Eds\.\),Online,pp\. 1901–1907\.External Links:[Link](https://aclanthology.org/2021.eacl-main.163),[Document](https://dx.doi.org/10.18653/v1/2021.eacl-main.163)Cited by:[§1](https://arxiv.org/html/2608.03452#S1.p2.1),[§3\.1](https://arxiv.org/html/2608.03452#S3.SS1.SSS0.Px1.p1.1)\.

## Appendix AL\-Shaped Test Lemmas by Split

Table[3](https://arxiv.org/html/2608.03452#A1.T3)lists the seven L\-shaped test lemmas of each of the three lemma splits, with the stem\-final alternation\.

Table 3:The seven L\-shaped test lemmas of each lemma split\. The alternation column gives the stem\-final segment outside the L\-shape pattern versus inside it\. The two /s/→\\rightarrow/s/ lemmas are L\-shaped in Spanish orthography \(*c/z*\) but disappears in the IPA transcription\.
## Appendix BLayerwise across architectures

![Refer to caption](https://arxiv.org/html/2608.03452v1/x2.png)Figure 3:The same distributions separately at each of the eight layers: linear\-probe balanced accuracy over the 12 trained models at every encoder and decoder layer, for all five architectures and all three properties\. Dashed lines mark the strongest surface baseline; dotted lines mark chance\.
## Appendix CCell clustering probe

Table 4:Cell\-clustering probe \(mean±\\pmstd over 12 models\): probability that a mood probe trained without1sg\.indclassifies held\-out1sg\.indinstances as subjunctive, split by lemma class, atdec2\.
## Appendix DL vs\. NL classification

Table 5:L\-shaped vs\. NL\-shaped classification \(best layer, linear probe, mean±\\pmstd over 12 models\) on the full instance set, the subset where the stem alternation is visible in the instance, and the subset where the stem\-final consonant is shared across all forms\. “Best surface” is the strongest n\-gram or n\-gram\-LM baseline for that subset\.
## Appendix EPositional readout

Table 6:Linear\-probe balanced accuracy on the hidden state at the stem\-final consonant position \(best layer, mean±\\pmstd over 12 models\): the surface\-cue property, morphome membership, and morphome membership restricted to the no\-alternation subset\. All values peak in the decoder exceptC\-Sep’s no\-alternation cell \(enc3\)\.

相似文章

Transformer模型中的词汇抽象泛化:功能词的案例

arXiv cs.CL

本文探讨了Transformer模型如何处理如代词和副词等功能词,分析了它们的嵌入表示和泛化能力。研究发现,在混合使用词汇化和功能句进行训练时,模型能学习到共享的句法-语义结构。