The Learning Objective Governs Perceptual Narrowing: A Cross-Lingual, Layer-Wise, Ten-Seed Study of Self-Supervised Speech Encoders

arXiv cs.CL Papers

Summary

A ten-seed study of self-supervised speech encoders shows that the learning objective (reconstruction vs. prediction) governs cross-lingual perceptual narrowing, with reconstruction degrading non-native phoneme discrimination and prediction improving it.

arXiv:2608.00507v1 Announce Type: new Abstract: Perceptual narrowing---the developmental loss of non-native phoneme discrimination in the first year of life \citep{werker1984}---is a canonical developmental finding, yet \emph{what learning objective produces it} remains open. We train a \(\sim\)7\,M-parameter Transformer encoder on child-directed and read speech and evaluate phoneme ABX in English, French, and Mandarin over ten seeds, the seed as the unit of replication. Six results. \textbf{(1)}~The objective sets the direction of cross-lingual transfer: reconstruction (masked mel-prediction) degrades non-native discrimination, prediction (frame-contrastive) improves it---a same-encoder, same-data gap of \(+0.051\) in first-layer Mandarin ABX (\(p=3\times10^{-8}\)), unanimous in sign across twenty runs. \textbf{(2)}~That decline combines a large arm-intrinsic difficulty gradient with a smaller language-specialization effect (matched vs.\ mismatched \(+0.022\), \(p=10^{-4}\), all four layers). \textbf{(3)}~Against a language-symmetric raw-mel floor, reconstruction pushes the first layer \emph{below} the discriminability of its input; prediction pushes it \emph{above}. \textbf{(4)}~Read speech gives a \(3.6\times\) steeper non-native decline than child-directed speech. \textbf{(5)}~The customary three-seed budget cannot see this reliably: an effect unambiguous at ten seeds is called significant by as few as 70\% of three-seed subsets. \textbf{(6)}~Six objective configurations---sharpening, compression, consolidation, their composition, and word-level semantic grounding in two forms---fail to produce the full developmental signature (native improves \emph{and} non-native declines): a single objective moves both languages the same way because it acts on a shared representation. We conclude that the objective, not the architecture, is the first-order determinant of narrowing-shaped representational change.
Original Article
View Cached Full Text

Cached at: 08/04/26, 07:42 AM

# The Learning Objective Governs Perceptual Narrowing: A Cross-Lingual, Layer-Wise, Ten-Seed Study of Self-Supervised Speech Encoders
Source: [https://arxiv.org/html/2608.00507](https://arxiv.org/html/2608.00507)
\(August 2026\)

###### Abstract

Perceptual narrowing—the developmental loss of non\-native phoneme discrimination in the first year of life\(Werker and Tees,[1984](https://arxiv.org/html/2608.00507#bib.bib20)\)—is a canonical developmental finding, yet*what learning objective produces it*remains open\. We train a∼\\sim7 M\-parameter Transformer encoder on child\-directed and read speech and evaluate phoneme ABX in English, French, and Mandarin over ten seeds, the seed as the unit of replication\. Six results\.\(1\)The objective sets the direction of cross\-lingual transfer: reconstruction \(masked mel\-prediction\) degrades non\-native discrimination, prediction \(frame\-contrastive\) improves it—a same\-encoder, same\-data gap of\+0\.051\+0\.051in first\-layer Mandarin ABX \(p=3×10−8p=3\\times 10^\{\-8\}\), unanimous in sign across twenty runs\.\(2\)That decline combines a large arm\-intrinsic difficulty gradient with a smaller language\-specialization effect \(matched vs\. mismatched\+0\.022\+0\.022,p=10−4p=10^\{\-4\}, all four layers\)\.\(3\)Against a language\-symmetric raw\-mel floor, reconstruction pushes the first layer*below*the discriminability of its input; prediction pushes it*above*\.\(4\)Read speech gives a3\.6×3\.6\\timessteeper non\-native decline than child\-directed speech\.\(5\)The customary three\-seed budget cannot see this reliably: an effect unambiguous at ten seeds is called significant by as few as 70% of three\-seed subsets\.\(6\)Six objective configurations—sharpening, compression, consolidation, their composition, and word\-level semantic grounding in two forms—fail to produce the full developmental signature \(native improves*and*non\-native declines\): a single objective moves both languages the same way because it acts on a shared representation\. We conclude that the objective, not the architecture, is the first\-order determinant of narrowing\-shaped representational change\.

## 1Introduction

By 6–12 months, infants lose the ability to discriminate phoneme contrasts absent from their native language while retaining native contrasts\(Werker and Tees,[1984](https://arxiv.org/html/2608.00507#bib.bib20);Kuhl et al\.,[1992](https://arxiv.org/html/2608.00507#bib.bib9)\)\. This perceptual narrowing is a canonical developmental\-neuroscience finding, and it poses a computational question that its neural description does not answer:what learning objective, operating on early speech input, produces experience\-driven loss of non\-native discrimination?

Self\-supervised learning \(SSL\) offers a controlled setting to ask it: a model with an explicit objective, trained on speech, whose internal discriminability can be read out at every layer\. The recent computational language\-acquisition literature has largely concluded that SSL speech models reproduce the*native\-language advantage*but not the developmental*decline*\(Schatz et al\.,[2021](https://arxiv.org/html/2608.00507#bib.bib19);Räsänen,[2026](https://arxiv.org/html/2608.00507#bib.bib17);Lavechin et al\.,[2025](https://arxiv.org/html/2608.00507#bib.bib10)\)\. We revisit that conclusion and find it under\-determined: the answer depends on measurement choices—the layer, the contrast set, the training language, the register, the seed count—that prior work has not varied together, and on the*objective*, which the literature treats as a fixed background rather than the variable of interest\.

Our central result is that the objective is the variable of interest\. A reconstruction objective and a prediction objective, holding the encoder, data, and seed identical, drive cross\-lingual transfer in opposite directions\. We dissect the reconstruction\-driven decline into two mechanisms, ground it in an absolute representational reference \(the input\-feature floor\), show its register\-sensitivity, and quantify how badly the field’s seed budget under\-powers it\. We test a brain\-circuit\-mapped architecture and find the objective, not the architecture, carries the effect\. Finally we search directly for the full signature with six objective configurations and map why none reproduces it\.

#### Contributions\.

\(1\) The learning objective sets the direction of cross\-lingual transfer \(ten seeds, unanimous\)\. \(2\) A training\-language crossover separating arm\-intrinsic difficulty from language specialization, with a contrast\-level confirmation\. \(3\) A representation\-geometry probe placing reconstruction below and prediction above the input\-feature floor\. \(4\) A register control and a seed\-budget fragility analysis, both methodological cautions with quantified effect\. \(5\) A brain\-circuit dual\-code region probe locating the negative architectural result\. \(6\) A divergence search: six objective configurations engineered to reproduce the full signature, a mechanism map showing why none does, and the selectivity requirement it implies\.

## 2Background

#### Perceptual narrowing and its neural description\.

Infants discriminate non\-native contrasts at birth and lose this by 10–12 months for contrasts absent from the native language, tested on exactly those contrasts \(Hindi retroflex for English learners; Mandarin tone for English learners\)\(Werker and Tees,[1984](https://arxiv.org/html/2608.00507#bib.bib20)\)\. The Native Language Magnet theory\(Kuhl et al\.,[1992](https://arxiv.org/html/2608.00507#bib.bib9)\)frames it as experience warping perceptual space\. The neural substrate is described by dual\-stream models of speech\(Hickok and Poeppel,[2007](https://arxiv.org/html/2608.00507#bib.bib8)\)and, in the verbal\-repetition circuit this project departs from, by a dual code—an articulatory representation \(left inferior frontal gyrus, LIFG\) and an acoustic\-phonetic representation \(left middle temporal gyrus, MTG\)\(Yoo,[2013](https://arxiv.org/html/2608.00507#bib.bib21)\)\. What none of these specify is the*learning objective*whose optimization produces the narrowing\.

#### Self\-supervised objectives\.

Two objective families dominate\.*Reconstruction*\(masked\-prediction, wav2vec2\-style;Baevski et al\.,[2020](https://arxiv.org/html/2608.00507#bib.bib1)\) predicts masked input detail from context, optimizing signal fidelity\.*Prediction*\(contrastive predictive coding, CPC;van den Oord et al\.,[2018](https://arxiv.org/html/2608.00507#bib.bib15)\) discriminates future frames via InfoNCE, optimizing invariant structure\.Räsänen\([2026](https://arxiv.org/html/2608.00507#bib.bib17)\)reviews their infant\-learning track record: the native advantage is reproduced, the decline is not\. The review treats the objective as fixed; we vary it\. Standard evaluation is final\-layer ABX, over all non\-native contrasts, under a single objective, at a small seed count; Sections[4\.2](https://arxiv.org/html/2608.00507#S4.SS2),[4\.3](https://arxiv.org/html/2608.00507#S4.SS3), and[4\.7](https://arxiv.org/html/2608.00507#S4.SS7)show each of these choices changes the conclusion\.

#### Prior computational tests of narrowing\.

Two lines bear most directly on ours\.Schatz et al\.\([2021](https://arxiv.org/html/2608.00507#bib.bib19)\)train models on realistic child\-centred audio and reproduce the native\-language advantage only partially and without a clean narrowing decline, arguing that phonetic*categories*need not be learned at all—perception may reorganise as a continuous space\.Millet and Dunbar\([2022](https://arxiv.org/html/2608.00507#bib.bib13)\)compare CPC, wav2vec2, and HuBERT against French\- and English\-listener perceptual spaces and report a small native\-language effect for CPC but a largely*language\-universal*space for the masked\-prediction models\. Both attach the objective to a model rather than isolating it \(one model per objective, full scale, whole\-model or final\-layer geometry\), and—read across the two objective families—their conclusions about which objective specializes do not obviously line up\. We make the objective the sole manipulated variable \(same encoder, data, and seed\) and read discriminability*change*layer by layer\. This is a different observable from their static perceptual\-space geometry: a model can raise non\-native ABX while still warping its space toward the native language, so our transfer\-direction result and their geometry results are complementary rather than directly commensurable\.

## 3Methods

#### Model\.

Single\-encoder Transformer: log\-mel→\\rightarrowLinear\(80,384\)→\\rightarrow4×\\timesTransformerEncoderLayer\(d=384, heads=6, ff=1536\)\. 7\.1 M parameters \(baseline\_mini\)\. A brain\-mappedDualCodeModel\(§[4\.8](https://arxiv.org/html/2608.00507#S4.SS8)\) shares the encoder scale and adds continuum\-memory\-system \(CMS;Behrouz et al\.,[2025](https://arxiv.org/html/2608.00507#bib.bib3)\) frequency\-gated regions mapped to STG/LIFG/dorsal/MTG\.

#### Data\.

Child\-directed speech: Providence \(CHILDES;Demuth et al\.,[2006](https://arxiv.org/html/2608.00507#bib.bib6)\), 176\.7 h\. Read speech: ZeroSpeech 2017 English/French/Mandarin 1 s clips\. All runs train 10 k steps at batch 4, 1 s crops==11\.1 h of exposure—4\.5×4\.5\\timesbelow the∼\\sim50 h at which large\-scale simulations first report robust phonemic ABX\(Lavechin et al\.,[2025](https://arxiv.org/html/2608.00507#bib.bib10)\)\.

#### Objectives\.

Selected by one config key so the encoder, data pipeline, and seeding are shared and only the top\-side supervision differs\. Reconstruction: mask 15% of mel frames, predict the original, MSE on masked positions\. Prediction: bilinear next\-frame InfoNCE, temperature 0\.1\.

#### Evaluation\.

Phoneme ABX\(Schatz et al\.,[2013](https://arxiv.org/html/2608.00507#bib.bib18)\), 5000 within\-context cross\-speaker triplets per arm \(English/French/Mandarin\), DTW\-aggregated per\-frame cosine\. Per\-triplet scores are stored so contrast\-level decomposition and uncertainty are recoverable without retraining\.

#### Statistics\.

Ten seeds for the objective comparison and the crossover; three for the geometry and dual\-code probes; single\-seed gates \(confirmed at additional seeds where cheap\) for the divergence search\. Theseed is the unit of replication: trajectory changes are aggregated across seeds with attinterval; within\-seed uncertainty uses a paired bootstrap over triplets\. One training run per seed yields all milestones\{0,1​k,5​k,10​k\}\\\{0,1\\mathrm\{k\},5\\mathrm\{k\},10\\mathrm\{k\}\\\}via checkpoints\. Reportedpp\-values are per\-test and uncorrected: the first\-order objective contrast \(Table[1](https://arxiv.org/html/2608.00507#S4.T1)\) survives any multiple\-comparison correction by orders of magnitude, while the borderline tests \(§[4\.3](https://arxiv.org/html/2608.00507#S4.SS3)contrast decomposition, §[4\.7](https://arxiv.org/html/2608.00507#S4.SS7)gap inversion\) should be read as nominal\.

#### Crossover\.

Train reconstruction on each of English/French/Mandarin read speech \(one register\), evaluate all three arms\.

## 4Results

### 4\.1The objective sets the direction of transfer

Encoder, data, and seed identical; only the objective varies\. At L1, over 10 k steps, native \(English\)Δ=−0\.006\\Delta=\-0\.006and non\-native \(Mandarin\)Δ=−0\.016\\Delta=\-0\.016under reconstruction, versus\+0\.019\+0\.019and\+0\.035\+0\.035under prediction\. The same\-loss contrast on the Mandarin arm—how much more it gains under prediction than reconstruction—is the paper’s strongest result \(Table[1](https://arxiv.org/html/2608.00507#S4.T1)\), significant at every layer and unanimous in sign across all ten seeds at every layer for both objectives \(40 of 40 layer\-condition cells\)\. The objective flips the direction of cross\-lingual transfer without exception \(Fig\.[1](https://arxiv.org/html/2608.00507#S4.F1)\)\. Neither objective reproduces the developmental signature: reconstruction gets the non\-native decline without native gain; prediction improves both arms, Mandarin more, from a lower start\.

Table 1:The objective contrast \(prediction−\-reconstruction\) on the Mandarin arm, by layer \(ten seeds, seed\-level 95% CI\)\.![Refer to caption](https://arxiv.org/html/2608.00507v1/x1.png)Figure 1:The learning objective sets the direction of cross\-lingual transfer \(L1, ten seeds\)\. Masked\-prediction \(left\): the native–non\-native gap grows via the Mandarin arm declining\. Frame\-contrastive \(right\): Mandarin crosses above English—the gap inverts\. Bands are seed\-level intervals\.
### 4\.2The decline is L1–L2 and the final layer inverts it

Reconstruction MandarinΔ\\Deltaby layer: L1−0\.016\-0\.016\(p=2×10−6p=2\\times 10^\{\-6\}\), L2−0\.014\-0\.014\(p=0\.003p=0\.003\), L3−0\.009\-0\.009\(n\.s\.\), L4\+0\.000\+0\.000\(n\.s\.\)\. The decline is a first\-half\-of\-network phenomenon, gone by L3 \(Fig\.[2](https://arxiv.org/html/2608.00507#S4.F2)\)\. At L1 the gap grows because the non\-native arm*declines*; at L4 a gap of similar size grows through parallel drift with no decline\. The field’s final\-layer ABX measures the second mechanism and misses the first\.

![Refer to caption](https://arxiv.org/html/2608.00507v1/x2.png)Figure 2:The masked\-prediction Mandarin decline is significant at L1–L2 \(∗\*\) and gone by L3–L4 \(n\.s\.\)\. Bars: seed\-level meanΔ\\Delta; whiskers: 95% CI\.
### 4\.3At the contrast level, the exclusively\-non\-native contrasts decline most

Decomposing the all\-contrasts Mandarin arm by whether a contrast exists in English \(a linguistic English\-absent set: tonal, retroflex, alveolo\-palatal, aspirated stops\), under reconstruction at L1: English\-shared−0\.014\-0\.014, English\-absent−0\.046\-0\.046\(difference−0\.032\-0\.032,p=0\.0017p=0\.0017\)\. This is not a floor effect—English\-absent starts*lower*\(0\.65 vs 0\.73\) yet declines more\. The contrasts that decline most are precisely those the infant paradigm tests\.

### 4\.4The crossover: two mechanisms, not one

Training reconstruction on each language and evaluating all three arms \(L1Δ\\Delta, ten seeds\) gives the matrix in Fig\.[4](https://arxiv.org/html/2608.00507#S4.F4)\.Arm\-intrinsic difficulty \(large\):the Mandarin arm declines under every training language, including Mandarin \(−0\.047\-0\.047\); pooled column means English−0\.015\-0\.015, French−0\.016\-0\.016, Mandarin−0\.051\-0\.051\.Language specialization \(smaller, significant\):pooled matched\-vs\-mismatched\+0\.022\+0\.022, significant at all four layers \(p=10−4p=10^\{\-4\}at L1–L3,p=10−3p=10^\{\-3\}at L4\); five of six per\-pair contrasts individually significant\.The decisive cell:training on Mandarin does not rescue the Mandarin arm—it still falls−0\.047\-0\.047\. The pattern is neither pure degradation \(specialization is significant\) nor pure specialization \(the hardest arm is not rescued\): reconstruction produces both\. At the contrast level \(Fig\.[4](https://arxiv.org/html/2608.00507#S4.F4)\) the extra decline of Mandarin\-specific contrasts halves as training moves toward Mandarin \(−0\.038→−0\.021\-0\.038\\rightarrow\-0\.021\) but does not vanish\.

![Refer to caption](https://arxiv.org/html/2608.00507v1/x3.png)Figure 3:Training\-language crossover \(L1Δ\\Delta\)\. The Mandarin column is uniformly negative \(arm\-intrinsic difficulty\); matched\-diagonal cells \(boxed\) fare best in their column \(specialization\)\.
![Refer to caption](https://arxiv.org/html/2608.00507v1/x4.png)Figure 4:The extra decline of Mandarin\-specific \(English\-absent\) contrasts by training language: steepest under English, roughly halved but not eliminated under Mandarin training\.

### 4\.5Reconstruction falls below its input floor; prediction rises above it

Against the raw\-mel ABX floor \(no encoder\), which is language\-symmetric \(native 0\.733, Mandarin 0\.731—no input\-level English advantage\), L1 ABX at 10 k: reconstruction native 0\.720 \(−0\.013\-0\.013vs floor\), Mandarin 0\.708 \(−0\.023\-0\.023\); prediction native 0\.742 \(\+0\.009\+0\.009\), Mandarin 0\.758 \(\+0\.026\+0\.026\)\. Reconstruction pushes L1 below the discriminability already in its input, more for non\-native; prediction pushes it above\. Effective rank*increases*under both \(reconstruction\+4\.6/\+4\.7\+4\.6/\+4\.7, prediction\+8\.5/\+7\.4\+8\.5/\+7\.4\), so this is not dimensional collapse but a reallocation of capacity\. The floor makes “degradation” and “improvement” absolute statements, and gives §[4\.1](https://arxiv.org/html/2608.00507#S4.SS1)its mechanism\.

### 4\.6The magnitude is register\-sensitive

Training reconstruction on read speech vs child\-directed speech \(English, both ten seeds\), the Mandarin decline is3\.6×3\.6\\timessteeper on read \(L1−0\.057\-0\.057vs−0\.016\-0\.016, difference−0\.041\-0\.041,p<10−4p<10^\{\-4\}\), while the native arm is register\-robust \(agreement within 0\.008\)\. The sign is register\-invariant; the magnitude is not\. This explains why a child\-directed\-only pilot saw a weak decline: CDS minimizes it\.

### 4\.7The three\-seed budget cannot see these effects reliably

Enumerating all\(103\)=120\\binom\{10\}\{3\}=120three\-seed subsets and re\-running the seed\-level test: the loss\-family contrast is called significant by 100% of subsets, the H1 Mandarin decline by 84%, but the gap inversion—also unambiguous at ten seeds \(p=8×10−6p=8\\times 10^\{\-6\}\)—by only 70%, subsetpp\-values spanning 0\.000–0\.209 \(Fig\.[5](https://arxiv.org/html/2608.00507#S4.F5)\)\. Our own original seed 0–2 sample landed atp=0\.073p=0\.073, in the 30% that miss\. The field’s customary seed budget under\-powers exactly the developmental effects it searches for\.

![Refer to caption](https://arxiv.org/html/2608.00507v1/x5.png)Figure 5:All three claims are significant at ten seeds; the bars show the fraction of three\-seed subsets that would also call each significant\. The gap inversion—significant atn=10n=10—is seen by only 70% of three\-seed draws\.
### 4\.8A brain\-circuit dual\-code architecture reproduces narrowing at no region

Probing all four Yoo \(2013\)\-mapped regions of a CMS\-gated DualCodeModel \(three seeds, reconstruction\): no region shows the native\-favoring narrowing gap \(every gap CI spans zero\)\. But the regions are not static—the fast LIFG \(articulatory\)*improves*\(\+0\.07\+0\.07to\+0\.09\+0\.09\) while the slow MTG \(acoustic\-phonetic\)*collapses*\(−0\.16\-0\.16to−0\.18\-0\.18\)\. The frequency gating produces a strong articulatory\-up / acoustic\-phonetic\-down dissociation, not a language specialization\. The dissertation maps MTG to the long\-term store, yet under reconstruction it degrades most—the slow region lags rather than consolidates, consistent with §[4\.5](https://arxiv.org/html/2608.00507#S4.SS5)\. So the architecture is far from inert: it produces a strong articulatory\-up / acoustic\-phonetic\-down dissociation, but that dissociation is not language\-selective: the*narrowing*\-shaped, native\-favouring effect is carried by the objective, not the region wiring\.

### 4\.9The divergence search: six configurations, none reproduce the signature

Results §[4\.1](https://arxiv.org/html/2608.00507#S4.SS1)–[4\.8](https://arxiv.org/html/2608.00507#S4.SS8)show the objective sets the*direction*but no objective produces the full developmental*signature*—native improves*and*non\-native declines \(divergence\)\. We searched for it directly with six objective configurations, each engineered to produce divergence and grounded in a distinct account of narrowing \(Table[2](https://arxiv.org/html/2608.00507#S4.T2), Fig\.[6](https://arxiv.org/html/2608.00507#S4.F6)\)\.

*Lexical top\-down*\(Feldman et al\.,[2013](https://arxiv.org/html/2608.00507#bib.bib7)\): a word\-supervised\-contrastive auxiliary lifts native\-word contrasts—but word discrimination is itself a discriminative objective, so it lifts both arms language\-agnostically\.*Perceptual magnet*\(Kuhl et al\.,[1992](https://arxiv.org/html/2608.00507#bib.bib9);Maye et al\.,[2002](https://arxiv.org/html/2608.00507#bib.bib11)\): a vector\-quantization codebook\(van den Oord et al\.,[2017](https://arxiv.org/html/2608.00507#bib.bib14)\)of native prototypes—but frame\-level quantization sinks both arms by collapsing sub\-phonemic detail\.*Self\-distillation*\(Baevski et al\.,[2022](https://arxiv.org/html/2608.00507#bib.bib2);Caron et al\.,[2021](https://arxiv.org/html/2608.00507#bib.bib4);McClelland et al\.,[1995](https://arxiv.org/html/2608.00507#bib.bib12)\): a data2vec EMA teacher supplies a native\-consolidated target—but the reconstruction\-like teacher\-prediction degrades native most\.*Composition*: lexical lift plus prototype compression cracks native only transiently before the compression drags it down\.*Semantic grounding*: regress each native word’s audio onto its GloVe embedding\(Pennington et al\.,[2014](https://arxiv.org/html/2608.00507#bib.bib16)\)—with a base it joins the both\-up family; alone it sinks both by collapsing phonetic detail to word\-meaning, which is coarser than phonetics\.

A single objective moves both languages the*same*way because it shapes a representation*shared*between them \(the §[4\.5](https://arxiv.org/html/2608.00507#S4.SS5)floor result in another guise\)\. Sharpening and word\-grounding lift both; compression, consolidation, and grounding\-alone sink both; composing a lift with a non\-selective compression gives a non\-selective net\. The signature requiresselectivity in the compression itself: to spare native, the objective must know which contrasts are native\-relevant—a signal finer than any single loss or word\-level target we tested carries\.

Table 2:The divergence\-search mechanism map: six configurations, L1Δ\\Delta\(seed\-0 gates, confirmed at further seeds where cheap\)\. None diverges \(native up*and*non\-native down\)\.![Refer to caption](https://arxiv.org/html/2608.00507v1/x6.png)Figure 6:No objective diverges: every configuration moves both arms the same way\. Divergence would be a native \(blue\) bar up beside a Mandarin \(orange\) bar down in the same group—absent everywhere\.

## 5Discussion

#### What produces narrowing: the objective\.

The results converge on one claim: the learning objective is the first\-order determinant of narrowing\-shaped representation change\. Reconstruction produces contrast\-specific, training\-language\-specific, layer\-localized loss of non\-native discriminability that falls below the input\-feature floor; prediction produces the opposite; the brain\-mapped architecture \(the dual code\) does not by itself carry the language\-selective effect—it yields an articulatory/acoustic dissociation instead\. For developmental neuroscience this is a testable dissociation: if early auditory learning approximates a reconstruction\-like objective, narrowing\-shaped loss should follow—most where the input register is most canonical \(§[4\.6](https://arxiv.org/html/2608.00507#S4.SS6)\); if a prediction\-like objective, it should not\. The dual\-stream mapping suggests a concrete hypothesis: the ventral acoustic\-phonetic stream \(MTG\) and the dorsal articulatory stream \(LIFG\) may implement different objectives, and §[4\.8](https://arxiv.org/html/2608.00507#S4.SS8)’s dissociation under a single objective is a first, indirect probe\.

#### The selectivity requirement, and the dorsal buffer\.

The divergence search turns a series of negatives into a positive statement: divergence needs*selectivity*—the force that lifts native must be distinct from the force that touches non\-native, which a single loss on a shared representation cannot supply\. This is not a scale problem \(prediction improves both arms*more*with more data, never diverging\); it is structural\. This is, we argue, the computational content of the Yoo \(2013\) dorsal buffer, whose role is semantic association—binding sound to meaning, a native\-relevance signal external to the acoustic objective\. Our tests sharpen*what*that signal must be: not a word\-contrastive push \(language\-agnostic\) nor a word\-meaning regression \(coarser than phonetics\), but something that reorganises phonetic categories toward native ones while sparing them—subtler than any single loss or word\-level target buildable at this scale\.

#### What we do not claim\.

No objective reproduces the full signature at 11\.1 h; the objective is necessary\-looking but not sufficient\. The gap inversion is significant at ten seeds but reported as secondary: it is the claim most sensitive to seed draw\. “Prediction” here is a simplified CPC without an autoregressive aggregator\. The divergence\-search gates are single\- to few\-seed with large, mechanistically distinct directions\.

#### Limitations, and what survives them\.

Four limitations bound the study, and it is worth being explicit about which touch the central claim and which touch only its scope\.*Scale*—11\.1 h/seed,4\.5×4\.5\\timesbelow the∼\\sim50 h at which large\-scale simulations first report robust phonemic ABX—cannot confound the first\-order result by construction, because the objective contrast is a same\-scale, same\-data, same\-seed difference in which scale cancels; it could overturn the conclusion only if the*sign*of the contrast reversed with more data, and the geometry result \(§[4\.5](https://arxiv.org/html/2608.00507#S4.SS5), a reallocation of capacity rather than an undertraining artifact\) gives no reason to expect that\. Scale therefore bounds the already\-negative*signature*claim, not the*direction*claim\.*Register entangled with corpus*\(Providence child\-directed vs ZeroSpeech read\) limits only the interpretation of the3\.6×3\.6\\timesmagnitude effect: the*sign*is register\- and corpus\-invariant \(§[4\.6](https://arxiv.org/html/2608.00507#S4.SS6)\), so the direction result holds within each corpus separately\.*Three\-seed geometry and dual\-code probes*are supporting rather than load\-bearing—the direction and crossover results are ten\-seed, and §[4\.7](https://arxiv.org/html/2608.00507#S4.SS7)directly measures that the load\-bearing contrast survives100%100\\%of three\-seed subsets—with one honest exception: the dual\-code “no region narrows” finding is a null atn=3n=3, the single place our defense rests on an argument \(a∼\\sim0\.02 language\-selective gap would surface in the majority of three\-seed draws\) rather than on measurement\.*Single architecture and language family*bounds generality, but the effect already reproduces across two architectures \(baseline and dual\-code\) and on a typologically distant, tonal non\-native target, and its mechanism—the input\-feature floor—is architecture\-agnostic in principle\. In sum, the direction result is immune to scale by design, robust to register and seed budget by measurement, and reproduced across two architectures; the one residual that remains an argument rather than a result is the three\-seed dual\-code null\.

#### Future work\.

These limitations convert cleanly into three experiments, each of which would turn a surviving argument into evidence: a ten\-seed re\-run of the geometry and dual\-code probes, to power the one null the defense currently rests on; a scale ladder \(11/45/176 h\) confirming the contrast keeps its sign and that prediction’s both\-arms gain*grows*rather than diverges; and a corpus\-matched register contrast de\-confounding the magnitude effect\. Beyond closing limitations, the substantive open lever for the full signature is a second, native\-relevance signal richer than a word label—utterance\-level visual grounding is its most natural form \(the dorsal\-buffer signal as such\)—since scale alone does not diverge and would have to be paired with a selective or grounded objective\. Per\-phoneme\-category geometry would localize where in the space the reallocation falls\.

## 6Conclusion

Perceptual narrowing is a developmental\-neuroscience phenomenon in search of a computational cause\. We show, in a controlled ten\-seed self\-supervised setting, that the learning objective is that cause at first order: reconstruction degrades non\-native phoneme discrimination—contrast\-, training\-language\-, and layer\-specifically, below the representation’s own input floor—while prediction improves it, a same\-encoder contrast of\+0\.051\+0\.051unanimous across twenty runs\. The decline is two mechanisms, its magnitude is register\-sensitive, and a brain\-circuit\-mapped architecture reproduces none of the language\-selective narrowing, yielding an articulatory/acoustic dissociation instead\. An effect unambiguous at ten seeds is missed by 30% of three\-seed studies\. And a direct search over six objective configurations shows that none reproduces the full signature, because an objective on a shared representation moves both languages together; the signature requires a selectivity finer than a word label, which our tests bound to utterance\-level grounding or a mechanism subtler than any single loss\. The objective sets the direction of narrowing; the second, native\-relevance signal is what would turn direction into the developmental signature, and that, not the architecture or the scale alone, is where the developmental question should be asked next\.

#### Reproducibility\.

Every result derives from public corpora and a fully specified setup\.*Data:*the Providence corpus of child\-directed speech \(CHILDES; 176\.7 h\) and the ZeroSpeech 2017 English/French/Mandarin read\-speech sets, with GloVe 6B \(50d\) word vectors for the semantic\-grounding probe—all publicly available\.*Model:*a 7\.1 M\-parameter single\-encoder Transformer \(log\-mel→\\rightarrowLinear\(80,384\)→\\rightarrow4×\\timesTransformerEncoderLayer\(dd=384, 6 heads, feed\-forward 1536\)\) and, for the architectural probe, a CMS frequency\-gated dual\-code variant of matched scale\.*Setup:*PyTorch on Apple\-silicon \(MPS\); per seed, 10 k optimiser steps at batch 4 over 1 s crops \(11\.1 h of exposure\), with fixed seeds and a single run per seed yielding all milestone checkpoints\{0,1​k,5​k,10​k\}\\\{0,1\\mathrm\{k\},5\\mathrm\{k\},10\\mathrm\{k\}\\\}\.*Evaluation:*phoneme ABX over 5000 within\-context, cross\-speaker triplets per arm, scored by DTW\-aggregated per\-frame cosine, with per\-triplet scores retained so that every contrast\-level decomposition and uncertainty estimate is recomputable without retraining\.

## References

- Baevski et al\. \(2020\)A\. Baevski, Y\. Zhou, A\. Mohamed, and M\. Auli\. wav2vec 2\.0: A framework for self\-supervised learning of speech representations\.*NeurIPS*, 2020\.
- Baevski et al\. \(2022\)A\. Baevski, W\.\-N\. Hsu, Q\. Xu, A\. Babu, J\. Gu, and M\. Auli\. data2vec: A general framework for self\-supervised learning in speech, vision and language\.*ICML*, 2022\.
- Behrouz et al\. \(2025\)A\. Behrouz, M\. Razaviyayn, P\. Zhong, and V\. Mirrokni\. Nested learning: The illusion of deep learning architectures\.*NeurIPS*, 2025\.
- Caron et al\. \(2021\)M\. Caron, H\. Touvron, I\. Misra, H\. Jégou, J\. Mairal, P\. Bojanowski, and A\. Joulin\. Emerging properties in self\-supervised vision transformers \(DINO\)\.*ICCV*, 2021\.
- Caucheteux et al\. \(2023\)C\. Caucheteux, A\. Gramfort, and J\.\-R\. King\. Evidence of a predictive\-coding hierarchy in the human brain listening to speech\.*Nature Human Behaviour*, 7, 2023\.
- Demuth et al\. \(2006\)K\. Demuth, J\. Culbertson, and J\. Alter\. Word\-minimality, epenthesis, and coda licensing in the acquisition of English\.*Language and Speech*, 49\(2\), 2006\.
- Feldman et al\. \(2013\)N\. H\. Feldman, T\. L\. Griffiths, and J\. L\. Morgan\. A role for the developing lexicon in phonetic category acquisition\.*Psychological Review*, 120\(4\), 2013\.
- Hickok and Poeppel \(2007\)G\. Hickok and D\. Poeppel\. The cortical organization of speech processing\.*Nature Reviews Neuroscience*, 8, 2007\.
- Kuhl et al\. \(1992\)P\. K\. Kuhl, K\. A\. Williams, F\. Lacerda, K\. N\. Stevens, and B\. Lindblom\. Linguistic experience alters phonetic perception in infants by 6 months of age\.*Science*, 255, 1992\.
- Lavechin et al\. \(2025\)M\. Lavechin, M\. de Seyssel, H\. Titeux, G\. Wisniewski, H\. Bredin, A\. Cristia, and E\. Dupoux\. Simulating early phonetic and word learning without linguistic categories\.*Developmental Science*, 28:e13606, 2025\.
- Maye et al\. \(2002\)J\. Maye, J\. F\. Werker, and L\. Gerken\. Infant sensitivity to distributional information can affect phonetic discrimination\.*Cognition*, 82\(3\), 2002\.
- McClelland et al\. \(1995\)J\. L\. McClelland, B\. L\. McNaughton, and R\. C\. O’Reilly\. Why there are complementary learning systems in the hippocampus and neocortex\.*Psychological Review*, 102\(3\), 1995\.
- Millet and Dunbar \(2022\)J\. Millet and E\. Dunbar\. Do self\-supervised speech models develop human\-like perception biases?*Proceedings of ACL*, 2022\.
- van den Oord et al\. \(2017\)A\. van den Oord, O\. Vinyals, and K\. Kavukcuoglu\. Neural discrete representation learning \(VQ\-VAE\)\.*NeurIPS*, 2017\.
- van den Oord et al\. \(2018\)A\. van den Oord, Y\. Li, and O\. Vinyals\. Representation learning with contrastive predictive coding\.*arXiv:1807\.03748*, 2018\.
- Pennington et al\. \(2014\)J\. Pennington, R\. Socher, and C\. D\. Manning\. GloVe: Global vectors for word representation\.*EMNLP*, 2014\.
- Räsänen \(2026\)O\. Räsänen\. Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors\. In*Advances in Child Development and Behavior*, vol\. 70\. Elsevier, 2026\. arXiv:2603\.08359\.
- Schatz et al\. \(2013\)T\. Schatz et al\. Evaluating speech features with the minimal\-pair ABX task\.*INTERSPEECH*, 2013\.
- Schatz et al\. \(2021\)T\. Schatz, N\. H\. Feldman, S\. Goldwater, X\.\-N\. Cao, and E\. Dupoux\. Early phonetic learning without phonetic categories: Insights from large\-scale simulations on realistic input\.*Proceedings of the National Academy of Sciences*, 118\(7\):e2001844118, 2021\.
- Werker and Tees \(1984\)J\. F\. Werker and R\. C\. Tees\. Cross\-language speech perception: Evidence for perceptual reorganization during the first year of life\.*Infant Behavior and Development*, 7, 1984\.
- Yoo \(2013\)S\. Yoo\.*Neural mechanism of verbal repetition: From sounds to speech\.*PhD dissertation, Seoul National University, 2013\.

Similar Articles

Representation Without Reward: A JEPA Audit for LLM Fine-Tuning

arXiv cs.LG

This paper audits Joint-embedding predictive architectures (JEPA) for LLM fine-tuning on a natural-language-to-regex task, testing twenty-two auxiliary objectives. The results show that hidden-state representation improvements are only weakly coupled to decoded-task accuracy, with no auxiliary surviving family-wise correction.