Decoding silent reading from non-invasive EEG

Hacker News Top Papers

Summary

This research uses non-invasive EEG and contrastive learning to decode words during silent reading, showing scalable lexical information recovery that scales with data volume.

No content available
Original Article
View Cached Full Text

Cached at: 08/23/26, 10:42 PM

# Decoding silent reading from non-invasive EEG
Source: [https://arxiv.org/html/2608.20186](https://arxiv.org/html/2608.20186)
Anthilia AlchanatPriyanka JainAffiliation:\[0\.5em\] nubrain

18 August 2026

###### Abstract

Non\-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person’s spontaneous inner monologue cannot be collected, and the available proxy paradigms \(cued*repetitive*and retrospectively reported*generative*inner speech\) are slow to acquire, poorly time\-locked, and subject compliance is unverifiable\. We therefore treat silent reading as a scalable proxy task and ask how much lexical and semantic information a contrastive decoder can extract from it\. We report an open\-vocabulary analysis of approximately 240,000 word presentations recorded from a single densely\-sampled participant across 393 runs \(ca\. 49 h\) of 19\-channel dry\-electrode EEG\. Words from continuous narrative text were presented in rapid serial visual presentation, with typography randomised on every trial to partially decorrelate word identity from low\-level visual form\. A convolutional EEG encoder, optionally followed by a causal transformer, was trained with a CLIP\-style contrastive objective to align short EEG windows with hidden\-state embeddings of the presented word taken from a large language model\. Decoding, evaluated as word\-grouped top\-10 retrieval against permutation baselines, was reliably above chance, extended to mid\-frequency and rare words, and scaled log\-linearly with training\-data volume with no sign of saturation\. Removing occipital and posterior\-temporal electrodes reduced the word\-level gain by roughly one third but left context tracking unchanged\. Control analyses separate word\-level decoding from narrative context tracking and from a non\-neural positional prior introduced by the transformer’s positional embedding\. These results establish that open\-vocabulary word\-level information is recoverable from EEG during silent reading, and that decoding is data\-limited rather than saturated\.

## 1Introduction

### 1\.1Motivation

Loss of speech following stroke, traumatic brain injury, amyotrophic lateral sclerosis or other neurodegenerative disease removes the most basic channel of human communication\. Brain\-computer interfaces \(BCIs\) that decode intended or imagined speech directly from neural activity have therefore been a long\-standing clinical goal, and the last five years have delivered striking progress, but mostly using invasive techniques\. Intracortical and electrocorticographic systems now decode attempted speech at conversational rates and near\-perfect accuracy over vocabularies of tens of thousands of words\([2](https://arxiv.org/html/2608.20186#bib.bib2);[17](https://arxiv.org/html/2608.20186#bib.bib17);[18](https://arxiv.org/html/2608.20186#bib.bib18)\), and recent work has shown that*inner*speech — imagined speech with no attempted articulation — is robustly represented in motor cortex as a scaled\-down version of attempted speech occupying a shared neural subspace, permitting real\-time decoding of imagined sentences from a 125,000\-word vocabulary\([11](https://arxiv.org/html/2608.20186#bib.bib11)\)\. Single\-neuron recordings in supramarginal gyrus tell a similar story, with internal speech decodable online and with representations shared across reading, listening, and vocalised speech\([22](https://arxiv.org/html/2608.20186#bib.bib22)\)\.

These results settle the scientific question of whether inner speech can be decoded from measured brain activity\. What they do not settle is whether that structure is accessible without neurosurgery\. Invasive devices carry surgical risk, have limited chronic longevity, and will not scale to the population of people who would benefit from restored or augmented communication\. The open question is whether non\-invasive recordings, and specifically EEG, can be pushed far enough to be useful\.

### 1\.2The state of non\-invasive language decoding

The non\-invasive language decoding literature divides into two distinct approaches with very different outcomes\.

The first approach decodes language that the participant*perceives*\.[7](https://arxiv.org/html/2608.20186#bib.bib7)trained a convolutional brain module with a contrastive \(CLIP\) objective to align M/EEG with self\-supervised speech representations from wav2vec 2\.0, and identified the correct 3\-second speech segment out of more than a thousand candidates with up to 41% top\-1 accuracy from MEG\. However, their EEG results were far weaker \(top\-10 accuracy of 17\.7% and 25\.7% on two datasets\)\.[21](https://arxiv.org/html/2608.20186#bib.bib21)recorded a staggering 175 h of EEG from a single participant reading aloud, and reached 48\.5% top\-1 and 76\.0% top\-10 classification accuracy over 512 candidate segments, and, more importantly than the headline number, demonstrated an unsaturated log\-linear scaling relationship between data volume and decoding accuracy\. When they subsampled their data to the∼3\{\\sim\}3h typical of the field, accuracy collapsed to near the levels previously reported\. Their conclusion, which motivates the present study, is that the binding constraint on EEG language decoding may be dataset size rather than a hard ceiling on signal quality\.

The second approach to non\-invasive language decoding attempts to decode*inner*speech directly, and the results are markedly worse\.[5](https://arxiv.org/html/2608.20186#bib.bib5)collected an unusually large single\-participant EEG and MEG dataset across three paradigms — silent reading, repetitive inner speech, and generative inner speech — using a five\-word vocabulary chosen for clinical relevance\. Silent reading decoded at 30–40% \(chance 20%\) across EEG, MEG and optically\-pumped magnetometers, with permutation feature importance localising the effect to visual cortex at roughly 150 ms\. Both inner\-speech paradigms were essentially at chance, across a wide range of decoding methods, and the failure prevented any test of transfer from reading to inner speech\. Earlier EEG studies of imagined speech report accuracies that are statistically above chance but far from usable, typically on small vocabularies and few electrodes\([3](https://arxiv.org/html/2608.20186#bib.bib3);[14](https://arxiv.org/html/2608.20186#bib.bib14)\), with occasional outliers whose paradigms invite confounds\([1](https://arxiv.org/html/2608.20186#bib.bib1)\)\. Direct MEG investigations of imagined phrases have reported high accuracies\([6](https://arxiv.org/html/2608.20186#bib.bib6)\), but with paradigms in which imagination immediately follows perception of the same phrase within a trial, leaving residual perceptual and motor\-preparatory activity as plausible contributors\.

### 1\.3The labelled\-data problem for inner monologue

The gap between these two approaches to non\-invasive language decoding is instructive\. Where large volumes of well\-time\-locked data are available \(as in perception and overt production\), non\-invasive decoding works, imperfectly but measurably\. Where they are not \(spontaneous inner speech\), decoding does not work\.

This discrepancy reflects a fundamental obstacle\. The ideal training corpus for an inner\-monologue decoder would pair short EEG segments with a transcription of what the person was thinking during that segment\. Such a corpus cannot exist\. To know a person’s inner monologue at a precise point in time, one needs a device that reads minds; to build that device one needs the corpus\. Every existing paradigm is a way of partially escaping this circle, and each pays a specific price\.

*Generative inner speech*, in which participants freely imagine an item and report it afterwards\([10](https://arxiv.org/html/2608.20186#bib.bib10);[5](https://arxiv.org/html/2608.20186#bib.bib5)\), preserves the endogenous character of the thought but destroys its timing\. Even a truthful report of having thought “I like bananas” does not tell us whether the thought occurred 400 ms or 1,500 ms before the report; a fatal ambiguity for a signal whose informative structure lies on a scale of tens of milliseconds\. It also demands sustained, unverifiable compliance: there is no objective check on whether a participant was on task, mind\-wandering, or reporting selectively\. Data quality and data quantity both suffer\.

*Repetitive inner speech*, in which participants are instructed which item to imagine and cued when to imagine it, restores timing but reintroduces the cue\. The recorded response contains the neural processing of the cue as well as the imagined item\. Careful designs mitigate this\. Presenting the item briefly, waiting a jittered interval, then delivering a content\-neutral go signal reduces sensory contamination, but does not eliminate it\. \(Perception\-based proxy tasks, like the one in the present study, are obviously also affected by sensory signals, as discussed in Section[4\.5](https://arxiv.org/html/2608.20186#S4.SS5)\.\)

But in our view, even more important than a lack of precise temporal information \(in generative inner speech\) and sensory contamination \(in repetitive inner speech\) is the fact that these paradigms impose a slow trial structure\. Slow, repetitive trials are boring, bored participants disengage, and disengagement is undetectable in the data\. The paradigm is therefore doubly limited: it does not scale, and its labels degrade in exactly the way that cannot be audited\. Neither approach can plausibly produce the tens or hundreds of hours per participant that the scaling results of[21](https://arxiv.org/html/2608.20186#bib.bib21)suggest are necessary\.

### 1\.4Our approach: reading \(and, later, listening\) as a scalable proxy

Any language decoding paradigm must trade off three objectives that pull against each other:

1. 1\.Validity: The task must induce language processing and activate semantic representations\. For example, a task that can be performed by paying attention to visual properties of word stimuli does not guarantee semantic processing\.
2. 2\.Scalability: The paradigm must cover a large amount of linguistic content per hour and remain viable over several hours\.
3. 3\.Compliance: The task must be engaging enough to sustain attention, easy enough to follow without excessive mental effort, and instrumented so that lapses can be detected\.

We concluded that the best available compromise for obtaining a training dataset for language decoding is a multimodal paradigm combining silent reading and passive listening\. When a word is read, early visual cortex first encodes letter shapes and positions\. Once the word form is recognised, higher\-level areas are engaged in encoding its meaning\. When the same word is heard, auditory cortex first encodes the phonetic structure, but the same semantic representation is ultimately evoked\. Reading and listening thus share a representational target while differing maximally in their input\-driven, modality\-specific components\. A decoder trained across both, or evaluated for transfer between them, offers a remedy against the confound that limits silent\-reading paradigms, namely that the decodable signal may be mostly driven by visual word\-form processing\([5](https://arxiv.org/html/2608.20186#bib.bib5);[13](https://arxiv.org/html/2608.20186#bib.bib13)\)\. Evidence for modality\-general semantic codes is well established\. Language\-independent semantic representations in anterior temporal lobe generalise across the two languages of bilingual listeners\([4](https://arxiv.org/html/2608.20186#bib.bib4)\), and fMRI classifiers trained on language perception successfully decode the content of language production\([12](https://arxiv.org/html/2608.20186#bib.bib12)\)\.

The present report covers only the silent\-reading half of the multimodal language decoding design\. The passive\-listening condition is planned for a follow\-up study, and a cross\-modal transfer analysis is therefore not yet possible\.

Relative to previous non\-invasive silent\-reading work, our design differs in three respects that matter for the interpretation of the results\. First, the vocabulary is open and natural\. Participants read continuous narrative prose rather than a small fixed set of clinically\-motivated words, so the model must operate over tens of thousands of word types rather than a handful\. Second, typography is randomised trial\-by\-trial \(font, size, colour, character spacing\), directly addressing the concern raised by[5](https://arxiv.org/html/2608.20186#bib.bib5)that low\-level visual form is confounded with word identity when each word is always rendered identically\. Third, the data volume is large enough to address the scaling question\.

### 1\.5Aims of this report

This report presents results from the first stage of an ongoing research project\. The present results are restricted to a single densely\-sampled participant\. We have separately collected one or two sessions of silent reading data across 60 participants, and plan to present cross\-subject results and multimodal results \(including a passive listening condition\) in future reports\. In the present report, we specifically ask:

1. 1\.Is word\-level information decodable at all from EEG during silent reading, in an open\-vocabulary contrastive setting?
2. 2\.How much of any apparent decoding is context tracking, i\.e\. following the ongoing narrative topic, rather than identifying the currently presented word? This is the central interpretive risk when target representations come from a causal language model, and we address it directly\. A related risk is that a sequence model’s learned positional embedding acts as a positional prior on contextual targets\. \(The positional embedding is the only trial\-varying, non\-neural input to the model\.\) We measure this contribution directly with masking probes\.
3. 3\.Is the signal confined to frequent \(largely function\) words, or does it extend to mid\-frequency and rare content words?
4. 4\.How does performance scale with training\-data volume?
5. 5\.Which analysis and optimisation choices maximise decoding accuracy, since the practical motivation for this work is BCI development rather than a purely descriptive characterisation\.

![Four-panel overview of the experimental paradigm, recording montage, decoding architecture and scoring procedure](https://arxiv.org/html/2608.20186v1/figures/figure_01_paradigm_and_pipeline.png)Figure 1:Paradigm and decoding pipeline\.Full caption on the following page\.Figure[1](https://arxiv.org/html/2608.20186#S1.F1)\(preceding page\): Paradigm and decoding pipeline\.\(a\)Words from continuous narrative prose are presented one at a time at the centre of a black screen \(rapid serial visual presentation\)\. Typography \(colour, font, size, and character spacing\) is drawn randomly on every trial, so that word identity is decorrelated from low\-level visual form\. Each word remains on screen for 500–900 ms, scaled by character count plus up to 100 ms of jitter\. The inter\-stimulus interval was 100 ms for the deep participant’s first 171 runs and 0 ms thereafter\. Runs contain∼600\{\\sim\}600words, and the narrative continues from one run to the next\.\(b\)Recording montage: 19 dry scalp electrodes in the 10\-20 layout, sampled at 600 Hz \(Supplementary[S0](https://arxiv.org/html/2608.20186#A0)\)\.\(c\)Decoding architecture\. A fixed\-length EEG window passes through a dual\-pathway convolutional encoder — a global pathway of convolutional blocks collapsed by learnable attention pooling, and a local pathway preserving six coarse temporal bins — and a fully\-connected head, giving a 256\-dimensional per\-trial feature\. An optional four\-layer causal transformer then operates over the sequence of features within a recording run\. Separately, the presented word is mapped to a hidden state of Llama\-3\.1\-8B, taken either at layer 0 \(non\-contextual: the input embedding, which depends only on word identity\) or at layer 20 \(contextual: dependent on all preceding text\), and projected linearly\. The two pathways meet only in a shared 256\-dimensional L2\-normalised space, where they are trained with the symmetric CLIP contrastive objective; the target pathway never sees the EEG\. The presence of the transformer and the choice of embedding layer \(0 or 20\) define the four configurations compared throughout the Results\.\(d\)How scores are computed\. Within a fixed pool of 512 validation trials, the model’s EEG prediction is scored against every candidate target, softmaxed, and probability mass is summed across trials sharing a core word \(case\-folded, punctuation stripped\) before the rank of the true word is read\. The metric reported throughout the present paper is the resulting top\-10 accuracy*minus*an empirical permutation baseline, so chance is always zero\.

## 2Methods

### 2\.1Participants and data collected

One contributor has completed 393 recording runs, each with different text\. The text read by the participant was drawn from a corpus of fictional books\. In total, the participant contributed 240,141 word presentations, corresponding to 48\.7 h of on\-task recording\. The number of trials passing the amplitude criterion described in Section[2\.4](https://arxiv.org/html/2608.20186#S2.SS4)depends on the analysis window \(roughly 89–93%\) and is therefore a property of the analysis configuration rather than of the dataset \(Supplementary[S8](https://arxiv.org/html/2608.20186#A8)\)\.

The present report analyses data from the aforementioned deep subject\. Additionally, we have collected silent reading data from 60 participants across one or two sessions each \(data collection ongoing\)\. In this first report, we deliberately focus on fundamental questions \(can language information be decoded non\-invasively at all?\) that are best addressed with a large single\-subject dataset\. Cross\-subject model training introduces additional complexity that will be addressed separately in a future report\. Here, the goal is generalisation to unseen EEG trials and unseen text \(but not unseen words\)\.

### 2\.2Stimuli and task

Text was drawn from fictional books \(Sherlock Holmes stories\), chosen to be enjoyable, engaging, yet not too cognitively demanding\. We think that in the context of language decoding, engaging stimulus material is not just a nice\-to\-have, but essential, to facilitate subject compliance and continual deep semantic processing of the stimuli\. Words were presented one at a time at the centre of a black screen \(rapid serial visual presentation\)\. Each run contained approximately 600 words; the exact count varies slightly because text was cut at paragraph boundaries\. Within a session, the narrative progressed continuously from run to run, so that the text of runn\+1n\{\+\}1continues the text of runnn\.

To decorrelate word identity from low\-level visual form, typography was randomised independently on every trial: 7 colours, 37 font families, 3 font sizes, and continuously varying character spacing\. This directly implements the recommendation of[5](https://arxiv.org/html/2608.20186#bib.bib5), who noted that using a single fixed rendering per word facilitates visual form as a confound in silent\-reading decoding\. Each word remained on screen for 500–900 ms, scaled by word length \(character count\) plus up to 100 ms of random jitter\. Specifically, words with up to nine characters were shown for 500 ms\. Words with ten characters or more were shown longer \(10 characters: 600 ms, 11 characters: 700 ms, 12 characters: 800 ms, 13 or more characters: 900 ms\)\. At the beginning of each run, a black screen was shown for 3 seconds\. After every 100 words, there was a break \(black screen\) for 6 seconds\.

Two protocol changes occurred early in data collection and are treated as nuisance factors in the analyses:

- •Attention task\.The first 184 runs used a one\-back task in which the participant pressed a button when a word repeated\. This was replaced by four multiple\-choice comprehension questions at the end of each run, which we judged a better guarantee of sustained attention\. In all analyses reported here, target events \(the second occurrence of a repeated word\) and behavioural false alarms are excluded at the trial level, so repeated words never enter the model\.
- •Inter\-stimulus interval \(ISI\)\.Words were initially separated by a 100 ms blank screen\. Participant feedback during pilot experiments indicated that most participants found text comprehension difficult at ISI=100=100ms, possibly resulting from attentional masking, so the ISI was set to 0 ms \(no blank screen between words\) and the affected pilot participants were excluded\. The deep participant \(on whose data the present report is based\) did not report this difficulty, possibly due to genuine inter\-individual variability, or caused by a training effect, so their first 171 runs, recorded at ISI=100=100ms, are retained\. All subsequent runs used ISI=0=0ms\. ISI was therefore included as an explicit factor in the first model training sweep \(Section[2\.10](https://arxiv.org/html/2608.20186#S2.SS10)\)\.

### 2\.3EEG acquisition

EEG was recorded from 19 scalp electrodes using a dry\-electrode system, at a sampling rate of 600 Hz\. For details, see Supplementary[S0](https://arxiv.org/html/2608.20186#A0)\.

### 2\.4Preprocessing

Preprocessing was deliberately minimal, following the finding of[7](https://arxiv.org/html/2608.20186#bib.bib7)that end\-to\-end architectures gain little from elaborate M/EEG artefact pipelines\. Per recording run, each channel was linearly detrended, band\-stop filtered around the mains frequency \(50 Hz±1\\pm 1Hz, 4th\-order Butterworth\), and band\-pass filtered between 1 and 80 Hz \(4th\-order Butterworth\)\.

Trials were then extracted as fixed\-length windows\. The window length has to be fixed, even though the trial duration was not\. Because presentation duration scales with word length, variable\-length epochs would let the decoder read word length directly off the epoch, a non\-neural cue that would inflate accuracy\. We tested the effect of locking the EEG time window to stimulus onset, stimulus offset, or the centre of the stimulus interval\. Moreover, window length was varied across sweeps \(0\.5, 0\.7, 0\.9 s\)\. Optionally, a random temporal shift of up to±50\\pm 50or±100\\pm 100ms was applied to the analysis window during training only, as a data\-augmentation strategy\.

Artefact rejection was a single amplitude criterion: a trial was excluded if any sample within its \(augmentation\-extended\) window exceeded 100μ\\muV in absolute value\. Amplitude checking was performed over the full extended segment so that temporal\-shift augmentation can never bring a rejected artefact into view\. Included trials were scaled by dividing by 100, so that the working unit is 100μ\\muV and values lie mainly in\[−1,1\]\[\-1,1\], an approach also adopted by recent EEG foundation models\([9](https://arxiv.org/html/2608.20186#bib.bib9)\)\. An alternative normalisation that additionally subtracted a pre\-stimulus baseline was tested in the first sweep\. Excluded trials \(artefacts, target events, false alarms\) are retained as placeholders in the trial sequence so that positional alignment with the language\-model targets is preserved, but are masked out of the loss and of all evaluation \(Section[2\.6](https://arxiv.org/html/2608.20186#S2.SS6)\)\.

### 2\.5Target representations

Each word was represented by a hidden state extracted from a pretrained Llama\-3\.1\-8B model \([https://huggingface\.co/meta\-llama/Llama\-3\.1\-8B](https://huggingface.co/meta-llama/Llama-3.1-8B);[8](https://arxiv.org/html/2608.20186#bib.bib8)\)\. The full text of a run was tokenised, a forward pass was run, and for each word the hidden state of the last sub\-word token of its “core” form \(leading and trailing punctuation stripped\) was taken, at a specified layer\. For example, consider a trial in which a subject read the word “corkscrew\.” \(including a full stop, because the word appeared at the end of a sentence\)\. To retrieve the target embedding for this word, punctuation was stripped, resulting in the core word “corkscrew”, which the Llama\-3\.1 tokenizer represents as three tokens: “Ġc”, “orks”, and “crew”\. In this case, the chosen target embedding is that of the last sub\-token, i\.e\. “crew”\. Because the model is causal, a hidden state at any layer above the input embedding incorporates all preceding words\.

Two layers were contrasted throughout:

- •Layer 0 — non\-contextual\.The input embedding\. Depends only on the word’s identity, not on the preceding text\. A decoder trained against layer\-0 targets cannot profit from tracking the narrative, because the target carries no narrative information\. \(In this case, for multi\-token words, only the last sub\-word token determines the embedding\.\)
- •Layer 20 — contextual\.A mid\-depth hidden state, semantically rich and heavily context\-dependent\. Targets for consecutive words within a run are strongly correlated\. A decoder trained against these targets can profit from tracking the discourse state as well as from identifying the current word\. \(Llama\-3\.1\-8B has 32 layers; we chose layer 20 rather than the final layer’s state because the latter presumably represents expectations about the next token to a larger degree than a middle layer\.\)

Target embeddings are 4,096\-dimensional and were projected linearly \(without bias\) into a shared, 256\-dimensional decoding space\.

### 2\.6Model and training objective

The decoder has three components \(Figure[1](https://arxiv.org/html/2608.20186#S1.F1)c; full specification in Supplementary[S2](https://arxiv.org/html/2608.20186#A2)\)\.

EEG encoder\.A per\-trial one\-dimensional convolutional network with two parallel pathways\. A*global*pathway applies four convolutional blocks \(48→96→192→38448\\rightarrow 96\\rightarrow 192\\rightarrow 384channels, kernel sizes11→511\\rightarrow 5, strides1→2→2→21\\rightarrow 2\\rightarrow 2\\rightarrow 2\) and collapses time with a learnable attention pooling\. A*local*pathway applies three stride\-1 blocks \(48→96→19248\\rightarrow 96\\rightarrow 192channels\) and pools into six temporal bins, preserving coarse within\-window timing\. The two pathways are concatenated and passed through a fully\-connected head to a 256\-dimensional per\-trial feature vector\. All blocks use ELU activations, batch normalisation and 20% dropout\.

Sequence model \(optional\)\.A four\-layer causal transformer \(256\-dimensional, 8 heads, 1,024\-dimensional feed\-forward, 40% dropout\) over the sequence of per\-trial features within a recording run\. The causal mask mirrors the causal attention of the language model that generated the targets: positioniiattends only to positionsj≤ij\\leq i\. Excluded trials are replaced by a learnable\[MASK\]token, so that the sequence stays aligned with the target sequence and the model is informed that a word occurred but its EEG is unusable\. Whether the transformer is present at all was a swept factor\.

Projection and objective\.Encoder \(or transformer\) outputs and projected LLM targets are mapped into a shared 256\-dimensional space, L2\-normalised, and trained with the symmetric CLIP contrastive loss\([20](https://arxiv.org/html/2608.20186#bib.bib20)\)using a learnable temperature initialised at the CLIP default\. One batch consists of 4 or 8 complete recording runs, so a run’s∼600\{\\sim\}600trials all contribute in\-batch negatives\. An option to mask*within\-run*negatives \(on the reasoning that consecutive contextual targets are near\-duplicates and make confusing negatives\) was tested and is reported as an optimisation factor\.

Models were trained for 100 epochs with AdamW\([15](https://arxiv.org/html/2608.20186#bib.bib15)\), weight decay 0\.1, and a cosine learning\-rate schedule with 5% warm\-up\. Checkpoints were selected on the within\-run retrieval gain \(Section[2\.8](https://arxiv.org/html/2608.20186#S2.SS8)\), not on validation loss \(because the contrastive loss can be confounded by an optionally learnable temperature parameter\)\.

### 2\.7Train/validation split

The present analysis is based on data from the deep subject only\. However, in preparation for the cross\-subject analysis \(where participants read partially identical text\), we had to ensure that no text leaked between training and validation\. We therefore used a content\-aware split based on clustering sets of 10\-grams\. Ca\. 20% of runs were assigned to the validation set, and the remaining 80% were assigned to the training set\. Consequently, no text chunk can appear on both sides of the split\. For the data\-scaling analysis, the split was computed first, and only training runs were then subsampled, so that the full validation set is scored identically at every data ratio\.

### 2\.8Evaluation metrics

Accounting for potential confounds in the evaluation required considerable care, because a contrastive decoder trained against contextual targets has several ways to achieve spuriously high decoding accuracy\. We report three regimes rather than a single accuracy\. All metrics are word\-grouped top\-10 retrieval within fixed candidate pools of 512 validation trials\. Within a pool, the model’s EEG prediction for a trial is compared \(cosine similarity, scaled by the learned temperature\) against every candidate target in the pool\. The resulting distribution is softmaxed and probability mass is summed across trials sharing the same word \(lower\-cased, punctuation\-stripped\), so that repeated presentations of the same word are treated as one candidate\. The rank of the true word is recorded\. Metrics are averaged over five independent pool draws\. A fixed pool size gives a well\-defined and stable reference level\.

Crucially, every reported quantity is a gain relative to an empirical permutation baseline, not a raw accuracy\. The baseline is obtained by permuting the EEG predictions within the pool \(10 permutations\) and recomputing the same metric\. This matters because word frequency alone confers a large advantage: a common word occupies many candidate slots and accumulates probability mass regardless of the EEG\. The permutation baseline absorbs that, so chance is zero for every gain reported below, and a positive value cannot be produced by the word\-frequency structure of the text\. Because a gain is a difference of two proportions \(model accuracy minus baseline accuracy\), we quote all gains in percentage points \(pp\): a within\-run gain of 6\.7 pp means the model’s top\-10 accuracy exceeds its empirical baseline by 6\.7 percentage points, and chance is 0 pp\. Percentage points \(absolute differences\) should not be confused with the relative changes also quoted in the Results \(e\.g\. “−32\-32%”\), which are percentages*of*a gain\.

The three evaluation regimes \(Figure[2](https://arxiv.org/html/2608.20186#S2.F2)\) are:

Overall retrieval gain\. Pools are random samples of validation trials, so a pool typically mixes trials from many recording runs\. The baseline in this regime is a permutation within the pool\. This is the total decoding signal, and it might be most directly comparable to published segment\-identification accuracies\. But on its own, the overall retrieval gain is uninterpretable, because in a cross\-run pool, the true target’s competitors mostly come from*other*runs \(with different text\), so simply knowing which narrative is being read can rank the target higher\. \(Run\-level EEG variance alone cannot inflate overall retrieval gain\. Ranking ispredi⋅tgtj\\mathrm\{pred\}\_\{i\}\\cdot\\mathrm\{tgt\}\_\{j\}across candidates, so a run\-identity component inpredi\\mathrm\{pred\}\_\{i\}only helps if the targets of that run’s words also cluster together in the shared space\.\)

Context\-tracking gain\. Identical to the overall gain, except that each trial’s EEG prediction is replaced by that of another trial from the*same*recording run\. This destroys the EEG\-to\-word pairing\. If the gain survives this swap, it was carried by run\-level context\. In contrast, if it collapses to zero, it required the correct EEG\-to\-word pairing\. This quantity therefore is the context\-tracking component of the overall gain, measured rather than assumed\.

Context\-independent gain==overall gain−\-context\-tracking gain\. The portion of the overall gain not explained by cross\-run context\. \(A derived quantity rather than a fourth pooling regime\.\)

Within\-run retrieval gain\. This is the primary metric and the checkpoint\-selection criterion\. Pools consist of consecutive trials from a*single*recording run, and the permutation baseline shuffles predictions*within*that run\. Cross\-run topic therefore does not help \(every candidate shares the run’s topic\), and the score reflects discrimination among words that occurred close together in the same text\. A positive within\-run gain cannot be produced by topic tracking, nor by slow EEG drift that merely tags run identity\.

There is one caveat regarding the within\-run retrieval gain\. In conditions using contextual targets \(from Llama layer 20\), “discrimination within a run above a within\-run shuffle” still conflates the*current word*with fine\-grained*position within the run*\(the slow trajectory of the discourse state\)\. The within\-run gain is therefore necessary, but not sufficient for lexical decoding\. It is cleanest in the non\-contextual, no\-transformer configuration, where the target carries no positional information\. Additionally, instead of relying on the non\-contextual configuration alone, we include position probes that measure the positional residual directly \(Section[2\.9](https://arxiv.org/html/2608.20186#S2.SS9)\)\.

Word\-frequency bins\. Validation words were split into three bins \(rare / mid / frequent\) by their trial count in the validation set, and all gains were additionally computed per bin, each against its own empirical chance level\.

Two caveats apply\. First, we have not used a held\-out test set or cross\-validation; all values are computed on the validation set, so absolute magnitudes are upper bounds and cross\-condition orderings are the most reliable output\. Second, reported scalars are read at the epoch that maximised the within\-run gain, so that metric carries winner’s\-curse optimism\.

![Schematic of the three evaluation regimes with pool construction, permutation nulls, and a worked example](https://arxiv.org/html/2608.20186v1/figures/figure_02_evaluation_regimes.png)Figure 2:The three evaluation regimes, and why the within\-run gain is the conservative one\.\(a\-c\)Pool construction and null for each regime; trials are coloured by recording run\.\(a\)*Overall retrieval gain*: pools are random samples of 512 validation trials, so most of a trial’s competitors come from other runs and therefore other text passages; the null permutes predictions within the pool\.\(b\)*Context\-tracking gain*: the same pools, but each trial’s EEG prediction is replaced by that of another trial from the same run\. Run identity is preserved while the EEG\-to\-word pairing is destroyed, so whatever score survives was carried by passage\-level context rather than by the current word\. Hence, the context component is measured, not assumed\.\(c\)*Within\-run retrieval gain*, the primary metric and the checkpoint\-selection criterion: pools are consecutive trials from a single run and the permutation null shuffles predictions within that run, so every candidate shares the passage’s topic\.\(d\)Worked example for a hypothetical decoder that recovers only which*text passage*is being read and nothing about the current word\. In a mixed pool, it ranks every same\-run candidate above the rest, and so scores above chance on the overall gain\. That score is unchanged by the run\-mate swap, which is what makes the swap a measurement of the context\-tracking component; and in a single\-run pool it has nothing left to exploit and falls back to the permutation null\. A positive within\-run gain therefore cannot be produced by topic tracking\. Pools are drawn with six candidates rather than 512 for legibility; the argument does not depend on pool size\.
### 2\.9Position probes

With a transformer, the position index enters the model through the learned positional embedding\. The positional embedding is the only trial\-varying input to the model that is not neural\. Because of the positional embedding, contextual \(layer\-20\) targets drift systematically along a run\. In principle, a positive within\-run gain could therefore be earned from position alone\. We measure this instead of assuming it, by re\-running the trained sequence model with per\-trial encoder outputs replaced by the learnable\[MASK\]token, the same substitution that artefact\-rejected trials receive \(Section[2\.6](https://arxiv.org/html/2608.20186#S2.SS6)\), so the model remains in distribution\.

Three nested information sets are evaluated: \(i\)*position only*, where every trial’s EEG is masked, leaving the model only its positional embedding; \(ii\) position plus the EEG of*preceding*words, i\.e\. the trial’s own EEG is masked; and \(iii\) the full model\. The differences between consecutive levels decompose the within\-run gain into aposition\-onlyterm, apreceding\-EEGterm, and acurrent\-trialterm\. The three terms sum to the within\-run gain\. We define theposition\-corrected within\-run gainas the within\-run gain minus the position\-only term; it is the portion of the within\-run gain that position alone cannot produce\.

As a validity check, we performed a control for representation collapse: a model that collapsed to a constant prediction under masking would yield a zero position\-only gain regardless of its decoding abilities\. This control is described in Supplementary[S6\.1](https://arxiv.org/html/2608.20186#A6.SS1)\.

The interpretation of the preceding\-EEG term deserves emphasis, because it shapes the reading of the results\. The model receives no text; apart from the positional embedding, its only input is EEG\. To exploit local context at all, the model must first have decoded information about preceding words from their EEG\. The preceding\-EEG term is therefore neural language decoding at a coarser granularity, not a shortcut, and combining it with the current word’s own response is what a deployed decoder should do\. Only the position\-only term is non\-neural, and that is the term the probes exist to bound\.

### 2\.10Experiments

Three sweeps are reported, comprising 788 model fits in total\. All were restricted to the deep participant, and all fits were additionally evaluated with the position probes of Section[2\.9](https://arxiv.org/html/2608.20186#S2.SS9)\.

Sweep 1\(576 fits\)\. A fully crossed grid over three scientifically\-motivated factors: LLM embedding layer \(0, 20\), window lock \(onset, offset, centre\), and presence of the transformer\. Additional optimisation factors: temporal\-shift data augmentation \(0,±100\\pm 100ms\), ISI subset \(0 ms only, 100 ms only, both\), normalisation \(scaling only vs\. baseline subtraction and scaling\), and two more technical factors \(masking of within\-run negatives; trainable vs\. fixed temperature\)\. The EEG window length was fixed at 0\.5 s\.

Sweep 2\(data scaling;2×102\\times 10fits\)\. A single configuration swept over the fraction of training runs retained \(10%, 20%, etc\. up to 100% of training runs\), separately for the contextual configuration \(layer 20 targets, transformer\) and the non\-contextual configuration \(layer 0 targets, no transformer\)\. One fit per ratio\.

Sweep 3\(192 fits\)\. The headline question was a*channel ablation*: all 19 channels versus a montage with O1, O2, T5 and T6 removed \(occipital and posterior\-temporal, i\.e\. the visual and ventral\-stream sites that drove silent\-reading decoding in[5](https://arxiv.org/html/2608.20186#bib.bib5)\)\. Also crossed: embedding layer \(0 or 20\), window length \(0\.5, 0\.7, 0\.9 s\), transformer presence, and three optimisation factors \(temporal\-shift data augmentation 0 vs\.±50\\pm 50ms; batch size 4 vs\. 8; learning rate 1e\-4 vs\. 1e\-5\)\. Window lock was fixed at onset and ISI at both \(i\.e\. including all trials irrespective of ISI\)\. Due to the ill\-defined spatial resolution of EEG, the channel ablation is no perfect ablation of visual signal, but we still consider its relative effect informative\.

Throughout, factors are described as*scientific*\(embedding layer, transformer, window lock and length, channel set, data volume\) or*optimisation*\(batch size, learning rate, normalisation, negative masking, temperature, augmentation\)\.

## 3Results

All gains reported here are top\-10 word\-grouped retrieval gains within 512\-trial pools, relative to an*empirical permutation baseline*, expressed in*percentage points*\(pp; Section[2\.8](https://arxiv.org/html/2608.20186#S2.SS8)\), so that*chance is 0 pp*\. Values are means±\\pmSD across the model fits contributing to each cell, and the number of fits \(nn\) is given for every cell in the Supplementary tables\.

### 3\.1Word\-level information is recoverable, and it is not narrative topic tracking

Across the 576 fits of Sweep 1, the within\-run retrieval gain, the primary and most conservative metric, averaged6\.7±2\.76\.7\\pm 2\.7pp \(median 6\.3 pp; range 1\.4 to 15\.4 pp\)\. Every one of the 576 fits produced a positive within\-run gain; the same was true of all 192 fits of Sweep 3 \(mean7\.8±2\.97\.8\\pm 2\.9pp, range 2\.0 to 16\.5 pp\) and all 20 fits of the scaling sweep\. Because the reported value is read at the epoch that maximised it, a positive minimum is not by itself evidence of signal; the relevant comparison is with the scale of epoch\-to\-epoch noise on this metric, which is on the order of 0\.1 to 0\.2 pp \(the pre\-training values at epoch 0 lie within±0\.2\\pm 0\.2pp of zero in every cell\)\. The observed minimum of 1\.4 pp is an order of magnitude above that, and the median of 6\.3 pp nearly two\. Decoding is therefore not marginal or configuration\-dependent in the sense of appearing only in favourable cells: it is present everywhere in the explored space, and the sweeps determine its magnitude rather than its existence\.

The decomposition is more informative than the aggregate\. Table[1](https://arxiv.org/html/2608.20186#S3.T1)and Figure[3](https://arxiv.org/html/2608.20186#S3.F3)give the four crossed configurations of embedding layer×\\timessequence model from Sweep 1\.

Table 1:Retrieval\-gain decomposition by target type and architecture \(Sweep 1;n=144n=144fits per cell; mean±\\pmSD; gains in percentage points, pp\)\.ConfigurationWithin\-rungain \(pp\)Overallgain \(pp\)Context\-trackinggain \(pp\)Context\-independentgain \(pp\)Non\-contextual target,no sequence model \(L0, no TF\)7\.5±2\.0\\mathbf\{7\.5\\pm 2\.0\}7\.9±1\.97\.9\\pm 1\.91\.0±0\.5\\mathbf\{1\.0\\pm 0\.5\}6\.9±1\.96\.9\\pm 1\.9Non\-contextual target,transformer \(L0, TF\)6\.2±2\.26\.2\\pm 2\.27\.9±2\.37\.9\\pm 2\.31\.9±0\.81\.9\\pm 0\.86\.0±2\.26\.0\\pm 2\.2Contextual target,no sequence model \(L20, no TF\)5\.5±2\.55\.5\\pm 2\.511\.0±3\.211\.0\\pm 3\.24\.6±1\.64\.6\\pm 1\.66\.3±3\.06\.3\\pm 3\.0Contextual target,transformer \(L20, TF\)7\.8±3\.3\\mathbf\{7\.8\\pm 3\.3\}19\.8±6\.1\\mathbf\{19\.8\\pm 6\.1\}5\.9±1\.55\.9\\pm 1\.513\.9±5\.4\\mathbf\{13\.9\\pm 5\.4\}We find:

The decomposition behaves exactly as designed, which validates it\.The context\-tracking gain increases monotonically with the amount of contextual information available to the model, from 1\.0 pp when neither the target nor the architecture carries context, through 1\.9 and 4\.6 pp when one of them does, to 5\.9 pp when both do\. This ordering was predicted by construction and is reproduced in Sweep 3 \(0\.7, 1\.5, 4\.8, 6\.7 pp for the same four cells\)\.

In the non\-contextual configuration, essentially all of the decoding is word\-level\.With non\-contextual layer\-0 targets and no sequence model, neither the target nor the model has any access to narrative context, and the context\-tracking gain is correspondingly close to zero \(1\.0 pp\) while the within\-run gain is 7\.5 pp, the second highest of the four cells\. Topic tracking cannot contribute here\. Short single\-trial EEG segments recorded during naturalistic silent reading carry information that discriminates among words presented within the same passage, above a within\-passage shuffle of the same EEG, in an open vocabulary\.

Contextual decoding is much larger overall, and its excess must be decomposed rather than assumed\.The contextual\-target\-plus\-transformer configuration \(L20, TF\) produces two and a half times the overall gain of the non\-contextual configuration \(19\.8 vs 7\.9 pp\) and the single best fits in both sweeps\. Of that overall gain, 5\.9 pp \(roughly 30%\) is recovered even when a trial’s EEG is swapped for a run\-mate’s, i\.e\. it reflects knowing*which text passage*is being read rather than*which word*\. This is a real and arguably useful capability \(see Section[4\.3](https://arxiv.org/html/2608.20186#S4.SS3)\), but it is not word\-level lexical decoding and should not be quoted as such\. A possible confounding factor, the transformer’s positional embedding acting as a non\-neural positional prior on the within\-run gain, is measured and discussed in Section[3\.2](https://arxiv.org/html/2608.20186#S3.SS2)\.

The interaction between target type and architecture is clean and consistent across sweeps:matchedconfigurations win\. Contextual targets require a sequence model to be exploited \(contextual embeddings from layer 20 \(L20\): 5\.5 pp without a transformer→\\rightarrow7\.8 pp with one\), whereas non\-contextual targets are better served without one \(layer 0 \(L0\) embeddings: 7\.5 pp without→\\rightarrow6\.2 pp with\)\. A causal transformer trained against a target that carries no context appears to spend capacity modelling sequence structure that the target does not reward, at the cost of per\-trial word discrimination\.

Selection\-optimistic best single fits were a within\-run gain of 15\.4 pp \(Sweep 1; contextual target, transformer, no augmentation\) and 16\.5 pp \(Sweep 3; contextual target, transformer, 0\.9 s window, all channels, no augmentation, batch 4, learning rate 1e\-4\)\. These are maxima over 576 and 192 fits respectively as well as over epochs, and should be read as upper bounds\.

![Grouped bar charts of within-run, overall and context-tracking gains for the four target-type by architecture configurations in Sweeps 1 and 3](https://arxiv.org/html/2608.20186v1/figures/figure_03_decomposition.png)Figure 3:Retrieval\-gain decomposition by target type and architecture\.Word\-level information is recoverable, and it is not restricted to narrative topic tracking\. Grouped bars give the three evaluation regimes of Figure[2](https://arxiv.org/html/2608.20186#S2.F2), i\.e\. within\-run \(the primary metric\), overall, and context\-tracking, for the four crossed configurations of target type×\\timesarchitecture, in\(a\)Sweep 1 \(576 fits, 144 per configuration\) and\(b\)Sweep 3 \(192 fits, 48 per configuration\)\. Configurations are ordered left to right by how much narrative context is available to the fit: neither the target nor the architecture carries it \(L0, no TF\), one of them does \(L0, TF and L20, no TF\), or both do \(L20, TF\)\. Bars are means over fits±1\\pm 1SD, points are individual fits, the number above each bar isnn, and every quantity is referenced to an empirical permutation baseline, so chance is 0 pp on the y\-axis\. The diamond marks the context\-independent gain \(overall−\-context\-tracking\); it is drawn as a derived marker rather than a fourth bar because the three measured gains come from different pooling regimes and therefore do not sum\. The decomposition behaves as designed: the context\-tracking bar rises with available context \(Sweep 1:1\.0→1\.9→4\.6→5\.91\.0\\rightarrow 1\.9\\rightarrow 4\.6\\rightarrow 5\.9pp; Sweep 3:0\.7→1\.5→4\.8→6\.70\.7\\rightarrow 1\.5\\rightarrow 4\.8\\rightarrow 6\.7pp\), an ordering predicted by construction\. Contextual targets and models \(L20, TF\) result in by far the largest overall gain, and some of this gain is context tracking\.
### 3\.2The within\-run gain is not produced by the positional embedding

The within\-run gain removes cross\-run topic and run\-constant drift, allowing us to make conclusions about word\-level decoding\. However, the within\-run gain might still be confounded by a non\-neural positional prior \(as described in Section[2\.9](https://arxiv.org/html/2608.20186#S2.SS9)\)\. The position probes \(Section[2\.9](https://arxiv.org/html/2608.20186#S2.SS9)\) further decompose the within\-run gain into a non\-neural positional prior, a preceding\-EEG contribution, and a current\-trial contribution \(Table[2](https://arxiv.org/html/2608.20186#S3.T2), Figure[4](https://arxiv.org/html/2608.20186#S3.F4)\)\.

Table 2:Position probes: the within\-run gain decomposed into a positional prior, a preceding\-EEG contribution, and a current\-trial contribution \(Sweep 1;n=144n=144fits per cell; mean±\\pmSD; gains in percentage points, pp\)\. Unlike Table[1](https://arxiv.org/html/2608.20186#S3.T1), these terms do sum: position\-only\+\+preceding\-EEG\+\+current\-trial==within\-run gain, and position\-corrected==within\-run gain−\-position\-only\. Position\-only is the only non\-neural term\. Dispersion is a validity check on the position\-only null; a dispersion of zero indicates representation collapse when the model receives masked EEG data and can only use the positional embedding for its prediction\. Hence, a near\-zero position\-only gain is evidence only where dispersion is above zero\. \(In the transformer\-free cells, a dispersion of exactly zero is expected by construction and does not indicate collapse\.\) For more details, see Supplementary[S6\.1](https://arxiv.org/html/2608.20186#A6.SS1)\.ConfigurationWithin\-rungain \(pp\)Position\-only \(pp\)Preceding\-EEG \(pp\)Current\-trial \(pp\)Position\-corrected \(pp\)DispersionL0, no TF7\.5±2\.07\.5\\pm 2\.00\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.07\.5±2\.07\.5\\pm 2\.07\.5±2\.07\.5\\pm 2\.00\.000±0\.0000\.000\\pm 0\.000L0, TF6\.2±2\.26\.2\\pm 2\.20\.0±0\.10\.0\\pm 0\.10\.0±0\.20\.0\\pm 0\.26\.1±2\.26\.1\\pm 2\.26\.1±2\.26\.1\\pm 2\.20\.167±0\.0410\.167\\pm 0\.041L20, no TF5\.5±2\.55\.5\\pm 2\.50\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.05\.5±2\.55\.5\\pm 2\.55\.5±2\.55\.5\\pm 2\.50\.000±0\.0000\.000\\pm 0\.000L20, TF7\.8±3\.37\.8\\pm 3\.32\.7±0\.72\.7\\pm 0\.71\.9±0\.81\.9\\pm 0\.83\.1±2\.73\.1\\pm 2\.75\.1±3\.05\.1\\pm 3\.00\.091±0\.0480\.091\\pm 0\.048The position probe allows the following conclusions:

The contextual model does exploit positional information\.For L20, TF, position alone earns 2\.7 pp of the 7\.8 pp within\-run gain, roughly a third of the cell mean\. This result demonstrates that when decoding language from neural signals with context\-aware models that have access to positional information, we have to account for the possibility of a non\-neural shortcut\. Without this control, about a third of the within\-run gain for the strongest configuration would have been quoted as neural decoding when it is attributable to a non\-neural positional prior\.

The remaining gain is neural, and it is split between the current and the preceding words\.The position\-corrected within\-run gain of the contextual configuration is 5\.1 pp, of which 3\.1 pp is contributed by the current trial’s own EEG and 1\.9 pp by the EEG of preceding trials \(for more details, see Supplementary[S6\.1](https://arxiv.org/html/2608.20186#A6.SS1)\)\. The preceding\-EEG term is not a confound: the model receives no text, so exploiting local context at all requires having decoded preceding words from their EEG\. It is neural language decoding at a coarser granularity, and combining it with the current word’s own response is what a deployed decoder should do\.

After position correction, the contextual and non\-contextual configurations decode the current word comparably\.The position\-corrected gain of L20, TF \(5\.1 pp\) does not exceed the non\-contextual cell’s 7\.5 pp in Sweep 1; in Sweep 3 the two are at parity \(7\.2 vs 7\.4 pp\)\. Sweep 3 reproduces the whole pattern \(L20, TF: within\-run 10\.0 pp==2\.9 position\-only\+\+2\.3 preceding\-EEG\+\+4\.9 current\-trial; Supplementary[S6](https://arxiv.org/html/2608.20186#A6)\)\. The contextual configuration’s real advantages lie elsewhere: in the overall and context\-independent gains \(Table[1](https://arxiv.org/html/2608.20186#S3.T1)\), in the rare and mid\-frequency vocabulary \(Section[3\.3](https://arxiv.org/html/2608.20186#S3.SS3)\), and in its scaling behaviour \(Section[3\.4](https://arxiv.org/html/2608.20186#S3.SS4)\)\.

The position probe is internally consistent\.In the transformer\-free cells, both probes return exactly zero, as they are expected to by construction, and the three terms sum to the within\-run gain\. In the one cell where a transformer is present but the target carries no context \(L0, TF\), the position\-only term is zero, despite a clearly positive dispersion \(0\.17\): the probe produces position\-dependent output and finds essentially nothing where there is nothing to find\.

Throughout the remainder of the paper, within\-run gains of transformer configurations are read against their position\-corrected values, and the position\-corrected value is the one we treat as quotable lexical decoding\.

![Four-panel figure of position-probe decompositions, dispersion diagnostic, per-frequency-bin corrections, and scaling of probe terms](https://arxiv.org/html/2608.20186v1/figures/figure_04_position_probes.png)Figure 4:Position probes: the within\-run gain is not produced by the positional embedding\.The transformer’s learned positional embedding is the only trial\-varying input to the model that is not neural, and contextual \(layer\-20\) targets drift systematically along a run, so in principle a within\-run gain could be earned from position alone\. The position probes measure this potential confound by re\-running the trained sequence model with per\-trial encoder outputs replaced by the learnable\[MASK\]token \(Section[2\.9](https://arxiv.org/html/2608.20186#S2.SS9)\)\.\(a\)and\(b\)The probe results for the four target\-type×\\timesarchitecture configurations, in Sweeps 1 \(\(a\)\) and 3 \(\(b\)\)\. The three terms are nested \(position only; position\+\+preceding\-word EEG;\+\+the current word’s own EEG\) and sum to the within\-run gain\. The preceding\-EEG segment is a signal, not a nuisance: The model receives no text, so exploiting local context at all requires having decoded preceding words from EEG\. Only the position\-only segment is non\-neural\.\(c\)Raw and position\-corrected within\-run gain per word\-frequency bin\. A potential concern is that the decoding performance in some frequency bins might be overwhelmingly driven by exploiting position information\. But the results survive position correction in all word frequency bins\.\(d\)Scaling of the position probe terms across a decade of training data \(Sweep 2\)\. The position\-only term stays flat near zero \(L0, no TF\) or grows at a lower rate than the corrected component \(L20, TF\)\. Bars are means over fits±1\\pm 1SD, and chance is zero on every gain axis\.
### 3\.3The signal extends beyond frequent words

If the decoder were exploiting only function words and other high\-frequency items, the frequency\-binned profile would show a gain confined to the frequent bins, but it does not \(Table[3](https://arxiv.org/html/2608.20186#S3.T3)\)\.

Table 3:Within\-run retrieval gain by word\-frequency bin \(Sweep 1;n=144n=144per cell; gains in percentage points, pp\)\.ConfigurationRareMidFrequentL0, no TF3\.2±1\.53\.2\\pm 1\.53\.2±1\.43\.2\\pm 1\.48\.1±2\.28\.1\\pm 2\.2L0, TF2\.4±1\.52\.4\\pm 1\.52\.4±1\.42\.4\\pm 1\.46\.7±2\.46\.7\\pm 2\.4L20, no TF3\.5±2\.03\.5\\pm 2\.03\.6±2\.13\.6\\pm 2\.15\.8±2\.55\.8\\pm 2\.5L20, TF6\.0±3\.56\.0\\pm 3\.56\.0±3\.36\.0\\pm 3\.38\.1±3\.38\.1\\pm 3\.3The frequent bin does carry the largest gain in every configuration, which is expected\. Word frequency confers two advantages: frequent words appear more often in training, so the encoder learns a better mapping for them, and they are represented by more validation trials, so both the model’s estimate and the metric’s per\-bin statistics are better resolved\. But the rare and mid bins are also positive throughout\. Sweep 3 reproduces this \(L0, no TF: rare 3\.6, mid 3\.8, frequent 7\.8 pp; L20, TF: rare 8\.9, mid 8\.5, frequent 10\.2 pp\)\.

The gap between the bins narrows in the contextual\-plus\-transformer configuration: the rare\-to\-frequent ratio rises from 0\.40 in the non\-contextual cell to 0\.74 in the contextual one \(Sweep 1\), and from 0\.46 to 0\.87 in Sweep 3\. Two readings are compatible with this, and they are not mutually exclusive\. First, on the output side, contextual targets may genuinely improve lexical decoding of rare words\. Alternatively, the narrowing may reflect the utility of contextual information at the input side\. Rare content words are the most topically distinctive items in a passage, and might therefore both be most indicative of the local discourse state and most predictable from it\. A model tracking the semantic trajectory of a run might use correctly decoded information on the topic of the current text passage, in addition to the words’ own neural response, to rank them well\. The position probes constrain a third reading: the narrowing is not a positional artefact\. After subtracting the per\-bin position\-only term, the rare\-to\-frequent ratio in the contextual cell remains 0\.70 in Sweep 1 \(3\.7/5\.3 pp\) and 0\.91 in Sweep 3 \(6\.6/7\.3 pp\), against 0\.40 and 0\.46 in the non\-contextual cell \(Supplementary[S6](https://arxiv.org/html/2608.20186#A6)\)\.

![Bar charts of within-run and overall gains split by rare, mid and frequent word bins, plus rare-to-frequent gain ratios](https://arxiv.org/html/2608.20186v1/figures/figure_05_frequency_bins.png)Figure 5:Word\-frequency profile\.The decodable signal is not confined to frequent words\.\(a, b\)Within\-run retrieval gain per word\-frequency bin\. Validation words were split into rare, mid and frequent thirds by their trial count in the validation set, each scored against its own empirical permutation baseline, for each of the four target\-type×\\timesarchitecture configurations, in\(a\)Sweep 1 \(144 fits per configuration\) and\(b\)Sweep 3 \(48 per configuration\)\. The solid inner bar is the within\-run gain \(the primary metric, word\-level decoding\); the wider, paler bar behind it is the overall gain from random cross\-run pools \(including a context\-tracking component\)\. Bars are means over fits±1\\pm 1SD, points are individual fits, and chance is zero for every bar\. The frequent bin carries the largest gain in every configuration, as expected; frequent words are seen more often in training, and are better resolved by the metric \(more validation trials\)\. But the rare and mid bins are also positive throughout\. A gain confined to function words would appear here as a rare and mid bin at zero, and it does not\.\(c\)Ratio of the rare\-bin to the frequent\-bin within\-run gain per configuration, computed from the cell means for both sweeps\. The dotted line at 1\.0 marks where the rare bin would decode as well as the frequent one\. The gap narrows where narrative context is available, from 0\.40 in the purely non\-contextual cell to 0\.74 in the contextual\-plus\-transformer cell in Sweep 1, and from 0\.46 to 0\.87 in Sweep 3\. Contextual targets may genuinely improve decoding of rare content words, and / or a model tracking the discourse state may have an advantage when ranking the topically distinctive rare words well, compared to a model that has no context and can only decode them from their own neural response\.
### 3\.4Decoding scales log\-linearly with training data and shows no saturation

Two configurations were swept across ten training\-data ratios, holding the validation set fixed \(Table[4](https://arxiv.org/html/2608.20186#S3.T4), Figure[6](https://arxiv.org/html/2608.20186#S3.F6)\)\. At full data, the training set comprised approximately 192,100 trials and the fixed validation set approximately 48,000 trials, before applying the amplitude\-threshold exclusion criterion \(ca\. 93% of trials survive the criterion in the 0\.5 s, no\-augmentation configuration used here; Supplementary[S8](https://arxiv.org/html/2608.20186#A8)\)\.

Table 4:Log\-linear scaling fits,gain∼a\+b⋅log10⁡\(training\-data ratio\)\\mathrm\{gain\}\\sim a\+b\\cdot\\log\_\{10\}\(\\text\{training\-data ratio\}\)\. Slopes \(bb\) are in top\-10 percentage points \(pp\) per decade, with theR2R^\{2\}of the log\-linear fit in parentheses; the endpoint columns give the gain at 10% and at 100% of training runs, in pp\. One fit per ratio\. The probe rows decompose the within\-run gain \(Section[2\.9](https://arxiv.org/html/2608.20186#S2.SS9)\); “flat” means the trace had no variance across the range, which is the desired outcome for the position\-only probe\.Contextual \(L20\+\+TF\)Non\-contextual \(L0, no TF\)MetricSlope, pp/decade\(R2R^\{2\}\)10%→\\rightarrow100%\(pp\)Slope, pp/decade\(R2R^\{2\}\)10%→\\rightarrow100%\(pp\)Within\-run gain \(primary\)\+8\.7\\mathbf\{\+8\.7\}\(0\.98\)6\.2→15\.06\.2\\rightarrow 15\.0\+4\.8\\mathbf\{\+4\.8\}\(0\.99\)5\.0→10\.15\.0\\rightarrow 10\.1Overall gain\+24\.5\+24\.5\(0\.96\)8\.9→33\.78\.9\\rightarrow 33\.7\+5\.0\+5\.0\(0\.99\)4\.7→9\.74\.7\\rightarrow 9\.7Context\-tracking gain\+5\.6\+5\.6\(0\.94\)1\.7→7\.31\.7\\rightarrow 7\.3\+0\.6\+0\.6\(0\.85\)0\.2→0\.70\.2\\rightarrow 0\.7Within\-run gain, rare bin\+9\.0\+9\.04\.5→13\.34\.5\\rightarrow 13\.3\+2\.3\+2\.31\.2→4\.41\.2\\rightarrow 4\.4Within\-run gain, mid bin\+8\.8\+8\.84\.0→13\.04\.0\\rightarrow 13\.0\+2\.6\+2\.60\.9→3\.60\.9\\rightarrow 3\.6Within\-run gain, frequent bin\+8\.6\+8\.66\.5→15\.36\.5\\rightarrow 15\.3\+5\.1\+5\.15\.5→10\.95\.5\\rightarrow 10\.9Probe: position\-only\+1\.5\+1\.5\(0\.51\)2\.4→3\.72\.4\\rightarrow 3\.70\.00\.0\(flat\)0\.0→0\.00\.0\\rightarrow 0\.0Probe: preceding\-EEG\+1\.3\+1\.3\(0\.45\)0\.8→2\.50\.8\\rightarrow 2\.50\.00\.0\(flat\)0\.0→0\.00\.0\\rightarrow 0\.0Probe: current\-trial\+5\.9\+5\.9\(0\.97\)3\.0→8\.83\.0\\rightarrow 8\.8\+4\.8\+4\.8\(0\.99\)5\.0→10\.15\.0\\rightarrow 10\.1Probe: position\-corrected\+7\.2\+7\.2\(0\.91\)3\.8→11\.33\.8\\rightarrow 11\.3\+4\.8\+4\.8\(0\.99\)5\.0→10\.15\.0\\rightarrow 10\.1Four observations:

Scaling is present in the primary metric\.Both configurations roughly double their within\-run gain across one decade of training data, and the log\-linear fit is close \(R2≥0\.98R^\{2\}\\geq 0\.98\)\. This is the pattern[21](https://arxiv.org/html/2608.20186#bib.bib21)reported for overt\-speech EEG decoding, reproduced here for silent reading and for a metric that has topic tracking removed by construction\.

Scaling reaches the rare and mid bins\.The rare and mid bins grow at least as fast as the frequent bin in the contextual configuration \(\+9\.0\+9\.0and\+8\.8\+8\.8vs\+8\.6\+8\.6pp per decade\) and at 45–51% of the frequent bin’s rate in the non\-contextual one, so additional data buys additional word decoding\. Hence, all frequency bins profit from more data, but the mid and rare bins improve much faster in the contextual condition \(L20, TF\) than without context \(L0, no TF\), similar to the ratios observed in Figure[5](https://arxiv.org/html/2608.20186#S3.F5)c\.

In the non\-contextual configuration the context\-tracking component stays flat and negligible\.It rises from 0\.2 to 0\.7 pp across the full decade \(slope\+0\.6\+0\.6pp/decade\)\. The gains in that configuration cannot be attributed to the model progressively learning to recognise passages\.

The positional prior does not grow with the result\.In the contextual configuration, the position\-only term grows at\+1\.5\+1\.5pp/decade \(and with the weakest fit in the table,R2=0\.51R^\{2\}=0\.51\) while the position\-corrected gain grows at\+7\.2\+7\.2pp/decade\. In the non\-contextual configuration the position\-only trace is essentially zero\. A confound that explained the scaling would have to scale with it\. In contrast, we observe that across a decade of training data, the additional data is buying neural decoding, not a better positional prior\.

Critically,none of the curves shows saturation\. At∼49\{\\sim\}49h of EEG data, the model is data\-limited\.

![Log-linear scaling curves of decoding gain against training-data volume for contextual and non-contextual configurations, overall and per frequency bin](https://arxiv.org/html/2608.20186v1/figures/figure_06_data_scaling.png)Figure 6:Data scaling: decoding grows log\-linearly with training\-data volume and shows no saturation at∼49\{\\sim\}49h\.Sweep 2, ten fits per configuration, using 10%, 20%, … up to 100% of the training data subset\. The lower and upper x\-axes show the same information on different metrics \(ratio of retained training runs and number of training trials after artefact rejection, respectively\)\. The same, full validation set is used at every ratio\.\(a\)Within\-run gain against training volume, with log\-linear fit annotated, separately for the most contextual \(L20, TF\) and least contextual \(L0, no TF\) conditions\. Both configurations roughly double their gain across one decade of data\.\(b\)All three evaluation regimes for most and least contextual configurations \(colour==configuration, line style==evaluation regime\)\. Importantly, in the non\-contextual condition, the context\-tracking trace stays flat against zero \(0\.2→0\.70\.2\\rightarrow 0\.7pp,\+0\.6\+0\.6pp/decade\) while its within\-run trace climbs, so the additional data buys word\-level decoding rather than progressively better passage recognition\. In the contextual condition, the overall gain grows fastest of all \(\+24\.5\+24\.5pp/decade\) but its context\-tracking component grows with it \(\+5\.6\+5\.6pp/decade\)\.\(c, d\)Per\-bin scaling, one panel per configuration\.\(c\)Non\-contextual,\(d\)contextual\. Growth reaches the rare and mid bins, at 45–51% of the frequent bin’s rate in the non\-contextual arm and at or slightly above the frequent bin’s rate in the contextual one\. Chance is zero on every gain axis\.
### 3\.5Removing visual and ventral\-stream channels reduces but does not abolish decoding

[5](https://arxiv.org/html/2608.20186#bib.bib5)found that silent\-reading decoding in EEG, MEG and OPM\-MEG was driven by visual processing, with spatial permutation feature importance peaking over occipital sensors and temporal importance peaking near 150 ms\. As an exploratory step, in Sweep 3 we removed electrodes O1, O2, T5 and T6 \(over visual cortex and ventral visual stream\), leaving 15 of 19 channels \(Table[5](https://arxiv.org/html/2608.20186#S3.T5), Figure[7](https://arxiv.org/html/2608.20186#S3.F7)\)\. Because of the ill\-defined spatial resolution of EEG, this channel ablation is not a perfect ablation of visual signal\.

Table 5:Effect of removing occipital/posterior\-temporal channels \(Sweep 3;n=96n=96per marginal cell, 48 per interaction cell; gains in percentage points, pp\)\.ContrastAll 19 channelsO1/O2/T5/T6 removedRelative changeWithin\-run gain \(marginal\)9\.2±2\.79\.2\\pm 2\.76\.3±2\.36\.3\\pm 2\.3−32%\-32\\%Overall gain \(marginal\)15\.2±8\.615\.2\\pm 8\.611\.8±8\.111\.8\\pm 8\.1−22%\-22\\%Context\-tracking gain \(marginal\)3\.5±2\.73\.5\\pm 2\.73\.3±2\.63\.3\\pm 2\.6−4%\-4\\%Position\-corrected within\-run gain8\.5±2\.48\.5\\pm 2\.45\.6±1\.85\.6\\pm 1\.8−34%\-34\\%Within\-run gain, non\-contextual target8\.8±1\.68\.8\\pm 1\.65\.5±1\.35\.5\\pm 1\.3−37%\-37\\%Within\-run gain, contextual target9\.7±3\.49\.7\\pm 3\.47\.0±2\.87\.0\\pm 2\.8−28%\-28\\%Within\-run gain, no transformer8\.6±2\.28\.6\\pm 2\.25\.5±1\.75\.5\\pm 1\.7−36%\-36\\%Within\-run gain, transformer9\.9±2\.99\.9\\pm 2\.97\.1±2\.67\.1\\pm 2\.6−29%\-29\\%Within\-run gain, rare bin6\.1±4\.06\.1\\pm 4\.04\.3±2\.94\.3\\pm 2\.9−29%\-29\\%Within\-run gain, frequent bin9\.6±2\.69\.6\\pm 2\.66\.5±2\.36\.5\\pm 2\.3−32%\-32\\%Removing occipital/posterior\-temporal channels does affect accuracy, but the decoder does not entirely depend on them\. Roughly two thirds of the within\-run gain survives their removal, in every configuration and in every frequency bin\. The ablation also does not change the ordering of any other factor: longer EEG windows still beat shorter ones, the transformer still helps, no data augmentation \(temporal shift\) still beats augmentation\.

Three features of the pattern are worth noting\. First, thecontext\-tracking component is almost unaffectedby the ablation \(−4%\-4\\%, well within the dispersion\), whereas the within\-run component drops by roughly a third\. Whatever supports passage\-level tracking is not primarily measured at occipital electrodes; whatever supports word\-level discrimination partly is\. Second, the ablation effect isslightly stronger in the non\-contextual and no\-transformer cells\(−37%\-37\\%and−36%\-36\\%\) than in the contextual and transformer cells \(−28%\-28\\%and−29%\-29\\%\), which is consistent with the contextual configurations having an additional, non\-visual source of signal to fall back on\. Third, theposition\-corrected within\-run gain falls by 34%, closely tracking the raw within\-run gain: the ablation cost is borne by the neural component of the signal, not by the positional prior, so the dissociation is not a positional artefact\.

Three limits of this ablation experiment are worth noting: \(1\) Removing 4 of 19 channels removes 21% of the input data, and performance loss would be expected from that alone\. \(2\) EEG has an ill\-defined spatial resolution, so activity generated in occipital and ventral\-temporal cortex still projects onto the remaining electrodes\. \(3\) The converse also holds: activity recorded at occipital electrodes is not necessarily entirely visual in nature\. Hence, the ablation is an interesting but weak test\. A decisive test would be whether a decoder trained on reading transfers to listening, which the planned passive\-listening condition is designed to provide\.

![Electrode layout with ablated channels marked, and bar charts of the ablation effect across evaluation regimes, configurations and frequency bins](https://arxiv.org/html/2608.20186v1/figures/figure_07_channel_ablation.png)Figure 7:Channel ablation: Removing occipital and posterior\-temporal electrodes reduces the word\-level component by about a third, but leaves passage\-level tracking untouched\.All panels are Sweep 3 \(192 fits; 96 per channel set, 48 per interaction bar\), the only sweep in which the channel set was varied\.\(a\)The 19\-electrode 10\-20 layout, with the four ablated electrodes \(O1, O2, T5, T6; visual cortex and ventral visual stream\) marked\.\(b\)The three evaluation regimes with the channel set in the hue, and the relative change of the ablated cell annotated above each pair\. The within\-run component falls by 32% \(9\.2→6\.39\.2\\rightarrow 6\.3pp\) and the overall gain by 22%, while the context\-tracking component moves by only−4%\-4\\%\(3\.5→3\.33\.5\\rightarrow 3\.3pp\), well within the dispersion\. Whatever supports passage\-level tracking is therefore not primarily measured at occipital electrodes, whatever supports word\-level discrimination partly is\.\(c\)The ablation cost by target type and by architecture, each bar marginal over the remaining swept factors\. The cost is slightly larger in the non\-contextual and no\-transformer cells \(−37%\-37\\%and−36%\-36\\%\) than in the contextual and transformer cells \(−28%\-28\\%and−29%\-29\\%\), consistent with the contextual configurations having an additional, non\-visual source of signal to fall back on\.\(d\)The ablation cost by word\-frequency bin: roughly two thirds of the gain survives in every bin \(rare−29%\-29\\%, frequent−32%\-32\\%\), so the loss is not concentrated in one part of the vocabulary\. Bars are means over fits±1\\pm 1SD, points are individual fits, numbers above bars arenn, and chance is zero throughout\. The channel ablation is not a clean localisation; removing 4 of 19 channels removes 21% of the input, and EEG has ill\-defined spatial resolution so occipital sources still project to the remaining electrodes \(Section[3\.5](https://arxiv.org/html/2608.20186#S3.SS5)\)\.
### 3\.6Temporal properties of the informative signal

Three analysis\-window factors were varied \(Table[6](https://arxiv.org/html/2608.20186#S3.T6), Figure[8](https://arxiv.org/html/2608.20186#S3.F8)\)\. First, the window lock, i\.e\. the EEG segments were aligned to the onset of a word, to the centre of the time window during which the word was shown, or to the word offset\. Words were shown for 0\.5 to 0\.9 seconds, depending on the length of the word, but the length of the model’s input EEG data was constant \(per hyperparameter condition\) irrespective of word length \(otherwise the model could have learned from a non\-neural signal\)\. Second, the window length of the EEG data used as model input\. Third, the temporal data augmentation, where the EEG data was shifted in time randomly by up to±50\\pm 50or±100\\pm 100ms\.

Table 6:Within\-run retrieval gain by analysis\-window parameters \(gains in percentage points, pp\)\.FactorLevelWithin\-run gain \(pp\)SweepWindow lockOnset6\.8±2\.96\.8\\pm 2\.91Centre6\.9±2\.86\.9\\pm 2\.81Offset6\.5±2\.56\.5\\pm 2\.51Window length0\.5 s6\.9±2\.76\.9\\pm 2\.730\.7 s8\.3±2\.8\\mathbf\{8\.3\\pm 2\.8\}30\.9 s8\.1±2\.98\.1\\pm 2\.93Temporal\-shift augmentationnone8\.4±2\.4\\mathbf\{8\.4\\pm 2\.4\}1±100\\pm 100ms5\.1±1\.95\.1\\pm 1\.91none8\.3±2\.9\\mathbf\{8\.3\\pm 2\.9\}3±50\\pm 50ms7\.2±2\.87\.2\\pm 2\.83Window lock made almost no difference\.Onset\-, centre\- and offset\-locked windows performed within 0\.4 pp of each other, with overlapping distributions\. This is perhaps not surprising\. With 0\.5 s windows and stimulus durations of 0\.5–0\.9 s, the three lock points define mostly overlapping segments, and in a continuous RSVP stream the response to one word overlaps the presentation of the next\. The result establishes that the decoder is not critically sensitive to window alignment at this granularity\.

Longer windows helped, up to a point\.Moving from 0\.5 s to 0\.7 s raised the within\-run gain by roughly 20%, with 0\.9 s no better than 0\.7 s\. Because the ISI was zero for most runs, a 0\.9 s window typically extends into the following word’s presentation; the absence of further improvement at 0\.9 s is therefore consistent with the informative, word\-specific activity being largely complete within∼0\.7\{\\sim\}0\.7s of word onset\.

Temporal jitter was consistently harmful, and dose\-dependently so\.Randomly shifting the analysis window by up to±100\\pm 100ms during training reduced the within\-run gain by 39% \(8\.4→5\.18\.4\\rightarrow 5\.1pp\);±50\\pm 50ms reduced it by 13% \(8\.3→7\.28\.3\\rightarrow 7\.2pp\)\. This was among the largest effects in Sweep 1 and it was uniform across every other factor\. Two readings are compatible with the data and are not mutually exclusive: the informative activity might be time\-locked to word onset, so jitter smears a phase\-locked response; and / or the model fails to learn a shift\-invariant decoding of the signal\. The finding parallels[19](https://arxiv.org/html/2608.20186#bib.bib19), who reported that ECoG phoneme classification degraded sharply once onset alignment was jittered beyond 100 ms\. For present purposes the practical conclusion is straightforward: temporal\-shift does not help as a data augmentation method in our current pipeline\.

![Bar charts of within-run gain by window lock, window length and training-time jitter, plus a stimulus timeline schematic](https://arxiv.org/html/2608.20186v1/figures/figure_08_temporal.png)Figure 8:Temporal properties of the informative signal\.The lock point barely matters, window length helps up to∼0\.7\{\\sim\}0\.7s, and training\-time jitter is harmful\.\(a\)Within\-run gain by window lock point \(Sweep 1, 0\.5 s windows; 192 fits per level\)\. Onset\-, centre\- and offset\-locked windows fall within 0\.4 pp of each other with overlapping distributions\. The decoder is not critically sensitive to window alignment at this granularity\.\(b\)Within\-run gain by window length \(Sweep 3; 64 fits per level\)\. Extending the window from 0\.5 s to 0\.7 s raises the gain by roughly 20% \(6\.9→8\.36\.9\\rightarrow 8\.3pp\); 0\.9 s is no better than 0\.7 s \(8\.1 pp\)\.\(c\)Within\-run gain by training\-time window jitter \(±100\\pm 100ms was crossed in Sweep 1,±50\\pm 50ms in Sweep 3\)\. Jitter is harmful and dose\-dependently so:−13%\-13\\%at±50\\pm 50ms and−39%\-39\\%at±100\\pm 100ms\. This is consistent either with the informative activity being precisely time\-locked to word onset, so that jitter smears a phase\-locked response, or with the model failing to learn a shift\-invariant decoding\. In either case, temporal\-shift augmentation does not help as a data augmentation strategy in this pipeline\.\(d\)Stimulus timeline and the three onset\-locked analysis windows\. Depending on the word duration \(0\.5 to 0\.9 s\) and on the window lock \(onset, offset, centre\), the 0\.7 s and 0\.9 s window lengths can contain the onset of the next word’s presentation\. The absence of any further improvement from 0\.7 s to 0\.9 s in panel \(b\) is consistent with the word\-specific activity being largely complete within∼0\.7\{\\sim\}0\.7s of onset\. Bars in \(a\), \(b\), and \(c\) are means over fits±1\\pm 1SD, points are individual fits, and chance is zero on every gain axis\.
### 3\.7Optimisation factors

The optimisation factors were swept to maximise decoding accuracy and to verify that they do not invert the scientific comparisons, and the results are reported briefly \(Figure[9](https://arxiv.org/html/2608.20186#S3.F9)\)\.

Data volume via ISI subsetting\.Training on runs from both ISI conditions gave the highest within\-run gain \(7\.2 pp\), followed by ISI=0=0ms alone \(6\.9 pp\) and ISI=100=100ms alone \(6\.1 pp\)\. The ordering tracks the number of available recording runs \(393, 222 and 171, respectively\) rather than any plausible ordering of ISI quality, so we read this as a data\-volume effect and not evidence that either ISI condition improves the signal\. The two ISI conditions are pooled in all subsequent sweeps\.

Learning rate and batch size\.A learning rate of 1e\-4 substantially outperformed 1e\-5 \(9\.4 vs 6\.1 pp\), and a batch of 4 runs outperformed 8 \(8\.2 vs 7\.3 pp\)\. Both differences are partly training\-budget artefacts under a fixed 100\-epoch schedule: the 1e\-5 fits selected their checkpoint at epoch91±791\\pm 7\(versus80±1580\\pm 15for 1e\-4\), i\.e\. they might still be improving at the end of training, and doubling the batch halves the number of optimiser steps per epoch\. Batch size also changes the contrastive task itself, since a batch is a set of complete recording runs and the number of in\-batch negatives scales with it\.

Normalisation\.Simple amplitude scaling outperformed scaling with pre\-stimulus baseline subtraction \(7\.2 vs 6\.2 pp\)\.

Masking of within\-run negatives\.Excluding same\-run pairs from the contrastive loss, motivated by the near\-duplicate nature of consecutive contextual targets,*reduced*the within\-run gain, and the penalty was largest in exactly the configuration the manipulation was designed to help \(contextual target\+\+transformer: 5\.0 pp masked vs 10\.6 pp unmasked\)\. Discriminating between adjacent words appears to be the useful part of the training signal rather than a source of confusion\. Note that masking changes the scale of the loss, so loss values are not comparable across this factor; all rankings here are on gains\.

Temperature\.Making the CLIP temperature learnable had no detectable effect on any gain\. No fit in any sweep exhibited temperature runaway \(maximum final temperature 19\.3 against a clamp at 100\), so the word\-grouped metrics \(which are temperature\-dependent\) are not distorted\.

Control check\.Neither technical factor inverted the ordering of any scientific factor\. The full control results are in Supplementary[S7](https://arxiv.org/html/2608.20186#A7)\.

![Five-column figure of optimisation-factor effects on within-run gain and the corresponding checkpoint-selection epochs](https://arxiv.org/html/2608.20186v1/figures/figure_09_optimisation.png)Figure 9:Optimisation factors, and how much of each difference is a training\-budget artefact rather than an optimum\.Five factors, one column each\.\(a\)Top row: within\-run gain per level\.\(b\)Bottom row: the epoch at which that fit’s checkpoint was selected, with the fixed 100\-epoch budget drawn as a dashed ceiling\. A bar sitting against the ceiling identifies a comparison in which the losing level was plausibly still improving when training stopped, so its gain is a lower bound and the difference should not be read as an intrinsic optimum\. Sweeps are separate, learning rate and batch size are marginals of Sweep 3, normalisation, within\-run negative masking and trainable temperature marginals of Sweep 1, so every comparison is a clean marginal of one fully crossed grid\. Bars are means over fits±1\\pm 1SD, points are individual fits, numbers above bars arenn, chance is zero\.*Learning rate*: 1e\-4 clearly beats 1e\-5 \(9\.4 vs 6\.1 pp\), but the 1e\-5 fits select their checkpoint at epoch91±791\\pm 7out of 100 against80±1580\\pm 15for 1e\-4, i\.e\. the difference is partly an incomplete budget\.*Batch size*: 4 runs beats 8 \(8\.2 vs 7\.3 pp\), with the same caveat \(83±1383\\pm 13vs87±1387\\pm 13epochs\) plus two structural confounds; doubling the batch halves the number of optimiser steps per epoch, and because a batch item is a complete recording run it also doubles the number of in\-batch contrastive negatives\.*Normalisation*: simple amplitude scaling beats scaling with pre\-stimulus baseline subtraction \(7\.2 vs 6\.2 pp\)\.*Within\-run negative masking*: masking same\-run pairs out of the contrastive loss*reduces*the gain \(8\.0→5\.58\.0\\rightarrow 5\.5pp\), the opposite of its motivation, so discriminating between adjacent words is the useful part of the training signal rather than a source of confusion\.*Trainable temperature*: no detectable effect on any gain, which is the expected and desired null\. The word\-grouped metrics are temperature\-dependent, so a null here is evidence they are undistorted\.

## 4Discussion

### 4\.1Summary

From roughly 49 h of 19\-channel EEG recorded from one participant reading continuous narrative prose, an open\-vocabulary contrastive decoder recovers information that discriminates among words presented within the same passage\. The effect is present in all 788 model fits we ran, survives a within\-passage permutation null, extends to mid\-frequency and rare words as well as frequent ones, and is partially but not entirely a result of tracking the narrative topic\. In the configuration where topic tracking is architecturally impossible, the within\-run gain is 7\.5 pp\. We have identified positional embeddings as a potential non\-neural confound\. Position alone accounts for 2\.7 pp of the contextual configuration’s 7\.8 pp within\-run gain, and the position\-corrected remainder is carried by neural signal from the current and the preceding words\. Performance grows log\-linearly with training data with no sign of saturation, and the growth reaches the rare and mid\-frequency bins rather than being confined to frequent words\. Removing occipital and posterior\-temporal electrodes costs about a third of the within\-run gain, but leaves the majority intact, and leaves the context\-tracking component essentially untouched\.

We regard five of these as substantive results: \(1\) open\-vocabulary word\-level decoding from EEG during silent reading of naturalistic text is possible; \(2\) the decoding can be decomposed into word\-level and context\-tracking components; \(3\) the positional prior that a causal sequence model introduces is an important potential confound that can be measured, and the within\-run gain is not reducible to it; \(4\) the system is in a data\-limited regime; \(5\) the signal is not exclusively measured at occipital and posterior\-temporal electrodes\.

### 4\.2Relation to prior non\-invasive work

The most direct comparison is[5](https://arxiv.org/html/2608.20186#bib.bib5), who also targeted silent reading with EEG and MEG in a densely\-sampled participant\. They decoded five words at 30–40% against a 20% chance level\. Our design differs in ways that make a numerical comparison meaningless but a qualitative one informative\. Our vocabulary is open and natural \(thousands of unique words drawn from continuous prose rather than five clinically\-chosen words\), our typography is randomised trial\-by\-trial, our decoder is contrastive against language\-model embeddings rather than a linear classifier over covariance features, and our evaluation is retrieval within 512\-candidate pools against a permutation baseline\. The convergent conclusion is that silent reading is decodable from EEG\. Our contribution is that it is decodable at the level of individual words drawn from an open vocabulary, and that the decodable information is not exhausted by the occipital and posterior\-temporal electrodes that their permutation feature importance identified as the driver\.

In the EEG component of their study,[7](https://arxiv.org/html/2608.20186#bib.bib7)achieved top\-10 segment accuracies of 17\.7% and 25\.7% out of 1,842 and 190 candidates respectively, for perceived speech\. Our task, metric, candidate\-set construction and chance level all differ, so we do not attempt an accuracy comparison\. What transfers is the methodological lesson their ablations established and our results reinforce: the contrastive objective and pretrained deep target representations enable non\-invasive decoding of natural, open\-vocabulary language\.

The most relevant comparison for the scaling result is[21](https://arxiv.org/html/2608.20186#bib.bib21)\. Their 175 h single\-participant overt\-speech EEG dataset reached 48\.5% top\-1 and 76\.0% top\-10 over 512 candidate segments, and exhibited an unsaturated log\-linear relationship between data volume and accuracy, with performance collapsing when subsampled to∼3\{\\sim\}3h\. Our current deep\-subject data volume is roughly a quarter of theirs, our unit is a single word rather than a 5 s multi\-word segment, our task is silent reading rather than reading aloud, and our metric is referenced to a permutation baseline rather than to uniform chance\. Within those differences, we reproduce the qualitative scaling result on a metric constructed to exclude topic tracking, which their design did not separate out\. We also avoid the concern that dominated their discussion: silent reading generates no speech\-related electromyographic activity, so the elaborate EMG\-decontamination they required \(adaptive filtering plus adversarial data augmentation\) has no analogue here\. We trade that concern for a different one, the visual confound, discussed in Section[4\.5](https://arxiv.org/html/2608.20186#S4.SS5)\.

Taken together, a coherent picture emerges\. Non\-invasive language decoding fails when data are scarce and succeeds, partially, when they are abundant\. The paradigms that yield abundant, well\-time\-locked data are perception and overt production\. The open question is whether anything learned in those regimes will generalise to inner speech\.

### 4\.3Context, position, and what counts as decoding

Every high\-performing invasive speech neuroprosthesis depends heavily on linguistic priors\.[18](https://arxiv.org/html/2608.20186#bib.bib18)reported that a language model reduced their median word error rate by 35 percentage points;[17](https://arxiv.org/html/2608.20186#bib.bib17)and[2](https://arxiv.org/html/2608.20186#bib.bib2)both rely on n\-gram language models and large\-language\-model rescoring to convert noisy phoneme posteriors into accurate text;[11](https://arxiv.org/html/2608.20186#bib.bib11)decode a 125,000\-word vocabulary in the same way\. A decoder that “knows” what is being talked about is a decoder whose language\-model prior is better conditioned\. In a deployed system, a component that tracks discourse state from EEG and a component that discriminates the current word would be complementary\.

The reason we separate context tracking and word\-level decoding is to gain a clearer understanding of the capabilities of our decoding model\. Contextual language\-model targets introduce the risk of obtaining large and impressive\-looking retrieval scores that contain little word\-level information, and the field has to be cautious of this class of problem\. In a previous example from the EEG\-to\-text literature, accuracies that appeared strong turned out to depend on teacher forcing at evaluation and dropped to chance without it \(as reviewed in[21](https://arxiv.org/html/2608.20186#bib.bib21)\)\. Our contextual\-plus\-transformer configuration produces an overall gain of 19\.8 pp, of which 5\.9 pp is measurably attributable to passage identity\.

The position probes extend this control from context to position, and we regard them as a central methodological strength of the present study\. With a sequence model, the learned positional embedding is the only trial\-varying input to the decoder that is not neural, and contextual targets drift systematically along a run \(Section[2\.9](https://arxiv.org/html/2608.20186#S2.SS9)\), so a model can earn an apparent within\-run gain merely by knowing*where in the run*a trial occurred\. The probes measure this shortcut directly\. In our strongest configuration, position alone accounts for 2\.7 pp of the 7\.8 pp within\-run gain, roughly a third of what would otherwise have been quoted as word\-level decoding\. All conclusions about contextual configurations are therefore drawn from position\-corrected values\. The same probes also isolate the contribution of the*preceding*words’ EEG, and this term is not a confound: the model receives no text, so any use of local context requires having decoded the preceding words from their EEG\. It is neural language decoding at a coarser granularity, and it mirrors the way the invasive systems mentioned above condition their language\-model priors on previously decoded history\.

The measured decomposition also sharpens the comparison between contextual and non\-contextual configurations\. The non\-contextual, no\-transformer cell serves as an important baseline, because there the positional and contextual routes are absent by construction rather than by measurement\. Once the positional prior is subtracted, the contextual configuration’s advantage is neural, but it does not lie in a larger current\-word signal at the present data volume \(position\-corrected within\-run gain: 5\.1 vs 7\.5 pp in Sweep 1, 7\.2 vs 7\.4 pp in Sweep 3\)\. It lies instead in the overall and context\-independent gains \(where decoded context is a legitimate additional discriminator\), in the rare and mid\-frequency vocabulary, and in a steeper scaling slope \(\+7\.2\+7\.2vs\+4\.8\+4\.8pp per decade, position\-corrected\)\. We would encourage future studies that align brain data to contextual language\-model embeddings to report word\-level and contextual components where applicable, and to safeguard against inflated accuracies from non\-neural positional\-embedding cues\.

### 4\.4What this does and does not imply for inner speech

Previous studies and our results support the view that non\-invasive language decoding requires large amounts of training data\. Subject compliance and attention cannot be sustained long enough when using slow\-paced, unnaturalistic, boring tasks\. Instead, silent reading and passive listening can act as proxy tasks to provide the bulk of training data for EEG\-to\-text decoders\. Whether a model trained on such a proxy task can transfer to spontaneous inner monologue remains an open question\.

The invasive literature is encouraging on the representational question\.[16](https://arxiv.org/html/2608.20186#bib.bib16)provided one of the first demonstrations that individual imagined words are decodable from direct cortical recordings, with discriminative information in superior temporal gyrus, inferior frontal gyrus and sensorimotor cortex\.[22](https://arxiv.org/html/2608.20186#bib.bib22)found strong shared representations in supramarginal gyrus between internal speech, visually presented word reading, and vocalised speech, and reported that decoders trained on written\-cue trials generalised better to internal and vocalised speech than decoders trained on auditory\-cue trials, a direct argument that reading\-based training data is a reasonable proxy for an inner\-speech decoder\.[11](https://arxiv.org/html/2608.20186#bib.bib11)showed that inner speech occupies a shared, scaled\-down subspace with attempted speech in motor cortex, so that training on attempted speech transfers\.[10](https://arxiv.org/html/2608.20186#bib.bib10)demonstrated zero\-shot transfer in the specific direction we care about: deep networks trained on elicited inner speech \(covert reading and repeating\) decoded self\-generated inner speech at an accuracy matching within\-task replication, in 7T fMRI\. The representational precondition for transfer therefore appears to hold\.

What is unresolved is whether it survives the signal\-to\-noise and spatial\-resolution penalty of scalp EEG\.[5](https://arxiv.org/html/2608.20186#bib.bib5)could decode silent reading but not inner speech, and the failure of the inner\-speech decoders prevented them from even testing transfer\. Our reading of that result is not that non\-invasive inner\-speech decoding is impossible, but that it was attempted with less data per participant than the scaling results suggest is needed, on a five\-word vocabulary, with paradigms whose compliance could not be verified\. Our strategy is to build a high\-volume, high\-compliance, well\-time\-locked proxy corpus first; establish that a decoder trained on it extracts language\-related rather than modality\-specific information \(the future reading/listening comparison\); and only then attempt transfer to inner speech\. The present results support the first step\.

### 4\.5The visual\-confound question

The single most important caveat on any silent\-reading result is that the decoder might be reading out visual word\-form processing instead of deep semantic language processing\.[13](https://arxiv.org/html/2608.20186#bib.bib13), decoding 80 words from EEG during silent reading, found that pairwise discriminability peaked around 170 ms and correlated with visual and orthographic similarity but not with semantic similarity, demonstrating that EEG word decoding*can*be entirely visual\.[5](https://arxiv.org/html/2608.20186#bib.bib5)localised their silent\-reading decoding to visual areas with an early temporal peak\.

We have taken three measures to address the visual confound\. First, typography is randomised on every trial, so word identity is decorrelated from any fixed visual template\. This partially removes a specific confound Csaky et al\. flagged, though it does not remove the systematic visual differences between different letter strings\. Second, we compared non\-contextual with contextual language\-model targets\. The layer\-20 targets are semantic and context\-dependent and bear no systematic relation to orthographic form, so the fact that they support the largest gains is at least consistent with non\-visual information contributing\. Third, we ablated the occipital and posterior\-temporal channels, and roughly two thirds of the within\-run gain survived\.

None of the three measures is decisive\. Volume conduction means occipital sources can reach frontal and central electrodes; a signal that survives occipital\-channel removal has not been shown to have a non\-visual generator\. Randomised typography removes template matching but leaves orthography correlated with word identity, which cannot be solved within a reading paradigm\.

Only a multimodal paradigm can ultimately answer the question of the visual confound, which is why we intend to follow up the present study with a passive\-listening paradigm\. If a decoder trained on reading transfers to listening, or if a jointly\-trained decoder shows a shared component, then the shared component cannot be visual word form, because listening has none\. Evidence for modality\-general semantic codes\([4](https://arxiv.org/html/2608.20186#bib.bib4);[12](https://arxiv.org/html/2608.20186#bib.bib12)\)provides cause for optimism, but it remains to be demonstrated at EEG signal\-to\-noise ratios\.

### 4\.6Limitations

#### Single\-subject analysis

All results presented here come from one densely\-sampled individual\. The cross\-subject analysis using a 60\-participant cohort is in preparation and will address both generalisation to unseen participants and whether pooling improves within\-participant decoding, as it did in[7](https://arxiv.org/html/2608.20186#bib.bib7)\.

#### Winner’s curse

Metrics presented in this report were computed on a validation set that also drove checkpoint selection and configuration selection\. Absolute magnitudes are therefore upper bounds; cross\-condition orderings are more reliable than the values themselves\.

#### Task transfer

Silent reading is a proxy for inner speech, and RSVP reading is not natural reading\. Presenting words at a fixed location one at a time eliminates saccades and parafoveal preview, which is a considerable methodological advantage \(precise timing, no eye\-movement artefacts confounded with word identity\), but simultaneously a departure from ecological reading\. The paradigm should be understood as a compliant, scalable proxy, not as a naturalistic one\.

#### Protocol heterogeneity

Two changes occurred during data collection: the attention task and the ISI\. Repeated words and behavioural responses are excluded at the trial level, so the one\-back manipulation cannot contaminate the decoded trials, but the change in task instruction remains a source of variance\. The ISI contrast is confounded with data volume and is not interpreted as an ISI effect\.

#### Evaluation is retrieval, not generation

Like all contrastive decoders\([7](https://arxiv.org/html/2608.20186#bib.bib7);[21](https://arxiv.org/html/2608.20186#bib.bib21)\), our model identifies the most likely candidate from a supplied set rather than generating text\. Given the current state of the EEG\-to\-text field \(data limited, low signal to noise ratio\) and challenges inherent to language modelling, contrastive learning is the practically preferable approach\. In the future, methods already employed in the invasive literature, such as CTC\-trained sequence models combined with a language model, might become useful in the non\-invasive field as well\([2](https://arxiv.org/html/2608.20186#bib.bib2);[17](https://arxiv.org/html/2608.20186#bib.bib17)\)\.

#### Hardware limitations

The 19\-electrode configuration used in the present study is a low\-density montage\. Moreover, we used dry electrodes and a comparatively low sampling rate \(600 Hz\)\. Higher\-density wet\-electrode EEG might reach higher accuracies\. However, our hardware choice aligns with our belief in an experimental paradigm that prioritises scalability and subject compliance\. We used an EEG device that is easier and quicker to set up and more comfortable to wear for extended sessions than a typical, wet\-electrode research EEG device\.

### 4\.7Ongoing and future work

#### Cross\-subject modelling

A cross\-subject analysis \(60 participants, data collection ongoing\) is in preparation\. We plan to test whether multi\-subject training improves decoding as it did for M/EEG speech perception\([7](https://arxiv.org/html/2608.20186#bib.bib7)\), and whether a model pretrained at scale \(either on deep subject data or cross\-subject pooled data\) can be calibrated to a new participant \(as[21](https://arxiv.org/html/2608.20186#bib.bib21)hypothesised\)\.

#### Generalisability to inner monologue

A multimodal dataset is required to test whether semantic language processing can be decoded independently of stimulus\-induced sensory processing \(see Section[4\.5](https://arxiv.org/html/2608.20186#S4.SS5)\)\. We suggest that an experimental paradigm that prioritises scalability and subject compliance is preferable, which is why we intend to use a passive listening condition, where participants listen to audiobooks \(using the same or similar text as in the present study\), interrupted by comprehension questions to ensure subject compliance\. We propose a multi\-stage research agenda towards non\-invasive brain\-computer interfaces\. First, modality\-specific decoding needs to be established \(using large silent reading and passive listening datasets\)\. Second, cross\-modal, non\-sensory language decoding needs to be achieved \(using combined reading and listening data\)\. Finally, if the first two steps succeed, a smaller inner\-monologue dataset \(using a repetitive or generative inner\-speech paradigm, as discussed by[5](https://arxiv.org/html/2608.20186#bib.bib5)\) can be used to fine\-tune a model for inner\-monologue decoding\. We suggest an analogy to the development of large language models \(LLMs\)\. Transformer\-based LLMs required an unprecedentedly large dataset to train\. The practical solution was to use almost the entire internet as a text corpus for pretraining\. To make LLMs actually useful, expensive\-to\-obtain fine\-tuning datasets were required \(e\.g\. human preference data\)\. Similarly, we propose to prioritise scalability and subject compliance for obtaining large proxy datasets, before fine\-tuning on a smaller, difficult\-to\-obtain inner\-monologue dataset\. Reading and listening are obvious choices for proxy tasks, but watching movies or interactive video games involving language might also be considered \(although the latter would struggle with motor activity confounds\)\.

### 4\.8Conclusion

For developing non\-invasive inner\-monologue decoding, a large, labelled inner\-monologue dataset would be ideal\. However, such a dataset cannot practically be obtained\. But there is a viable proxy\. We have shown that using a modestly\-sized dataset \(∼49\{\\sim\}49h\) from a single participant, an open\-vocabulary contrastive decoder extracts word\-level information from EEG during naturalistic silent reading\. Our result is not attributable to narrative topic tracking or to positional priors, is not confined to frequent words, and is still improving log\-linearly with the amount of data collected\. The present results are the first step in a multi\-stage research agenda aimed towards inner\-monologue decoding\. For developing non\-invasive language decoding, we suggest the use of proxy tasks that prioritise scalability and subject compliance\.

## References

- Abdulghaniet al\.\(2024\)M\. M\. Abdulghani, W\. L\. Walters, and H\. K\. AbedEnhancing the classification accuracy of EEG\-informed inner speech decoder using multi\-wavelet feature and support vector machine\.IEEE Access12,pp\. 147929–147941\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2024.3474854)Cited by:[§1\.2](https://arxiv.org/html/2608.20186#S1.SS2.p3.1)\.
- Cardet al\.\(2024\)N\. S\. Card, M\. Wairagkar, C\. Iacobacci, X\. Hou, T\. Singer\-Clark, F\. R\. Willett, E\. M\. Kunz, C\. Fan, M\. Vahdati Nia, D\. R\. Deo,et al\.An accurate and rapidly calibrating speech neuroprosthesis\.New England Journal of Medicine391\(7\),pp\. 609–618\.External Links:[Document](https://dx.doi.org/10.1056/NEJMoa2314132)Cited by:[§1\.1](https://arxiv.org/html/2608.20186#S1.SS1.p1.1),[§4\.3](https://arxiv.org/html/2608.20186#S4.SS3.p1.1),[§4\.6](https://arxiv.org/html/2608.20186#S4.SS6.SSS0.Px5.p1.1)\.
- Cooneyet al\.\(2019\)C\. Cooney, A\. Korik, F\. Raffaella, and D\. CoyleClassification of imagined spoken word\-pairs using convolutional neural networks\.InThe 8th Graz BCI Conference, 2019,pp\. 338–343\.External Links:[Document](https://dx.doi.org/10.3217/978-3-85125-682-6-62)Cited by:[§1\.2](https://arxiv.org/html/2608.20186#S1.SS2.p3.1)\.
- Correiaet al\.\(2014\)J\. Correia, E\. Formisano, G\. Valente, L\. Hausfeld, B\. Jansma, and M\. BonteBrain\-based translation: fMRI decoding of spoken words in bilinguals reveals language\-independent semantic representations in anterior temporal lobe\.Journal of Neuroscience34\(1\),pp\. 332–338\.External Links:[Document](https://dx.doi.org/10.1523/JNEUROSCI.1302-13.2014)Cited by:[§1\.4](https://arxiv.org/html/2608.20186#S1.SS4.p3.1),[§4\.5](https://arxiv.org/html/2608.20186#S4.SS5.p4.1)\.
- Csakyet al\.\(2025\)R\. Csaky, M\. W\. van Es, O\. P\. Jones, and M\. WoolrichTowards decoding inner speech from EEG and MEG\.bioRxiv,pp\. 2025–10\.External Links:[Document](https://dx.doi.org/10.1101/2025.10.13.682161)Cited by:[§1\.2](https://arxiv.org/html/2608.20186#S1.SS2.p3.1),[§1\.3](https://arxiv.org/html/2608.20186#S1.SS3.p3.1),[§1\.4](https://arxiv.org/html/2608.20186#S1.SS4.p3.1),[§1\.4](https://arxiv.org/html/2608.20186#S1.SS4.p5.1),[§2\.10](https://arxiv.org/html/2608.20186#S2.SS10.p4.1),[§2\.2](https://arxiv.org/html/2608.20186#S2.SS2.p2.1),[§3\.5](https://arxiv.org/html/2608.20186#S3.SS5.p1.1),[§4\.2](https://arxiv.org/html/2608.20186#S4.SS2.p1.1),[§4\.4](https://arxiv.org/html/2608.20186#S4.SS4.p3.1),[§4\.5](https://arxiv.org/html/2608.20186#S4.SS5.p1.1),[§4\.7](https://arxiv.org/html/2608.20186#S4.SS7.SSS0.Px2.p1.1)\.
- Dashet al\.\(2020\)D\. Dash, P\. Ferrari, and J\. WangDecoding imagined and spoken phrases from non\-invasive neural \(MEG\) signals\.Frontiers in Neuroscience14,pp\. 290\.External Links:[Document](https://dx.doi.org/10.3389/fnins.2020.00290)Cited by:[§1\.2](https://arxiv.org/html/2608.20186#S1.SS2.p3.1)\.
- Défossezet al\.\(2023\)A\. Défossez, C\. Caucheteux, J\. Rapin, O\. Kabeli, and J\. KingDecoding speech perception from non\-invasive brain recordings\.Nature Machine Intelligence5\(10\),pp\. 1097–1107\.External Links:[Document](https://dx.doi.org/10.1038/s42256-023-00714-5)Cited by:[§1\.2](https://arxiv.org/html/2608.20186#S1.SS2.p2.1),[§2\.4](https://arxiv.org/html/2608.20186#S2.SS4.p1.1),[§4\.2](https://arxiv.org/html/2608.20186#S4.SS2.p2.1),[§4\.6](https://arxiv.org/html/2608.20186#S4.SS6.SSS0.Px1.p1.1),[§4\.6](https://arxiv.org/html/2608.20186#S4.SS6.SSS0.Px5.p1.1),[§4\.7](https://arxiv.org/html/2608.20186#S4.SS7.SSS0.Px1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The Llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Note:arXiv:2407\.21783Cited by:[§2\.5](https://arxiv.org/html/2608.20186#S2.SS5.p1.1)\.
- Jianget al\.\(2024\)W\. Jiang, L\. Zhao, and B\. LuLarge brain model for learning generic representations with tremendous EEG data in BCI\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 16405–16426\.Cited by:[§2\.4](https://arxiv.org/html/2608.20186#S2.SS4.p3.1)\.
- Jones and Voets \(2021\)O\. P\. Jones and N\. L\. VoetsA note on decoding elicited and self\-generated inner speech\.bioRxiv,pp\. 2021–05\.External Links:[Document](https://dx.doi.org/10.1101/2021.05.23.445249)Cited by:[§1\.3](https://arxiv.org/html/2608.20186#S1.SS3.p3.1),[§4\.4](https://arxiv.org/html/2608.20186#S4.SS4.p2.1)\.
- Kunzet al\.\(2025\)E\. M\. Kunz, B\. A\. Krasa, F\. Kamdar, D\. T\. Avansino, N\. Hahn, S\. Yoon, A\. Singh, S\. R\. Nason\-Tomaszewski, N\. S\. Card, J\. J\. Jude,et al\.Inner speech in motor cortex and implications for speech neuroprostheses\.Cell188\(17\),pp\. 4658–4673\.External Links:[Document](https://dx.doi.org/10.1016/j.cell.2025.06.015)Cited by:[§1\.1](https://arxiv.org/html/2608.20186#S1.SS1.p1.1),[§4\.3](https://arxiv.org/html/2608.20186#S4.SS3.p1.1),[§4\.4](https://arxiv.org/html/2608.20186#S4.SS4.p2.1)\.
- Lin and Hsieh \(2022\)Y\. Lin and P\. HsiehNeural decoding of speech with semantic\-based classification\.Cortex154,pp\. 231–240\.External Links:[Document](https://dx.doi.org/10.1016/j.cortex.2022.05.018)Cited by:[§1\.4](https://arxiv.org/html/2608.20186#S1.SS4.p3.1),[§4\.5](https://arxiv.org/html/2608.20186#S4.SS5.p4.1)\.
- Linget al\.\(2019\)S\. Ling, A\. C\. Lee, B\. C\. Armstrong, and A\. NestorHow are visual words represented? Insights from EEG\-based visual word decoding, feature derivation and image reconstruction\.Human Brain Mapping40\(17\),pp\. 5056–5068\.External Links:[Document](https://dx.doi.org/10.1002/hbm.24757)Cited by:[§1\.4](https://arxiv.org/html/2608.20186#S1.SS4.p3.1),[§4\.5](https://arxiv.org/html/2608.20186#S4.SS5.p1.1)\.
- Liwickiet al\.\(2022\)F\. S\. Liwicki, V\. Gupta, R\. Saini, K\. De, and M\. LiwickiRethinking the methods and algorithms for inner speech decoding and making them reproducible\.NeuroSci3\(2\),pp\. 226–244\.External Links:[Document](https://dx.doi.org/10.3390/neurosci3020017)Cited by:[§1\.2](https://arxiv.org/html/2608.20186#S1.SS2.p3.1)\.
- Loshchilov and Hutter \(2017\)I\. Loshchilov and F\. HutterDecoupled weight decay regularization\.arXiv preprint arXiv:1711\.05101\.Note:arXiv:1711\.05101Cited by:[§2\.6](https://arxiv.org/html/2608.20186#S2.SS6.p5.1)\.
- Martinet al\.\(2016\)S\. Martin, P\. Brunner, I\. Iturrate, J\. d\. R\. Millán, G\. Schalk, R\. T\. Knight, and B\. N\. PasleyWord pair classification during imagined speech using direct brain recordings\.Scientific Reports6\(1\),pp\. 25803\.External Links:[Document](https://dx.doi.org/10.1038/srep25803)Cited by:[§4\.4](https://arxiv.org/html/2608.20186#S4.SS4.p2.1)\.
- Metzgeret al\.\(2023\)S\. L\. Metzger, K\. T\. Littlejohn, A\. B\. Silva, D\. A\. Moses, M\. P\. Seaton, R\. Wang, M\. E\. Dougherty, J\. R\. Liu, P\. Wu, M\. A\. Berger,et al\.A high\-performance neuroprosthesis for speech decoding and avatar control\.Nature620\(7976\),pp\. 1037–1046\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06443-4)Cited by:[§1\.1](https://arxiv.org/html/2608.20186#S1.SS1.p1.1),[§4\.3](https://arxiv.org/html/2608.20186#S4.SS3.p1.1),[§4\.6](https://arxiv.org/html/2608.20186#S4.SS6.SSS0.Px5.p1.1)\.
- Moseset al\.\(2021\)D\. A\. Moses, S\. L\. Metzger, J\. R\. Liu, G\. K\. Anumanchipalli, J\. G\. Makin, P\. F\. Sun, J\. Chartier, M\. E\. Dougherty, P\. M\. Liu, G\. M\. Abrams,et al\.Neuroprosthesis for decoding speech in a paralyzed person with anarthria\.New England Journal of Medicine385\(3\),pp\. 217–227\.External Links:[Document](https://dx.doi.org/10.1056/NEJMoa2027540)Cited by:[§1\.1](https://arxiv.org/html/2608.20186#S1.SS1.p1.1),[§4\.3](https://arxiv.org/html/2608.20186#S4.SS3.p1.1)\.
- Mugleret al\.\(2014\)E\. M\. Mugler, J\. L\. Patton, R\. D\. Flint, Z\. A\. Wright, S\. U\. Schuele, J\. Rosenow, J\. J\. Shih, D\. J\. Krusienski, and M\. W\. SlutzkyDirect classification of all American English phonemes using signals from functional speech motor cortex\.Journal of Neural Engineering11\(3\),pp\. 035015\.External Links:[Document](https://dx.doi.org/10.1088/1741-2560/11/3/035015)Cited by:[§3\.6](https://arxiv.org/html/2608.20186#S3.SS6.p4.1)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.Learning transferable visual models from natural language supervision\.InInternational Conference on Machine Learning,pp\. 8748–8763\.Cited by:[Appendix S2](https://arxiv.org/html/2608.20186#A2.SS0.SSS0.Px4.p1.1),[§2\.6](https://arxiv.org/html/2608.20186#S2.SS6.p4.1)\.
- Satoet al\.\(2024\)M\. Sato, K\. Tomeoka, I\. Horiguchi, K\. Arulkumaran, R\. Kanai, and S\. SasaiScaling law in neural data: non\-invasive speech decoding with 175 hours of EEG data\.arXiv preprint arXiv:2407\.07595\.Note:arXiv:2407\.07595Cited by:[§1\.2](https://arxiv.org/html/2608.20186#S1.SS2.p2.1),[§1\.3](https://arxiv.org/html/2608.20186#S1.SS3.p5.1),[§3\.4](https://arxiv.org/html/2608.20186#S3.SS4.p3.1),[§4\.2](https://arxiv.org/html/2608.20186#S4.SS2.p3.1),[§4\.3](https://arxiv.org/html/2608.20186#S4.SS3.p2.1),[§4\.6](https://arxiv.org/html/2608.20186#S4.SS6.SSS0.Px5.p1.1),[§4\.7](https://arxiv.org/html/2608.20186#S4.SS7.SSS0.Px1.p1.1)\.
- Wandeltet al\.\(2024\)S\. K\. Wandelt, D\. A\. Bjånes, K\. Pejsa, B\. Lee, C\. Liu, and R\. A\. AndersenRepresentation of internal speech by single neurons in human supramarginal gyrus\.Nature Human Behaviour8\(6\),pp\. 1136–1149\.External Links:[Document](https://dx.doi.org/10.1038/s41562-024-01867-y)Cited by:[§1\.1](https://arxiv.org/html/2608.20186#S1.SS1.p1.1),[§4\.4](https://arxiv.org/html/2608.20186#S4.SS4.p2.1)\.

Supplementary Material

## Appendix S0EEG hardware

EEG was recorded from 19 scalp electrodes using a dry\-electrode system \(DSI\-24 from Wearable Sensing\), at a sampling rate of 600 Hz\. Electrode positions were: Fp1, Fp2, F3, F4, F7, F8, Fz, C3, C4, Cz, T3, T4, T5, T6, P3, P4, Pz, O1, O2 of the 10\-20 layout\. The DSI\-24 system also measures from two earclip sensors placed on the earlobes \(A1, A2\)\.

## Appendix S1Metric glossary

All metrics are word\-grouped top\-10 retrieval within fixed 512\-trial candidate pools, averaged over five independent pool draws, and all “gains” are model score minus an empirical permutation baseline \(10 permutations\), so that chance is zero\. Gains are quoted in percentage points \(pp\) throughout the report \(Section[2\.8](https://arxiv.org/html/2608.20186#S2.SS8)\)\.

Table S1:Metric glossary\.Paper termPool constructionNull / baselineWhat it measuresOverall retrieval gainRandom samples of 512 validation trials \(mixes recording runs\)Permutation of predictions within the poolTotal decoding signal, including passage\-level contextWithin\-run retrieval gain*\(primary\)*Consecutive trials from a single recording run, chunked to≤512\\leq 512Permutation within the same run \(within\-run EEG shuffle\)Discrimination among words in the same passage; excludes cross\-run topic and run\-identity driftContext\-tracking gainRandom cross\-run pools \(as overall\)Each trial’s EEG prediction replaced by a run\-mate’sThe passage\-tracking component of the overall gainContext\-independent gainn/an/aOverall gain−\-context\-tracking gainHeadroom\-normalised accuracyas overallas overall\(accuracy−\-baseline\) / \(1−\-baseline\)Rare / mid / frequent bin gainsas parent regimeper\-bin empirical permutation rateFrequency\-resolved profilePosition probes \(Sections[2\.9](https://arxiv.org/html/2608.20186#S2.SS9),[3\.2](https://arxiv.org/html/2608.20186#S3.SS2)\):position\-only,preceding\-EEGandcurrent\-trialare the nested decomposition of the within\-run gain and sum to it \(unlike the pooling regimes above, which must not be summed\);position\-corrected==within\-run gain−\-position\-only\. Only the position\-only term is non\-neural\. The dispersion diagnostic validates the position\-only null \(Supplementary[S6\.1](https://arxiv.org/html/2608.20186#A6.SS1)\)\.

Word grouping: before evaluation, each stimulus word is reduced to a “core” form by stripping leading and trailing punctuation and case\-folding, so that “Horse”, “horse” and “horse,” are treated as one candidate\. Probability mass from all trials sharing a core word is summed before ranking\.

Checkpoint selection uses the within\-run gain, averaged over the five pool draws\. Validation loss is logged but not used for selection; loss\-based selection would systematically under\-select the configurations of scientific interest\. Loss is not comparable across the within\-run negative\-masking factor, which changes the number of terms in the softmax\.

## Appendix S2Model architecture

#### EEG encoder

Input is a fixed\-length window ofC×TC\\times Tsamples \(C=19C=19or 15 channels;T=300T=300, 420 or 540 samples at 600 Hz for 0\.5/0\.7/0\.9 s windows\)\.

- •*Global pathway*: four masked 1\-D convolutional blocks with output channels \(48, 96, 192, 384\), kernel sizes \(11, 9, 7, 5\) and strides \(1, 2, 2, 2\), each block being Conv1d→\\rightarrowmasked BatchNorm→\\rightarrowELU→\\rightarrow20% dropout\. Time is then collapsed by a learnable attention pooling \(a linear query producing softmax weights over valid timepoints\)\.
- •*Local pathway*: three stride\-1 blocks with output channels \(48, 96, 192\) and kernel sizes \(11, 9, 7\), followed by masked adaptive average pooling into 6 temporal bins, preserving coarse within\-window timing\.
- •*Head*: concatenation of the two pathway outputs→\\rightarrowtwo fully\-connected layers of 512 units \(BatchNorm, ELU, 20% dropout\)→\\rightarrowlinear projection to a 256\-dimensional per\-trial feature vector\.

All normalisation and pooling operations are mask\-aware, computed over valid timepoints only, so zero\-padding cannot leak into a trial’s feature vector\.

#### Sequence model \(optional\)

Fourtorch\.nn\.TransformerEncoderLayerblocks with d\_model==256, 8 attention heads, feed\-forward dimension 1,024, dropout 0\.4, GELU activation, pre\-norm, learned positional embeddings \(maximum sequence length 1,400\), final LayerNorm\. A causal mask restricts positioniito attend to positionsj≤ij\\leq i, mirroring the causal attention of the language model that produced the targets\.

#### Excluded trials

Trials failing the amplitude criterion and target/false\-alarm trials \(in the early data subset using a one\-back attention task instead of comprehension questions\) are never passed through the encoder \(so they cannot contaminate BatchNorm statistics\) and their position in the sequence is filled with a learnable\[MASK\]token, trained end\-to\-end\. The transformer therefore “knows” that a word occurred at that position but that its EEG is unavailable, and the alignment with the target sequence is preserved\. These positions are excluded from the loss and from all evaluation\.

#### Projections and loss

A bias\-free linear layer maps encoder/transformer output to a 256\-dimensional shared space; a second bias\-free linear layer maps the 4,096\-dimensional LLM target into the same space\. Both are L2\-normalised\. The symmetric CLIP cross\-entropy\([20](https://arxiv.org/html/2608.20186#bib.bib20)\)is computed over all valid positions in the batch, with a learnable log\-temperature initialised atlog⁡\(1/0\.07\)\\log\(1/0\.07\)and clamped atexp≤100\\exp\\leq 100\.

#### Batching

One batch item is a complete recording run \(∼600\{\\sim\}600trials\), so a batch of 4 runs contributes ca\. 2,200 valid in\-batch negatives \(assuming ca\. 90% valid trials\)\.

## Appendix S3Hyperparameter grids

Fixed across all reported sweeps:deep participant only; target embeddings from Llama\-3\.1\-8B, 4,096\-dimensional; projection dimension 256; 100 epochs; AdamW with weight decay 0\.1; cosine schedule with 5% warm\-up; no gradient clipping; band\-pass 1–80 Hz \(4th\-order Butterworth\); band\-stop±1\\pm 1Hz around 50 Hz mains \(4th\-order Butterworth\); amplitude rejection at 100μ\\muV; target/false\-alarm trials excluded; validation fraction 0\.2; content\-aware text split with 10\-gram overlap threshold; evaluation pool size 512; 10 permutations; 5 evaluation seeds; 3 frequency bins; position probes \(single\-pass all\-masked, and grouped leave\-one\-out mask\-current\) evaluated for every fit\.

Table S2:Sweep 1, 576 fits=2×3×3×2×2×2×2×2=2\\times 3\\times 3\\times 2\\times 2\\times 2\\times 2\\times 2\.FactorLevelsLLM embedding layer0, 20Window lockonset, offset, centreISI subset\{0 ms\}, \{100 ms\}, \{0, 100 ms\}Sequence modelnone, transformerNormalisationscale\-only, baseline subtraction\+\+scaleTemporal\-shift augmentation0,±100\\pm 100msMask within\-run negativesFalse, TrueTrainable temperatureFalse, True*\(fixed\)*window length0\.5 s*\(fixed\)*batch size / learning rate4 / 1e\-4Sweep 2 \(data scaling\),2×102\\times 10fits\.Fixed configuration: onset lock, 0\.5 s window, no augmentation, scale\-only normalisation, both ISIs, no negative masking, trainable temperature, batch 4, learning rate 1e\-4\. Swept: the fraction of training runs retained,∈\{0\.1,…,1\.0\}\\in\\\{0\.1,\\ldots,1\.0\\\}, separately for \(layer 20\+\+transformer\) and \(layer 0, no transformer\)\. Only training runs are subsampled; the content\-aware split is computed beforehand, so the same validation trials are scored at every ratio\.

Table S3:Sweep 3, 192 fits=2×3×2×2×2×2×2=2\\times 3\\times 2\\times 2\\times 2\\times 2\\times 2\.FactorLevelsLLM embedding layer0, 20Window length0\.5, 0\.7, 0\.9 sTemporal\-shift augmentation0,±50\\pm 50msChannel setall 19; O1/O2/T5/T6 removedSequence modelnone, transformerBatch size4, 8Learning rate1e\-4, 1e\-5*\(fixed\)*window lock / ISI / normalisationonset / both / scale\-only
## Appendix S4Absolute accuracies

We report gains rather than raw accuracies because the raw top\-10 accuracy in a 512\-trial pool is dominated by the frequency structure of English prose: a frequent word occupies many candidate slots and accumulates probability mass irrespective of the EEG\. The permutation baseline absorbs this, which is why chance is zero for every gain\. The below raw accuracy numbers are provided for completeness, and are not directly comparable across studies without matching the pool construction\.

VV==mean number ofdistinct core wordsper evaluation pool, which is not the pool size: words are grouped before ranking, soVVis smaller than the pool size of 512\.10/V10/Vis the ‘uniform over candidate words’ reference and understates chance, because a word filling many trial slots draws mass from all of them; the measured permutation baseline \(the ‘implied empirical chance’ column\) sits above it and is the figure to quote\. The gain column is in percentage points, so gain==implied model top\-10−\-implied empirical chance \(up to rounding\); norm\-acc is the headroom\-normalised accuracy, a ratio rather than a percentage\.

Table S4:Sweep 1, random cross\-run pools: absolute accuracies\. Across all 576 fits: meanV=299V=299distinct core words per pool, so the naive uniform\-over\-words reference is10/V=3\.34%10/V=3\.34\\%, against a measured empirical chance of 14\.31% and a model top\-10 of 25\.96%\.Configurationnn\(gain\)nn\(abs\)Gain\(pp\)norm\-accImpliedempiricalchanceImpliedmodeltop\-10VV\(distinctwords / pool\)10/V10/Vreferenceall fits57657611\.640\.135214\.31%25\.96%2993\.34%noTF\_L01441447\.920\.093015\.09%23\.01%2993\.34%noTF\_L2014414410\.950\.127214\.19%25\.14%3003\.34%TF\_L01441447\.900\.092114\.42%22\.32%2993\.34%TF\_L2014414419\.810\.228613\.54%33\.35%3003\.34%Table S5:Sweep 1, within\-run / strict \(by\-run pools\): absolute accuracies\.Configurationnn\(gain\)nn\(abs\)Gain\(pp\)norm\-accImpliedempiricalchanceImpliedmodeltop\-10VV\(distinctwords / pool\)10/V10/Vreferenceall fits5765766\.740\.082218\.30%25\.03%2464\.07%noTF\_L01441447\.480\.090417\.43%24\.91%2464\.07%noTF\_L201441445\.470\.067719\.69%25\.16%2464\.07%TF\_L01441446\.180\.074016\.53%22\.71%2464\.07%TF\_L201441447\.810\.096519\.54%27\.35%2464\.07%Table S6:Sweep 3, random cross\-run pools: absolute accuracies\. Across all 192 fits: meanV=301V=301distinct core words per pool, so the naive uniform\-over\-words reference is10/V=3\.32%10/V=3\.32\\%, against a measured empirical chance of 15\.75% and a model top\-10 of 29\.28%\.Configurationnn\(gain\)nn\(abs\)Gain\(pp\)norm\-accImpliedempiricalchanceImpliedmodeltop\-10VV\(distinctwords / pool\)10/V10/Vreferenceall fits19219213\.540\.158315\.75%29\.28%3013\.32%noTF\_L048487\.110\.085717\.78%24\.89%3013\.32%noTF\_L20484813\.690\.160415\.32%29\.01%3023\.31%TF\_L048487\.930\.094516\.46%24\.38%3013\.32%TF\_L20484825\.420\.292513\.44%38\.86%3013\.32%all\_channels969615\.240\.177615\.21%30\.45%3013\.32%no\_visual\_channels969611\.830\.139016\.29%28\.12%3023\.32%Table S7:Sweep 3, within\-run / strict \(by\-run pools\): absolute accuracies\.Configurationnn\(gain\)nn\(abs\)Gain\(pp\)norm\-accImpliedempiricalchanceImpliedmodeltop\-10VV\(distinctwords / pool\)10/V10/Vreferenceall fits1921927\.760\.096319\.96%27\.72%2444\.09%noTF\_L048487\.370\.091920\.34%27\.71%2444\.09%noTF\_L2048486\.710\.083920\.88%27\.59%2454\.09%TF\_L048486\.930\.084918\.73%25\.66%2444\.09%TF\_L20484810\.040\.124419\.89%29\.93%2444\.09%all\_channels96969\.240\.114019\.39%28\.63%2444\.09%no\_visual\_channels96966\.280\.078520\.53%26\.82%2444\.09%
## Appendix S5Complete result tables

Table S8:Sweep 1: decomposition by factor \(mean±\\pmSD\)\.n=288n=288per level for two\-level factors, 192 for three\-level factors\. Gains in percentage points \(pp\); norm\-acc is a ratio\.FactorLevelWithin\-rungain \(pp\)Overallgain \(pp\)Context\-trackinggain \(pp\)Context\-independentgain \(pp\)norm\-accval\-lossEmbedding layer06\.83±2\.226\.83\\pm 2\.227\.91±2\.087\.91\\pm 2\.081\.46±0\.811\.46\\pm 0\.816\.45±2\.076\.45\\pm 2\.070\.09257\.645206\.64±3\.136\.64\\pm 3\.1315\.38±6\.5715\.38\\pm 6\.575\.24±1\.715\.24\\pm 1\.7110\.14±5\.8010\.14\\pm 5\.800\.17796\.648Window lockonset6\.80±2\.866\.80\\pm 2\.8611\.74±6\.3411\.74\\pm 6\.343\.38±2\.403\.38\\pm 2\.408\.37±4\.888\.37\\pm 4\.880\.13637\.138offset6\.50±2\.486\.50\\pm 2\.4811\.42±6\.0011\.42\\pm 6\.003\.37±2\.313\.37\\pm 2\.318\.05±4\.578\.05\\pm 4\.570\.13277\.169centre6\.91±2\.786\.91\\pm 2\.7811\.77±6\.1011\.77\\pm 6\.103\.30±2\.253\.30\\pm 2\.258\.47±4\.748\.47\\pm 4\.740\.13677\.133ISI subset0 ms6\.86±2\.676\.86\\pm 2\.6711\.67±5\.5511\.67\\pm 5\.553\.37±2\.183\.37\\pm 2\.188\.30±4\.388\.30\\pm 4\.380\.13567\.162100 ms6\.14±2\.516\.14\\pm 2\.519\.33±4\.559\.33\\pm 4\.552\.38±1\.472\.38\\pm 1\.476\.96±3\.686\.96\\pm 3\.680\.10937\.364both7\.21±2\.847\.21\\pm 2\.8413\.93±7\.1613\.93\\pm 7\.164\.30±2\.724\.30\\pm 2\.729\.63±5\.569\.63\\pm 5\.560\.16096\.914Sequence modelnone6\.48±2\.466\.48\\pm 2\.469\.43±3\.009\.43\\pm 3\.002\.82±2\.172\.82\\pm 2\.176\.62±2\.496\.62\\pm 2\.490\.11017\.397transformer7\.00±2\.937\.00\\pm 2\.9313\.86±7\.5313\.86\\pm 7\.533\.88±2\.353\.88\\pm 2\.359\.98±5\.739\.98\\pm 5\.730\.16036\.896Normalisationbaseline\+\+scale6\.23±2\.636\.23\\pm 2\.6311\.07±6\.0411\.07\\pm 6\.043\.33±2\.313\.33\\pm 2\.317\.74±4\.627\.74\\pm 4\.620\.12877\.187scale only7\.24±2\.717\.24\\pm 2\.7112\.21±6\.2012\.21\\pm 6\.203\.36±2\.333\.36\\pm 2\.338\.85±4\.788\.85\\pm 4\.780\.14177\.106Temporal shift08\.38±2\.368\.38\\pm 2\.3613\.07±5\.9613\.07\\pm 5\.963\.16±2\.253\.16\\pm 2\.259\.91±4\.539\.91\\pm 4\.530\.15127\.131±100\\pm 100ms5\.09±1\.935\.09\\pm 1\.9310\.22±5\.9910\.22\\pm 5\.993\.54±2\.373\.54\\pm 2\.376\.69±4\.366\.69\\pm 4\.360\.11937\.162Mask within\-run neg\.False8\.02±2\.658\.02\\pm 2\.6512\.91±7\.4912\.91\\pm 7\.492\.93±2\.232\.93\\pm 2\.239\.97±5\.619\.97\\pm 5\.610\.15017\.502True5\.45±2\.105\.45\\pm 2\.1010\.38±4\.0310\.38\\pm 4\.033\.76±2\.333\.76\\pm 2\.336\.62±2\.776\.62\\pm 2\.770\.12046\.791Trainable temperatureFalse6\.76±2\.716\.76\\pm 2\.7111\.77±6\.2811\.77\\pm 6\.283\.39±2\.323\.39\\pm 2\.328\.38±4\.828\.38\\pm 4\.820\.13717\.097True6\.71±2\.716\.71\\pm 2\.7111\.52±6\.0111\.52\\pm 6\.013\.30±2\.323\.30\\pm 2\.328\.22±4\.638\.22\\pm 4\.630\.13347\.196Context conditionL0, no TF7\.48±2\.027\.48\\pm 2\.027\.92±1\.867\.92\\pm 1\.861\.02±0\.481\.02\\pm 0\.486\.89±1\.866\.89\\pm 1\.860\.09307\.652L0, TF6\.18±2\.236\.18\\pm 2\.237\.90±2\.287\.90\\pm 2\.281\.89±0\.851\.89\\pm 0\.856\.01±2\.186\.01\\pm 2\.180\.09217\.638L20, no TF5\.47±2\.455\.47\\pm 2\.4510\.95±3\.1510\.95\\pm 3\.154\.61±1\.654\.61\\pm 1\.656\.34±2\.986\.34\\pm 2\.980\.12727\.142L20, TF7\.81±3\.307\.81\\pm 3\.3019\.81±6\.1019\.81\\pm 6\.105\.86±1\.545\.86\\pm 1\.5413\.94±5\.4313\.94\\pm 5\.430\.22866\.154Table S9:Sweep 1: within\-run gain by frequency bin\. Gains in percentage points \(pp\)\.FactorLevelRareMidFrequentEmbedding layer02\.81±1\.532\.81\\pm 1\.532\.80±1\.442\.80\\pm 1\.447\.43±2\.397\.43\\pm 2\.39204\.75±3\.104\.75\\pm 3\.104\.81±2\.964\.81\\pm 2\.966\.92±3\.166\.92\\pm 3\.16Window lockonset3\.27±2\.583\.27\\pm 2\.583\.41±2\.513\.41\\pm 2\.517\.31±3\.007\.31\\pm 3\.00offset3\.98±2\.523\.98\\pm 2\.523\.92±2\.373\.92\\pm 2\.376\.88±2\.566\.88\\pm 2\.56centre4\.08±2\.724\.08\\pm 2\.724\.07±2\.684\.07\\pm 2\.687\.34±2\.857\.34\\pm 2\.85ISI subset0 ms4\.30±2\.564\.30\\pm 2\.564\.29±2\.534\.29\\pm 2\.537\.26±2\.797\.26\\pm 2\.79100 ms2\.75±2\.072\.75\\pm 2\.072\.83±2\.032\.83\\pm 2\.036\.67±2\.676\.67\\pm 2\.67both4\.29±2\.894\.29\\pm 2\.894\.28±2\.724\.28\\pm 2\.727\.59±2\.907\.59\\pm 2\.90Sequence modelnone3\.36±1\.773\.36\\pm 1\.773\.41±1\.783\.41\\pm 1\.786\.94±2\.646\.94\\pm 2\.64transformer4\.19±3\.224\.19\\pm 3\.224\.19±3\.064\.19\\pm 3\.067\.41±2\.967\.41\\pm 2\.96Normalisationbaseline\+\+scale3\.47±2\.443\.47\\pm 2\.443\.50±2\.423\.50\\pm 2\.426\.64±2\.736\.64\\pm 2\.73scale only4\.09±2\.784\.09\\pm 2\.784\.10±2\.624\.10\\pm 2\.627\.71±2\.797\.71\\pm 2\.79Temporal shift04\.76±2\.814\.76\\pm 2\.814\.74±2\.654\.74\\pm 2\.658\.92±2\.428\.92\\pm 2\.42±100\\pm 100ms2\.80±2\.002\.80\\pm 2\.002\.87±2\.022\.87\\pm 2\.025\.43±1\.965\.43\\pm 1\.96Context conditionL0, no TF3\.22±1\.493\.22\\pm 1\.493\.17±1\.423\.17\\pm 1\.428\.12±2\.188\.12\\pm 2\.18L0, TF2\.40±1\.462\.40\\pm 1\.462\.42±1\.362\.42\\pm 1\.366\.74±2\.406\.74\\pm 2\.40L20, no TF3\.51±2\.003\.51\\pm 2\.003\.64±2\.053\.64\\pm 2\.055\.75±2\.535\.75\\pm 2\.53L20, TF5\.98±3\.505\.98\\pm 3\.505\.97±3\.265\.97\\pm 3\.268\.08±3\.308\.08\\pm 3\.30#### Sweep 2 — data scaling

Per\-ratio selected\-epoch scalars for the two scaling arms\. The log\-linear fits over the ten ratios, including the position\-probe traces, are reported in Table[4](https://arxiv.org/html/2608.20186#S3.T4)\. One fit per ratio: the slope is the reportable quantity, and the per\-ratio levels are selection\-optimistic upper bounds\. In fits trained on small fractions of the data, the validation loss reached its minimum early and rose thereafter \(a sign of over\-fitting\), while the within\-run gain continued to improve; gain\-based checkpoint selection \(Section[2\.6](https://arxiv.org/html/2608.20186#S2.SS6)\) therefore matters most at small data volumes\.

Table S10:Sweep 3: decomposition by factor \(mean±\\pmSD\)\. Gains in percentage points \(pp\); norm\-acc is a ratio\. Distribution across all 192 fits: within\-run gain7\.76±2\.897\.76\\pm 2\.89pp \(median 7\.52, min 1\.99, max 16\.53\); overall gain13\.54±8\.5213\.54\\pm 8\.52pp; context\-tracking gain3\.41±2\.643\.41\\pm 2\.64pp\.FactorLevelWithin\-rungain \(pp\)Overallgain \(pp\)Context\-trackinggain \(pp\)Context\-independentgain \(pp\)norm\-accval\-lossEmbedding layer07\.15±2\.177\.15\\pm 2\.177\.52±2\.357\.52\\pm 2\.351\.06±0\.481\.06\\pm 0\.486\.46±2\.196\.46\\pm 2\.190\.09017\.969208\.38±3\.368\.38\\pm 3\.3619\.55±8\.2019\.55\\pm 8\.205\.76±1\.605\.76\\pm 1\.6013\.79±6\.8113\.79\\pm 6\.810\.22647\.134Window length0\.5 s6\.86±2\.736\.86\\pm 2\.7312\.24±8\.0612\.24\\pm 8\.063\.26±2\.523\.26\\pm 2\.528\.97±5\.898\.97\\pm 5\.890\.14347\.6120\.7 s8\.32±2\.858\.32\\pm 2\.8514\.14±8\.7814\.14\\pm 8\.783\.41±2\.683\.41\\pm 2\.6810\.74±6\.4410\.74\\pm 6\.440\.16517\.5420\.9 s8\.11±2\.908\.11\\pm 2\.9014\.22±8\.6914\.22\\pm 8\.693\.57±2\.733\.57\\pm 2\.7310\.66±6\.3110\.66\\pm 6\.310\.16647\.499Temporal shift08\.31±2\.888\.31\\pm 2\.8813\.93±8\.4113\.93\\pm 8\.413\.30±2\.563\.30\\pm 2\.5610\.63±6\.1910\.63\\pm 6\.190\.16257\.565±50\\pm 50ms7\.22±2\.807\.22\\pm 2\.8013\.14±8\.6613\.14\\pm 8\.663\.52±2\.723\.52\\pm 2\.729\.61±6\.289\.61\\pm 6\.280\.15407\.537Channel setall 199\.24±2\.669\.24\\pm 2\.6615\.24±8\.6315\.24\\pm 8\.633\.48±2\.683\.48\\pm 2\.6811\.75±6\.2911\.75\\pm 6\.290\.17767\.480O1/O2/T5/T6 removed6\.28±2\.296\.28\\pm 2\.2911\.83±8\.0911\.83\\pm 8\.093\.34±2\.613\.34\\pm 2\.618\.49±5\.788\.49\\pm 5\.780\.13907\.622Sequence modelnone7\.04±2\.497\.04\\pm 2\.4910\.40±4\.9110\.40\\pm 4\.912\.74±2\.292\.74\\pm 2\.297\.66±3\.217\.66\\pm 3\.210\.12307\.741transformer8\.49±3\.088\.49\\pm 3\.0816\.67±10\.1016\.67\\pm 10\.104\.08±2\.804\.08\\pm 2\.8012\.59±7\.4612\.59\\pm 7\.460\.19357\.362Batch size48\.21±2\.888\.21\\pm 2\.8813\.66±8\.2613\.66\\pm 8\.263\.14±2\.333\.14\\pm 2\.3310\.51±6\.2310\.51\\pm 6\.230\.15967\.31287\.31±2\.847\.31\\pm 2\.8413\.41±8\.8113\.41\\pm 8\.813\.68±2\.903\.68\\pm 2\.909\.73±6\.269\.73\\pm 6\.260\.15697\.790Learning rate1e\-49\.39±2\.819\.39\\pm 2\.8116\.66±9\.6616\.66\\pm 9\.664\.04±2\.994\.04\\pm 2\.9912\.62±7\.0412\.62\\pm 7\.040\.19187\.6001e\-56\.14±1\.896\.14\\pm 1\.8910\.41±5\.7310\.41\\pm 5\.732\.79±2\.062\.79\\pm 2\.067\.62±4\.027\.62\\pm 4\.020\.12487\.502Context conditionL0, no TF7\.37±2\.337\.37\\pm 2\.337\.11±2\.577\.11\\pm 2\.570\.65±0\.290\.65\\pm 0\.296\.46±2\.386\.46\\pm 2\.380\.08577\.952L0, TF6\.93±1\.996\.93\\pm 1\.997\.93±2\.057\.93\\pm 2\.051\.48±0\.191\.48\\pm 0\.196\.45±2\.006\.45\\pm 2\.000\.09457\.986L20, no TF6\.71±2\.626\.71\\pm 2\.6213\.69±4\.4613\.69\\pm 4\.464\.84±1\.224\.84\\pm 1\.228\.85±3\.518\.85\\pm 3\.510\.16047\.529L20, TF10\.04±3\.2110\.04\\pm 3\.2125\.42±6\.7625\.42\\pm 6\.766\.68±1\.406\.68\\pm 1\.4018\.73±5\.6118\.73\\pm 5\.610\.29256\.738Table S11:Sweep 3: within\-run gain by frequency bin\. Gains in percentage points \(pp\)\.FactorLevelRareMidFrequentEmbedding layer03\.39±1\.263\.39\\pm 1\.263\.51±1\.273\.51\\pm 1\.277\.63±2\.327\.63\\pm 2\.32207\.05±4\.177\.05\\pm 4\.176\.85±3\.816\.85\\pm 3\.818\.56±3\.308\.56\\pm 3\.30Window length0\.5 s3\.64±2\.753\.64\\pm 2\.753\.80±2\.713\.80\\pm 2\.717\.26±2\.787\.26\\pm 2\.780\.7 s5\.90±3\.565\.90\\pm 3\.565\.70±3\.285\.70\\pm 3\.288\.65±2\.858\.65\\pm 2\.850\.9 s6\.11±3\.846\.11\\pm 3\.846\.04±3\.416\.04\\pm 3\.418\.38±2\.878\.38\\pm 2\.87Temporal shift05\.62±3\.745\.62\\pm 3\.745\.61±3\.415\.61\\pm 3\.418\.65±2\.878\.65\\pm 2\.87±50\\pm 50ms4\.81±3\.384\.81\\pm 3\.384\.75±3\.114\.75\\pm 3\.117\.54±2\.807\.54\\pm 2\.80Channel setall 196\.11±3\.966\.11\\pm 3\.966\.11±3\.496\.11\\pm 3\.499\.65±2\.609\.65\\pm 2\.60O1/O2/T5/T6 removed4\.33±2\.904\.33\\pm 2\.904\.25±2\.794\.25\\pm 2\.796\.54±2\.256\.54\\pm 2\.25Sequence modelnone4\.43±2\.314\.43\\pm 2\.314\.49±2\.184\.49\\pm 2\.187\.37±2\.597\.37\\pm 2\.59transformer6\.00±4\.376\.00\\pm 4\.375\.87±4\.005\.87\\pm 4\.008\.82±2\.998\.82\\pm 2\.99Context conditionL0, no TF3\.63±1\.333\.63\\pm 1\.333\.82±1\.383\.82\\pm 1\.387\.84±2\.497\.84\\pm 2\.49L0, TF3\.15±1\.143\.15\\pm 1\.143\.21±1\.083\.21\\pm 1\.087\.41±2\.137\.41\\pm 2\.13L20, no TF5\.24±2\.785\.24\\pm 2\.785\.17±2\.605\.17\\pm 2\.606\.91±2\.646\.91\\pm 2\.64L20, TF8\.86±4\.548\.86\\pm 4\.548\.52±4\.108\.52\\pm 4\.1010\.22±3\.0810\.22\\pm 3\.08Table S12:Sweep 3: channel\-ablation interactions \(within\-run gain\)\. Gains in percentage points \(pp\)\. The ordering of every crossed factor is preserved across channel sets; the ablation is additive rather than interactive in sign, though its relative magnitude is larger in the non\-contextual and no\-transformer cells\.Crossed factorLevelAll 19 channelsO1/O2/T5/T6 removedEmbedding layer08\.76±1\.618\.76\\pm 1\.615\.54±1\.265\.54\\pm 1\.26209\.72±3\.369\.72\\pm 3\.367\.03±2\.817\.03\\pm 2\.81Sequence modelnone8\.58±2\.228\.58\\pm 2\.225\.51±1\.675\.51\\pm 1\.67transformer9\.91±2\.919\.91\\pm 2\.917\.06±2\.577\.06\\pm 2\.57Window length0\.5 s8\.36±2\.498\.36\\pm 2\.495\.35±2\.085\.35\\pm 2\.080\.7 s9\.77±2\.649\.77\\pm 2\.646\.87±2\.286\.87\\pm 2\.280\.9 s9\.59±2\.719\.59\\pm 2\.716\.63±2\.296\.63\\pm 2\.29Batch size49\.72±2\.629\.72\\pm 2\.626\.71±2\.296\.71\\pm 2\.2988\.77±2\.648\.77\\pm 2\.645\.86±2\.245\.86\\pm 2\.24Learning rate1e\-411\.12±2\.2211\.12\\pm 2\.227\.65±2\.197\.65\\pm 2\.191e\-57\.36±1\.477\.36\\pm 1\.474\.92±1\.424\.92\\pm 1\.42

## Appendix S6Position probes — implementation caveats and per\-factor tables

All conventions as in Sections[2\.9](https://arxiv.org/html/2608.20186#S2.SS9)and[3\.2](https://arxiv.org/html/2608.20186#S3.SS2): gains in percentage points \(pp\); the three probe terms are nested and sum to the within\-run gain; position\-corrected==within\-run−\-position\-only; a near\-zero position\-only gain is evidence only where dispersion is clearly above zero \([S6\.1](https://arxiv.org/html/2608.20186#A6.SS1); dispersion is a cosine distance, not a percentage\); transformer\-free cells are zero by construction and serve as a built\-in correctness check on the probe machinery\.

### S6\.1Implementation caveats and validity of the position\-only null

The preceding\-EEG level of the probe ladder \(Section[2\.9](https://arxiv.org/html/2608.20186#S2.SS9), level ii\) requires each trial’s prediction to be calculated while that trial’s own EEG is masked and all other trials’ EEG is intact\. Computing this exactly would require one transformer pass per position \(∼600\{\\sim\}600per run\)\. We approximate it with a grouped leave\-one\-out: positions are partitioned into 16 interleaved groups \(position index modulo 16\), one pass is run per group with the whole group masked, and each position’s prediction is read from the pass in which it was itself masked\. The property required for the probe to be valid \(that the read\-out position never sees its own EEG\) holds\. The approximation is that the same pass also masks the preceding positions congruent to the read\-out position modulo 16, so∼1/16\{\\sim\}1/16\(≈6%\\approx 6\\%\) of the preceding context is masked that the unablated model would have seen\. The position\-only probe masks every position in a single pass and involves no approximation\.

Two caveats apply to the reading of the probe terms\. First, because the grouped leave\-one\-out also masks a small fraction of the preceding positions, the preceding\-EEG term is biased down and the current\-trial term up by the same amount; the position\-only and position\-corrected terms are exact and unaffected, which is why the load\-bearing claims rest on those two\. Second, masking every trial is out of distribution, so a near\-zero position\-only gain is only evidence if the all\-masked model still produces position\-dependent output rather than collapsing to a constant \(a constant prediction would score zero against a within\-run shuffle, regardless of the model’s decoding capability\)\. We therefore report a dispersion diagnostic, the mean cosine distance of the position\-only predictions from their own centroid, alongside the position\-only gain\. A near\-zero gain is informative only where dispersion is above zero\. For transformer\-free configurations, both probes return exactly zero by construction \(every position is processed independently, so both probes yield the same constant vector at every position\); this serves as a built\-in correctness check on the probe machinery rather than as an empirical result\.

Table S13:Sweep 1: probes by architecture \(mean over fits\)\.LevelnnWithin\-rungain \(pp\)Position\-only \(pp\)Preceding\-EEG \(pp\)Current\-trial \(pp\)Position\-corrected \(pp\)Dispersionall fits5766\.740\.700\.495\.556\.040\.0645no transformer2886\.480\.000\.006\.486\.480\.0000transformer2887\.001\.400\.984\.625\.600\.1290Table S14:Sweep 1: position\-corrected within\-run gain per word\-frequency bin\.LevelnnRareMidFrequentall fits5763\.203\.246\.46L0, no TF1443\.223\.178\.12L20, no TF1443\.513\.645\.75L0, TF1442\.402\.426\.69L20, TF1443\.693\.725\.26Table S15:Sweep 3: probes by target type and architecture\.LevelnnWithin\-rungain \(pp\)Position\-only \(pp\)Preceding\-EEG \(pp\)Current\-trial \(pp\)Position\-corrected \(pp\)Dispersionall fits1927\.760\.730\.586\.467\.030\.0642L0, no TF487\.370\.000\.007\.377\.370\.0000L20, no TF486\.710\.000\.006\.716\.710\.0000L0, TF486\.930\.07−0\.02\-0\.026\.886\.860\.1495L20, TF4810\.042\.852\.344\.867\.200\.1071Table S16:Sweep 3: probes by channel set\. The position\-only term is essentially unchanged by the ablation while the current\-trial term drops by∼37%\{\\sim\}37\\%: the ablation removes neural signal, not positional information, which is why the position\-corrected gain falls in step with the raw within\-run gain \(Table[5](https://arxiv.org/html/2608.20186#S3.T5)\)\.LevelnnWithin\-rungain \(pp\)Position\-only \(pp\)Preceding\-EEG \(pp\)Current\-trial \(pp\)Position\-corrected \(pp\)Dispersionall 19 channels969\.240\.770\.567\.928\.480\.0659O1/O2/T5/T6 removed966\.280\.690\.604\.995\.590\.0624Table S17:Sweep 3: position\-corrected within\-run gain per word\-frequency bin\.LevelnnRareMidFrequentall fits1924\.674\.667\.34L0, no TF483\.633\.827\.84L20, no TF485\.245\.176\.91L0, TF483\.193\.207\.33L20, TF486\.626\.467\.28all 19 channels965\.535\.588\.85O1/O2/T5/T6 removed963\.803\.745\.83

## Appendix S7Control checks on technical factors

#### Within\-run negative masking \(Sweep 1\)

Masking reduced the within\-run gain in every crossed cell, and the ordering of the scientific factors was preserved in both masking conditions with one exception: For the non\-contextual target, the overall gain was slightly*higher*under masking \(8\.37 vs 7\.45 pp\), reversing the direction seen for contextual targets \(12\.39 vs 18\.37 pp\)\. The primary metric \(within\-run gain\) did not reverse anywhere\. Because masking changes the number of terms in the contrastive softmax, validation losses are not comparable across this factor\.

Within\-run gain in percentage points \(pp\)\.

Table S18:Sweep 1: within\-run gain by within\-run negative masking\.Crossed factorLevelMasking offMasking onEmbedding layer07\.17±2\.217\.17\\pm 2\.216\.49±2\.186\.49\\pm 2\.18208\.86±2\.798\.86\\pm 2\.794\.42±1\.394\.42\\pm 1\.39Sequence modelnone7\.41±2\.027\.41\\pm 2\.025\.54±2\.515\.54\\pm 2\.51transformer8\.62±3\.058\.62\\pm 3\.055\.37±1\.595\.37\\pm 1\.59Temporal shift09\.88±1\.939\.88\\pm 1\.936\.88±1\.726\.88\\pm 1\.72±100\\pm 100ms6\.15±1\.846\.15\\pm 1\.844\.03±1\.344\.03\\pm 1\.34Context conditionL0, no TF7\.70±2\.037\.70\\pm 2\.037\.25±1\.997\.25\\pm 1\.99L20, TF10\.61±2\.3710\.61\\pm 2\.375\.01±0\.655\.01\\pm 0\.65
#### Trainable temperature \(Sweep 1\)

No effect on any metric in any cell; all differences were at or below∼0\.1\{\\sim\}0\.1pp on the within\-run gain\.

#### Temperature runaway

Final values of the learnable temperature \(logit scale\): Sweep 1 mean 15\.4 \(max 19\.3\); Sweep 3 mean 15\.7 \(max 19\.3\), against a clamp of 100\. No fit in any sweep approached runaway, so the temperature\-dependent word\-grouped metrics are undistorted throughout\.

#### Training\-budget diagnostics \(Sweep 3\)

Selected checkpoint epoch was80±1580\\pm 15at learning rate 1e\-4 and91±791\\pm 7at 1e\-5, and83±1383\\pm 13at batch 4 versus87±1387\\pm 13at batch 8, out of 100 epochs\. Both the learning\-rate and the batch\-size effect are therefore partly attributable to an incomplete training budget rather than to an intrinsic optimum\.

#### Training dynamics

Validation loss frequently showed an extremum\-then\-reversal pattern in the non\-contextual configurations \(minimum around epoch 7–13, rising thereafter\) while the within\-run gain continued to improve, which is why checkpoint selection is not performed on loss\.

![Training curves of within-run gain and validation loss across epochs for Sweeps 1 and 3](https://arxiv.org/html/2608.20186v1/figures/figure_S1_training_dynamics.png)Figure S1:Training dynamics: The validation\-loss minimum and the within\-run\-gain maximum do not coincide, which is why checkpoints are selected on the gain\.\(a, b\)Sweep 1;\(c, d\)Sweep 3\. Left column: within\-run retrieval gain against epoch\. Right column: validation loss against epoch\. Lines are means across fits within each of the four target\-type×\\timesarchitecture configurations, bands are±1\\pm 1SD, and the extremum of each group\-mean curve is marked \(▲\\blacktrianglegain maximum,▼\\blacktriangledownloss minimum\)\. Epoch 0 is the pre\-training validation pass, so every gain curve starts at∼0\{\\sim\}0by construction\. The Sweep 1 row is restricted to fits with within\-run negative masking off: masking changes the number of terms in the contrastive softmax, so validation loss is on a different scale in the two masking conditions and averaging across them would be meaningless\. The gain panel is restricted identically so both panels in the row describe the same set of fits\. Sweep 3 holds masking off throughout, so no filter applies there\. In the non\-contextual configurations the validation loss reaches a minimum around epoch 7–13 and rises thereafter while the within\-run gain keeps improving\. Selecting checkpoints on validation loss would therefore pick early, weak checkpoints\.![Histogram of within-run gains for all fits against the pre-training null, and box plots split by configuration](https://arxiv.org/html/2608.20186v1/figures/figure_S2_gain_distribution.png)Figure S2:Distribution of the primary metric across all fits, against its pre\-training null\.\(a\)Within\-run retrieval gain at the selected checkpoint for all 768 fits of Sweeps 1 and 3, stacked by sweep \(colour\), overlaid on the pre\-training \(epoch 0\) values of the same fits \(grey\)\. The grey distribution is the empirical null for this metric: the model has been initialised but not trained, so any departure from zero is the metric’s own noise given five evaluation pool draws and ten permutations per pool\. The two distributions do not overlap: the pre\-training values lie within±0\.2\\pm 0\.2pp of zero, while the selected\-checkpoint values run from 1\.4 to 16\.5 pp, with medians of 6\.3 pp \(Sweep 1\) and 7\.5 pp \(Sweep 3\)\. The best single fit of each sweep is marked \(15\.4 pp in Sweep 1, 16\.5 pp in Sweep 3\); both are maxima over fits as well as over epochs and should be read as upper bounds\.\(b\)The same values split by target type×\\timesarchitecture as box plots \(median, quartiles, whiskers at1\.5×1\.5\\timesIQR\) with individual fits jittered behind\. No fit in any of the four configurations falls at or below chance, so the decoding effect is present everywhere in the explored hyperparameter space and the sweeps determine its magnitude rather than its existence\. Chance is zero for every value shown\.

## Appendix S8Word presentations surviving the amplitude criterion

The 100μ\\muV amplitude criterion \(Section[2\.4](https://arxiv.org/html/2608.20186#S2.SS4)\) is applied to the analysis window plus the temporal\-shift margin \(for the data augmentation hyperparameter\), so a longer window, a different lock point \(onset, offset, centre\) or a larger augmentation shift can result in a different number of trials being excluded\. The inclusion count is therefore a property of the analysis cell\.

Table S19:Word presentations surviving the 100μ\\muV amplitude criterion, by analysis\-window parameter \(mean across the fits in each cell; range in parentheses where it varies\)\.FactorLevelFitsPresentedIncludedExcludedInclusionrateSweepWindow lockonset64240,141220,787 \(218,530–223,044\)19,35491\.0–92\.9%Sweep 1centre64240,141221,334 \(219,018–223,651\)18,80691\.2–93\.1%Sweep 1offset64240,141221,832 \(219,670–223,995\)18,30891\.5–93\.3%Sweep 1Temporal shift096240,141223,563 \(223,044–223,995\)16,57892\.9–93\.3%Sweep 1±100\\pm 100ms96240,141219,073 \(218,530–219,670\)21,06891\.0–91\.5%Sweep 1Window length0\.5 s64240,141222,157 \(220,498–223,790\)17,98491\.8–93\.2%Sweep 30\.7 s64240,141221,056 \(219,651–222,443\)19,08591\.5–92\.6%Sweep 30\.9 s64240,141216,122 \(213,808–218,429\)24,01989\.0–91\.0%Sweep 3Temporal shift096240,141221,134 \(217,487–223,790\)19,00790\.6–93\.2%Sweep 3±50\\pm 50ms96240,141218,422 \(213,808–221,295\)21,71889\.0–92\.2%Sweep 3Channel setall 1996240,141219,350 \(213,808–223,044\)20,79189\.0–92\.9%Sweep 3O1/O2/T5/T6 removed96240,141220,207 \(214,764–223,790\)19,93489\.4–93\.2%Sweep 3

Similar Articles

Brain-CLIPLM: Decoding Compressed Semantic Representations in EEG for Language Reconstruction

arXiv cs.CL

Researchers propose Brain-CLIPLM, a two-stage EEG-to-text decoding framework using contrastive learning for semantic anchor extraction and a retrieval-grounded LLM with Chain-of-Thought reasoning, achieving 67.55% top-5 sentence retrieval accuracy and suggesting EEG-to-text decoding should focus on recovering compressed semantic content rather than full sentence reconstruction.

Silent speech with ultrasound

Hacker News Top

A model trained on ultrasound recordings of the tongue can decode silent speech with a 15.6% word error rate, approaching lip-reading performance despite a smaller dataset.

End-to-End Intracortical Speech Decoding from Neural Activity

arXiv cs.CL

This paper proposes an end-to-end Conformer-based neural decoder for intracortical speech decoding from a participant with ALS, achieving a 23.80% character error rate without any external language model. It demonstrates that meaningful character-level decoding is possible in a fully end-to-end framework.