LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
Summary
The paper introduces LibriBrain100, a large-scale MEG dataset with over 100 hours of data for neural speech decoding, achieving state-of-the-art performance on word classification benchmarks.
View Cached Full Text
Cached at: 08/27/26, 09:36 AM
# LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
Source: [https://arxiv.org/html/2608.25204](https://arxiv.org/html/2608.25204)
Francesco MantegnaAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKEmail:[francesco@robots\.ox\.ac\.uk](mailto:)Dulhan JayalathAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKEmail:[oiwi@robots\.ox\.ac\.uk](mailto:)Tasha KimAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKBenjamin BallykAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKAlex FungAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKAffiliation:FMRIB, Oxford Centre for Integrative Neuroimaging, University of Oxford, UKSungJun ChoAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKAffiliation:OHBA, Oxford Centre for Integrative Neuroimaging, University of Oxford, UKTeyun KwonAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKLuisa KurthAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKMiran ÖzdoganAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKGilad LandauAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKPratik SomaiyaAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UKNatalie VoetsAffiliation:FMRIB, Oxford Centre for Integrative Neuroimaging, University of Oxford, UKMark WoolrichAffiliation:OHBA, Oxford Centre for Integrative Neuroimaging, University of Oxford, UKOiwi Parker JonesAffiliation:PNPL![[Uncaptioned image]](https://arxiv.org/html/2608.25204v1/logos/PNPL_logos_final_Pineapple.png), Department of Engineering Science, University of Oxford, UK
###### Abstract
We introduce LibriBrain100, a large\-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation\. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting inover 100 hours of high\-quality MEGacquired while subjects listened to naturalistic continuous speech\. With∼\\sim80 hours from a single subject, LibriBrain100 sets a new record for deep, within\-subject neural data \(8×\\timesmore than the next comparable dataset and roughly 80×\\timesmore than other datasets\)\. To demonstrate the payoff of this depth\-first design, we evaluate on a word\-classification benchmark—an increasingly well\-established stepping stone towards the open challenge of non\-invasive brain\-to\-text decoding\. Using an existing decoding model, we achieve state\-of\-the\-art performance—validating both the quality of the recordings and the value of within\-subject data at scale\. Because collecting 80 hours of data per user is impractical for real\-world applications, we also collected∼\\sim40 minutes of additional data from each of 32 subjects\. Using the same word\-classification benchmark, we demonstrate the value of broad multi\-subject data: supervised fine\-tuning of a pre\-trained model can substantially compensate for limited per\-subject data\. We providestandard train, validation, and test splits, all reproducible through an open\-sourced Python librarythat supports easy downloading, optional preprocessing, and data loading for common deep learning frameworks\. In addition, the dataset and evaluation infrastructure are being released alongside an open machine\-learning competition with a public leaderboard for standardised benchmarking\. Ultimately, our hope is that LibriBrain100 will accelerate progress towards practical non\-invasive brain\-computer interfaces, capable of restoring communication to people living with severe paralysis\.
## 1Introduction
Deep learning had a transformational effect on computer vision once ImageNet gave it a benchmark: data at scale and standard evaluation protocols\. Neural speech decoding is at a similar inflection point — but with an added complexity\. In invasive settings, systems decoding from surgically implanted electrodes now outperform automatic speech recognition in some conditions, achieving word error rates below 5% in paralysed patients\([Card et al\., 2024](https://arxiv.org/html/2608.25204#bib.bib5)\)\. But invasive recording requires brain surgery, limiting deployment to clinical trials and excluding the vast majority of people who might benefit\. Non\-invasive alternatives are therefore the real prize, and deep learning is beginning to show real gains there too\([Tang et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib13);[Défossez et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib7);[d’Ascoli et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib6);[Jayalath et al\., 2025a](https://arxiv.org/html/2608.25204#bib.bib34)\)\. Yet unlike computer vision, the field has lacked the shared infrastructure to know how much progress is real, what can be built on reliably, and how quickly to anticipate advances\. Without standard benchmarks, shared data, and common evaluation protocols, it is difficult to know what works or by how much\.
The LibriBrain dataset\([Özdogan et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib14)\)— and its open ML competition\([Landau et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib16)\)— represented a significant step in the right direction\. Taking the insight that within\-subject data pushes results fastest\([d’Ascoli et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib6)\), LibriBrain provided the largest within\-subject magnetoencephalography \(MEG\) dataset for speech decoding from brain recordings to date, together with standard splits and a curriculum of benchmark tasks\. Building on the LibriBrain infrastructure, subsequent work has advanced the state of the art across speech detection, phoneme classification, word classification, keyword spotting, and even full brain\-to\-text decoding\([Elvers et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib59);[Elvers et al\., 2026](https://arxiv.org/html/2608.25204#bib.bib55);[Jayalath et al\., 2025a](https://arxiv.org/html/2608.25204#bib.bib34);[Jayalath and Parker Jones, 2026](https://arxiv.org/html/2608.25204#bib.bib36)\)\.
But LibriBrain has its limitations\. It covers only a single subject listening to a single genre \(detective fiction\) and so does not speak directly to cross\-subject generalisation\. This is of a critical importance for real\-world BCIs, where per\-user data collection must be minimal\. The limited stimulus diversity inherent in a tight focus on the series of Sherlock Holmes books, leaves nuanced questions about phonetic and semantic decoding underexplored\. Finally, even with its unprecedented scale of within\-subject data, analyses of LibriBrain have shown no signs of plateauing\([Özdogan et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib14), see, e\.g\.,\), suggesting that more data should yield even bigger gains\.
LibriBrain100 addresses these limitations directly\. It extends LibriBrain both indepthand inbreadth\. Along the depth axis, LibriBrain100 extends the within\-subject data from∼\\sim50 to∼\\sim80 hours, completing the entire canon of Sherlock Holmes and adding new stimuli designed with both phonetics and semantics in mind: TIMIT\([Garofolo et al\., 1993a](https://arxiv.org/html/2608.25204#bib.bib18)\)and MOCHA\-TIMIT\([Wrench, 1999](https://arxiv.org/html/2608.25204#bib.bib20);[Wrench, 2000](https://arxiv.org/html/2608.25204#bib.bib21)\)are standard corpora in both ASR research\([Mohamed et al\., 2009](https://arxiv.org/html/2608.25204#bib.bib53);[Hinton et al\., 2012](https://arxiv.org/html/2608.25204#bib.bib54);[Graves et al\., 2013](https://arxiv.org/html/2608.25204#bib.bib51), e\.g\.,\)and invasive speech BCIs\([Mesgarani et al\., 2014](https://arxiv.org/html/2608.25204#bib.bib42);[Cheung et al\., 2016](https://arxiv.org/html/2608.25204#bib.bib15);[Anumanchipalli et al\., 2019](https://arxiv.org/html/2608.25204#bib.bib2);[Makin et al\., 2020](https://arxiv.org/html/2608.25204#bib.bib3);[Metzger et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib11), e\.g\.,\), which control for phoneme and phoneme\-pair distributions and enable direct comparison with prior studies using the same stimuli; and podcast narratives spanning a wide range of semantic domains, which were previously used for fMRI\-based semantic decoding\([Tang et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib13)\)enabling cross\-modal comparison\. Along the breadth axis, LibriBrain100 extends data collection to 32 new subjects, each contributing∼\\sim40 minutes of MEG, designed to support supervised fine\-tuning and data\-efficient cross\-subject generalisation\. Together, these corpora span the acoustic–semantic axis, positioning LibriBrain100 to address the questions of how MEG\-based speech decoding may be driven by representations of sound versus meaning in the brain\.
On the evaluation side, LibriBrain100 advances the curriculum of tasks from speech detection and phoneme classification to word classification — a task capable of utilising both phonetic and semantic representations in the brain, and that sits closer to the brain\-to\-text goal that motivates much of the field \(Figure[1](https://arxiv.org/html/2608.25204#S1.F1)\)\. We evaluate word classification in two settings: within\-subject, targeting state\-of\-the\-art performance by scaling up training data for subject 0; and cross\-subject, focusing on data\-efficient generalisation via supervised fine\-tuning from a single high\-data subject to new subjects with limited recordings\. Standard splits are provided across all sub\-datasets, and baseline results using MEG\-XL\([Jayalath and Parker Jones, 2026](https://arxiv.org/html/2608.25204#bib.bib36)\)are released to support reproducible community benchmarking\.
LibriBrain100 makes the following contributions:
- •A large\-scale, stimulus\-diverse within\-subject MEG dataset:∼80\{\\sim\}80hours of recordings from subject 0 — the largest within\-subject speech decoding dataset to date — completing the full Sherlock Holmes canon and adding TIMIT and MOCHA\-TIMIT for phonetically controlled analyses and podcast narratives for semantic decoding — enabling cross\-modal comparison with the fMRI literature and direct comparison with influential invasive BCI studies
- •A multi\-subject MEG dataset for cross\-subject generalisation: recordings from 32 new subjects \(∼40\{\\sim\}40minutes each\), designed to support supervised fine\-tuning and data\-efficient generalisation benchmarking
- •A word classification benchmark with standard splits and baselines: an extended evaluation curriculum with train/validation/test splits across all sub\-datasets, together with reproducible baseline results for within\-subject and cross\-subject settings using MEG\-XL
- •Open data release and tooling: raw and preprocessed data on HuggingFace, rich linguistic annotations spanning sub\-lexical to lexical levels, and a Python library for streamlined deep learning integration — all resources to support a planned open ML competition
Figure 1:Decoding words from LibriBrain100\. LibriBrain100 supports word\-classification benchmarks across both deep within\-subject data and broad cross\-subject data\. \(a\)Within\-subject decoding: example speech stimuli and MEG responses from subject 0 across multiple corpora\. Target words are highlighted in the text, aligned to the acoustic spectrogram, and marked in the corresponding MEG time series\. MEG time series were averaged across temporal sensors illustrated in the sensor layout\. \(b\)Cross\-subject decoding: the same target\-word classification task can be evaluated across the 32 additional subjects with limited per\-subject data, enabling benchmarking for data\-efficient generalisation\.
## 2Related Work
In recent decoding work on MEG acquired during naturalistic speech listening\([Défossez et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib7);[d’Ascoli et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib6);[Jayalath et al\., 2025a](https://arxiv.org/html/2608.25204#bib.bib34);[Jayalath et al\., 2025b](https://arxiv.org/html/2608.25204#bib.bib37);[Jayalath and Parker Jones, 2026](https://arxiv.org/html/2608.25204#bib.bib36), e\.g\.,\), a recurring set of large\-scale MEG datasets have emerged\([King et al\., 2026](https://arxiv.org/html/2608.25204#bib.bib52), also see\)\. The usual suspects include MOUS\([Schoffelen et al\., 2019](https://arxiv.org/html/2608.25204#bib.bib23)\), MEG\-MASC\([Gwilliams et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib8)\),[Armeni et al\. \(2022\)](https://arxiv.org/html/2608.25204#bib.bib22), Le Petit Prince\([d’Ascoli et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib6)\), and LibriBrain\([Özdogan et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib14)\)\. Table[2](https://arxiv.org/html/2608.25204#A1.T2)\(Appendix[A\.2](https://arxiv.org/html/2608.25204#A1.SS2)\) summarises these datasets along with LibriBrain100, listing details such as the language of stimuli and dimensions in which we can measure their size\. In terms of ‘deep’, within\-subject data, LibriBrain100 is∼\\sim8×\\timesbigger than the next non\-LibriBrain dataset\([Armeni et al\., 2022](https://arxiv.org/html/2608.25204#bib.bib22)\), and∼\\sim80×\\timesbigger than the rest\.
We can compare the existing datasets in terms ofbreadth—the number of subjects—anddepth—the hours of data per subject\. On the one hand, MOUS, MEG\-MASC, and Le Petit Prince represent breadth\-first designs, spanning 27–96 subjects but with under 2 hours per subject\. Armeni and LibriBrain contrast depth\-first designs, with 10 and 52\.3 hours per subject respectively, but only 1–3 subjects\. This distinction matters for decoding\. Despite other datasets having greater total hours,[d’Ascoli et al\. \(2025\)](https://arxiv.org/html/2608.25204#bib.bib6)find word\-classification performance to be strongest by a substantial margin for Armeni—the deepest dataset in their study, LibriBrain having not come out yet\. When LibriBrain is included\([Jayalath et al\., 2025a](https://arxiv.org/html/2608.25204#bib.bib34), as in\), decoding performance is even stronger with it than with Armeni, further underscoring the importance of within\-subject depth\.
The obvious trade\-off for LibriBrain, which is the closest dataset to LibriBrain100, is that its depth comes at the cost of any breadth—the original LibriBrain release contained just one subject\. LibriBrain100 addresses this limitation while continuing to prioritise depth\. It deepens the single\-subject recording to∼\\sim80 hours while adding∼\\sim40 minutes of data from each of 32 additional subjects—resulting in a hrs/subject range of 0\.6–80\.0\. This reflects both a depth\-first design and respectable provisions for studying generalisation to new subjects with more realistic amounts of data—of the six datasets in Table[2](https://arxiv.org/html/2608.25204#A1.T2), LibriBrain100 has the 3rd highest number of subjects\. The nature of the audio stimuli in these datasets is also interesting\. Unlike LibriBrain, which focused only on readings of Sherlock Holmes \(seven out of nine books\), LibriBrain100’s stimulus set—which includes TIMIT\([Garofolo et al\., 1993a](https://arxiv.org/html/2608.25204#bib.bib18)\), MOCHA\-TIMIT\([Wrench, 1999](https://arxiv.org/html/2608.25204#bib.bib20);[Wrench, 2000](https://arxiv.org/html/2608.25204#bib.bib21)\), and 30 podcasts\-fromThe Moth\([https://themoth\.org](https://themoth.org/)\)–is designed to better control for sound and meaning\-based brain activity\.
## 3The LibriBrain100 Dataset
### 3\.1Overview
The LibriBrain100 dataset contains over 100 hours of non\-invasive magnetoencephalography \(MEG\) recordings, which were acquired while subjects listened to connected speech \(Table[1](https://arxiv.org/html/2608.25204#S3.T1)\)\. In addition, the dataset includes paired event files with time\-locked annotations for linguistic events \(e\.g\. speech/non\-speech, phonemes, words\)\. These annotations were designed for use as labels for decoding tasks such as word classification \(Section[4](https://arxiv.org/html/2608.25204#S4)\)\. The MEG data are being released in both raw and preprocessed/serialised formats, where both may be thought of as matrices with 306 sensor dimensions×\\timesTTtime samples\. MEG recordings were originally sampled at 1 kHz but then downsampled to 250 Hz during preprocessing to preserve oscillations into the high\-gamma range \(70–125 Hz\)\. Consequently, each samplet∈\[1,T\]t\\in\[1,T\]represents 1 ms in the raw data and 4 ms in the preprocessed data\.
Table 1:LibriBrain100 data sources and standard splits\. The dataset comprises adeepsingle\-subject component \(∼\\sim80 hours\) and abroadmulti\-subject component \(∼\\sim40 minutes per subject across 32 subjects\)\. All durations are shown in hours and rounded to one decimal place; totals are computed before rounding\. For a visual summary of the dataset, see Figure[5](https://arxiv.org/html/2608.25204#A1.F5)\(Appendix[A\.1](https://arxiv.org/html/2608.25204#A1.SS1)\)\.ComponentSource\# SubjAll DataTrainValTestDeepSubject 0Sherlock \(all books\)168\.167\.30\.40\.4TIMIT15\.34\.90\.20\.2MOCHA\-TIMIT11\.10\.80\.20\.2The Moth \(30 podcasts\)16\.05\.70\.20\.2Subtotal \(hours\)180\.578\.61\.00\.9BroadSubjects 1–32Sherlock \(2 chapters\)3223\.7–11\.512\.2\(∼\\sim44 minutes/subject\)Total \(hours\)33104\.278\.612\.513\.1
### 3\.2Data Collection and Structure
All MEG data were collected at the Oxford Centre for Human Brain Activity \(OHBA\) on a MEGIN TRIUX™ Neo system with 306 sensors known as superconducting quantum interference devices \(SQUIDs\)\([Hämäläinen et al\., 1993](https://arxiv.org/html/2608.25204#bib.bib50)\)\. The SQUIDs in this system partition into 102 magnetometers and 204 gradiometers\. In a magnetically shielded room, these sensors are sufficiently sensitive to detect electromagnetic fields generated by cortical activity throughout the whole brain\([Bénar et al\., 2021](https://arxiv.org/html/2608.25204#bib.bib49)\)\. The present release includes data from 33 subjects, all of whom were proficient English speakers and provided informed consent for pseudonymised data to be shared; this study was approved by the University of Oxford Medical Sciences Interdivisional Research Ethics Committee \(R90053/RE003\)\.
In terms of structure, the majority of the data \(∼\\sim80 hours\) comes from a single subject, referred to as ‘subject zero’ \(sub\-0\)\. This was a deliberate choice, as hours of data per subject is known to drive performance more than total number of hours across subjects\([d’Ascoli et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib6)\)\. At the same time, for downstream applications where it would be preferable not to require such large data from every new subject, we also include more practical amounts of data \(∼\\sim20 minutes×\\times2 sessions per subject\) from a cohort of 32 additional subjects \(sub\-1tosub\-32\)\. In total, the dataset includes more than 100 hours of MEG data with annotations generated from the audio stimuli that the subjects listened to\.
The audio stimuli include LibriVox recordings for every book in the canon of Sherlock Holmes\([Doyle, 1888](https://arxiv.org/html/2608.25204#bib.bib25);[Doyle, 1890](https://arxiv.org/html/2608.25204#bib.bib26);[Doyle, 1892](https://arxiv.org/html/2608.25204#bib.bib27);[Doyle, 1893](https://arxiv.org/html/2608.25204#bib.bib28);[Doyle, 1902](https://arxiv.org/html/2608.25204#bib.bib29);[Doyle, 1905](https://arxiv.org/html/2608.25204#bib.bib30);[Doyle, 1915](https://arxiv.org/html/2608.25204#bib.bib31);[Doyle, 1917](https://arxiv.org/html/2608.25204#bib.bib32);[Doyle, 1927](https://arxiv.org/html/2608.25204#bib.bib33)\)\. While the first seven of these books \(∼\\sim50 hours\) featured in the original LibriBrain dataset\([Özdogan et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib14)\), completing the canon here was not only completionist but aimed to maximise data quantity while minimising variability\. For example, all but one book were read by the same narrator and both narrators selected to have similar Southern British accents \(see Appendix[B\.2](https://arxiv.org/html/2608.25204#A2.SS2)for more details\)\. Our aim in limiting the variability of these stimuli was pragmatic\. By analogy, if you wanted to train automatic speech recognition \(ASR\) from scratch on about 68 hours of audio recordings, it would be easier to concentrate on just one or two speakers than attempt to model every accent or dialect\.
To complement the Sherlock Holmes audiobooks, we looked to the literature for inspiration\. For better phonetic coverage, we used TIMIT\([Garofolo et al\., 1993a](https://arxiv.org/html/2608.25204#bib.bib18)\)and MOCHA\-TIMIT\([Wrench, 1999](https://arxiv.org/html/2608.25204#bib.bib20);[Wrench, 2000](https://arxiv.org/html/2608.25204#bib.bib21)\)\. Not only has TIMIT long been a cornerstone for the development of ASR technology\([Graves et al\., 2013](https://arxiv.org/html/2608.25204#bib.bib51), e\.g\.,\), but it has played a foundational role in our understanding of speech comprehension in the brain\([Mesgarani et al\., 2014](https://arxiv.org/html/2608.25204#bib.bib42);[Cheung et al\., 2016](https://arxiv.org/html/2608.25204#bib.bib15), e\.g\.,\)\. One feature of TIMIT is that the sentences in it were designed to balance the distribution of phonemes and phoneme\-bigrams\. Unlike Sherlock, TIMIT also contains audio recordings from 630 speakers across 8 dialect groups in the USA; MOCHA\-TIMIT adds male and female speakers from the UK \(see Appendix[B\.2](https://arxiv.org/html/2608.25204#A2.SS2)\)\.
For semantic coverage, we looked to a landmark semantic decoding study\([Tang et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib13)\)which used multiple short \(∼\\sim12 minute\) podcast stories to cover a broad range of topics\. We include 30 podcast stories fromThe Moth\(for details, see Appendix[B\.2](https://arxiv.org/html/2608.25204#A2.SS2)\)\.
For reproducible evaluation, the data are split into standard training, validation, and testing sets\. Following both[Özdogan et al\. \(2025\)](https://arxiv.org/html/2608.25204#bib.bib14)and[Landau et al\. \(2025\)](https://arxiv.org/html/2608.25204#bib.bib16), sessions 11 and 12—which correspond to chapters 11 and 12—from the first Sherlock book are held out for validation and testing, respectively\. Sessions 13 and 14, which were previously held\-out for competition evaluation\([Landau et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib16)\), have now been made public and can be used for training, together with the rest of the Sherlock data\. As there is a standard split for TIMIT, we use that in the MEG data for comparability\. Concretely, we use the 24\-speaker core test set, which consists of 2 male and 1 female speaker from each dialect region, totalling 192 utterances \(24 speakers×\\times8 utterances which exclude the sharedSA1andSA2sentences\)\. For validation, we follow the standard Kaldi TIMIT recipe and use its 50\-speaker development set, again excluding the sharedSA1andSA2sentences\. Although surgical decoding studies have used MOCHA\-TIMIT\([Anumanchipalli et al\., 2019](https://arxiv.org/html/2608.25204#bib.bib2);[Makin et al\., 2020](https://arxiv.org/html/2608.25204#bib.bib3), e\.g\.,\), we are unaware of a precedent for a standard split; therefore, we created our own split\. To do this, we note that MOCHA\-TIMIT is organised into four sets, A–D, where sentences repeat between sets A and D and between sets B and C\. We use A and D for training\. For validation and testing, we split the sentences in sets B and C into equally sized subsets of unique sentences\. This results in unique sentences for train, validation, and test sets, minimising the risk of information leakage due to sentence type\. Table[1](https://arxiv.org/html/2608.25204#S3.T1)summarises the data and splits\.
### 3\.3Data Formats, Access, and Supporting Python Library
We are releasing LibriBrain100 in two formats\. First, we are releasing the raw data \(with no preprocessing\) in BIDS format\([Niso et al\., 2018](https://arxiv.org/html/2608.25204#bib.bib24)\)\. BIDS structures are directories that follow guidelines on file naming and organisation, with MEG data in FIF format\. Technically, LibriBrain100 is a collection of BIDS datasets, one for each Sherlock book, TIMIT, MOCHA\-TIMIT, and The Moth podcasts\. Second, we are releasing the data in a ready serialised format\. These data have been minimally preprocessed \(i\.e\., corrected for head movement, filtered, downsampled\) and then converted to float32 HDF5 format, for easy machine learning\.
The MEG data in both FIF and HDF5 files can be thought of as a 306 channel×\\timesTTtime\-sample matrix, with time sampled at 1,000 Hz for the raw data or 250 Hz for the serialised data\. Each FIF or HDF5 file is paired with an annotation file in TSV format\. The annotation files contain time\-stamps for linguistic events like word onsets, which can be used to define labels for supervised learning\.
The recommended way to access the data is through our custom Python library,pnpl\(`pip install pnpl`\)\. The library provides task\-driventorch\.utils\.data\.Datasetclasses whose constructors return ready\-to\-iterate windows of MEG, time\-locked to events such as word or phoneme onsets:
frompnpl\.datasetsimportLibriBrain100
frompnpl\.tasksimportWordClassification
ds=LibriBrain100\(
data\_path="\./data/LibriBrain100",
task=WordClassification\(tmin=0\.2,tmax=0\.6\),
partition="train",
\)
x,y=ds\[0\]
If the requested files are not atdata\_path, the library downloads them on demand from Hugging Face \(theLibriBrain100loader transparently fetches each record from whichever underlying repository owns it\)\. Given the size of the full dataset \(∼\\sim104 hours of MEG, hundreds of GB even after serialisation\), the constructor exposes selectors so users can scope the download to the part they care about:
ds=LibriBrain100\(data\_path=…,task=…,partition="train",
subjects="deep"\)
ds=LibriBrain100\(data\_path=…,task=…,
subjects="deep",corpus="timit"\)
ds=LibriBrain100\(data\_path=…,task=…,
subjects="deep",corpus=\["timit","mocha"\]\)
Thesubjects=argument accepts"all"\(default\),"deep"\(subject 0\),"broad"\(subjects 1–32\), an integer, a single subject id \("0"or"sub\-0"\), or any iterable of those\. Thecorpus=argument accepts"all"\(default\),"sherlock","timit","mocha","podcasts", or a list of those \(the canonical token for “MOCHA\-TIMIT” is"mocha"but aliases like"mocha\-timit"or"the\_moth"are accepted\)\. As the broad component \(subjects 1–32\) was only collected with the Sherlock stimuli, the loader rejects combinations likesubjects="broad", corpus="timit"up front rather than silently returning an empty or incomplete dataset\. Per\-task wrappers \(LibriBrain100Speech,LibriBrain100Phoneme,LibriBrain100Word\) are also available for convenience\. For finer control, users can pass a list of explicitinclude\_run\_keysorexclude\_run\_keysas\(subject, session, task, run\)tuples, mirroring the run\-key convention from the original LibriBrain release\.
While we recommend accessing the data through the Python library, with its high\-level API which does not require an understanding of the underlying repository structure, the data are also available for direct download from two Hugging Face repositories: the original LibriBrain repository \([https://huggingface\.co/datasets/pnpl/LibriBrain](https://huggingface.co/datasets/pnpl/LibriBrain)\) and a new LibriBrain2 repository \([https://huggingface\.co/datasets/pnpl/LibriBrain2](https://huggingface.co/datasets/pnpl/LibriBrain2)\)\. For philosophers and fans of category mistakes\([Ryle, 1949](https://arxiv.org/html/2608.25204#bib.bib56)\), we might point out that LibriBrain100 is not a third repository alongside LibriBrain and LibriBrain2, but rather their union — as the University of Oxford is not a building alongside its colleges, departments, and other facilities, but rather the collection of them all\.
## 4Decoding Experiments
To evaluate LibriBrain100, we focus here on the task of word classification which is defined as follows\. Given a multi\-channel MEG recording𝐗∈ℝC×T\\mathbf\{X\}\\in\\mathbb\{R\}^\{C\\times T\}withCCchannels andTTtime samples, the word classification task maps to a target wordy∈𝒱y\\in\\mathcal\{V\}where𝒱\\mathcal\{V\}is a fixed vocabulary\. Word classification has previously been studied in MEG using the most frequent 50 and 250 words\([d’Ascoli et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib6);[Jayalath et al\., 2025a](https://arxiv.org/html/2608.25204#bib.bib34);[Jayalath and Parker Jones, 2026](https://arxiv.org/html/2608.25204#bib.bib36)\)defined over various corpora and using top\-10 balanced accuracy, computed by averaging per\-class top\-1010accuracy across the vocabulary𝒱\\mathcal\{V\}to account for imbalances in word frequency\([Zipf, 1949](https://arxiv.org/html/2608.25204#bib.bib48)\)\. In the following experiments, we use a vocabulary of 50 words \(see Appendix[F\.2](https://arxiv.org/html/2608.25204#A6.SS2)\)\.
As a backbone, we use MEG\-XL\([Jayalath and Parker Jones, 2026](https://arxiv.org/html/2608.25204#bib.bib36)\)\. This pre\-trained model has achieved state\-of\-the\-art word classification results in cross\-subject MEG decoding using unsupervised pre\-training followed by supervised fine\-tuning\. MEG\-XL draws on insights from LLMs to significantly improve downstream performance by extending the temporal context \(5–300×\\times\) of previous foundation models of the brain\([Jiang et al\., 2024](https://arxiv.org/html/2608.25204#bib.bib43);[Avramidis et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib44);[Wang et al\., 2024](https://arxiv.org/html/2608.25204#bib.bib46);[Xiao et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib45);[Wang et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib47), cf\. \)\. MEG\-XL has been pre\-trained on approximately 300 hours from 800 subjects over multiple speech and non\-speech datasets—i\.e\., MOUS\([Schoffelen et al\., 2019](https://arxiv.org/html/2608.25204#bib.bib23)\), Cam\-CAN\([Shafto et al\., 2014](https://arxiv.org/html/2608.25204#bib.bib9)\), and SMN4Lang\([Wang et al\., 2022](https://arxiv.org/html/2608.25204#bib.bib10)\)—before being fine\-tuned on LibriBrain\([Özdogan et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib14)\)\. In this paper, we fine\-tune the base MEG\-XL checkpoint \([https://huggingface\.co/pnpl/MEG\-XL](https://huggingface.co/pnpl/MEG-XL)\) on LibriBrain100 for the following tasks\.
### 4\.1Task 1: From Many\-to\-One — Improving Within\-Subject Word Classification
Figure[2](https://arxiv.org/html/2608.25204#S4.F2)describes the results on the held\-out test data of Subject 0 when fine\-tuning MEG\-XL in two settings\. In the first, we fine\-tune MEG\-XL on all of the available training data in LibriBrain100; in the second, we exclude the training data of subjects 1–32\. Training with the broad subject data appears to be helpful for generalisation to different speech distributions in the form of TIMIT, Mocha\-TIMIT, and the podcast data\. However, when evaluating on Sherlock, which forms the majority of the data, training without subjects 1\-32 led to the best test performance, though only marginally\. In general, even when we have a large amount of data for a single subject \(Subject 0\), additional data from 32 other subjects appears to improve downstream performance on the single subject \(Subject 0\)\.
Figure 2:Subject 0 performance on held\-out test data\. Performance improves when training jointly with Subject 1\-32 training data, except on the Sherlock subset\. Here, training only on Sherlock training data leads to the best performance\. Random chance accuracy is0\.20\.2\. \* indicatesp<\.05p<\.05and \*\* indicatesp<\.01p<\.01under Mann\-Whitney U\-tests\.
### 4\.2Task 2: From One\-to\-Many — Improving Between\-Subject Word Classification
In Figure[3](https://arxiv.org/html/2608.25204#S4.F3), we assess generalisation to the held\-out test data on subjects 1–32\. Similar to the previous section, we fine\-tune MEG\-XL in two settings\. Here, we analyse training on the multi\-subject training data jointly with Subject 0’s training data \(“Training w/ Subject 0”\), and also the effect of excluding the Subject 0 training data \(“Training w/o Subject 0”\)\. All subjects generalise similarly, and including Subject 0’s training data improves generalisation across subjects by about 15 percentage points\. This shows that a large amount of data from a single subject \(∼\\sim80 hours\) can be used to improve performance on many individuals for whom we have∼\\sim120×\\timesless data \(∼\\sim40 minutes each\)\. In the next section, we investigate pushing this discrepancy further, reducing the amount of fine\-tuning data needed for subjects 1–32 down to∼\\sim10 minutes apiece\.
Figure 3:Fine\-tuning MEG\-XL on subjects 1\-32, with and without training alongside Subject 0\. We report performance on the Sherlock test set\. Training jointly with Subject 0 significantly improves generalisation on the broad and shallow data of subjects 1\-32\. Random chance accuracy is0\.20\.2\. \*\*\* indicatesp<\.001p<\.001under a Mann\-Whitney U\-test\. The per\-subject results in the right panel illustrate that difference is not driven by any specific subject, or proper subset of subjects, but is rather consistent across all 32 subjects\.
### 4\.3Task 3: Exploring Efficiency — Maintaining Improvements in Between\-Subject Word Classification with Decreasing amounts of Supervised Fine\-Tuning Data
Here we investigate the impact of reducing the available fine\-tuning data for subjects 1–32, building on the results from Section[4\.2](https://arxiv.org/html/2608.25204#S4.SS2)\. This is an important test for future applications, as even 40 minutes of brain recordings may be too much for patients to use a speech BCI\. A speech brain–computer interface should generalise to clinical patients with minimal subject\-specific training\.
In Figure[4](https://arxiv.org/html/2608.25204#S4.F4), we compare the use of 100% of the available training data for 12 subjects, with 50% for the next 10 subjects, and 25% for the remaining 10 subjects\. In all cases, we see minimal degradation on generalisation, suggesting that cross\-subject transfer is responsible for most of the present decoding performance\. Our strategy of pre\-training on deep single\-subject data and fine\-tuning to new individuals appears to be sound, and may even be practical for future clinical applications — though we emphasise that all LibriBrain100 data was acquired from healthy volunteers\.
Figure 4:Performance as a function of fine\-tuning data size\. Generalisation remains robust, even with only 25% training data, equivalent to about 10 minutes of brain recordings \(i\.e\., a small enough amount of data to be interesting for clinical applications\)\. The largest difference comes from including data from Subject 0\. Random chance accuracy is0\.20\.2\. None of the differences in performance under different training data percentages were significant under Mann\-Whitney U\-tests\.
## 5Discussion
LibriBrain100 substantially extends the LibriBrain ecosystem in both scale and scope\. With∼80\{\\sim\}80hours from a single subject and∼40\{\\sim\}40minutes from each of 32 additional subjects, it provides the largest within\-subject MEG dataset for speech decoding to date while simultaneously opening the door to cross\-subject generalisation research\.
##### Limitations addressed\.
The original LibriBrain paper closed with a list of limitations\. LibriBrain100 directly addresses several of them\. First, thesingle\-subject design: the addition of 32 new subjects makes cross\-subject generalisation a first\-class benchmark task for the first time in this dataset series\. Second,focused language content: whereas LibriBrain drew exclusively from books in the Sherlock Holmes series, LibriBrain100 adds stimuli deliberately spanning the phonetic–semantic axis — TIMIT and MOCHA\-TIMIT for controlled phonetic coverage, and podcast narratives for broader semantic diversity\([Tang et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib13)\)\. We expect these sub\-datasets will be useful in their own right, enabling targeted studies of phonetic and semantic decoding independently of the collective LibriBrain100 datasets\. Third,preprocessing: the serialised HDF5 release lowers the barrier to entry for ML researchers by removing the need for domain\-specific signal processing knowledge, while the raw BIDS release preserves full flexibility for neuroscientists who wish to apply their own pipelines\.
##### Limitations remaining\.
Two limitations from the original LibriBrain list remain open\. The first is thelistening paradigm focus\. All data in LibriBrain100 were collected during passive listening to continuous speech\. Listening provides a natural starting point for speech decoding, as alignments between stimulus and neural response are straightforward to derive, and signal\-to\-noise is comparatively high\. However, it is not the paradigm that ultimately matters for BCIs, where users must attempt to produce or imagine speech\. We have been actively collecting inner speech data in parallel and plan to release multiple inner speech datasets going forward\. The second remaining limitation isbrain\-to\-text\. We made a deliberate decision not to include a brain\-to\-text baseline in this release\. Progress on the open\-ended decoding problem has been rapid but uneven, and we judged that the field is not yet at a point where a single baseline result would be stable enough to anchor community benchmarking — however, we will revisit this in future releases as methods mature\.
##### Conclusion\.
LibriBrain100 follows the logic of ImageNet: the value of a shared benchmark compounds over time as more methods are trained and evaluated against it\. The open ML competition, public leaderboard, and standardised Python library are all designed with cumulative progress in mind\. We hope that LibriBrain100, like its predecessor, will serve not just as a dataset but as shared infrastructure — lowering the cost of entry for ML researchers, enabling rigorous comparison across methods, and ultimately accelerating progress toward practical non\-invasive brain\-computer interfaces capable of restoring communication to people living with severe paralysis\.
## Acknowledgments and Disclosure of Funding
We are grateful to our anonymous reviewers for their time and attention, which helped to make this paper better, and to the NeurIPS Evaluations & Datasets organisers for all of their efforts in making the track possible\. We acknowledge the use of the University of Oxford Advanced Research Computing \(ARC\) facility \([http://dx\.doi\.org/10\.5281/zenodo\.22558](http://dx.doi.org/10.5281/zenodo.22558)\), Hartree Centre resources, and the NVIDIA Corporation for donating additional GPUs\. PNPL is supported by the MRC \(MR/X00757X/1\), Royal Society \(RG\\\\backslashR1\\\\backslash241267\), NSF \(2314493\), NFRF \(NFRFT\-2022\-00241\), SSHRC \(895\-2023\-1022\), and ARIA \(SCNI\-SE01\-P004\)\.
## References
- Anumanchipalliet al\.\(2019\)G\. K\. Anumanchipalli, J\. Chartier, and E\. F\. ChangSpeech synthesis from neural decoding of spoken sentences\.Nature568\(7753\),pp\. 493–498\.External Links:[Document](https://dx.doi.org/10.1038/s41586-019-1119-1)Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p4.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p6.1)\.
- Armeniet al\.\(2022\)K\. Armeni, U\. Güçlü, M\. van Gerven, and J\.\-M\. SchoffelenA 10\-hour within\-participant magnetoencephalography narrative dataset to test models of language comprehension\.Scientific Data9,pp\. 278\.Cited by:[Table 2](https://arxiv.org/html/2608.25204#A1.T2.3.1.4.2),[§2](https://arxiv.org/html/2608.25204#S2.p1.1)\.
- Avramidiset al\.\(2025\)K\. Avramidis, T\. Feng, W\. Jeong, J\. Lee, W\. Cui, R\. M\. Leahy, and S\. NarayananNeural codecs as biosignal tokenizers\.arXiv preprint arXiv:2510\.09095\.Cited by:[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Bénaret al\.\(2021\)C\. Bénar, J\. Velmurugan, V\. J\. López\-Madrona, F\. Pizzo, and J\. BadierDetection and localization of deep sources in magnetoencephalography: a review\.Current Opinion in Biomedical Engineering18,pp\. 100285\.External Links:[Document](https://dx.doi.org/10.1016/j.cobme.2021.100285)Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p1.1)\.
- Boersma \(2001\)P\. BoersmaPraat, a system for doing phonetics by computer\.Glot International5\(9/10\),pp\. 341–345\.Cited by:[§B\.2](https://arxiv.org/html/2608.25204#A2.SS2.p1.1),[§B\.6](https://arxiv.org/html/2608.25204#A2.SS6.p4.1)\.
- Cardet al\.\(2024\)N\. S\. Card, M\. Wairagkar, C\. Iacobacci, X\. Hou, T\. Singer\-Clark, F\. R\. Willett, E\. M\. Kunz, C\. Fan, M\. Vahdati Nia, D\. R\. Deo, A\. Srinivasan, E\. Y\. Choi, M\. F\. Glasser, L\. R\. Hochberg, J\. M\. Henderson, K\. Shahlaie, S\. D\. Stavisky, and D\. M\. BrandmanAn accurate and rapidly calibrating speech neuroprosthesis\.New England Journal of Medicine391\(7\),pp\. 609–618\.Cited by:[§E\.2](https://arxiv.org/html/2608.25204#A5.SS2.p1.1),[§1](https://arxiv.org/html/2608.25204#S1.p1.1)\.
- Cheunget al\.\(2016\)C\. Cheung, L\. S\. Hamilton, K\. Johnson, and E\. F\. ChangThe auditory representation of speech sounds in human motor cortex\.eLife5,pp\. e12577\.External Links:[Document](https://dx.doi.org/10.7554/eLife.12577),[Link](https://doi.org/10.7554/eLife.12577)Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p4.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p4.1)\.
- Défossezet al\.\(2023\)A\. Défossez, C\. Caucheteux, J\. Rapin, O\. Kabeli, and J\. KingDecoding speech perception from non\-invasive brain recordings\.Nature Machine Intelligence5\(10\),pp\. 1097–1107\.External Links:[Document](https://dx.doi.org/10.1038/s42256-023-00714-5),[Link](https://www.nature.com/articles/s42256-023-00714-5)Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p1.1),[§2](https://arxiv.org/html/2608.25204#S2.p1.1)\.
- Desplanqueset al\.\(2020\)B\. Desplanques, J\. Thienpondt, and K\. DemuynckECAPA\-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification\.Interspeech,pp\. 3830–3834\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2020-2650)Cited by:[§C\.1](https://arxiv.org/html/2608.25204#A3.SS1.p2.1)\.
- Doyle \(1888\)A\. C\. DoyleA study in scarlet\.Ward, Lock & Co\.Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1)\.
- Doyle \(1890\)A\. C\. DoyleThe sign of the four\.Spencer Blackett\.Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1)\.
- Doyle \(1892\)A\. C\. DoyleThe adventures of sherlock holmes\.George Newnes\.Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1)\.
- Doyle \(1893\)A\. C\. DoyleThe memoirs of sherlock holmes\.George Newnes\.Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1)\.
- Doyle \(1902\)A\. C\. DoyleThe hound of the baskervilles\.George Newnes\.Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1)\.
- Doyle \(1905\)A\. C\. DoyleThe return of sherlock holmes\.George Newnes\.Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1)\.
- Doyle \(1915\)A\. C\. DoyleThe valley of fear\.George H\. Doran Company\.Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1)\.
- Doyle \(1917\)A\. C\. DoyleHis last bow\.John Murray\.Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1)\.
- Doyle \(1927\)A\. C\. DoyleThe case\-book of sherlock holmes\.John Murray\.Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1)\.
- d’Ascoliet al\.\(2025\)S\. d’Ascoli, C\. Bel, J\. Rapin, H\. Banville, Y\. Benchetrit, C\. Pallier, and J\. KingTowards decoding individual words from non\-invasive brain recordings\.Nature Communications16,pp\. 10521\.External Links:[Document](https://dx.doi.org/10.1038/s41467-025-65499-0),[Link](https://www.nature.com/articles/s41467-025-65499-0)Cited by:[Table 2](https://arxiv.org/html/2608.25204#A1.T2.3.1.5.2),[Figure 14](https://arxiv.org/html/2608.25204#A5.F14),[Figure 14](https://arxiv.org/html/2608.25204#A5.F14.5.1),[§1](https://arxiv.org/html/2608.25204#S1.p1.1),[§1](https://arxiv.org/html/2608.25204#S1.p2.1),[§2](https://arxiv.org/html/2608.25204#S2.p1.1),[§2](https://arxiv.org/html/2608.25204#S2.p2.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p2.1),[§4](https://arxiv.org/html/2608.25204#S4.p1.1)\.
- Elverset al\.\(2026\)G\. Elvers, G\. Landau, F\. Mantegna, M\. Özdogan, T\. Kim, T\. Kwon, S\. Cho, B\. Ballyk, L\. Kurth, D\. Jayalath, P\. Somaiya, X\. de Zuazo, B\. Shillingford, G\. Farquhar, M\. Jiang, K\. Jerbi, H\. Abdelhedi, Y\. Mantilla Ramos, C\. Gulcehre, M\. Woolrich, N\. Voets, and O\. Parker JonesBenchmarking non\-invasive speech BCIs: lessons learned from the 2025 PNPL competition\.Journal of Machine Learning Research \(JMLR\)\.Note:In pressCited by:[§1](https://arxiv.org/html/2608.25204#S1.p2.1)\.
- Elverset al\.\(2025\)G\. Elvers, G\. Landau, and O\. Parker JonesElementary, my dear Watson: non\-invasive neural keyword spotting in the LibriBrain dataset\.Advances in Neural Information Processing Systems \(NeurIPS\), Workshop on Data on the Brain & Mind\.Note:arXiv preprint arXiv:2510\.21038External Links:[Link](https://openreview.net/forum?id=gRJ9dd07QF)Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p2.1)\.
- Garofoloet al\.\(1993a\)J\. S\. Garofolo, L\. F\. Lamel, W\. M\. Fisher, J\. G\. Fiscus, D\. S\. Pallett, and N\. L\. DahlgrenTIMIT acoustic\-phonetic continuous speech corpus\.Note:Linguistic Data Consortium[https://catalog\.ldc\.upenn\.edu/LDC93S1](https://catalog.ldc.upenn.edu/LDC93S1)Cited by:[§B\.2](https://arxiv.org/html/2608.25204#A2.SS2.SSSx2.p1.1),[§B\.7](https://arxiv.org/html/2608.25204#A2.SS7.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2608.25204#S1.p4.1),[§2](https://arxiv.org/html/2608.25204#S2.p3.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p4.1)\.
- Garofoloet al\.\(1993b\)J\. S\. Garofolo, L\. F\. Lamel, W\. M\. Fisher, J\. G\. Fiscus, D\. S\. Pallett, and N\. L\. DahlgrenDARPA TIMIT acoustic\-phonetic continuous speech corpus CD\-ROM\.Technical reportTechnical ReportNISTIR 4930,National Institute of Standards and Technology,Gaithersburg, MD\.Cited by:[§B\.2](https://arxiv.org/html/2608.25204#A2.SS2.SSSx2.p1.1)\.
- Graveset al\.\(2013\)A\. Graves, A\. Mohamed, and G\. HintonSpeech recognition with deep recurrent neural networks\.InProceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 6645–6649\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP.2013.6638947)Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p4.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p4.1)\.
- Gwilliamset al\.\(2023\)L\. Gwilliams, G\. Flick, A\. Marantz, L\. Pylkkanen, D\. Poeppel, and J\. KingIntroducing MEG\-MASC: a high\-quality magneto\-encephalography dataset for evaluating natural speech processing\.Scientific Data10\(1\),pp\. 862\.External Links:[Document](https://dx.doi.org/10.1038/s41597-023-02752-5),[Link](https://www.nature.com/articles/s41597-023-02752-5)Cited by:[Table 2](https://arxiv.org/html/2608.25204#A1.T2.3.1.3.2),[§2](https://arxiv.org/html/2608.25204#S2.p1.1)\.
- Hämäläinenet al\.\(1993\)M\. Hämäläinen, R\. Hari, R\. J\. Ilmoniemi, J\. Knuutila, and O\. V\. LounasmaaMagnetoencephalography—theory, instrumentation, and applications to noninvasive studies of the working human brain\.Reviews of Modern Physics65\(2\),pp\. 413–497\.External Links:[Document](https://dx.doi.org/10.1103/RevModPhys.65.413)Cited by:[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p1.1)\.
- Hintonet al\.\(2012\)G\. Hinton, L\. Deng, D\. Yu, G\. E\. Dahl, A\. Mohamed, N\. Jaitly, A\. Senior, V\. Vanhoucke, P\. Nguyen, T\. N\. Sainath, and B\. KingsburyDeep neural networks for acoustic modeling in speech recognition: the shared views of four research groups\.IEEE Signal Processing Magazine29\(6\),pp\. 82–97\.External Links:[Document](https://dx.doi.org/10.1109/MSP.2012.2205597)Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p4.1)\.
- Jayalathet al\.\(2026\)D\. Jayalath, B\. Ballyk, and O\. Parker JonesA common measure of communication for speech brain–computer interfaces\.Note:Manuscript in preparationCited by:[§E\.2](https://arxiv.org/html/2608.25204#A5.SS2.p1.1)\.
- Jayalathet al\.\(2025a\)D\. Jayalath, G\. Landau, and O\. Parker JonesUnlocking non\-invasive brain\-to\-text\.International Conference on Machine Learning \(ICML\), Workshop on Generative AI and Biology\.Note:arXiv preprint arXiv:2505\.13446Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p1.1),[§1](https://arxiv.org/html/2608.25204#S1.p2.1),[§2](https://arxiv.org/html/2608.25204#S2.p1.1),[§2](https://arxiv.org/html/2608.25204#S2.p2.1),[§4](https://arxiv.org/html/2608.25204#S4.p1.1)\.
- Jayalathet al\.\(2025b\)D\. Jayalath, G\. Landau, B\. Shillingford, M\. W\. Woolrich, and O\. Parker JonesThe Brain’s Bitter Lesson: scaling speech decoding with self\-supervised learning\.International Conference on Machine Learning \(ICML\)\.Note:arXiv preprint arXiv:2406\.04328Cited by:[§2](https://arxiv.org/html/2608.25204#S2.p1.1)\.
- Jayalath and Parker Jones \(2026\)D\. Jayalath and O\. Parker JonesMEG\-XL: data\-efficient brain\-to\-text via long\-context pre\-training\.International Conference on Machine Learning \(ICML\)\.Note:arXiv preprint arXiv:2602\.02494Cited by:[Figure 14](https://arxiv.org/html/2608.25204#A5.F14),[Figure 14](https://arxiv.org/html/2608.25204#A5.F14.5.1),[§E\.1](https://arxiv.org/html/2608.25204#A5.SS1.p1.1),[§1](https://arxiv.org/html/2608.25204#S1.p2.1),[§1](https://arxiv.org/html/2608.25204#S1.p5.1),[§2](https://arxiv.org/html/2608.25204#S2.p1.1),[§4](https://arxiv.org/html/2608.25204#S4.p1.1),[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Jianget al\.\(2024\)W\. Jiang, L\. Zhao, and B\. LuLarge brain model for learning generic representations with tremendous EEG data in BCI\.International Conference on Learning Representations \(ICLR\)\.External Links:[Link](https://openreview.net/forum?id=QzTpTRVtrP)Cited by:[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Kinget al\.\(2026\)J\. King, C\. Bel, L\. Evanson, J\. Gadonneix, S\. Houhamdi, J\. Lévy, J\. Raugel, A\. Santos Revilla, M\. Zhang, J\. Bonnaire, C\. Caucheteux, A\. Défossez, T\. Desbordes, P\. Diego\-Simón, S\. Khanna, J\. Millet, P\. Orhan, S\. Panchavati, A\. Ratouchniak, A\. Thual, T\. L\. Brooks, K\. Begany, Y\. Benchetrit, M\. Careil, H\. Banville, S\. d’Ascoli, S\. Dahan, and J\. RapinNeuralSet: a high\-performing Python package for Neuro\-AI\.arXiv preprint arXiv:2605\.03169\.External Links:[Link](https://arxiv.org/abs/2605.03169)Cited by:[§2](https://arxiv.org/html/2608.25204#S2.p1.1)\.
- Landauet al\.\(2025\)G\. Landau, M\. Özdogan, G\. Elvers, F\. Mantegna, P\. Somaiya, D\. Jayalath, L\. Kurth, T\. Kwon, B\. Shillingford, G\. Farquhar, M\. Jiang, K\. Jerbi, H\. Abdelhedi, Y\. Mantilla Ramos, C\. Gulcehre, M\. Woolrich, N\. Voets, and O\. Parker JonesThe 2025 PNPL competition: speech detection and phoneme classification in the LibriBrain dataset\.Advances in Neural Information Processing Systems \(NeurIPS\), Competition Track\.Note:arXiv preprint arXiv:2506\.10165Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p2.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p6.1)\.
- Makinet al\.\(2020\)J\. G\. Makin, D\. A\. Moses, and E\. F\. ChangMachine translation of cortical activity to text with an encoder\-decoder framework\.Nature Neuroscience23\(4\),pp\. 575–582\.External Links:[Document](https://dx.doi.org/10.1038/s41593-020-0608-8)Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p4.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p6.1)\.
- Mesgaraniet al\.\(2014\)N\. Mesgarani, C\. Cheung, K\. Johnson, and E\. F\. ChangPhonetic feature encoding in human superior temporal gyrus\.Science343\(6174\),pp\. 1006–1010\.External Links:[Document](https://dx.doi.org/10.1126/science.1245994)Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p4.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p4.1)\.
- Metzgeret al\.\(2023\)S\. L\. Metzger, K\. T\. Littlejohn, A\. B\. Silva, D\. A\. Moses, M\. P\. Seaton, R\. Wang, M\. E\. Dougherty, J\. R\. Liu, P\. Wu, M\. A\. Berger, I\. Zhuravleva, A\. Tu\-Chan, K\. Ganguly, G\. K\. Anumanchipalli, and E\. F\. ChangA high\-performance neuroprosthesis for speech decoding and avatar control\.Nature620,pp\. 1037–1046\.Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p4.1)\.
- Mikolovet al\.\(2013\)T\. Mikolov, I\. Sutskever, K\. Chen, G\. S\. Corrado, and J\. DeanDistributed representations of words and phrases and their compositionality\.Advances in Neural Information Processing Systems \(NeurIPS\)26,pp\. 3111–3119\.Cited by:[§C\.1](https://arxiv.org/html/2608.25204#A3.SS1.p4.1)\.
- Mohamedet al\.\(2009\)A\. Mohamed, G\. E\. Dahl, and G\. HintonDeep belief networks for phone recognition\.InNIPS Workshop on Deep Learning for Speech Recognition and Related Applications,Cited by:[§1](https://arxiv.org/html/2608.25204#S1.p4.1)\.
- Moseset al\.\(2021\)D\. A\. Moses, S\. L\. Metzger, J\. R\. Liu, G\. K\. Anumanchipalli, J\. G\. Makin, P\. F\. Sun, J\. Chartier, M\. E\. Dougherty, P\. M\. Liu, G\. M\. Abrams, A\. Tu\-Chan, K\. Ganguly, and E\. F\. ChangNeuroprosthesis for decoding speech in a paralyzed person with anarthria\.New England Journal of Medicine385\(3\),pp\. 217–227\.External Links:[Document](https://dx.doi.org/10.1056/NEJMoa2027540),[Link](https://doi.org/10.1056/NEJMoa2027540)Cited by:[Figure 15](https://arxiv.org/html/2608.25204#A5.F15),[Figure 15](https://arxiv.org/html/2608.25204#A5.F15.5.1),[§E\.2](https://arxiv.org/html/2608.25204#A5.SS2.p1.1)\.
- Nisoet al\.\(2018\)G\. Niso, K\. J\. Gorgolewski, E\. Bock, T\. L\. Brooks, G\. Flandin, A\. Gramfort, R\. N\. Henson, M\. Jas, V\. Litvak, J\. T\. Moreau, R\. Oostenveld, J\. Schoffelen, F\. Tadel, J\. Wexler, and S\. BailletMEG\-BIDS, the brain imaging data structure extended to magnetoencephalography\.Scientific Data5,pp\. 180110\.External Links:[Document](https://dx.doi.org/10.1038/sdata.2018.110)Cited by:[§3\.3](https://arxiv.org/html/2608.25204#S3.SS3.p1.1)\.
- Ochshorn and Hawkins \(2015\)R\. M\. Ochshorn and M\. HawkinsGentle: a robust yet lenient forced aligner built on Kaldi\.External Links:[Link](https://lowerquality.com/gentle/)Cited by:[§B\.6](https://arxiv.org/html/2608.25204#A2.SS6.p2.1)\.
- Özdoganet al\.\(2025\)M\. Özdogan, G\. Landau, G\. Elvers, D\. Jayalath, P\. Somaiya, F\. Mantegna, M\. Woolrich, and O\. Parker JonesLibriBrain: over 50 hours of within\-subject MEG to improve speech decoding methods at scale\.Advances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track\.Note:arXiv preprint arXiv:2506\.02098Cited by:[Table 2](https://arxiv.org/html/2608.25204#A1.T2.3.1.6.2),[§B\.7](https://arxiv.org/html/2608.25204#A2.SS7.SSS0.Px1.p1.1),[Table 3](https://arxiv.org/html/2608.25204#A2.T3),[Table 3](https://arxiv.org/html/2608.25204#A2.T3.6.1),[§1](https://arxiv.org/html/2608.25204#S1.p2.1),[§1](https://arxiv.org/html/2608.25204#S1.p3.1),[§2](https://arxiv.org/html/2608.25204#S2.p1.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p3.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p6.1),[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Peelle and Davis \(2012\)J\. E\. Peelle and M\. H\. DavisNeural oscillations carry speech rhythm through to comprehension\.Frontiers in Psychology3,pp\. 320\.Cited by:[§D\.1](https://arxiv.org/html/2608.25204#A4.SS1.p2.1)\.
- Peirce \(2007\)J\. W\. PeircePsychoPy—psychophysics software in Python\.Journal of Neuroscience Methods162\(1\-2\),pp\. 8–13\.Cited by:[§B\.3](https://arxiv.org/html/2608.25204#A2.SS3.p1.1)\.
- Poveyet al\.\(2011\)D\. Povey, A\. Ghoshal, G\. Boulianne, L\. Burget, O\. Glembek, N\. Goel, M\. Hannemann, P\. Motlicek, Y\. Qian, P\. Schwarz, J\. Silovsky, G\. Stemmer, and K\. VeselyThe Kaldi speech recognition toolkit\.InIEEE 2011 Workshop on Automatic Speech Recognition and Understanding,Hilton Waikoloa Village, Big Island, Hawaii, US\.External Links:[Link](http://kaldi-asr.org/)Cited by:[§B\.6](https://arxiv.org/html/2608.25204#A2.SS6.p1.1)\.
- Ryle \(1949\)G\. RyleThe concept of mind\.Hutchinson,London\.Cited by:[§3\.3](https://arxiv.org/html/2608.25204#S3.SS3.p6.1)\.
- Schoffelenet al\.\(2019\)J\. Schoffelen, R\. Oostenveld, N\. H\. L\. Lam, J\. Uddén, A\. Hultén, and P\. HagoortA 204\-subject multimodal neuroimaging dataset to study language processing\.Scientific Data6\(1\),pp\. 17\.External Links:[Document](https://dx.doi.org/10.1038/s41597-019-0020-y)Cited by:[Table 2](https://arxiv.org/html/2608.25204#A1.T2.3.1.2.2),[§2](https://arxiv.org/html/2608.25204#S2.p1.1),[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Shaftoet al\.\(2014\)M\. A\. Shafto, L\. K\. Tyler, M\. Dixon, J\. R\. Taylor, J\. B\. Rowe, R\. Cusack, A\. J\. Calder, W\. D\. Marslen\-Wilson, J\. Duncan, T\. Dalgleish, R\. N\. Henson, C\. Brayne, F\. E\. Matthews, and Cam\-CANThe Cambridge Centre for Ageing and Neuroscience \(Cam\-CAN\) study protocol: a cross\-sectional, lifespan, multidisciplinary examination of healthy cognitive ageing\.BMC Neurology14,pp\. 204\.External Links:[Document](https://dx.doi.org/10.1186/s12883-014-0204-1)Cited by:[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Tanget al\.\(2023\)J\. Tang, A\. LeBel, S\. Jain, and A\. G\. HuthSemantic reconstruction of continuous language from non\-invasive brain recordings\.Nature Neuroscience26\(5\),pp\. 858–866\.External Links:[Document](https://dx.doi.org/10.1038/s41593-023-01304-9),[Link](https://www.nature.com/articles/s41593-023-01304-9)Cited by:[§B\.2](https://arxiv.org/html/2608.25204#A2.SS2.SSSx4.p1.1),[§B\.2](https://arxiv.org/html/2608.25204#A2.SS2.SSSx4.p3.1),[§B\.7](https://arxiv.org/html/2608.25204#A2.SS7.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2608.25204#S1.p1.1),[§1](https://arxiv.org/html/2608.25204#S1.p4.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p5.1),[§5](https://arxiv.org/html/2608.25204#S5.SS0.SSS0.Px1.p1.1)\.
- Taulu and Simola \(2006\)S\. Taulu and J\. SimolaSpatiotemporal signal space separation method for rejecting nearby interference in MEG measurements\.Physics in Medicine & Biology51\(7\),pp\. 1759–1768\.External Links:[Document](https://dx.doi.org/10.1088/0031-9155/51/7/008)Cited by:[§B\.5](https://arxiv.org/html/2608.25204#A2.SS5.p1.1)\.
- Wanget al\.\(2024\)G\. Wang, W\. Liu, Y\. He, C\. Xu, L\. Ma, and H\. LiEEGPT: pretrained transformer for universal and reliable representation of EEG signals\.InThe Thirty\-Eighth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=lvS2b8CjG5)Cited by:[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Wanget al\.\(2025\)J\. Wang, S\. Zhao, Z\. Luo, Y\. Zhou, H\. Jiang, S\. Li, T\. Li, and G\. PanCBraMod: a criss\-cross brain foundation model for EEG decoding\.International Conference on Learning Representations \(ICLR\)\.External Links:[Link](https://openreview.net/forum?id=NPNUHgHF2w)Cited by:[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Wanget al\.\(2022\)S\. Wang, X\. Zhang, J\. Zhang, and C\. ZongA synchronized multimodal neuroimaging dataset for studying brain language processing\.Scientific Data9,pp\. 590\.External Links:[Document](https://dx.doi.org/10.1038/s41597-022-01708-5)Cited by:[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Willettet al\.\(2023\)F\. R\. Willett, E\. M\. Kunz, C\. Fan, D\. T\. Avansino, G\. H\. Wilson, E\. Y\. Choi, F\. Kamdar, M\. F\. Glasser, L\. R\. Hochberg, S\. Druckmann, K\. V\. Shenoy, and J\. M\. HendersonA high\-performance speech neuroprosthesis\.Nature620,pp\. 1031–1036\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06377-x)Cited by:[§E\.2](https://arxiv.org/html/2608.25204#A5.SS2.p1.1)\.
- Wrench \(2000\)A\. A\. WrenchA multi\-channel/multi\-speaker articulatory database for continuous speech recognition research\.Phonus5,pp\. 1–13\.Cited by:[§B\.2](https://arxiv.org/html/2608.25204#A2.SS2.SSSx3.p1.1),[§1](https://arxiv.org/html/2608.25204#S1.p4.1),[§2](https://arxiv.org/html/2608.25204#S2.p3.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p4.1)\.
- Wrench \(1999\)A\. WrenchThe MOCHA\-TIMIT articulatory database\.Note:[https://www\.cstr\.ed\.ac\.uk/research/projects/artic/mocha\.html](https://www.cstr.ed.ac.uk/research/projects/artic/mocha.html)Created November 1999; accessed 2026\-02\-07Cited by:[§B\.2](https://arxiv.org/html/2608.25204#A2.SS2.SSSx3.p1.1),[§1](https://arxiv.org/html/2608.25204#S1.p4.1),[§2](https://arxiv.org/html/2608.25204#S2.p3.1),[§3\.2](https://arxiv.org/html/2608.25204#S3.SS2.p4.1)\.
- Xiaoet al\.\(2025\)Q\. Xiao, Z\. Cui, C\. Zhang, S\. Chen, W\. Wu, A\. Thwaites, A\. Woolgar, B\. Zhou, and C\. ZhangBrainOmni: a brain foundation model for unified EEG and MEG signals\.InThe Thirty\-Ninth Annual Conference on Neural Information Processing Systems,External Links:[Link](https://openreview.net/forum?id=cjHQj0tCy6)Cited by:[§4](https://arxiv.org/html/2608.25204#S4.p2.1)\.
- Zipf \(1949\)G\. K\. ZipfHuman behavior and the principle of least effort: an introduction to human ecology\.Addison\-Wesley Press,Cambridge, MA\.Cited by:[§4](https://arxiv.org/html/2608.25204#S4.p1.1)\.
## Appendices / Supplemental Materials
## Appendix AAdditional Dataset Details
### A\.1Visual Summary
Figure[5](https://arxiv.org/html/2608.25204#A1.F5)provides two complementary views of the LibriBrain100 dataset composition\. Panel \(a\) shows recording hours by subject and corpus; panel \(b\) shows a treemap of the full dataset with rectangle area proportional to duration\. See Table[1](https://arxiv.org/html/2608.25204#S3.T1)for a complementary perspective\.
Figure 5:Dataset proportions\.\(a\) Recording hours by subject and corpus\. The stacked bar plot shows total duration per subject, with sub\-0 comprising multiple linguistic materials and other subjects a subset of audiobook data\. The bubble plot represents the same information \(area proportional to hours\); the larger sub\-0 circle is subdivided into audiobooks, sentences, and podcasts, while smaller circles represent other subjects\. Colours indicate data type, blue shades distinguish between audiobooks, and circle outlines distinguish between subjects\. Black concentric circles provide a size reference \(~1 h, ~10 h, ~80 h\)\. \(b\) Treemap of the full dataset, with rectangle area proportional to duration\. Categories are subdivided into corpora and then into MEG recording sessions\. The multicolour outline marks the only two audiobook sessions recorded from multiple subjects\.
### A\.2Comparison with Existing Datasets
Table[2](https://arxiv.org/html/2608.25204#A1.T2)lists existing MEG datasets for speech comprehension at the time of writing\. LibriBrain100 compares favourably in total hours \(104\) and hours per subject \(0\.6–80\), with the highest scores \(by far, in the case of maximum hours per subject\)\. It sits comfortably in the middle of the pack in terms of number of subjects, with 33 \(range 1–96\)\. Thepnpllibrary contains data loaders for all of these datasets with support for reproducible data splits and for dataset downloading where possible, to support machine learning at scale\.
NameDatasetLanguageSensorsTotal hrs\# SubjHrs/SubjStimulusPublicPNPL DataloaderMOUS[Schoffelen et al\. \(2019\)](https://arxiv.org/html/2608.25204#bib.bib23)Dutch27581960\.8Spoken sentences[Yes](https://data.ru.nl/collections/di/dccn/DSC_3011020.09_236)YesMEG\-MASC[Gwilliams et al\. \(2023\)](https://arxiv.org/html/2608.25204#bib.bib8)English20849271\.0MASC stories \(synthetic voice\)[Yes](https://doi.org/10.17605/OSF.IO/AG3KJ)YesArmeni[Armeni et al\. \(2022\)](https://arxiv.org/html/2608.25204#bib.bib22)English26930310\.0LibriVox \(Sherlock Book 3\)[Yes](https://data.ru.nl/collections/di/dccn/DSC_3011085.05_995)YesLe Petit Prince[d’Ascoli et al\. \(2025\)](https://arxiv.org/html/2608.25204#bib.bib6)French30693581\.6Le Petit Prince[Yes](https://openneuro.org/datasets/ds007523/versions/1.0.1)YesLibriBrain[Özdogan et al\. \(2025\)](https://arxiv.org/html/2608.25204#bib.bib14)English30652152\.3LibriVox \(Sherlock Book 1–7\)[Yes](https://huggingface.co/datasets/pnpl/LibriBrain)YesLibriBrain100\(ours\)English306104330\.6–80\.0LibriVox \(Sherlock Book 1–9\),TIMIT, MOCHA\-TIMIT,The Moth \(30 Podcasts\)[Yes](https://huggingface.co/datasets/pnpl/LibriBrain2)Yes
Table 2:Breadth and depth of speech comprehension MEG datasets\.
## Appendix BData Collection Methods
### B\.1Subjects
MEG recordings were acquired from 33 volunteers \(10 female, 23 male; age range 19–50 years, median = 28, interquartile range = 8\)\. All participants reported normal hearing and normal or corrected\-to\-normal vision, and none had a history of neurological disorders\. Sixteen participants were native speakers of English, while the remaining seventeen acquired English as a second language but were highly proficient and reported daily use for work or study\. Prior to participation, all individuals provided informed consent for the use of anonymised data for research purposes\. The study was approved by the University of Oxford Medical Sciences Interdivisional Research Ethics Committee \(R90053/RE003\)\.
### B\.2Stimulus Materials
To prepare the audio stimuli for the MEG experiments, all audio files were converted to uncompressed WAV format and resampled to 48 kHz using SoX\. The goal was to segment the audio into natural sentences or phrases, typically defined by pauses, and to align each segment with its corresponding textual content\. This enabled the delivery of precise triggers to the MEG system during stimulus presentation, allowing continuous verification of temporal alignment between MEG recordings, audio signals, and linguistic annotations\. Audio segmentation was carried out in two stages: an initial automated phase followed by manual refinement\. The automated segmentation relied on voice activity detection \(VAD\), while the manual stage corrected segment boundaries in cases where VAD failed, for example by truncating low\-intensity sounds such as utterance\-initial voiceless plosives\. In addition, manual refinement involved checking and reconciling consistency between data\-driven audio segmentation and text\-based boundaries \(e\.g\., punctuation and sentence structure\)\. VAD was implemented using custom scripts in Praat\([Boersma, 2001](https://arxiv.org/html/2608.25204#bib.bib1)\), identifying speech segments based on an intensity threshold of 59 dB and a minimum duration of 600 ms\. The resulting TextGrid annotations were then populated with the corresponding text and carefully reviewed and corrected by hand\. This process required over 200 hours of expert human effort, exceeding standard practices in the design and curation of comparable datasets, and was undertaken to ensure precise alignment across MEG recordings, audio signals, and linguistic annotations\.
#### Audiobooks
Audio recordings for all nine books in the canon of Sherlock Holmes were sourced from LibriVox \([https://librivox\.org/](https://librivox.org/)\)\. The books \(including both novels and short stories\) were presented in chronological order of publication\. To minimise variance, we utilised recordings from the same reader \(David Clarke\) for books 1–8\. As audio recordings were not available from this speaker for book 9, we selected a second reader with similar characteristics \(Thomas A\. Copeland; British English male\)\. Each MEG recording session corresponded to a standalone chapter from the audiobooks\. All texts are in the public domain and were obtained from Project Gutenberg \([https://www\.gutenberg\.org/](https://www.gutenberg.org/)\)\. Manual correction was used to fix typos and, where text and audio diverged, to align the text with the spoken audio\. Ambiguous text, such as numbers, was normalised \(e\.g\., “1492” rendered as “fourteen ninety\-two” if spoken that way\)\. Details for the audiobooks are provided in Table[3](https://arxiv.org/html/2608.25204#A2.T3)\.
The Sherlock Holmes audiobooks comprise multiple detective fiction stories focused on solving mysteries through observation and logical deduction\. Narratives typically involve identifying and interpreting clues, with cases ranging from concrete, crime\-based investigations to stories that initially present more unusual or seemingly supernatural elements before being resolved through reasoning\. The prose includes a mixture of narration and both direct and indirect dialogue\. The thematics are relatively homogeneous \(e\.g\., crime, investigation\), but variation arises across books and chapters through differences in setting, narrative context, and secondary characters\. The recordings are read by professional narrators, resulting in a controlled delivery: pauses tend to align with punctuation and sentence structure, diction is clear and consistent, and breathing and other non\-linguistic artefacts are minimised\.
BookNameTypeSessionsHours1[A Study in Scarlet](https://librivox.org/a-study-in-scarlet-version-6-by-sir-arthur-conan-doyle/)\(1888\)Novel1404:37:342[The Sign of the Four](https://librivox.org/the-sign-of-the-four-version-3-by-sir-arthur-conan-doyle/)\(1890\)Novel1204:27:313[The Adventures of Sherlock Holmes](https://librivox.org/the-adventures-of-sherlock-holmes-version-4-by-sir-arthur-conan-doyle/)\(1892\)Short Stories1210:56:134[The Memoirs of Sherlock Holmes](https://librivox.org/the-memoirs-of-sherlock-holmes-by-sir-arthur-conan-doyle-2/)\(1893\)Short Stories1208:53:175[The Hound of the Baskervilles](https://librivox.org/the-hound-of-the-baskervilles-version-4-by-sir-arthur-conan-doyle/)\(1901–1902\)Novel1506:10:326[Return of Sherlock Holmes](https://librivox.org/the-return-of-sherlock-holmes-by-sir-arthur-conan-doyle-2/)\(1905\)Short Stories1411:51:177[The Valley of Fear](https://librivox.org/the-valley-of-fear-version-3-by-sir-arthur-conan-doyle/)\(1914–1915\)Novel1406:06:178[His Last Bow](https://librivox.org/his-last-bow-version-3-by-sir-arthur-conan-doyle/)\(1917\)Short Stories1007:10:299[The Case\-Book of Sherlock Holmes](https://librivox.org/the-case-book-of-sherlock-holmes-by-sir-arthur-conan-doyle/)\* \(1927\)Short Stories1308:27:0168:40:11\*Read by Thomas A\. CopelandTable 3:Audiobooks available in LibriBrain100\. Books 1–7 appeared in the original LibriBrain release\([Özdogan et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib14)\); the rest are new here\. Text in the name column link to the source audio on LibriVox\.
#### TIMIT
TIMIT is a corpus of read American English speech designed for acoustic–phonetic research and the development and evaluation of automatic speech recognition systems\([Garofolo et al\., 1993b](https://arxiv.org/html/2608.25204#bib.bib19);[Garofolo et al\., 1993a](https://arxiv.org/html/2608.25204#bib.bib18)\)\. It contains recordings from 630 speakers across 8 major dialect regions of the United States, each producing 10 sentences, for a total of 6300 utterances \(∼\\sim5 hours of speech\)\. All speakers were native speakers of American English and were screened to exclude clinically significant speech pathology\. Speaker recruitment aimed to cover major dialect regions based on the region in which speakers lived during childhood\. While the corpus achieves broad dialectal coverage, it is not perfectly balanced: some regions \(e\.g\., New England, New York City, and “Army Brat”\) are underrepresented due to practical constraints\. The corpus is also sex\-imbalanced \(438 male, 192 female; approximately 70%/30%\)\.
The sentence prompts were designed to provide controlled phonetic coverage\. Each speaker reads 2 shared SA \(“dialect”\) sentences, 5 SX sentences selected from a set of 450 phonetically compact sentences \(each repeated across 7 speakers\), and 3 SI sentences drawn from a set of 1890 phonetically diverse sentences \(each produced by a single speaker\), yielding 2342 distinct sentence texts in total\. The SX and SI sentences were constructed to maximise coverage of phonetic contexts and allophonic variation, rather than semantic diversity\. Speech was recorded at 16 kHz using 16\-bit linear PCM encoding\. In this dataset release, we provide MEG recordings for the full set of 6300 utterances\.
The TIMIT corpus consists of short, isolated sentences with no shared discourse context across utterances\. The recordings are read speech, produced in a controlled setting, with clear articulation, limited disfluencies, and pauses aligned with sentence boundaries\. At the same time, variability is introduced through the large number of speakers and dialect regions, providing diversity in pronunciation, accent, and voice characteristics\. This combination of tightly controlled phonetic design and speaker variability makes TIMIT particularly well\-suited for analyses focused on acoustic–phonetic representations\.
#### MOCHA\-TIMIT
MOCHA\-TIMIT is a multichannel articulatory–acoustic corpus designed to support research on speech production and acoustic–articulatory modelling\([Wrench, 1999](https://arxiv.org/html/2608.25204#bib.bib20);[Wrench, 2000](https://arxiv.org/html/2608.25204#bib.bib21)\)\. The publicly distributed release contains data from two speakers: one male \(msak0\) and one female \(fsew0\), both Southern British English speakers\. In this dataset release, we provide all 460 utterances from each speaker \(920 utterances total\)\. The sentence set was constructed to provide broad phonetic coverage while explicitly targeting connected\-speech processes in English, such as assimilation and weak forms, and was designed to parallel the phonetic coverage of TIMIT while reflecting British English pronunciation\. Although the MOCHA\-TIMIT corpus was originally developed for studying speech production, the tight control over articulation and phonetic content also makes it well\-suited for perception studies, as it provides highly structured and phonetically balanced stimuli with well\-characterised acoustic realisations\.
Each utterance is a short, read sentence produced in a controlled recording environment\. As in TIMIT, there is no extended narrative context across sentences; instead, the corpus emphasises phonetic and articulatory diversity within isolated utterances\. The recordings are carefully produced, with clear diction, minimal disfluencies, and pauses aligned with sentence boundaries\. Compared to TIMIT, however, the speech more systematically reflects connected\-speech phenomena, resulting in more natural coarticulation patterns despite the controlled setting\. For example, assimilation processes may occur across word boundaries \(e\.g\., “good boy” realised with a more bilabial /d/ influenced by the following /b/\), and weak forms are frequently used for function words \(e\.g\., “and”→\\rightarrow/@n/, “to”→\\rightarrow/t@/, “of”→\\rightarrow/@v/\)\. Vowel reduction and consonant lenition are also present in unstressed positions, and segmental realisations are influenced by surrounding phonetic context, yielding more continuous and context\-dependent articulatory patterns\.
#### Podcasts
We used 30 stories \(6 hours\) from The Moth podcast, drawn from a larger set of 77 stories \(∼\\sim15 hours\) previously used for semantic decoding from fMRI data\([Tang et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib13)\)\. This subset was selected using a data\-driven criterion aimed at maximising semantic coverage\. Specifically, we estimated the spread of each story in semantic space using sentence embedding models, which map sentences to dense vector representations that capture their semantic content\. For each story, we computed the centroid of its sentence embeddings and quantified dispersion as the mean cosine distance of individual sentence vectors to this centroid in the high\-dimensional embedding space\. Stories with greater dispersion \(i\.e\., covering a wider range of semantic content, as opposed to being concentrated in a narrow region of the embedding space\) were preferentially selected\. Details for the selected stories are provided in Table[4](https://arxiv.org/html/2608.25204#A2.T4)\.
The Moth podcast stories provide an effective set of naturalistic stimuli for neural decoding, in particular for semantic representations, because they combine ecological validity with high semantic diversity and coherent narrative structure\. Each stimulus consists of a single speaker telling an autobiographical narrative, ensuring continuity in voice, perspective, and discourse structure\. Crucially, the corpus spans a wide range of topics and contexts, from historically and geographically specific settings \(e\.g\., Zimbabwe during independence, wartime Baghdad, travel in the USSR, scientific expeditions\) to personal and emotional experiences \(e\.g\., bereavement, parenthood, moral dilemmas\), as well as humorous, reflective, and self\-development narratives\. This breadth yields rich variation in entities, events, and abstract themes, providing dense coverage of semantic space\. At the same time, individual stories are internally coherent and contextually well\-defined, allowing models to track how meaning unfolds over extended timescales\.
These narratives are delivered as spontaneous speech rather than read text, and therefore include natural disfluencies such as interjections \(e\.g\., “like,” “you know,” “I mean”\), repetitions \(e\.g\., “and and then…,” “I was I was…”\), and repairs \(e\.g\., “on Tuesday—uh, Wednesday,” “her brother—no, her cousin”\), which are largely absent in scripted speech\. Temporal dynamics also differ from read speech: pauses can reflect on\-the\-fly planning, while other segments are produced in dense, continuous stretches, a characteristic feature of spontaneous narration\. In addition, prosodic variation is more pronounced, as speakers may raise their voice or modulate intonation for emphasis or comedic effect\. Recordings also include audience responses such as laughter and applause, which introduce acoustic variability but can carry contextual and pragmatic information tied to the narrative\. Together, these properties enhance ecological validity and introduce variability that better reflects real\-world language processing\. The combination of rich thematic diversity, extended context, and naturalistic delivery, as leveraged in prior work\([Tang et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib13), e\.g\.,\), makes Moth stories particularly well\-suited for training and evaluating neural decoders that map brain activity to continuous, context\-sensitive semantic representations\.
Table 4:Podcasts available in LibriBrain100\. Story titles are hyperlinked to the source audio and transcript at[https://themoth\.org/stories/](https://themoth.org/stories/)\. Validation and test set stories are highlightedyellowandred, respectively\. Durations are given in minutes and seconds for the audio files used in our stimulus set\.\#StoryAuthorDuration1[Alternate Ithaca Tom](https://themoth.org/stories/alternate-ithaca-tom)Tom Weiser11:472[Thumbs Up\!](https://themoth.org/stories/thumbs-up)Nathan Englander14:093[Goldie, the Goldfish](https://themoth.org/stories/goldie-the-goldfish)Becca Stevens10:544[Under the Influence](https://themoth.org/stories/under-the-influence)Jeffery Rudell10:285[Treasure Island](https://themoth.org/stories/treasure-island)Boots Riley13:306[That Thing on My Arm](https://themoth.org/stories/that-thing-on-my-arm)Padma Lakshmi14:497[Breaking Up in the Age of Google](https://themoth.org/stories/breaking-up-in-the-age-of-google)Jessi Klein17:438[The Triangle Shirtwaist Connection](https://themoth.org/stories/the-triangle-shirtwaist-connection)Michelle Fecteau7:069[Stage Fright](https://themoth.org/stories/stage-fright)Suzanne Vega10:0710[Not on the Usual Tour](https://themoth.org/stories/not-on-the-usual-tour)Jon Lovett8:3111[Life Reimagined](https://themoth.org/stories/life-reimagined)Raymond Christian11:1512[Only One Way To Find Out](https://themoth.org/stories/only-one-way-to-find-out)Sarah Gray13:0413[The Postman Always Calls](https://themoth.org/stories/the-postman-always-calls)Elizabeth Browning15:2914[Blue Hope](https://themoth.org/stories/blue-hope)Sylvia Earle14:0015[The Curse](https://themoth.org/stories/the-curse)Dame Wilburn13:5616[Beneath the Mushroom Cloud](https://themoth.org/stories/clifton-truman-daniel)Clifton Truman Daniel11:4617[Going the Liberty Way](https://themoth.org/stories/going-the-liberty-way)Kevin Roose13:2818[Caution: Eating](https://themoth.org/stories/caution-eating)Evan Kleiman9:4019[Birth of a Nation](https://themoth.org/stories/birth-of-a-nation)Petina Gappah9:0920[Life and Death on the Oregon Trail](https://themoth.org/stories/life-and-death-on-the-oregon-trail)Micaela Blei12:2421[Fire Test For Love](https://themoth.org/stories/fire-test-for-love)Lauren Slater10:5822[Leaving Baghdad](https://themoth.org/stories/leaving-baghdad)Abbas Mousa11:1523[Coming of Age on Death Row](https://themoth.org/stories/coming-of-age-on-death-row)Gautam Narula11:5824[How To Draw A Nekkid Man](https://themoth.org/stories/how-to-draw-a-nekkid-man-storyslam-version)Tricia Rose Burt12:0825[The Mayor of the Freaks](https://themoth.org/stories/mayor-of-freaks)David Crabb16:1126[Swimming with Astronauts](https://themoth.org/stories/swimming-with-astronauts)Michael J\. Massimino13:1127[Wild Womxn and Dancing Queens](https://themoth.org/stories/wild-womxn-and-dancing-queens)Lex Jade6:4028[Vixen and the USSR](https://themoth.org/stories/vixen-and-the-ussr)Sue Steinacher13:2429[From Boyhood to Fatherhood](https://music.apple.com/pa/album/the-best-of-the-moth-vol-7/488876794?l=en)Jonathan Ames11:5730[Where There’s Smoke](https://themoth.org/stories/where-theres-smoke)Jenifer Hixson10:026:00:59
### B\.3Experimental Design & Procedure
Each recording session began with visually presented instructions projected onto a translucent whiteboard using a DLP LED projector \(ProPixx, VPixx Technologies Inc\., Saint\-Bruno, Canada\)\. Participants were seated inside the MEG scanner and initiated the experiment via button press\. Auditory stimuli were delivered binaurally through non\-metallic air tube earphones \(Aero Technologies\) at approximately 70 dB SPL, with minor adjustments to bass and treble based on participant preference\. Stimulus delivery was controlled using the PsychoPy toolbox\([Peirce, 2007](https://arxiv.org/html/2608.25204#bib.bib39)\), and the stimulus computer synchronized with the MEG acquisition system via a parallel port to send precise event triggers marking stimulus onset with millisecond temporal accuracy\.
For the multiple subjects cohort \(sub\-01 to sub\-32\), all participants listened to Chapters 11 and 12 of A Study in Scarlet\. Both chapters were presented within a single recording session lasting approximately one hour, with a short break between chapters\. To assess attention and comprehension, participants answered five questions at the end of each chapter\. Each question was presented in a four\-alternative multiple\-choice format with distractors \(e\.g\., “What does Jefferson Hope take from Lucy Ferrier after her death?” with answer options “A: A locket, B: A necklace, C: A wedding ring, D: A bracelet”\)\. Responses were recorded via button press using a MEG\-compatible ResponsePixx Dual Handheld system \(VPixx Technologies Inc\., Saint\-Bruno, Canada\)\.
Subject 0 participated in multiple recording sessions and was exposed to a broader set of linguistic materials, including all books in the Sherlock Holmes canon, as well as additional speech corpora \(e\.g\., TIMIT, MOCHA\-TIMIT, and The Moth podcasts\)\. Comprehension assessment varied by stimulus type\. For audiobook stimuli, a single comprehension question was presented at the end of each chapter in a two alternative options format \(e\.g\., “Where is the body of the murder victim found?” with answer options “A: in the bedroom”, “B: in the garden”\)\. For TIMIT and MOCHA\-TIMIT, questions were presented immediately after selected sentences \(approximately 20 questions per session; 5% of trials for TIMIT, 10% for MOCHA\-TIMIT\)\. These questions were presented in a four alternative multiple options format \(e\.g\., “\_\_\_\_\_\_ is a pop singer\.” with answer options “A: Michael Jackson, B: Tina Turner, C: Elton John, D: Madonna”; “Clear \_\_\_\_\_\_ is appreciated\.” with answer options “A: grammar, B: diction, C: articulation, D: pronunciation”\) and typically targeted a key word \(often a noun\), with distractors that were either acoustically or semantically similar\. For the podcast corpus, five comprehension questions were presented at the end of each episode, probing key details, event sequences, and overall meaning \(e\.g\., “How did thinking about his alternate life affect him?” with answer options “A: It made him confident, B: It made him confused and unhappy, C: It motivated him to study, D: It made him excited”\)\. For subject 0, each recording session lasted approximately three hours and typically included either 3–5 audiobook chapters, 1000–1200 sentences, or 8–10 podcast episodes\. Short breaks were allowed between recording blocks\. Recording sessions were spaced at least one day apart, with no more than two months between sessions, depending on participant and experimenter availability\.
### B\.4Data Acquisition
Prior to data acquisition, each participant’s head shape was digitised using a Polhemus Fastrak 3D digitiser \(Polhemus, Vermont, USA\)\. This procedure included the localisation of fiducial landmarks \(nasion, left and right pre\-auricular points\), along with approximately 300 additional points sampled across the scalp, forehead, and nose\. Five Head Position Indicator \(HPI\) coils were positioned on the mastoid and forehead to enable continuous monitoring of head position during MEG recording via electromagnetic induction\. MEG recordings were acquired using a MEGIN Triux™ Neo system \(York Instruments Ltd\., Heslington, UK\), consisting of 102 magnetometers and 204 orthogonal planar gradiometers\. The system was installed in a magnetically shielded room to reduce environmental interference\. Prior to entering the recording environment, participants were screened for metallic objects and other potential sources of electromagnetic noise\. During acquisition, participants were seated with their head positioned close to the dewar\. Participants were instructed to minimise head, body, and limb movements throughout the recording\. Data were sampled at 1000 Hz with an online band\-pass filter of 0\.01–330 Hz\. Ocular activity was monitored using bipolar electrooculogram \(EOG\) electrodes, with one pair placed at the outer canthi \(horizontal EOG\) and another above and below the left eye \(vertical EOG\)\. Cardiac signals were recorded via bipolar electrocardiogram \(ECG\) electrodes positioned on the clavicle and hip\. Articulatory muscle activity was continuously tracked using electromyography \(EMG\), with electrodes placed below the cheekbone \(jaw movement\), below the lower lip \(lip movement\), and beneath the chin to capture potential laryngeal activity\.
### B\.5Minimal Preprocessing Pipeline
The MEG data were minimally preprocessed to remove measurement artefacts while preserving flexibility for additional preprocessing steps for downstream analyses\. Head position was estimated using continuous recordings from the Head Position Indicator \(HPI\) coils throughout each recording session\. This information was used to correct for head movements, and data from all participants were realigned to a common reference head position\. This procedure was applied consistently across subjects and is expected to reduce variability due to head position differences and within\-session head movement\. Noisy channels were identified, excluded, and subsequently reconstructed via interpolation from neighbouring sensors\. Environmental noise was attenuated using Maxwell filtering\. Specifically, a temporally non\-extended Signal Space Separation \(SSS\) algorithm\([Taulu and Simola, 2006](https://arxiv.org/html/2608.25204#bib.bib40)\)was applied to suppress sources originating outside the head\. In contrast, no explicit steps were taken to remove physiological artefacts \(e\.g\., eye blinks, cardiac activity, or muscle contractions\), allowing users to apply additional artefact correction procedures as appropriate for their specific downstream analyses\. Power line noise at 50 Hz and its 100 Hz harmonic was removed using notch filters\. The data were then band\-pass filtered between 0\.1 and 125 Hz using zero\-phase, two\-pass Butterworth filters to reduce slow drifts and prevent aliasing effects prior to downsampling\. Finally, the data were downsampled to 250 Hz, yielding recordings with 306 channels and a temporal resolution of 4 ms per sample\.
### B\.6Event Files/Labels
Annotations were generated for each session, providing onset times and durations \(in seconds\) for multiple event types, including silence \(i\.e\., non\-speech segments, which may include breathing and occasional laughter or applause\), word\-level units \(e\.g\., A, Study, in, Scarlet\), and phoneme\-level units \(e\.g\., ah, s, t, ah, d, iy, ih, n, s, k, aa, r, l, ah, t\)\. Additional annotation fields include phoneme position within words, encoded using conventions from Kaldi\([Povey et al\., 2011](https://arxiv.org/html/2608.25204#bib.bib38)\): B \(beginning\), I \(inside\), E \(end\), and S \(singleton\)\. Each MEG recording session \(available as a FIF file\) is accompanied by a corresponding event file in TSV format, which can be used to define labels for supervised learning tasks\. These event files were derived from the transcripts and their corresponding audio segments\.
To obtain precise temporal alignment between text and audio, forced alignment was performed using Gentle\([Ochshorn and Hawkins, 2015](https://arxiv.org/html/2608.25204#bib.bib17)\), a Kaldi\-based aligner that employs Gaussian Mixture Model–Hidden Markov Model \(GMM\-HMM\) architectures to combine acoustic and language models and determine the most likely alignment path between audio and transcript\. Because the audio had already been segmented into short utterances using VAD, as described above, the forced alignment step was simplified\. Instead of processing entire chapters, the aligner was applied to short, pre\-segmented audio–transcript pairs, improving both robustness and alignment accuracy\.
Gentle produces phoneme\-level annotations in the ARPABET format, consisting of a set of 39 phoneme categories\. ARPABET provides a discrete, standardized phonemic representation that captures the more abstract, relatively invariant characteristics of speech sounds, but does not capture finer phonetic detail such as allophonic variation or diphthong structure\. The same forced\-alignment procedure was applied consistently across all linguistic materials \(audiobooks, TIMIT, MOCHA\-TIMIT, and podcasts\) to ensure internal consistency within the dataset\. Although alternative phonetic annotations \(e\.g\., IPA transcriptions\) are available in the standard releases of the TIMIT and MOCHA\-TIMIT corpora, they were not used to generate event files in this dataset\. Instead, a unified ARPABET representation was adopted to maintain consistency across heterogeneous linguistic materials\.
The forced aligner occasionally produced suboptimal outputs, particularly in challenging cases\. Failures were most frequent for proper names and other out\-of\-vocabulary \(OOV\) items, quotations in different languages, and segments with atypical prosody \(e\.g\., voice imitation or exaggerated emphasis\)\. Additional difficulties arose in more spontaneous speech \(e\.g\., podcasts\), including repetitions, repairs, and unclear or reduced pronunciations\. These cases were identified and corrected through manual inspection\. Audio segments were reviewed by human experts using Praat\([Boersma, 2001](https://arxiv.org/html/2608.25204#bib.bib1)\), with spectrograms and waveforms examined to refine word and phoneme boundaries and recover missed lexical items\. As a result of this intervention, all word\-level annotations were recovered\. Following manual correction, the proportion of OOV annotations was substantially reduced for audiobooks and podcasts\. Overall, this process resulted in high\-quality, densely annotated data suitable for both word\-level and phoneme\-level analyses\.
### B\.7Standard Data Splits
##### Audiobooks\.
For Subject 0, we reuse the same validation and test splits defined in LibriBrain\([Özdogan et al\., 2025](https://arxiv.org/html/2608.25204#bib.bib14)\)\. Here, Sherlock book 1, session 11 is used for validation and session 12 is used for test\. As subjects 1\-32 have no official training data, when training on these subjects, we use the first half of book 1, session 11 for training, the second half for validation, and keep session 12 independent for testing\. When training multi\-subject data jointly with Subject 0, we also use only the second half of Subject 0’s session 11 for validation, adding the first half to the training set\.
##### TIMIT\.
We use the official TIMIT core test speaker set of 24 speakers for our test set and 50 speakers from the dev split\([Garofolo et al\., 1993a](https://arxiv.org/html/2608.25204#bib.bib18)\)\. We exclude allSAlabelled sentence IDs from testing and validation as these are repeated by all speakers \(including those in training\)\. We also exclude any speakers from training who have sentences that overlap with the speakers in the test set\. See Table[5](https://arxiv.org/html/2608.25204#A2.T5)for a complete list of the speakers we use in validation and test\.
Table 5:TIMIT speaker IDs used for the test and validation splits\.Test set\(24 speakers\)mdab0mwbt0felc0mtas1mwew0fpas0mjmp0mlnt0fpkt0mlll0mtls0fjlm0mbpm0mklt0fnlp0mcmj0mjdh0fmgd0mgrt0mnjm0fdhc0mjln0mpam0fmld0Validation set\(50 speakers, from dev split\)faks0fdac1fjem0mgwt0mjar0mmdb1mmdm2mpdf0fcmh0fkms0mbdg0mbwm0mcsh0fadg0fdms0fedw0mgjf0mglb0mrtk0mtaa0mtdt0mthc0mwjg0fnmr0frew0fsem0mbns0mmjr0mdls0mdlf0mdvc0mers0fmah0fdrw0mrcs0mrjm4fcal1mmwh0fjsj0majc0mjsw0mreb0fgjd0fjmg0mroa0mteb0mjfc0mrjr0fmml0mrws1
##### MOCHA\-TIMIT\.
As there is no standard split for MOCHA\-TIMIT, we define our own split, being careful to avoid any sentence stimulus overlap while maintaining an independent test session\. We note that MOCHA\-TIMIT contains four different sets: A, B, C, and D where sentences are repeated between A and D, and between B and C in shuffled orders\. We use all of set A and D for training, then split B and C into validation and test as follows\. We take the first half of the sentences in set B as the validation set\. We then use the sentences in set C that are not in the first half of set B as test data, noting that these test sentences will not be consecutive within set C\. In the Hugging Face repository, sets A, B, C, D correspond to sessions 1, 2, 3, 4\.
##### Podcasts\.
Following[Tang et al\. \(2023\)](https://arxiv.org/html/2608.25204#bib.bib13), two stories were designated to be held\-out from training\. For reproducibility we thus includedFrom Boyhood to Fatherhoodin our validation set andWhere There’s Smokein our test set\. All other stories should be included in training\.
## Appendix CAdditional Stimulus Details
### C\.1Linguistic Variability
One of the central objectives of this dataset is to enable rigorous tests of generalization in neural decoding of speech perception\. In the context of MEG, this is particularly important because models are exposed to substantial variability in natural speech, spanning differences in speakers, phonetic realisations, and semantic content\. Rather than aiming to isolate invariant representations, the goal is to evaluate how well models can achieve robust performance in the presence of such variability, potentially leveraging both variable and more stable features of the signal\. To this end, the dataset was designed to span multiple dimensions of variability that are known to shape speech processing, including speaker\-dependent acoustic properties, phonetic composition, and higher\-level semantic structure\. By systematically varying these factors across complementary corpora, we can assess the extent to which decoding models remain resilient across changes in linguistic context and input statistics\. This is especially relevant for tasks such as phoneme and word classification, where strong performance should be maintained despite differences in speakers, phonetic distributions, and semantic environments\. In addition, the combination of variability and scale allows us to examine how exposure to diverse speech inputs influences learning dynamics, for instance whether it supports faster convergence toward more robust representations that better reflect the range of conditions encountered in natural speech perception\.
At the acoustic level, one of the main goals of the dataset is to introduce substantial variability in speaker characteristics, thereby enabling a more realistic assessment of robustness in speech perception\. To quantify this variability, we performed a speaker embedding analysis using a pretrained ECAPA\-TDNN model\([Desplanques et al\., 2020](https://arxiv.org/html/2608.25204#bib.bib58)\)\. These embeddings capture voice\-related acoustic properties such as pitch, timbre, and vocal tract characteristics, providing a compact representation of speaker differences\. When projected into a low\-dimensional space \(Figure[6](https://arxiv.org/html/2608.25204#A3.F6)\), the embeddings reveal a clear separation between male and female speakers, indicating that they capture meaningful acoustic structure\. This separation is expected, as one of the most salient differences in speech arises from fundamental frequency \(F0\) and vocal tract size: on average, male speakers have larger vocal tracts and produce lower\-pitched sounds, whereas female speakers tend to have smaller vocal tracts and higher\-pitched voices\. Beyond this dominant axis, the embeddings also reflect finer\-grained differences between individual speakers, consistent with their intended role in capturing speaker identity\. Notably, TIMIT spans a broader region of this space compared to the other corpora, indicating wider coverage of speaker variability, in part due to its large number of speakers\. Combining all corpora further expands this coverage, resulting in a richer sampling of pitch\- and timbre\-related dimensions\. This is particularly relevant for MEG neural decoding, as such acoustic features are represented in auditory cortex and contribute to the variability of the neural signal\. By exposing models to this range of speaker\-dependent variation, the dataset allows us to assess how well decoding performance is maintained across acoustic differences, and establishes a foundation for variability not only at the acoustic level but also in the phonetic realisations of speech\.
Figure 6:Speaker embeddings\. Two\-dimensional t\-SNE projection of ECAPA\-TDNN speaker embeddings, which capture acoustic properties such as timbre, pitch, and vocal tract characteristics\. Each point represents a single speaker, obtained by averaging embeddings across multiple utterances\. Marker fill colour indicates the speech corpus \(TIMIT, Podcasts, MOCHA\-TIMIT, Audiobook\), while marker edge colour denotes speaker sex \(male, female\)\.At the phonetic level, this variability extends to the distribution and realisation of speech sounds, providing a complementary axis along which robustness can be evaluated\. TIMIT, podcasts, and audiobooks offer distinct phonetic properties, with TIMIT in particular engineered to achieve a more balanced coverage of phonemes, including relatively rare sounds\. This is accomplished by repeatedly including sentences that contain less frequent phonemes, thereby increasing their occurrence while preserving the structure of natural speech\. As shown in Fig\.[7](https://arxiv.org/html/2608.25204#A3.F7), using approximately 200,000 phoneme tokens per corpus, noticeable differences emerge in the frequency of rare phonemes such as G, Y, SH, OY, and ZH compared to the other corpora\. This pattern is also reflected in the rank–frequency analysis \(Fig\.[8](https://arxiv.org/html/2608.25204#A3.F8)\), where the fitted slopes indicate a flatter distribution for TIMIT \(−0\.7271\-0\.7271\) relative to Podcasts \(−0\.7796\-0\.7796\) and Audiobook \(−0\.8254\-0\.8254\), implying reduced dominance of high\-frequency phonemes and increased representation of low\-frequency ones\. Crucially, the large number of speakers in TIMIT further amplifies phonetic variability by introducing a wide range of accents, speaking styles, and speaker\-specific articulations\. As a result, individual phonemes are realised through a richer set of allophonic variants, and co\-articulatory patterns vary more extensively across contexts\. This combination of balanced phoneme frequencies and diverse phonetic realisations leads to a substantially broader coverage of acoustic–phonetic patterns\. From the perspective of MEG\-based neural decoding, this increased variability provides a stronger test of whether models can maintain performance across both shifts in phoneme statistics and differences in their realisation, and whether increased diversity and scale facilitate the learning of representations that remain effective across heterogeneous linguistic inputs\.
Figure 7:Phoneme frequency distribution per corpus\. Bar plots show normalised phoneme frequencies for each corpus, sorted by decreasing frequency in TIMIT\. Each phoneme category on the x\-axis contains three adjacent bars corresponding to TIMIT \(orange\), Podcasts \(green\), and Audiobook \(blue\)\. The y\-axis is displayed on a logarithmic scale to enhance visibility of low\-frequency phonemes\. Selected low\-frequency phonemes \(G, Y, SH, OY, ZH\) are highlighted with downward\-pointing arrows and thicker bar outlines to facilitate comparison across corpora\.Finally, at the semantic level, the three corpora differ in the richness and diversity of their lexical content, providing a higher\-level dimension of variability\. The audiobook corpus is relatively constrained, as it centres on Sherlock Holmes narratives, leading to redundancy in settings, characters, and thematic structure\. TIMIT, by contrast, spans a broad range of topics but consists of short, largely unrelated sentences, limiting coherence across samples\. The podcast corpus provides the greatest semantic diversity, comprising multiple extended narratives on distinct topics\. These differences are illustrated in Fig\.[9](https://arxiv.org/html/2608.25204#A3.F9), where a pretrained Word2Vec model\([Mikolov et al\., 2013](https://arxiv.org/html/2608.25204#bib.bib57)\)was used to extract word embeddings for representative keywords\. The embeddings were normalised and projected into two dimensions using cosine distance\. In semantic space, keywords show a compact cluster for the audiobook corpus and more dispersed distributions for the TIMIT and, especially, the podcasts corpus\. More generally, the corpora occupy partially distinct regions of the semantic embedding space, reflecting differences in lexical content\. Taken together, they provide broader semantic coverage than any individual corpus alone\. This diversity allows us to examine how decoding models perform across variations in meaning and discourse context, and whether exposure to semantically richer and more varied inputs supports more robust performance\. In combination with the other sources of variability, this also enables us to investigate how scale and diversity jointly influence the stability of learned neural representations in naturalistic speech perception\.
Figure 8:Phoneme rank\-frequency distribution per corpus\. Each point represents a phoneme, with its position determined by its rank \(x\-axis\) and normalised frequency \(y\-axis\), both shown in logarithmic scale\. Phonemes are ranked in decreasing order of frequency based on TIMIT\. Dashed lines indicate linear fits computed over ranks\. Different marker shapes \(circles, squares, triangles\) and different colors \(orange, green, blue\) distinguish the three corpora\.Figure 9:Keyword embeddings per corpus\. Two\-dimensional t\-SNE projection of keyword embeddings from the Audiobook \(blue\), TIMIT \(orange\), and Podcasts \(green\) corpus \. Each point corresponds to a keyword, and proximity between points reflects similarity in the embedding space, capturing lexical\-semantic relationships between words\.
## Appendix DAdditional MEG Details
### D\.1Neural Variability
A challenge in decoding speech perception from MEG lies in the variability of the neural signal itself\. Unlike controlled experimental paradigms, naturalistic listening introduces fluctuations that arise not only from the stimulus, but also from the listener\. In the present dataset, this variability is intrinsic to the design: recordings span multiple participants and multiple sessions per participant, each associated with differences in attention, engagement, fatigue, and individual neurophysiology\. Participants vary in language background \(native vs\. non\-native\), in their level of sustained attention to the narrative, and in cognitive states such as drowsiness or mind wandering\. Even within the same individual, these factors can fluctuate across recording sessions, especially when sessions are conducted consecutively, leading to changes in alertness and engagement\. In addition, inter\-individual differences in brain anatomy, head shape, and age\-related factors influence the spatial configuration and magnitude of the measured MEG signals\. Together, these sources of variability shape the observed neural responses and constitute a fundamental challenge for any model aiming to extract robust neural representations of speech\.
To characterise this variability, we examined three complementary dimensions of the MEG signal — amplitude, phase, and power — contrasting speech and non\-speech segments\. These measures capture distinct aspects of auditory cortical processing\. The root mean square \(RMS\) of the evoked response reflects the strength of stimulus\-locked activity, which is typically enhanced over bilateral temporal sensors during speech perception\. Inter\-trial coherence \(ITC\) quantifies the consistency of phase alignment across trials, and is known to increase at low frequencies when neural activity entrains to the temporal structure of speech\. Finally, power spectral density \(PSD\) provides a frequency\-resolved measure of oscillatory activity, with speech perception commonly associated with increased power in the delta–theta frequency band \(~2–7 Hz\), corresponding to prosodic and syllabic rhythms\([Peelle and Davis, 2012](https://arxiv.org/html/2608.25204#bib.bib41)\)\.
We first assessed the consistency of these neural signatures within a single participant across multiple recording sessions \(Figure[10](https://arxiv.org/html/2608.25204#A4.F10)\)\. The evoked responses reveal clear commonalities: speech segments elicit stronger amplitude over bilateral temporal regions, increased phase alignment at low frequencies, and elevated low\-frequency power for speech segments as compared to non\-speech segments\. These effects are robust at the group level and consistent with established findings in auditory neuroscience\. However, when examining individual sessions in detail \(Figure[12](https://arxiv.org/html/2608.25204#A4.F12)\), substantial variability emerges\. The magnitude of amplitude differences, the degree of phase locking, and the extent of low\-frequency power enhancement vary noticeably across sessions\. Some sessions exhibit pronounced and clean separations between speech and non\-speech, whereas others show attenuated or noisier patterns\. This within\-subject variability likely reflects fluctuations in attention, fatigue, and engagement with the narrative\.
A similar pattern is observed across participants \(Figure[11](https://arxiv.org/html/2608.25204#A4.F11),[13](https://arxiv.org/html/2608.25204#A4.F13)\), where shared neural signatures coexist with inter\-individual variability\. At the group level, the expected response patterns associated with speech perception are preserved: increased evoked amplitude over bilateral auditory cortices, stronger low\-frequency phase coherence, and enhanced delta–theta power for speech relative to non\-speech segments\. Yet, the expression of these effects differs considerably between individuals\. Some participants show strong and spatially focal responses, while others exhibit weaker or more diffuse patterns\. These differences can arise from a combination of anatomical variability, differences in signal\-to\-noise ratio, age\-related factors, and variability in cognitive engagement\. Behavioural measures, such as responses to comprehension questions, further suggest that not all participants maintain the same level of attention throughout the recordings, which likely contributes to the observed heterogeneity in neural responses\.
Figure 10:MEG evoked response within subject\. Differences between speech and non\-speech segments are shown as a proxy for auditory and linguistic processing\. \(a\) Topographic map of the root mean square \(RMS\) difference between speech and non\-speech evoked responses, averaged across sessions, showing stronger amplitude for speech over bilateral temporal sensors\. Statistically significant sensors \(p<0\.001\) are marked with white dots\. \(b\) Time–frequency representation of the difference in inter\-trial coherence \(ITC\) between speech and non\-speech segments, averaged across sessions\. Speech elicits greater phase consistency in low\-frequency bands \(~2–7 Hz; delta–theta range\)\. Statistically significant clusters \(p<0\.001\) are outlined with a black contour\. \(c\) Power spectral density \(PSD\) for speech \(red\) and non\-speech \(blue\) segments, averaged across sessions, indicating increased low\-frequency power during speech\. Statistically significant frequency ranges \(p<0\.001\) are highlighted with a semi\-transparent grey region\. For each panel, data are shown for a single representative participant \(sub\-0\), averaged across recording sessions, where each session corresponds to one chapter \(1\-14\) ofA Study in Scarletfrom Sherlock Holmes audiobook\.Figure 11:MEG evoked response between subjects\. Data are shown for all participants \(sub\-0 to sub\-32\), averaged across subjects, for chapters 11–12 ofStudy in Scarletfrom Sherlock Holmes audiobook\. Differences between speech and non\-speech segments are shown as a proxy for auditory and linguistic processing\. \(a\) Topographic map of the root mean square \(RMS\) difference between speech and non\-speech evoked responses, averaged across subjects, showing stronger amplitude for speech over bilateral temporal sensors\. Statistically significant sensors \(p<0\.001\) are marked with white dots\. \(b\) Time–frequency representation of the difference in inter\-trial coherence \(ITC\) between speech and non\-speech segments, averaged across subjects\. Speech elicits greater phase consistency in low\-frequency bands \(∼\\sim2–7 Hz; delta–theta range\)\. Statistically significant clusters \(p<0\.001\) are outlined with a black contour\. \(c\) Power spectral density \(PSD\) for speech \(red\) and non\-speech \(blue\) segments, averaged across subjects, indicating increased low\-frequency power during speech\. Statistically significant frequency ranges \(p<0\.001\) are highlighted with a semi\-transparent grey region\.Taken together, these analyses show that neural variability operates at multiple levels — both within and across subjects — and directly impacts the structure of the MEG signal used for decoding\. Importantly, despite this variability, \(Figures[10](https://arxiv.org/html/2608.25204#A4.F10)and[11](https://arxiv.org/html/2608.25204#A4.F11)\) reveal consistent and robust neural signatures across sessions and participants, including increased evoked amplitude over temporal regions, stronger low\-frequency phase coherence, and enhanced delta–theta power during speech\. These shared patterns indicate that a stable and meaningful signal is present across conditions, providing a common substrate that neural decoding models can exploit\. At the same time, variability in the strength, spatial distribution, and reliability of these effects reflects fluctuations in attention, physiology, and recording conditions\. As such, this variability is not merely noise to be eliminated, but a defining characteristic of the naturalistic speech perception task\. It provides a critical testbed for evaluating the robustness of deep learning models, which must learn to extract stable and behaviourally relevant features while remaining invariant to these sources of variation\. The dataset therefore enables two complementary forms of generalisation: within\-subject generalisation across sessions and linguistic contexts, and between\-subject generalisation across individuals\. By explicitly incorporating both shared structure and variability, it offers a realistic and challenging benchmark for developing models that can scale to the complexity of naturalistic speech perception\.
Figure 12:Variability across recording sessions\. Spatial, temporal, and spectral components showing strong discrimination between speech and non\-speech segments are illustrated across recording sessions to assess variability in amplitude, phase, and power\. \(a\) Amplitude is quantified as the root mean square \(RMS\) of the evoked response\. \(b\) Phase consistency is measured using inter\-trial coherence \(ITC\)\. \(c\) Oscillatory power is measured using power spectral density \(PSD\)\. For each metric, the mean is computed within functionally relevant masks derived from statistical testing on aggregate data, shown as grey\-shaded insets in each panel: a sensor mask for amplitude, a time–frequency mask for phase, and a frequency mask for power\. Error bars indicate variability within the corresponding mask \(across sensors, time–frequency points, or frequencies, depending on the panel\)\. Data are shown for a single representative participant \(sub\-0\), with each bar corresponding to one recording session, i\.e\., one chapter \(1–14\) ofA Study in Scarletfrom Sherlock Holmes audiobook\.Figure 13:Variability across subjects\. Spatial, temporal, and spectral components showing strong discrimination between speech and non\-speech segments are illustrated across participants to assess variability in amplitude, phase, and power\. \(a\) Amplitude is quantified as the root mean square \(RMS\) of the evoked response\. \(b\) Phase consistency is measured using inter\-trial coherence \(ITC\)\. \(c\) Oscillatory power is measured using power spectral density \(PSD\)\. For each metric, the mean is computed within functionally relevant masks derived from statistical testing on aggregate data, shown as grey\-shaded insets in each panel: a sensor mask for amplitude, a time–frequency mask for phase, and a frequency mask for power\. Error bars indicate variability within the corresponding mask \(across sensors, time–frequency points, or frequencies, depending on the panel\)\. Data are shown across participants \(sub\-0 to sub\-32\), with each bar corresponding to one subject\. Colours differentiate participants\.
## Appendix EAdditional Decoding Experiments
### E\.1Generalisation of Supervised and Pre\-trained Models to Deep and Broad Data
Building on the deep within\-subject data of LibriBrain, LibriBrain100 introduces a newbreadthaxis of scaling through 32 new subjects with relatively little within\-subject data compared to Subject 0\. These two regimes appear to favour different kinds of models, with supervised models performing best on the deep data portion \(Subject 0\), while fine\-tuning a self\-supervised model pre\-trained over many subjects performs best for generalisation over Subjects 1\-32 \(Figure[14](https://arxiv.org/html/2608.25204#A5.F14)\)\. This reflects the findings in[Jayalath and Parker Jones \(2026\)](https://arxiv.org/html/2608.25204#bib.bib36), who observe that with sufficient \(deep\) data, the statistical priors of pre\-trained methods are superseded by rich in\-domain data, whereas on shallow multi\-subject data, the pre\-trained multi\-subject priors assist generalisation\.
Figure 14:d’Ascoli \(Supervised\) vs MEG\-XL \(Self\-supervised\)\.We compare the supervised word decoding model proposed by[d’Ascoli et al\. \(2025\)](https://arxiv.org/html/2608.25204#bib.bib6)to fine\-tuning the pre\-trained self\-supervised MEG\-XL model proposed by[Jayalath and Parker Jones \(2026\)](https://arxiv.org/html/2608.25204#bib.bib36)\. On deep Subject 0 data, the supervised model performs better, while on the shallow data of subjects 1\-32, MEG\-XL generalises better owing to its data\-efficiency\. Random chance is0\.20\.2\. \*\* indicatesp<\.01p<\.01under a Mann\-Whitney U\-test\.
### E\.2Measuring Information Transfer with LibriBrain100
The eventual hope in releasing LibriBrain100 is that it will accelerate progress towards non\-invasive brain–computer interfaces that restore speech to paralysed patients\. To this end, we benchmark the information throughput of models trained on LibriBrain100\. We measure OVMI\([Jayalath et al\., 2026](https://arxiv.org/html/2608.25204#bib.bib35)\), an information\-theoretic quantity for the mutual information between a user’s intent and a decoding model, indexed to a particular communication distribution\. We compare results on LibriBrain100 to that of a prior surgically implanted speech BCI used on a patient with anarthria\([Moses et al\., 2021](https://arxiv.org/html/2608.25204#bib.bib12)\)\. The early landmark result in[Moses et al\. \(2021\)](https://arxiv.org/html/2608.25204#bib.bib12), with a simple 50\-word vocabulary, paved the way for significant advances in following years\([Willett et al\., 2023](https://arxiv.org/html/2608.25204#bib.bib4);[Card et al\., 2024](https://arxiv.org/html/2608.25204#bib.bib5)\)\. By measuring performance against[Moses et al\. \(2021\)](https://arxiv.org/html/2608.25204#bib.bib12), we aim to track how far the non\-invasive decoding field is from reaching a similar breakthrough moment\. The results in Figure[15](https://arxiv.org/html/2608.25204#A5.F15)should be read with caution, however\. LibriBrain100 is a heard speech dataset rather than an intended speech dataset, so the reported throughput does not represent that of a practical communication interface which would require evaluation against attempted or imagined speech decoding\. Nevertheless, the results are indicative of significant progress being made towards strong performance on heard speech decoding through non\-invasive methods\.
Figure 15:Information transfer comparison\.We measure the OVMI achieved by[Moses et al\. \(2021\)](https://arxiv.org/html/2608.25204#bib.bib12), calculating the quantity from their reported47\.1%47\.1\\%decoder accuracy and 50\-word vocabulary, against MEG\-XL fine\-tuned on LibriBrain with our target vocabulary and against an OVMI\-optimised vocabulary\. We use SUBTLEX\-UK as the communication distributionppthat parameterises OVMI\. Higher OVMI scores are better\. The results here are provided as references for future benchmarks\.
## Appendix FAdditional Decoding Experiment Details
### F\.1Computational Requirements
All experiments were performed on individual NVIDIA H100 GPUs on a system with 64 GiB of CPU memory\. Fine\-tuning MEG\-XL with 100 hours of data took approximately 20 hours per run\.
### F\.2Target Words
The 50\-word vocabulary was designed to span both high\-frequency content and function words\. It includes pronouns \(e\.g\., “she”, “him”, “i”, “we”\), conjunctions and prepositions \(e\.g\., “and”, “but”, “on”, “at”\), determiners and quantifiers \(e\.g\., “the”, “a”, “any”\), negation \(e\.g\., “not”\), auxiliary and modal verbs \(e\.g\., “was”, “is”, “will”, “can”\), common lexical verbs \(e\.g\., “think”, “do”\), and a small number of content words \(e\.g\., “people”, “time”, “good”, and “new”\) \(Table[6](https://arxiv.org/html/2608.25204#A6.T6)\)\.
Even with only 50 words, the vocabulary can support a range of short, practically useful utterances in the context of assistive communication\. Examples include state reports \(“i am good”, “i am not good”, “i think this is good”, “i can do that”, “i am out of it”\), requests \(“can i have that”, “can he be there”, “do not do that”\), location or attention cues \(“i will be there”, “is it time”\), and social reference \(“she was good to me”\)\.
\#Word\#Word\#Word\#Word\#Word1is11he21but31at41do2the12that22will32out42can3a13have23so33our43time4to14this24all34am44think5it15they25my35it’s45good6i16of26for36had46always7not17there27she37him47new8was18and28were38an48people9we19are29any39very49as10be20in30really40has50onTable 6:The 50\-word vocabulary used in the word classification task\.Figure 16:Model predictions on full vocabulary\. \(a\) Top\-10 accuracy across all word occurrences in Chapters 11 and 12 of A Study in Scarlet and across all 32 subjects, grouped by Part\-of\-Speech \(POS\) category \(purple function, red adverb, green adjective, orange verb, blue noun\)\. Predictions were evaluated against the full vocabulary, and accuracy is defined as the proportion of events for which the true word appears within the top\-10 predicted words\. \(b\) Word cloud of the 100 best\-predicted words, ranked by top\-10 balanced accuracy\. Word size is proportional to top\-10 balanced accuracy, and word colours indicate their corresponding Part\-of\-Speech categories\.
## Appendix GEthical Considerations
The development and release of the LibriBrain100 dataset raise several important ethical considerations, which we have addressed throughout the research process:
Informed consent and participants privacy\. All participants provided informed consent for data collection and explicitly approved the sharing of pseudoanonymised data for research purposes, in accordance with the University of Oxford’s ethical oversight procedures\. To safeguard privacy, all released data are subject to anonymisation procedures, including the removal or replacement of any information that could directly or indirectly identify individual participants\. These steps are designed to minimise re\-identification risk while preserving the utility of the dataset\.
Open science and reproducibility\. We intentionally rely where possible on public\-domain materials \(e\.g\., LibriVox audiobooks and Project Gutenberg texts\), and use established speech and podcast corpora subject to their respective terms and licences \(TIMIT, MOCHA\-TIMIT, and The Moth\)\. This commitment to openness extends to our open\-source tooling and publicly released data, which are documented to support reproducibility\.
Dual\-use considerations\. While the primary motivation of this work is to support assistive communication technologies for clinical populations, brain decoding methods could be misused in ways that compromise privacy if applied without consent\. Although current non\-invasive approaches remain far from enabling such uses, we stress the importance of maintaining strong ethical standards, including informed consent and respect for participant autonomy\. By releasing data and methods, we aim to support the development of shared ethical guidelines and responsible practices within the research community\.
Long\-term data stewardship\. We are committed to ensuring the long\-term availability and integrity of the dataset\. The BIDS\-formatted version will be hosted on established public platforms \(e\.g\., OSF\), and the dataset is also distributed via Hugging Face\. Together with code hosted on GitHub, this provides redundancy and helps ensure sustained accessibility\.Similar Articles
The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
The article announces the 2026 PNPL Competition focused on word classification and efficient cross-subject generalization in the LibriBrain100 dataset, advancing non-invasive speech decoding for brain-computer interfaces.
@HuggingPapers: LISA: a compact, interpretable MEG decoder for heard speech It retrieves perceived speech from non-invasive brain recor…
LISA is a compact, interpretable MEG decoder that retrieves perceived speech from non-invasive brain recordings with 39.75% Top-1 accuracy among 1,005 candidates using ~20x fewer parameters, while revealing cortical sources and speech features.
BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
BrainBench is a new unified benchmark for evaluating large language models on comprehensive, instruction-conditioned EEG understanding, covering 17 datasets, 172 tasks, and over 4K real-data instances. The paper evaluates 13 LLMs across two execution paradigms, showing that EEG competence varies by model and operationalization.
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
This paper presents an interpretable MEG-to-audio retrieval model for perceived speech, redesigned with spherical-harmonics spatial attention and source mapping, achieving 39.75% Top-1 accuracy with far fewer decoder parameters while revealing which speech features drive retrieval.
The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
This paper introduces Brain2Semantics2Text, a non-invasive speech decoding method that maps MEG responses to semantic embeddings to reconstruct sentence-level text without word-level alignment.