Contextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech Comprehension
Summary
This study demonstrates that contextual semantic relevance, measuring how strongly an incoming word relates to its recent semantic context, reliably predicts fMRI BOLD responses during naturalistic speech comprehension across two datasets, whereas surprisal (local probabilistic expectation) does not. The findings support that slow hemodynamic responses are especially sensitive to contextual semantic integration rather than local prediction.
View Cached Full Text
Cached at: 07/20/26, 09:35 AM
# Contextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech Comprehension
Source: [https://arxiv.org/html/2607.15856](https://arxiv.org/html/2607.15856)
Rong WangDepartment of Computational Linguistics, Tuebingen University††thanks:rong\.wang@uni\-tuebingen\.de
###### Abstract
Naturalistic language comprehension requires listeners to process both local probabilistic expectations and contextual semantic relations\. Surprisal has been widely used to quantify local word unexpectedness, but evidence that it robustly predicts fMRI BOLD responses during continuous comprehension has been mixed\. This study investigates whether contextual semantic relevance, defined as how strongly an incoming word relates to its recent semantic context, predicts BOLD responses during naturalistic speech comprehension\. We analyzed two public fMRI datasets, the Alice dataset and the Moth dataset, treating them as complementary rather than identical replications\. Transformed BOLD responses were modeled with generalized additive mixed models \(GAMMs\) and original continuous BOLD time series were tested with FIR/deconvolution analyses\. In Alice, semantic relevance was significant across all 12 ROIs \(region of interest\), whereas surprisal was not significant after FDR correction\. In Moth, semantic relevance showed consistent negative effects across all 30 ROIs, while surprisal showed no comparable pattern\. These findings suggest that semantic relevance is a promising BOLD\-sensitive metric of contextual semantic fit\. More broadly, our findings support the view that slow hemodynamic responses during naturalistic speech comprehension may be especially sensitive to contextual semantic integration, whereas local probabilistic prediction error may be more difficult to detect reliably with fMRI\. In this sense, semantic relevance extends computational models of language comprehension from prediction alone toward context\-sensitive semantic integration\.
Key Words: surprisal; semantic integration; hemodynamic response; FIR/deconvolution; neural timescale
## 1Introduction
Understanding speech in natural contexts requires listeners to process a rapidly unfolding acoustic signal while building a coherent representation of events, entities, and relations over time\. This process is incremental\. Each incoming word is interpreted against prior context, but it can also update the listener’s current interpretation\. Naturalistic language paradigms are therefore useful because they preserve the temporal continuity, contextual richness, and semantic diversity of everyday comprehension\(Hamilton and Huth,[2020](https://arxiv.org/html/2607.15856#bib.bib23); Huthet al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib19)\)\. They also make it possible to ask whether different computational dimensions of comprehension are expressed differently in neural data\.
A central debate concerns the relation between prediction and integration\. Information\-theoretic accounts propose that the processing cost of a word is related to its negative log probability given the preceding context, or*surprisal*\(Hale,[2001](https://arxiv.org/html/2607.15856#bib.bib3); Levy,[2008](https://arxiv.org/html/2607.15856#bib.bib4)\)\. Prediction\-based accounts further propose that comprehenders can use context to pre\-activate likely upcoming linguistic input, although the strength, level, and necessity of prediction remain debated\(Kuperberg and Jaeger,[2016](https://arxiv.org/html/2607.15856#bib.bib24); Pickering and Gambi,[2018](https://arxiv.org/html/2607.15856#bib.bib25)\)\. Electrophysiological \(EEG\) evidence has often linked predictability and surprisal to N400\-family responses, consistent with rapid lexical\-semantic access or expectation\-related processing\(Kutas and Hillyard,[1984](https://arxiv.org/html/2607.15856#bib.bib5); Kutas and Federmeier,[2011](https://arxiv.org/html/2607.15856#bib.bib6); DeLonget al\.,[2005](https://arxiv.org/html/2607.15856#bib.bib7); Franket al\.,[2015](https://arxiv.org/html/2607.15856#bib.bib8); Lauet al\.,[2008](https://arxiv.org/html/2607.15856#bib.bib26)\)\. Several fMRI studies have also reported surprisal\-related BOLD effects, particularly in temporal and inferior frontal regions\(Willemset al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib10); Brennanet al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib43); Hendersonet al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib57); Shainet al\.,[2020](https://arxiv.org/html/2607.15856#bib.bib44); Songet al\.,[2024](https://arxiv.org/html/2607.15856#bib.bib67)\), although these findings have not been consistently replicated across stimuli, language models, and analysis approaches\. Recent neuroimaging and electrophysiological studies have identified hierarchical prediction\-related responses during natural speech comprehension, although the detectability and spatial distribution of such effects depend on the linguistic level, computational model, and neural measurement employed\(Caucheteuxet al\.,[2023](https://arxiv.org/html/2607.15856#bib.bib54)\)\. Several studies have further shown that incorporating semantic information into surprisal\-based models can improve the modeling of BOLD responses during naturalistic narrative comprehension\(Russoet al\.,[2020](https://arxiv.org/html/2607.15856#bib.bib48)\)\.
However, naturalistic speech comprehension also requires semantic integration\. Incoming words are related to recently processed words, combined with semantic memory, and incorporated into an evolving interpretation of the narrative\. Neurobiological models of language emphasize that comprehension depends on multiple interacting operations, including memory retrieval, unification or integration, prediction, and control\(Hagoort,[2013](https://arxiv.org/html/2607.15856#bib.bib41); Lauet al\.,[2008](https://arxiv.org/html/2607.15856#bib.bib26)\)\. Retrieval\-based accounts further highlight the role of memory access mechanisms in sentence processing, whereby each incoming word triggers cue\-based retrieval from a decaying memory representation\(Lewis and Vasishth,[2005](https://arxiv.org/html/2607.15856#bib.bib14)\)\. In addition, semantic cognition is supported by distributed temporal, inferior parietal, frontal, and medial systems rather than by a single classical language area\(Binder and Desai,[2011](https://arxiv.org/html/2607.15856#bib.bib31); Lambon Ralphet al\.,[2017](https://arxiv.org/html/2607.15856#bib.bib32); Seghier,[2013](https://arxiv.org/html/2607.15856#bib.bib33)\)\. Naturalistic fMRI studies are especially relevant here because they show that semantic information during story comprehension is represented across broad cortical systems\(Wehbeet al\.,[2014](https://arxiv.org/html/2607.15856#bib.bib20); Huthet al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib19); Jain and Huth,[2018](https://arxiv.org/html/2607.15856#bib.bib22); Kaufet al\.,[2024](https://arxiv.org/html/2607.15856#bib.bib58); Sassenhagen and Fiebach,[2020](https://arxiv.org/html/2607.15856#bib.bib59)\)\. Recent evidence suggests that the brain may combine relatively short incoming contexts with incrementally accumulated long\-term contextual representations, rather than processing the entire preceding discourse in parallel\(Tikochinskiet al\.,[2025](https://arxiv.org/html/2607.15856#bib.bib51)\), and embedding\-based transformation could predict human neural activities\(Goldsteinet al\.,[2022](https://arxiv.org/html/2607.15856#bib.bib53); Kumaret al\.,[2024](https://arxiv.org/html/2607.15856#bib.bib52)\)\.
These theoretical debates are closely tied to measurement timescale\. Surprisal may index a fast local prediction\-error\-like signal, which is well suited to EEG and eye\-movement measures\. By contrast, fMRI measures a delayed and temporally smoothed BOLD response\(Boyntonet al\.,[1996](https://arxiv.org/html/2607.15856#bib.bib11); Logothetis,[2003](https://arxiv.org/html/2607.15856#bib.bib18)\)\. In continuous speech, words arrive rapidly and adjacent hemodynamic responses overlap, potentially making brief word\-level prediction\-error effects difficult to isolate\. At the same time, higher\-order cortical regions integrate information over longer temporal windows during naturalistic comprehension\(Hassonet al\.,[2008](https://arxiv.org/html/2607.15856#bib.bib27); Lerneret al\.,[2011](https://arxiv.org/html/2607.15856#bib.bib28); Hassonet al\.,[2015](https://arxiv.org/html/2607.15856#bib.bib29)\)\. More recent evidence further suggests that narrative information flows across a cortical hierarchy characterized by progressively longer processing timescales, with temporally lagged interactions between shorter\- and longer\-timescale regions\(Changet al\.,[2022](https://arxiv.org/html/2607.15856#bib.bib56)\)\. BOLD may therefore be more sensitive to contextual semantic variables that unfold over several seconds than to brief word\-level fluctuations in probabilistic unexpectedness\.
To understand the role of contextual semantic information, the current study examines a computational metric that we define*\(contextual\) semantic relevance*\. Operationally, semantic relevance is computed as a distance\-weighted combination of cosine similarities between a target word’s embedding and those of its three preceding context words, supplemented by pairwise context\-context similarity \(see Section[2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2)for the formal definition\)\. Rather than asking how unexpected a word is, semantic relevance asks how strongly an incoming word relates to, fits with, or contributes to the recent semantic context\. A word can be locally unexpected but still semantically appropriate, and a predictable word can contribute little additional semantic structure\. This distinction motivates the hypothesis that surprisal and semantic relevance are not simply alternative labels for the same process\. Instead, they may index partially different computations and timescales: surprisal reflects local probabilistic expectation, whereas semantic relevance reflects contextual semantic fit or integration\.
The present study examines this account using two public naturalistic speech comprehension fMRI datasets with different statistical structures\. The Alice dataset provides strong group alignment because many participants listened to the same short literary narrative\(Bhattasaliet al\.,[2020](https://arxiv.org/html/2607.15856#bib.bib1)\)\. The Moth dataset provides richer and more heterogeneous natural speech because fewer participants listened to multiple autobiographical podcast stories\(LeBelet al\.,[2023](https://arxiv.org/html/2607.15856#bib.bib2)\)\. The current study addresses three research questions: \(1\) Does semantic relevance predict BOLD responses after controlling for lexical and timing variables, and does it show HRF\-consistent delayed response profiles? \(2\) Is semantic relevance more robust than surprisal across datasets and modeling frameworks? \(3\) Do the two metrics show distinguishable patterns across ROI\-level analyses, consistent with partially different underlying computations?
## 2Materials and Methods
### 2\.1Datasets
##### Alice fMRI dataset\.
The Alice dataset \(abbreviated asAlice\) contains fMRI data from participants listening to the first chapter of*Alice’s Adventures in Wonderland*\(Bhattasaliet al\.,[2020](https://arxiv.org/html/2607.15856#bib.bib1)\)\. The fMRI sample includes 26 participants who listened passively to the same 12\.4\-minute audiobook stimulus\. The stimulus contains 2,129 words and was acquired with TR = 2 s, yielding 372 time points\. This dataset provides a compact and highly aligned naturalistic design: all participants heard the same literary narrative, the same word sequence, and the same discourse progression\. This makes Alice especially useful for group\-level ROI analyses and for estimating shared delayed response profiles with FIR/deconvolution\. In the present study, the Alice dataset serves as the higher\-participant, tightly aligned test case for whether semantic relevance is associated with BOLD responses during naturalistic story comprehension\.
##### Moth fMRI dataset\.
The Moth dataset \(abbreviated asMoth\) contains fMRI responses recorded while 8 participants listened to 27 natural autobiographical stories from the Moth podcasts\(LeBelet al\.,[2023](https://arxiv.org/html/2607.15856#bib.bib2)\)\. Small N \(= 8\) still works for detecting within\-dataset time\-series associations, but it limits population\-level generalization\. Functional data were acquired with TR = 2 s\. Compared with the Alice dataset, the Moth stories are longer, more heterogeneous, and more naturalistic, with variation in speaker identity, speaking style, topic, prosody, and discourse structure\. The Moth dataset therefore provides a complementary test case, and it samples a broader semantic space and contains substantially more speech per participant, but it has fewer participants and weaker group\-level stimulus alignment\. This design is valuable for testing whether semantic relevance effects generalize beyond a single literary narrative to more variable real\-world speech, while also making the statistical detection of ROI\-level effects more challenging\.
##### Complementary role of the two datasets\.
The two datasets allow a stronger test than either dataset alone\. Alice provides sensitivity for shared group\-level effects because many participants heard the same short stimulus\. Moth provides ecological breadth because the stimuli include many natural podcast narratives\. If semantic relevance is associated with BOLD in both datasets, then the effect is unlikely to be specific to one story, one speaker, or one preprocessing pipeline\. At the same time, differences between Alice and Moth are informative: a robust effect in Alice but a more constrained effect in Moth would suggest that semantic relevance is a plausible BOLD\-related metric whose detectability depends on stimulus alignment, participant number, and story\-level heterogeneity\.
### 2\.2Computational Metrics
We compared two word\-level computational metrics: word surprisal and semantic relevance\. The central theoretical contrast is that surprisal measures local probabilistic unexpectedness, whereas semantic relevance measures contextual semantic fit\.
##### Surprisal\.
Surprisal was treated as a metric of local probabilistic unexpectedness\. For a target wordwtw\_\{t\}, surprisal is defined as:
Surprisal\(wt\)=−logP\(wt∣w<t\)\.\\mathrm\{Surprisal\}\(w\_\{t\}\)=\-\\log P\(w\_\{t\}\\mid w\_\{<t\}\)\.\(1\)
In the Alice dataset, we computed word\-level surprisal usingGPT\-2rather than relying on the available surprisal annotations\. In the Moth dataset, surprisal was likewise derived primarily fromGPT\-2, withGPT\-Neoused as a robustness check\. Surprisal was interpreted as a measure of local probabilistic unexpectedness: higher values indicate words that are less expected given the preceding context\.
##### Semantic relevance\.
Semantic relevance quantifies how strongly a target word is semantically connected to its recent local context\. Unlike surprisal, which estimates how unexpected a word is, semantic relevance estimates the degree of semantic fit between the target word and the words that have been processed\. The metric was motivated by the assumption that local contextual information is not equally available during incremental comprehension\. Recent words tend to have stronger influence on current processing than more distant words\(Cowan,[2001](https://arxiv.org/html/2607.15856#bib.bib12); Oberauer,[2002](https://arxiv.org/html/2607.15856#bib.bib13); Lewis and Vasishth,[2005](https://arxiv.org/html/2607.15856#bib.bib14); Christiansen and Chater,[2016](https://arxiv.org/html/2607.15856#bib.bib15)\)\. This local\-context view is also consistent with EEG and eye\-tracking evidence that short\-range semantic context predicts neural and reading\-time responses during naturalistic comprehension\(Frank and Willems,[2017](https://arxiv.org/html/2607.15856#bib.bib9); Brodericket al\.,[2018](https://arxiv.org/html/2607.15856#bib.bib17); Sun and Wang,[2026c](https://arxiv.org/html/2607.15856#bib.bib40); Sunet al\.,[2023](https://arxiv.org/html/2607.15856#bib.bib34); Sun and Wang,[2026a](https://arxiv.org/html/2607.15856#bib.bib39); Sun and Liu,[2024](https://arxiv.org/html/2607.15856#bib.bib38); Sun and Wang,[2022](https://arxiv.org/html/2607.15856#bib.bib37),[2026b](https://arxiv.org/html/2607.15856#bib.bib36); Sunet al\.,[2026](https://arxiv.org/html/2607.15856#bib.bib35)\)\.
The general semantic relevance algorithm combines two sources of information: target\-context semantic fit and context\-context semantic coherence\. For a target wordwtw\_\{t\}, we represent the target and its three preceding context words,wt−3w\_\{t\-3\},wt−2w\_\{t\-2\}, andwt−1w\_\{t\-1\}, using embedding vectors\. In static embedding implementations, these vectors were derived fromModel2vecword embeddings\(Tulkens and van Dongen,[2024](https://arxiv.org/html/2607.15856#bib.bib30)\)\. We first compute the cosine similarity between the target vector and each context\-word vector\. These target\-context similarities estimate how strongly the target word fits the recent semantic context\. To reflect graded contextual availability, the three preceding words receive distance\-based weights, with closer words receiving larger weights\.
The metric also includes pairwise semantic similarities among the three preceding context words\. These context\-context similarities capture the local semantic coherence of the material already processed before the target word is encountered\. In the present metric, target\-context similarity was treated as the direct semantic fit between the incoming word and the recent context, whereas context\-context similarity was treated as the background coherence already present in the preceding context\. Thus, semantic relevance combines two sources of information: the semantic fit between the target word and the recent context, and the semantic coherence within that context\.
Formally, semantic relevance for target wordwtw\_\{t\}is:
SemRel\(wt\)=∑k∈\{t−3,t−2,t−1\}αk⋅Sim\(vt,vk\)\+∑\(k,l\)∈𝒫β⋅Sim\(vk,vl\),\\mathrm\{SemRel\}\(w\_\{t\}\)=\\sum\_\{k\\in\\\{t\-3,t\-2,t\-1\\\}\}\\alpha\_\{k\}\\cdot\\mathrm\{Sim\}\(v\_\{t\},v\_\{k\}\)\+\\sum\_\{\(k,l\)\\in\\mathcal\{P\}\}\\beta\\cdot\\mathrm\{Sim\}\(v\_\{k\},v\_\{l\}\),\(2\)
whereviv\_\{i\}denotes the embedding vector for wordwiw\_\{i\},𝒫=\{\(t−3,t−2\),\(t−3,t−1\),\(t−2,t−1\)\}\\mathcal\{P\}=\\\{\(t\-3,t\-2\),\(t\-3,t\-1\),\(t\-2,t\-1\)\\\}, andβ\\betais the fixed weight assigned to pairwise context\-context similarity\. The second term captures the local semantic coherence already present among the preceding context words\. Cosine similarity is defined as:
Sim\(vi,vj\)=vi⋅vj∥vi∥∥vj∥\.\\mathrm\{Sim\}\(v\_\{i\},v\_\{j\}\)=\\frac\{v\_\{i\}\\cdot v\_\{j\}\}\{\\lVert v\_\{i\}\\rVert\\lVert v\_\{j\}\\rVert\}\.\(3\)
The distance\-based target\-context weights were generated from a linear recency\-decay function\. Motivated by time\-based working\-memory decay models, we approximated recency\-based contextual availability as a word\-distance decay function\(Oberauer and Lewandowsky,[2011](https://arxiv.org/html/2607.15856#bib.bib49); Gauvrit and Mathy,[2018](https://arxiv.org/html/2607.15856#bib.bib50)\)\. During online comprehension, recently encountered words are assumed to remain more accessible in working memory and to exert stronger influence on the interpretation of the target word, whereas more distant words become less available as memory traces decay or are displaced by intervening material\(Cowan,[2001](https://arxiv.org/html/2607.15856#bib.bib12); Oberauer,[2002](https://arxiv.org/html/2607.15856#bib.bib13); Lewis and Vasishth,[2005](https://arxiv.org/html/2607.15856#bib.bib14); Christiansen and Chater,[2016](https://arxiv.org/html/2607.15856#bib.bib15)\)\. We therefore assigned larger weights to more recent context words using the following function:
αt−d=αmax⋅m−d\+1m,\\alpha\_\{t\-d\}=\\alpha\_\{\\max\}\\cdot\\frac\{m\-d\+1\}\{m\},\(4\)
wheredddenotes the distance between the target word and the context word,mmis the size of the local context window, andαmax\\alpha\_\{\\max\}is the maximum weight assigned to the immediately preceding word\. In the present study, we used a three\-word context window \(m=3m=3\) and setαmax=0\.9\\alpha\_\{\\max\}=0\.9, yieldingαt−1=0\.9\\alpha\_\{t\-1\}=0\.9,αt−2=0\.6\\alpha\_\{t\-2\}=0\.6, andαt−3=0\.3\\alpha\_\{t\-3\}=0\.3\. The context\-context weight was fixed atβ=0\.2\\beta=0\.2, lower than the target\-context weights, because this term captures background semantic coherence among the preceding words rather than the direct fit of the target word itself\. All weights and the window size were fixed before any fMRI modeling and should not be interpreted as parameters optimized on brain data\. Alternative weighting functions, such as exponential decay or equal weighting, and alternative window sizes would be useful targets for future confirmatory analyses\. A higher semantic relevance value indicates stronger semantic fit between the target word and its recent context, together with stronger semantic coherence within that context\. The computation of semantic relevance is illustrated in Figure[1](https://arxiv.org/html/2607.15856#S2.F1)\.
Figure 1:Schematic illustration of the semantic relevance metric in a reduced semantic\-vector space\. The example is “… went applebee slice sizzling apple pie put candle”, with “pie” as the target word\. The gray line shows the word sequence path\. The three immediately preceding words, “slice”, “sizzling”, and “apple”, form the local context\. Purple links indicate target\-context semantic fit, computed as cosine similarity between the target\-word vector and each context\-word vector, weighted by recency weightsαk\\alpha\_\{k\}\. Cyan dashed links indicate the context\-coherence component, computed from pairwise similarities among the local context words and weighted byβ\\beta\. The orange arrow schematically indicates that the word\-level semantic relevance predictor is later mapped to BOLD responses through hemodynamic modeling\. The figure is intended as a conceptual visualization of the metric and is not a statistical activation map\.
##### Control predictors\.
Models included lexical and timing controls where available\. Alice models included log \(word frequency\), word length, word duration, and random effects\(Desaiet al\.,[2020](https://arxiv.org/html/2607.15856#bib.bib64); Schusteret al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib65); Gilliset al\.,[2023](https://arxiv.org/html/2607.15856#bib.bib66)\)\. Moth GAMMs included log \(word frequency\), word length, and random effects; the Moth FIR/deconvolution models additionally included run random effects\. FIR/deconvolution analyses also included lagged control predictors and slow time effects\. We note that several potentially relevant nuisance variables, including acoustic envelope, speech rate, phonological complexity, word position in sentence, and semantic diversity, were not included in the current models\. The absence of these controls is a limitation, and future analyses should incorporate them to more precisely isolate the contribution of semantic relevance \(see Section[4\.6](https://arxiv.org/html/2607.15856#S4.SS6)\)\.
### 2\.3FIR/deconvolution Analysis of Original BOLD
To test delayed hemodynamic response profiles, we modeled original continuous BOLD time series using FIR/deconvolution\. FIR/deconvolution preserves the continuous time series and estimates separate effects at multiple post\-word\-onset delays\. This makes it possible to ask whether a predictor has the delayed temporal profile expected of a hemodynamic response\.
For each predictor, lagged regressors were constructed across post\-word\-onset delays\. Alice FIR models tested semantic relevance and surprisal while controlling for FIR\-lagged log \(word frequency\), word length, duration, slow time, and subject effects\. Grouped FIR tests assessed whether each predictor explained BOLD variance across the full lag profile\. This standard FIR/deconvolution approach was appropriate for Alice because all participants listened to the same short narrative, providing strong temporal alignment across subjects\.
For the Moth dataset, ordinary grouped FIR tests were less sensitive to semantic relevance effects\. We therefore used an HRF\-weighted 4–12 s directional contrast as an*exploratory, Alice\-informed*follow\-up analysis, rather than as a fully independent confirmatory test\. This contrast weighted FIR lags corresponding to 4, 6, 8, 10, and 12 s \(i\.e\., lags 2–6 at TR = 2 s\) and tested whether semantic relevance showed a consistent negative delayed effect\. The time window was motivated by the expected timing of the canonical BOLD response and by the delayed semantic relevance profile observed in Alice\. Because the analysis window was selected after observing the Alice results, the Moth directional test should be interpreted as a constrained exploratory analysis, and its results require validation with a preregistered metric and independent held\-out data\. The Moth FIR model included the semantic relevance HRF\-weighted contrast, a surprisal HRF\-weighted contrast, FIR\-lagged log frequency and word length controls, slow time, subject random effects, and run random effects\.
### 2\.4Transformed BOLD and HRF\-convolved Predictors
We used transformed BOLD data for further statistical analyses\. Transformed BOLD refers to derived ROI\-level response tables in which the BOLD signal has been aligned with word\-level or TR\-level predictors and expressed as an analyzable response variable for each ROI, participant, and stimulus time point\. Since such tables can depend on alignment, interpolation, HRF shifting, and the mapping between words and TRs, they are not used as the primary basis for timing\-sensitive inference\.
The Alice dataset included precomputed word\-level transformed BOLD tables from the original dataset release\(Bhattasaliet al\.,[2020](https://arxiv.org/html/2607.15856#bib.bib1)\), in which ROI\-level responses were already aligned to the word sequence\. The Moth dataset did not provide equivalent word\-aligned ROI tables, so we constructed HRF\-convolved predictor time series by convolving word\-level metrics with a canonical HRF and sampling them at the fMRI TR resolution\(LeBelet al\.,[2023](https://arxiv.org/html/2607.15856#bib.bib2)\)\.
Specifically, for Alice, transformed BOLD tables were available at the word/ROI level and included the linguistic metrics, control predictors, subject labels, and transformed BOLD response values\. As word duration was available in Alice, duration was included as a covariate\. For Moth, word\-level metrics were first converted into BOLD\-like predictor time series\. Each word\-level metric was convolved with a canonical HRF and then sampled or aligned at the fMRI TR resolution\. This procedure produces HRF\-convolved predictor series that approximate the delayed BOLD response expected from each linguistic variable\. Duration was omitted from the Moth GAMMs when unavailable\. The HRF\-weighted 4–12 s window was selected based on the delayed semantic\-relevance profile observed in the Alice FIR/deconvolution analysis and on the expected timing of the canonical hemodynamic response\. Therefore, the Moth FIR/deconvolution analysis was treated as an Alice\-informed exploratory convergence analysis rather than a fully independent confirmatory replication\.
As an additional stability check, we performed a split\-half analysis within the Moth dataset by dividing stories into two non\-overlapping subsets and applying the same fixed 4–12 s HRF\-weighted directional test separately to each subset\. In addition, as the GAMM inputs were constructed differently across datasets, differences in GAMM partial\-effect shapes between Alice and Moth should be interpreted cautiously and may reflect preprocessing and alignment choices as well as dataset\-specific neural responses\.
This transformed\-BOLD strategy is complementary to FIR/deconvolution but has important limitations\. Transformed BOLD models are useful for visualizing the shape of partial effects, especially nonlinear associations, but they do not by themselves recover the full hemodynamic response profile and may inflate apparent precision if multiple word\-level observations inherit information from the same TR\-level BOLD sample\. Therefore, we treat GAMMs on transformed BOLD as exploratory effect\-shape analyses, and FIR/deconvolution on original BOLD as the primary analysis for delayed hemodynamic timing\.
### 2\.5GAMM Analysis of Transformed BOLD
We modeled transformed BOLD data using generalized additive mixed models \(GAMMs\)\. GAMMs are useful because they allow continuous predictors to have potentially nonlinear effects while also accounting for repeated observations from the same participants\. However, because the transformed\-BOLD analyses were not the primary basis for hemodynamic timing inference, these models were interpreted as supplementary analyses of effect presence and shape\.
For the Alice dataset, a separate model was fitted for each ROI:We modeled transformed BOLD data using generalized additive mixed models \(GAMMs\)\. GAMMs are useful because they allow smooth nonlinear effects of continuous predictors while also including random\-effect terms for participants\. However, because the transformed\-BOLD tables are not the primary unit for hemodynamic timing inference, these GAMMs were interpreted as supplementary effect\-shape models\. For Alice, each ROI was fit using:
bold∼s\(log\_freq,k=5\)\+s\(wordlen,k=5\)\+s\(duration,k=5\)\+s\(surprisal,k=5\)\+s\(SemRel,k=5\)\+s\(subj,bs=“re”\)\.\\begin\{split\}\\text\{bold\}\\sim&\\ s\(\\text\{log\\\_freq\},k=5\)\+s\(\\text\{wordlen\},k=5\)\+s\(\\text\{duration\},k=5\)\\\\ &\+s\(\\text\{surprisal\},k=5\)\+s\(\\text\{SemRel\},k=5\)\+s\(\\text\{subj\},\\text\{bs\}=\\text\{\`\`re''\}\)\.\\end\{split\}\(5\)
For Moth, duration was omitted when unavailable, and HRF\-convolved predictor time series were used:
bold∼s\(log\_freqHRF,k=5\)\+s\(wordlenHRF,k=5\)\+s\(surprisalHRF,k=5\)\+s\(SemRelHRF,k=5\)\+s\(subj,bs=“re”\)\.\\begin\{split\}\\text\{bold\}\\sim&\\ s\(\\text\{log\\\_freq\}\_\{HRF\},k=5\)\+s\(\\text\{wordlen\}\_\{HRF\},k=5\)\\\\ &\+s\(\\text\{surprisal\}\_\{HRF\},k=5\)\+s\(\\text\{SemRel\}\_\{HRF\},k=5\)\\\\ &\+s\(\\text\{subj\},\\text\{bs\}=\\text\{\`\`re''\}\)\.\\end\{split\}\(6\)
Here,SemRel\\mathrm\{SemRel\}denotes semantic relevance\. The notations\(x\)s\(x\)represents a smooth function of predictorxx, allowing its association with BOLD amplitude to be nonlinear rather than imposing a strictly linear relationship\. The parameterkkspecifies the basis dimension, which sets the maximum complexity available to the smooth\. We usedk=5k=5to permit moderately flexible effect shapes while limiting overfitting; the effective complexity of each smooth was determined during model fitting through penalization\. The terms\(subj,bs=“re”\)s\(\\text\{subj\},\\text\{bs\}=\\text\{\`\`re''\}\)specifies a participant\-level random effect\. It allows each participant to have a different baseline BOLD level, thereby accounting for repeated observations and between\-participant variability\.
Accordingly,s\(SemRel\)s\(\\mathrm\{SemRel\}\)estimates the partial association between semantic relevance and transformed BOLD while holding the other predictors constant, whereass\(surprisal\)s\(\\mathrm\{surprisal\}\)estimates the corresponding partial association for surprisal\. These models were used to determine whether semantic relevance and surprisal were associated with BOLD amplitude and to visualize the shapes of these associations across ROIs\.
##### Complementary role of GAMM and FIR/deconvolution\.
The two statistical approaches answer related but different questions\. GAMMs on transformed BOLD ask whether a predictor is associated with BOLD amplitude and what the partial\-effect curve looks like after controlling for other variables\. FIR/deconvolution asks when the effect appears in the original BOLD time series and whether it follows a plausible delayed HRF profile\. Thus, GAMM provides an interpretable but supplementary effect\-shape analysis, whereas FIR/deconvolution provides timing\-sensitive evidence about delayed BOLD responses\. Because the transformed\-BOLD and FIR analyses can differ in scaling, centering, and temporal alignment, apparent agreement or disagreement between their effect directions should be interpreted cautiously\.
### 2\.6Multiple Comparisons
For ROI\-level tests,p\-values were corrected using Benjamini–Hochberg false discovery rate \(FDR\) correction within explicitly defined test families\. Alice grouped FIR tests were corrected separately for semantic relevance across the 12 ROIs and for surprisal across the 12 ROIs\. Alice GAMM smooth terms were corrected separately within each smooth predictor across the 12 ROIs\. Moth HRF\-weighted FIR tests were corrected separately within each candidate semantic relevance metric across the 30 analyzable ROIs; directional negative tests and two\-sided LRT tests were reported as separate exploratory families\. Moth GAMM smooth terms were corrected across the 30 analyzable ROIs for the plotted predictor\. We report the number of ROIs surviving FDR correction atq<\.05q<\.05\.
## 3Results
### 3\.1Analysis overview
Speech comprehension relies on interacting temporal and frontal systems, including ventral and dorsal processing pathways, rather than on a single classical language region\(Specht,[2014](https://arxiv.org/html/2607.15856#bib.bib68)\)\. To maximize cross\-dataset comparability, the primary analyses focused on the 12 ROIs available in both datasets: AG, CG, FOC, HG, IC, IFG, MTG, PT, SG, SMG, STG, and vmPFC\. These common ROIs provided the basis for evaluating whether semantic relevance and surprisal showed conceptually convergent effects across the Alice and Moth datasets\.
Because the Moth dataset contained substantially more naturalistic speech and provided analyzable transformed\-BOLD and original\-BOLD FIR/deconvolution time series for a broader set of regions, we additionally conducted an exploratory extended analysis across 30 ROIs\. In addition to the 12 common ROIs, the extended Moth analysis included ACC, AI, COC, DLPFC, FUS, HPC, ITG, LING, LOC, OFG, PCG, PCUN, PHG, PP, preSMA, RSC, SFG, and TP\. Accordingly, cross\-dataset comparisons are based primarily on the 12 common ROIs, whereas the expanded 30\-ROI Moth analysis is used to examine whether the observed effects extend to broader auditory, language, memory, control, and narrative\-processing systems\. The 30\-ROI set was smaller than the full collection of anatomical masks available for the Moth dataset because several masks did not have corresponding combined ROI\-level BOLD tables in the current analysis pipeline\.
Collectively, these ROIs support the distributed neural processes involved in naturalistic story listening\. Temporal and auditory regions, including HG, STG, MTG, PT, and ITG, contribute to acoustic, phonological, lexical, and semantic processing\. Frontal, opercular, and insular regions, including IFG, FOC, DLPFC, SFG, preSMA, IC, and AI, support speech\-related processing, working memory, cognitive control, salience detection, and semantic selection\. Parietal regions, including AG, SMG, PP, and PCUN, contribute to attention, multimodal integration, and the construction of coherent situation models\. Medial temporal and default\-mode regions, including HPC, PHG, RSC, vmPFC, and TP, support memory, contextual association, emotional evaluation, and narrative\-level integration\. In short, these regions form a broad network involved in speech perception, language comprehension, memory, cognitive control, and the integration of information across an unfolding narrative\. The commonly associated functions of the individual ROIs are summarized in Table[1](https://arxiv.org/html/2607.15856#A1.T1)in Appendix A\.
Moreover, Pearson correlation analyses showed that semantic relevance and surprisal were only weakly related in both datasets\. In Alice, semantic relevance was weakly negatively correlated with surprisal \(rr= \-0\.18\), and in Moth the correlation was near zero \(rr= \-0\.01\)\. This indicates that semantic relevance and surprisal capture largely distinct word\-level properties rather than two versions of the same metric\. Thus, the stronger BOLD association for semantic relevance cannot be explained simply by high collinearity with surprisal\. The correlation heatmap \(Figure[7](https://arxiv.org/html/2607.15856#A1.F7)\) is presented in Appendix A\.
### 3\.2Results of the Alice Dataset
#### 3\.2\.1Semantic relevance for FIR/deconvolution effects
In the Alice original\-BOLD FIR/deconvolution analysis, the available semantic similarity metric was significant in all 12 ROIs after FDR correction\. In contrast, surprisal was not significant in any ROI after FDR correction\. The corrected semantic relevance effects were strong across the ROI set, with the smallest within\-predictor FDR values observed in STG and MTG\. The lag\-level coefficients showed that semantic relevance effects were concentrated at delayed lags, consistent with a hemodynamic response profile rather than an instantaneous word\-level artifact\. As shown in Figure[2](https://arxiv.org/html/2607.15856#S3.F2), Panel \(a\) provides the anatomical summary of significant ROIs, panel \(b\) ranks ROIs by corrected significance, and panel \(c\) shows the temporal FIR lag profile that reveals when the semantic relevance effect emerges after word onset\. the three panels show that semantic relevance was significantly associated with delayed BOLD responses across Alice ROIs, with effects concentrated in the expected hemodynamic\-response window\. Thus, in Alice, the dataset\-provided semantic relevance metric was robust in both transformed\-BOLD GAMMs and original\-BOLD FIR/deconvolution, whereas surprisal was not robust in the original\-BOLD timing analysis\.
Figure 2:Alice FIR/deconvolution results for the dataset\-provided semantic similarity metric\. \(a\) Brain map of ROI\-level significance, shown as−log10\(FDRq\)\-\\log\_\{10\}\(\\mathrm\{FDR\}\\ q\)\. \(b\) ROI\-ranked significance, with the dashed line markingq=\.05q=\.05and labels showing ROI\-levelttvalues\. \(c\) FIR lag profile from 0 to 16 s after word onset; color shows lag\-specificttvalues, and open circles mark FDR\-significant lag coefficients\. Semantic similarity was significant in all 12 ROIs, whereas surprisal was not significant in any ROI after FDR correction\.
#### 3\.2\.2Transformed\-BOLD GAMMs Results
In the Alice transformed\-BOLD GAMMs, the dataset\-provided semantic similarity metric was a significant smooth term in all 12 ROIs\. The smooth term forsemsimwas significant in AG, CG, FOC, HG, IC, IFG, MTG, PT, SG, SMG, STG, and vmPFC\. Surprisal showed significant smooth effects in 6 of the 12 ROIs \(FOC, HG, IFG, MTG, PT, and STG\), indicating that transformed BOLD can contain local surprisal\-related structure\. Because transformed\-BOLD observations may not be independent at the word level, these GAMM findings are used primarily to describe effect shapes and to motivate the original\-BOLD analyses rather than to support the main inferential claim\. The partial effect of semantic relevance on BOLD responses is visualized in Figure[3](https://arxiv.org/html/2607.15856#S3.F3)\.
Figure 3:Partial effects of the dataset\-provided semantic similarity metric on transformed BOLD in the Alice dataset\. Each panel shows the GAMM\-estimated smooth effect while controlling for log \(word frequency\), word length, duration, surprisal, and random effects\.It is noted that as the Alice transformed\-BOLD GAMMs explained nearly all deviance in several ROI models, these results were treated as descriptive effect\-shape analyses rather than primary inferential evidence\. The high deviance explained likely reflects the strongly transformed and temporally smoothed nature of the dependent variable, which may reduce residual variance and inflate nominal smooth\-term significance\. Accordingly, the main Alice inference is based on the FIR/deconvolution analysis of original BOLD time series\.
### 3\.3The Results of the Moth Dataset
#### 3\.3\.1Moth FIR/deconvolution analysis
The Moth dataset showed a different inferential pattern\. The current Moth FIR/deconvolution analysis included 30 analyzable ROIs, rather than the larger set of anatomical masks available in the project directory\. Because the dataset contains fewer participants and greater story\-level heterogeneity than Alice, ordinary grouped FIR/deconvolution tests were less sensitive to semantic relevance\. We therefore examined an HRF\-weighted 4–12 s directional follow\-up contrast motivated by the Alice delayed\-effect profile and by canonical BOLD timing\.
The split\-half analysis showed qualitatively consistent negative semantic\-relevance effects in both Moth story subsets using the fixed 4–12 s directional FIR/deconvolution test\. The effect was observed in 30/30 ROIs in the first subset and 29/30 ROIs in the second subset after FDR correction\. This supports the stability of the exploratory Moth result, although it should not be interpreted as a fully independent confirmatory replication because the tested window and direction were informed by the Alice analysis\.
Overall, semantic relevance was significant in all 30 analyzable ROIs using the two\-sided LRT FDR criterion and in all 30 analyzable ROIs using the directional negative FDR criterion\. All 30 estimates were negative, with a median estimate of−0\.0238\-0\.0238and a mean absolute t\-value of 2\.94\. The effects are summarized in Figure[4](https://arxiv.org/html/2607.15856#S3.F4)\. In contrast, surprisal did not survive in any ROIs, which is consistent with the Alice dataset\.
Figure 4:Moth FIR/deconvolution results for semantic relevance in original BOLD\. Panel \(a\): Brain\-region visualization of HRF\-weighted 4–12 s effects; warmer colors indicate stronger negative delayed effects, and highlighted regions survived directional FDR correction\. Panel \(b\): ROI\-level summary\. The top plot ranks all 30 analyzable ROIs by effect strength \(−t\-t\), with FDR\-adjustedqqvalues shown beside each ROI; the bottom plot summarizes the number of significant ROIs across the directional FDR, LRT FDR, and negative\-direction tests\. semantic relevance survived correction in all 30 analyzable ROIs in the exploratory Moth analysis\.
#### 3\.3\.2Moth transformed\-BOLD GAMM Analysis
In the Moth transformed\-BOLD analyses, word\-level metrics were first converted into HRF\-convolved predictor time series and then entered into ROI\-level GAMMs\. The plotted Moth GAMM figure shows all 30 analyzable ROIs for the semantic relevance metric after controlling for lexical variables and surprisal\. As the Moth dataset contains fewer participants and greater story\-level heterogeneity, the main inferential claims for Moth are based on the FIR/deconvolution results in the original BOLD time series\. The partial effect of semantic relevance on BOLD responses in each ROI is illustrated in Figure[5](https://arxiv.org/html/2607.15856#S3.F5)\.
Figure 5:Moth GAMM partial effects for semantic relevance on transformed BOLD across the 30 analyzable ROIs\. Thex\-axis shows semantic relevance and they\-axis shows the estimated partial effect on transformed BOLD\.Since semantic relevance showed non\-negligible temporal autocorrelation \(lag\-1 ACF = \.337 in Moth\), the nominal GAMM smooth\-termppvalues should be interpreted cautiously\. Residual temporal autocorrelation can make standard parametric significance tests anti\-conservative\. We therefore treat the GAMM results primarily as effect\-shape analyses and rely on the FIR/deconvolution and time\-structure\-preserving surrogate analyses as stricter tests of temporal specificity\.
The GAMM partial\-effect curves should also be interpreted in light of the different transformed\-BOLD constructions across the two datasets\. In Alice, the available transformed\-BOLD tables expressed ROI responses at the word level, allowing each ROI to retain its own word\-aligned response profile and therefore permitting more ROI\-specific nonlinear smooth shapes\. In Moth, by contrast, word\-level predictors were first convolved with a canonical HRF and sampled at the TR level before entering the GAMMs\. This HRF convolution and TR\-level standardization likely smoothed the predictor\-response relationship and may have contributed to the more uniform partial\-effect shapes across Moth ROIs\. In this sense, differences in the partial\-effect curve shape between Alice and Moth should not be interpreted as purely neural differences\.
The Moth GAMM analysis showed significant nonlinear partial effects of semantic relevance in all 30 ROIs, whereas surprisal was significant in only 5 of the 30 ROIs\. Across these regions, transformed BOLD generally increased from low to moderate\-to\-high surprisal values\. However, these GAMM results should be interpreted separately from the original\-BOLD FIR/deconvolution analyses, in which surprisal did not show robust delayed HRF\-weighted effects\. The partial effect of surprisal is visualized in Figure[8](https://arxiv.org/html/2607.15856#A3.F8)of Appendix[C](https://arxiv.org/html/2607.15856#A3)\.
Across both fMRI datasets, surprisal showed weaker and less consistent BOLD effects than the semantic relevance measures used here\. In Alice, surprisal showed smooth effects in some transformed\-BOLD GAMMs, but it was not significant in any ROI in the original\-BOLD FIR/deconvolution group tests after FDR correction\. In Moth,surprisalwas included as a control in the HRF\-weighted FIR models, but it did not show a robust pattern comparable to semantic relevance metric\. This does not imply that surprisal cannot predict BOLD responses under any condition\. Rather, the present ROI\-level analyses suggest that surprisal effects on BOLD are conditional and less robust than the semantic relevance measures examined in these two naturalistic datasets\. The next section compares the effects of the two metrics directly\.
#### 3\.3\.3Direct comparison
Figure[6](https://arxiv.org/html/2607.15856#S3.F6)directly compares semantic relevance and surprisal effects across the two fMRI datasets\. In Alice, semantic relevance showed a consistent FIR profile across 12 ROIs, with the strongest negative effects appearing approximately 6–16 s after word onset\. By contrast, surprisal showed weak and inconsistent lag\-specific effects, with no clear delayed hemodynamic profile\.
The Moth dataset showed the same qualitative dissociation\. Across 30 ROIs, HRF\-weighted 4–12 s effects for semantic relevance were consistently negative and significant after FDR correction, whereas surprisal effects clustered near zero and were not significant in any ROI\. Both datasets therefore showed the same core pattern\. Semantic relevance was robustly associated with BOLD responses, whereas surprisal showed weaker and less reliable BOLD effects\.
Figure 6:Direct comparison of semantic relevance and surprisal effects on BOLD responses\. \(a\) FIR/deconvolution lag profiles in the Alice dataset\. Heatmaps show lag\-specificttvalues for semantic relevance and surprisal across 12 ROIs from 0 to 16 s after word onset\. Semantic relevance showed a consistent delayed negative effect across ROIs, strongest at approximately 6–16 s, whereas surprisal showed weak and inconsistent lag\-specific effects\. Asterisks mark ROIs with significant semantic\-relevance effects after FDR correction\. \(b\) ROI\-level comparison in the Moth dataset\. Each point represents one ROI\. Thex\-axis shows the HRF\-weighted 4–12 s effect of surprisal, and they\-axis shows the HRF\-weighted 4–12 s effect of semantic relevance\. Semantic relevance was significant in 30/30 ROIs after FDR correction, whereas surprisal was significant in 0/30 ROIs\. The two panels show that semantic relevance, but not surprisal, was robustly associated with BOLD responses during naturalistic speech comprehension\.
### 3\.4Temporal\-structure Robustness Analysis
To assess whether semantic relevance appeared stronger than surprisal partly because of its temporal structure, we conducted a set of robustness analyses comparing the temporal properties of the predictors\. Semantic relevance was temporally smoother and contained more low\-frequency variation than surprisal\. We therefore smoothed surprisal to approximate the lag\-1 autocorrelation of semantic relevance and also tested time\-structure\-preserving semantic\-relevance nulls using circular shifts, phase randomization, and block permutation within stories or runs\. Matched smoothing did not produce surprisal effects comparable to semantic relevance, suggesting that temporal smoothness alone does not explain the stronger BOLD association\.
However, the surrogate analyses warrant caution\. Although the directional FIR/deconvolution analysis showed significant semantic relevance effects across ROIs, the conservative familywise max\-\|t\|\|t\|surrogate tests did not reject the temporal\-structure null\. This indicates that the strongest ROI\-level effect was not larger than expected under surrogate predictors preserving substantial temporal structure\. These analyses provide an important boundary on interpretation, that is, the results support a strong association between semantic relevance and BOLD dynamics, but they do not fully establish semantic specificity beyond temporal structure\. Future work should use larger null distributions, independently fixed metrics, and cross\-validated encoding models to better separate semantic information from temporal\-spectrum matching\. Detailed analyses are reported in Appendix[B](https://arxiv.org/html/2607.15856#A2)\.
### 3\.5Additional Results
First, to provide a quantitative summary of GAMM effect strength, we summarized the smooth\-term test statistics for the main predictors\. In the Alice transformed\-BOLD GAMMs, semantic relevance was significant in all 12 ROIs, with a median smooth\-termFvalue of 14\.19 \(range: 3\.95–200\.98\)\. Surprisal showed weaker and less consistent GAMM effects, with a medianFvalue of 4\.21 \(range: 0\.60–45\.71\) and FDR\-significant smooth terms in 6/12 ROIs\. The model\-level deviance explained was high in the Alice transformed\-BOLD models \(median = 0\.999\), which should be interpreted cautiously because the dependent measure was already transformed/aggregated\.
In the Moth HRF\-convolved GAMMs, semantic relevance showed FDR\-significant smooth terms in 14/30 ROIs \(medianF= 2\.86, range: 1\.05–6\.61\), whereas surprisal showed FDR\-significant smooth terms in 5/30 ROIs \(medianF= 3\.06, range: 1\.16–15\.39\)\. The corresponding model\-level deviance explained was modest\(deviance explained = 0\.0028; adjustedR2R^\{2\}= 0\.0023\), consistent with the small effect sizes typically expected in ROI\-level naturalistic fMRI analyses\. Although the model\-level deviance explained in the Moth HRF\-convolved GAMMs was small \(Rdev2=\.0028R^\{2\}\_\{\\mathrm\{dev\}\}=\.0028\), such small incremental effects are expected for single word\-level predictors in naturalistic fMRI, where BOLD variance is distributed across many linguistic, acoustic, attentional, and physiological sources\(Huthet al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib19)\)\.
Second, semantic relevance was more temporally autocorrelated than surprisal in the Moth dataset, with higher lag\-1 autocorrelation \(0\.337 vs\. 0\.099\) and smoother local trajectories in the representative story segment\. However, this difference was concentrated mainly at short word lags, indicating that semantic relevance is smoother than surprisal but not simply a uniformly low\-frequency predictor, as illustrated in Figure[9](https://arxiv.org/html/2607.15856#A3.F9)in Appendix[C](https://arxiv.org/html/2607.15856#A3)\.
Third, across the six typical ROIs shared by Alice and Moth, semantic relevance showed consistent negative BOLD response profiles, whereas surprisal remained weak and close to zero\. This pattern suggests that the semantic relevance effect generalizes across multiple language\-related regions rather than being driven by a single selected ROI\. The results are illustrated in Figure[10](https://arxiv.org/html/2607.15856#A3.F10)in the Appendix[C](https://arxiv.org/html/2607.15856#A3)\.
Finally, as a supplementary descriptive analysis, we visualized ROI\-to\-ROI BOLD correlations in the Moth dataset to contextualize the predictor\-response findings within broader functional networks\. The results showed strong coupling among language\-related and memory/contextual\-integration regions during narrative listening, suggesting that semantic relevance effects could be interpreted as operating within a distributed comprehension network rather than isolated ROIs\. Our findings are consistent withDenizet al\.\([2019](https://arxiv.org/html/2607.15856#bib.bib60)\)\. These connectivity analyses were descriptive and were not used as statistical tests of semantic relevance or surprisal effects\. The figures are provided in the Appendix[D](https://arxiv.org/html/2607.15856#A4)
## 4Discussion
The present study examined whether semantic relevance could predict BOLD responses during naturalistic speech comprehension, and whether its effects differ from those of surprisal\. Across the Alice and Moth fMRI datasets, semantic relevance showed broader and more consistent associations with BOLD than surprisal in the present ROI\-level analyses\. This pattern supports the view that contextual semantic fit may be especially visible in slow hemodynamic signals, whereas local probabilistic error may be more variable in fMRI\.
### 4\.1Semantic relevance as sustained semantic integration
Semantic relevance differs from surprisal in both computational meaning and expected neural timescale\. Surprisal measures how unexpected a word is given its preceding context\. Semantic relevance instead measures how strongly the current word fits with, relates to, or contributes to the recent semantic context\. In this way, a given word can be highly surprising but still semantically appropriate, and a predictable word may contribute little to the developing discourse representation\.
This distinction helps explain why semantic relevance could be robust in BOLD\. The BOLD response is slow, delayed, and temporally smoothed\(Boyntonet al\.,[1996](https://arxiv.org/html/2607.15856#bib.bib11); Logothetis,[2003](https://arxiv.org/html/2607.15856#bib.bib18)\)\. It is therefore less suited to isolating brief word\-by\-word prediction errors, but well suited to detecting processes that unfold over several seconds\. The temporally overlapping local\-context estimates may track contextual semantic structure that changes continuously during comprehension\. This kind of sustained contextual processing is compatible with previous naturalistic fMRI work showing that semantic information is represented across distributed cortical systems during story comprehension\(Huthet al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib19); Wehbeet al\.,[2014](https://arxiv.org/html/2607.15856#bib.bib20); Jain and Huth,[2018](https://arxiv.org/html/2607.15856#bib.bib22)\)\.
The same account also links the present fMRI findings to EEG and eye\-movement evidence\. In EEG, semantic fit and contextual integration are often reflected in N400\-family responses, which are sensitive to how easily a word can be integrated with prior context\(Kutas and Hillyard,[1984](https://arxiv.org/html/2607.15856#bib.bib5); Kutas and Federmeier,[2011](https://arxiv.org/html/2607.15856#bib.bib6); Franket al\.,[2015](https://arxiv.org/html/2607.15856#bib.bib8)\)\. In eye movements and reading\-time measures, semantically well\-fitting words are expected to reduce processing difficulty, whereas weakly fitting words may increase fixation time, regressions, or later integration cost\. These measures capture faster behavioral and EEG consequences of semantic fit, and BOLD captures the slower hemodynamic consequence of related integration and updating processes\. As a result, semantic relevance can plausibly affect EEG, eye movements, and BOLD, but with different temporal signatures\.
Despite theses, the sign of the BOLD effect should be interpreted cautiously\. In the Moth FIR analyses, stronger semantic relevance was consistently associated with reduced delayed BOLD amplitude across all 30 ROIs\. This uniformly negative direction deserves consideration\. A negative association may indicate easier integration, reduced updating demands, or more efficient use of an existing context representation which is analogous to the repetition suppression or predictability facilitation effects observed in other domains\(Grill\-Spectoret al\.,[2006](https://arxiv.org/html/2607.15856#bib.bib45)\)\. That is, when an incoming word fits well with the recent semantic context, the neural system may require less metabolic effort to integrate it, producing lower BOLD amplitude\. Alternatively, in default\-mode regions \(e\.g\., vmPFC, PCUN, RSC\), reduced BOLD during high semantic relevance could reflect reduced default\-mode deactivation when contextual fit is strong\. The fact that the strongest effects in Moth were observed in medial and default\-mode regions \(CG, ACC, vmPFC, PCUN\) rather than in classical perisylvian language areas is consistent with the view that semantic relevance tracks discourse\-level integration and situation\-model updating rather than purely lexical\-semantic access\. In other paradigms or datasets, positive effects may reflect stronger semantic binding, contextual updating, or sustained discourse\-level processing\. Therefore, semantic relevance should be interpreted as an index of contextual semantic fit whose neural expression may vary by region, task structure, and temporal model\.
### 4\.2Why surprisal effects on BOLD may be mixed
The weaker and less consistent surprisal effects should not be taken as evidence that surprisal is unimportant for language comprehension\. Surprisal remains a central metric of local probabilistic prediction and processing difficulty\. It is also closely related to EEG findings, especially N400 effects, where millisecond\-level temporal resolution can capture rapid lexical\-semantic prediction and access processes\(Kutas and Federmeier,[2011](https://arxiv.org/html/2607.15856#bib.bib6); Franket al\.,[2015](https://arxiv.org/html/2607.15856#bib.bib8); Willemset al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib10)\)\.
However, several factors may make surprisal effects harder to detect in BOLD\. First, surprisal is an event\-level signal tied to the moment a word is encountered, whereas BOLD integrates neural activity over a much longer time window\. In continuous speech, words arrive quickly, their hemodynamic responses overlap, and local prediction\-error effects may be blurred across adjacent words\. Second, surprisal estimates depend strongly on the language model, tokenizer, context length, and corpus used to estimate word probability\. Different models may capture different kinds of expectation, making fMRI effects sensitive to predictor construction\. Third, surprisal is correlated with lexical frequency, word length, and acoustic timing\. Once these covariates are included, the unique variance left for surprisal may be small\. Fourth, surprisal may be regionally specific rather than broadly distributed, and it may appear in regions involved in lexical access, syntactic prediction, or control, but not necessarily across all semantic\-language ROIs\. If surprisal effects are spatially focal, then ROI\-level analyses averaged across large anatomical regions may dilute these effects, and voxel\-wise encoding models would provide a more sensitive test\(Huthet al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib19); Jain and Huth,[2018](https://arxiv.org/html/2607.15856#bib.bib22)\)\.
These considerations may explain why previous fMRI findings on surprisal have been mixed\. Studies using different stimuli, language models, HRF assumptions, control variables, and analysis levels may reasonably obtain different results\. The present findings therefore suggest a cautious conclusion: surprisal may affect BOLD under some conditions, but its hemodynamic signature is likely weaker, more model\-dependent, and more temporally fragile than the effect of semantic relevance\.
Additionally, the two datasets also differed in how easily semantic relevance was detected\. In Alice, all participants listened to the same well\-known short literary narrative, producing strong temporal and semantic alignment across participants\(Bhattasaliet al\.,[2020](https://arxiv.org/html/2607.15856#bib.bib1)\)\. This alignment likely increased group\-level sensitivity, allowing ordinary FIR/deconvolution analyses to detect semantic relevance effects across the available ROIs\. In Moth, participants listened to multiple podcast stories with greater variability in speaker, topic, prosody, narrative structure, and discourse context\(LeBelet al\.,[2023](https://arxiv.org/html/2607.15856#bib.bib2)\)\. This design is richer and more naturalistic, but it is also less aligned for ROI\-level group inference\. Under these conditions, semantic relevance effects were clearest when the analysis used an HRF\-weighted 4–12 s directional FIR/deconvolution contrast\. This difference does not necessarily imply that the Moth effect is weaker in the brain; rather, it suggests that heterogeneous naturalistic datasets may require more temporally targeted models to recover the relevant hemodynamic profile\.
### 4\.3Prediction and semantic integration as complementary computations
The present findings speak to a central debate in language comprehension: whether neural responses during comprehension primarily reflect prediction, integration, or the interaction between the two\. Surprisal provides an information\-theoretic metric of local probabilistic expectation, indexing how unexpected a word is given the preceding context\(Hale,[2001](https://arxiv.org/html/2607.15856#bib.bib3); Levy,[2008](https://arxiv.org/html/2607.15856#bib.bib4)\)\. Prediction\-based accounts propose that comprehenders use context to pre\-activate likely upcoming linguistic input, but the extent and level of prediction remain debated\(Kuperberg and Jaeger,[2016](https://arxiv.org/html/2607.15856#bib.bib24); Pickering and Gambi,[2018](https://arxiv.org/html/2607.15856#bib.bib25)\)\. Semantic relevance, in contrast, indexes how well the incoming word fits the recent semantic context\. It is therefore closer to semantic integration or contextual fit than to local probability error\.
This distinction helps clarify why surprisal and semantic relevance should not be treated as competing versions of the same predictor\. Recent fMRI evidence further suggests that linguistic prediction and contextual updating are hierarchically organized across representational levels and timescales\(Zhouet al\.,[2026](https://arxiv.org/html/2607.15856#bib.bib63)\)\. Prediction and integration should therefore be viewed as interacting rather than strictly separable computations\. A word can be locally unexpected but still semantically coherent, and a predictable word can add little to the current semantic representation\. Surprisal may therefore capture the cost of local expectation violation or lexical access, whereas semantic relevance may capture how easily a word can be incorporated into an evolving contextual representation\. This view is consistent with models in which language comprehension depends on multiple interacting operations, including memory retrieval, unification, prediction, and control\(Hagoort,[2013](https://arxiv.org/html/2607.15856#bib.bib41); Lauet al\.,[2008](https://arxiv.org/html/2607.15856#bib.bib26)\)\.
Additionally, previous fMRI studies have sometimes reported significant surprisal effects, but these findings were often restricted to a small number of ROIs rather than distributed across the broader language network\. In some analyses, limited control for lexical, acoustic, timing, or discourse\-related variables may also have allowed surprisal to absorb variance attributable to correlated factors, thereby maximizing its apparent contribution\. Also, surprisal effects are sensitive to the language model, tokenization, context length, HRF specification, and statistical framework used\. Differences in stimulus structure, regional averaging, sample size, and temporal autocorrelation may therefore explain why surprisal is significant in some datasets or brain regions but not consistently across studies\.
The present results therefore support a complementary\-computation interpretation\. They do not show that prediction is unimportant, nor do they imply that semantic relevance replaces surprisal\. Rather, they suggest that contextual semantic fit may be more robustly visible in slow BOLD responses than local probabilistic error\. This provides a possible bridge between psycholinguistic theories of prediction and neurobiological accounts of semantic integration\.
### 4\.4Fast EEG responses and slow hemodynamic timescales
Another implication concerns the temporal scales at which lexical expectation and contextual semantic fit become detectable\. Surprisal and predictability effects are well established in temporally precise measures such as EEG and eye movements\. In EEG, predictable or semantically fitting words typically elicit reduced N400 amplitudes, linking contextual expectation and semantic access to neural responses within a few hundred milliseconds after word onset\(Kutas and Hillyard,[1984](https://arxiv.org/html/2607.15856#bib.bib5); Kutas and Federmeier,[2011](https://arxiv.org/html/2607.15856#bib.bib6); Lauet al\.,[2008](https://arxiv.org/html/2607.15856#bib.bib26); Franket al\.,[2015](https://arxiv.org/html/2607.15856#bib.bib8)\)\. Eye\-movement studies similarly show that word predictability and contextual constraint influence fixation durations and other reading behaviors during incremental comprehension\(Rayneret al\.,[1996](https://arxiv.org/html/2607.15856#bib.bib46); Smith and Levy,[2013](https://arxiv.org/html/2607.15856#bib.bib47); Sunet al\.,[2026](https://arxiv.org/html/2607.15856#bib.bib35)\)\.
A companion naturalistic\-reading EEG analysis provides converging evidence\(Sun and Wang,[2026c](https://arxiv.org/html/2607.15856#bib.bib40)\)\. Across 22 participants and 32 channels, both semantic relevance and surprisal were associated with N400 and P600 responses, but with partly distinguishable temporal profiles\. Semantic relevance showed broader effects in the N400 window, whereas surprisal exhibited comparatively stronger modulation around the P600 interval\. The weak correlation between the two metrics in that dataset further suggests that they capture partly distinct aspects of word\-level processing\.
These EEG results indicate that semantic relevance should not be understood exclusively as a slow process\. Recent ECoG evidence similarly shows that earlier and deeper language\-model representations correspond to earlier and later stages of cortical language processing, respectively\(Goldsteinet al\.,[2025](https://arxiv.org/html/2607.15856#bib.bib69)\)\. Computational measures differing in contextual depth may therefore be most detectable at different neural timescales\. Contextual semantic fit can influence neural activity within several hundred milliseconds, affecting both early semantic access or integration and later updating or reanalysis\. At the same time, EEG preserves rapid word\-locked fluctuations that may be attenuated in fMRI\. BOLD responses are delayed and temporally smoothed\(Boyntonet al\.,[1996](https://arxiv.org/html/2607.15856#bib.bib11); Logothetis,[2003](https://arxiv.org/html/2607.15856#bib.bib18)\), and in continuous speech the hemodynamic responses elicited by adjacent words overlap substantially\. A brief and rapidly varying surprisal\-related response may therefore be difficult to isolate in BOLD, particularly after controlling for correlated lexical variables such as frequency, word length, and duration\.
Semantic relevance may be comparatively well aligned with the temporal characteristics of BOLD because it is computed from overlapping relations between each target word and its recent context\. Although each estimate is word\-specific, adjacent estimates share contextual information and may track semantic structure that evolves continuously over several words\. Such temporally extended contextual variation may survive hemodynamic smoothing more readily than highly local fluctuations in probabilistic unexpectedness\. This interpretation is compatible with evidence for hierarchical temporal receptive windows across cortex, whereby early auditory and language regions respond to relatively short\-timescale information, whereas higher\-order temporal, parietal, and default\-mode regions integrate information over longer periods during naturalistic comprehension\(Hassonet al\.,[2008](https://arxiv.org/html/2607.15856#bib.bib27); Lerneret al\.,[2011](https://arxiv.org/html/2607.15856#bib.bib28); Hassonet al\.,[2015](https://arxiv.org/html/2607.15856#bib.bib29)\)\.Changet al\.\([2022](https://arxiv.org/html/2607.15856#bib.bib56)\)further showed that information during narrative comprehension propagates across this cortical timescale hierarchy, providing a systems\-level framework for understanding why context\-sensitive semantic variables may be expressed as delayed and spatially distributed BOLD responses\.
In sum, the EEG and fMRI findings suggest a difference in measurement sensitivity rather than a strict separation between fast and slow linguistic computations\. Both surprisal and semantic relevance can modulate rapid EEG responses, but semantic relevance may also generate temporally sustained variation that is more readily detectable in the slow BOLD signal\. Surprisal\-related BOLD effects may therefore be more dependent on stimulus structure, language\-model specification, temporal alignment, and the control variables included in the analysis\.
### 4\.5Naturalistic semantic networks and situation\-model updating
The use of naturalistic spoken narratives is central to the present study\. Real story comprehension requires listeners to maintain and update characters, events, goals, causal relations, and discourse context over time\. Naturalistic fMRI studies have shown that semantic information is represented across distributed cortical systems rather than being confined to classical language areas\(Huthet al\.,[2016](https://arxiv.org/html/2607.15856#bib.bib19); Wehbeet al\.,[2014](https://arxiv.org/html/2607.15856#bib.bib20); Jain and Huth,[2018](https://arxiv.org/html/2607.15856#bib.bib22)\)\. This broader organization makes naturalistic datasets especially useful for testing predictors related to contextual semantic fit\.
Semantic relevance may therefore relate to a distributed comprehension network\. The present ROI\-level results are broadly consistent with this view but also reveal an informative gradient\. In the Moth dataset, the strongest semantic relevance effects were observed in CG, ACC, vmPFC, and PCUN \(these regions associated with cognitive control, contextual integration, and situation\-model construction \) rather than in classical perisylvian language areas such as STG, MTG, or IFG, which showed weaker \(though still significant\) effects\. This pattern is consistent with the idea that semantic relevance primarily indexes discourse\-level contextual fit and situation\-model updating\(Ferstlet al\.,[2008](https://arxiv.org/html/2607.15856#bib.bib42)\), rather than lower\-level lexical\-semantic access\. In Alice, by contrast, the strongest FIR effects were in STG and MTG, possibly reflecting the stronger temporal alignment and stimulus control in that dataset\. These regional differences across datasets underscore the need for future voxel\-wise analyses that can more precisely localize the neural sources of semantic relevance effects\(Binder and Desai,[2011](https://arxiv.org/html/2607.15856#bib.bib31); Lambon Ralphet al\.,[2017](https://arxiv.org/html/2607.15856#bib.bib32); Seghier,[2013](https://arxiv.org/html/2607.15856#bib.bib33)\)\.
At the same time, widespread effects should be interpreted cautiously\. Broad ROI\-level significance, especially in the Moth dataset, does not imply that all regions encode semantic relevance in the same way\. Some effects may reflect core semantic processing, whereas others may reflect attention, memory, narrative structure, global temporal dynamics, or unmodeled acoustic and discourse\-level covariates\. The safe conclusion is that semantic relevance tracks a naturalistic comprehension signal that plausibly relates to contextual semantic integration and situation\-model updating, but future work should use nuisance controls, temporal surrogate analyses, and held\-out encoding models to determine which components of this distributed response are specifically semantic\.
### 4\.6Methodological implications
The results have broader implications for computational neuroscience of language\. Computational predictors should not be treated as interchangeable regressors simply because they are derived from the same words\. Surprisal and semantic relevance make different claims about neural computation\. Surprisal indexes local probabilistic unexpectedness, whereas semantic relevance indexes contextual semantic fit and integration\. These computations may operate over partly different timescales and may therefore require different statistical models\.
Recent work has demonstrated that GPT\-based encoding models can predict sentence\-level responses in the human language network and can identify novel sentences that reliably drive or suppress its activity, with surprisal and linguistic well\-formedness emerging as important determinants of response strength\(Tuckuteet al\.,[2024](https://arxiv.org/html/2607.15856#bib.bib61)\)\. These findings illustrate the power of computational representations for predicting and manipulating neural responses to language\. Nevertheless, successful brain prediction by a computational representation does not by itself establish that the brain implements the model’s training objective or computational mechanism\(Antonello and Huth,[2024](https://arxiv.org/html/2607.15856#bib.bib55)\)\. Accordingly, the stronger association between semantic relevance and BOLD in the present analyses should be interpreted as evidence for the neural sensitivity of the metric, rather than as proof that the brain explicitly computes semantic relevance in the precise form implemented here\.
The complementary use of GAMMs and FIR/deconvolution is important for this reason\. GAMMs applied to transformed or HRF\-convolved BOLD are useful for testing nonlinear predictor\-response relationships while controlling lexical and subject\-level variability\. FIR/deconvolution applied to original continuous BOLD is useful for estimating the delayed response profile without forcing a single canonical HRF shape\. In short, these methods allow one to ask both whether a predictor explains BOLD variance and when its effect emerges after word onset\. This issue is especially important because BOLD is a low\-pass signal and semantic relevance is itself temporally smoother than surprisal\.
Recent work has introduced methods for estimating semantic relevance and shown that these metrics predict eye movements during reading across multiple languages\(Sun and Liu,[2024](https://arxiv.org/html/2607.15856#bib.bib38); Sunet al\.,[2026](https://arxiv.org/html/2607.15856#bib.bib35)\)\. Related studies further suggest that semantic relevance may contribute to modeling speech production, spontaneous speech characteristics, EEG responses, and visual attention\(Sun and Wang,[2026b](https://arxiv.org/html/2607.15856#bib.bib36),[c](https://arxiv.org/html/2607.15856#bib.bib40),[a](https://arxiv.org/html/2607.15856#bib.bib39)\)\. The present study extends this line of research by demonstrating that semantic relevance also predicts BOLD responses during language comprehension\.
In spite of these, several limitations should be considered\. As semantic relevance is temporally smoother than surprisal, some effects may reflect better alignment with BOLD’s low\-frequency structure rather than semantic processing alone\. In addition, unmodeled factors such as acoustic properties, speech rate, story progression, motion, scanner drift, global signal, and temporal autocorrelation may contribute to the observed effects\. The Moth dataset also included only eight participants, limiting group\-level generalizability despite its rich within\-participant data\. Finally, the fMRI findings are related to companion EEG evidence only at the interpretive level rather than through a joint cross\-modal analysis\.
## 5Conclusion
Across two naturalistic fMRI datasets, semantic relevance was associated with BOLD responses during spoken\-language comprehension, whereas surprisal showed weaker and less consistent effects in these ROI\-level analyses\. These findings suggest that contextual semantic fit may be more readily expressed in slow hemodynamic responses than local word\-level unexpectedness\. The contribution of semantic relevance to computational neuroscience is that it provides a computable bridge between word\-level language models and slower neural dynamics measured with fMRI\. Unlike surprisal, which captures forward prediction error, semantic relevance quantifies how an incoming word fits into the recent semantic context, making it useful for modeling sustained integration processes that unfold across discourse and engage distributed cortical systems\.
This distinction links different computational descriptions of language to different neural measurement timescales\. Surprisal may primarily capture fast prediction\-related processing, whereas semantic relevance may capture more sustained contextual integration that is better matched to the temporal dynamics of BOLD\. More broadly, the results argue that models of naturalistic language comprehension should move beyond forward prediction alone and explicitly incorporate contextual coherence, semantic fit, and discourse\-level integration as core components of neural language processing\. Semantic relevance may therefore provide a useful bridge between computational models of lexical processing and neural systems supporting semantic integration, narrative comprehension, and situation\-model updating\.
## Data and Code Availability
All analyses were conducted using publicly available fMRI datasets\. The Alice data were obtained from the Alice dataset release, and the Moth data were obtained from the natural language fMRI dataset for voxelwise encoding models\. Analysis scripts, intermediate result tables, and generated figures are currently stored in the local project directory\. The public repository is available at:\.
## Acknowledgments
We thank the creators of the Alice and Moth datasets for making their naturalistic fMRI resources publicly available\.
## References
- R\. Antonello and A\. Huth \(2024\)Predictive coding or just feature discovery? an alternative account of why language models fit brain data\.Neurobiology of Language5\(1\),pp\. 64–79\.Cited by:[§4\.6](https://arxiv.org/html/2607.15856#S4.SS6.p2.1)\.
- S\. Bhattasali, J\. R\. Brennan, W\. Luh, B\. Franzluebbers, and J\. T\. Hale \(2020\)The alice datasets: fmri & eeg observations of natural language comprehension\.InProceedings of the 12th Conference on Language Resources and Evaluation \(LREC 2020\),pp\. 120–125\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p6.1),[§2\.1](https://arxiv.org/html/2607.15856#S2.SS1.SSS0.Px1.p1.1),[§2\.4](https://arxiv.org/html/2607.15856#S2.SS4.p2.1),[§4\.2](https://arxiv.org/html/2607.15856#S4.SS2.p4.1)\.
- J\. R\. Binder and R\. H\. Desai \(2011\)The neurobiology of semantic memory\.Trends in Cognitive Sciences15\(11\),pp\. 527–536\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1),[§4\.5](https://arxiv.org/html/2607.15856#S4.SS5.p2.1)\.
- G\. M\. Boynton, S\. A\. Engel, G\. H\. Glover, and D\. J\. Heeger \(1996\)Linear systems analysis of functional magnetic resonance imaging in human v1\.Journal of Neuroscience16\(13\),pp\. 4207–4221\.External Links:[Document](https://dx.doi.org/10.1523/JNEUROSCI.16-13-04207.1996)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p4.1),[§4\.1](https://arxiv.org/html/2607.15856#S4.SS1.p2.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p3.1)\.
- J\. R\. Brennan, E\. P\. Stabler, S\. E\. Van Wagenen, W\. Luh, and J\. T\. Hale \(2016\)Abstract linguistic structure correlates with temporal activity during naturalistic comprehension\.InBrain and Language,Vol\.157–158,pp\. 81–94\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1)\.
- M\. P\. Broderick, A\. J\. Anderson, G\. M\. Di Liberto, M\. J\. Crosse, and E\. C\. Lalor \(2018\)Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narrative speech\.Current Biology28\(5\),pp\. 803–809\.e3\.External Links:[Document](https://dx.doi.org/10.1016/j.cub.2018.01.080)Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1)\.
- C\. Caucheteux, A\. Gramfort, and J\. King \(2023\)Evidence of a predictive coding hierarchy in the human brain listening to speech\.Nature Human Behaviour7\(3\),pp\. 430–441\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1)\.
- C\. H\. C\. Chang, S\. A\. Nastase, and U\. Hasson \(2022\)Information flow across the cortical timescale hierarchy during narrative construction\.Proceedings of the National Academy of Sciences119\(51\),pp\. e2209307119\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p4.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p4.1)\.
- M\. H\. Christiansen and N\. Chater \(2016\)The now\-or\-never bottleneck: a fundamental constraint on language\.Behavioral and Brain Sciences39,pp\. e62\.External Links:[Document](https://dx.doi.org/10.1017/S0140525X1500031X)Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p8.1)\.
- N\. Cowan \(2001\)The magical number 4 in short\-term memory: a reconsideration of mental storage capacity\.Behavioral and Brain Sciences24\(1\),pp\. 87–114\.External Links:[Document](https://dx.doi.org/10.1017/S0140525X01003922)Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p8.1)\.
- K\. A\. DeLong, T\. P\. Urbach, and M\. Kutas \(2005\)Probabilistic word pre\-activation during language comprehension inferred from electrical brain activity\.Nature Neuroscience8,pp\. 1117–1121\.External Links:[Document](https://dx.doi.org/10.1038/nn1504)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1)\.
- F\. Deniz, A\. O\. Nunez\-Elizalde, A\. G\. Huth, and J\. L\. Gallant \(2019\)The representation of semantic information across human cerebral cortex during listening versus reading is invariant to stimulus modality\.Journal of Neuroscience39\(39\),pp\. 7722–7736\.Cited by:[§3\.5](https://arxiv.org/html/2607.15856#S3.SS5.p5.1)\.
- R\. H\. Desai, W\. Choi, and J\. M\. Henderson \(2020\)Word frequency effects in naturalistic reading\.Language, Cognition and Neuroscience35\(5\),pp\. 583–594\.Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px3.p1.1)\.
- E\. C\. Ferstl, J\. Neumann, C\. Bogler, and D\. Y\. von Cramon \(2008\)The role of the medial frontal cortex in text comprehension\.NeuroImage40\(4\),pp\. 1661–1673\.Cited by:[§4\.5](https://arxiv.org/html/2607.15856#S4.SS5.p2.1)\.
- S\. L\. Frank, L\. J\. Otten, G\. Galli, and G\. Vigliocco \(2015\)The erp response to the amount of information conveyed by words in sentences\.Brain and Language140,pp\. 1–11\.External Links:[Document](https://dx.doi.org/10.1016/j.bandl.2014.10.006)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1),[§4\.1](https://arxiv.org/html/2607.15856#S4.SS1.p3.1),[§4\.2](https://arxiv.org/html/2607.15856#S4.SS2.p1.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p1.1)\.
- S\. L\. Frank and R\. M\. Willems \(2017\)Word predictability and semantic similarity show distinct patterns of brain activity during language comprehension\.Language, Cognition and Neuroscience32\(9\),pp\. 1192–1203\.External Links:[Document](https://dx.doi.org/10.1080/23273798.2017.1323109)Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1)\.
- N\. Gauvrit and F\. Mathy \(2018\)Mathematical transcription of the time\-based resource sharing theory of working memory\.British Journal of Mathematical and Statistical Psychology71\(1\),pp\. 146–166\.Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p8.1)\.
- M\. Gillis, J\. Vanthornhout, and T\. Francart \(2023\)Heard or understood? neural tracking of language features in a comprehensible story, an incomprehensible story and a word list\.eNeuro10\(7\),pp\. ENEURO\.0075–23\.2023\.Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px3.p1.1)\.
- A\. Goldstein, E\. Ham, M\. Schain, S\. A\. Nastase, B\. Aubrey, Z\. Zada,et al\.\(2025\)Temporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models\.Nature Communications16\(1\),pp\. 10529\.Cited by:[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p3.1)\.
- A\. Goldstein, Z\. Zada, E\. Buchnik, M\. Schain, A\. Price, B\. Aubrey, S\. A\. Nastase, A\. Feder, D\. Emanuel, A\. Cohen, A\. Jansen, H\. Gazula, G\. Choe, A\. Rao, C\. Kim, C\. Casto, L\. Fanda, W\. Doyle, D\. Friedman, P\. Dugan, L\. Melloni, R\. Reichart, S\. Devore, A\. Flinker, U\. Hasson,et al\.\(2022\)Shared computational principles for language processing in humans and deep language models\.Nature Neuroscience25\(3\),pp\. 369–380\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1)\.
- K\. Grill\-Spector, R\. Henson, and A\. Martin \(2006\)Repetition and the brain: neural models of stimulus\-specific effects\.Trends in Cognitive Sciences10\(1\),pp\. 14–23\.Cited by:[§4\.1](https://arxiv.org/html/2607.15856#S4.SS1.p4.1)\.
- P\. Hagoort \(2013\)MUC \(memory, unification, control\) and beyond\.Frontiers in Psychology4,pp\. 416\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1),[§4\.3](https://arxiv.org/html/2607.15856#S4.SS3.p2.1)\.
- J\. Hale \(2001\)A probabilistic earley parser as a psycholinguistic model\.InProceedings of the Second Meeting of the North American Chapter of the Association for Computational Linguistics on Language Technologies,pp\. 1–8\.External Links:[Document](https://dx.doi.org/10.3115/1073336.1073357)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1),[§4\.3](https://arxiv.org/html/2607.15856#S4.SS3.p1.1)\.
- L\. S\. Hamilton and A\. G\. Huth \(2020\)The revolution will not be controlled: natural stimuli in speech neuroscience\.Language, Cognition and Neuroscience35\(5\),pp\. 573–582\.External Links:[Document](https://dx.doi.org/10.1080/23273798.2018.1499946)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p1.1)\.
- U\. Hasson, J\. Chen, and C\. J\. Honey \(2015\)Hierarchical process memory: memory as an integral component of information processing\.Trends in Cognitive Sciences19\(6\),pp\. 304–313\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p4.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p4.1)\.
- U\. Hasson, E\. Yang, I\. Vallines, D\. J\. Heeger, and N\. Rubin \(2008\)A hierarchy of temporal receptive windows in human cortex\.Journal of Neuroscience28\(10\),pp\. 2539–2550\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p4.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p4.1)\.
- J\. M\. Henderson, W\. Choi, M\. W\. Lowder, and Ferreira,Fernanda \(2016\)Language structure in the brain: a fixation\-related fmri study of syntactic surprisal in reading\.NeuroImage132,pp\. 293–300\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1)\.
- A\. G\. Huth, W\. A\. de Heer, T\. L\. Griffiths, F\. E\. Theunissen, and J\. L\. Gallant \(2016\)Natural speech reveals the semantic maps that tile human cerebral cortex\.Nature532,pp\. 453–458\.External Links:[Document](https://dx.doi.org/10.1038/nature17637)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p1.1),[§1](https://arxiv.org/html/2607.15856#S1.p3.1),[§3\.5](https://arxiv.org/html/2607.15856#S3.SS5.p2.2),[§4\.1](https://arxiv.org/html/2607.15856#S4.SS1.p2.1),[§4\.2](https://arxiv.org/html/2607.15856#S4.SS2.p2.1),[§4\.5](https://arxiv.org/html/2607.15856#S4.SS5.p1.1)\.
- S\. Jain and A\. G\. Huth \(2018\)Incorporating context into language encoding models for fmri\.InAdvances in Neural Information Processing Systems,Vol\.31\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1),[§4\.1](https://arxiv.org/html/2607.15856#S4.SS1.p2.1),[§4\.2](https://arxiv.org/html/2607.15856#S4.SS2.p2.1),[§4\.5](https://arxiv.org/html/2607.15856#S4.SS5.p1.1)\.
- C\. Kauf, G\. Tuckute, R\. Levy, J\. Andreas, and E\. Fedorenko \(2024\)Lexical\-semantic content, not syntactic structure, is the main contributor to ann\-brain similarity of fmri responses in the language network\.Neurobiology of Language5\(1\),pp\. 7–42\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1)\.
- S\. Kumar, T\. R\. Sumers, T\. Yamakoshi, A\. Goldstein, U\. Hasson, K\. A\. Norman, T\. L\. Griffiths, R\. D\. Hawkins, and S\. A\. Nastase \(2024\)Shared functional specialization in transformer\-based language models and the human brain\.Nature Communications15\(1\),pp\. 5523\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1)\.
- G\. R\. Kuperberg and T\. F\. Jaeger \(2016\)What do we mean by prediction in language comprehension?\.Language, Cognition and Neuroscience31\(1\),pp\. 32–59\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1),[§4\.3](https://arxiv.org/html/2607.15856#S4.SS3.p1.1)\.
- M\. Kutas and K\. D\. Federmeier \(2011\)Thirty years and counting: finding meaning in the n400 component of the event\-related brain potential \(erp\)\.Annual Review of Psychology62,pp\. 621–647\.External Links:[Document](https://dx.doi.org/10.1146/annurev.psych.093008.131123)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1),[§4\.1](https://arxiv.org/html/2607.15856#S4.SS1.p3.1),[§4\.2](https://arxiv.org/html/2607.15856#S4.SS2.p1.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p1.1)\.
- M\. Kutas and S\. A\. Hillyard \(1984\)Brain potentials during reading reflect word expectancy and semantic association\.Nature307,pp\. 161–163\.External Links:[Document](https://dx.doi.org/10.1038/307161a0)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1),[§4\.1](https://arxiv.org/html/2607.15856#S4.SS1.p3.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p1.1)\.
- M\. A\. Lambon Ralph, E\. Jefferies, K\. Patterson, and T\. T\. Rogers \(2017\)The neural and computational bases of semantic cognition\.Nature Reviews Neuroscience18\(1\),pp\. 42–55\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1),[§4\.5](https://arxiv.org/html/2607.15856#S4.SS5.p2.1)\.
- E\. F\. Lau, C\. Phillips, and D\. Poeppel \(2008\)A cortical network for semantics: \(de\)constructing the n400\.Nature Reviews Neuroscience9\(12\),pp\. 920–933\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1),[§1](https://arxiv.org/html/2607.15856#S1.p3.1),[§4\.3](https://arxiv.org/html/2607.15856#S4.SS3.p2.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p1.1)\.
- A\. LeBel, L\. Wagner, S\. Jain, A\. Adhikari\-Desai, B\. Gupta, A\. Morgenthal, J\. Tang, L\. Xu, and A\. G\. Huth \(2023\)A natural language fmri dataset for voxelwise encoding models\.Scientific Data10,pp\. 555\.External Links:[Document](https://dx.doi.org/10.1038/s41597-023-02437-z)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p6.1),[§2\.1](https://arxiv.org/html/2607.15856#S2.SS1.SSS0.Px2.p1.1),[§2\.4](https://arxiv.org/html/2607.15856#S2.SS4.p2.1),[§4\.2](https://arxiv.org/html/2607.15856#S4.SS2.p4.1)\.
- Y\. Lerner, C\. J\. Honey, L\. J\. Silbert, and U\. Hasson \(2011\)Topographic mapping of a hierarchy of temporal receptive windows using a narrated story\.Journal of Neuroscience31\(8\),pp\. 2906–2915\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p4.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p4.1)\.
- R\. Levy \(2008\)Expectation\-based syntactic comprehension\.Cognition106\(3\),pp\. 1126–1177\.External Links:[Document](https://dx.doi.org/10.1016/j.cognition.2007.05.006)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1),[§4\.3](https://arxiv.org/html/2607.15856#S4.SS3.p1.1)\.
- R\. L\. Lewis and S\. Vasishth \(2005\)An activation\-based model of sentence processing as skilled memory retrieval\.Cognitive Science29\(3\),pp\. 375–419\.External Links:[Document](https://dx.doi.org/10.1207/s15516709cog0000%5F25)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p8.1)\.
- N\. K\. Logothetis \(2003\)The underpinnings of the bold functional magnetic resonance imaging signal\.Journal of Neuroscience23\(10\),pp\. 3963–3971\.External Links:[Document](https://dx.doi.org/10.1523/JNEUROSCI.23-10-03963.2003)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p4.1),[§4\.1](https://arxiv.org/html/2607.15856#S4.SS1.p2.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p3.1)\.
- K\. Oberauer and S\. Lewandowsky \(2011\)Modeling working memory: a computational implementation of the time\-based resource\-sharing theory\.Psychonomic Bulletin & Review18\(1\),pp\. 10–45\.External Links:[Document](https://dx.doi.org/10.3758/s13423-010-0020-6)Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p8.1)\.
- K\. Oberauer \(2002\)Access to information in working memory: exploring the focus of attention\.Journal of Experimental Psychology: Learning, Memory, and Cognition28\(3\),pp\. 411–421\.External Links:[Document](https://dx.doi.org/10.1037/0278-7393.28.3.411)Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p8.1)\.
- M\. J\. Pickering and C\. Gambi \(2018\)Predicting while comprehending language: a theory and review\.Psychological Bulletin144\(10\),pp\. 1002–1044\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1),[§4\.3](https://arxiv.org/html/2607.15856#S4.SS3.p1.1)\.
- K\. Rayner, S\. C\. Sereno, and G\. E\. Raney \(1996\)Eye movement control in reading: a comparison of two types of models\.Journal of Experimental Psychology: Human Perception and Performance22\(5\),pp\. 1188–1200\.Cited by:[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p1.1)\.
- A\. G\. Russo, M\. De Martino, A\. Mancuso, G\. Iaconetta, R\. Manara, A\. Elia, A\. Laudanna, F\. Di Salle, and F\. Esposito \(2020\)Semantics\-weighted lexical surprisal modeling of naturalistic functional mri time\-series during spoken narrative listening\.NeuroImage222,pp\. 117281\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1)\.
- J\. Sassenhagen and C\. J\. Fiebach \(2020\)Traces of meaning itself: encoding distributional word vectors in brain activity\.Neurobiology of Language1\(1\),pp\. 54–76\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1)\.
- S\. Schuster, S\. Hawelka, F\. Hutzler, M\. Kronbichler, and F\. Richlan \(2016\)Words in context: the effects of length, frequency, and predictability on brain responses during natural reading\.Cerebral Cortex26\(10\),pp\. 3889–3904\.Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px3.p1.1)\.
- M\. L\. Seghier \(2013\)The angular gyrus: multiple functions and multiple subdivisions\.The Neuroscientist19\(1\),pp\. 43–61\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1),[§4\.5](https://arxiv.org/html/2607.15856#S4.SS5.p2.1)\.
- C\. Shain, I\. A\. Blank, M\. van Schijndel, W\. Schuler, and E\. Fedorenko \(2020\)fMRI reveals language\-specific predictive coding during naturalistic sentence comprehension\.Neuropsychologia138,pp\. 107307\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1)\.
- N\. J\. Smith and R\. Levy \(2013\)The effect of word predictability on reading time is logarithmic\.Cognition128\(3\),pp\. 302–319\.Cited by:[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p1.1)\.
- M\. Song, J\. Wang, and Q\. Cai \(2024\)The unique contribution of uncertainty reduction during naturalistic language comprehension\.Cortex181,pp\. 12–25\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1)\.
- K\. Specht \(2014\)Neuronal basis of speech comprehension\.Hearing Research307,pp\. 121–135\.Cited by:[§3\.1](https://arxiv.org/html/2607.15856#S3.SS1.p1.1)\.
- K\. Sun and H\. Liu \(2024\)Attention\-aware semantic relevance predicting chinese sentence reading\.Cognition\.Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1),[§4\.6](https://arxiv.org/html/2607.15856#S4.SS6.p4.1)\.
- K\. Sun, Q\. Wang, and X\. Lu \(2023\)An interpretable measure of semantic similarity for predicting eye movements in reading\.Psychonomic Bulletin & Review,pp\. 1–16\.Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1)\.
- K\. Sun, R\. Wang, and H\. Baayen \(2026\)Semantic coherence predicts reading fixation durations across languages beyond surprisal and lexical factors\.Linguistics\.Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p1.1),[§4\.6](https://arxiv.org/html/2607.15856#S4.SS6.p4.1)\.
- K\. Sun and R\. Wang \(2022\)Semantic similarity and mutual information predicting sentence comprehension: the case of dangling topic construction in chinese\.Journal of Cognitive Psychology,pp\. 1–24\.Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1)\.
- K\. Sun and R\. Wang \(2026a\)Contextual semantic relevance predicting human visual attention\.IEEE Transactions on Cognitive and Developmental Systems\.External Links:[Document](https://dx.doi.org/10.1109/TCDS.2026.3683873)Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1),[§4\.6](https://arxiv.org/html/2607.15856#S4.SS6.p4.1)\.
- K\. Sun and R\. Wang \(2026b\)Non\-linear effects of semantic relevance on word duration in spontaneous speech\.InterSpeech\.Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1),[§4\.6](https://arxiv.org/html/2607.15856#S4.SS6.p4.1)\.
- K\. Sun and R\. Wang \(2026c\)Semantic integration and lexical expectation shape n400 and p600 dynamics during naturalistic reading\.arXiv\.Note:arXiv:2607\.04107Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p1.1),[§4\.4](https://arxiv.org/html/2607.15856#S4.SS4.p2.1),[§4\.6](https://arxiv.org/html/2607.15856#S4.SS6.p4.1)\.
- R\. Tikochinski, A\. Goldstein, Y\. Meiri, U\. Hasson, and R\. Reichart \(2025\)Incremental accumulation of linguistic context in artificial and biological neural networks\.Nature Communications16\(1\),pp\. 803\.Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1)\.
- G\. Tuckute, A\. Sathe, S\. Srikant, M\. Taliaferro, M\. Wang, M\. Schrimpf, K\. Kay, and E\. Fedorenko \(2024\)Driving and suppressing the human language network using large language models\.Nature Human Behaviour8\(3\),pp\. 544–561\.Cited by:[§4\.6](https://arxiv.org/html/2607.15856#S4.SS6.p2.1)\.
- S\. Tulkens and T\. van Dongen \(2024\)Model2Vec: fast state\-of\-the\-art static embeddings\.Zenodo\.External Links:[Document](https://dx.doi.org/10.5281/zenodo.17270888),[Link](https://github.com/MinishLab/model2vec)Cited by:[§2\.2](https://arxiv.org/html/2607.15856#S2.SS2.SSS0.Px2.p2.4)\.
- L\. Wehbe, B\. Murphy, P\. Talukdar, A\. Fyshe, A\. Ramdas, and T\. Mitchell \(2014\)Simultaneously uncovering the patterns of brain regions involved in different story reading subprocesses\.PLoS ONE9\(11\),pp\. e112575\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0112575)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p3.1),[§4\.1](https://arxiv.org/html/2607.15856#S4.SS1.p2.1),[§4\.5](https://arxiv.org/html/2607.15856#S4.SS5.p1.1)\.
- R\. M\. Willems, S\. L\. Frank, A\. D\. Nijhof, P\. Hagoort, and A\. van den Bosch \(2016\)Prediction during natural language comprehension\.Cerebral Cortex26\(6\),pp\. 2506–2516\.External Links:[Document](https://dx.doi.org/10.1093/cercor/bhv075)Cited by:[§1](https://arxiv.org/html/2607.15856#S1.p2.1),[§4\.2](https://arxiv.org/html/2607.15856#S4.SS2.p1.1)\.
- F\. Zhou, S\. Zhou, Y\. Long, A\. Flinker, and C\. Lu \(2026\)Hierarchical linguistic predictions and cross\-level information updating during narrative comprehension\.Communications Biology9\(107\)\.Cited by:[§4\.3](https://arxiv.org/html/2607.15856#S4.SS3.p2.1)\.
Appendix
## Appendix AROI Definitions and Correlations for Predictors
Table 1:Brain regions included in the Alice and Moth ROI analyses and their commonly associated functions\. The functional descriptions are intentionally broad because each region contributes to multiple cognitive processes\.ROIFull nameCommonly associated functionsROIs included in both the Alice and Moth analysesAGAngular gyrusSemantic and conceptual integration, discourse comprehension, and multimodal information processing\.CGCingulate gyrusAttention, cognitive control, conflict monitoring, and internally directed processing\.FOCFrontal opercular cortexSpeech production, articulation, phonological processing, and language\-related cognitive control\.HGHeschl’s gyrusPrimary auditory processing and early analysis of acoustic information\.ICInsular cortexSalience detection, auditory–motor integration, speech processing, and cognitive control\.IFGInferior frontal gyrusSyntactic processing, semantic selection, language control, and speech production\.MTGMiddle temporal gyrusLexical–semantic processing, conceptual representation, and sentence comprehension\.PTPlanum temporaleHigher\-order auditory analysis, phonological processing, and speech perception\.SGSuperior gyrusHigher\-order sensory, auditory, or language\-related processing, depending on the anatomical definition\.SMGSupramarginal gyrusPhonological processing, verbal working memory, and multimodal integration\.STGSuperior temporal gyrusSpeech perception, auditory\-language comprehension, and semantic processing\.vmPFCVentromedial prefrontal cortexSchema\-based interpretation, valuation, contextual integration, and internally guided cognition\.Additional ROIs included in the Moth analysisACCAnterior cingulate cortexConflict monitoring, attentional control, salience processing, and response regulation\.AIAnterior insulaSalience detection, interoceptive processing, and coordination of cognitive control\.COCCentral opercular cortexSensorimotor integration, somatosensory processing, and speech\-related motor functions\.DLPFCDorsolateral prefrontal cortexWorking memory, executive control, planning, and goal\-directed processing\.FUSFusiform gyrusVisual\-form processing, object recognition, and higher\-level conceptual representation\.HPCHippocampusEpisodic memory, contextual association, and formation of event representations\.ITGInferior temporal gyrusObject recognition, lexical\-semantic representation, and conceptual processing\.LINGLingual gyrusVisual processing, visual imagery, and recognition of complex visual patterns\.LOCLateral occipital cortexVisual object recognition and processing of shape and form\.OFGOrbitofrontal gyrusValuation, expectation, decision\-making, and integration of contextual information\.PCGPrecentral gyrusMotor planning, speech articulation, and preparation of voluntary movements\.PCUNPrecuneusEpisodic imagery, internally directed cognition, and construction of situation models\.PHGParahippocampal gyrusContextual memory, scene processing, and associations between events and environments\.PPPosterior parietal regionAttention, working memory, spatial processing, and multimodal integration\.preSMAPre\-supplementary motor areaSequencing, response selection, cognitive control, and speech\-motor planning\.RSCRetrosplenial cortexContextual memory, scene representation, and integration of narrative or spatial information\.SFGSuperior frontal gyrusExecutive control, working memory, attention, and internally directed cognition\.TPTemporal poleSemantic, social\-conceptual, emotional, and narrative\-level integration\.
Note\.Functional descriptions summarize commonly reported associations and should not be interpreted as exclusive functions\. The precise definitions of SG and PP should be verified against the anatomical atlas used in the present analyses\.
As shown in Figure[7](https://arxiv.org/html/2607.15856#A1.F7), surprisal is lowly correlated with semantic relevance in both datasets\. Specifically,rr= \-0\.18 Alice,rr= \-0\.01 Moth, and these indicate that surprisal and semantic relevance are distinct metrics used in the present study\.
Figure 7:Correlations for predictors in the two datasets
## Appendix BRobustness Analysis: Temporal Smoothness of Semantic Relevance
One potential concern is that the semantic\-relevance predictor may appear stronger than surprisal partly because of its temporal structure\. Since the BOLD signal is itself a slow, low\-pass physiological response, this raises an important alternative explanation: semantic relevance may predict BOLD more strongly not because it captures semantic integration, but because its temporal spectrum is better matched to BOLD\.
To evaluate this possibility, we conducted three supplementary analyses\. First, we compared the temporal autocorrelation and spectral properties of semantic relevance and surprisal\. Second, we smoothed surprisal to approximately match the lag\-1 autocorrelation of semantic relevance and then reran the HRF\-weighted 4–12 s FIR/deconvolution\-style analysis\. Third, we constructed time\-structure\-preserving null predictors for semantic relevance using circular shifts, phase randomization, and block permutation within story/run structure\.
### B\.1Temporal autocorrelation and spectral properties
For each predictor, we computed the lag\-1 autocorrelation, an autocorrelation\-based effective temporal sample size, the proportion of power in the lowest 10% of the frequency range, and the spectral centroid\. The effective temporal sample size was computed as
Neff=N1−r11\+r1,N\_\{\\mathrm\{eff\}\}=N\\frac\{1\-r\_\{1\}\}\{1\+r\_\{1\}\},\(7\)wherer1r\_\{1\}is the lag\-1 autocorrelation\. This value provides a simple index of the effective temporal degrees of freedom after accounting for first\-order autocorrelation\.
Table 2:Temporal smoothness diagnostics for semantic relevance and surprisal\. In the Moth dataset, semantic relevance refers to semantic relevance\. Smoothed surprisal was exponentially smoothed to approximately match the lag\-1 autocorrelation of semantic relevance\.DatasetPredictorLag\-1 ACFNeffN\_\{\\mathrm\{eff\}\}Low\-frequency powerSpectral centroidAliceSemantic\-similarity metric0\.485740\.30\.2750\.149AliceSurprisal0\.0312001\.90\.1450\.241AliceSmoothed surprisal0\.486737\.50\.3240\.145MothSemantic relevance0\.337278\.10\.2070\.181MothSurprisal0\.099434\.90\.1660\.228MothSmoothed surprisal0\.337278\.30\.2540\.178
These diagnostics confirm that semantic relevance is more temporally autocorrelated than surprisal and has relatively stronger low\-frequency power\. Thus, the smoothness concern is empirically valid and should not be dismissed\.
### B\.2Frequency\-matched surprisal control
We next tested whether smoothing surprisal to match semantic relevance would make surprisal perform similarly\. Surprisal was exponentially smoothed to approximate the lag\-1 autocorrelation of semantic relevance\. We then reran an HRF\-weighted 4–12 s linear FIR approximation that controlled for the competing predictor, lagged lexical covariates, a polynomial time trend, and run fixed effects\.
Table 3:Comparison between semantic relevance and autocorrelation\-matched smoothed surprisal in the Moth HRF\-weighted 4–12 s analysis\.PredictorROIsFDR\-significantROIsMedian\|t\|\|t\|Mean\|t\|\|t\|Semantic relevance30212\.3772\.533Smoothed surprisal3000\.9270\.993The smoothed\-surprisal control did not reproduce the semantic\-relevance effect\. Although smoothed surprisal was made similar to semantic relevance in lag\-1 autocorrelation, it was not significant in any ROI after FDR correction\. This result argues against the strongest version of the smoothness\-only explanation\.
### B\.3Time\-structure\-preserving null predictors
Finally, we replaced semantic relevance with surrogate predictors that preserved aspects of temporal structure while disrupting the original word\-to\-BOLD alignment\. We used three null procedures: circular shift, phase randomization, and block permutation\. Each procedure was applied within story structure rather than by fully random global permutation\. Each null method used 40 surrogate repetitions per ROI\. This number of repetitions gives only coarse empiricalppvalues, with a minimum possible value of1/\(40\+1\)=\.0241/\(40\+1\)=\.024\. Thus, these tests should be interpreted as sensitivity analyses rather than definitive permutation tests\.
Table 4:Time\-structure\-preserving surrogate tests for semantic relevance in the Moth dataset\. The familywise statistic compares the observed maximum absolutettvalue across ROIs with the surrogate distribution of maximum absolutettvalues\.Null methodObserved max\|t\|\|t\|Mean null max\|t\|\|t\|FamilywiseppObserved FDR ROIsBlock permutation4\.1224\.8960\.80521Circular shift4\.1224\.6460\.65921Phase randomization4\.1224\.6390\.68321Table 5:ROI\-wise empirical surrogate results for semantic relevance in the Moth dataset\. The table reports ROIs with same\-direction empiricalp<\.05p<\.05, where the observed effect was negative\. Empiricalppvalues are coarse because each null procedure used 40 repetitions\. Two\-sided absolute\-value tests were more conservative, with two ROIs reachingp<\.05p<\.05for each null method\.Null methodROIs withsame\-direction empiricalp<\.05p<\.05Number of ROIsBlock permutationACC \(\.024\), PCUN \(\.024\), vmPFC \(\.024\), RSC \(\.049\)4Circular shiftRSC \(\.024\), ACC \(\.049\), FUS \(\.049\), CG \(\.049\), SFG \(\.049\), preSMA \(\.049\)6Phase randomizationCG \(\.024\), ACC \(\.049\), FUS \(\.049\), LING \(\.049\)4
The surrogate analyses produced an important limitation on interpretation\. The observed semantic\-relevance effects were not reproduced by frequency\-matched surprisal, but the strict familywise max\-\|t\|\|t\|surrogate tests did not reject the time\-structure null\. Moreover, the mean maximum absolutettvalue under the surrogate distributions was larger than the observed maximum absolutettvalue for all three null methods\. This means that the familywise surrogate results cannot be treated merely as an underpowered positive trend; instead, they indicate that temporally structured surrogate predictors can sometimes produce ROI\-level maxima as large as, or larger than, the observed effect\. ROI\-wise empirical tests provided limited local support in a small subset of ROIs, but these effects were not broad enough to overturn the familywise null result\.
In sum, these analyses support a narrower and more cautious interpretation\. Semantic relevance is smoother and more low\-frequency than surprisal, but the semantic\-relevance advantage is not explained by simple lag\-1 smoothing because frequency\-matched surprisal did not yield comparable effects\. However, the time\-structure\-preserving surrogate tests do not establish that the observed effect is uniquely semantic rather than partly driven by temporally structured discourse\-level variation\. We therefore treat this analysis as a robustness boundary\. It weakens a pure smoothness\-only explanation, but it does not provide a confirmatory semantic\-specificity test\. A stronger future test should use substantially larger surrogate distributions, preregistered ROI\-wise and familywise statistics, cross\-validated encoding models, and independently fixed semantic\-relevance metrics to further separate semantic content from temporal\-spectrum matching\.
## Appendix COther Results
The Moth HRF\-convolved GAMM analysis showed significant partial effects of surprisal in five ROIs when semantic relevance was modeled as semantic relevance in the same candidate\-pair specification\. Across these regions, the fitted smooths showed a broadly similar near\-linear positive pattern: transformed BOLD was lower at low surprisal values and higher at moderate\-to\-high surprisal values\. Thus, surprisal showed limited but detectable associations with transformed BOLD in this GAMM framework, although this pattern contrasts with the FIR/deconvolution analyses of original BOLD, where surprisal did not show robust HRF\-weighted effects\. The partial effect of surprisal in the Moth dataset is illustrated in Figure[8](https://arxiv.org/html/2607.15856#A3.F8)\.
Figure 8:Partial effects of surprisal on transformed BOLD responses in the Moth dataset\. Panels show the 5 ROIs from the GAMM analysis\. For each ROI, thex\-axis shows surprisal, and they\-axis shows the estimated partial effect on transformed BOLD after controlling for word frequency, word length, semantic relevance, and subject\-level random effects\. The GAMM formula wasbold∼s\(log\_freq,k=5\)\+s\(wordlen,k=5\)\+s\(SemRel,k=5\)\+s\(Surprisal\_GPT2,k=5\)\+s\(subj,bs=‘‘re′′\)\\mathrm\{bold\}\\sim s\(\\mathrm\{log\\\_freq\},k=5\)\+s\(\\mathrm\{wordlen\},k=5\)\+s\(\\mathrm\{SemRel\},k=5\)\+s\(\\mathrm\{Surprisal\\\_GPT2\},k=5\)\+s\(\\mathrm\{subj\},bs=\\mathrm\{\`\`re^\{\\prime\\prime\}\}\)\. Panels show the five ROIs in which the surprisal smooth term remained significant after FDR correction in the candidate\-pair GAMM specification \(preSMA, SFG, PCUN, RSC, and vmPFC\)\. Purple curves show the fitted partial effect of surprisal on transformed BOLD, and lavender shaded bands indicate approximately 95% confidence intervals\. These results indicate that surprisal showed limited but detectable transformed\-BOLD effects in a small subset of ROIs when semantic relevance was included in the same model\.Further, the temporal\-smoothness analysis showed that semantic relevance and surprisal differ in their short\-range temporal structure\. In a representative 60\-s Moth segment, semantic relevance, measured as semantic relevance, varied more smoothly across adjacent words than surprisal, whereas surprisal showed sharper word\-to\-word fluctuations\. This pattern was confirmed by the autocorrelation analysis across Moth stories: semantic relevance had a higher mean lag\-1 autocorrelation than surprisal \(0\.337 vs\. 0\.099\), although the difference became small after the first few word lags\. Thus, semantic relevance is indeed more temporally autocorrelated than surprisal, supporting the need for smoothness\-control analyses, but the difference is concentrated mainly at short lags rather than reflecting a uniformly slow predictor\. The result is summarized in Figure[9](https://arxiv.org/html/2607.15856#A3.F9)\.
Figure 9:Temporal smoothness of semantic relevance and surprisal in the Moth dataset\. \(a\) Representative 60\-s story segment showing z\-scored semantic relevance and surprisal values\. Faint lines show word\-level values, and thick lines show locally smoothed trajectories for visualization\. As exact word onset times were not available for all Moth stories, thex\-axis uses approximate story time based on the word\-to\-story alignment used in the FIR analysis\. \(b\) Autocorrelation functions averaged across Moth stories\. Semantic relevance showed higher lag\-1 autocorrelation than surprisal \(0\.337 vs\. 0\.099\), indicating greater short\-range temporal smoothness\. The difference decreased rapidly after the first few word lags, suggesting that semantic relevance is smoother than surprisal but not simply a uniformly low\-frequency predictor\.Additionally, across the six typical ROIs shared by the Alice and Moth analyses, the response profiles showed a consistent separation between semantic relevance and surprisal, as shown in Figure[10](https://arxiv.org/html/2607.15856#A3.F10)\. In Alice, semantic relevance produced sustained negative FIR estimates that became most apparent in the delayed hemodynamic range, whereas the surprisal curves fluctuated weakly around zero and did not show a systematic delayed profile\. In Moth, the HRF\-weighted profiles showed the same qualitative pattern across the matched ROIs: semantic relevance yielded negative effects in the expected 4\-12 s window, while surprisal remained small and close to baseline\. Thus, the figure suggests that the semantic relevance effect is not driven by a single selected region, but appears across several canonical temporal, frontal, and temporoparietal language\-related ROIs\.
Figure 10:Common\-ROI response profiles for semantic relevance and surprisal in the Alice and Moth datasets\. Each row shows one ROI that was available in both datasets \(STG, MTG, IFG, AG, PT, and HG\)\. Left panels show FIR lag\-response curves from the Alice dataset\. Thex\-axis indicates time after word onset from 0 to 16 s, and they\-axis shows FIR lag coefficient estimates\. Shaded bands indicate approximately 95% confidence intervals\. The light\-blue background marks the delayed hemodynamic window\. Right panels show Moth HRF\-weighted 4–12 s response profiles\. As the Moth main analysis used an HRF\-weighted directional FIR/deconvolution test, the right panels show scaled HRF\-weighted profiles rather than unconstrained FIR coefficients\. Across the common ROIs, semantic relevance showed consistent delayed negative effects, whereas surprisal effects remained close to zero\. These profiles support the interpretation that semantic relevance, more than surprisal, is associated with delayed BOLD responses in naturalistic speech comprehension\.
## Appendix DSupplementary Descriptive Connectivity Analyses
To provide additional anatomical and network\-level context for the ROI analyses, we visualized pairwise BOLD correlations among selected Moth ROIs during narrative listening\. These analyses were descriptive and were not used as inferential tests of semantic relevance or surprisal effects\. Instead, they illustrate how the analyzed ROIs were embedded in broader language\- and memory\-related functional networks during continuous spoken\-language comprehension\.
The language\-network visualization \(Figure[11](https://arxiv.org/html/2607.15856#A4.F11)\) showed strong functional coupling among temporal, inferior frontal, angular, and auditory\-related regions\. This pattern is consistent with the idea that naturalistic speech comprehension recruits a distributed language system rather than isolated cortical sites\. Importantly, the connectivity pattern should not be interpreted as evidence that semantic relevance specifically modulates these connections; rather, it provides a network\-level background for interpreting the ROI\-based GAMM and FIR/deconvolution results\.
Figure 11:Descriptive language\-network connectivity in the Moth dataset\. Nodes represent selected language\-related ROIs, and edges represent pairwise BOLD correlations above the visualization threshold \(r\>\.85r\>\.85\)\. Edge color indicates correlation strength\. This figure provides network\-level context for the ROI\-based analyses, showing that temporal, inferior frontal, angular, and auditory\-related regions were strongly coupled during naturalistic narrative listening\. The figure is descriptive and does not test semantic relevance or surprisal effects directly\.As illustrated in Figure[12](https://arxiv.org/html/2607.15856#A4.F12), the memory\-network visualization provides complementary context for the semantic relevance findings\. Because semantic relevance indexes the fit between an incoming word and the preceding semantic context, its effect may depend not only on lexical\-semantic processing but also on discourse\-level integration, contextual updating, and situation\-model construction\. The strong correlations among memory\- and integration\-related ROIs therefore support the plausibility of interpreting semantic relevance as a sustained contextual integration signal\.
Figure 12:Descriptive memory\- and contextual\-integration\-network connectivity in the Moth dataset\. Nodes represent selected ROIs associated with memory, contextual integration, and discourse\-level processing, and edges represent pairwise BOLD correlations above the visualization threshold \(r\>\.85r\>\.85\)\. The strong coupling among these regions provides anatomical context for interpreting semantic relevance as a measure related to contextual semantic fit and discourse integration\. This visualization is descriptive and should not be interpreted as a direct statistical test of semantic relevance\.Finally, the ROI correlation heatmap summarizes the pairwise correlation structure among a subset of analyzed regions, as shown in Figure[13](https://arxiv.org/html/2607.15856#A4.F13)\. The generally high correlations indicate that many ROIs shared substantial low\-frequency BOLD dynamics during narrative listening\.
Figure 13:Pairwise ROI correlation heatmap for selected Moth regions during narrative listening\. Cell values indicate Pearson correlations between ROI\-level BOLD time series\. The high correlations suggest that the analyzed regions shared substantial task\-related and low\-frequency BOLD dynamics during naturalistic speech comprehension\. This descriptive result motivates caution in interpreting ROI effects as independent regional activations and supports treating the main GAMM and FIR/deconvolution analyses as predictor\-response tests embedded within a strongly coupled language\-comprehension network\.Similar Articles
Heterogeneous Neural Predictivity from Language Models During Naturalistic Comprehension
This paper investigates how language model representations predict neural activity during naturalistic language comprehension across MEG, ECoG, and other recordings. The findings demonstrate that language model features serve as useful neural predictors, but caution against overinterpreting predictive success as evidence for shared neural organization.
Beyond Parallel Tracking: Interactive Multi-Feature Fusion Drives Semantic Reconstruction from Non-invasive Brain Recordings
This paper introduces a multi-feature fusion framework for semantic reconstruction from non-invasive brain recordings, combining static lexical (Word2Vec) and dynamic contextual (GPT) representations via cross-attention, achieving state-of-the-art performance in brain-to-text decoding.
Sentence-Level Contextual Entrainment in Large Language Models
This paper extends contextual entrainment from token-level to sentence-level, showing that even counterfactual sentences in prompts increase their probability during inference. The effect decreases with model size and is driven by 2-4% of attention heads, which can be ablated without performance loss.
Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding
This paper introduces a meta-optimized approach for semantic visual decoding from fMRI signals that generalizes to novel subjects without fine-tuning, using in-context learning to infer unique neural encoding patterns from a small set of image-brain activation examples. The method achieves strong cross-subject and cross-scanner generalization without requiring anatomical alignment or stimulus overlap.
Brain Score Tracks Shared Properties of Languages: Evidence from Many Natural Languages and Structured Sequences
This paper investigates whether Brain Score, a metric comparing language model representations to human fMRI activations during reading, is truly capturing human-like language processing or merely structural similarity. The researchers train language models on diverse natural languages and non-linguistic structured data (genome, Python, nested parentheses), finding that models trained on different languages and even non-linguistic sequences achieve similar Brain Score performance, suggesting the metric may not be sensitive enough to distinguish human-specific processing.