Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models

arXiv cs.CL Papers

Summary

This paper investigates whether language models' next-word prediction aligns with human cognitive processing by analyzing EEG signals and event-related potentials, finding that only surprisal correlates with human brain responses, especially for open-class words.

arXiv:2607.16549v1 Announce Type: new Abstract: Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neuroscience research reveals that next-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography (EEG). While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension? Here, we generate regressors for both humans and LMs based on two information measures, including top-1 prediction and surprisal, to predict event-related potential (ERP) elicited from EEG recordings which reflect different stages of cognitive processing during reading. We argue that modelling ERP patterns offers fine-grained analysis of the cognitive plausibility of various LMs during reading. Our results indicate that only surprisal potentially correlates with language-processing ERPs, especially for open-class words with high semantic content. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human-like linguistic processing.
Original Article
View Cached Full Text

Cached at: 07/21/26, 06:42 AM

# Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models
Source: [https://arxiv.org/html/2607.16549](https://arxiv.org/html/2607.16549)
Boi Mai QuachBinh T\. NguyenCathal Gurrinmai\.quach3@mail\.dcu\.iengtbinh@hcmus\.edu\.vncathal\.gurrin@dcu\.ieSchool of Computing, ML\-Labs,Department of Computer Science,School of Computing, ADAPT Centre,Dublin City University, IrelandVNUHCM – University of ScienceDublin City University, Ireland

Graham Healygraham\.healy@dcu\.ieSchool of Computing, ADAPT Centre,Dublin City University, Ireland

###### Abstract

Language models \(LMs\) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension\. Neuroscience research reveals that next\-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography \(EEG\)\. While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next\-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension? Here, we generate regressors for both humans and LMs based on two information measures, including top\-1 prediction and surprisal, to predict event\-related potential \(ERP\) elicited from EEG recordings which reflect different stages of cognitive processing during reading\. We argue that modelling ERP patterns offers fine\-grained analysis of the cognitive plausibility of various LMs during reading\. Our results indicate that only surprisal potentially correlates with language\-processing ERPs, especially for open\-class words with high semantic content\. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human\-like linguistic processing\.

Keywords:Cognitive Neuroscience; Reading Comprehension; Language Models; GPT; Electroencephalography \(EEG\)

## Introduction

Word predictability reflects a broader cognitive process in which individuals continuously integrate context to anticipate upcoming events and test those predictions against perceptual input from the utterances they hear or read\\@BBOPcitep\\@BAP\\@BBN\(Bar,[2007](https://arxiv.org/html/2607.16549#bib.bib31)\)\\@BBCP\. During reading, predictability effects are hypothesized to reflect the cognitive demands associated with probabilistic inference, whereby the brain incrementally evaluates and updates possible upcoming word interpretations\\@BBOPcitep\\@BAP\\@BBN\(Shainet al\.,[2024](https://arxiv.org/html/2607.16549#bib.bib30)\)\\@BBCP\. From an information\-theoretic perspective\\@BBOPcitep\\@BAP\\@BBN\(Shannon,[1948](https://arxiv.org/html/2607.16549#bib.bib32)\)\\@BBCP, prediction serves as a core function of a probabilistic, generative cognitive system that incrementally processes the words of an unfolding sentence\. Linguistic units convey quantifiable information, with measures such as surprisal \(the unexpectedness of a word given its prior context\)\. Thus, surprisal serves as a useful metric for quantifying word\-by\-word predictability during sentence processing\\@BBOPcitep\\@BAP\\@BBN\(Hale,[2016](https://arxiv.org/html/2607.16549#bib.bib33)\)\\@BBCP\.

Research has indicated that surprisal is a reliable predictor of neural responses during reading, particularly in relation to the N400 component\.\\@BBOPcitet\\@BAP\\@BBNMichaelov and Bergen \([2020](https://arxiv.org/html/2607.16549#bib.bib41)\)\\@BBCPfound that surprisal effectively predicts variations in N400 amplitude, a neural indicator of processing difficulty during language comprehension\.\\@BBOPcitet\\@BAP\\@BBNFranket al\.\([2013](https://arxiv.org/html/2607.16549#bib.bib42)\)\\@BBCPfurther supported these findings by analysing EEG data from participants reading identical sentences and examining four distinct ERP components\. Their results highlighted that surprisal estimates significantly predict N400 amplitude, with more surprising words eliciting larger negative N400 responses\.

Surprisal modelling from LMs has been widely used to explain neural responses during language comprehension, with early work showing that trigram\-based surprisal correlates with the N400 in reading studies\\@BBOPcitep\\@BAP\\@BBN\(Franket al\.,[2015](https://arxiv.org/html/2607.16549#bib.bib51); Willemset al\.,[2016](https://arxiv.org/html/2607.16549#bib.bib52); Armeniet al\.,[2019](https://arxiv.org/html/2607.16549#bib.bib53)\)\\@BBCP\. More recent work with transformer\-based LMs reports similar effects:\\@BBOPcitet\\@BAP\\@BBNHeilbronet al\.\([2019](https://arxiv.org/html/2607.16549#bib.bib54)\)\\@BBCPshowed that GPT\-2’s word\-by\-word unpredictability aligns with N400\-like responses during audiobook listening, while\\@BBOPcitet\\@BAP\\@BBNMichaelovet al\.\([2024](https://arxiv.org/html/2607.16549#bib.bib55)\)\\@BBCPfound that GPT\-3 surprisal best predicted N400 amplitude across models, suggesting that effects of expectancy, plausibility, and contextual semantic similarity can be explained by variations in word predictability\. Although GPT\-2 outperforms TransformerXL and XLNet in psycholinguistic prediction\\@BBOPcitep\\@BAP\\@BBN\(Haoet al\.,[2020](https://arxiv.org/html/2607.16549#bib.bib48)\)\\@BBCP, evidence that larger models yield stronger predictability–processing relationships is mixed\. Within the GPT family,\\@BBOPcitet\\@BAP\\@BBNShainet al\.\([2024](https://arxiv.org/html/2607.16549#bib.bib30)\)\\@BBCPchallenged claims that more advanced LMs should exhibit stronger logarithmic relationships between contextual predictability and processing difficulty\. They found that surprisal estimates from GPT\-3 were not more “super\-logarithmic” than those from smaller models like GPT\-2, despite GPT\-3’s greater size and computational power\.

Despite major advances in Natural Language Processing \(NLP\), LMs still lack an interpretable, mechanistic account aligned with how the human brain processes language\. This raises debate about whether LMs capture aspects of human intelligence or simply produce outputs that mimic human thought\\@BBOPcitep\\@BAP\\@BBN\(Mitchell and Krakauer,[2023](https://arxiv.org/html/2607.16549#bib.bib1)\)\\@BBCP\. Next\-word predictability is a fundamental aspect of human language processing, which importantly supports the cognitive plausibility of LMs\\@BBOPcitep\\@BAP\\@BBN\(Keller,[2010](https://arxiv.org/html/2607.16549#bib.bib2)\)\\@BBCP\. When it comes to thought, we should examine brain activity\. During language comprehension, the human brain exhibits systematic patterns of neural activity that reflect ongoing predictive processing\\@BBOPcitep\\@BAP\\@BBN\(Fitz and Chang,[2019](https://arxiv.org/html/2607.16549#bib.bib3)\)\\@BBCP\. Therefore, rather than examining next\-word prediction performance across various LMs, we investigate the relationship between next\-word predictability and neural responses in natural reading contexts, especially in longer narratives\.

To investigate this, we run multiple experiments\. First, we calculate top\-1 prediction and lexical surprisal at the word level for content and function words across three predictors: human subjects, n\-gram models, and GPT\-family models, using the DERCO dataset, a language resource combining EEG and next\-word prediction data\\@BBOPcitep\\@BAP\\@BBN\(Quachet al\.,[2024](https://arxiv.org/html/2607.16549#bib.bib59)\)\\@BBCP\. Next, we encode neural responses using regression\-based deconvolution to estimate predictability effects on neural activity\. We then compare the correlations between neural response predictions derived from top\-1 prediction and surprisal estimates of language models and those obtained from human cloze responses\. The purpose of this comparison is to identify which model most closely mirrors human\-like predictability in reading behaviour\. To provide deeper insights, these correlations will be visualised within significant time windows and across significant electrode clusters\.

## Methodology

### Stimuli and EEG Data Preparation

We utilised the DERCo dataset\\@BBOPcitep\\@BAP\\@BBN\(Quachet al\.,[2024](https://arxiv.org/html/2607.16549#bib.bib59)\)\\@BBCP, which contains EEG recordings from 22 native English speakers while they were reading The Grimm Brothers’ Fairy Tales\. Two participants \(“QPF42” and “USQ95”\) were excluded due to excessive eye movements\. Additionally, word\-by\-word cloze probabilities were collected through a cloze procedure on Mechanical Turk crowdsourcing platform\.

EEG data were recorded using a 32\-channel electrode scalp following the international 10–20 system\\@BBOPcitep\\@BAP\\@BBN\(Klem,[1999](https://arxiv.org/html/2607.16549#bib.bib87)\)\\@BBCP\. Since the analysis used preprocessed data, the number of word\-level EEG trials in the DERCo dataset’s transcript was reduced due to artifact removal\. All remaining words, after preprocessing, served as stimuli for encoding brain signals\. The Python library IPA was used to extract parts of speech, which were then grouped into content and function words\.

### Information\-theoretic measures

To investigate next\-word predictability, we used two measures: top\-1 prediction and surprisal\. These measures serve as proxies for human and computational models’ expectations and processing effort in reading comprehension, capturing different aspects of cognitive load associated with prediction\.

#### Top\-1 Prediction Estimation

Most LMs aim to estimate a probability distribution over the vocabularyVV, for the likely next\-wordwi∈Vw\_\{i\}\\in Vat positionii, given the contextw1,w2​…​wi−1w\_\{1\},w\_\{2\}\\ldots w\_\{i\-1\}containing the sequence of preceding words in text\. The highest probability, also known as top\-1 prediction, for the next token, is defined as follows:

Pwi=maxwi∈V⁡P​\(wi∣w1,w2,…,wi−1\)P\_\{w\_\{i\}\}=\\max\_\{w\_\{i\}\\in V\}P\(w\_\{i\}\\mid w\_\{1\},w\_\{2\},\\ldots,w\_\{i\-1\}\)\(1\)

#### Surprisal Estimation

Surprisal is a measure of the unexpectedness of a target word\.\\@BBOPcitet\\@BAP\\@BBNHale \([2001](https://arxiv.org/html/2607.16549#bib.bib34)\)\\@BBCPand\\@BBOPcitet\\@BAP\\@BBNLevy \([2008](https://arxiv.org/html/2607.16549#bib.bib35)\)\\@BBCPargued that the less expected a word is in a given context, the higher its surprisal\. For example, “Peter won the championship\. Afterward, he was in seventh …”\. If readers recognise the idiom, they can guess that the missing word is “heaven\.” Since the word is highly predictable, it has low surprisal and conveys minimal new information\.

After the first t words of the sentence,w1​…​tw\_\{1\.\.\.t\}, will be processed, the identity of the upcoming word,wt\+1w\_\{t\+1\}, is still unknown and can therefore be viewed as a random variable\. The surprisal is defined as the negative log probability of the actual next word, given its preceding context:

Surprisal​Swt\+1=−l​o​g​P​\(wt\+1\|w1​…​t\)\\textit\{Surprisal \}S\_\{w\_\{t\+1\}\}=\-logP\(w\_\{t\+1\}\|w\_\{1\.\.\.t\}\)\(2\)

### Predictors

#### Human Prediction

Top\-1 prediction refers to the highest percentage of participants who guessed the same next word\. A top\-1 prediction value of 100% indicates that all participants guessed the same next word, whereas 0% indicates that no participant predicted that upcoming word\. This can be simply defined as the maximum cloze probability among the possible words that could appear in the upcoming position\.

Lexical surprisal, by contrast, is the cloze probability of the correct next word in the transcript\. It is important to note that the cloze value of the correct next word does not necessarily equal the top\-1 prediction value\. These values are equal only if the word is exceptionally easy to predict, meaning that the word predicted by all participants is also the correct word in the transcript\.

A major issue with cloze procedures is that zero\-probability responses yield undefined surprisal values\. This occurs when no participant predicts the accurate target wordwiw\_\{i\}given the preceding contextw1,w2,…,wi−1w\_\{1\},w\_\{2\},\\ldots,w\_\{i\-1\}\. With realistic sample sizes, words withP​\(wi\|w1,w2,…,wi−1\)<0\.001P\(w\_\{i\}\\allowbreak\|\\allowbreak w\_\{1\},\\allowbreak w\_\{2\},\\allowbreak\\ldots,\\allowbreak w\_\{i\-1\}\)<0\.001are absent from responses, motivating an expanded probability distribution to include more words\. Accordingly, following\\@BBOPcitet\\@BAP\\@BBNLowderet al\.\([2018](https://arxiv.org/html/2607.16549#bib.bib64)\)\\@BBCP, we replace zero cloze probabilities with half the smallest nonzero value before computing surprisal\.

#### N\-gram models

In this study, we trained bigram to quadgram models using the NLTK Python package111[https://www\.nltk\.org/](https://www.nltk.org/)\. Unlike advanced language models such as transformer\-based LMs\\@BBOPcitep\\@BAP\\@BBN\(Amaratunga,[2023](https://arxiv.org/html/2607.16549#bib.bib60); Desaiet al\.,[2023](https://arxiv.org/html/2607.16549#bib.bib61)\)\\@BBCP, n\-gram models have a limitation in capturing very long\-range dependencies\. To mitigate this, we trained our n\-gram models on the Fairy Tale Corpus\\@BBOPcitep\\@BAP\\@BBN\(Lobo and de Matos,[2010](https://arxiv.org/html/2607.16549#bib.bib62)\)\\@BBCP, a domain\-aligned dataset to DERCo, comprising 453 fairy tales from Project Gutenberg\\@BBOPcitep\\@BAP\\@BBN\(Klein and Manning,[2002](https://arxiv.org/html/2607.16549#bib.bib63)\)\\@BBCP\.

To mitigate overfitting, we excluded the five Grimm Brothers’ Fairy Tales used in the DERCo dataset’s transcript\. All punctuation was removed, and the letters were converted to lowercase\. To address data sparsity, we trained separate n\-gram models with no smoothing, Laplace smoothing, and Kneser\-Ney smoothing\. Each model then used a fixed\-sized context window ofn \- 1words, locating all matching windows within the problem instance and counting the number of occurrences of each possible next token to calculate word probabilities for top\-1 prediction and surprisal estimations\.

#### GPT\-2 and GPT\-Neo Families

Generative Pre\-trained Transformer \(GPT\)\\@BBOPcitep\\@BAP\\@BBN\(Radford,[2018](https://arxiv.org/html/2607.16549#bib.bib88)\)\\@BBCPis a transformer\-based autoregressive language model that uses a multi\-head “attention” mechanism based on an encoder\-decoder architecture\\@BBOPcitep\\@BAP\\@BBN\(Vaswani,[2017](https://arxiv.org/html/2607.16549#bib.bib89)\)\\@BBCP\. We selected GPT\-2 and GPT\-Neo families for investigation because they are mainly trained to predict the upcoming tokens based on probability in a left\-to\-right manner, similar to the cloze procedure conducted in next\-word prediction tasks\\@BBOPcitep\\@BAP\\@BBN\(Taylor,[1953](https://arxiv.org/html/2607.16549#bib.bib36)\)\\@BBCP\.

The transformer models take token sequences \(e\.g\., words, phonemes, or punctuation\) as input, with context window depends on the story length\. If a story exceeds one context window \(e\.g\., 1024 tokens\), probability estimates for subsequent tokens are conditioned on the second half of the previous window\. Predictions were performed separately for each story, with context reset between stories rather than carrying over the context from the previous one\. The story’s topic was included in the prediction, as it was disclosed to participants at the beginning of the experiment in the DERCo dataset\.

For words spanning multiple tokens, the word probability was calculated as the joint probability of the tokens using the chain rule\. Surprisal was obtained by summing token\-level surprisal values, reflecting the online experiment in which participants viewed words together with associated punctuation\. In contrast, top\-1 prediction focused only on correct word identification; since participants typed only the word without punctuation, the corresponding probability was averaged across the constituent tokens\.

Models were implemented in PyTorch using the transformer modules from the HuggingFace Hub222https://huggingface\.co/\. The examined GPT family variants differed primarily in their size, with the specific hyperparameters outlined in Table[1](https://arxiv.org/html/2607.16549#Sx2.T1)\.

Model Namenlayersn\_\{\\text\{layers\}\}nheadn\_\{\\text\{head\}\}dmodeld\_\{\\text\{model\}\}nparamsn\_\{\\text\{params\}\}GPT\-2 Small1212768∼\\sim124MGPT\-2 Medium24161024∼\\sim355MGPT\-2 Large36201280∼\\sim774MGPT\-Neo 125M1212768∼\\sim125MGPT\-Neo 1\.3B24162048∼\\sim1\.3BGPT\-Neo 2\.7B32162560∼\\sim2\.7B

Table 1:Hyper\-parameters of GPT\-2, GPT\-Neo families\.nlayersn\_\{\\text\{layers\}\},nheadn\_\{\\text\{head\}\},dmodeld\_\{\\text\{model\}\}, andnparamsn\_\{\\text\{params\}\}respectively refer to the number of layers, number of attention heads per layer, embedding size, and number of parameters\.

### Brain Encoding Models

![Refer to caption](https://arxiv.org/html/2607.16549v1/figure_01.jpg)Figure 1:Brain encoding model was used to predict the neural responses to each word in the context\.Brain encoding models\\@BBOPcitep\\@BAP\\@BBN\(Heilbronet al\.,[2019](https://arxiv.org/html/2607.16549#bib.bib54); Goldsteinet al\.,[2022](https://arxiv.org/html/2607.16549#bib.bib27)\)\\@BBCPentail fitting a regressor to predict neural responses at each time point based on measures such as lexical surprisal and top\-1 prediction for each participant, as used in this study\. However, running separate mixed effects models for each time point, as in prior studies\\@BBOPcitep\\@BAP\\@BBN\(Franket al\.,[2013](https://arxiv.org/html/2607.16549#bib.bib42); Haoet al\.,[2020](https://arxiv.org/html/2607.16549#bib.bib48); Ohet al\.,[2024](https://arxiv.org/html/2607.16549#bib.bib66)\)\\@BBCP, would require estimating an excessive number of parameters, potentially leading to model overfitting through the capture of noise rather than meaningful patterns\. To reduce this, we use ridge regression, introduced by\\@BBOPcitet\\@BAP\\@BBNHoerl and Kennard \([1970](https://arxiv.org/html/2607.16549#bib.bib68)\)\\@BBCP, which regularizes the model to control its variability and improve prediction reliability\\@BBOPcitep\\@BAP\\@BBN\(Bishop and Nasrabadi,[2006](https://arxiv.org/html/2607.16549#bib.bib67)\)\\@BBCP\.

Figure[1](https://arxiv.org/html/2607.16549#Sx2.F1)illustrates the procedure for predicting EEG data using information\-theoretic measures from the cloze experiment and LMs\. In brief, words from the DERCo transcript were aligned with the EEG recordings from the EEG\-based reading experiment \(serving as the ground truth\), with each word onset marked as time point 0 ms to standardise timing across trials\. Two measures, top\-1 prediction and surprisal, were used separately as independent variables to build regressors for predicting the EEG signal in a word\-by\-word, time\-resolved manner\. After running regressions for each measure, we estimated a series of predicted EEG amplitudes fromtm​i​nt\_\{min\}\(\-100 ms\) totm​a​xt\_\{max\}\(500 ms\) for each subject per electrode\.

We quantified the similarity between the predicted and actual EEG responses using Pearson’s correlation coefficient \(rr\) at the word level to validate the model’s ability to capture neural representations\. Using 5\-fold cross\-validation, each regressor was trained on 80% of the data to learn 600 coefficients for each combination of time point and sensor, and then tested on the remaining 20% of held\-out words\.

To ensure consistent regression parameters across subjects, we selected the regularisation hyperparameterλ\\lambdaby fitting a single model to pooled data from all participants across a log\-spaced range \[10−5,10−4,…,10510^\{\-5\},10^\{\-4\},…,10^\{5\}\]\. The optimalλ\\lambdawas chosen yielding the lowest generalisation error, as estimated via 5\-fold cross\-validation shuffled trials, using the Scikit\-learn library\\@BBOPcitep\\@BAP\\@BBN\(Pedregosaet al\.,[2011](https://arxiv.org/html/2607.16549#bib.bib69)\)\\@BBCP\.

### Significance Tests

In EEG research, conducting more statistical tests increases the probability of getting a false positive result due to random chance\\@BBOPcitep\\@BAP\\@BBN\(Greenlandet al\.,[2016](https://arxiv.org/html/2607.16549#bib.bib71)\)\\@BBCP\. This issue is significantly amplified in this work, with 32 EEG sensors and 600\-time points resulting in 19,200tt\-values per subject\. To address this, we implemented cluster\-based permutation tests\\@BBOPcitep\\@BAP\\@BBN\(Maris and Oostenveld,[2007](https://arxiv.org/html/2607.16549#bib.bib72)\)\\@BBCP, using 5,000 permutations per test to identify significant time windows for the encoding models\.

A mass\-univariate testing approach applies one\-samplett\-tests with a “hat” variance adjustment to compensate implausibly small variances \(σ=10−3\\sigma=10^\{\-3\}\)\\@BBOPcitep\\@BAP\\@BBN\(Ridgwayet al\.,[2012](https://arxiv.org/html/2607.16549#bib.bib91)\)\\@BBCP\. The resultingtt\-statistics were used to computepp\-values, which were then corrected for multiple comparisons using the false discovery rate \(FDR\)\\@BBOPcitep\\@BAP\\@BBN\(Genoveseet al\.,[2002](https://arxiv.org/html/2607.16549#bib.bib90)\)\\@BBCPfor each time point and electrode to ensure statistically consistent effects\.

## Results

### Performance and Human Alignment in Different Language Models

Prior studies quantified neural pattern similarity between word pairs using Pearson correlation\\@BBOPcitep\\@BAP\\@BBN\(Heet al\.,[2022](https://arxiv.org/html/2607.16549#bib.bib76); Goldsteinet al\.,[2022](https://arxiv.org/html/2607.16549#bib.bib27)\)\\@BBCP, as it captures neural patterns similarity regardless of their amplitude\. To strengthen our analysis, we also examined the Spearman correlation to provide more robust evidence for the observed relationships\.

Table[2](https://arxiv.org/html/2607.16549#Sx3.T2)presents the correlation between LMs and human predictions in terms of top\-1 prediction and surprisal, along with their performance in the next\-word prediction task, as measured by accuracy\. Figure[2](https://arxiv.org/html/2607.16549#Sx3.F2)further visualises differences in accuracy and joint accuracy between LMs and human predictions\. These results yield several important observations\.

Model VariantsTop\-1 Prediction\(vs\. Human\)Surprisal\(vs\. Human\)Top\-1 Prediction\(vs\. Human\)Surprisal\(vs\. Human\)Accuracy\(%\)Pearson rpPearson rpSpearman rpSpearman rpBigrams \(KN\)0\.09<p∗<p^\{\*\}0\.41<p∗<p^\{\*\}0\.040\.020\.020\.41<p∗<p^\{\*\}14\.92Trigrams \(KN\)0\.14<p∗<p^\{\*\}0\.26<p∗<p^\{\*\}0\.12<p∗<p^\{\*\}0\.28<p∗<p^\{\*\}19\.23Quadgrams \(KN\)0\.050\.010\.010\.15<p∗<p^\{\*\}0\.020\.220\.15<p∗<p^\{\*\}19\.36GPT\-2 Small0\.52<p∗<p^\{\*\}0\.73<p∗<p^\{\*\}0\.51<p∗<p^\{\*\}0\.78<p∗<p^\{\*\}30\.62GPT\-2 Medium0\.56<p∗<p^\{\*\}0\.75<p∗<p^\{\*\}0\.55<p∗<p^\{\*\}0\.80<p∗<p^\{\*\}33\.64GPT\-2 Large0\.59<p∗<p^\{\*\}0\.77<p∗<p^\{\*\}0\.57<p∗<p^\{\*\}0\.82<p∗<p^\{\*\}34\.59GPT\-Neo 125M0\.52<p∗<p^\{\*\}0\.74<p∗<p^\{\*\}0\.51<p∗<p^\{\*\}0\.78<p∗<p^\{\*\}29\.33GPT\-Neo 1\.3B0\.59<p∗<p^\{\*\}0\.77<p∗<p^\{\*\}0\.58<p∗<p^\{\*\}0\.82<p∗<p^\{\*\}35\.77GPT\-Neo 2\.7B0\.61<p∗<p^\{\*\}0\.77<p∗<p^\{\*\}0\.59<p∗<p^\{\*\}0\.83<p∗<p^\{\*\}37\.20Human1\.000\.000\.001\.000\.000\.001\.000\.000\.001\.000\.000\.0045\.17

Table 2:Correlation comparisons of various language model families against human benchmarks in a next\-word prediction task\. Metrics include accuracy, Pearson and Spearman correlation coefficientsrrfor top\-1 prediction and surprisal\. Statistically significant correlations withp∗=0\.001p^\{\*\}=0\.001are indicated in a two\-sided test\.![Refer to caption](https://arxiv.org/html/2607.16549v1/figure_06.jpg)Figure 2:Performance comparison of LM families \(N\-grams, GPT\-2, GPT\-Neo\) and human benchmarks in a next\-word prediction task, evaluated by per\-model accuracy and joint prediction percentages “both correct” \(green\) and “both incorrect” \(red\)\.First, humans still outperformed these LMs, achieving the highest accuracy \(45\.17%\), highlighting the gap between LMs and human predictive capabilities\. The results indicate that LMs predict next words similarly to humans in narrative contexts, with their performance gradually approaching that of humans; larger LMs are more accurate than smaller ones, but they have not surpassed human performance\.

Among n\-gram models, quadgrams achieved the highest accuracy, but trigrams showed the strongest correlation with humans in top\-1 predictions, while bigrams had the highest surprisal correlation\. Overall, correlations between n\-gram models and human predictions were generally weak and inconsistent\. These results show that while historically important, n\-gram models struggle to capture the complexity of human\-like next\-word prediction\.

GPT\-Neo models outperformed GPT\-2, with GPT\-Neo 2\.7B achieving the best accuracy \(37\.20%\) and the highest percentage of joint correct predictions \(29\.91%\) among these transformer\-based LMs\. Both model families demonstrated strong correlations with human behaviour in top\-1 predictions and surprisal metrics, with performance improving as the model size increases\. Consistent increasing in both Pearson’s and Spearman’srrcorrelations indicate a robust, distribution\-independent linear relationship, suggesting larger models better capture human\-like linguistic processing as evidenced by improved predictive accuracy and joint performance metrics\.

Although advanced LMs improve significantly with scale, the mechanisms underlying their correlation with human behaviour remain unclear\.Do these models truly reflect human reading processes, or do they just exhibit surface\-level convergence in next\-word prediction?To investigate this, we analyse and compare results from brain encoding models, using information\-theoretic measures estimated by these LMs, as detailed in the following sections\.

### Neural Encoding Using Predictive Metrics from Human Cloze

![Refer to caption](https://arxiv.org/html/2607.16549v1/figure_02.jpg)Figure 3:Correlations between human cloze–based regressors and neural responses for content and function words over time, shown for lexical surprisal \(left\) and top\-1 prediction \(right\)\. Encoding analysis was conducted per electrode and then averaged across electrodes\. Asterisks denote significant time windows \(p<0\.001p<0\.001\) based on cluster\-based permutation tests\. The shaded regions represent the between\-subject standard error of the mean \(SEM\) of the encoding models\.Figure[3](https://arxiv.org/html/2607.16549#Sx3.F3)shows that lexical surprisal derived from human cloze probabilities is a stronger predictor of neural responses, particularly within the N400 time window, compared to the top\-1 prediction measure\. These results align with the well\-established semantic effects on N400 amplitude\\@BBOPcitep\\@BAP\\@BBN\(Franket al\.,[2013](https://arxiv.org/html/2607.16549#bib.bib42); Michaelov and Bergen,[2020](https://arxiv.org/html/2607.16549#bib.bib41)\)\\@BBCP\. Additionally, encoding correlations are stronger for content words than function words, suggesting that surprisal more effectively captures neural processes associated with semantically rich lexical words\\@BBOPcitep\\@BAP\\@BBN\(Münteet al\.,[2001](https://arxiv.org/html/2607.16549#bib.bib93); Heet al\.,[2022](https://arxiv.org/html/2607.16549#bib.bib76)\)\\@BBCP\.

Therefore, we use human cloze results as the baseline for quantifying LMs’ next\-word predictability, lexical surprisal as the most informative metric, and mainly focus on content words to maximise analytical sensitivity\.

![Refer to caption](https://arxiv.org/html/2607.16549v1/figure_04.png)Figure 4:Radar plots of the mean \(top\) and standard deviation \(bottom\) of cross\-subject surprisal correlations for content words, averaged over electrodes\.![Refer to caption](https://arxiv.org/html/2607.16549v1/figure_05.jpg)Figure 5:Full EEG Topographies of grand\-averaged encoding correlations \(Pearson’srr\) for lexical surprisal, computed by human predictive modelling and representative LMs over time\.
### Encoding Performance Comparison

We examined encoding performance at the individual\-subject level and selected Bigrams, GPT\-2 Large, and GPT\-Neo 2\.7B as representative models for each LM family based on their top performance within their groups \(see Figure[2](https://arxiv.org/html/2607.16549#Sx3.F2)\)\. As shown in Figure[4](https://arxiv.org/html/2607.16549#Sx3.F4), the GPT\-2 Large regression model shows the strongest correlations, with its mean and standard deviation for individual subjects closely aligning with human predictive models\. Both the overall shape and the pattern of encoding correlation variances \(i\.e\., the increase or decrease in values\) are similar between GPT\-2 Large and the human regressor\.

### Topographic EEG analysis

Significance levels of observed differences are reported in Figure[5](https://arxiv.org/html/2607.16549#Sx3.F5)\. The transformer\-based LMs can effectively track human neural signals using surprisal metric, particularly in predicting the content words compared to n\-grams\. Additionally, surprisal\-based GPT\-2 shows the best alignment with the encoding results from the cloze procedure\.

In the early time window \(\-100\-100 ms\), no significant spatial pattern emerges, suggesting diffuse neural processing\. However, in the200–300 mswindow, changes in correlation are observed in the centro\-frontal and parieto\-occipital regions, in line with the view that the P200 component reflects the processing of unexpected or affectively salient linguistic information\\@BBOPcitep\\@BAP\\@BBN\(Raney,[1993](https://arxiv.org/html/2607.16549#bib.bib98); Leutholdet al\.,[2015](https://arxiv.org/html/2607.16549#bib.bib96)\)\\@BBCP\. In the300–500ms range, stronger encoding correlations are observed, aligning with the N400 effect, which is associated with the “expectedness” level representation on the processing of upcoming words\\@BBOPcitep\\@BAP\\@BBN\(Hoekset al\.,[2004](https://arxiv.org/html/2607.16549#bib.bib97); Yeet al\.,[2022](https://arxiv.org/html/2607.16549#bib.bib94)\)\\@BBCP\.

## Discussion and Conclusion

This study investigated how LM predictions relate to neural responses during reading using surprisal and top\-1 prediction, and their alignment with human next\-word predictability\. The findings indicated that surprisal\-based predictors showed significant differences in neural responses for these two lexical categories, whereas top\-1 prediction failed to capture this difference\. Furthermore, although larger and more advanced language models typically showed a close correspondence to human productions in next\-word prediction in terms of information\-theoretic measures, our results demonstrated that larger model size and increased computational resources may not reliably produce more human\-like language processing\. Notably, surprisal\-based GPT\-2 Large regression substantially outperformed larger and more advanced language models in both overall and individual subject\-level analyses\.

Despite these insights, several limitations must be acknowledged to guide future improvements\. First, our analysis focused mainly on GPT\-family transformer models; however, the rapidly evolving LM landscape includes other unidirectional architectures, such as LLaMA\\@BBOPcitep\\@BAP\\@BBN\(Touvronet al\.,[2023](https://arxiv.org/html/2607.16549#bib.bib80)\)\\@BBCPand DeepSeekMoE\\@BBOPcitep\\@BAP\\@BBN\(Daiet al\.,[2024](https://arxiv.org/html/2607.16549#bib.bib99)\)\\@BBCP, whose distinct training objectives may impact cognitive modelling\. Future work should therefore examine a broader range of transformer\-based models\. Second, we examined next\-word predictability only via lexical groups, capturing a limited aspect of human reading behaviour\. Prior work suggests that surprisal sensitivity also varies with word frequency and predictability\\@BBOPcitep\\@BAP\\@BBN\(van Schijndel and Linzen,[2019](https://arxiv.org/html/2607.16549#bib.bib82); Michaelov and Bergen,[2020](https://arxiv.org/html/2607.16549#bib.bib41); Xiaet al\.,[2023](https://arxiv.org/html/2607.16549#bib.bib83); Ohet al\.,[2024](https://arxiv.org/html/2607.16549#bib.bib66)\)\\@BBCP, potentially limiting generalisability across word difficulty levels\. Finally, our analyses emphasised top\-down semantic processing, whereas predictability estimates derived from language models reflect a combination of semantic and syntactic information\\@BBOPcitep\\@BAP\\@BBN\(Qian and Levy,[2019](https://arxiv.org/html/2607.16549#bib.bib84); Wilcoxet al\.,[2021](https://arxiv.org/html/2607.16549#bib.bib86); Arehalliet al\.,[2022](https://arxiv.org/html/2607.16549#bib.bib85)\)\\@BBCP; future work should incorporate syntactic contributions\.

## Acknowledgements

This publication has emanated from research conducted with the financial support of Science Foundation Ireland under Grant number 18/CRT/6183 and 13/RC/2106\_P2\. We would like to thank the anonymous reviewers for their helpful remarks\.

## Code Availability

## References

- NLP through the ages\.InUnderstanding Large Language Models: Learning Their Underlying Concepts and Technologies,pp\. 9–54\.Cited by:[N\-gram models](https://arxiv.org/html/2607.16549#Sx2.SSx3.SSSx2.p1.1)\.
- S\. Arehalli, B\. W\. Dillon, and T\. Linzen \(2022\)Syntactic surprisal from neural models predicts, but underestimates, human processing difficulty from syntactic ambiguities\.InProceedings of the 26th Conference on Computational Natural Language Learning \(CoNLL\),pp\. 301–313\.Cited by:[Discussion and Conclusion](https://arxiv.org/html/2607.16549#Sx4.p2.1)\.
- K\. Armeni, R\. M\. Willems, A\. Van den Bosch, and J\. Schoffelen \(2019\)Frequency\-specific brain dynamics related to prediction during language comprehension\.NeuroImage198,pp\. 283–295\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p3.1)\.
- M\. Bar \(2007\)The proactive brain: using analogies and associations to generate predictions\.Trends in cognitive sciences11\(7\),pp\. 280–289\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p1.1)\.
- C\. M\. Bishop and N\. M\. Nasrabadi \(2006\)Pattern recognition and machine learning\.Vol\.4,Springer\.Cited by:[Brain Encoding Models](https://arxiv.org/html/2607.16549#Sx2.SSx4.p1.1)\.
- D\. Dai, C\. Deng, C\. Zhao, R\.x\. Xu, H\. Gao, D\. Chen, J\. Li, W\. Zeng, X\. Yu, Y\. Wu, Z\. Xie, Y\.k\. Li, P\. Huang, F\. Luo, C\. Ruan, Z\. Sui, and W\. Liang \(2024\)DeepSeekMoE: towards ultimate expert specialization in mixture\-of\-experts language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 1280–1297\.External Links:[Link](https://aclanthology.org/2024.acl-long.70/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.70)Cited by:[Discussion and Conclusion](https://arxiv.org/html/2607.16549#Sx4.p2.1)\.
- B\. Desai, K\. Patil, A\. Patil, and I\. Mehta \(2023\)Large language models: a comprehensive exploration of modern ai’s potential and pitfalls\.Journal of Innovative Technologies6\(1\)\.Cited by:[N\-gram models](https://arxiv.org/html/2607.16549#Sx2.SSx3.SSSx2.p1.1)\.
- H\. Fitz and F\. Chang \(2019\)Language erps reflect learning through prediction error propagation\.Cognitive Psychology111,pp\. 15–52\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p4.1)\.
- S\. L\. Frank, L\. J\. Otten, G\. Galli, and G\. Vigliocco \(2013\)Word surprisal predicts n400 amplitude during reading\.InProceedings of the 51st Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 878–883\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p2.1),[Brain Encoding Models](https://arxiv.org/html/2607.16549#Sx2.SSx4.p1.1),[Neural Encoding Using Predictive Metrics from Human Cloze](https://arxiv.org/html/2607.16549#Sx3.SSx2.p1.1)\.
- S\. L\. Frank, L\. J\. Otten, G\. Galli, and G\. Vigliocco \(2015\)The erp response to the amount of information conveyed by words in sentences\.Brain and language140,pp\. 1–11\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p3.1)\.
- C\. R\. Genovese, N\. A\. Lazar, and T\. Nichols \(2002\)Thresholding of statistical maps in functional neuroimaging using the false discovery rate\.Neuroimage15\(4\),pp\. 870–878\.Cited by:[Significance Tests](https://arxiv.org/html/2607.16549#Sx2.SSx5.p2.4)\.
- A\. Goldstein, Z\. Zada, E\. Buchnik, M\. Schain, A\. Price, B\. Aubrey, S\. A\. Nastase, A\. Feder, D\. Emanuel, A\. Cohen,et al\.\(2022\)Shared computational principles for language processing in humans and deep language models\.Nature neuroscience25\(3\),pp\. 369–380\.Cited by:[Brain Encoding Models](https://arxiv.org/html/2607.16549#Sx2.SSx4.p1.1),[Performance and Human Alignment in Different Language Models](https://arxiv.org/html/2607.16549#Sx3.SSx1.p1.1)\.
- S\. Greenland, S\. J\. Senn, K\. J\. Rothman, J\. B\. Carlin, C\. Poole, S\. N\. Goodman, and D\. G\. Altman \(2016\)Statistical tests, p values, confidence intervals, and power: a guide to misinterpretations\.European journal of epidemiology31\(4\),pp\. 337–350\.Cited by:[Significance Tests](https://arxiv.org/html/2607.16549#Sx2.SSx5.p1.1)\.
- J\. Hale \(2001\)A probabilistic earley parser as a psycholinguistic model\.InSecond meeting of the north american chapter of the association for computational linguistics,Cited by:[Surprisal Estimation](https://arxiv.org/html/2607.16549#Sx2.SSx2.SSSx2.p1.1)\.
- J\. Hale \(2016\)Information\-theoretical complexity metrics\.Language and Linguistics Compass10\(9\),pp\. 397–412\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p1.1)\.
- Y\. Hao, S\. Mendelsohn, R\. Sterneck, R\. Martinez, and R\. Frank \(2020\)Probabilistic predictions of people perusing: evaluating metrics of language model performance for psycholinguistic modeling\.InProceedings of the Workshop on Cognitive Modeling and Computational Linguistics,pp\. 75–86\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p3.1),[Brain Encoding Models](https://arxiv.org/html/2607.16549#Sx2.SSx4.p1.1)\.
- T\. He, M\. A\. Boudewyn, J\. E\. Kiat, K\. Sagae, and S\. J\. Luck \(2022\)Neural correlates of word representation vectors in natural language processing models: evidence from representational similarity analysis of event\-related brain potentials\.Psychophysiology59\(3\),pp\. e13976\.Cited by:[Performance and Human Alignment in Different Language Models](https://arxiv.org/html/2607.16549#Sx3.SSx1.p1.1),[Neural Encoding Using Predictive Metrics from Human Cloze](https://arxiv.org/html/2607.16549#Sx3.SSx2.p1.1)\.
- M\. Heilbron, B\. Ehinger, P\. Hagoort, and F\. P\. De Lange \(2019\)Tracking naturalistic linguistic predictions with deep neural language models\.In2019 Conference on Cognitive Computational Neuroscience \(CCN 2019\),pp\. 424–427\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p3.1),[Brain Encoding Models](https://arxiv.org/html/2607.16549#Sx2.SSx4.p1.1)\.
- J\. C\. Hoeks, L\. A\. Stowe, and G\. Doedens \(2004\)Seeing words in context: the interaction of lexical and sentence level information during reading\.Cognitive brain research19\(1\),pp\. 59–73\.Cited by:[Topographic EEG analysis](https://arxiv.org/html/2607.16549#Sx3.SSx4.p2.1)\.
- A\. E\. Hoerl and R\. W\. Kennard \(1970\)Ridge regression: applications to nonorthogonal problems\.Technometrics12\(1\),pp\. 69–82\.Cited by:[Brain Encoding Models](https://arxiv.org/html/2607.16549#Sx2.SSx4.p1.1)\.
- F\. Keller \(2010\)Cognitively plausible models of human language processing\.InProceedings of the ACL 2010 Conference Short Papers,pp\. 60–67\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p4.1)\.
- D\. Klein and C\. D\. Manning \(2002\)Fast exact inference with a factored model for natural language parsing\.Advances in neural information processing systems15\.Cited by:[N\-gram models](https://arxiv.org/html/2607.16549#Sx2.SSx3.SSSx2.p1.1)\.
- G\. H\. Klem \(1999\)The ten\-twenty electrode system of the international federation\. The international federation of clinical neurophysiology\.Electroencephalogr\. Clin\. Neurophysiol\. Suppl\.52,pp\. 3–6\.Cited by:[Stimuli and EEG Data Preparation](https://arxiv.org/html/2607.16549#Sx2.SSx1.p2.1)\.
- H\. Leuthold, A\. Kunkel, I\. G\. Mackenzie, and R\. Filik \(2015\)Online processing of moral transgressions: erp evidence for spontaneous evaluation\.Social cognitive and affective neuroscience10\(8\),pp\. 1021–1029\.Cited by:[Topographic EEG analysis](https://arxiv.org/html/2607.16549#Sx3.SSx4.p2.1)\.
- R\. Levy \(2008\)Expectation\-based syntactic comprehension\.Cognition106\(3\),pp\. 1126–1177\.Cited by:[Surprisal Estimation](https://arxiv.org/html/2607.16549#Sx2.SSx2.SSSx2.p1.1)\.
- P\. V\. Lobo and D\. M\. de Matos \(2010\)Fairy tale corpus organization using latent semantic mapping and an item\-to\-item top\-n recommendation algorithm\.InProceedings of the Seventh International Conference on Language Resources and Evaluation \(LREC’10\),Cited by:[N\-gram models](https://arxiv.org/html/2607.16549#Sx2.SSx3.SSSx2.p1.1)\.
- M\. W\. Lowder, W\. Choi, F\. Ferreira, and J\. M\. Henderson \(2018\)Lexical predictability during natural reading: effects of surprisal and entropy reduction\.Cognitive science42,pp\. 1166–1183\.Cited by:[Human Prediction](https://arxiv.org/html/2607.16549#Sx2.SSx3.SSSx1.p3.3)\.
- E\. Maris and R\. Oostenveld \(2007\)Nonparametric statistical testing of eeg\-and meg\-data\.Journal of neuroscience methods164\(1\),pp\. 177–190\.Cited by:[Significance Tests](https://arxiv.org/html/2607.16549#Sx2.SSx5.p1.1)\.
- J\. A\. Michaelov, M\. D\. Bardolph, C\. K\. Van Petten, B\. K\. Bergen, and S\. Coulson \(2024\)Strong prediction: language model surprisal explains multiple n400 effects\.Neurobiology of language5\(1\),pp\. 107–135\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p3.1)\.
- J\. A\. Michaelov and B\. K\. Bergen \(2020\)How well does surprisal explain n400 amplitude under different experimental conditions?\.InProceedings of the 24th Conference on Computational Natural Language Learning,pp\. 652–663\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p2.1),[Neural Encoding Using Predictive Metrics from Human Cloze](https://arxiv.org/html/2607.16549#Sx3.SSx2.p1.1),[Discussion and Conclusion](https://arxiv.org/html/2607.16549#Sx4.p2.1)\.
- M\. Mitchell and D\. C\. Krakauer \(2023\)The debate over understanding in ai’s large language models\.Proceedings of the National Academy of Sciences120\(13\),pp\. e2215907120\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p4.1)\.
- T\. F\. Münte, B\. M\. Wieringa, H\. Weyerts, A\. Szentkuti, M\. Matzke, and S\. Johannes \(2001\)Differences in brain potentials to open and closed class words: class and frequency effects\.Neuropsychologia39\(1\),pp\. 91–102\.Cited by:[Neural Encoding Using Predictive Metrics from Human Cloze](https://arxiv.org/html/2607.16549#Sx3.SSx2.p1.1)\.
- B\. Oh, S\. Yue, and W\. Schuler \(2024\)Frequency explains the inverse correlation of large language models’ size, training data amount, and surprisal’s fit to reading times\.InProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 2644–2663\.Cited by:[Brain Encoding Models](https://arxiv.org/html/2607.16549#Sx2.SSx4.p1.1),[Discussion and Conclusion](https://arxiv.org/html/2607.16549#Sx4.p2.1)\.
- F\. Pedregosa, G\. Varoquaux, A\. Gramfort, V\. Michel, B\. Thirion, O\. Grisel, M\. Blondel, P\. Prettenhofer, R\. Weiss, V\. Dubourg,et al\.\(2011\)Scikit\-learn: machine learning in python\.the Journal of machine Learning research12,pp\. 2825–2830\.Cited by:[Brain Encoding Models](https://arxiv.org/html/2607.16549#Sx2.SSx4.p4.3)\.
- P\. Qian and R\. P\. Levy \(2019\)Neural language models as psycholinguistic subjects: representations of syntactic state\.Cited by:[Discussion and Conclusion](https://arxiv.org/html/2607.16549#Sx4.p2.1)\.
- B\. M\. Quach, C\. Gurrin, and G\. Healy \(2024\)DERCo: a dataset for human behaviour in reading comprehension using eeg\.Scientific Data11\(1\),pp\. 1104\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p5.1),[Stimuli and EEG Data Preparation](https://arxiv.org/html/2607.16549#Sx2.SSx1.p1.1)\.
- A\. Radford \(2018\)Improving language understanding by generative pre\-training\.Cited by:[GPT\-2 and GPT\-Neo Families](https://arxiv.org/html/2607.16549#Sx2.SSx3.SSSx3.p1.1)\.
- G\. E\. Raney \(1993\)Monitoring changes in cognitive load during reading: an event\-related brain potential and reaction time analysis\.\.Journal of Experimental Psychology: Learning, Memory, and Cognition19\(1\),pp\. 51\.Cited by:[Topographic EEG analysis](https://arxiv.org/html/2607.16549#Sx3.SSx4.p2.1)\.
- G\. R\. Ridgway, V\. Litvak, G\. Flandin, K\. J\. Friston, and W\. D\. Penny \(2012\)The problem of low variance voxels in statistical parametric mapping; a new hat avoids a ‘haircut’\.Neuroimage59\(3\),pp\. 2131–2141\.Cited by:[Significance Tests](https://arxiv.org/html/2607.16549#Sx2.SSx5.p2.4)\.
- C\. Shain, C\. Meister, T\. Pimentel, R\. Cotterell, and R\. Levy \(2024\)Large\-scale evidence for logarithmic effects of word predictability on reading time\.Proceedings of the National Academy of Sciences121\(10\),pp\. e2307876121\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p1.1),[Introduction](https://arxiv.org/html/2607.16549#Sx1.p3.1)\.
- C\. E\. Shannon \(1948\)A mathematical theory of communication\.The Bell system technical journal27\(3\),pp\. 379–423\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p1.1)\.
- W\. L\. Taylor \(1953\)“Cloze procedure”: A new tool for measuring readability\.Journalism quarterly30\(4\),pp\. 415–433\.Cited by:[GPT\-2 and GPT\-Neo Families](https://arxiv.org/html/2607.16549#Sx2.SSx3.SSSx3.p1.1)\.
- H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet, M\. Lachaux, T\. Lacroix, B\. Rozière, N\. Goyal, E\. Hambro, F\. Azhar,et al\.\(2023\)Llama: open and efficient foundation language models\.arXiv preprint arXiv:2302\.13971\.Cited by:[Discussion and Conclusion](https://arxiv.org/html/2607.16549#Sx4.p2.1)\.
- M\. van Schijndel and T\. Linzen \(2019\)Can entropy explain successor surprisal effects in reading?\.InProceedings of the Society for Computation in Linguistics \(SCiL\) 2019,pp\. 1–7\.Cited by:[Discussion and Conclusion](https://arxiv.org/html/2607.16549#Sx4.p2.1)\.
- A\. Vaswani \(2017\)Attention is all you need\.Advances in Neural Information Processing Systems\.Cited by:[GPT\-2 and GPT\-Neo Families](https://arxiv.org/html/2607.16549#Sx2.SSx3.SSSx3.p1.1)\.
- E\. Wilcox, P\. Vani, and R\. Levy \(2021\)A targeted assessment of incremental processing in neural language models and humans\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 939–952\.Cited by:[Discussion and Conclusion](https://arxiv.org/html/2607.16549#Sx4.p2.1)\.
- R\. M\. Willems, S\. L\. Frank, A\. D\. Nijhof, P\. Hagoort, and A\. Van den Bosch \(2016\)Prediction during natural language comprehension\.Cerebral cortex26\(6\),pp\. 2506–2516\.Cited by:[Introduction](https://arxiv.org/html/2607.16549#Sx1.p3.1)\.
- M\. Xia, M\. Artetxe, C\. Zhou, X\. V\. Lin, R\. Pasunuru, D\. Chen, L\. Zettlemoyer, and V\. Stoyanov \(2023\)Training trajectories of language models across scales\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13711–13738\.Cited by:[Discussion and Conclusion](https://arxiv.org/html/2607.16549#Sx4.p2.1)\.
- Z\. Ye, X\. Xie, Y\. Liu, Z\. Wang, X\. Chen, M\. Zhang, and S\. Ma \(2022\)Towards a better understanding of human reading comprehension with brain signals\.InProceedings of the ACM Web Conference 2022,pp\. 380–391\.Cited by:[Topographic EEG analysis](https://arxiv.org/html/2607.16549#Sx3.SSx4.p2.1)\.

Similar Articles

Heterogeneous Neural Predictivity from Language Models During Naturalistic Comprehension

arXiv cs.CL

This paper investigates how language model representations predict neural activity during naturalistic language comprehension across MEG, ECoG, and other recordings. The findings demonstrate that language model features serve as useful neural predictors, but caution against overinterpreting predictive success as evidence for shared neural organization.

Human-Like Anaphor Resolution in Large Language Models

arXiv cs.CL

This paper investigates whether five open-weight LLMs exhibit human-like sensitivity to psycholinguistic factors in anaphor resolution, using surprisal and comprehension accuracy as behavioral measures. Results show selective cognitive alignment, with some models matching human discourse sensitivity but not semantic interference effects.