Human-Like Anaphor Resolution in Large Language Models

arXiv cs.CL Papers

Summary

This paper investigates whether five open-weight LLMs exhibit human-like sensitivity to psycholinguistic factors in anaphor resolution, using surprisal and comprehension accuracy as behavioral measures. Results show selective cognitive alignment, with some models matching human discourse sensitivity but not semantic interference effects.

arXiv:2608.05630v1 Announce Type: new Abstract: Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:51 AM

# Human-Like Anaphor Resolution in Large Language Models
Source: [https://arxiv.org/html/2608.05630](https://arxiv.org/html/2608.05630)
vchinta6rajsanjayshahvarma\}@gatech\.edu Georgia Institute of Technology![[Uncaptioned image]](https://arxiv.org/html/2608.05630v1/cropped-buzz-logo.png)

###### Abstract

*Anaphors*are expressions that refer to other expressions, called*antecedents*\. The process of connecting the two is called*resolution*\. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation\-model properties, and semantic factors\. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models \(LLMs\) with open weights: GPT\-2\-XL, Llama\-3\.1\-8B, Pythia\-12B, Mistral\-7B, and Mistral\-24B\. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor\. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors\. The results show selective cognitive alignment: some LLMs exhibit human\-like sensitivity to discourse prominence and distance\-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects\. These findings delimit the conditions under which LLMs approximate human anaphor resolution\.

Keywords:anaphor resolution; situation models; Large Language Models; surprisal; cognitive alignment

††footnotetext:Code for reproducing all model evaluations, analyses, and figures is available at:[github\.com/wristy/anaphor](https://github.com/wristy/anaphor)\.## Introduction

Large Language Models \(LLMs\) are deep neural networks with millions or billions of parameters trained on large text corpora\. Transformer\-based architectures, beginning with BERT\[[7](https://arxiv.org/html/2608.05630#bib.bib10)\]and GPT\-2\[[27](https://arxiv.org/html/2608.05630#bib.bib38)\]and extending to more recent families such as GPT\[[24](https://arxiv.org/html/2608.05630#bib.bib55)\], Llama\[[10](https://arxiv.org/html/2608.05630#bib.bib15)\], and Mistral\[[20](https://arxiv.org/html/2608.05630#bib.bib56)\], are highly performant on language comprehension tasks\[[35](https://arxiv.org/html/2608.05630#bib.bib52)\]\. Beyond task performance, these models are increasingly evaluated as candidate models of human cognition\[[13](https://arxiv.org/html/2608.05630#bib.bib19),[26](https://arxiv.org/html/2608.05630#bib.bib37),[30](https://arxiv.org/html/2608.05630#bib.bib41)\], and more specifically as models of human language processing\[[2](https://arxiv.org/html/2608.05630#bib.bib57),[16](https://arxiv.org/html/2608.05630#bib.bib24)\]\.

Much of the existing work evaluating LLMs as cognitive models has focused on phenomena at the word and sentence levels, examining model behavior in relatively local and decontextualized linguistic settings\. This leaves open the question of whether LLMs capture discourse\-level processes that unfold over extended text\. Anaphor resolution is one such process: it requires maintaining and retrieving information across sentences, integrating linguistic input with a developing situation model, and resolving interference from competing referents\. Modeling these demands goes beyond local prediction and instead requires sensitivity to connected text\. In sequential language models, shifts in expected word probabilities, typically operationalized as surprisal, have been shown to track human reading times, providing a process\-level link between model predictions and human comprehension behavior\[[36](https://arxiv.org/html/2608.05630#bib.bib53),[16](https://arxiv.org/html/2608.05630#bib.bib24)\]\.

Despite growing interest in cognitive alignment, comparatively little attention has been paid to text\- and discourse\-level phenomena\. The current study addresses this gap by focusing on anaphor resolution\. It focuses on*anaphors*, which are expressions that refer to other expressions, called*antecedents*\. Pronouns are a familiar class of anaphors; more generally, language is rife with referential expressions\. The process of connecting an anaphor to its antecedent during online comprehension is called*resolution*\. This can be easy when an anaphor and its antecedent occur within the same sentence or within a few sentences of each other; in this case, resolution requires only searching working memory\[[6](https://arxiv.org/html/2608.05630#bib.bib9)\]\. However, when reading longer texts, anaphors can be separated from their antecedents by many sentences\. In this case, resolution requires using an anaphor as a cue to memory and attempting to retrieve its antecedent from the reader’s situation model, which is the evolving representation of the overall meaning of the text\[[19](https://arxiv.org/html/2608.05630#bib.bib29),[33](https://arxiv.org/html/2608.05630#bib.bib46)\]\.

Cognitive science research has investigated the factors that affect the speed and accuracy of anaphor resolution for humans\. For example, the greater the number of sentences between an anaphor and its antecedent, the slower it is read, presumably because readers must “search" farther back in their memory for the text\. Reading time is one measure; another is the accuracy of comprehension questions\. To continue the example, the greater the number of sentences, the less accurate people are when answering a comprehension question about the antecedent, presumably because there is a greater chance that the resolution failed\.Here, we ask whether LLMs are sensitive to the same factors as human comprehenders during anaphor resolution\.To the extent that they do, LLMs gain credibility as candidate models of discourse\-level language processing in cognitive science\[[13](https://arxiv.org/html/2608.05630#bib.bib19),[30](https://arxiv.org/html/2608.05630#bib.bib41)\]\.

### Cognitive Science Studies of Anaphor Resolution

The current study focuses on six factors identified by cognitive science as affecting the speed and accuracy of anaphor resolution\.

The first factor is a thematic feature of texts: If the antecedent istopicalized, then it should feature more prominently in a reader’s situation model, and therefore it should be more accessible when an anaphor to it is encountered\. In this case, resolution should be relatively fast and accurate\. There are various ways to topicalize an antecedent\. For example,\[[23](https://arxiv.org/html/2608.05630#bib.bib33)\]manipulated whether or not it was focused by the title of the text\. The second factor concerns a surface feature of text itself: The greater thesentential distance\(i\.e\., number of sentences\) between an anaphor and its antecedent, the slower and less accurate resolution is\[[4](https://arxiv.org/html/2608.05630#bib.bib6)\]\.\[[23](https://arxiv.org/html/2608.05630#bib.bib33)\]orthogonally varied these two factors, antecedent topicality and sentential distance, and Experiment 1 runs LLMs on the materials of this study\.

The third and fourth factors concern not surface distance butcontextualdistance\. As readers comprehend a text, they build a model of the story world it describes, variously called the situation model, mental model, or world model\[[19](https://arxiv.org/html/2608.05630#bib.bib29),[33](https://arxiv.org/html/2608.05630#bib.bib46)\]\. The third factor is thespatial distance\(i\.e\., perceived physical distance\) between the anaphor and antecedent in the reader’s situation model\[[21](https://arxiv.org/html/2608.05630#bib.bib31),[28](https://arxiv.org/html/2608.05630#bib.bib39)\]\. The fourth factor is thetemporal duration\(i\.e\., perceived elapsed time\) between the two\[[1](https://arxiv.org/html/2608.05630#bib.bib1),[34](https://arxiv.org/html/2608.05630#bib.bib48),[37](https://arxiv.org/html/2608.05630#bib.bib54)\]\. The greater the spatial distance and the longer the temporal duration, the slower and less successful anaphor resolution becomes\. The following text illustrates a relatively long temporal duration:

> Antecedent: The mechanic kickedthe metal chair\. Duration \(long\): It took2 hoursto drive back home and return with the critical tool\. Anaphor: Afterwards, he felt foolish for having kicked themetal furniturewhen he had been frustrated\.

\[[1](https://arxiv.org/html/2608.05630#bib.bib1)\]found that readers are relatively slow to read the anaphor and less accurate in answering a follow\-up comprehension question about the antecedent:

> What piece ofmetal furnituredid the mechanic kick out of frustration?

compared to a text with a relatively short temporal duration:

> Duration \(short\): It took10 minutesto drive back home and return with the critical tool\.

The fifth and sixth factors concern the semantics of the situation models that readers build\. Thesemantic similaritybetween the anaphor and antecedent is how much they resemble each other: the greater the similarity, the better the anaphor is as a cue to memory to retrieve the antecedent from the situation model\. Returning to the example above,chairis a typical member of thefurniturecategory\[[29](https://arxiv.org/html/2608.05630#bib.bib40)\]\. Thus, the two have high semantic similarity, and this will speed resolution\. By contrast, if the antecedent had been atypical \(e\.g\.,stool\), this would have slowed resolution\[[9](https://arxiv.org/html/2608.05630#bib.bib14),[34](https://arxiv.org/html/2608.05630#bib.bib48)\]\. Next, consider if a second member of thefurniturecategory had been present in the text, one that is not the antecedent:

> Distractor: The mechanic walked past thewooden deskand headed to the door\.

If it is a typical member of the same category, as in the example above, then this causessemantic interferencein the search for the correct antecedent \(i\.e\.,chair\) in the situation model, and this will slow resolution and undermine its accuracy\. By contrast, if the non\-antecedent distractor is atypical \(e\.g\.,wooden shelf\), then this should have only a small deleterious effect\[[5](https://arxiv.org/html/2608.05630#bib.bib8),[34](https://arxiv.org/html/2608.05630#bib.bib48)\]\.

### Anaphor Resolution in LLMs

Process\-level links between language models and human comprehension are often established using surprisal, which has been shown to correlate with human reading times across a range of syntactic and semantic phenomena\[[11](https://arxiv.org/html/2608.05630#bib.bib16),[15](https://arxiv.org/html/2608.05630#bib.bib22)\]\. While surprisal\-based analyses have been widely applied at the word and sentence levels, their use in evaluating discourse\-level processes such as anaphor resolution remains limited\[[36](https://arxiv.org/html/2608.05630#bib.bib53),[16](https://arxiv.org/html/2608.05630#bib.bib24)\]\.

In early work,\[[12](https://arxiv.org/html/2608.05630#bib.bib59)\]and\[[14](https://arxiv.org/html/2608.05630#bib.bib60)\]evaluated the ability of early LLMs \(e\.g\., BERT, GPT\-2\) to resolve reflexive pronominal references to antecedents within the same sentence, with mixed results even for this simple case\. More promisingly,\[[25](https://arxiv.org/html/2608.05630#bib.bib61)\]evaluated BERT’s ability to resolve anaphors over increasing function of sentential distance\. They identified attention heads at higher layers that bridge between anaphors and their antecedents, though these connections were weaker at longer distances \(6−106\-10sentences\)\. They also used a cloze procedure to probe, at the anaphor position, the model’s best guess of the antecedent, finding low accuracy even at close sentential distances \(i\.e\.,≥3\\geq 3sentences\)\.

Recent work has examined the ability of LLMs to resolve anaphors and coreference relations, either by prompting them directly or by adapting them to existing benchmarks\. These studies show substantial variability across models, datasets, and prompt formulations, and general\-purpose LLMs often underperform specialized coreference systems\[[3](https://arxiv.org/html/2608.05630#bib.bib4),[17](https://arxiv.org/html/2608.05630#bib.bib26)\]\. Moreover, reported gains can be sensitive to dataset artifacts, lexical heuristics, and recency biases, raising questions about whether high accuracy reflects robust discourse understanding or shallow statistical cues\[[22](https://arxiv.org/html/2608.05630#bib.bib32),[8](https://arxiv.org/html/2608.05630#bib.bib12)\]\.

More importantly, high coreference accuracy does not imply human\-like processing\. Existing evaluation practices rely on inconsistent metrics and lack a standard framework for assessing LLM\-generated responses, particularly across datasets and task formulations\[[31](https://arxiv.org/html/2608.05630#bib.bib44)\]\. Standard NLP benchmarks do not manipulate or isolate the factors that cognitive science has shown to govern anaphor resolution, such as discourse prominence, distance\-based accessibility, semantic similarity, or interference from competing referents\. As a result, existing evaluations cannot distinguish between models that resolve anaphors using surface heuristics and those that exhibit sensitivity to the cognitive constraints shaping human comprehension\[[13](https://arxiv.org/html/2608.05630#bib.bib19),[30](https://arxiv.org/html/2608.05630#bib.bib41)\]\.

*Rather than proposing a new benchmark or optimizing coreference performance, the present study adopts classic cognitive science paradigms as the evaluation framework\. We ask whether LLMs exhibit sensitivity to the same discourse, situational, and semantic factors that shape human anaphor resolution, using surprisal and comprehension accuracy as complementary behavioral proxies\.*

This framing allows us to evaluate cognitive alignment beyond correct referent identification, focusing instead on whether models are sensitive to the same processing factors that affect human readers\.

![Refer to caption](https://arxiv.org/html/2608.05630v1/x1.png)Figure 1:Mean surprisal on the anaphor as a function of antecedent topicality and sentential distance \(Exp\. 1\)\. The error bars are standard errors computed across the 16 texts\.![Refer to caption](https://arxiv.org/html/2608.05630v1/x2.png)Figure 2:Mean comprehension question accuracy as a function of antecedent topicality and sentential distance \(Exp\. 1\)\. The error bars are standard errors\.
### Research Questions

We evaluate whether LLMs are sensitive to the six factors delineated above that have been shown to affect anaphor resolution in humans\. In particular, we test whether resolution is facilitated when antecedents are more accessible, due to greater discourse prominence, shorter sentential, spatial, or temporal distance, higher semantic similarity, and reduced interference from competing referents\.

We address this question using five open\-weight LLMs \(GPT\-2\-XL, LLaMa\-3\.1\-8B, Pythia\-12B, Mistral\-7B, and Mistral\-24B\)\. To assess process\-level sensitivity, we adopt the standard linking hypothesis that relates model surprisal at the anaphor to human reading times\[[11](https://arxiv.org/html/2608.05630#bib.bib16),[15](https://arxiv.org/html/2608.05630#bib.bib22),[36](https://arxiv.org/html/2608.05630#bib.bib53),[16](https://arxiv.org/html/2608.05630#bib.bib24)\]\. To assess resolution success, we also measure model accuracy on comprehension questions that probe the antecedents of anaphors after text processing\.

## Experiment 1

Experiment 1 investigated the effect of \(1\) a thematic feature, whether the antecedent is topicalized by the text’s title, and \(2\) a textual feature, the sentential distance \(i\.e\., number of intervening sentences\) between anaphors and antecedents, on anaphor resolution time and accuracy\.

### Design and Materials

The materials were from\[[23](https://arxiv.org/html/2608.05630#bib.bib33)\]\. There were 16 texts, each occurring in four versions formed by orthogonally varying two factors\. One factor was sentential distance \(near, far\), with the antecedent occurring eitherM= 5\.6 orM= 15\.2 sentences prior to the anaphor\. The other factor was antecedent topicality \(high, low\), with the antecedent either focused by the title of the text or not\. Thus, the four versions were:

- A\.near sentential distance, high antecedent topicality
- B\.near sentential distance, low antecedent topicality
- C\.far sentential distance, high antecedent topicality
- D\.far sentential distance, low antecedent topicality

### Procedure and Dependent Measures

#### Anaphor Processing\.

The surprisal of a probabilistic model in predicting the next tokeniiis−log2⁡\(pi\)\-\\log\_\{2\}\(p\_\{i\}\)wherepip\_\{i\}is the probability ofii\. A common linking hypothesis is that the greater a model’s surprisal for the next word, the longer the predicted reading time\[[11](https://arxiv.org/html/2608.05630#bib.bib16),[15](https://arxiv.org/html/2608.05630#bib.bib22),[36](https://arxiv.org/html/2608.05630#bib.bib53),[16](https://arxiv.org/html/2608.05630#bib.bib24)\]\. The surprisal for a text/version was computed as the average surprisal across the tokens making up the anaphor\. Because these values were highly variable across the 16 texts, we normalized them within each text using:

s​u​r​p​r​i​s​a​l−min⁡\(s​u​r​p​r​i​s​a​lj\)max⁡\(s​u​r​p​r​i​s​a​lj\)−min⁡\(s​u​r​p​r​i​s​a​lj\)\\frac\{surprisal\-\\min\(surprisal\_\{j\}\)\}\{\\max\(surprisal\_\{j\}\)\-\\min\(surprisal\_\{j\}\)\}wheres​u​r​p​r​i​s​a​ljsurprisal\_\{j\}denotes the set of the four surprisals for versions A\-D\. Thus, the normalized surprisals ranged from 0 to 1\. For each of the four versions A\-D, we averaged the normalized surprisals across the 16 texts, resulting in four mean surprisals\. We carried out this process for each LLM\.

#### Comprehension Question Answering\.

For each of the 16 texts, there is a comprehension question asking for the antecedent of the anaphor\. After an LLM processed a text, it was prompted with this question\. We code the generated responses by adopting an "LLM\-as\-a\-judge" paradigm using Gemini\-2\.5\-flash\-preview\-09\-2025 as the judge\. Each question had a "gold" answer provided by us\. The judge was prompted to first provide a brief justification comparing the candidate response to the gold answer and then output a binary accuracy label \(1 = accurate, 0 = inaccurate\)\. As with surprisal, for each of the four versions A\-D, we averaged the accuracy score across the 16 texts, resulting in four mean accuracies\. We repeated this process for each LLM\.

### Results and Discussion

#### Surprisal Prediction of Reading Times\.

Figure[1](https://arxiv.org/html/2608.05630#Sx1.F1)shows, for each of the LLMs, the average surprisal for each of the four text versions\. The prediction is that this should be lowest \(i\.e\., that anaphor reading times should be fastest\) when the title topicalizes the antecedent, leading readers to focus their attention on it during situation model construction, and when the sentential distance between the anaphor and antecedent is near\. Thus, average surprisal should be lowest for version A, highest for version D, and intermediate for versions B and C\. GPT\-2\-XL, Llama\-3\.1\-8B, Pythia\-12B, and Mistral\-7B showed this pattern\.

#### Comprehension Question Accuracy\.

Figure[2](https://arxiv.org/html/2608.05630#Sx1.F2)shows, for each of the LLMs, the average accuracy for each of the four text versions\. The predictions mirror those for surprisal, i\.e\., that accuracy should be highest when the title topicalizes the antecedent and when the sentential distance between the anaphor and antecedent is near\. Thus, accuracy should be highest for version A, lowest for version D, and intermediate for the other versions B and C\. GPT\-2\-XL showed this pattern\. Among the other models, only Mistral\-7B correctly ordered versions A and D\. We note that because Llama\-3\.1\-8B and Mistral\-24B performed at ceiling on the comprehension questions, this might have limited our ability to detect differences between conditions\.

## Experiment 2

![Refer to caption](https://arxiv.org/html/2608.05630v1/x3.png)Figure 3:Mean surprisal on the anaphor as a function of spatial distance and temporal duration \(Exp\. 2\)\. The error bars are standard errors\.![Refer to caption](https://arxiv.org/html/2608.05630v1/x4.png)Figure 4:Mean comprehension question accuracy as a function of spatial distance and temporal duration \(Exp\. 2\)\. The error bars are standard errors\.Experiment 2 investigated whether two contextual factors, spatial distance and temporal duration between anaphors and antecedents in readers’ situation models, affect resolution time and accuracy\. The details were the same as Experiment 1, except where otherwise noted\.

### Design and Materials

The materials were those of Experiment 1 of\[[34](https://arxiv.org/html/2608.05630#bib.bib48)\]\. There were 19 texts, and each occurred in four versions formed by orthogonally varying the spatial distance \(near, far\) and temporal duration \(short, long\) between the anaphor and antecedent in the reader’s situation model:

- A\.near spatial distance, short temporal duration
- B\.near spatial distance, long temporal duration
- C\.far spatial distance, short temporal duration
- D\.far spatial distance, long temporal duration

These factors were manipulated in theM= 15\.1 \(SD= 2\.4\) sentences that separated anaphors from their antecedents\. The same categorical anaphor \(e\.g\.,metal furniture\) and typical antecedent \(e\.g\.,metal chair\) were used for each of the four versions A\-D\.

### Results and Discussion

#### Surprisal Prediction of Reading Times\.

Figure[3](https://arxiv.org/html/2608.05630#Sx3.F3)shows the average surprisal of each of the five models on each of the four text versions\. The prediction is that this should be lowest \(i\.e\., anaphor resolution fastest\) for version A, highest \(i\.e\., anaphor resolution slowest\) for version D, and intermediate for versions B and C\. Llama\-3\.1\-8B and Mistral\-7B show the predicted pattern\.

#### Comprehension Question Accuracy\.

Figure[4](https://arxiv.org/html/2608.05630#Sx3.F4)shows the average accuracies\. The predictions parallel those for surprisal: the highest average accuracy is expected for version A, the lowest for version D, and intermediate values for versions B and C\. Only GPT\-2\-XL shows this pattern\. That said, all of the other models correctly order versions A and D\.

## Experiment 3

![Refer to caption](https://arxiv.org/html/2608.05630v1/x5.png)Figure 5:Mean surprisal on the anaphor as a function of the semantic overlap and semantic interference factors \(Exp\. 3\)\. The error bars are standard errors\.![Refer to caption](https://arxiv.org/html/2608.05630v1/x6.png)Figure 6:Mean comprehension question accuracy as a function of the semantic overlap and semantic interference factors \(Exp\. 3\)\. The error bars are standard errors\.Experiment 3 investigated whether two semantic factors, the semantic overlap between the anaphor and antecedent and the semantic interference from non\-antecedent distractors, affect resolution time and accuracy\. Except where otherwise noted, the details are the same as previous experiments\.

### Design and Materials

The materials were those of Experiment 2 of\[[34](https://arxiv.org/html/2608.05630#bib.bib48)\]: 19 texts, each occurring in four versions formed by orthogonally varying the semantic overlap between the antecedent and anaphor \(high, low\) and the semantic interference caused by a non\-antecedent distractor \(low, high\)\. The anaphors were categorical \(e\.g\.,metal furniture\)\. Following prior work\[[5](https://arxiv.org/html/2608.05630#bib.bib8),[9](https://arxiv.org/html/2608.05630#bib.bib14),[34](https://arxiv.org/html/2608.05630#bib.bib48)\], high semantic overlap was operationalized by antecedents that were typical members of the category \(e\.g\.,chair\) and low semantic overlap by antecedents that were atypical \(e\.g\.,stool\)\. Low semantic interference was operationalized by non\-antecedent distractors that are atypical members and high semantic interference by non\-antecedent distractors that are typical\. Thus, the four versions were:

- A\.typical antecedent, atypical distractor
- B\.typical antecedent, typical distractor
- C\.atypical antecedent, atypical distractor
- D\.atypical antecedent, typical distractor

The base texts were the version D \(far spatial distance, long temporal duration\) texts from Experiment 2\. The four versions of each text shared the same categorical anaphor\. Only the antecedent and the non\-antecedent distractors were varied\.

### Results and Discussion

#### Surprisal Prediction of Reading Times\.

The average surprisal of the five models on the four text versions is shown in Figure[5](https://arxiv.org/html/2608.05630#Sx4.F5)\. The prediction is that the anaphor resolution is fastest for version A, slowest for D, and intermediate for versions B and C\. None of the models show the predicted pattern, although GPT\-2\-XL correctly orders versions A and D\.

#### Comprehension Question Accuracy\.

The average accuracy of the five models on the four text versions is shown in Figure[6](https://arxiv.org/html/2608.05630#Sx4.F6)\. The predictions mirror those for surprisal on the anaphor, with the highest average accuracy expected for version A, the lowest for version D, and intermediate values for versions B and C\. Only Llama\-3\.1\-8B shows the predicted pattern, although its overall accuracies are quite low\.

## General Discussion

Table 1:Summary of where LLMs show human\-like directional patterns in anaphor resolution across discourse, situation\-model, and semantic factors\.Cognitive science studies have identified factors that affect the speed and success of anaphor resolution\. This study examined whether the same factors affect anaphor resolution in LLMs\. Positive results would constitute evidence that LLMs can serve as cognitive models of human anaphor resolution, and perhaps of language understanding more generally\.

Experiment 1 investigated the effects of \(1\) antecedent topicality and \(2\) sentential distance\. The remaining experiments focused on the reader’s situation model of the text\. Experiment 2 investigated the contextual effects of \(3\) spatial distance and \(4\) temporal duration; Experiment 3 investigated the semantic effects of \(5\) antecedent similarity and \(6\) interference from non\-antecedent distractors\. Two indices of model performance, surprisal on the anaphor and accuracy on a comprehension question about the antecedent, were compared with the corresponding human measures of reading time and accuracy\. The models considered were GPT\-2\-XL, Llama\-3\.1\-8B, Pythia, Mistral\-24B, and Mistral\-7B\. To varying degrees, the models showed the six effects documented in the literature, whether in their surprisal values or comprehension accuracies\. The model with the greatest overall alignment to human anaphor resolution was perhaps Mistral\-7B, although it failed to account for the effects of \(5\) antecedent similarity and \(6\) interference in Experiment 3\. GPT\-2\-XL achieved surprisingly respectable overall alignment given that it is the oldest and smallest of the models\.

*A notable pattern across experiments is that LLMs more reliably exhibited human\-like sensitivity to discourse prominence and distance\-based factors than to semantic similarity and interference\.*This asymmetry is informative\. Effects of sentential, spatial, and temporal distance reflect graded accessibility, which may be approximated by recency biases and attention dynamics in sequential language models\. In contrast, semantic interference effects require structured competition among partially overlapping representations during memory retrieval, a process central to cognitive theories of anaphor resolution but not explicitly implemented in current LLM architectures\. From this perspective, partial alignment does not undermine the present findings; rather, it helps localize where LLMs diverge from human discourse processing\.

Thus, we have some evidence of cognitive alignment between LLMs, especially Mistral\-7B and GPT\-2\-XL, and human performance, suggesting their potential utility as cognitive science models\. Still, work remains to be done\. For example, although the profiles of these models’ comprehension question accuracies across passage versions were generally human\-like, their absolute performance was quite low\.

The largest limitation of the current study is the lack of items \(i\.e\., texts\), which precluded running statistical analyses to better establish the informal trends present in the model results\. Here, we were limited by the relatively small size of cognitive science studies relative to ML studies and the general unavailability of materials and human datasets for classic studies of anaphor resolution\. It would have been easy to generate multiple ’participants’ by increasing the temperature parameter, running each model multiple times, and computing statistics across these samples\. However, we question the validity of this approach\. There is no reason to believe that such simulated "participants" vary in the same ways that humans do: in their working memory capacity, executive function, domain knowledge, reading skill, etc\. For this reason, we limited ourselves to providing descriptive, summary statistics and cautiously interpreting the observed trends\. We also recognize the limitations of our LLM\-as\-a\-judge technique for scoring comprehension accuracy\. Although it provides a reasonable heuristic, the binary scoring and reliance on an automatic judge introduce known biases and can be different from human judgments\. It is best to treat the reported accuracies as approximations\.

One goal for future research is to examine whether other factors that affect anaphor resolution performance in humans also affect LLMs\. Some of these are relatively subtle, such as whether an antecedent and anaphor belong to the same event within a narrative text or whether they are in different events separated by a boundary\[[32](https://arxiv.org/html/2608.05630#bib.bib45)\]\. Whether or not LLMs are sensitive to such factors is an important question for cognitive scientists evaluating their potential as models of human language understanding\. Examination of a broad range of factors known to affect human anaphor resolution might also guide NLP researchers as they design the new benchmarks for measuring the coreference resolution abilities of LLMs\[[8](https://arxiv.org/html/2608.05630#bib.bib12),[18](https://arxiv.org/html/2608.05630#bib.bib28)\]\.

## References

- \[1\]A\. Anderson, S\. C\. Garrod, and A\. J\. Sanford\(1983\)The accessibility of pronominal antecedents as a function of episode shifts in narrative text\.35,pp\. 427–440\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p3.1),[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p5.1)\.
- \[2\]BabyLM Organizers\(2025\-11\)Findings of the third BabyLM challenge: accelerating language modeling research with cognitively plausible data\.InProceedings of the First BabyLM Workshop,L\. Charpentier, L\. Choshen, R\. Cotterell, M\. O\. Gul, M\. Y\. Hu, J\. Liu, J\. Jumelet, T\. Linzen, A\. Mueller, C\. Ross, R\. S\. Shah, A\. Warstadt, E\. G\. Wilcox, and A\. Williams \(Eds\.\),Suzhou, China,pp\. 399–420\.External Links:[Link](https://aclanthology.org/2025.babylm-main.28/),[Document](https://dx.doi.org/10.18653/v1/2025.babylm-main.28)Cited by:[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1)\.
- \[3\]E\. Cambria\(2025\)Semantics processing\.Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p3.1)\.
- \[4\]H\. H\. Clark and C\. J\. Sengul\(1979\)In search of referents for nouns and pronouns\.7,pp\. 35–41\.External Links:[Document](https://dx.doi.org/10.3758/BF03196932)Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p2.1)\.
- \[5\]A\. T\. Corbett\(1984\)Prenominal adjectives and the disambiguation of anaphoric nouns\.23,pp\. 683–695\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p9.1),[Design and Materials](https://arxiv.org/html/2608.05630#Sx4.SSx1.p1.1)\.
- \[6\]M\. Daneman and P\. A\. Carpenter\(1980\)Individual differences in working memory and reading\.19,pp\. 450–466\.Cited by:[Introduction](https://arxiv.org/html/2608.05630#Sx1.p3.1)\.
- \[7\]J\. Devlin, M\.\-W\. Chang, K\. Lee, and K\. Toutanova\(2019\)BERT: pre\-training of deep bidirectional transformers for language understanding\.Cited by:[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1)\.
- \[8\]Y\. Gan, M\. Poesio, and J\. Yu\(2024\-05\)Assessing the capabilities of large language models in coreference: an evaluation\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),N\. Calzolari, M\. Kan, V\. Hoste, A\. Lenci, S\. Sakti, and N\. Xue \(Eds\.\),Torino, Italia,pp\. 1645–1665\.External Links:[Link](https://aclanthology.org/2024.lrec-main.145/)Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p3.1),[General Discussion](https://arxiv.org/html/2608.05630#Sx5.p6.1)\.
- \[9\]S\. Garrod and A\. J\. Sanford\(1977\)Interpreting anaphoric relations: the integration of semantic information while reading\.16,pp\. 77–90\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p7.1),[Design and Materials](https://arxiv.org/html/2608.05630#Sx4.SSx1.p1.1)\.
- \[10\]A\. Grattafioriet al\.\(2024\)The llama 3 herd of models\.Note:arXivExternal Links:[Link](https://doi.org/10.48550/arXiv.2407.21783)Cited by:[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1)\.
- \[11\]J\. Hale\(2001\)A probabilistic earley parser as a psycholinguistic model\.InProceedings of the 2nd Meeting of NAACL,Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p1.1),[Research Questions](https://arxiv.org/html/2608.05630#Sx1.SSx3.p2.1),[Anaphor Processing\.](https://arxiv.org/html/2608.05630#Sx2.SSx2.SSSx1.p1.4)\.
- \[12\]J\. Hu, S\. Y\. Chen, and R\. Levy\(2020\-01\)A closer look at the performance of neural language models on reflexive anaphor licensing\.InProceedings of the Society for Computation in Linguistics 2020,New York, New York,pp\. 323–333\.External Links:[Link](https://aclanthology.org/2020.scil-1.39/)Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p2.2)\.
- \[13\]A\. A\. Ivanova\(2025\)How to evaluate the cognitive abilities of llms\.Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p4.1),[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1),[Introduction](https://arxiv.org/html/2608.05630#Sx1.p4.1)\.
- \[14\]S\. Lee and S\. Schuster\(2022\-02\)Can language models capture syntactic associations without surface cues? a case study of reflexive anaphor licensing in English control constructions\.InProceedings of the Society for Computation in Linguistics 2022,online,pp\. 206–211\.External Links:[Link](https://aclanthology.org/2022.scil-1.18/)Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p2.2)\.
- \[15\]R\. Levy\(2008\)Expectation\-based syntactic comprehension\.106,pp\. 1126–1177\.Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p1.1),[Research Questions](https://arxiv.org/html/2608.05630#Sx1.SSx3.p2.1),[Anaphor Processing\.](https://arxiv.org/html/2608.05630#Sx2.SSx2.SSSx1.p1.4)\.
- \[16\]A\. Li, T\. Cai, X\. Feng, S\. Narang, A\. Peng, R\. S\. Shah, and S\. Varma\(2024\)Incremental comprehension of garden\-path sentences by large language models: semantic interpretation, syntactic re\-analysis, and attention\.InProceedings of the 46th Annual Conference of the Cognitive Science Society,pp\. 6069–6076\.Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p1.1),[Research Questions](https://arxiv.org/html/2608.05630#Sx1.SSx3.p2.1),[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1),[Introduction](https://arxiv.org/html/2608.05630#Sx1.p2.1),[Anaphor Processing\.](https://arxiv.org/html/2608.05630#Sx2.SSx2.SSSx1.p1.4)\.
- \[17\]X\. Liu, S\. Deng, M\. Wang, Z\. Dong, L\. Dai, J\. Li, and R\. Nong\(2025\)Enhancing coreference resolution with pretrained language models: bridging the gap between syntax and semantics\.Note:arXivExternal Links:[Link](https://doi.org/10.48550/ARXIV.2504.05855)Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p3.1)\.
- \[18\]K\. Manikantan, M\. Tapaswi, V\. Gandhi, and S\. Toshniwal\(2024\)IdentifyMe: a challenging long\-context mention resolution benchmark\.Note:arXivCited by:[General Discussion](https://arxiv.org/html/2608.05630#Sx5.p6.1)\.
- \[19\]D\. S\. McNamara and J\. Magliano\(2009\)Toward a comprehensive model of comprehension\.Note:In B\. Ross \(Ed\.\), The psychology of learning and motivationCited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p3.1),[Introduction](https://arxiv.org/html/2608.05630#Sx1.p3.1)\.
- \[20\]Mistral AI Team\(2025\-01\-30\)Mistral small 3\.Note:Accessed: 2026\-01\-30External Links:[Link](https://mistral.ai/news/mistral-small-3)Cited by:[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1)\.
- \[21\]D\. G\. Morrow, S\. L\. Greenspan, and G\. H\. Bower\(1987\)Accessibility and situation models in narrative comprehension\.26,pp\. 165–187\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p3.1)\.
- \[22\]M\. Nováket al\.\(2025\)Findings of the fourth shared task on multilingual coreference resolution\.InProceedings of the CRAC Shared Task,External Links:[Document](https://dx.doi.org/10.18653/v1/2025.crac-1.9)Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p3.1)\.
- \[23\]E\. J\. O’Brien\(1987\)Antecedent search processes and the structure of text\.13,pp\. 278–290\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p2.1),[Design and Materials](https://arxiv.org/html/2608.05630#Sx2.SSx1.p1.1)\.
- \[24\]OpenAI\(2025\-08\)Introducing gpt\-5\.Note:Accessed: 2026\-01\-30External Links:[Link](https://openai.com/index/introducing-gpt-5/)Cited by:[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1)\.
- \[25\]O\. Pandit and Y\. Hou\(2021\-06\)Probing for bridging inference in transformer language models\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,Online,pp\. 4153–4163\.External Links:[Link](https://aclanthology.org/2021.naacl-main.327/),[Document](https://dx.doi.org/10.18653/v1/2021.naacl-main.327)Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p2.2)\.
- \[26\]S\. T\. Piantadosiet al\.\(2024\)Why concepts are \(probably\) vectors\.28,pp\. 844–856\.Cited by:[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1)\.
- \[27\]A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. Sutskever\(2019\)Language models are unsupervised multitask learners\.Cited by:[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1)\.
- \[28\]M\. Rinck and G\. H\. Bower\(1995\)Anaphora resolution and the focus of attention in situation models\.34,pp\. 110–131\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p3.1)\.
- \[29\]E\. Rosch, C\. B\. Mervis, W\. D\. Gray, D\. M\. Johnson, and P\. Boyes\-Braem\(1976\)Basic objects and natural categories\.9,pp\. 382–440\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p7.1)\.
- \[30\]R\. S\. Shah and S\. Varma\(2025\)The potential and the pitfalls of using pre\-trained language models as cognitive science theories\.Note:arXivExternal Links:[Link](https://arxiv.org/abs/2501.12651)Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p4.1),[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1),[Introduction](https://arxiv.org/html/2608.05630#Sx1.p4.1)\.
- \[31\]C\. Talukdar and M\. Rahman\(2025\)Coreference resolution in machine learning: a survey\.In2025 IEEE Guwahati Subsection Conference \(GCON\),External Links:[Document](https://dx.doi.org/10.1109/GCON65540.2025.11173320)Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p4.1)\.
- \[32\]A\. N\. Thompson and G\. A\. Radvansky\(2016\)Event boundaries and anaphoric reference\.23,pp\. 849–856\.Cited by:[General Discussion](https://arxiv.org/html/2608.05630#Sx5.p6.1)\.
- \[33\]T\. A\. van Dijk and W\. Kintsch\(1983\)Strategies of discourse comprehension\.Academic Press\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p3.1),[Introduction](https://arxiv.org/html/2608.05630#Sx1.p3.1)\.
- \[34\]S\. Varma and A\. Janssen\(2019\)The structure of situation models as revealed by anaphor resolution\.72,pp\. 104–115\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p3.1),[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p7.1),[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p9.1),[Design and Materials](https://arxiv.org/html/2608.05630#Sx3.SSx1.p1.1),[Design and Materials](https://arxiv.org/html/2608.05630#Sx4.SSx1.p1.1)\.
- \[35\]Y\. Wanget al\.\(2024\)MMLU\-pro: a more robust and challenging multi\-task language understanding benchmark\.Note:arXivExternal Links:[Link](https://doi.org/10.48550/arXiv.2406.01574)Cited by:[Introduction](https://arxiv.org/html/2608.05630#Sx1.p1.1)\.
- \[36\]E\. Wilcox, J\. Gauthier, J\. Hu, P\. Qian, and R\. P\. Levy\(2020\)On the predictive power of neural language models for human real\-time comprehension behavior\.InProceedings of the 42nd Annual Meeting of the Cognitive Science Society,pp\. 1707–1713\.Cited by:[Anaphor Resolution in LLMs](https://arxiv.org/html/2608.05630#Sx1.SSx2.p1.1),[Research Questions](https://arxiv.org/html/2608.05630#Sx1.SSx3.p2.1),[Introduction](https://arxiv.org/html/2608.05630#Sx1.p2.1),[Anaphor Processing\.](https://arxiv.org/html/2608.05630#Sx2.SSx2.SSSx1.p1.4)\.
- \[37\]R\. A\. Zwaan\(1996\)Processing narrative time shifts\.22,pp\. 1196–1207\.Cited by:[Cognitive Science Studies of Anaphor Resolution](https://arxiv.org/html/2608.05630#Sx1.SSx1.p3.1)\.

Similar Articles

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

arXiv cs.CL

This paper introduces a register-aware linguistic evaluation framework to assess how human-like large language models (LLMs) are by comparing the distribution of 67 lexico-grammatical features between human and LLM-generated texts using Maximum Mean Discrepancy. Experiments across seven instruction-tuned open-source models and five registers show that no model perfectly matches human baselines, and closeness to human language varies by register rather than model size.

Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning

arXiv cs.CL

This paper investigates whether assigning personas to large language models induces human-like motivated reasoning, finding that persona-assigned LLMs show up to 9% reduced veracity discernment and are up to 90% more likely to evaluate scientific evidence in ways congruent with their induced political identity, with prompt-based debiasing largely ineffective.