The Machine's Internal Clock: Do LLMs Share Human Temporal Illusions?
Summary
This paper explores if large language models exhibit human-like temporal illusions by testing them on a narrative benchmark, finding that LLMs often choose scenarios consistent with psychology literature rather than mimicking human perception.
View Cached Full Text
Cached at: 08/18/26, 10:07 AM
# The Machine’s Internal Clock: Do LLMs Share Human Temporal Illusions?
Source: [https://arxiv.org/html/2608.15394](https://arxiv.org/html/2608.15394)
###### Abstract
Human perception of time is subjective\. Well\-documented temporal illusions show that the brain relies on context and relational cues for judging duration instead of tracking elapsed time directly\. Prior studies established these effects with visual and auditory stimuli\. Existing LLM evaluations of temporal perception focus on estimating event durations or multi\-step temporal reasoning\. In this work, we investigate whether written narratives alone can evoke human temporal illusions, using a new benchmark of 6,684 narrative pairs spanning five illusions\. We find that human readers \(60 participants\) prefer expected scenarios in only two of the five illusions, those where the manipulation is directly visible in text rather than requiring readers to internally simulate duration\. We evaluate 14 LLMs on the same benchmark\. Surprisingly, we find that models pick the literature\-predicted scenario across four of the five illusions, diverging from human behavior\. Reasoning traces show that∼\\sim70% of responses explicitly evoke psychology research, suggesting that this alignment is consistent with retrieval of published findings rather than human\-like temporal biases\.
## 1Introduction
Figure 1:We created narrative comparisons to elicit temporal illusions discovered in psychology literature\. The LLM thinking trace retrieves literature about the temporal illusion and agrees with the literature\. Yet, human annotators do not have a consensus on the illusion in the narrative form\. The example demonstrates the filled\-duration illusion\.Our perception of time is subjective; event density, novelty, and association can affect the perceived speed of events\. The theory of an internal clock does not explain such human\-like temporal understanding\([37](https://arxiv.org/html/2608.15394#bib.bib1);[7](https://arxiv.org/html/2608.15394#bib.bib32)\)\. Alternative frameworks, such as retrospective timing or purely cognitive models, discount an internal clock, and instead suggest that we construct our sense of time from experience\([7](https://arxiv.org/html/2608.15394#bib.bib32);[5](https://arxiv.org/html/2608.15394#bib.bib2)\)\. Temporal illusions, such as emotional time dilation and the oddball effect, illustrate the gaps between the objective passage of time, and its subjective experience\. They show that prospective judgments \(made while an event unfolds\) rely on expectation and tend to be longer and more variable than retrospective ones\([6](https://arxiv.org/html/2608.15394#bib.bib8)\)\.*We ask: given the purported human\-like capabilities of large language models \(LLMs\), do their responses reflect a subjective or an objective view of time?*
Time\-based questions are generally difficult for LLMs\. Recent work shows that they struggle with both mathematical and relative time\-based tasks, scoring between 60\-80% depending on the question format\([31](https://arxiv.org/html/2608.15394#bib.bib31);[14](https://arxiv.org/html/2608.15394#bib.bib29);[39](https://arxiv.org/html/2608.15394#bib.bib27),*inter alia*\)\. However, these evaluations are about objective temporal reasoning, and do not address the subjective and context\-dependent nature of temporal perception\. A critical gap exists in current AI research about the subjectivity of temporal comprehension, specifically the expansion or compression of time perception\.
To address this gap, we present a new benchmark that uses temporal illusions to help study the contextual impact on temporal perception\. The psychology literature studies temporal illusions using visual and auditory stimuli rather than textual ones\. We construct a standard benchmark to assess LLMs by adapting these tests to a narrative question\-answer format\. \(See Figure[1](https://arxiv.org/html/2608.15394#S1.F1)\.\) Our benchmark spans five illusions, translated into 6,684 narrative pairs generated from 111 templates\. Narrative framing can separate what a character in a situation experiences from how a reader processes it\. To capture this, the benchmark evaluates both perspectives separately, allowing us to test whether models track a character’s subjective time, or a readers, or both\.
Using this benchmark, we conduct a human study to investigate whether written narratives alone preserve temporal illusions\. Human readers reliably perceive only two of the five illusions through text\. We also examine whether 14 LLMs produce temporal judgments that align with the annotator behavior or with the published temporal distortions\. We found that the evaluated LLMs replicate the pattern predicted in the literature across four of the five illusions\. Our analysis of their reasoning traces suggests that the effect is consistent with explicit retrieval of published literature\.
In summary, our contributions are:
- •We present a new narrative benchmark that uses five temporal illusions to study whether LLMs produce temporal judgments that align with previously studied human distortions of time\.
- •Our human study with the benchmark reveals that only two of the five illusions transfer to a purely textual narrative format\.
- •We find that LLMs align with the expected distortions in the literature, and reasoning traces frequently invoke psychology literature explicitly, suggesting that their alignment reflects literature retrieval rather than human\-like temporal perception\.
## 2Background and Related Work
### 2\.1Human Temporal Perception in Psychology
Psychological research on temporal illusions has studied visual or auditory stimuli over short durations\([37](https://arxiv.org/html/2608.15394#bib.bib1);[15](https://arxiv.org/html/2608.15394#bib.bib7);[36](https://arxiv.org/html/2608.15394#bib.bib9);[40](https://arxiv.org/html/2608.15394#bib.bib11);[9](https://arxiv.org/html/2608.15394#bib.bib12)\)\. In this work, we focus on five illusions, selected for their reliance on attention, density, novelty, and association\([13](https://arxiv.org/html/2608.15394#bib.bib22)\), which make them suitable for a narrative representation\. They fall into two categories:
\(a\)Arousal\-based illusions \(Emotional Time Dilation and the Oddball Effect\), and\(b\)Information\-processing illusions \(Filled\-Duration, Temporal Order, and Familiarity\-Duration\)\.
*Emotional Time Dilation*arises from physiological arousal and attention; researchers typically study it using visual cues of high\-arousal stimuli, such as pictures of phobias or physical mutilation\. For example, individuals with spider phobias tend to underestimate neutral stimuli versus spider\-related imagery\([37](https://arxiv.org/html/2608.15394#bib.bib1)\)\. Since the distortion is tied to arousal and attention, the phenomenon extends to non\-emotional contexts\. The*Oddball Effect*demonstrates this: when an unexpected stimulus appears within a series of repeated expected stimuli, the duration of the unique item seems longer\([26](https://arxiv.org/html/2608.15394#bib.bib21)\)\. The novelty of the oddball relative to the repeating standard items drives the effect\([4](https://arxiv.org/html/2608.15394#bib.bib20);[26](https://arxiv.org/html/2608.15394#bib.bib21)\)\.
Stimulus density and complexity affect illusions based on processing information\. Intervals containing discrete elements, such as tones and visual markers, are perceived as longer than empty intervals of the same duration\([36](https://arxiv.org/html/2608.15394#bib.bib9);[40](https://arxiv.org/html/2608.15394#bib.bib11)\)\. This is the*Filled\-Duration Effect*\. As the discrete components within an interval increase, the subjective duration also increases\([9](https://arxiv.org/html/2608.15394#bib.bib12);[28](https://arxiv.org/html/2608.15394#bib.bib16)\)\. These studies covered short durations\. For long periods, when these elements are structured, such as music, time may be thought of as shorter, suggesting that the "flow" of organized complexity can compress time over long durations\([12](https://arxiv.org/html/2608.15394#bib.bib10)\)\.
Neural encoding causes the*Familiarity\-Duration Effect*\. Familiar items cause a shorter neural response and subjective time\([20](https://arxiv.org/html/2608.15394#bib.bib15);[2](https://arxiv.org/html/2608.15394#bib.bib17);[28](https://arxiv.org/html/2608.15394#bib.bib16)\); the brain processes familiar items more efficiently\([30](https://arxiv.org/html/2608.15394#bib.bib19)\)\. As a person’s physical space becomes more familiar, the mental map of space expands while the estimate of time taken to traverse it contracts\([17](https://arxiv.org/html/2608.15394#bib.bib18)\)\.*Temporal Order Effects*prioritize the sequence of events over the durations\. As the brain diverts resources, the mental effort required to process sequence events decreases the accuracy of duration judgment\([8](https://arxiv.org/html/2608.15394#bib.bib14)\)\.
### 2\.2Existing LLM Benchmarks
Recent work has drawn from cognitive psychology to study LLM behavior\([3](https://arxiv.org/html/2608.15394#bib.bib6)\)\. Especially relevant is the line of work on theory of mind, which examines whether models can reason about beliefs, perspectives, and mental states\([18](https://arxiv.org/html/2608.15394#bib.bib5);[29](https://arxiv.org/html/2608.15394#bib.bib4)\)\.
Existing temporal benchmarks are either factual, mathematical, or reasoning\-based\. Factual temporal questions treat time as an attribute of historical knowledge\. They require models to “look up” timestamped data, such as identifying “George Washington’s specific role in 1777”\([11](https://arxiv.org/html/2608.15394#bib.bib25)\)or “the professional team Messi played for in 2010”\([33](https://arxiv.org/html/2608.15394#bib.bib26)\)\. Other benchmarks are based on temporal factual extractions\([14](https://arxiv.org/html/2608.15394#bib.bib29);[42](https://arxiv.org/html/2608.15394#bib.bib30)\), but measure a model’s retention and not its perception of duration\. Math\-based temporal tasks focus on arithmetic operations within calendar or clock systems, such as modular arithmetic on hours and minutes\([31](https://arxiv.org/html/2608.15394#bib.bib31)\)\.
Reasoning\-based temporal questions require models to infer the frequency, duration, or plausibility of events from their temporal context\. They test “temporal common sense” and the identification of plausible causal relationships\([43](https://arxiv.org/html/2608.15394#bib.bib24);[31](https://arxiv.org/html/2608.15394#bib.bib31);[39](https://arxiv.org/html/2608.15394#bib.bib27)\)\. These ask for a guess of objective time to subjective questions \(e\.g\., ‘The chairman said the deal remains of substantial benefit\. How long did the chairman speak?’\), or ask the model to understand likely social constraints \(e\.g\., determine why a character like Amy would choose to ‘start laundry early in the morning every weekend’\)\([31](https://arxiv.org/html/2608.15394#bib.bib31)\)\.
The most closely related work on subjective measures of time is by[10](https://arxiv.org/html/2608.15394#bib.bib28), who systematically evaluate temporal relativity in LLMs\. Unlike other benchmarks, the study recognizes the relationship between time perception and context using probes such as ‘How long does a stressful event feel?’\. This changes the evaluation from objective toward subjective judgments\. Perceived duration depends on contextual factors: people may experience the same interval differently depending on their goals, expectations, and level of engagement\. For example, an hour\-long exam may feel longer to an unprepared student than to one who is well prepared and deeply engaged\.
All these datasets assume a real\-world ground truth, but many subjective, time\-based questions lack fixed answers\. Human reasoning is inherently relative, and comparisons between experiences, rather than absolute values, shape our perception of time\. Instead, in this work, we ask for comparisons of two objectively equal durations\. Psychological literature establishes our ground truth, which more closely studies how temporal distortions are triggered and measured\. Our human evaluation addresses whether readers can experience temporal illusions through written narratives\. Our LLM evaluation determines whether models align with the human time distortions discussed in existing literature\.
## 3Benchmark Design
### 3\.1Design Principles
Our benchmark translates psychology experiments into narrative prompts, mimicking how the human brain stretches and compresses intervals through writing\. Each template consists of a paired control and experimental condition that differ in a single targeted manipulation\. The experimental condition corresponds to the scenario that, based on prior psychological literature, is expected to be perceived as longer for the character\. Figure[1](https://arxiv.org/html/2608.15394#S1.F1)shows an example\.
Since text length could affect perceived duration, we write the paired conditions to match closely in length and narrative structure\. Where the question design allows, templates include both longer and shorter versions to test for this effect\. We ask the reader to judge duration both from the character’s perspective \(lived time\) and the reader’s perspective \(observed/processed time\)\.
We split each illusion into subtypes, and each subtype targets one mechanism identified in the source literature\. We separate illusions as described in §[2](https://arxiv.org/html/2608.15394#S2)\. For example, we define Emotional Time Dilation as a meaningful, high\-arousal experience that sustains top\-down attention, whereas the Oddball Effect results from brief, bottom\-up attentional capture by a stimulus that is merely different from the preceding standards without emotional weight\.
Appendix[A\.1](https://arxiv.org/html/2608.15394#A1.SS1)provides full details of the design process for each illusion \(representative prompt templates and examples\)\. Here we describe one illusion\.
#### Arousal\-Based Illusions: A Worked Example
We use Emotional Time Dilation as a running example of our design process, including how we distinguish it from the closely related Oddball Effect below\.
The difference between emotional time dilation and the oddball effect is what captures attention and how strongly\. Heightened arousal \(events that feel important, stressful, or unexpected\) drives emotional time dilation\. These situations create top\-down attention, meaning internal states such as concern, urgency, or significance guide attention\. Our design of the oddball effect is bottom\-up novelty\. A stimulus stands out within a predictable sequence and briefly captures attention simply because it is different, not because it carries emotional weight or consequence\. We define emotional time dilation as meaningful, high\-arousal experiences that sustain attention, whereas the oddball effect results from brief, stimulus\-driven attention without deeper emotional impact\.
Within this framing, we define emotional time dilation via three mechanisms that cause sustained top\-down attention:
- AStakes and Expectations: we present a neutral stimulus \(e\.g\., a flickering reflection\) in both low\-stakes and high\-stakes contexts, testing how potential consequences alter engagement with otherwise identical stimuli\.
- BInterval Intensity: we vary cognitive demand in an empty interval by comparing high\-focus tasks with low\-demand activities\. When cognitive resources are heavily engaged, fewer remain available for tracking time, altering how long the interval feels\([8](https://arxiv.org/html/2608.15394#bib.bib14)\)\.
- CEmotional Arousal: we pair a neutral baseline with a sudden or emotional event while holding context and stakes constant, testing whether involuntary attentional capture expands perceived duration\.
By holding the objective action constant while varying the emotional context, we test whether readers perceive the same event as lasting longer when the scene feels more intense\. The remaining four illusions follow the same pattern: the mechanism from the source literature determines the subtypes, and each subtype becomes a matched, length\-controlled pair\.
Figure 2:This example shows a template and instantiation of a pair of the Stakes and Expectations subcategory of Emotional Time Dilation\.
### 3\.2Dataset Creation
Using the design principles created for each illusion type, we create and populate slot\-based templates with fillers using a Jinja\-based generation framework\.111https://github\.com/utahnlp/madlibsWe design our collection of fillers \(stimuli\) for contextual variation while preserving the structure of each illusion type\. Refer to Figure[2](https://arxiv.org/html/2608.15394#S3.F2)as an example\. We wrote the initial set of 10 fillers per category manually, expanded it to 20 using GPT\-4o\([22](https://arxiv.org/html/2608.15394#bib.bib34)\), and subsequently reviewed and edited the set by hand\. We sampled character names from the top 100 most common names in the United States\([38](https://arxiv.org/html/2608.15394#bib.bib3)\)\. We generated 2000 examples per illusion type by randomly sampling from the available templates, for 10000 candidate questions\. We also manually created a subset of 250 “golden” cases \(50 per illusion type\), edited by hand to ensure clarity and representation of each illusion\.
To reduce stylistic noise and increase grammatical coherence, we paraphrased each example three times using Qwen3\-32B\([35](https://arxiv.org/html/2608.15394#bib.bib33)\), few\-shot prompted with the “golden” cases \(Appendix[A](https://arxiv.org/html/2608.15394#A1)\)\. From these variants, we selected the version that most closely matched in length between control and experimental conditions, minimizing word\-count differences that could bias perceived duration judgments\. We removed any pair with a length difference greater than 7 words, creating 6684 examples \(Table[2](https://arxiv.org/html/2608.15394#A1.T2), appendix\)\.
While the dataset is not perfectly balanced after filtering, each category retains a sufficiently large sample size for analysis \(Table[3](https://arxiv.org/html/2608.15394#A1.T3)\)\. Because we generate items from templates, examples within a template family are not statistically independent\. So the effective sample size is smaller than the raw counts\. We will release the complete slot fillers and associated scripts publicly upon publication\.
## 4Experimental Setup
With both human and LLM evaluations, we examine whether behavior aligns with human sensitivity to narrative time perception\. Specifically, we ask:
- RQ1:Do humans perceive differences consistent with known time dilation effects across narrative templates?
- RQ2:Do language models replicate known temporal illusions in psychology literature? If so, do they show stronger alignment with some illusions than others?
- RQ3:Do models distinguish between character and reader perspectives in time perception?
We address these questions via two complementary evaluations, as described below\.
### 4\.1Human Evaluation
We recruited participants for the human evaluation using Prolific\([25](https://arxiv.org/html/2608.15394#bib.bib35)\)\. Eligibility criteria, enforced using Prolific’s prescreening tools, required participants to be at least 18 years old, reside in the United States, identify English as their primary language, have completed at least one prior study on Prolific, and maintain an approval rate of 90% or higher\. Data collected through Prolific were fully anonymous\. We administered an optional demographic questionnaire at the end of the survey \(age range and highest education level\), shown in Figure[23](https://arxiv.org/html/2608.15394#A2.F23)of Appendix[B\.1](https://arxiv.org/html/2608.15394#A2.SS1)\. We compensated participants $10\.04 per hour\. We recruited 60 participants completing 25 questions in∼8\\sim 8minutes\. We sampled questions from a curated set of 250 “golden” examples\. We chose the sample size to balance statistical power \(given multiple judgments per participant\) against participant burden and cost, giving60×25=150060\\times 25=1500responses across the design\.
We instructed participants to rely on their first intuition and respond quickly\. For each question, participants evaluated randomized scenarios from one of two perspectives\. The first, the character’s perspective, asked:“Which event felt longer for the character experiencing it?”The second, the reader’s perspective, asked:“Which event felt longer to read and process?”\. Our institution’s Institutional Review Board reviewed and approved this study\. For additional details on survey instructions, see Appendix[A](https://arxiv.org/html/2608.15394#A1)\.
### 4\.2LLM Evaluation
ModelSizesOpen modelsGemma 3\([34](https://arxiv.org/html/2608.15394#bib.bib36)\)4B, 12B, 27BGPT\-OSS\([23](https://arxiv.org/html/2608.15394#bib.bib37)\)20B, 120BLlama3\([21](https://arxiv.org/html/2608.15394#bib.bib40)\)8B, 70BQwen3\([35](https://arxiv.org/html/2608.15394#bib.bib33)\)8B, 32BProprietary modelsGPT 5\.4\-mini, 5\.4, 5\.5\([24](https://arxiv.org/html/2608.15394#bib.bib38)\)undisclosedClaude Sonnet 4\.6\([1](https://arxiv.org/html/2608.15394#bib.bib39)\)undisclosedClaude Haiku 4\.5undisclosedTable 1:LLMs used in our evaluation\.We evaluated a diverse set of models \(Table[1](https://arxiv.org/html/2608.15394#S4.T1)\) on both the curated golden set and a broader filtered dataset\. To ensure consistency with human evaluation, we prompt each model using the same instructions given to participants \(Appendix[B\.1](https://arxiv.org/html/2608.15394#A2.SS1)\)\. We conduct both a 2\-way evaluation, in which the model selects between two randomized scenarios as shown in Figure[1](https://arxiv.org/html/2608.15394#S1.F1), and a 4\-way evaluation, which additionally allows responses of*same*or*cannot tell*\. In the former, models choose between Scenario A and Scenario B\. The latter adds*same*and*cannot tell*as options to help measure decisiveness and confidence\.
We ran open\-weight models locally using Hugging Face Transformers with greedy decoding and a 512\-token generation budget, using each model’s native chat template\. We evaluated proprietary models through the OpenAI Batch API and Anthropic Message Batches using default sampling settings and a 1024\-token completion budget\. We randomized scenario order per item using a fixed seed \(42\)\.
## 5Results
In this section, we will first see the LLM evaluation results \(§[5\.1](https://arxiv.org/html/2608.15394#S5.SS1), addressingRQ2,RQ3\), followed by human evaluation \(§[5\.2](https://arxiv.org/html/2608.15394#S5.SS2), addressingRQ1\)\. Two statistical issues affect both\. First, responses can depend on the order in which we present the scenarios\. Following[19](https://arxiv.org/html/2608.15394#bib.bib43), we correct positional bias by symmetrizing results across both orderings for each pair\. Second, since we instantiate all 6,684 items in the data from 111 templates, examples in a template family are not statistically independent\. Treating them as such could understate uncertainty\. To address this, we cluster all standard errors in the originating template, with a Student\-t small\-sample adjustment for illusions with fewer template varieties\. Every confidence interval and comparison below reflects these corrections\.
### 5\.1LLM Evaluation
We first validate that the paraphrased dataset reproduces the hand\-curated gold set\. Agreement between the two is high for the character perspective \(Spearman rank correlationρ=0\.85/0\.79\\rho=0\.85/0\.79for 2\-way/4\-way resp\) and the reader perspective \(Spearman rank correlationρ=0\.84/0\.86\\rho=0\.84/0\.86for 2\-way/4\-way resp\)\. The 4\-way protocol is also more sensitive to illusion\-specific differences than the 2\-way protocol\. In general, we found that both the 2\-way and 4\-way evaluation settings have almost equivalent analysis results\. Appendix[B](https://arxiv.org/html/2608.15394#A2)includes the full concordance, rank\-preservation and detailed model\-specific results for both settings\.
#### Differences across Illusion Types\.
Figure 3:P\(Experimental Chosen\) representing model agreement with the literature, averaged across the 14 model results\.Illusion strength varies across the five illusion types \(RQ2, Figure[3](https://arxiv.org/html/2608.15394#S5.F3)\)\. The bars report the proportionpexpp\_\{\\text\{exp\}\}of*decisive*responses \(excludingsame/unsure\) in which models select the literature\-predicted scenario, with whiskers showing 95% CI of the cross\-model mean\. The size of the departure from chance is more important than its existence\. So we read effects off the intervals rather than significance stars\.
Oddball, Familiarity, and Temporal Order reach highpexpp\_\{\\text\{exp\}\}across nearly all models \(0\.85–1\.00, confidence intervals well clear of chance\)\. Emotional Time Dilation is markedly weaker\. Both Emotional Time Dilation and Oddball are arousal\-related illusions, but Emotional Time Dilation uses emotionally salient contexts \(e\.g\., surprise or stress\) while Oddball relies on more structured and systematic changes\. The comparatively weaker performance on Emotional Time Dilation may indicate that LLMs find emotionally grounded temporal reasoning more challenging than structured, pattern\-based distinctions\.
#### Character vs\. Reader Perspectives\.
Figure 4:P\(Experimental Chosen\), showing the change in model agreement with the literature when asking from the Character Perspective versus the Reader Perspective, averaged across 14 model results\.There are clear differences between character and reader\-based evaluations \(Figure[4](https://arxiv.org/html/2608.15394#S5.F4),RQ3\) for Filled Duration: the model’s selection changes from the experimental option when moving from the character to the reader perspective\. To explain this, we conjecture a relationship between the reader/character split and the previously studied split between lived versus retrospective experiences\. In shorter experimental settings, intervals containing tones or visual markers are typically perceived as longer than silent gaps\([36](https://arxiv.org/html/2608.15394#bib.bib9)\)\. In contrast, over longer periods, “melodic” sequences can be perceived as shorter than unfilled intervals, reflecting the intuition that time appears to pass more quickly during continuous stimuli\([12](https://arxiv.org/html/2608.15394#bib.bib10)\)\. This means shorter, lived intervals should favor the filled condition, while longer, retrospective ones should favor the empty condition\.
Our results mirror this distinction\. For the character perspective, where the narrative frames the experience as lived through, models favor the empty interval as longer\. The reader’s perspective, evaluated retrospectively after processing the narrative, shows opposite results\. This shift is not an inconsistency, but reflects a framing effect that is consistent with the character/reader distinction we aimed to capture\.
Other illusion types show relatively smaller shifts between perspectives\. Rather than reflecting “correct” or “incorrect” responses, these differences indicate that some illusions depend on how the narrative contextualizes the temporal experience\. The appendix has detailed model\-specific results\.
#### Differences across Subtypes within Illusion Types\.
Figure 5:Each value in the heatmap represents the P\(Experimental Chosen\) for a subcategory within each illusion type, computed separately for the Character and Reader perspectives\.Within each illusion type, Figure[5](https://arxiv.org/html/2608.15394#S5.F5)shows the relative differences across its subcategories, illustrating how individual components contribute to the overall effect\. We excluded non\-decisive selection options for the 4\-way setting\.
Within Emotional Time Dilation, Interval Intensity consistently yields higher values compared to Emotional Arousal and Stakes\-and\-Expectations\. This contrast suggests that LLMs capture directly observable temporal cues more reliably than subcategories that require interpreting emotional context\. The strong and consistent performance across Oddball Effect subtypes, driven primarily by clear, structured temporal deviations, further supports this interpretation\.
For Temporal Order, Narrative Disruption exhibits the largest difference between reader and character perspectives\. This subtype retains a relatively strong signal, potentially because it follows a linear progression with a clear deviation\. In contrast, other Temporal Order subcategories, such as those involving repetition or subtle ordering constraints, show more variability\. These subtypes require higher\-level contextual or social reasoning \(e\.g\., determining whether repeated or reordered events are meaningful or irregular\)\.
Overall, these results reinforce our previous observation that situations grounded in clear, structured temporal deviations tend to produce more consistent model behavior\.
#### Thinking Tokens
Figure 6:The most common words for all 5 illusion types split by 2\-way and 4\-way answer options within the thinking tokens of Qwen3\-32b\. Refer to Figure[14](https://arxiv.org/html/2608.15394#A2.F14)for a by illusion\-type breakdown\.Figure 7:Human evaluation across illusion types following a 2\-way evaluation with positional bias correction\.To better understand the strong alignment with the literature, we manually analyzed the thinking traces of a representative model in our suite, Qwen3\-32B\. We chose Qwen3\-32B as one of the largest models with exposed thinking tokens that we evaluated\. \(GPT\-OSS\-120B’s exposed reasoning predominantly consisted of compliance\-related meta text rather than task\-specific temporal reasoning\.\) We observed a recurring tendency to frame reasoning in terms of psychological phenomena\. Using the NLTK package to extract and analyze word frequencies, we found that terms such asmentallyandpsychologicallyappeared across four of the five illusion categories \(Figure[6](https://arxiv.org/html/2608.15394#S5.F6)\)\. Moreover,∼\\sim70% of the traces contained explicit research\-oriented framing, including phrases such as ‘Psychological research suggests’, ‘Research indicates people perceive time differently’, and ‘In psychology research’\. This framing emerged even when the prompt did not explicitly request a psychological explanation, suggesting that the model classified these questions as instances of established psychological phenomena\. Consistent with this framing, many generated responses mirrored explanations from psychology literature rather than relying solely on the specific details of the presented scenario\.
The repeated use of the termmentallyfurther suggests a distinction between subjective and objective time within the model’s reasoning process\. Phrases such as ‘would make him pause mentally longer,’ and ‘mentally stretches perception,’ indicate that the model frequently interpreted duration in terms of internal cognitive experience rather than elapsed clock time\. These patterns are consistent with the model associating the prompts with known psychology experiments\.
The evaluation prompts did not explicitly reference psychological research\. But the prevalence of such framing suggests that associations with psychological literature may influence the model’s reasoning process and final answers\.
#### Model Response Behavior\.
While we used Qwen\-32B to generate paraphrased options for the large\-scale dataset, its behavior does not substantially differ from Qwen\-8B, suggesting limited generation\-induced bias\. We analyze additional factors affecting model behavior, including positional bias, prompt complexity, and non\-committal responses\.
Models exhibit systematic differences in uncertainty: smaller open\-source models show stronger positional bias\. Models with greater uncertainty also demonstrate higher rates of non\-committal responses\. Across model families, GPT models tend to avoid decisions under ambiguity, Qwen models more explicitly express uncertainty, and Gemma models typically force a selection\. Prompt length has limited impact across most illusion types, with the exception of filled duration, where longer narratives strengthen the expected effect\. Appendix[B](https://arxiv.org/html/2608.15394#A2)has additional analyses of positional bias, prompt length, and response patterns\.
### 5\.2Human Evaluation
We used the 2\-way experimental setup to make weak preferences more observable in aggregated statistics\. Human evaluation shows little evidence that a literary narrative can cause temporal illusions or differences in character/reader framings \(RQ1, Figures[7](https://arxiv.org/html/2608.15394#S5.F7)and[22](https://arxiv.org/html/2608.15394#A2.F22)in Appendix[B](https://arxiv.org/html/2608.15394#A2)\)\.
Only Oddball and Temporal Order have a statistically significant preference towards the experimental sentence\. Oddball and Temporal Order disrupt the text directly, creating a task similar to a ‘spot the difference’ question\. Thus, the reader can detect them without internally simulating duration\. The other illusions require a more demanding inferential step\. The reader must construct a mental model to estimate how long an experience would feel\. There are no significant differences between reader and character questions\.
In short, only manipulations detectable directly in the text produce reliable effects\. These results are in sharp contrast to the robust LLM effects we saw in §[5\.1](https://arxiv.org/html/2608.15394#S5.SS1)\.
## 6Discussion, Limitations and Future Work
While we filtered for word count difference, some asymmetry remains across illusion types\. This is unavoidable because certain illusion templates inherently require longer descriptions as a structural limitation of translating the modality of experiments to narrative form\. Relatedly, some illusions, e\.g\., Familiarity, depend on assumptions about experiences that are not universal\. A task that is routine for one reader may be novel for another\.
We also observe substantial variation in positional bias across models\. While debiasing strategies such as UNQOVER\-style symmetrization can mitigate moderate bias, they are less reliable in cases dominated by positional heuristics\. In such settings, evaluation results may reflect model\-specific artifacts rather than a meaningful signal\.
A limitation of this work is that we only study the model outputs and reasoning traces, not the internal representations\. To better test the underlying mechanisms responsible for these behaviors, future work would require probing the model’s internal representations during decision\-making\. Isolating activation vectors associated with emotion, attention, novelty, or other psychological factors would allow for a more direct comparison to human psychology experiments\.
Our evaluation focuses on deciding if LLMs exhibit these illusions, not if they should\. They show that models and humans diverge: readers show a reliable effect for only oddball and temporal order, while models for four of the five illusions\. Model behavior aligns with the psychology literature rather than with our human participants\. Understanding how models interpret and represent temporal illusions can inform the design of agents\. More broadly, this work contributes to behavior modeling in LLMs, where capturing subjective time may be important for building systems that interact with humans\. In other cases, such as safety\-critical systems, temporal biases may affect decision\-making\.
## 7Conclusion
We introduced a narrative benchmark of 6,684 paired scenarios, generated from 111 templates across five temporal illusions, to ask whether written text can evoke illusions\. Our human study finds no reliable evidence that they transfer, except where the manipulation is directly visible in the text\. Models select the literature\-predicted scenario for four of the five illusions\. Roughly 70% of the analyzed reasoning traces explicitly invoke psychological research even when the prompt does not request it\. This suggests that model alignment is consistent with retrieval of published findings rather than human\-like temporal biases\.
## 8Acknowledgments
This work was supported by the Parent Fund Scholarship through the University of Utah Undergraduate Research Opportunity Program\. The support and resources from the Center for High Performance Computing at the University of Utah are gratefully acknowledged\.
## References
- Anthropic \(2026\)AnthropicSystem card: claude sonnet 4\.6\.Technical reportAnthropic\.External Links:[Link](https://www-cdn.anthropic.com/bbd8ef16d70b7a1665f14f306ee88b53f686aa75.pdf)Cited by:[Table 1](https://arxiv.org/html/2608.15394#S4.T1.1.9.1)\.
- Avant and Lyman \(1975\)L\. L\. Avant and P\. J\. LymanStimulus familiarity influences perceived duration in prerecognition visual processing\.\.Journal of Experimental Psychology: Human Perception and Performance1\(3\),pp\. 205\.Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx4.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p4.1)\.
- Binz and Schulz \(2023\)M\. Binz and E\. SchulzUsing cognitive psychology to understand gpt\-3\.Proceedings of the National Academy of Sciences120\(6\),pp\. e2218523120\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2218523120),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.2218523120),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.2218523120Cited by:[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p1.1)\.
- Birngruberet al\.\(2015\)T\. Birngruber, H\. Schröter, and R\. UlrichIntroducing a control condition in the classic oddball paradigm: oddballs are overestimated in duration not only because of their oddness\.Attention, Perception, & Psychophysics77\(5\),pp\. 1737–1749\.Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx2.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p2.1)\.
- Block and Grondin \(2014\)R\. A\. Block and S\. GrondinTiming and time perception: a selective review and commentary on recent reviews\.Frontiers in PsychologyVolume 5 \- 2014\.External Links:[Link](https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2014.00648),[Document](https://dx.doi.org/10.3389/fpsyg.2014.00648),ISSN 1664\-1078Cited by:[§1](https://arxiv.org/html/2608.15394#S1.p1.1)\.
- Block and Zakay \(1997\)R\. A\. Block and D\. ZakayProspective and retrospective duration judgments: a meta\-analytic review\.Psychonomic Bulletin & Review4\(2\),pp\. 184–197\.External Links:[Document](https://dx.doi.org/10.3758/BF03209393)Cited by:[§1](https://arxiv.org/html/2608.15394#S1.p1.1)\.
- Boltz \(1998\)M\. G\. BoltzThe processing of temporal and nontemporal information in the remembering of event durations and musical structure\.\.Journal of experimental psychology: human perception and performance24\(4\),pp\. 1087\.Cited by:[§1](https://arxiv.org/html/2608.15394#S1.p1.1)\.
- Brown and Smith\-Petersen \(2014\)S\. W\. Brown and G\. A\. Smith\-PetersenTime perception and temporal order memory\.Acta psychologica148,pp\. 173–180\.Cited by:[item B \-](https://arxiv.org/html/2608.15394#A1.I1.ix2.p1.1),[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx5.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p4.1),[item B](https://arxiv.org/html/2608.15394#S3.I1.ix2.p1.1)\.
- Buffardi \(1971\)L\. BuffardiFactors affecting the filled\-duration illusion in the auditory, tactual, and visual modalities\.Perception & Psychophysics10\(4\),pp\. 292–294\.Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx3.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p3.1)\.
- Chenet al\.\(2025\)S\. Chen, Y\. Zheng, S\. Li, Q\. Cheng, and X\. QiuPerceive the passage of time: a systematic evaluation of large language model in temporal relativity\.InProceedings of the 31st International Conference on Computational Linguistics,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),Abu Dhabi, UAE,pp\. 8304–8313\.External Links:[Link](https://aclanthology.org/2025.coling-main.554/)Cited by:[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p4.1)\.
- Chenet al\.\(2021\)W\. Chen, X\. Wang, W\. Y\. Wang, and W\. Y\. WangA dataset for answering time\-sensitive questions\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks,J\. Vanschoren and S\. Yeung \(Eds\.\),Vol\.1,pp\.\.External Links:[Link](https://datasets-benchmarks-proceedings.neurips.cc/paper_files/paper/2021/file/1f0e3dad99908345f7439f8ffabdffc4-Paper-round2.pdf)Cited by:[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p2.1)\.
- Droit\-Voletet al\.\(2010\)S\. Droit\-Volet, E\. Bigand, D\. Ramos, and J\. L\. O\. BuenoTime flies with music whatever its emotional valence\.Acta psychologica135\(2\),pp\. 226–232\.Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx3.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p3.1),[§5\.1](https://arxiv.org/html/2608.15394#S5.SS1.SSSx2.p1.1)\.
- Eagleman \(2008\)D\. M\. EaglemanHuman time perception and its illusions\.Current opinion in neurobiology18\(2\),pp\. 131–136\.Cited by:[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p1.1)\.
- Fatemiet al\.\(2025\)B\. Fatemi, S\. M\. Kazemi, A\. Tsitsulin, K\. Malkan, J\. Yim, J\. Palowitch, S\. Seo, J\. Halcrow, and B\. PerozziTest of time: a benchmark for evaluating llms on temporal reasoning\.InInternational Conference on Learning Representations,Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(Eds\.\),Vol\.2025,pp\. 94426–94447\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/file/eb7295a8bc613b375726659c2ecd6f14-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2608.15394#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p2.1)\.
- Gil and Droit\-Volet \(2012\)S\. Gil and S\. Droit\-VoletEmotional time distortions: the fundamental role of arousal\.Cognition & emotion26\(5\),pp\. 847–862\.Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx1.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p1.1)\.
- Hurlemannet al\.\(2005\)R\. Hurlemann, B\. Hawellek, A\. Matusch, H\. Kolsch, H\. Wollersen, B\. Madea, K\. Vogeley, W\. Maier, and R\. J\. DolanNoradrenergic modulation of emotion\-induced forgetting and remembering\.Journal of Neuroscience25\(27\),pp\. 6343–6349\.Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx1.p1.1)\.
- Jafarpour and Spiers \(2017\)A\. Jafarpour and H\. SpiersFamiliarity expands space and contracts time\.Hippocampus27\(1\),pp\. 12–16\.Cited by:[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p4.1)\.
- Kosinski \(2024\)M\. KosinskiEvaluating large language models in theory of mind tasks\.Proceedings of the National Academy of Sciences121\(45\),pp\. e2405460121\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2405460121),[Link](https://www.pnas.org/doi/abs/10.1073/pnas.2405460121),https://www\.pnas\.org/doi/pdf/10\.1073/pnas\.2405460121Cited by:[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p1.1)\.
- Liet al\.\(2020\)T\. Li, D\. Khashabi, T\. Khot, A\. Sabharwal, and V\. SrikumarUNQOVERing stereotyping biases via underspecified questions\.InFindings of the Association for Computational Linguistics: EMNLP 2020,T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 3475–3489\.External Links:[Link](https://aclanthology.org/2020.findings-emnlp.311/),[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.311)Cited by:[§B\.4](https://arxiv.org/html/2608.15394#A2.SS4.p2.1),[§5](https://arxiv.org/html/2608.15394#S5.p1.1)\.
- Manahovaet al\.\(2020\)M\. E\. Manahova, E\. Spaak, and F\. P\. de LangeFamiliarity increases processing speed in the visual system\.Journal of cognitive neuroscience32\(4\),pp\. 722–733\.Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx4.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p4.1)\.
- Meta \(2024\)MetaThe llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[Table 1](https://arxiv.org/html/2608.15394#S4.T1.1.5.1)\.
- OpenAI \(2024\)OpenAIGPT\-4o system card\.arXiv preprint arXiv:2410\.21276\.External Links:[Link](https://arxiv.org/abs/2410.21276)Cited by:[§A\.2](https://arxiv.org/html/2608.15394#A1.SS2.p2.1),[§3\.2](https://arxiv.org/html/2608.15394#S3.SS2.p1.1)\.
- OpenAI \(2025a\)OpenAIGpt\-oss\-120b & gpt\-oss\-20b model card\.External Links:2508\.10925,[Link](https://arxiv.org/abs/2508.10925)Cited by:[Table 1](https://arxiv.org/html/2608.15394#S4.T1.1.4.1)\.
- OpenAI \(2025b\)OpenAIOpenAI gpt\-5 system card\.External Links:[Link](https://arxiv.org/abs/2601.03267)Cited by:[Table 1](https://arxiv.org/html/2608.15394#S4.T1.1.8.1)\.
- Palan and Schitter \(2018\)S\. Palan and C\. SchitterProlific\. ac—a subject pool for online experiments\.Journal of behavioral and experimental finance17,pp\. 22–27\.Cited by:[§4\.1](https://arxiv.org/html/2608.15394#S4.SS1.p1.1)\.
- Pariyadath and Eagleman \(2007\)V\. Pariyadath and D\. EaglemanThe effect of predictability on subjective duration\.PloS one2\(11\),pp\. e1264\.Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx2.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p2.1)\.
- Pezeshkpour and Hruschka \(2024\)P\. Pezeshkpour and E\. HruschkaLarge language models sensitivity to the order of options in multiple\-choice questions\.InFindings of the Association for Computational Linguistics: NAACL 2024,pp\. 2006–2017\.Cited by:[§B\.4](https://arxiv.org/html/2608.15394#A2.SS4.p3.1)\.
- Schiffman and Bobko \(1977\)H\. Schiffman and D\. J\. BobkoThe role of number and familiarity of stimuli in the perception of brief temporal intervals\.The American journal of psychology,pp\. 85–93\.Cited by:[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p3.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p4.1)\.
- Schreweet al\.\(2024\)L\. Schrewe, C\. Nickel, and L\. FlekProbing the robustness of theory of mind in large language models\.InEighth Widening NLP Workshop \(WiNLP 2024\) Phase II,External Links:[Link](https://openreview.net/forum?id=X8Mdv9qLOS)Cited by:[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p1.1)\.
- Skylark and Gheorghiu \(2017\)W\. J\. Skylark and A\. I\. GheorghiuFurther evidence that the effects of repetition on subjective time depend on repetition probability\.Frontiers in Psychology8,pp\. 1915\.Cited by:[item A \-](https://arxiv.org/html/2608.15394#A1.I5.ix1.p1.1),[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx4.p1.1),[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx5.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p4.1)\.
- Suet al\.\(2024\)Z\. Su, J\. Zhang, T\. Zhu, X\. Qu, J\. Li, M\. Zhang, and Y\. ChengTimo: towards better temporal reasoning for language models\.CoRRabs/2406\.14192\.External Links:[Link](https://doi.org/10.48550/arXiv.2406.14192)Cited by:[§1](https://arxiv.org/html/2608.15394#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p3.1)\.
- Tamet al\.\(2025\)Z\. R\. Tam, C\. Wu, C\. Lin, and Y\. ChenNone of the above, less of the right parallel patterns in human and LLM performance on multi\-choice questions answering\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 20112–20134\.External Links:[Link](https://aclanthology.org/2025.findings-acl.1031/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.1031),ISBN 979\-8\-89176\-256\-5Cited by:[§B\.4](https://arxiv.org/html/2608.15394#A2.SS4.p3.1)\.
- Tanet al\.\(2023\)Q\. Tan, H\. T\. Ng, and L\. BingTowards benchmarking and improving the temporal reasoning capability of large language models\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 14820–14835\.External Links:[Link](https://aclanthology.org/2023.acl-long.828/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.828)Cited by:[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p2.1)\.
- Team \(2025a\)G\. TeamGemma 3\.External Links:[Link](https://goo.gle/Gemma3Report)Cited by:[Table 1](https://arxiv.org/html/2608.15394#S4.T1.1.3.1)\.
- Team \(2025b\)Q\. TeamQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§A\.2](https://arxiv.org/html/2608.15394#A1.SS2.p3.1),[§3\.2](https://arxiv.org/html/2608.15394#S3.SS2.p2.1),[Table 1](https://arxiv.org/html/2608.15394#S4.T1.1.6.1)\.
- Thomas and Brown \(1974\)E\. C\. Thomas and I\. BrownTime perception and the filled\-duration illusion\.Perception & Psychophysics16\(\),pp\. 449–458\.External Links:[Document](https://dx.doi.org/10.3758/BF03198571)Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx3.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p3.1),[§5\.1](https://arxiv.org/html/2608.15394#S5.SS1.SSSx2.p1.1)\.
- Tipples \(2008\)J\. TipplesNegative emotionality influences the effects of emotion on time perception\.Emotion8\(1\),pp\. 127–131\.External Links:[Document](https://dx.doi.org/10.1037/1528-3542.8.1.127)Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx1.p1.1),[§1](https://arxiv.org/html/2608.15394#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p2.1)\.
- U\.S\. Social Security Administration \(2024\)U\.S\. Social Security AdministrationPopular baby names by decade and century\.Note:https://www\.ssa\.gov/oact/babynames/decades/century\.htmlAccessed April 2026Cited by:[§A\.2](https://arxiv.org/html/2608.15394#A1.SS2.p2.1),[§3\.2](https://arxiv.org/html/2608.15394#S3.SS2.p1.1)\.
- Wang and Zhao \(2024\)Y\. Wang and Y\. ZhaoTRAM: benchmarking temporal reasoning for large language models\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 6389–6415\.External Links:[Link](https://aclanthology.org/2024.findings-acl.382/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.382)Cited by:[§1](https://arxiv.org/html/2608.15394#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p3.1)\.
- Weardenet al\.\(2007\)J\. H\. Wearden, R\. Norton, S\. Martin, and O\. Montford\-BebbInternal clock processes and the filled\-duration illusion\.\.Journal of Experimental Psychology: Human Perception and Performance33\(3\),pp\. 716\.Cited by:[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.15394#S2.SS1.p3.1)\.
- Wehrmanet al\.\(2020\)J\. J\. Wehrman, J\. Wearden, and P\. SowmanThe expected oddball: effects of implicit and explicit positional expectation on duration perception\.Psychological research84\(3\),pp\. 713–727\.Cited by:[§A\.1](https://arxiv.org/html/2608.15394#A1.SS1.SSSx2.p1.1)\.
- Weiet al\.\(2023\)Y\. Wei, Y\. Su, H\. Ma, X\. Yu, F\. Lei, Y\. Zhang, J\. Zhao, and K\. LiuMenatQA: a new dataset for testing the temporal comprehension and reasoning abilities of large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 1434–1447\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.100/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.100)Cited by:[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p2.1)\.
- Zhouet al\.\(2019\)B\. Zhou, D\. Khashabi, Q\. Ning, and D\. Roth“Going on a vacation” takes longer than “going for a walk”: a study of temporal commonsense understanding\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 3363–3369\.External Links:[Link](https://aclanthology.org/D19-1332/),[Document](https://dx.doi.org/10.18653/v1/D19-1332)Cited by:[§2\.2](https://arxiv.org/html/2608.15394#S2.SS2.p3.1)\.
## Appendix AAppendix: Additional Benchmark Information
### A\.1Full Benchmark Design
By translating classic psychology experiments into narrative prompts, we’re able to mimic how the human brain stretches and compresses intervals through writing\. This lets us easily evaluate how LLMs mirror or diverge from human\-like temporal biases\. Each of the five illusions has multiple question styles mimicking different types of evaluations and components in classic psychology\. Most templates contain both longer and shorter versions of the question to account for the effects of sentence length on perceived time if allowed by the question design\. We outline a detailed explanation of each template type and its reasoning below; for concrete examples, refer to Table[4](https://arxiv.org/html/2608.15394#A1.T4)in Appendix[A](https://arxiv.org/html/2608.15394#A1)\.
#### Emotional Time Dilation
Emotional time dilation, which occurs with high\-arousal stimuli, causes an individual to perceive an interval as longer than the truth\. For example, people with arachnophobia consistently overestimate the duration a spider image is presented in comparison to neutral stimuli\([37](https://arxiv.org/html/2608.15394#bib.bib1)\)\. Other studies targeting this illusion used mutilated images to stimulate arousal\. The mutilated images showed significant overestimation compared to neutral faces\([15](https://arxiv.org/html/2608.15394#bib.bib7)\)\. The pleasantness of an event and arousal influence both real\-time perception and retrospective memory encoding\([16](https://arxiv.org/html/2608.15394#bib.bib13)\)\.
We translated the experiments into a narrative framework that uses two variables: situational stakes and narrative explication\. We create conditions that isolate the factors\. While the literature generally associates emotional time dilation with stimulus\-driven surprise or threat, we adopt a broader definition that considers an arousal–attention system\. Our conditions also examine context\-induced arousal \(via stakes\) and attentional load as emotional mechanisms known to influence perceived duration\.
The difference between emotional time dilation and the oddball effect is what captures attention and how strongly\. Heightened arousal—events that feel important, stressful, or unexpected—drives emotional time dilation\. These situations create top\-down attention, meaning internal states such as concern, urgency, or significance guide attention\. Our design of the oddball effect is bottom\-up novelty\. A stimulus stands out within a predictable sequence and briefly captures attention simply because it is different, not because it carries emotional weight or consequence\. We define emotional time dilation as meaningful, high\-arousal experiences that sustain attention, whereas the oddball effect results from brief, stimulus\-driven attention without deeper emotional impact\.
- A \-Stakes and Expectations: We vary the level of situational stakes, presenting a neutral stimulus \(e\.g\., a flickering reflection\) in both low\-stakes and high\-stakes contexts\. This condition tests how potential consequences or pressure alter engagement with otherwise identical stimuli, which influences subjective time\.
- B \-Interval Intensity: We vary the cognitive demand in an empty interval by comparing high\-focus tasks with low\-demand activities\. This induces cognitive arousal\. When cognitive resources are heavily engaged, fewer resources remain available for tracking time, altering how long the interval feels\([8](https://arxiv.org/html/2608.15394#bib.bib14)\)\.
- C \-Emotional Arousal: We pair a neutral baseline with a sudden, unexpected, or emotional event while holding context and stakes constant\. This tests whether involuntary attentional capture from surprise or emotional salience expands perceived duration\.
By introducing high\-stakes or threatening elements, we simulate the arousal in emotional time dilation studies through narrative events\. By holding the objective action constant while varying the emotional context, we can test whether readers perceive the same event as lasting longer when the scene feels more intense\.
In laboratory settings, researchers simulate emotional time dilation in the participant directly by inducing an increase in heart rate\. In contrast, our narrative\-based questions ask the reader to imagine these states rather than experience them directly\. As a result, responses may reflect reasoning about what should feel longer, rather than reproducing how an internal timing mechanism would actually behave under real arousal\.
#### Oddball Effect
The Oddball Effect happens when a novel stimulus appears within a series of identical stimuli\. This makes the unique item seem to last longer\([26](https://arxiv.org/html/2608.15394#bib.bib21)\)\. Duration overestimation is more pronounced for the oddball\([4](https://arxiv.org/html/2608.15394#bib.bib20)\)\. Two contrasting mechanisms drive this illusion: a top\-down process, where the closer an individual gets to expecting a change, the more significant that change feels when it occurs; and a bottom\-up process, where the rarity of the oddball automatically captures attention and interrupts the brain’s habituation to the repeated sequence\([41](https://arxiv.org/html/2608.15394#bib.bib23)\)\. We mimicked the effects of the experiments by using sequential predictability and novel stimuli\.
- A \-Sequence Disruption: We create this illusion with a predictable sequence, and then introduce a deviation within that pattern\. The disruption captures attention through bottom\-up novelty, mirroring classic oddball paradigms where a distinct stimulus within a uniform sequence is perceived as lasting longer\. We divide this illusion template into three subsections: - 1\.Containing: The oddball occurs within the repeating sequence itself, embedded directly in the interval being evaluated\. - 2\.Preceding: The oddball appears before the target interval and influences how the following interval is perceived\. - 3\.Following: The oddball occurs after the target interval, testing how a later disruption retroactively affects the perceived duration\.
To separate the overlap between the Oddball Effect and Emotional Time Dilation, we distinguish structural surprise and arousal\. The Oddball Effect relies on perceptual novelty within a sequence\. It does not require the "oddball" to be scary, exciting, or high\-stakes; it simply needs to be different from the preceding standards\.
#### Filled\-Duration Effect
The Filled\-Duration Effect describes when a time interval containing discrete elements is perceived as longer than an empty interval of an objectively identical duration\. In experimental settings, designs consistently show that intervals containing brief tones or visual markers appear longer than silent gaps\([36](https://arxiv.org/html/2608.15394#bib.bib9)\)\. The number of elements affects this\. The more distinct events the brain processes within the window, the longer the window of time feels\([9](https://arxiv.org/html/2608.15394#bib.bib12)\)\. In parallel, continuous structured or “melodic” sequences are shorter than silent ones, showing that time flies when listening to music tested over longer durations\([12](https://arxiv.org/html/2608.15394#bib.bib10)\)\.
To translate these findings into a narrative framework, we fill an interval by comparing a sequence of distinct actions to a continuous period of time:
- A \-Information Density: We contrast an interval filled with events with an empty waiting period, reflecting findings that filled intervals are often perceived as longer than empty ones for short durations\.
From the reader’s perspective, we expect the filled interval to feel longer\. However, from the character’s perspective, the empty interval may feel longer, as the absence of stimulation can heighten awareness of time passing \(similar to how engaging activities, like music, can make longer durations feel shorter\)\. This structure mirrors existing literature on short\- versus long\-duration judgments by asking the reader to evaluate duration from both an external and internal perspective\.
#### Familiarity Duration Effect
The subjective duration of an interval depends on the brain recognizes the task\. This phenomenon is known as the familiarity\-duration effect\. Generally, novel events are perceived as lasting longer than familiar ones, where the newness of an experience needs more cognitive processing and more frequent neural sampling\([20](https://arxiv.org/html/2608.15394#bib.bib15);[2](https://arxiv.org/html/2608.15394#bib.bib17)\)\. When a task becomes routine, the brain processes certain recurring activities more effectively, and time will begin to pass more quickly\([30](https://arxiv.org/html/2608.15394#bib.bib19)\)\.
We focus on the cognitive differences between novelty and routine to demonstrate this illusion\. We define familiarity in two ways: 1\) the general perceived familiarity by the public, and 2\) the declared expertise of the character\.
- A \-Labeled Expertise: We present identical steps but explicitly frame the event as either “familiar” or “unfamiliar\.” This tests whether labeling alone influences perceived duration, targeting whether the reader infers temporal perception without changing the core actions\.
- B \-General Novelty: We compare a character performing an objectively familiar, everyday task with an objectively unfamiliar task, while holding the character’s stated level of familiarity constant across both\. This tests whether intrinsic task novelty affects the reader’s perceived duration\.
A significant challenge and possible shortcoming in translating familiarity effects into a test question set is that "familiarity" is not universal\. It can be culturally and demographically dependent\. A laboratory study can ensure novelty by using controlled stimuli; narrative prompts rely on real\-world concepts\. What is a routine, "compressed" task for a digital native might be a high\-effort, "expanded" task for an older adult or someone from a different professional background\.
#### Temporal Order Effects
Temporal Order Effects describe how the structure and predictability of event order influence subjective duration\. We distinguish between two complementary processes: order violation and order redundancy\. A person needs to actively pay attention to the sequence to detect an explicit order violation\. This can cause duration judgments to become overestimated\([8](https://arxiv.org/html/2608.15394#bib.bib14)\)\. Actively anticipating or monitoring for an upcoming event can lengthen perceived duration\([30](https://arxiv.org/html/2608.15394#bib.bib19)\)\. Order redundancy happens when repeated stimuli compress subjective time\. To translate these sequence effects, we isolate the order of a character’s workflow:
- A \-Scrambled Workflow: We present a character performing a task in the wrong order, explicitly stating it at the start\. This forces the reader to be aware of the incorrect sequence throughout\. This mirrors experimental setups showing that violations of expected order reduce temporal accuracy and alter perceived duration\. This structure follows[30](https://arxiv.org/html/2608.15394#bib.bib19)\.
- B \-Event Repetition: We compare a narrative sequence containing repetitive actions to a sequence with distinct events\. This tests whether repetition within a sequence compresses perceived duration, as repeated events become more predictable and require less processing, leading to a denser, more compressed memory representation of the interval\. This manipulation does not introduce an outlier; It reduces novelty across the sequence with repetition\.
- C \-Narrative Disruption: We present a sequential, coherent narrative that an unexpected event suddenly interrupts\. This tests whether an unexpected event triggers retrospective expansion of the surrounding interval, as the reader re\-evaluates prior events in light of the disruption\. The effect comes from a structural violation of the sequence, leading to reinterpretation of what came before, rather than from physiological or emotional arousal\.
### A\.2Dataset Creation
Each template consists of a paired control and experimental condition that differs in a single targeted manipulation\. The experimental condition corresponds to the scenario that, based on prior psychological literature, is expected to be perceived as longer for the character\. We design the templates so that the content closely matches in length, wording, and narrative structure\.
We design our collection of fillers \(stimuli\) for contextual variation while preserving the structure of each illusion type\. We wrote the initial set of 10 fillers per category manually and then expanded it to 20 using GPT\-4o\([22](https://arxiv.org/html/2608.15394#bib.bib34)\)\. We subsequently reviewed and edited all generated fillers by hand\. We sampled character names from the top 100 most common names in the United States over the last one hundred years\([38](https://arxiv.org/html/2608.15394#bib.bib3)\)\. Using these fillers, we generated 2000 examples for each illusion type by randomly sampling from the available templates, resulting in a total of 10000 candidate questions\. In addition to the filtered dataset, we manually curated a subset of 250 “golden” cases \(50 per illusion type\)\. We selected and lightly edited these examples by hand to ensure clarity, consistency, and faithful representation of each illusion\.
To reduce stylistic noise and increase grammatical coherence, we paraphrased each example three times using Qwen3\-32B\([35](https://arxiv.org/html/2608.15394#bib.bib33)\)with examples from the “golden” cases for few\-shot prompting \(Appendix[A](https://arxiv.org/html/2608.15394#A1)\)\. From these variants, we selected the version that most closely matched the desired length constraints between control and experimental conditions for evaluation\. We designed this step to minimize differences in word count that could bias perceived duration judgments\. We removed any example pair with a length difference greater than seven words\. After filtering, 6684 examples remained, with 3316 removed across all categories \(Table[2](https://arxiv.org/html/2608.15394#A1.T2)\)\.
Because we generate items from a small number of templates, examples within a template family are not statistically independent; accounting for this, the effective sample size is smaller\.
Table 2:Dataset size before and after filtering\.While the dataset is not perfectly balanced after filtering, each category retains a sufficiently large sample size for analysis \(Table[3](https://arxiv.org/html/2608.15394#A1.T3)\)\. Although this imbalance may introduce minor differences in statistical power across conditions, it does not pose a significant limitation for our evaluation\. Future work may re\-balance categories\.
We will make the complete generated dataset and slot fillers available in a public repository\.222https://github\.com/utahnlp/temporal\-illusions
Illusion Type/SubtypeCount%Emotional Time DilationInterval Intensity72144\.2%Stakes and Expectations62938\.6%Emotional Arousal28017\.2%Oddball EffectContaining42735\.3%Following39832\.9%Preceding38331\.7%Filled Duration EffectInformation Density1563100\.0%Familiarity Duration EffectGeneral Novelty61951\.1%Labeled Expertise59348\.9%Temporal Order EffectScrambled Workflow48945\.7%Narrative Disruption35533\.1%Event Repetition22721\.2%Table 3:Distribution of examples across illusion types and subtypes after filtering\.Table 4:Example of template structures for each illusion type and corresponding control and experimental conditionsTemplateControlExperimentalEmotional Time DilationStakes and ExpectationsAndrew wascasually checking his emailwhen he noticed a flickering reflection in the puddle outside\. He continued calmly\.Andrew wasworking on a time\-sensitive, crucial project when he saw a flickering reflection in the puddle outside\. He kept working\.Interval IntensityWhile waiting for a delayed train,Sarah looked over some unread text messages\.While waiting for a delayed train,Sarah tried to complete a difficult assignment\.Emotional ArousalSitting on the bleachers watching a game alone, Richard observed adry leaf rolling along the pavement, scraped by a light breeze that spun it and paused before it tumbled again toward the gutter\. He continued with his own business\.Sitting on the bleachers watching a game alone, Richard noticed aman wearing bright post\-it notes walking down the sidewalk\. A few post\-its flew in the wind before he vanished around the corner\. He continued with his own business\.Oddball EffectSequence DisruptionJames systematically processed badge activation processes, ensuring each setup matched perfectly\.Nothing was different\.He finished all tasks\.James systematically processed badge activation processes, but apeculiar channel pattern emerged midway\. He continued and finished all tasks\.Filled\-Duration EffectInformation DensityEmily spent the evening at home with various tasks\. Throughout it,she polished silverware, ironed tablecloths, arranged centerpieces, and set place settings\. She grabbed a snack after\.Emily spent the evening at home with little to do\. Throughout it,she lay on the floor watching the ceiling fan spin\. Eventually, she was hungry enough to grab a snack\.Familiarity Duration EffectLabeled ExpertiseDavid started building a detailed model airplane during his leisure time, naturallyfollowing familiar steps in order: first sorting pieces, then assembling the fuselage and wings, next adding small components, and finally inspecting the model\.David began constructing a detailed model airplane during his free time, carefullyfollowing each unfamiliar step: first sorting pieces, then assembling the fuselage and wings, next adding small components, and finally inspecting the model\.General NoveltyAt home, Linda beganboiling waterusing these unfamiliar steps: filling a kettle with water, placing it on the stove, heating until it boils, then turning off once boiling is confirmed\.Lindaassembled a complex bookshelffor the first time: unpacking components, arranging parts to identify screws and connectors, checking stability, and making adjustments\.Temporal Order EffectsScrambled WorkflowMelissa tied her shoelaces methodically and efficiently\. First, she grabbed the two laces\. Next, crossing and pulling them tight\. Then, form a loop and thread the other lace\. Finally, pulling the knot tight and securing it\. Finishing each step\.Melissa was tying her shoelaces but realizedshe was doing it incorrectly\. First, Melissa crossed and pulled the laces tight\.Next, she grabbed them\. Then, pulling the knot tight and securing it\. Finally, she formed a loop and threaded the other lace\.Event RepetitionJoshua made a list of recent tasks: running,bathroom break, resting after work, anda second bathroom visit\.Joshua noted down activities like going to the bathroom, running, and relaxing after a long day by watching a basketball game\.Narrative DisruptionStarting the day, Richard makes his bed\. Then, he reads a book\. Afterwards, he goes on a run\. Finally, he continues to read the book he read earlier\.Starting the day, Richard makes his bed\. Next, he reads a book\. Suddenly,he sees a man covered in Post\-it notes walking silently\.Finally, he goes for a run\.Full Paraphrasing PromptWe insert the examples after the system prompt and before the final user query, as alternating USER/ASSISTANT turns \(i\.e\., standard few\-shot formatting\)\.``` SYSTEM: You are an expert editor specializing in paraphrasing sentence pairs for psychology research. PRIORITIES (strictly in this order): 1. The experimental and control fields must each contain a full, fluent sentence. - Never use placeholder words like experimental or control as the sentence text. 2. Grammatical correctness and sentence fluency output must sound completely natural 3. Matched lengths: experimental and control sentences must be as close in word count as possible 4. Core meaning: preserve the key semantic contrast between sentences. - Details may be omitted when needed for fluency or length Return ONLY a valid JSON array of exactly 3 paraphrased variants. No preamble, no explanation. USER: Paraphrase these sentences. Target ˜N words each. Experimental: "..." Control: "..." ASSISTANT: [{"experimental": "...", "control": "..."}] USER: Paraphrase these sentences. Target ˜N words each. Experimental: "..." Control: "..." ASSISTANT: [{"experimental": "...", "control": "..."}] USER: Paraphrase these sentences. Target ˜{N} words each. Experimental ({ew} words): "{exp sentence}" Control ({cw} words): "{ctrl sentence}" Return JSON array of 3 diverse variants: ASSISTANT: ```
## Appendix BAppendix: Additional Evaluation Information
### B\.1Evaluation Instructions and Prompts
Survey InstructionsIn this survey, our goal is to understand how people perceive and experience the passage of time\. You will be presented with 25 pairs of scenarios, each describing a character undergoing a particular event\. For each pair, you will be asked to answer a question comparing the two scenarios\.We are interested in your intuitive, first\-instinct responses\. Please answer each question as quickly as possible, ideally within 15 seconds\. There are no correct or incorrect answers\.Each pair of scenarios describes events that lasted the same objective amount of time\. However, your task is to judge how long theyfelt, depending on the perspective specified in the question\.Character’s Perspective\.“Given the two events below that lasted the same amount of time, which one do you think felt longer for the character experiencing it?”Imagine yourself as the character\. Which event would feel longer to experience within the story?Reader’s Perspective\.“Given the two events below that lasted the same amount of time, which one felt longer to you while reading?”Answer from your perspective as the reader\. Which event felt longer or more time\-consuming to read and process?Please base your answers on your immediate impression rather than careful deliberation\.
Evaluation PromptsCharacter Perspective — System``` You are reading two short stories. Imagine you are the character. Which event felt longer to experience within the story? Given the two following events below that lasted the same amount of time, which one do you think felt longer for the character in the events? ``` Reader Perspective — System``` You are reading two short stories. As the person reading the text, which event felt longer or more time-consuming to read and process? Given the two following events below that lasted the same amount of time, which one felt longer to you while reading? ``` User Prompt \(shared across perspectives\)``` Sentence One: "{sentence}" Sentence Two: "{sentence}" Please respond with only "Sentence One" or "Sentence Two". ``` The model selects one sentence\. We map responses \("Sentence One"/"Sentence Two"\) back to experimental or control conditions based on presentation order\. We randomize sentence order once per pair and hold it constant across both perspectives\.
### B\.2Configuration Details
We ran open\-weight models locally with HuggingFace Transformers 5\.8\.0 \(PyTorch 2\.13\.0\) on NVIDIA L40S GPUs, loaded in 4\-bit NF4 precision via bitsandbytes with bfloat16 compute \(gpt\-oss models retain their native MXFP4 quantization\)\. We used a 512\-token generation budget, using each model’s own chat template\. We ran proprietary models through the OpenAI Batch API and Anthropic Message Batches at default sampling settings with a 1024\-token completion budget\. We randomize scenario order per item under a fixed seed \(42\)\.
### B\.3Evaluation Framework and Dataset Validation

\(a\) 4\-way evaluation
\(b\) 2\-way evaluationFigure 8:Concordance between Gold and paraphrased datasets measured by the Pearson correlation coefficientrrand the Spearman rank correlationρ\\rho\. Each point represents a model–illusion pair; the diagonal indicates perfect agreement\.
\(a\) 4\-way evaluation
\(b\) 2\-way evaluationFigure 9:Model rank preservation between Gold and paraphrased datasets measured by the Spearman rank correlationρ\\rhoacross illusion types\.The gold dataset contains 50 hand\-edited examples per illusion, rewritten to most accurately reflect the illusion design, and designed from the character’s perspective\. We also create a large automated dataset based on the gold dataset to enable scalable evaluation\.
To assess the consistency between these datasets, we compare model and illusion\-level results using two complementary measures: the Pearson correlation coefficientrrto assess the linear agreement in effect sizes \(i\.e\., whether the debiased proportionspdebiasedp\_\{\\text\{debiased\}\}track proportionally across datasets \(Figure[8](https://arxiv.org/html/2608.15394#A2.F8)\)\), and the Spearman rank correlationρ\\rhoboth to confirm this under rank\-based assumptions and to evaluate model rank preservation within each illusion \(Figure[9](https://arxiv.org/html/2608.15394#A2.F9)\)\. Overall, the paraphrased dataset provides a strong approximation of the Gold dataset, particularly for character\-based evaluations\. In this setting, both illusion\-level effect sizes and model rankings are highly consistent across both measures for both evaluation types \(2\-way and 4\-way\)\.
In contrast, reader\-based evaluations show weaker agreement between the Gold and paraphrased datasets\. One possible explanation is that modeling temporal perception from a reader’s perspective is inherently more challenging and less consistent for LLMs than reasoning from a character’s point of view; however, this hypothesis requires further validation, for example, through targeted human evaluation\.
Comparing evaluation protocols, the 4\-way evaluation consistently demonstrates higher fidelity to the Gold dataset and greater sensitivity to illusion\-specific differences\. By contrast, the 2\-way evaluation provides a simpler approach but reduces sensitivity in more challenging settings\. Taken together, these results suggest that 2\-way is sufficient for coarse comparisons, whereas 4\-way evaluation remains preferable for fine\-grained analysis and for capturing subtle behavioral differences across models and illusion types\. We perform all analyses for both the 4\-way and the 2\-way evaluation\. \(Appendix[B](https://arxiv.org/html/2608.15394#A2)\)\.
### B\.4Effects of Additional Factors: Positional Bias, Complexity, and Non\-Committal Answers
Beyond illusion type and perspective, we analyze how additional factors, such as positional bias and prompt complexity \(short vs\. long\), affect model behavior\. Although we used Qwen\-32B to generate paraphrased options for the large\-scale dataset, its behavior does not appear to systematically differ from that of Qwen\-8B \(Figure[15](https://arxiv.org/html/2608.15394#A2.F15)\)\. This suggests that the use of a Qwen model for data generation does not introduce a strong model bias in the observed results\.
We observe substantialpositional bias, particularly in the smaller open\-source models\. The smallest model within each open\-source family consistently has the largest positional bias\. Larger and proprietary models show substantially smaller gaps \(Table[5](https://arxiv.org/html/2608.15394#A2.T5)\)\. To mitigate these effects, we apply a debiasing strategy used in[19](https://arxiv.org/html/2608.15394#bib.bib43)that symmetrizes results across option orderings\. A paired McNemar test confirms substantial position bias for many \(mainly open\-source\) models\.
Positional influence appears more pronounced in models when there is greater uncertainty among the leading candidate choices\([27](https://arxiv.org/html/2608.15394#bib.bib41)\), demonstrated throughnoncommittal answers\. Prior research indicates that introducing "None of the Above" \(NA\) as a correct option can result in performance decrements of 30–50%, as models often struggle to reject all provided choices\([32](https://arxiv.org/html/2608.15394#bib.bib42)\)\. In our evaluation, smaller models exhibit both greater positional sensitivity \(Table[5](https://arxiv.org/html/2608.15394#A2.T5)\) and higher rates of non\-committal responses \(Figure[10](https://arxiv.org/html/2608.15394#A2.F10)\), suggesting that underlying uncertainty may modulate positional bias\. Conversely, larger, more capable models demonstrate less dependence on option ordering and a lower propensity to abstain from a definitive choice\. Taken together, these findings suggest that positional bias and non\-committal behaviors may be linked to model uncertainty, potentially serving as heuristic strategies when a clear answer cannot be determined\.
Table 5:We measure positional bias as the difference in selecting the experimental option when it appears first versus second, reported for the 4\-way \(top block\) and 2\-way \(bottom block\) evaluations\.GPT\-20B, especially, and the GPT\-OSS family as a whole, show a strong tendency toward non\-committal responses \(Figure[10](https://arxiv.org/html/2608.15394#A2.F10)\), with up to 85% of answers falling into the same category at times\. In contrast, Qwen models are far less likely to avoid a decision, staying around∼\\sim10% non\-selective overall and primarily selecting can’t tell rather than defaulting to same\. In fact, Qwen is the only model family with a substantial number of responses where it explicitly refuses to select an answer\. Gemma models, by comparison, almost always commit to one of the available options, rarely using either non\-selective category\. Models were generally more likely to select non\-committal responses when evaluating events from the character’s perspective\. Smaller models were more likely to select a non\-committal answer\. Additionally, models tended to default to only one of the two non\-committal options, highlighting differences in whether they prefer to force a selection or admit uncertainty\. This reveals a clear behavioral distinction: GPT models tend to resolve ambiguity by avoiding a decision, Qwen prefers to admit uncertainty, and Gemma consistently forces a choice\.
Across illusion types,prompt lengthhas a relatively limited and inconsistent effect on model behavior \(Figure[12](https://arxiv.org/html/2608.15394#A2.F12)\) across most illusion types\. In most cases, both short and long prompts preserve the same overall pattern of illusion strength, suggesting that prompt length does not substantially alter the design\. The insignificance of length affects all illusions except for filled duration, which shows that the models respond to length as we expect: the more text there is, the more filling\. This suggests that the length\-matching condition, although not exact, was effective enough for the other illusions\.
One exception is the Filled Duration effect, in which longer narratives are more likely to yield selections aligned with the experimental option\. This is consistent with the nature of the illusion: longer textual descriptions have more context, filling a small duration of time, thereby strengthening the effect\. In reverse, as the character, longer descriptions or more items may seem to elongate the duration further\. As a result, increased narrative length may directly impact perceived duration\. Overall, while prompt length can modulate specific illusion types, it does not substantially change the relative ordering or general pattern of effects across conditions\.
Figure 10:Model response distributions across illusion types and perspectives for 4\-way analysis\.Figure 11:Model response distributions across illusion types and perspectives for 2\-way evaluation\.Figure 12:Comparison of model performance between short and long prompts across illusion types under 4\-way evaluation\. While certain models and illusion types show small differences, the overall pattern of effects remains largely consistent, indicating that prompt length has limited impact on the underlying signal captured by the model\.Figure 13:Model performance across different prompt lengths for two\-way analysis\.
### B\.5Additional LLM Analysis Visualizations
Figure 14:The most common words shown by models in their thinking tokens for each illusion category\.Figure 15:Model agreement with the literature across illusion types under 4\-way evaluation\. Bars represent the decisive\-only proportionpexpp\_\{\\text\{exp\}\}of responses selecting the literature\-predicted scenario; whiskers are 95% cluster\-robust confidence intervals\. Cells marked×\\timesare not significant under Benjamini–Hochberg correction at a false discovery rate ofq<0\.05q<0\.05\.Figure 16:Model performance across illusion types for 2\-way evaluation with positional\-bias correction\. Bars represent the debiased proportionpdebiasedp\_\{\\text\{debiased\}\}for each model and perspective \(character vs\. reader\); whiskers are 95% cluster\-robust confidence intervals\. Cells marked×\\timesare not significant under Benjamini–Hochberg correction at a false discovery rate ofq<0\.05q<0\.05\.Figure 17:Comparison of reader and character perspectives under 4\-way evaluation across illusion types\. Each line connects the same model evaluated under the two perspectives\.Figure 18:Comparison of reader and character perspectives under 2\-way evaluation across illusion types\. Each line connects the same model evaluated under the two perspectives\.Figure 19:Comparison of reader and character perspectives under 2\-way evaluation across illusion sub\-categories\. Specific sub\-categories show differences in agreement with the literature within each illusion\.Figure 20:Comparison of reader and character perspectives across illusion sub\-categories\. Specific sub\-categories show differences in agreement with the literature\.Figure 21:The most common words shown by models in their thinking tokens for each illusion category for 2\-way evaluation\.
### B\.6Additional Human Analysis Visualizations
Figure 22:Comparison of reader and character perspectives under 2\-way evaluation across illusion types for human evaluation\.Figure 23:Demographic distribution of human evaluation\.Figure 24:Position bias of human evaluation\.Similar Articles
Temporal Preference Concepts and their Functions in a Large Language Model
This paper causally localizes a subgraph for temporal preference in a distilled LLM, finding that the model discounts the future less steeply than humans and that steering vectors can shift temporal preference, highlighting the need for explicit control mechanisms.
Do Large Language Models Always Tell The Same Stories?
This paper investigates whether large language models generate diverse stories. Using narrative similarity analysis, the authors find that LLM-generated narratives are consistently more similar to each other than human-written stories, and that common mitigation strategies like negative prompting and temperature scaling fail to address this homogeneity.
Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models
This paper investigates whether long-context language models exhibit episodic-like order memory similar to humans, using a temporal order memory task based on a novel. The authors find that models show the same distance effect and uncover a single attention head reinstating temporal context, supporting temporal context reinstatement as a mechanism for episodic-like memory in LLMs.
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning
The paper constructs a benchmark to evaluate LLMs on temporal legal reasoning, revealing biases towards applying the most recently enacted laws and an inverse relationship between general reasoning ability and temporal performance.
Human-Like Anaphor Resolution in Large Language Models
This paper investigates whether five open-weight LLMs exhibit human-like sensitivity to psycholinguistic factors in anaphor resolution, using surprisal and comprehension accuracy as behavioral measures. Results show selective cognitive alignment, with some models matching human discourse sensitivity but not semantic interference effects.