Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark

arXiv cs.CL Papers

Summary

This paper introduces a new benchmark for evaluating LLMs on Polish high school history exit exams (Matura), showing that models outperform humans on aggregate scores but exhibit distinct failure modes in source interpretation and temporal reasoning.

arXiv:2608.12343v1 Announce Type: new Abstract: AI chatbots are widely used by students as knowledge sources, yet LLM benchmarks rarely assess interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exams (Matura) in history - three official papers from 2023-2025, comprising short-answer questions and extended essays - comparing model performance against the human examinee population. Every model dramatically outperforms human examinees, yet aggregate scores mask distinct competency profiles: rankings are unstable across task type, source modality, and geographical scope, with a consistent penalty on Polish versus Global history content. Qualitative error analysis reveals two recurring failure modes - source conflation, in which models reason from source content rather than treating it as an object of analysis, and temporal disorientation, in which responses are historically misplaced. This study introduces the first LLM history benchmark grounded in Polish national curriculum.
Original Article
View Cached Full Text

Cached at: 08/14/26, 09:25 AM

# Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark
Source: [https://arxiv.org/html/2608.12343](https://arxiv.org/html/2608.12343)
Adrian Trzoss1,6Kacper Dudzic1,2,3Wiktor Werner1Marcin Moskalewicz1,4,5

1Adam Mickiewicz University, Poznań, Poland 2IDEAS Research Institute, Warsaw, Poland 3AMU Center for Artificial Intelligence, Poznań, Poland 4Poznań University of Medical Sciences, Poznań, Poland 5Maria Curie\-Skłodowska University, Lublin, Poland 6WSB Merito University, Poznań, Poland

###### Abstract

AI chatbots are widely used by students as knowledge sources, yet LLM benchmarks rarely assess interpretative historical reasoning\. We evaluate eight leading LLMs on the Polish high school exit exams \(Matura\) in history—three official papers from 2023–2025, comprising short\-answer questions and extended essays—comparing model performance against the human examinee population\. Every model dramatically outperforms human examinees, yet aggregate scores mask distinct competency profiles: rankings are unstable across task type, source modality, and geographical scope, with a consistent penalty on Polish versus Global history content\. Qualitative error analysis reveals two recurring failure modes—source conflation, in which models reason from source content rather than treating it as an object of analysis, andtemporal disorientation, in which responses are historically misplaced\. This study introduces the first LLM history benchmark grounded in Polish national curriculum111We release our code at:[https://github\.com/kdudzic/history\-matura\-llm\-evaluation](https://github.com/kdudzic/history-matura-llm-evaluation)\.

Large Language Models Pass the History Exam But Miss the <<History\>\>: A Polish High School Exit ExamMaturaBenchmark

Adrian Trzoss1,6††thanks:Correspondence:adrian\.trzoss@amu\.edu\.plKacper Dudzic1,2,3Wiktor Werner1Marcin Moskalewicz1,4,51Adam Mickiewicz University, Poznań, Poland2IDEAS Research Institute, Warsaw, Poland3AMU Center for Artificial Intelligence, Poznań, Poland4Poznań University of Medical Sciences, Poznań, Poland5Maria Curie\-Skłodowska University, Lublin, Poland6WSB Merito University, Poznań, Poland

## 1Introduction

AI chatbots powered by Large Language Models have become a significant source of knowledge in educational settings, widely adopted by secondary school and university students alike\. A growing body of researchGilsonet al\.\([2023](https://arxiv.org/html/2608.12343#bib.bib22)\); Locatelliet al\.\([2024](https://arxiv.org/html/2608.12343#bib.bib24)\); Hendryckset al\.\([2020](https://arxiv.org/html/2608.12343#bib.bib23)\); Darg̀iset al\.\([2024](https://arxiv.org/html/2608.12343#bib.bib25)\)evaluates these models with respect to knowledge and pedagogical value using standardized benchmarks\. However, more interpretative assessments of historical topics, which rarely admit clear\-cut answers, remain substantially underrepresented\. This paper evaluates the performance of leading LLMs on the Polish high school exit exams \(Matura\) in history, prepared by the state\-run Central Examination Board \(Centralna Komisja Egzaminacyjna, CKE\)222[https://cke\.gov\.pl/egzamin\-maturalny/egzamin\-maturalny\-w\-formule\-2023/](https://cke.gov.pl/egzamin-maturalny/egzamin-maturalny-w-formule-2023/)\. We test eight models on official examination papers from 2023–2025 and assess performance along two dimensions: comparing models against one another, and against human examinees\. Matriculation examinations in history are a methodologically valid benchmark of historical skills because they assess both factual knowledge and historical reasoning within the national curriculum framework\. Moreover, the question pool changes annually, and tasks are not designed with LLMs in mind, reducing research design bias\. Finally, each paper is accompanied by a standardized marking scheme with model answers and grading criteria, enabling objective scoring\. As a result, the benchmark represents a rare instance of an open\-ended humanities task with an externally defined and relatively objective evaluation framework\.

We investigate whether LLMs outperform human examinees overall, whether performance varies systematically by question type, source modality, and historical scope—including a predicted penalty on Polish versus Global history contentDadaset al\.\([2025](https://arxiv.org/html/2608.12343#bib.bib15)\)\. Additionally, we examine whether certain task formulations prove systematically incomprehensible under baseline prompting conditions\. We situate this study in relation to two bodies of prior work: global history benchmarks, which have not addressed non\-Anglophone educational content, and Polish\-language NLP benchmarks, which have not addressed open\-ended humanities tasks\.

## 2Related Work

Hauser introduced HiST\-LLM benchmarkHauseret al\.\([2024](https://arxiv.org/html/2608.12343#bib.bib17)\), demonstrating substantial variation in model performance across historical regions and periods\. Chartier extended this line of work with HiBenchLLMChartieret al\.\([2025](https://arxiv.org/html/2608.12343#bib.bib18)\), confirming that models perform less well on non\-Anglophone content—including French\-language material\. These benchmarks, however, rely on arbitrarily constructed question sets and manually designed evaluation criteria, and neither relates model performance to human baselines\. Within the Polish context, LLMzSzŁ \(LLMs Behind the School Desk\) by Jassem is the first large\-scale LLM benchmark for high school examsJassemet al\.\([2025](https://arxiv.org/html/2608.12343#bib.bib16)\), built from archival examinations across multiple subjects \(apart from history\)\. The authors found that multilingual models frequently outperform monolingual ones, though the latter can be competitive under size constraints\. These findings do not transfer directly to the present study: LLMzSzŁ relies exclusively on closed multiple\-choice questions, which are far easier to evaluate automatically and at scale\. Our study addresses both gaps simultaneously: it is the first benchmark to evaluate LLMs on Polish historical content using authentic open\-ended examination tasks, and the first to compare model performance against a human distributional baseline—the two limitations that characterize all prior work in this space\.

## 3Dataset

The benchmark consists of three official historyMaturapapers from 2023, 2024, and 2025\. Each contains*tasks*—thematic units built around source materials—where each may contain one or more*questions*\. The 2023 paper included 36 questions across 25 tasks; the 2024 paper, 39 questions across 25 tasks; and the 2025 paper, 37 questions across 24 tasks\. The final task in each paper required an essay of at least 600 words on one of three proposed topics\. All papers covered a broad chronological range from antiquity to the late twentieth century, with roughly equal coverage of Polish and Global history\. Short\-answer questions varied widely in format: single\-word or single\-phrase responses \(naming a figure or event\), multiple\-choice items, binary judgment tasks \(true/false\) with brief justificatory reasoning, and source\-analysis exercises \(3–5 sentences\)\. Each*task*—except the essay—included source materials, which appeared in varying combinations: historical texts, iconographic materials \(maps, photographs\), and tables \(economic or genealogical\)\. Most questions were worth one or two points; the maximum total score per paper was 60 points, with 15 allocated to the essay\. Responses were scored according to official CKE grading guidelines and cross\-checked by a high\-level human expert—a CKE\-trained and experiencedMaturaexaminer\.

## 4Methodology

### 4\.1Model Selection

We evaluated a total ofN=8N=8models from 4 leading providers: GPT\-4oOpenAI \([2024](https://arxiv.org/html/2608.12343#bib.bib8)\)and GPT\-5\.4OpenAI \([2026](https://arxiv.org/html/2608.12343#bib.bib12)\)from OpenAI, Claude Sonnet 3\.7Anthropic \([2025](https://arxiv.org/html/2608.12343#bib.bib13)\)and Claude Sonnet 4\.6Anthropic \([2026](https://arxiv.org/html/2608.12343#bib.bib14)\)from Anthropic, Gemini 2\.5 ProGoogle \([2025](https://arxiv.org/html/2608.12343#bib.bib9)\)and Gemini 3\.1 ProGoogle \([2026](https://arxiv.org/html/2608.12343#bib.bib10)\)from Google, as well as Grok 4xAI \([2025](https://arxiv.org/html/2608.12343#bib.bib11)\)and Grok 4\.20 from xAI\.

Our model selection criteria encompassed public interest and general performance, the ability to handle image inputs, availability on OpenRouter333[https://openrouter\.ai/](https://openrouter.ai/), as well as availability through the web interface on free plans of the respective providers\. The last criterion was methodologically motivated by the potential real\-world use of LLMs by high school students\. Similarly, the choice of the two models from each provider was intended to investigate any substantial performance differences between the newest available version and its older equivalent\.

### 4\.2Model Inference Procedure

Model outputs were obtained via the OpenRouter API\. Each question was passed in a separate API call\. The question text was passed along with the prompt \(all in Polish\), whereas the image inputs were included in an additional API call payload\.

Each short\-answer question was passed to each model 3 separate times with the default temperature setting to evaluate potential inconsistencies arising from the non\-deterministic nature of the models\. Similarly, each essay topic was passed 3 times, with all possible topics from a sheet evaluated for each model \(in contrast to a student who must choose one\); this adjustment allowed for a more comprehensive evaluation of model performance on the essay section—topics concern various time periods and problems, and a model is not guaranteed to perform equally well on each\. In total, 2688 API calls for short answer questions \(36 questions×\\times25 tasks from the 2023 sheet, 39×\\times25 from 2024, and 37×\\times24 from 2025; each question set×\\times3 attempts×\\times8 models each\) and 216 for essays \(3 essay topics×\\timessheet×\\times3 attempts×\\times8 models each\) have been made\. No in\-context learning paradigm was employed in the inference protocol\.

### 4\.3Data Preparation and Annotation

Each question was manually annotated along three dimensions: geographical scope \(Polish vs\. Global history\); historical period \(antiquity, medieval, early modern, nineteenth century, twentieth century, PRL\-communist Poland post\-1945\); and source material type \(text only, photo only, photo and text, table\-based\)\.

### 4\.4Quantitative Analysis

Three overall model rankings were computed—all tasks combined, short\-answer only, and essays only—based on mean normalized scores aggregated across all years and runs, with uncertainty estimated via 95% bootstrap confidence intervals \(5,000 replications\)\. Rankings were further aggregated by each annotation category\. For human vs\. model comparisons, we calculated the Wasserstein distance \([D\.3](https://arxiv.org/html/2608.12343#A4.F3),[D\.4](https://arxiv.org/html/2608.12343#A4.F4)\)\. Because all three essay topics are evaluated for each model—compared with only one topic for human examinees—a complete model run yields a maximum of 90 points, whereas the human is 60\. All comparisons, therefore, use normalized scores to account for this asymmetry\.

## 5Results

### 5\.1Short\-answer questions vs\. Essays

Figure[1](https://arxiv.org/html/2608.12343#S5.F1)presents model rankings across three score dimensions, revealing a three\-tier structure: Claude Sonnet 4\.6 and Gemini 3\.1 Pro lead with overlapping confidence intervals \(96\.6% & 96\.2%\), Grok 4 and GPT\-5\.4 form a distinct second tier, and the remaining models cluster tightly with indistinguishable confidence intervals\. Short\-answer performance closely tracks overall scores, confirming it as the primary discriminator between models\. Essay scores are uniformly high for seven models, offering no discriminative power across the benchmark; GPT\-4o is the sole exception at 90\.1%\. The aggregate ranking, however, masks substantial instability across topical and source\-type categories—and remains far above the human examinee average of 44\.1% across all analyzed years \(Table[D\.3](https://arxiv.org/html/2608.12343#A4.T3)\)\.

![Refer to caption](https://arxiv.org/html/2608.12343v1/images/fig_1_task_comparison.png.png)Figure 1:Overall performance of LLMs based on normalized scores aggregated across all years and runs: all tasks combined, short\-answer questions, and essays only\. Error bars represent 95% bootstrap CI\.
### 5\.2Polish vs\. Global History

Figure[2](https://arxiv.org/html/2608.12343#S5.F2)shows model rankings split by geographical scope\. Almost all models score higher on Global than on Polish history tasks, with the penalty on Polish content ranging up to 10\.8 percentage points\. The gap is negligible for the top two models but substantial for the remaining six, where it ranges from 2\.3 to 10\.8 percentage points—suggesting that weaker models are disproportionately disadvantaged by nationally specific content\. The split produces non\-trivial rank reordering: Grok 4, ranked third overall, scores 97\.6% on Global but drops to 91\.7% on Polish, falling behind GPT\-5\.4\. The epoch breakdown \(Figure[A\.1](https://arxiv.org/html/2608.12343#A1.F1)\) mirrors this pattern, with twentieth\-century and PRL\-period showing the largest score variance\.

![Refer to caption](https://arxiv.org/html/2608.12343v1/images/fig_2_polish_global.png)Figure 2:Model rankings by geographical scope \(short\-answer questions\)\. Error bars represent 95% bootstrap confidence intervals\.
### 5\.3Source Type for Question Results

Figure[3](https://arxiv.org/html/2608.12343#S5.F3)presents mean normalized scores across the model×\\timessource\-type matrix\. Photo \+ Text tasks yield the highest and most consistent scores, suggesting a ceiling or redundant\-cue effect\. Text\-only tasks show the widest cross\-model variance \(69\.4%–96\.8%\), with Gemini 2\.5 Pro collapsing to the lowest cell in the matrix\. Gemini 3\.1 Pro leads on Photo\-only \(99\.0%\), Claude Sonnet 4\.6 on Table\-based tasks \(100%\), and Grok 4 on Text\-only \(96\.8%\) despite ranking third overall\. GPT\-4o underperforms frontier models across all types\.

![Refer to caption](https://arxiv.org/html/2608.12343v1/images/fig_3_heatmap.png)Figure 3:Mean normalized scores in the model×\\timessource type matrix\. Photo \+ Text tasks yield the highest and most consistent scores; Text\-only tasks show the widest variance\. Error bars omitted for clarity\.
### 5\.4Hardest Questions

Table[C\.1](https://arxiv.org/html/2608.12343#A3.T1)lists the three hardest and most discriminating questions for models\. They are situated mainly in the 2024 and 2025 papers and concern either temporal\-oriented reasoning or culturally specific Polish content from the twentieth century and the PRL period\. Question11\_02\_2025is both in the hardest and most model\-discriminating items in the benchmark: "Determine which of the documents cited in fragments A–C was created first\. Justify your answer by referring to the sources and your own knowledge" \(See Appendix:[E](https://arxiv.org/html/2608.12343#A5)\)\. Models correctly recognized sources names but failed to order them in chronological sequence\. Both Gemini and Grok 4 scored max points across all runs, GPT\-5\.4 scored a point only in one run, while others scored 0 across all three tests\. A cross\-population comparison reveals an inversion: per CKE reports, human\-hardest items disproportionately involve Photo \+ Text combinations \(Table[C\.2](https://arxiv.org/html/2608.12343#A3.T2)\)—precisely the question type on which models perform most consistently \(Figure[3](https://arxiv.org/html/2608.12343#S5.F3)\)\.

### 5\.5Failure Modes

Manual inspection of all model outputs reveals two recurring failure modes\. The first, which we deemsource conflation, occurs when a model reasons from the semantic content of a provided source rather than treating it as an object of historical analysisBasmovet al\.\([2024](https://arxiv.org/html/2608.12343#bib.bib21)\)\. In question07\_2024\([E\.4](https://arxiv.org/html/2608.12343#A5.T4)\), models consistently inferred chronological order from content comparison rather than authorial context, effectively treating sources as evidence about the world rather than as historically situated documents\. This was observed across six models on all test runs; the only exceptions were Gemini 3\.1 Pro and GPT\-5\.4\. The second, we calltemporal disorientation, concerns questions requiring models to identify or sequence causally linked events, or locate them within a specific Polish historical period\. On task22\_2025, models produced partially plausible responses—correctly identifying relevant actors or concepts—but placed them in the wrong period or orderHerelet al\.\([2024](https://arxiv.org/html/2608.12343#bib.bib20)\); Fatemiet al\.\([2025](https://arxiv.org/html/2608.12343#bib.bib19)\)\. This pattern was particularly pronounced on PRL\-period questions \(questions23\_2024,25\_2024\)\. The essay component reveals differences in rhetorical strategies\. While all models achieve broadly comparable essay scores \(see Figure[1](https://arxiv.org/html/2608.12343#S5.F1)\), their argumentative strategies diverge\. The majority of models align with the proposed thesis, constructing confirmatory arguments\. The Grok family is the sole exception: both systematically adopt a counter\-argumentative position, particularly in the 2023 exam \([F\.5](https://arxiv.org/html/2608.12343#A6.T5)\)\.

## 6Conclusions

This study presents the first benchmark evaluation of LLMs on the Polish historyMatura\. Every model dramatically outperforms human examinees, yet aggregate scores mask distinct competency profiles: rankings are unstable across task type, source modality, and geographical scope\. The Polish versus Global history split produces systematic rank reordering, extending Chartier’s findings on non\-Anglophone content to a Central\-Eastern European contextChartieret al\.\([2025](https://arxiv.org/html/2608.12343#bib.bib18)\)\. The two identified failure modes suggest that what models lack on the hardest questions is not factual coverage but historically situated reasoning—the ability to reason within a period rather than merely about it\. Essay argumentation further reveals rhetorical divergence: while most models align with the proposed thesis, Grok models systematically adopt a counter\-argumentative position\. Nationally grounded, human\-referenced benchmarks are necessary to complement synthetic evaluation frameworks, particularly for languages and domains where global models remain undertested\. After all, passing an examination is not the same as understanding its subject\.

## 7Limitations

The benchmark comprises three examination papers, which limits statistical power and may not capture the full range of question types across years\. Official CKE reports provide only aggregate human score distributions rather than individual\-level data, constraining the precision of human\-model comparisons\. All evaluated models are closed\-source, precluding analysis of the behavioral and architectural factors underlying observed failure modes\. Potential contamination cannot be ruled out, as examination papers and official model answers are publicly available online and may appear in pretraining corpora\. All models were prompted under baseline conditions without chain\-of\-thought or few\-shot examples, meaning results reflect but one point in a broader prompting space\. Finally, no Polish monolingual model was included: Bielik, the most capable Polish\-language modelDadaset al\.\([2025](https://arxiv.org/html/2608.12343#bib.bib15)\), lacks multimodal support and could not be evaluated on multimodal tasks\.

## 8Ethical Considerations

This work presents a methodological contribution and does not involve human subjects or personally identifiable information\. All experiments were conducted using publicly available resources under their respective licenses\. The human expert—a CKE\-trained and experiencedMaturaexaminer—worked voluntarily\.

## Acknowledgments

This research was supported by the National Science Centre under the Miniatura\-9 project \(Grant No\. 2025/09/X/HS3/00230\)\. We sincerely thank Mrs\. Magdalena Łysakowska, who served as an external expert and a CKE\-trained matura examiner\. We would like to express our deepest gratitude to Mr\. Łukasz Nierzwicki PhD, and Mr\. Cyprian Kleist MSc for their Data Science consultations\.

## References

- Anthropic \(2025\)Claude 3\.7 Sonnet System Card\.External Links:[Link](https://www-cdn.anthropic.com/9ff93dfa8f445c932415d335c88852ef47f1201e.pdf)Cited by:[§4\.1](https://arxiv.org/html/2608.12343#S4.SS1.p1.1)\.
- Anthropic \(2026\)Claude 4\.6 Sonnet System Card\.External Links:[Link](https://www-cdn.anthropic.com/bbd8ef16d70b7a1665f14f306ee88b53f686aa75.pdf)Cited by:[§4\.1](https://arxiv.org/html/2608.12343#S4.SS1.p1.1)\.
- V\. Basmov, Y\. Goldberg, and R\. Tsarfaty \(2024\)LLMs’ reading comprehension is affected by parametric knowledge and struggles with hypothetical statements\.arXiv preprint arXiv:2404\.06283\.Cited by:[§5\.5](https://arxiv.org/html/2608.12343#S5.SS5.p1.1)\.
- M\. Chartier, N\. Dakkoune, G\. Bourgeois, and S\. Jean \(2025\)HiBenchLLM: historical inquiry benchmarking for large language models\.Data & Knowledge Engineering156,pp\. 102383\.Cited by:[§2](https://arxiv.org/html/2608.12343#S2.p1.1),[§6](https://arxiv.org/html/2608.12343#S6.p1.1)\.
- S\. Dadas, M\. Grębowiec, M\. Perełkiewicz, and R\. Poświata \(2025\)Evaluating polish linguistic and cultural competency in large language models\.InInternational Conference on Artificial Intelligence and Soft Computing,pp\. 60–71\.Cited by:[§1](https://arxiv.org/html/2608.12343#S1.p2.1),[§7](https://arxiv.org/html/2608.12343#S7.p1.1)\.
- R\. Darg̀is, G\. Barzdins, I\. Skadiņa, N\. Gruzitis, and B\. Saulīte \(2024\)Evaluating open\-source llms in low\-resource languages: insights from latvian high school exams\.InProceedings of the 4th International Conference on Natural Language Processing for Digital Humanities,pp\. 289–293\.Cited by:[§1](https://arxiv.org/html/2608.12343#S1.p1.1)\.
- B\. Fatemi, S\. M\. Kazemi, A\. Tsitsulin, K\. Malkan, J\. Yim, J\. Palowitch, S\. Seo, J\. Halcrow, and B\. Perozzi \(2025\)Test of time: a benchmark for evaluating llms on temporal reasoning\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 94426–94447\.Cited by:[§5\.5](https://arxiv.org/html/2608.12343#S5.SS5.p1.1)\.
- A\. Gilson, C\. W\. Safranek, T\. Huang, V\. Socrates, L\. Chi, R\. A\. Taylor, and D\. Chartash \(2023\)How does chatgpt perform on the united states medical licensing examination \(usmle\)? the implications of large language models for medical education and knowledge assessment\.JMIR medical education9,pp\. e45312\.Cited by:[§1](https://arxiv.org/html/2608.12343#S1.p1.1)\.
- Google \(2025\)Gemini 2\.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities\.External Links:2507\.06261,[Link](https://arxiv.org/abs/2507.06261)Cited by:[§4\.1](https://arxiv.org/html/2608.12343#S4.SS1.p1.1)\.
- Google \(2026\)Gemini 3\.1 Pro Model Card\.External Links:[Link](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card.pdf)Cited by:[§4\.1](https://arxiv.org/html/2608.12343#S4.SS1.p1.1)\.
- J\. Hauser, D\. Kondor, J\. Reddish, M\. Benam, E\. Cioni, F\. Villa, J\. S\. Bennett, D\. Hoyer, P\. Francois, P\. Turchin,et al\.\(2024\)Large language models’ expert\-level global history knowledge benchmark \(hist\-llm\)\.Advances in Neural Information Processing Systems37,pp\. 32336–32369\.Cited by:[§2](https://arxiv.org/html/2608.12343#S2.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2020\)Measuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.Cited by:[§1](https://arxiv.org/html/2608.12343#S1.p1.1)\.
- D\. Herel, V\. Bartek, J\. Jirak, and T\. Mikolov \(2024\)Time awareness in large language models: benchmarking fact recall across time\.arXiv preprint arXiv:2409\.13338\.Cited by:[§5\.5](https://arxiv.org/html/2608.12343#S5.SS5.p1.1)\.
- K\. Jassem, M\. Ciesiółka, F\. Graliński, P\. Jabłoński, J\. Pokrywka, M\. Kubis, M\. Jabłońska, and R\. Staruch \(2025\)LLMzSzł: a comprehensive llm benchmark for polish\.arXiv preprint arXiv:2501\.02266\.Cited by:[§2](https://arxiv.org/html/2608.12343#S2.p1.1)\.
- M\. S\. Locatelli, M\. P\. Miranda, I\. J\. da Silva Costa, M\. T\. Prates, V\. Thomé, M\. Z\. Monteiro, T\. Lacerda, A\. Pagano, E\. R\. Neto, W\. Meira Jr,et al\.\(2024\)Examining the behavior of llm architectures within the framework of standardized national exams in brazil\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,Vol\.7,pp\. 879–890\.Cited by:[§1](https://arxiv.org/html/2608.12343#S1.p1.1)\.
- OpenAI \(2024\)GPT\-4o System Card\.External Links:2410\.21276,[Link](https://arxiv.org/abs/2410.21276)Cited by:[§4\.1](https://arxiv.org/html/2608.12343#S4.SS1.p1.1)\.
- OpenAI \(2026\)GPT\-5\.4 Thinking System Card\.External Links:[Link](https://deploymentsafety.openai.com/gpt-5-4-thinking/gpt-5-4-thinking.pdf)Cited by:[§4\.1](https://arxiv.org/html/2608.12343#S4.SS1.p1.1)\.
- xAI \(2025\)Grok 4 Model Card\.External Links:[Link](https://data.x.ai/2025-08-20-grok-4-model-card.pdf)Cited by:[§4\.1](https://arxiv.org/html/2608.12343#S4.SS1.p1.1)\.

## Appendix AModel Performance by Historical Epoch

![Refer to caption](https://arxiv.org/html/2608.12343v1/images/epochs_chronological_overall_appendix.png)Figure A\.1:Model rankings by historical epoch \(short\-answer questions only\)\. Aggregate benchmark rankings conceal substantial chronological specialization: performance is near\-ceiling on antiquity tasks, but diverges markedly on twentieth\-century and PRL\-period content\. The largest variance occurs on modern Polish history\. Error bars represent 95% bootstrap confidence intervals\.
## Appendix BRun\-Level Score Distributions by Model and Year

![Refer to caption](https://arxiv.org/html/2608.12343v1/images/level_run_score_appendix.png)Figure B\.2:Run\-level normalized score distributions by model across all examination years \(2023–2025\)\. The highest\-ranked models combine near\-ceiling median performance with comparatively tight inter\-run variance, suggesting stable behavior across historical topics and exam formulations\. Lower\-ranked models exhibit wider dispersion and greater year sensitivity, particularly GPT\-4o\. The overlap among mid\-tier models contrasts with the clearer separation of the top frontier systems\.
## Appendix CHardest and Most Discriminating Questions

Table C\.1:Hardest and most discriminating questions for models\. Mean and SD are normalized scores across all models and runs\.Table C\.2:Hardest questions for human examinees based on official CKE reports, together with source\-type\.
## Appendix DHuman–Model Score Distribution Comparison

Table D\.3:Bootstrap comparison between human and model score distributions across examination years\. Columns report normalized mean scores \(%\), bootstrap mean differences \(Diff\), Wasserstein distances \(W\-dist\), Kolmogorov–Smirnov \(KS\) statistics, and the number of model observations\.![Refer to caption](https://arxiv.org/html/2608.12343v1/images/human_vs_models_matrix_appendix.png)Figure D\.3:Wasserstein distance matrix for normalized score distributions of humans and models across all examination years\. Frontier models cluster into several low\-distance groups despite moderate ranking differences, indicating broadly similar distributional behavior\. The human distribution forms a clearly isolated cluster, with substantially larger distances to every evaluated model than any inter\-model comparison\.![Refer to caption](https://arxiv.org/html/2608.12343v1/images/fig_wasserstein.png)Figure D\.4:Model performance versus Wasserstein distance from the human examinee distribution\. Human examinees occupy a distinct low\-score reference position \(44\.1%\), whereas all evaluated models cluster between approximately 87% and 98% normalized performance\. Bubble size corresponds to standard deviation across runs\.
## Appendix ESelected Benchmark Questions with English Translations

Table E\.4:Selected benchmark questions in Polish and English translation\.
## Appendix FGrok Models Narrative Strategies

Table F\.5:Comparative stance alignment of Grok 4 and Grok 4\.20 across repeated essay generations

Similar Articles

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

arXiv cs.CL

This paper proposes a multidimensional evaluation framework for assessing statistical reasoning in large language models, combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis across 15 LLMs and 90 exam questions. It finds that accuracy alone is insufficient to characterize LLM statistical reasoning and that vendor-specific stylistic differences exist.