LLM Performance on a Real, Double-Marked GCSE Benchmark

arXiv cs.CL Papers

Summary

Introduces a dataset of 32,534 double-marked GCSE student responses across five subjects, finding that top-performing LLMs agree with examiners more closely than examiners agree with each other, including on handwritten and subjective tasks.

arXiv:2606.24973v1 Announce Type: new Abstract: We introduce a dataset of 32,534 double-marked real student responses to GCSE mock exams (GCSEs are the UK's national exams, taken at age ~16), spanning 328 questions across five subjects and including handwritten work. We test whether off-the-shelf large language models agree with examiners as closely as the two examiners agree with each other. We find that models overwhelmingly agree well with the examiner consensus across subjects, with the top performing models agreeing more closely with examiners than examiners agree with each other. Models achieve high scores for subjective tasks like English essay marking, as well as handling complex and messy handwritten Maths paper scripts. Agreement is uniform near the examiner line, and not massively discriminated by model size, providing cost-effective automated marking solutions.
Original Article
View Cached Full Text

Cached at: 06/25/26, 05:09 AM

# LLM Performance on a Real, Double-Marked GCSE Benchmark
Source: [https://arxiv.org/html/2606.24973](https://arxiv.org/html/2606.24973)
Malachy Fox Medly AI malachy@medlyai\.com&Kavi Samra Medly AI kavi@medlyai\.com&Paul Jung Medly AI paul@medlyai\.com

\(9 June 2026\)

###### Abstract

We introduce a dataset of 32,534 double\-marked real student responses to GCSE mock exams \(GCSEs are the UK’s national exams, taken at age 16\), spanning 328 questions across five subjects and including handwritten work\. We test whether off\-the\-shelf large language models agree with examiners as closely as the two examiners agree with each other\. We find that models overwhelmingly agree well with the examiner consensus across subjects, with the top performing models agreeing more closely with examiners than examiners agree with each other\. Models achieve high scores for subjective tasks like English essay marking, as well as handling complex and messy handwritten Maths paper scripts\. Agreement is uniform near the examiner line, and not massively discriminated by model size, providing cost\-effective automated marking solutions\.

*Keywords*Automated essay scoring⋅\\cdotLarge language models⋅\\cdotEducational assessment⋅\\cdotLLM evaluation⋅\\cdotQuadratic weighted kappa

## 1Introduction

Exam boards and schools require reliable marking\[[9](https://arxiv.org/html/2606.24973#bib.bib9)\], but manual assessment imposes a heavy workload on teachers\[[10](https://arxiv.org/html/2606.24973#bib.bib10)\]\. Automated essay scoring has a long history, progressing from early feature\-based and neural systems\[[3](https://arxiv.org/html/2606.24973#bib.bib3),[4](https://arxiv.org/html/2606.24973#bib.bib4),[6](https://arxiv.org/html/2606.24973#bib.bib6),[7](https://arxiv.org/html/2606.24973#bib.bib7)\]to recent large language model approaches\[[8](https://arxiv.org/html/2606.24973#bib.bib8),[5](https://arxiv.org/html/2606.24973#bib.bib5),[13](https://arxiv.org/html/2606.24973#bib.bib13),[14](https://arxiv.org/html/2606.24973#bib.bib14)\], including subject\-specific scoring such as science assessment\[[15](https://arxiv.org/html/2606.24973#bib.bib15)\]\. Early results were mixed: prompted off\-the\-shelf models fell short of task\-specific systems on standard essay benchmarks\[[13](https://arxiv.org/html/2606.24973#bib.bib13)\], and later work recovered ground by using the model to generate rubric\-grounded features rather than to score directly\[[14](https://arxiv.org/html/2606.24973#bib.bib14)\]\. Most school examinations, though, span multiple subjects and include handwritten equations, short text answers, and diagrams, which automated systems find hard to mark reliably\[[11](https://arxiv.org/html/2606.24973#bib.bib11),[12](https://arxiv.org/html/2606.24973#bib.bib12)\]; even on born\-digital multimodal mathematics, leading models still trail humans\[[17](https://arxiv.org/html/2606.24973#bib.bib17)\]\. We ask whether commercially available LLMs, run under a generic prompt at minimum reasoning effort, are already a reliable second marker across these response types, and where they fail\.

We present a benchmark of 32,534 double\-marked GCSE mock responses across 328 questions in English Language, Maths, Biology, Chemistry, and Physics\. GCSEs are the national subject examinations taken by students in England, Wales and Northern Ireland at age∼16\\sim 16, marked by qualified examiners against published mark schemes; the “Higher” tier labelled in the tables and worked examples is the more demanding of the two GCSE tiers in Maths and the sciences\. To our knowledge it is the first publicly described double\-marked, multi\-subject GCSE marking benchmark to include handwritten student work alongside typed text\. We evaluate models by their average agreement with each examiner, measured against the agreement between the two examiners themselves\. This isolates whether a model marks as consistently as a human marker\.

We report agreement subject by subject, and investigate how models are differentiated by bias when marking English essays\.

## 2The Dataset

The dataset comprises 32,534 student responses across 328 questions in five subjects \(English Language, Maths, Biology, Chemistry, and Physics\), sampled to span the full attainment range\. Each answer was independently marked by two qualified examiners\. Answers combine typed text, on\-screen text boxes, and freehand handwriting and drawings captured as strokes, so a marker must be multimodal\. The two worked examples below show the range: a typed English essay, and a handwritten Maths item\. Table[1](https://arxiv.org/html/2606.24973#S2.T1)gives the share of handwritten input per subject\.

English Language \(creative writing, 40 marks\)Prompt to the model\.“Mark the student’s answer using the mark scheme provided\. Respond with JSON only, in the form\{"mark": <integer\>\}\.”This is sent with the question, mark scheme and student answer below\.Question\.Your local newspaper is running a creative writing competition\.*Either:*write a description of a storm at sea as suggested by a picture;*or:*write a story about an unexpected visitor\.Mark scheme\.Two assessment objectives, marked separately on best\-fit level descriptors\.AO5, content and organisation\(24 marks\): communicate clearly and imaginatively, matching tone, style and register to purpose and audience, with coherent structure\.AO6, technical accuracy\(16 marks\): a range of vocabulary and sentence structures, accurate spelling and punctuation\.Student answer\.It was a quiet room, I was doing my regular daily routines as normality invaded me… Then I heard it, a loud knock that invaded my conscience; the sound didn’t belong and my chest was beating faster than ever\. \[…\] For a second, I didn’t recognise him: he looked thinner somehow, sharper around the edges like time had stripped something away\. \[…\] He stepped into the dark and whispered, “some consequences follow you home\.”Examiner marks: 19/40 \(both examiners; AO5 \+ AO6 combined\)\.

Maths Higher \(highest common factor, 2 marks\)Prompt to the model\.“Mark the student’s answer using the mark scheme provided\. Respond with JSON only, in the form\{"mark": <integer\>\}\.”This is sent with the question, mark scheme and student answer below\. The rendered canvas image is attached alongside it\.Question\.Find the highest common factor \(HCF\) of 84 and 126\.Mark scheme\.M1for a correct method, e\.g\. the prime factorisation of either number \(84=2×2×3×784=2\\times 2\\times 3\\times 7or126=2×3×3×7126=2\\times 3\\times 3\\times 7\)\.A1for4242\(or2×3×72\\times 3\\times 7\)\.Student answer\.The student worked by hand on the canvas: trial division and factor trees to find the prime factors of each number, the shared factors combined, and the answer4242circled \(see below\)\.![[Uncaptioned image]](https://arxiv.org/html/2606.24973v1/charts/example_maths_canvas.png)Examiner marks: 2/2 \(both examiners\)\. AI mark: 2/2\.

Table 1:Share of responses containing handwritten input \(on\-screen handwriting captured as strokes\)\.
## 3Method

Each response was marked with a generic prompt: the question, mark scheme, and student answer, with a one\-line instruction to mark it out of the maximum \(Figure[1](https://arxiv.org/html/2606.24973#S3.F1)\)\. For handwritten work the rendered canvas image is attached, so the model must read the handwriting itself\. The model is constrained to structured output, a JSON schema with a single integermarkfield\. Models were run at minimum reasoning effort\.111“Minimum” is the lowest reasoning setting each model exposes: this disables reasoning entirely for every model except Gemini 3\.1 Pro, whose lowest available setting is ‘low’ rather than ‘off’\.

Marks are ordinal, so agreement is measured with Quadratic Weighted Kappa \(QWK\)\[[1](https://arxiv.org/html/2606.24973#bib.bib1)\], computed per question and then averaged across questions*weighted by each question’s maximum mark*\. This matches how a student’s grade is formed, a sum of marks in which a 6\-mark question counts six times a 1\-mark one, and it damps the noise of the many near\-binary low\-mark items\. QWK runs from 0 \(chance agreement\) to 1 \(identical marks\)\. On the standard scale\[[2](https://arxiv.org/html/2606.24973#bib.bib2)\]0\.21–0\.40 is fair, 0\.41–0\.60 moderate, 0\.61–0\.80 substantial, and 0\.81–1\.00 almost perfect, so a QWK around 0\.8 already indicates close agreement between two raters\.

generic prompt \(one multimodal message\)Instruction“Mark the answer using themark scheme; reply\{"mark": n\}”QuestionMark schemeStudent answer\(typed text\)Canvas image\(if handwritten\)LLM\(minimumreasoning\)Structured output\{"mark": n\}Figure 1:The generic marking prompt: a single multimodal message \(instruction, question, mark scheme, student answer, and any canvas image\) returns a structured integer mark\.We report the average of examiner\-model agreement against examiner\-examiner agreement\. For a model the mean agreement with the two examiners is

RA=12​\[QWK​\(AI,E1\)\+QWK​\(AI,E2\)\],R\_\{A\}=\\tfrac\{1\}\{2\}\\left\[\\mathrm\{QWK\}\(\\mathrm\{AI\},E\_\{1\}\)\+\\mathrm\{QWK\}\(\\mathrm\{AI\},E\_\{2\}\)\\right\],the mean QWK between the model and each examiner\. The reference is the agreement between the two examiners,RH=QWK​\(E1,E2\)R\_\{H\}=\\mathrm\{QWK\}\(E\_\{1\},E\_\{2\}\), whereE1,E2E\_\{1\},E\_\{2\}are the two examiners\. SoRAR\_\{A\}is the model’s mean agreement with the two examiners, andRHR\_\{H\}is how closely the two examiners agree with each other\. An automated marker should land on the mark the two examiners would settle on\. EachRAR\_\{A\}andRHR\_\{H\}carry a 95% confidence interval from a cluster\-over\-questions bootstrap \(2,000 resamples\)\. We report the differenceΔ=RA−RH\\Delta=R\_\{A\}\-R\_\{H\}with its interval \(Table[2](https://arxiv.org/html/2606.24973#S4.T2)\)\.

For essays we also reportsigned marking bias: mean\(predicted−consensus\)\(\\text\{predicted\}\-\\text\{consensus\}\)as a fraction of the maximum mark, where consensus is the two\-examiner mean \(positive = lenient, negative = harsh\)\.

## 4Agreement with the Examiners

We investigate whether a model agrees with two examiners as closely as the two examiners agree with each other\. Figure[2](https://arxiv.org/html/2606.24973#S4.F2)shows each best\-in\-subject model’s mean agreement,RAR\_\{A\}, against the agreement between the two examiners,RHR\_\{H\}\. Table[2](https://arxiv.org/html/2606.24973#S4.T2)gives the per\-subject values and the differenceΔ=RA−RH\\Delta=R\_\{A\}\-R\_\{H\}with its 95% confidence interval\.

Across all subjects, the leading models agree with an examiner more closely than the two examiners agree with each other\. English shows the best performance, with models of all sizes reaching examiner agreement and top performing models greatly exceeding it\. In Maths, although the examiner agreement is very high at around 0\.84, the top performing model achieves a delta of \+0\.02\. Science also shows examiner\-level marking ability, with a highest delta of \+0\.06\.

Differences between models are small and fall within the confidence intervals, and model size does not predict agreement: small models mark as consistently as large ones\. The leading models per subject are GPT\-5\.5 in English Language, Gemini 3\.5 Flash in Maths, and Gemini 3\.1 Pro across the pooled sciences, with more cost\-effective models scoring very similarly\. The sections that follow break these results down\.

![Refer to caption](https://arxiv.org/html/2606.24973v1/x1.png)Figure 2:Best model per subject \(RAR\_\{A\}, coloured\) against the examiner\-examiner agreement \(RHR\_\{H\}, grey\)\. A coloured bar reaching the grey marks as consistently as a second examiner\.Table 2:Numeric values for Figure[2](https://arxiv.org/html/2606.24973#S4.F2): the best model’s agreementRAR\_\{A\}, the examiner lineRHR\_\{H\}, andΔ=RA−RH\\Delta=R\_\{A\}\-R\_\{H\}with its 95% CI\.Δ\>0\\Delta\>0\(CI clear of zero\) means the model marks closer to the examiners than they do to each other\.
## 5Essay Marking

Essays are the most subjective marking task\. Examiner\-examiner agreement is much lower than the STEM subject agreement rates\. Additionally, the English assessment rests on only eight questions, so the per\-subject QWK in Figure[3](https://arxiv.org/html/2606.24973#S5.F3)is comparatively noisy\. Prior work applying GPT\-4 to second\-language essay assessment found significant correlations with human scores on a single annotated dataset\[[16](https://arxiv.org/html/2606.24973#bib.bib16)\]; here we mark real GCSE essays against a double\-marked examiner baseline, asking not whether the model correlates with humans but how it sits relative to the spread between two markers\.

The models still agree with an examiner more closely than the two examiners agree with each other \(RA\>RHR\_\{A\}\>R\_\{H\}\)\. The spread of errors is similar across models, and does not correlate with model size/cost\. One distinguishing factor, however, is marking bias, shown in Figure[4](https://arxiv.org/html/2606.24973#S5.F4)\. The difference is a harshness offset\. Claude Opus 4\.8, Claude Haiku 4\.5, and GPT\-5\.5 are the most neutral, while Claude Sonnet 4\.6 and Gemma 4 26B are the harshest\.

![Refer to caption](https://arxiv.org/html/2606.24973v1/x2.png)Figure 3:English: each model’s mean QWK with the two examiners \(higher = better\), best first\. Dashed line and band are the examiner\-examiner QWK and its 95% CI; a bar reaching the band marks as consistently as a second examiner\. Figures[5](https://arxiv.org/html/2606.24973#S6.F5)and[6](https://arxiv.org/html/2606.24973#S6.F6)share this layout\.![Refer to caption](https://arxiv.org/html/2606.24973v1/x3.png)Figure 4:English essays, signed error per model \(AI−\-consensus, % of max mark\), most accurate at the top\. Dashed line is neutral; left is harsh, right lenient\.
## 6Science and Maths Marking

The majority of tested models were found to mark the sciences and Maths as consistently as a second examiner, with Maths showing larger gaps in performance between models\. Top models in Science achieve positive deltas approaching \+0\.06, showing an improvement in consistency over examiners\.

In Maths the examiner line is high, anRHR\_\{H\}of 0\.84, because the two examiners agree very tightly\. Google’s Gemini models seem to be more aligned with the expectation of the GCSE marking criteria than OpenAI’s models, perhaps an artefact of their training data, appearing to be a stronger indicator than model size\. \(Figure[5](https://arxiv.org/html/2606.24973#S6.F5)\)\.

![Refer to caption](https://arxiv.org/html/2606.24973v1/x4.png)Figure 5:Maths: model mean QWK with the two examiners against the examiner\-examiner line \(cf\. Figure[3](https://arxiv.org/html/2606.24973#S5.F3)\)\.![Refer to caption](https://arxiv.org/html/2606.24973v1/x5.png)Figure 6:Science \(Biology, Chemistry, Physics combined\): model mean QWK with the two examiners against the examiner\-examiner line \(cf\. Figure[3](https://arxiv.org/html/2606.24973#S5.F3)\)\.
## 7Cost and Model Selection

Marking 1,000 student papers costs between $1 to $120 at list price, as shown in Figure[7](https://arxiv.org/html/2606.24973#S7.F7)\. Cost correlates only weakly with agreement; several cheaper models such as Claude Haiku 4\.5 for English, and Gemini 3\.1 Flash Lite for Science/Maths, achieve better performance than Claude Opus 4\.8 for a fraction of the cost\.

![Refer to caption](https://arxiv.org/html/2606.24973v1/x6.png)Figure 7:Mean QWK vs cost to mark 1,000 papers \(USD list price, log scale\), for English and Maths\+Science\. Each point is a model; top\-left is best \(high agreement, low cost\)\. Dashed line and band are the examiner line and its 95% CI\.
## 8Discussion and Conclusion

This benchmark shows that the LLMs available today, run under a single generic prompt at minimum reasoning effort, work as a reliable second marker\. Additionally, there is scope for improved performance under improved prompting or even fine\-tuning\. A double\-marked, multi\-subject dataset that includes handwriting tests this under realistic conditions, overcoming limitations of other benchmarks\.

We note several limitations\. The dataset comes from Medly’s own mock examinations, so the question types mimic official GCSE exam papers rather than using them directly\. We also evaluate models only at minimum reasoning effort, to explore a best\-case cost and latency scenario for each model\.

Finally, the two\-examiner consensus carries a variance caveat\. With only two raters, between\-examiner variance is one degree of freedom\. It does not mean models are compared to the broader examiner population\. Characterising the entire examiner population would require more than two raters\.

## Data Availability

## References

- \[1\]Cohen, J\. \(1968\)\. Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit\.Psychological Bulletin, 70\(4\), 213–220\.
- \[2\]Landis, J\. R\. & Koch, G\. G\. \(1977\)\. The Measurement of Observer Agreement for Categorical Data\.Biometrics, 33\(1\), 159–174\.
- \[3\]Shermis, M\. D\. & Burstein, J\. \(Eds\.\) \(2013\)\.Handbook of Automated Essay Evaluation: Current Applications and New Directions\. Routledge\.
- \[4\]Hewlett Foundation \(2012\)\. Automated Student Assessment Prize \(ASAP\)\. Kaggle Competition\.[https://www\.kaggle\.com/c/asap\-aes](https://www.kaggle.com/c/asap-aes)
- \[5\]Huang, Y\. & Wilson, J\. \(2025\)\. Evaluating LLM\-Based Automated Essay Scoring: Accuracy, Fairness, and Validity\.Proceedings of AIME\-CON \(Works in Progress\), 71–83\.
- \[6\]Taghipour, K\. & Ng, H\. T\. \(2016\)\. A Neural Approach to Automated Essay Scoring\.EMNLP 2016, 1882–1891\.
- \[7\]Dong, F\., Zhang, Y\. & Yang, J\. \(2017\)\. Attention\-based Recurrent Convolutional Neural Network for Automatic Essay Scoring\.CoNLL 2017, 153–162\.
- \[8\]Xiao, C\., Ma, W\., Song, Q\., Xu, S\. X\., Zhang, K\., Wang, Y\. & Fu, Q\. \(2024\)\. Human\-AI Collaborative Essay Scoring: A Dual\-Process Framework with LLMs\.arXiv:2401\.06431\.
- \[9\]Ofqual \(2018\)\. Marking Consistency Metrics: An Update\.[https://assets\.publishing\.service\.gov\.uk/media/5bfbfd70e5274a0fb775cca3/Marking\_consistency\_metrics\_\-\_an\_update\_\-\_FINAL64492\.pdf](https://assets.publishing.service.gov.uk/media/5bfbfd70e5274a0fb775cca3/Marking_consistency_metrics_-_an_update_-_FINAL64492.pdf)
- \[10\]Department for Education \(2019\)\. Teacher Workload Survey 2019: Research Report\.[https://www\.gov\.uk/government/publications/teacher\-workload\-survey\-2019](https://www.gov.uk/government/publications/teacher-workload-survey-2019)
- \[11\]Kortemeyer, G\., Nöhl, J\. & Onishchuk, D\. \(2024\)\. Grading Assistance for a Handwritten Thermodynamics Exam using AI: An Exploratory Study\.Phys\. Rev\. Phys\. Educ\. Res\., 20, 020144\.
- \[12\]Caraeni, A\., Scarlatos, A\. & Lan, A\. \(2024\)\. Evaluating GPT\-4 at Grading Handwritten Solutions in Math Exams\.arXiv:2411\.05231\.
- \[13\]Mansour, W\., Albatarni, S\., Eltanbouly, S\. & Elsayed, T\. \(2024\)\. Can Large Language Models Automatically Score Proficiency of Written Essays?arXiv:2403\.06149\.
- \[14\]Eltanbouly, S\., Albatarni, S\. & Elsayed, T\. \(2025\)\. TRATES: Trait\-Specific Rubric\-Assisted Cross\-Prompt Essay Scoring\.arXiv:2505\.14577\.
- \[15\]Latif, E\., Fang, L\., Ma, P\. & Zhai, X\. \(2024\)\. Knowledge Distillation of Large Language Models for Automatic Scoring of Science Assessments\.arXiv:2312\.15842\.
- \[16\]Bannò, S\., Vydana, H\. K\., Knill, K\. M\. & Gales, M\. J\. F\. \(2024\)\. Can GPT\-4 do L2 analytic assessment?arXiv:2404\.18557\.
- \[17\]Lu, P\., Bansal, H\., Xia, T\., Liu, J\., Li, C\., Hajishirzi, H\., Cheng, H\., Chang, K\.\-W\., Galley, M\. & Gao, J\. \(2024\)\. MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts\.ICLR 2024;arXiv:2310\.02255\.

Similar Articles

Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

arXiv cs.CL

This paper investigates whether standard benchmarks underestimate LLM performance by re-evaluating hallucination detection datasets using an LLM-first, human-adjudicated assessment method. The study finds that incorporating LLM reasoning into the adjudication process improves agreement and suggests that model-assisted re-evaluation yields more reliable benchmarks for ambiguity-prone tasks.

Can LLM Teams Play What? Where? When?

arXiv cs.CL

This paper investigates whether team-based interaction improves LLM performance in the quiz game 'What? Where? When?' (ChGK). Using six recent open LLMs on a 2025 dataset of 572 questions, they show that team strategies (voting, silent captain, talkative captain) outperform single models by up to 20 percentage points, with the best team achieving 44.23% accuracy, approaching human performance.

Evaluating LLMs as Human Surrogates in Controlled Experiments

arXiv cs.CL

This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.