The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty
Summary
This paper investigates how LLM-generated difficulty ratings for math items align with actual student performance, finding that LLMs systematically underestimate difficulty for items driven by learner misconceptions, a phenomenon termed the 'Easy Trap'.
View Cached Full Text
Cached at: 07/31/26, 04:01 AM
# The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty Source: [https://arxiv.org/abs/2607.26067](https://arxiv.org/abs/2607.26067) [View PDF](https://arxiv.org/pdf/2607.26067) > Abstract:Large language models \(LLMs\) are increasingly used for estimating item difficulty in educational assessment\. However, it remains unclear whether such estimates reflect how learners actually experience difficulty\. This study investigates the alignment between LLM\-generated difficulty ratings and empirical student performance on basic mathematics tasks\. Four widely used LLM\-based systems generated difficulty ratings on a 1\-100 scale for 32 arithmetic items across multiple runs \(N = 640 ratings\)\. These were compared with empirical difficulty derived from responses of 770 Indonesian undergraduates using Classical Test Theory \(CTT\) and Item Response Theory \(2PL\)\. Results show moderate rank correlations \(Spearman's rho = 0\.52\-0\.70\), indicating that LLMs capture coarse ordering of item difficulty\. However, substantial and systematic misalignment emerges in fraction items\. Several items consistently rated as easy by LLMs were among the most difficult for students, such as an item with only 34\.16% correct for 100 : 1/2\. We argue that LLMs approximate curricular difficulty, or what should be easy based on instructional sequencing, rather than cognitive difficulty driven by learner misconceptions\. This leads to systematic underestimation of misconception\-driven items, a phenomenon we term the Easy Trap\. These findings highlight a critical limitation of LLM\-based difficulty estimation and suggest that relying on such estimates without empirical grounding may introduce bias in assessment design and adaptive systems\. ## Submission history From: Amanda La Hadi \[[view email](https://arxiv.org/show-email/bfe706fe/2607.26067)\] **\[v1\]**Mon, 22 Jun 2026 09:53:25 UTC \(769 KB\)
Similar Articles
Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs
This paper investigates whether LLMs can accurately predict item difficulty levels in large-scale reading and writing tests, finding that GPT-4.1 achieves moderate accuracy but is outperformed by ConvBERT, and that LLMs tend to underestimate difficulty for hard items.
The Problem with “Mathematically Proven” Claims About LLMs (15 minute read)
This article critiques the sensationalized media coverage of mathematical proofs regarding LLM limitations, specifically highlighting how conditional results about self-improvement are often misrepresented as universal impossibilities.
Error as a Lens: Probing LLM Reasoning through Synthetic Misconception Generation
This paper presents a framework using LLMs to generate targeted synthetic misconceptions aligned to a five-class taxonomy adapted from Bloom's taxonomy, addressing the scarcity of labeled student error data in education research.
The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
This paper introduces the 'Agentic Formalism Trap' and an Evaluative Dissonance Index, showing how LLM-as-a-Judge systems can be misled by structural formalism and consensus mimicry rather than semantic truth, based on 22,500 trajectories across multiple domains.
Investigating LLM's Problem Solving Capability -- a Study on Statics Questions
This paper evaluates LLM performance on statics problems, finding that while text-only questions are handled well, accuracy drops with diagrams and multi-step reasoning, suggesting difficulties in applying visual information consistently.