Tag
The study uses exploratory factor analysis to compare latent structures in human and LLM responses on assessments, revealing that LLMs rely on statistically opaque mechanisms unlike human reasoning.
This paper proposes a multidimensional evaluation framework for assessing statistical reasoning in large language models, combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis across 15 LLMs and 90 exam questions. It finds that accuracy alone is insufficient to characterize LLM statistical reasoning and that vendor-specific stylistic differences exist.
This paper proposes an evaluation framework for predicting item parameters from text embeddings using regularized regression and reliability/design ceilings. Results show that difficulty is substantially predictable from text, while discrimination and pseudo-guessing are limited by reliability, not text signal.