Tag
CSTutorBench is a benchmark for evaluating small language models as tutors for block-based programming, focusing on pedagogical behaviors. Preliminary results show models struggle with deeper tutoring skills like avoiding answer leakage, and prompt revision improves scores.
Researchers from Arizona State University present a framework for evaluating adaptive personalization of educational reading materials using theory-grounded simulated learners, incorporating memory models, misconception revision, and Bayesian Knowledge Tracing. Experiments across three subjects show adaptive reading significantly improved outcomes in computer science but had mixed results in chemistry and biology.