Tag
Ai2 introduces TutorMoments, a replay-based evaluation framework and dataset for measuring whether LLMs can balance when to help and when to hold back in one-on-one math tutoring. Preliminary results show models tend to over-help, and prompt engineering only partially closes the gap to human tutors.
This paper introduces the Pedagogical Suitability Index (PSI), a composite metric for evaluating and improving how well LLM-based AI tutors align responses with learner readiness and curricular progression, showing that PSI-guided feedback improves weak tutoring cases.
Wealthy families in the US are turning to AI-powered private schools like Forge Prep and Alpha School, paying tens of thousands of dollars for AI tutors and project-based learning, despite broader public distrust of AI and concerns about educational outcomes.
A Stanford Law School study found that law professors rated LLM-generated answers higher than peer answers in a blinded evaluation of short-answer tutoring in contracts courses, with LLMs winning 75.33% of comparisons and being flagged as harmful less often.
Praktika is an AI-powered language learning app that uses a multi-agent system with GPT-5.2 and GPT-5 Pro to provide personalized conversational lessons that mimic real human tutors through adaptive lesson delivery, progress tracking, and dynamic learning planning.