Tag
StudentSim trains personalized LLM-based student simulators from sparse data to mirror learner responses and adapt to tutor guidance, outperforming GPT-5.4 across chess, writing, and math domains.
The paper introduces a novel approach using Poly-Encoders for computationally efficient automated creativity assessment, achieving performance comparable to Large Language Models with significantly reduced computational demands.
OmniPhys is a large-scale multimodal benchmark for physics understanding and generation, covering middle school to university-level problems from Chinese educational corpora, aimed at evaluating and advancing multimodal large language models in scientific domains.
The study examines the semantic consistency of LLM-generated replies across different models and conversational contexts, highlighting the need for infrastructure and design strategies to maintain stable responses for conversation-based assessments.
The paper presents a dual gatekeeping system for AI-generated educational videos that combines educator input and automated metrics to enhance pedagogical quality, demonstrating that principled resistance to AI outputs improves instructional design.
TRACE-BN introduces a curriculum-guided dataset for structured Bangla-English tutoring and demonstrates transferring this behavior to a sub-1B language model using LoRA, achieving significant improvements in tutoring quality for resource-constrained offline environments.
TeachMateGPT presents a multi-agent framework that improves retrieval-augmented generation for creating pedagogical assessments from science textbooks, achieving higher faithfulness and answer relevancy compared to baseline systems.
This article presents LittleLearner, a controlled sandbox for studying LLM knowledge acquisition using a K-5 curriculum-filtered dataset, finding that interventions like scaling and post-training enhance in-scope performance but do not improve out-of-scope capabilities.
ProPRL introduces a property-aware framework for prerequisite relation learning in educational knowledge graphs, combining concept-resource hypergraph and directed behavior graph with adaptive pair-conditioned fusion and an irreversibility constraint to achieve state-of-the-art performance.
This paper proposes a conditional generalizability framework to evaluate nonuniform dependability across response conditions in automated essay scoring.
This paper presents VectorizationLLM, a specialized LLM built on Google open-weight models and a RAG knowledge base, designed to help students learn smart vectorization, Fourier analysis, and differential equations in MATLAB without providing direct answers.
This paper evaluates cross-dataset generalization of supervised ML/DL models and prompted LLMs for automatic Bloom's taxonomy classification of assessment questions, finding that LLMs are more robust across diverse educational contexts.
IntElicit is a framework that uses dialogue policy optimization with a decomposed process reward mechanism to elicit and assess contextualized creativity through adaptive AI interviewing, reducing confounders like domain knowledge and engagement. Experiments show it improves creative outcomes over static assessment methods.
A pre-registered trial in Sierra Leone found that AI-powered Guided Learning significantly improved math scores, achieving 1.2 to 1.7 years of progress in eight weeks, while teachers reported enhanced professional growth and a shift toward facilitation roles.
This paper introduces Elmes+, an automated framework for constructing fine-grained evaluation rubrics for LLMs in long-tail educational scenarios, and presents the Edu-330 benchmark covering 330 scenarios across 11 subjects. The framework uses a multi-agent engine and self-evolving module to co-optimize evaluation criteria and test data, revealing multidimensional educational capability differences among top LLMs.
TeachObs introduces a human-validated benchmark for multimodal teaching observation, consisting of 30 classroom videos annotated with segment-level binary codes and lesson-level expert ratings, and evaluates five frontier LLMs across three tracks, finding no single model consistently outperforms and that model evaluations overrate procedurally clear lessons.
This paper presents a framework that uses domain-specific expert knowledge to ground large language models for providing Just-in-Time adaptive feedback to students based on their written reasoning, achieving over 80% improvement in student performance in a large university course.
This paper presents a modular pipeline for educational analogy generation, decomposing the task into four stages and evaluating 12 LLMs and 7 embedding models. Results show that sub-concept grounding improves explanation quality and retrieval precision, with a novel LLM-as-a-judge evaluation validated against human annotations.
This paper presents a forward-looking perspective on agentic multi-agent AI platforms in higher education, addressing the need for integrated, inclusive systems that support learning, teaching, and institutional operations. It identifies gaps in current fragmented AI tools and proposes directions for scalable, human-aligned multi-agent ecosystems.
This paper details the RETUYT-INCO team's participation in the BEA 2026 Shared Task 2, introducing a meta-prompting approach for rubric-based scoring of German short answers.