Examining Variation in How Guided AI Tutors Resolve Student Impasses

arXiv cs.AI Papers

Summary

A study analyzing 20,462 student turns in a guided LLM chemistry tutor finds that pedagogical guardrails change what tutors withhold but not how they adapt, showing questioning helps early in impasses while direct error feedback becomes more effective as impasses persist.

arXiv:2609.38346v1 Announce Type: new Abstract: When a student is stuck, a tutor faces the assistance dilemma: help given too early can hinder productive struggle, while help withheld too long leaves the student in a frustrating, persistent impasse (i.e., wheel spinning). Generative AI tutors increasingly use guardrails restricting answer-giving, yet little is known about how such tutors behave once an impasse persists. We analyze 20,462 student turns from 1,260 authentic sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns of three major types: conceptual errors, expressed uncertainty, or help-seeking. We then used these impasses to simulate three tutoring conditions to study variation in AI tutor guidance through impasses: baseline, no-direct-answer, and guided tutor. For a sample of 150 impasses, prompt specificity changed pedagogy: a baseline tutor provided the answer directly in 50.7% of responses, a no-direct-answer tutor asked a follow-up question every time, and the guided tutor responded in a wide variety of ways depending on the context. We then analyzed impasse trajectories in authentic interactions, finding that each additional impasse turn lowered the odds of next-turn recovery by 12.7% (AOR = 0.873, p < .001), and early dropouts were caught in recursive concept elicitation before reaching execution. The benefit of questioning decayed as impasses persisted (scripted question x depth AOR = 0.78; follow-up x depth AOR = 0.83), whereas addressing the student's error grew more beneficial (AOR = 1.14); after a failed scripted question, repeating it was followed by recovery in 28.1% of cases, compared with 39.8% when the tutor addressed the error instead. For learning analytics, these findings identify impasse depth and type as observable, turn-level dialogue signals that analytics can use to trigger graduated, state-sensitive assistance in real time.
Original Article
View Cached Full Text

Cached at: 10/01/26, 09:40 AM

# 1Introduction
Source: [https://arxiv.org/html/2609.38346](https://arxiv.org/html/2609.38346)
\\PaperTitle

Examining Variation in How Guided AI Tutors Resolve Student Impasses\\AuthorsBakhtawar Ahtisham1, Kirk Vanacore1, Alessandra Napoli2, Josh Arens2, Ksenia Ionova1, Clayton Cohn1, Shima Salehi2, Rene Kizilcec1\\KeywordsGenerative AI tutoring, Student impasses, Assistance dilemma, Pedagogical guardrails, Large language models, Dialogue analytics, Learning Analytics\\Submitted09/28/2026\\Accepted\-\\Published\-\\Volume1\\Number1\\Pages1—10\\Doixxx\-xxx\-xxx\\NotesnameNotes for Practice\\notePedagogical guardrails in LLM tutors \(e\.g\., “do not give the answer”\) change*what*the tutor withholds, but do not by themselves specify*how*it should adapt when a student stays stuck\.\\noteIn authentic dialogues with a guided tutor, questioning helped most at the first sign of an impasse; its benefit shrank with every additional turn the student remained stuck, whereas directly addressing the student’s error became relatively more helpful\.\\noteStudents who abandoned sessions early were caught in repeated concept\-elicitation loops before ever reaching problem execution\.\\noteLearning analytics systems can track impasse depth and type in tutoring dialogues in real time to flag students who remain stuck, and support graduated escalation, moving from questioning to targeted error feedback and, when needed, worked sub\-steps\.\\AbstractWhen a student is stuck, a tutor faces the assistance dilemma: help given too early can hinder productive struggle, while help withheld too long leaves the student in a frustrating, persistent impasse \(i\.e\., wheel spinning\)\. Generative AI tutors increasingly use guardrails restricting answer\-giving, yet little is known about how such tutors behave once an impasse persists\. We analyze 20,462 student turns from 1,260 authentic sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns of three major types: conceptual errors, expressed uncertainty, or help\-seeking\. We then used these impasses to simulate three tutoring conditions to study variation in AI tutor guidance through impasses: baseline, no\-direct\-answer, and guided tutor\. For a sample of 150 impasses, prompt specificity changed pedagogy: a baseline tutor provided the answer directly in 50\.7% of responses, a no\-direct\-answer tutor asked a follow\-up question every time, and the guided tutor responded in a wide variety of ways depending on the context\. We then analyzed impasse trajectories in authentic interactions, finding that each additional impasse turn lowered the odds of next\-turn recovery by 12\.7% \(AOR=0\.873\\mathrm\{AOR\}=0\.873,p<\.001p<\.001\), and early dropouts were caught in recursive concept elicitation before reaching execution\. The benefit of questioning decayed as impasses persisted \(scripted question×\\timesdepthAOR=0\.78\\mathrm\{AOR\}=0\.78; follow\-up×\\timesdepthAOR=0\.83\\mathrm\{AOR\}=0\.83\), whereas addressing the student’s error grew more beneficial \(AOR=1\.14\\mathrm\{AOR\}=1\.14\); after a failed scripted question, repeating it was followed by recovery in 28\.1% of cases, compared with 39\.8% when the tutor addressed the error instead\. For learning analytics, these findings identify impasse depth and type as observable, turn\-level dialogue signals that analytics can use to trigger graduated, state\-sensitive assistance in real time\.

## 1Introduction

Being stuck is a normal, and often productive, part of learning to solve problems\. Theories of impasse\-driven learning describe these moments asimpasses: states in which a learner’s current knowledge is insufficient to produce a correct next step and, precisely because of this failure, new knowledge may be constructed[VanLehn \(\(1988\)\)](https://arxiv.org/html/2609.38346#bib.bib44);[VanLehn et al\. \(\(2003\)\)](https://arxiv.org/html/2609.38346#bib.bib47)\. However, impasses are productive only when they are resolved\. When students cannot repair them, productive struggle becomes unproductive failure[Kapur \(\(2016\)\)](https://arxiv.org/html/2609.38346#bib.bib21)or*wheel\-spinning*—in which students repeatedly fail to master a skill despite extended practice[Beck & Gong \(\(2013\)\)](https://arxiv.org/html/2609.38346#bib.bib7);[Gong & Beck \(\(2015\)\)](https://arxiv.org/html/2609.38346#bib.bib16);[Kai et al\. \(\(2018\)\)](https://arxiv.org/html/2609.38346#bib.bib19)—and may lead to disengagement\. Deciding when and how much to help is therefore a central problem for any tutor, a problem[Koedinger & Aleven \(\(2007\)\)](https://arxiv.org/html/2609.38346#bib.bib24)named the*assistance dilemma*: both too much and too little assistance can impair learning, and the optimal level depends on the learner’s state\.

Large language models \(LLMs\) have made it cheap to deploy conversational tutors at scale, and early evidence is mixed\. Well\-designed AI tutors can produce substantial learning gains[Kestin et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib22);[Pardos & Bhandari \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib31), but unrestricted access to a general\-purpose model can improve practice performance while harming later unassisted learning[Bastani et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib6)\. The prevailing response is to add*pedagogical guardrails*: system prompts that forbid giving away answers and ask the model to guide students with questions and hints[Liffiton et al\. \(\(2023\)\)](https://arxiv.org/html/2609.38346#bib.bib28);[Bastani et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib6);[Jurenka et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib18)\. Such guardrails address the risks of excessive assistance but provide less guidance about what a tutor should do when its initial support is insufficient, and the student remains stuck\.

This distinction matters because LLM tutors are typically evaluated one response at a time, for example by rating whether a single turn avoids revealing the answer or identifies the student’s mistake[Tack & Piech \(\(2022\)\)](https://arxiv.org/html/2609.38346#bib.bib42);[Maurya et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib30)\. Such evaluations cannot reveal whether a tutor*adapts*across turns as evidence accumulates that its assistance is failing\. Human tutoring research, by contrast, has long emphasized contingency: effective tutors increase the specificity of their help after a failure and reduce it after success[D\. Wood et al\. \(\(1976\)\)](https://arxiv.org/html/2609.38346#bib.bib50);[H\. Wood & Wood \(\(1999\)\)](https://arxiv.org/html/2609.38346#bib.bib51)\. Whether guardrailed LLM tutors behave contingently, and whether contingency matters for students in authentic use, is an open empirical question for learning analytics\.

We address this question with log data from a constrained LLM tutor that supports chemistry students in planning solutions to multi\-step problems, organized around expert problem\-solving practices[A\.M\. Price et al\. \(\(2021\)\)](https://arxiv.org/html/2609.38346#bib.bib34);[Burkholder et al\. \(\(2020\)\)](https://arxiv.org/html/2609.38346#bib.bib9);[Schwartz Poehlmann et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib39)\. We combine a controlled replay of student impasses through tutors with different levels of constraint with sequential and multivariable analyses of the tutor’s authentic dialogues\. We ask:

- RQ1\.How do student impasses manifest across problem\-solving practices in tutoring dialogues?
- RQ2\.How does the degree of prompt\-based guidance change the way an AI tutor responds to student impasses?
- RQ3\.How do students who complete the session differ from those who drop out in how they progress through problem\-solving practices?
- RQ4\.Which tutor moves are associated with recovery from student impasses?

This paper makes three contributions to learning analytics\.*Conceptually*, it extends the evaluation of guardrailed LLM tutors beyond single\-turn compliance by examining how tutor support adapts across persistent student impasses\.*Empirically*, it shows, in 1,260 authentic sessions, that the benefit of the tutor’s dominant move of questioning decays as an impasse persists, while error\-directed feedback grows relatively more useful, and that early attrition is concentrated in concept\-elicitation loops\.*Methodologically*, it offers a reusable operationalization of impasse episodes, depth, and recovery from codebook\-annotated dialogue logs\. Together, these findings motivate*graduated assistance*: tutor policies that condition the type and degree of support on how long a student has been stuck\.

## 2Literature Review

### 2\.1Impasses, Productive Struggle, and the Assistance Dilemma

[VanLehn \(\(1988\)\)](https://arxiv.org/html/2609.38346#bib.bib44)proposed that learning is driven by impasses: when a learner’s procedures cannot produce the next step, they must repair the gap, and the repair may be generalized into new knowledge\. Studies of problem solving[VanLehn \(\(1999\)\)](https://arxiv.org/html/2609.38346#bib.bib45)and human tutoring[VanLehn et al\. \(\(2003\)\)](https://arxiv.org/html/2609.38346#bib.bib47)found that learning events concentrate at impasses, and that tutor explanations are most effective when they follow a student’s impasse rather than precede it\. Related work on affect shows that confusion, the affective signature of an impasse, can benefit learning when it is resolved, but becomes harmful when it persists into frustration or boredom[D’Mello et al\. \(\(2014\)\)](https://arxiv.org/html/2609.38346#bib.bib14);[Baker et al\. \(\(2010\)\)](https://arxiv.org/html/2609.38346#bib.bib5)\.

Research on productive failure extends this idea to instructional design: allowing students to struggle before receiving instruction can deepen conceptual understanding[Kapur \(\(2008\)\)](https://arxiv.org/html/2609.38346#bib.bib20), but the benefits depend on how that struggle is supported[Kapur \(\(2016\)\)](https://arxiv.org/html/2609.38346#bib.bib21)\. This creates a practical version of the assistance dilemma[Koedinger & Aleven \(\(2007\)\)](https://arxiv.org/html/2609.38346#bib.bib24): support must be sufficient to help learners progress without eliminating opportunities for productive problem solving\. Our work examines this tradeoff turn by turn, asking how the value of different forms of assistance changes as an impasse persists\.

### 2\.2Help\-Seeking, Contingent Tutoring, and Graduated Assistance

Classic accounts of scaffolding describe tutors who adjust their support to the learner’s success or failure[D\. Wood et al\. \(\(1976\)\)](https://arxiv.org/html/2609.38346#bib.bib50)\.[H\. Wood & Wood \(\(1999\)\)](https://arxiv.org/html/2609.38346#bib.bib51)formalized this as*contingent tutoring*: increase the specificity of help after a failure and reduce it after success\. Studies of naturalistic human tutoring show that tutors rely heavily on questioning, feedback, and collaborative repair of errors[Graesser et al\. \(\(1995\)\)](https://arxiv.org/html/2609.38346#bib.bib17), and that eliciting student constructions can be as effective as tutor explanations[Chi et al\. \(\(2001\)\)](https://arxiv.org/html/2609.38346#bib.bib11)\. Intelligent tutoring systems implement a version of contingency through hint sequences that move from general prompts to bottom\-out hints[VanLehn \(\(2011\)\)](https://arxiv.org/html/2609.38346#bib.bib46), and a substantial literature models when and how students seek such help[Aleven et al\. \(\(2003\)\)](https://arxiv.org/html/2609.38346#bib.bib2);[Aleven et al\. \(\(2006\)\)](https://arxiv.org/html/2609.38346#bib.bib1);[T\.W\. Price et al\. \(\(2017\)\)](https://arxiv.org/html/2609.38346#bib.bib35)\. Help\-seeking is a meaningful signal in its own right: explicit requests indicate metacognitive awareness of an impasse, whereas hedged answers signal partial knowledge\. We therefore treat conceptual errors, expressed uncertainty, and help requests as distinct manifestations of an impasse rather than a single failure state\.

### 2\.3Generative AI Tutors and Pedagogical Guardrails

LLM tutors differ from rule\-based tutoring systems in that their pedagogical behavior is specified largely through natural\-language instructions rather than an explicit tutoring model\. Field and laboratory studies report both promise and risk[Yan et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib52)\. ChatGPT\-generated hints produced learning gains comparable to human\-authored hints[Pardos & Bhandari \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib31), and a research\-designed AI tutor outperformed in\-class active learning in a university physics course[Kestin et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib22)\. However, unrestricted access to GPT\-4 harmed high\-school students’ subsequent unassisted performance, while a version prompted to give hints rather than answers largely removed this harm[Bastani et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib6)\. Similarly, how students use LLMs, as a substitute for or complement to their own work, shapes whether they learn[Lehmann et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib26)\. Indeed, surveys of STEM college students show a near even split between these two practices, with 54% of students using AI as a scaffold for their own reasoning and 46% as a shortcut to answers[K\.D\. Wang et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib48)\.

In response, researchers and developers have built guardrailed tutors that withhold solutions[Liffiton et al\. \(\(2023\)\)](https://arxiv.org/html/2609.38346#bib.bib28);[Sheese et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib40), fine\-tuned models for pedagogy[Jurenka et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib18);[Macina et al\. \(\(2023\)\)](https://arxiv.org/html/2609.38346#bib.bib29);[Scarlatos, Liu et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib38), and developed methods for diagnosing and remediating specific student errors[Daheim et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib13)\. Evaluation frameworks rate tutor responses along pedagogical dimensions such as mistake identification, answer revelation, and actionability[Tack & Piech \(\(2022\)\)](https://arxiv.org/html/2609.38346#bib.bib42);[Maurya et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib30)\. Human–AI approaches such as Tutor CoPilot show that LLM suggestions can shift human tutors toward more guiding questions and fewer revealed answers[R\.E\. Wang et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib49);[Thomas et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib43)\. Most of this work evaluates responses in isolation or uses simulated students\. Fewer studies analyze authentic student–LLM dialogues over time, and, to our knowledge, none examines whether a guardrailed tutor’s assistance adapts as a student’s impasse persists\.

### 2\.4Temporal Analytics of Tutoring Dialogue

Learning analytics has increasingly argued that the order and timing of events, not only their frequency, carry information about learning processes[Reimann \(\(2009\)\)](https://arxiv.org/html/2609.38346#bib.bib36);[Chen et al\. \(\(2018\)\)](https://arxiv.org/html/2609.38346#bib.bib10);[Knight et al\. \(\(2017\)\)](https://arxiv.org/html/2609.38346#bib.bib23)\. Sequential methods such as transition matrices and lag analysis have been used to characterize tutoring and collaborative dialogue[Bakeman & Quera \(\(1995\)\)](https://arxiv.org/html/2609.38346#bib.bib4);[Borchers et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib8), and recent work applies LLMs to trace knowledge directly from tutor–student dialogue[Scarlatos, Baker & Lan \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib37)\. Discrete\-time event models provide a natural framework for asking how the probability of an event, here, resolving an impasse, changes with elapsed time and with time\-varying covariates such as the tutor’s most recent move[Singer & Willett \(\(1993\)\)](https://arxiv.org/html/2609.38346#bib.bib41)\. We draw on these traditions to model impasse episodes as turn\-level processes nested within sessions\.

Taken together, prior work points to a common problem: assistance is often evaluated as a property of a single tutor response or as a fixed instructional policy, rather than as something that should adapt as evidence accumulates that a student remains stuck\. Research on impasses and contingent tutoring suggests that support should become more specific after unsuccessful attempts, while work on LLM guardrails shows how prompts can constrain tutor behavior but says much less about whether that behavior changes appropriately across turns\. Learning analytics provides methods for modeling sequential interaction processes, yet these methods have rarely been used to examine how the value of specific LLM tutor moves changes within persistent impasses\. We address this gap by combining a controlled comparison of tutors with different levels of prompt guidance and turn\-level analysis of authentic student–LLM dialogues, using impasse type, depth, tutor moves, and next\-turn recovery to examine how assistance changes as students remain stuck\.

## 3Methods

### 3\.1Guided AI Tutor and Context

The dialogue data used in this study were collected using the Guided AI Tutor \(name withheld for review\), an LLM\-based tutor \(GPT\-4o via the OpenAI API\) that guides students to plan a solution to a multistep problem without giving them the solution[Anonymous \(\(2026\)\)](https://arxiv.org/html/2609.38346#bib.bib3)\. The tutor’s behavior is controlled entirely by its prompt, which is based on validated problem\-solving templates for chemistry[Schwartz Poehlmann et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib39)and physics[Burkholder et al\. \(\(2020\)\)](https://arxiv.org/html/2609.38346#bib.bib9)\. The prompt covers five parts: \(1\) a*role*as a problem\-solving coach; \(2\)*objectives*to help students build their own plan, reflect metacognitively, and stay engaged; \(3\)*behavioral guidelines*: never solve the problem, ask one question at a time, keep responses to three sentences or fewer, do not praise incorrect answers, and after an incorrect or incomplete answer give “increasingly specific, targeted hints”; \(4\) the*problem and four scripted questions*, each paired with its correct answer so the tutor can check responses against a reference; and \(5\)*response\-handling rules*for answer requests and off\-topic turns\. Two features matter for our study\. First, the scripted sequence determines which practice each exchange targets, so the tutor, not the student, decides when the dialogue moves on\. Second, escalation after errors is requested only as a natural\-language instruction\. This deployed configuration is called “Guided AI Tutor” in the analyses addressing RQ2\. The tutor’s planning sequence is organized around*problem\-solving practices*, the decisions experts make when solving authentic problems[A\.M\. Price et al\. \(\(2021\)\)](https://arxiv.org/html/2609.38346#bib.bib34), adapted into a planning template for chemistry[Schwartz Poehlmann et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib39)\. We code eight practices: problem identification, relevant concepts, similar problems, information needed, assumptions, strategy, execution, and answer checking \(defined in Table[1](https://arxiv.org/html/2609.38346#S3.T1)\)\.

##### Data and ethics\.

Data were collected in an introductory general chemistry course at a large research university in the United States, where students used the tutor through a stand\-alone web application to plan solutions to 8 assigned chemistry problems\. The model version and configuration were held fixed throughout data collection\. The corpus comprises 1,260 sessions \(20,462 student turns; mean 16\.2 per session\)\. This study is a secondary analysis of existing tutoring logs; all transcripts were de\-identified before analysis and contained no personally identifiable information\. The study was determined exempt by the authors’ Institutional Review Board \[protocol number withheld for review\]\.

### 3\.2Codebook Development

The codebook was first developed deductively for a physics implementation of the tutor, building on existing frameworks for evaluating AI tutors[Tack & Piech \(\(2022\)\)](https://arxiv.org/html/2609.38346#bib.bib42);[Petukhova & Kochmar \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib33);[Maurya et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib30);[Pauzi et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib32)\. It initially described tutor moves \(e\.g\., asking questions, giving feedback\); codes for student behavior \(e\.g\., answering, help\-seeking\) and for the correctness of student and tutor statements were added to follow the conversational flow\. For the chemistry implementation, the problem\-solving practices of the planning template[A\.M\. Price et al\. \(\(2021\)\)](https://arxiv.org/html/2609.38346#bib.bib34);[Schwartz Poehlmann et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib39)were added as deductive codes, alongside inductive codes that emerged from the data\. Two researchers refined the codebook iteratively, coding the same two to four conversations, discussing disagreements, and splitting or merging codes until definitions stabilized\. The full codebook also includes categories not analyzed here \(e\.g\., tutor correctness identification, answer leakage\) and is available as supplementary material; the codes used in this study are listed in Table[1](https://arxiv.org/html/2609.38346#S3.T1)\.

Table 1:Codebook categories and codes used in the analyses\. Student\-input and correctness codes define impasse types and recovery \(Section[3\.4](https://arxiv.org/html/2609.38346#S3.SS4)\); tutor\-output codes are the moves analyzed in RQ2 and RQ4\.CategoryCodeDefinitionStudent inputDirect answer or explanationResponse to a question and/or explanation of thinking\.Answer with uncertaintyResponse to a question expressed with doubt or uncertainty\.Specific questionQuestion seeking specific information or asking about a specific step\.General help\-seekingExpresses confusion or general uncertainty without referencing specific information or a step\.CorrectnessCompletely correctFully accurate and complete, with no errors or missing key information \(defines recovery\)\.Partially incorrectSome components are correct, but at least one key idea, step, or piece of information is incorrect or missing\.Completely incorrectContains fundamental factual errors that make the reasoning or result invalid\.Not specific enoughBroadly correct but too vague to determine whether the student can apply it in context\.Tutor outputScripted questionAsks a question explicitly included in the prompt’s scripted sequence\.Follow\-up questionAsks additional questions to clarify, extend, or probe the student’s thinking\.Check for understandingExplicitly asks whether the student understands a concept, explanation, or step\.Direct answerDirectly responds to the student’s explicit question or request for information\.Confirm / elaborateAcknowledges a correct response and may provide additional explanation or context\.Address incorrect answerResponds to incorrect or partially incorrect work by correcting, guiding, or explaining\.Partial answer / hintGives partial solution information or scaffolding without fully solving the problem\.Praise / encouragementGives supportive, motivating, or affirming feedback\.Remind exercise purposeRestates the goals or intended learning purpose of the activity\.Problem\-solving practiceProblem identificationIdentifies or restates the type, goal, or key components of the problem\.Relevant conceptsIdentifies concepts, principles, or skills needed to solve the problem, separate from applying them\.Similar problemsIdentifies and/or compares a similar problem\.Information neededIdentifies or gathers specific information, values, equations, or reference data, separate from using them\.AssumptionsMakes assumptions or simplifications that help define or solve the problem\.StrategyPlans the steps or sequence of actions needed to solve the problem\.ExecutionCarries out solution steps or calculations to obtain the result\.Answer checkingChecks or evaluates whether a solution, step, or answer is correct\.
### 3\.3Ground Truth Construction and Annotation

We annotated the corpus using a human\-validated LLM annotation pipeline built around the Guided AI Tutor dialogue codebook \(Table[1](https://arxiv.org/html/2609.38346#S3.T1)\)\. First, two annotators with chemistry domain expertise independently coded 166 utterances from six randomly selected sessions on all codebook dimensions with an initial inter\-rater reliability of 83%; remaining disagreements were resolved through discussion, and code definitions were refined where they proved ambiguous to provide a single agreed\-upon ground truth for an LLM annotator\. We then applied an LLM annotator \(Gemini 3\.1 Pro\), prompted with the refined codebook definitions and examples, to 230 utterances from 10 further randomly sampled sessions\. We refined the prompt until reaching high agreement between the LLM and human annotations \(meanκ=0\.86\\kappa=0\.86\)\. Because prior work has found that LLMs can match or exceed human annotators on some text\-annotation tasks[Gilardi et al\. \(\(2023\)\)](https://arxiv.org/html/2609.38346#bib.bib15), we did not assume that disagreements necessarily reflected errors by the LLM\. Instead, two human annotators independently reviewed every AI\-assigned label, either confirming it or replacing it with the appropriate code\. These adjudicated labels served as the gold standard\.

To scale the annotation across the entire corpus, we used the final prompt on this gold\-standard set, which yielded substantial agreement[Cohen \(\(1960\)\)](https://arxiv.org/html/2609.38346#bib.bib12);[Landis & Koch \(\(1977\)\)](https://arxiv.org/html/2609.38346#bib.bib25)across the codebook dimensions used in our analyses: student input type \(meanκ=\.84\\kappa=\.84\), tutor moves \(meanκ=\.79\\kappa=\.79\), and problem\-solving practice \(meanκ=\.80\\kappa=\.80\)\. For correctness, the code that defines recovery,*completely correct*, was identified reliably \(κ=\.81\\kappa=\.81\)\. Rare or unreliable codes \(e\.g\.,*ignore student question*,*unintelligible*\) were excluded, and the validated prompt was then applied to the full corpus of 1,260 sessions\. The same validated prompt was used to code the tutor responses in the controlled comparison \(Section[3\.5](https://arxiv.org/html/2609.38346#S3.SS5.SSS0.Px2)\)\.

### 3\.4Operationalizing Student Impasses

Following impasse\-driven learning theory[VanLehn \(\(1988\)\)](https://arxiv.org/html/2609.38346#bib.bib44);[VanLehn et al\. \(\(2003\)\)](https://arxiv.org/html/2609.38346#bib.bib47), we treat an impasse as a state in which the student’s current knowledge is insufficient to produce a correct next step, visible in dialogue as an error, hesitation, or a request for help\. We define four constructs over the annotated logs:

1. 1\.Impasse turn\.A student turn coded as \(a\) a*conceptual error*: a direct answer or explanation coded as completely incorrect, partially incorrect, or not specific enough; \(b\)*expressed uncertainty*: an answer with explicit hedging \(e\.g\., “I think”, “maybe”, “is it…?”\); or \(c\)*help\-seeking*: a general help request \(“I don’t understand”, “idk”\) or a specific question that proposes no answer\. The three types are mutually exclusive by construction\.
2. 2\.Impasse episode\.A maximal run of consecutive impasse turns within a session\.
3. 3\.Impasse depth\.The position of a turn within its episode \(1,2,…,k1,2,\\dots,k\); depth 1 is the onset\.
4. 4\.Recovery\.The student’s next turn \(t\+1t\+1\) is coded as a completely correct answer, which ends the episode\. Episodes that end without a correct response \(e\.g\., the session ends\) are unresolved\.

Across the corpus we identified 6,630 impasse turns in 4,170 episodes within 1,204 sessions: 2,311 conceptual errors \(34\.9%\), 1,970 expressions of uncertainty \(29\.7%\), and 2,349 help requests \(35\.4%\)\.

Table 2:Distribution of impasse turns across problem\-solving practices and the share of each impasse type within a practice \(N=6,630N=6\{,\}630\)\.
### 3\.5Analyses

##### RQ1: Manifestation and persistence\.

To evaluate when specific impasses occur, we cross\-tabulated impasse type by problem\-solving practice and tested independence with a Pearsonχ2\\chi^\{2\}test \(Cramér’sVVas effect size\)\. We then computed next\-turn recovery, whether the impasse is recovered on the subsequent turn, by depth and by impasse type, with 95% confidence intervals from a session\-level cluster bootstrap\. This allows us to assess how persistent each impasse type is, and how the likelihood of recovery changes as impasse depth increases\.

##### RQ2: Controlled comparison of tutor conditions\.

To evaluate how prompt\-based guidance influences an AI tutor’s responses to student impasses, we ran a small simulation study\. First, we drew a stratified random sample of 150 impasse turns \(n=50n=50per impasse type\)\. For each of the sampled impasse turn, we prompted an AI model with three experimental tutoring conditions:*Baseline AI Tutor*,*No\-Direct\-Answer Baseline AI Tutor*,*Guided AI Tutor*\. Each prompt included the problem on which the student had the impasse, preceding dialogue, and student utterance\. The*Baseline AI Tutor*was instructed only to respond as a tutor\. The*No\-Direct\-Answer Baseline AI Tutor*was additionally instructed to guide the student toward the solution rather than give the answer\. The*Guided AI Tutor*is the response the deployed system actually produced, governed by a detailed protocol: elicit a structured plan through sequential questions organized around expert problem\-solving practices, ask one question at a time, avoid supplying correct answers after incorrect or incomplete responses, give increasingly specific hints when needed, keep responses concise, and encourage without praising incorrect answers\. All comparison conditions were generated with GPT 5\.6, and all responses were coded with the same multi\-label set of tutor moves \(scripted question, follow\-up question, check for understanding, direct answer, confirm/elaborate, address incorrect answer, partial answer or hint, praise/encouragement, remind of exercise purpose\)\. We compare conditions on move prevalence and on*response diversity*, the number of distinct move combinations across the 150 responses\. We tested whether the prevalence of each move differed across the three conditions using Pearsonχ2\\chi^\{2\}tests \(Bonferroni\-corrected across the nine moves; Cramér’sVVas effect size\), and quantified response diversity as the number of distinct move combinations and their Shannon entropy\.

##### RQ3: Session outcomes and practice transitions\.

Because a short session may reflect either fast success or early abandonment, we classified sessions using both length and progress\.*Early dropouts*\(N=252N=252, 20\.0%\) ended with fewer than 10 student turns \(corpus mean 16\.2\) without ever reaching the execution or answer\-checking practices\.*Completers*\(N=635N=635, 50\.4%\) reached execution or answer checking, including 68 short sessions \(<10<10turns\) that did so\. The remaining 374 sessions \(29\.7%\) exceeded 10 turns without reaching execution and were excluded to keep the contrast unambiguous\. For each group, we estimated a first\-order transition matrix over the six core practices,P⁡\(PSPt\+1=j∣PSPt=i\)=ni​j/∑kni​kP\(\\mathrm\{PSP\}\_\{t\+1\}=j\\mid\\mathrm\{PSP\}\_\{t\}=i\)=n\_\{ij\}/\\sum\_\{k\}n\_\{ik\}, and the differenceΔ​𝐏=𝐏Completers−𝐏Dropouts\\Delta\\mathbf\{P\}=\\mathbf\{P\}\_\{\\text\{Completers\}\}\-\\mathbf\{P\}\_\{\\text\{Dropouts\}\}\. Two rare practices \(similar problems,N=13N=13; assumptions,N=31N=31\) were excluded\. Because dropouts are defined by never reaching execution, we interpret only transitions among the four pre\-execution practices\.

##### RQ4: Tutor moves and recovery\.

We fit turn\-level logistic regressions predicting recovery att\+1t\+1from the tutor’s response to impasse turntt, using all impasse turns with an observed next student turn \(N=6,163N=6\{,\}163turns, 3,917 episodes,K=1,151K=1\{,\}151sessions\)\. Standard errors are cluster\-robust \(Huber–White\) by session\. All models control for impasse depth \(the turn’s position within its episode\), impasse type \(reference: conceptual error\), and problem\-solving practice \(reference: relevant concepts\)\. Tutor\-move indicators are not mutually exclusive; most tutor turns combine several moves\.

Model 1 \(current and preceding moves\)includes indicators for seven tutor moves in the response to turntt\(𝐌t\\mathbf\{M\}\_\{t\}\), indicators for five moves in the tutor’s response to the preceding impasse turn of the same episode \(𝐌t−1\\mathbf\{M\}\_\{t\-1\}; zero at onset\), and an indicator for repeating the same dominant move on consecutive turns:

logit⁡P⁡\(Yt\+1=1\)=β0\+βd​Deptht\+𝜸⊤​𝐗t\+𝜷⊤​𝐌t\+𝜼⊤​𝐌t−1\+βr​Repeatt,\\operatorname\{logit\}P\(Y\_\{t\+1\}=1\)=\\beta\_\{0\}\+\\beta\_\{d\}\\,\\mathrm\{Depth\}\_\{t\}\+\\boldsymbol\{\\gamma\}^\{\\top\}\\mathbf\{X\}\_\{t\}\+\\boldsymbol\{\\beta\}^\{\\top\}\\mathbf\{M\}\_\{t\}\+\\boldsymbol\{\\eta\}^\{\\top\}\\mathbf\{M\}\_\{t\-1\}\+\\beta\_\{r\}\\,\\mathrm\{Repeat\}\_\{t\},\(1\)where𝐗t\\mathbf\{X\}\_\{t\}holds impasse type and practice\.

Model 2 \(dominant two\-turn sequences\)asks whether specific pairings of consecutive tutor moves matter, which Model 1 cannot show because it estimates the current and preceding moves as separate, additive effects; for example, Model 2 can compare repeating a scripted question after it has failed with switching to addressing the student’s error\. It assigns each tutor response a dominant move \(priority: address incorrect\>\>check for understanding\>\>follow\-up question\>\>scripted question\>\>partial answer\>\>confirm/elaborate\>\>direct answer\) and replaces the move indicators with the sequenceMt−1→MtM\_\{t\-1\}\\rightarrow M\_\{t\}within the episode \(sequences with≥25\\geq 25occurrences; others pooled\), withOnset→\\rightarrowScripted questionas the reference\. Because most tutor responses combine several moves, the priority order ranks moves from most to least diagnostic of the tutor’s strategy, so that supportive moves that accompany almost every response \(confirmation, praise\) do not mask the move that directs the dialogue\. We verify that this choice does not drive the results \(Section[4\.4\.1](https://arxiv.org/html/2609.38346#S4.SS4.SSS1)\)\.

Model 3 \(depth×\\timesmove\)adds to the current\-move specification interactions between depth and five moves central to the assistance dilemma \(scripted question, follow\-up question, address incorrect answer, partial answer, direct answer\), testing whether a move’s association with recovery changes as an impasse persists:

logit⁡P⁡\(Yt\+1=1\)=β0\+βd​Deptht\+𝜸⊤​𝐗t\+𝜷⊤​𝐌t\+𝜹⊤​\(Deptht×𝐌t\)\.\\operatorname\{logit\}P\(Y\_\{t\+1\}=1\)=\\beta\_\{0\}\+\\beta\_\{d\}\\,\\mathrm\{Depth\}\_\{t\}\+\\boldsymbol\{\\gamma\}^\{\\top\}\\mathbf\{X\}\_\{t\}\+\\boldsymbol\{\\beta\}^\{\\top\}\\mathbf\{M\}\_\{t\}\+\\boldsymbol\{\\delta\}^\{\\top\}\(\\mathrm\{Depth\}\_\{t\}\\times\\mathbf\{M\}\_\{t\}\)\.\(2\)This is a discrete\-time event model of episode resolution[Singer & Willett \(\(1993\)\)](https://arxiv.org/html/2609.38346#bib.bib41)\. We report adjusted odds ratios \(AOR\) with 95% confidence intervals and, for Model 3, average predicted recovery with each move present versus absent at depths 1–4\. As robustness checks \(Section[4\.4\.1](https://arxiv.org/html/2609.38346#S4.SS4.SSS1)\), we re\-estimated Model 3 with generalized estimating equations \(exchangeable correlation\)[Liang & Zeger \(\(1986\)\)](https://arxiv.org/html/2609.38346#bib.bib27)and Model 1 with depth as a categorical variable \(1–5, 6\+\)\.

## 4Results

![Refer to caption](https://arxiv.org/html/2609.38346v1/impasse_depth_type_n.png)Figure 1:Next\-turn recovery by impasse depth, for all impasses and by impasse type\. Error bars show 95% confidence intervals from a session\-level cluster bootstrap \(2,000 resamples\); the number of impasse turns at each depth appears in the bottom panel\. Hollow markers indicate fewer than 30 turns\.Table 3:Prevalence of tutor moves across the three tutoring conditions for the same 150 impasses, with tests of between\-conditions differences\. Cells give the percentage of each condition’s responses that contained the move; because a single response could contain several moves, columns do not sum to 100%\.*Response diversity*counts the distinct combinations of moves each condition produced and their Shannon entropy \(higher values indicate more varied responses\)\.### 4\.1RQ1: Where Impasses Occur and How Recovery Declines

Impasses were concentrated in the planning phases of problem solving \(Table[2](https://arxiv.org/html/2609.38346#S3.T2)\): 52\.3% occurred while identifying relevant concepts, 18\.1% during strategy development, and 14\.5% during execution\. These are counts of where impasses were observed, not practice\-specific rates, since practices differ in exposure\. Impasse type depended strongly on practice \(χ2​\(10\)=2313\.9\\chi^\{2\}\(10\)=2313\.9,p<\.001p<\.001, Cramér’sV=\.42V=\.42\)\. Help requests made up most of the impasses in relevant\-concepts phase \(58\.1%\), conceptual errors and uncertainty were roughly balanced during strategy and execution \(45–51% errors\), and nearly all answer\-checking impasses were expressions of uncertainty \(95\.1%\)\.

Across all impasse types, next\-turn recovery fell steadily as impasses persisted, from 39\.7% at onset to 32\.6% at depth 2, 26\.2% at depth 3, and 18\.8% at depth 6 or more, less than half the onset rate \(Figure[1](https://arxiv.org/html/2609.38346#S4.F1)\)\. However, trajectories differed by type\. Help\-seeking was initially the most recoverable state \(47\.0% at onset\) but declined most sharply once initial assistance had failed \(25\.3% at depth 3; 12\.5% at depth≥\\geq6\)\. Conceptual errors dropped to about 21–25% by depth 3 and remained there \(20\.6% at depth≥\\geq6\), whereas uncertainty remained comparatively recoverable through depth 4 \(31\.6% at onset; 33\.3% at depth 4\), beyond which few uncertainty episodes persisted\.

![Refer to caption](https://arxiv.org/html/2609.38346v1/Tutor_Response_Profiles-1.png)Figure 2:Pedagogical response profiles across tutor conditions and impasse types for the same 150 impasses \(n=50n=50per type\)\. Cells show the percentage of responses containing each move; responses can contain several moves, so rows need not sum to 100%\.
### 4\.2RQ2: Prompt Specificity Shapes Tutor Responses to Impasses

The three conditions produced significantly different response profiles for the same 150 impasses \(Figure[2](https://arxiv.org/html/2609.38346#S4.F2)\)\. All nine moves differed significantly across conditions \(all Bonferroni\-correctedp<\.001p<\.001; Cramér’sV=\.20V=\.20–\.56\.56\), with the largest differences for follow\-up questions \(V=\.56V=\.56\), praise \(V=\.52V=\.52\), direct answers \(V=\.49V=\.49\), and scripted questions \(V=\.48V=\.48\)\. Response diversity followed the same ordering: the Guided AI Tutor’s move combinations had an entropy of 5\.13 bits, compared with 3\.11 for the Baseline AI Tutor and 1\.25 for the No\-Direct\-Answer Baseline AI Tutor\. These differences held when the analysis was restricted to the 104 impasses with an unambiguous problem context \(all correctedp<\.05p<\.05\)\.

TheBaseline AI Tutorprovided direct answers in 50\.7% of responses \(vs\. 19\.3% for the Guided AI Tutor and 0% for the No\-Direct\-Answer Baseline AI Tutor\)\. This tendency was stronger for conceptual errors \(70%\), where it also addressed the incorrect answer in every case\. It nevertheless asked follow\-up questions in 62\.7% of responses\. TheNo\-Direct\-Answer Baseline AI Tutoradopted a uniform questioning policy: it asked a follow\-up question in 100% of responses, never gave a direct or partial answer, and produced only 3 distinct move combinations\. TheGuided AI Tutorused the broadest repertoire, including moves absent from both alternatives: scripted questions \(31\.3%\), checks for understanding \(23\.3%\), praise or encouragement \(35\.3%\), and reminders of the exercise purpose \(12\.0%\)\. It also combined questioning with partial answers or hints \(39\.3%\) and confirmation or elaboration \(42\.0%\)\.

The Guided AI Tutor’s responses also varied with the type of impasse\. Help requests most often elicited scripted questions \(56%\); conceptual errors elicited follow\-up questions \(50%\), confirmation or elaboration \(48%\), hints \(40%\), and error\-directed feedback \(38%\); and uncertainty most often elicited confirmation or elaboration \(74%\)\. Table[4](https://arxiv.org/html/2609.38346#S4.T4)illustrates the contrast for one conceptual error: the Baseline AI Tutor produces the full solution procedure, the No\-Direct\-Answer Baseline AI Tutor redirects with questions, and the Guided AI Tutor names the misconception \(the apparent mass is a weighted average and must lie between the two isotope masses\) before asking a follow\-up question\.

In short, a limited prompt instructions not to give answers did not produce a nuanced pedagogy; it produced uniform questioning\. A detailed protocol produced varied, impasse\-sensitive responses\. Constraining answer\-giving is therefore not equivalent to specifying a pedagogical policy\.

Table 4:Responses of the three tutor conditions to the same conceptual\-error impasse\. Coded tutor moves are shown in brackets and italicized\.![Refer to caption](https://arxiv.org/html/2609.38346v1/fig_markov_compact.png)Figure 3:First\-order practice\-transition probabilities for \(A\) completers, \(B\) early dropouts, and \(C\) their differenceΔ​𝐏\\Delta\\mathbf\{P\}\. Transitions into execution and answer checking are zero for dropouts by definition and are not interpreted\.
### 4\.3RQ3: Early Dropouts Are Trapped in Concept Elicitation

Figure[3](https://arxiv.org/html/2609.38346#S4.F3)shows practice\-transition matrices for completers and early dropouts\. Overall, students who completed the problem showed a broader distribution of transitions and were generally less likely to transition back to relevant concepts and more likely to transition to strategy\-focused practices\. Both groups spent much of their time on relevant concepts, but completers moved forward through the problem\-solving sequence: from problem definition, they moved either to relevant concepts \(P=\.42P=\.42\) or directly to strategy \(P=\.29P=\.29\); from strategy, they progressed to execution \(P=\.13P=\.13\); and from execution, to answer checking \(P=\.11P=\.11\)\.

Early dropouts showed a different structure before execution\. From problem definition, 71% of their transitions led to relevant concepts \(vs\. 42% for completers\)\. Once there, they stayed there: the concepts self\-transition wasP=\.74P=\.74\. Even when they reached strategy, 60% of their next turns returned to relevant concepts, compared with 18% for completers \(Δ=−\.42\\Delta=\-\.42\); the same backward pull appeared from information needed \(Δ=−\.28\\Delta=\-\.28\)\. Because the tutor sets the practice of each exchange through its scripted sequence, these loops reflect the joint behavior of student and tutor: when students could not supply the concepts the tutor elicited, the tutor continued to elicit them\. Combined with RQ1, where 52\.3% of impasses occurred in the concepts phase, this suggests that early attrition is associated with unresolved loops at the start of planning\. Because dropouts are defined by never reaching execution, this finding describes where their sessions stalled rather than ruling out difficulties they might have met later\.

### 4\.4RQ4: The Value of Questioning Decays as Impasses Persist

##### Impasse depth and state\.

In Model 1, each additional turn at an impasse was associated with 13% lower odds of recovery \(AOR=0\.87=0\.87, 95% CI \[0\.81, 0\.94\],p<\.001p<\.001\), controlling for impasse type, practice, and tutor moves\. Students expressing uncertainty \(AOR=1\.37=1\.37, \[1\.17, 1\.60\]\) or asking for help \(AOR=1\.44=1\.44, \[1\.17, 1\.77\]\) were more likely to recover than those making conceptual errors\. Relative to relevant concepts, impasses during answer checking \(AOR=0\.57=0\.57, \[0\.41, 0\.80\]\) and strategy \(AOR=0\.81=0\.81, \[0\.68, 0\.96\]\) were less recoverable; other practices did not differ significantly\.

Table 5:Model 2: Recovery following selected two\-turn scaffolding sequences \(Mt−1→MtM\_\{t\-1\}\\rightarrow M\_\{t\}\), relative to a first scripted question at onset, adjusted for depth, impasse type, and practice\. Bold rows mark the reference and the two sequences compared in the text \(repeating vs\. switching after a scripted question\)\.
##### Current and preceding moves\.

In the tutor’s response to the current impasse turn, scripted questions \(AOR=2\.24=2\.24, \[1\.77, 2\.85\]\), follow\-up questions \(AOR=1\.95=1\.95, \[1\.57, 2\.44\]\), and addressing the incorrect answer \(AOR=1\.36=1\.36, \[1\.11, 1\.65\]\) were associated with higher recovery, whereas partial answers \(AOR=0\.81=0\.81, \[0\.70, 0\.94\]\) and checks for understanding \(AOR=0\.37=0\.37, \[0\.28, 0\.48\]\) were associated with lower recovery; direct answers and confirmation did not differ significantly\. The latter result is likely partly a measurement artifact, because questions such as “Does that make sense?” invite brief affirmations that cannot be coded as completely correct\. The tutor’s preceding move also mattered: a scripted question on the preceding impasse turn was associated with lower recovery \(AOR=0\.78=0\.78, \[0\.62, 0\.98\],p=\.032p=\.032\), whereas preceding error\-directed feedback was marginally associated with higher recovery \(AOR=1\.26=1\.26, \[0\.99, 1\.61\],p=\.062p=\.062\)\. Repeating the same dominant move had no additional effect \(AOR=0\.99=0\.99\)\. Full model estimates are provided in the supplementary material\.

##### Scaffolding sequences\.

Model 2 \(Table[5](https://arxiv.org/html/2609.38346#S4.T5)\) shows the same pattern at the level of sequences\. A first scripted question at onset was followed by recovery in 55\.5% of cases\. When a scripted question followed another scripted question, recovery fell to 28\.1%, halving the odds \(AOR=0\.48=0\.48, \[0\.34, 0\.67\],p<\.001p<\.001\)\. Moving from a scripted question to addressing the student’s error preserved recovery \(39\.8%; AOR=0\.78=0\.78,p=\.29p=\.29, not different from onset\)\. Sequences ending in a check for understanding had the lowest recovery \(7\.9–13\.6%; AOR0\.100\.10–0\.190\.19, allp<\.001p<\.001\)\. Model 2’s dominant\-move assignment rarely overrode the moves of interest: 90% of tutor responses containing a scripted question were assigned that move\. Recomputing the key sequences without any priority order, counting a sequence whenever both responses contained the move, reproduced the pattern \(first scripted question: 54\.4%; scripted question→\\rightarrowscripted question: 30\.5%; scripted question→\\rightarrowaddress incorrect: 37\.5%\)\. An order\-free specification that added a \(scripted question att−1t\-1\)×\\times\(scripted question attt\) interaction to Model 1 likewise showed that a repeated scripted question carried substantially less benefit \(AOR=0\.57=0\.57, 95% CI \[0\.40, 0\.82\],p=\.002p=\.002\)\.

![Refer to caption](https://arxiv.org/html/2609.38346v1/fig_depth_move_interaction.png)Figure 4:Average predicted next\-turn recovery with each tutor move present versus absent, by impasse depth \(Model 3\)\. Labels give the difference in percentage points at depths 1 and 4\.
##### Depth\-dependent move effects\.

Model 3 tests directly whether tutor move effects change as an impasse persists \(Figure[4](https://arxiv.org/html/2609.38346#S4.F4)\)\. The benefit of scripted questions \(depth×\\timesscripted question AOR=0\.78=0\.78, \[0\.66, 0\.91\],p=\.002p=\.002\) and of follow\-up questions \(AOR=0\.83=0\.83, \[0\.73, 0\.95\],p=\.008p=\.008\) declined with each additional impasse turn, whereas the benefit of addressing the student’s incorrect answer increased \(AOR=1\.14=1\.14, \[1\.02, 1\.28\],p=\.026p=\.026\)\. In predicted terms, a scripted question was associated with a 20 percentage\-point \(pp\) higher recovery rate at onset \(51\.0% vs\. 31\.1%\) but only 3 pp by the fourth impasse turn; follow\-up questions fell from\+15\+15to\+3\+3pp; and addressing the error rose from\+5\+5to\+12\+12pp\. Direct answers were associated with slightly lower recovery at every depth \(−3\-3to−2\-2pp\); their interaction with depth was not significant\.

#### 4\.4\.1Robustness

The depth\-by\-move interaction effects from Model 3 remain statistically significant even after accounting for clustering\(scripted question0\.810\.81,p=\.006p=\.006; follow\-up question0\.860\.86,p=\.028p=\.028; address incorrect1\.131\.13,p=\.030p=\.030\)\. Because the continuous depth term assumes a constant per\-turn decline, we also modeled depth categorically \(1–5 and 6\+, matching Figure[1](https://arxiv.org/html/2609.38346#S4.F1)\)\. Relative to onset, recovery odds were lower at every depth and lowest for the longest impasses \(depth 2:0\.830\.83; depth 3:0\.610\.61; depth 4:0\.640\.64; depth 5:0\.550\.55; depth≥\\geq6:0\.430\.43, 95% CI \[0\.25, 0\.74\]; allp≤\.01p\\leq\.01\)\.

## 5Discussion

This study examined how a guardrailed LLM tutor responds as student impasses persist across turns\. Across the controlled comparison and authentic dialogue analyses, a consistent distinction emerged between constraining tutor behavior and adapting assistance to an evolving interaction state\. A more targeted tutoring protocol produced more varied tutor responses, but persistent impasses still exposed limits in how support changed after earlier interventions failed\.

Our findings suggest that constraining an AI tutor and making it adaptive are distinct design problems\. In the controlled comparison, simply instructing a tutor not to provide answers produced a relatively rigid questioning policy, whereas a more targeted tutoring protocol produced more varied responses across impasse contexts\. However, even the deployed tutor, which was explicitly instructed to provide increasingly specific hints after incorrect or incomplete responses, sometimes continued with questioning after earlier questioning had failed to resolve the impasse\. These findings suggest that pedagogical guardrails should specify not only constraints on undesirable behaviors such as answer revelation, but also how support should adapt when an initial intervention fails\. This extends work on guardrailed LLM tutoring[Liffiton et al\. \(\(2023\)\)](https://arxiv.org/html/2609.38346#bib.bib28);[Bastani et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib6);[Jurenka et al\. \(\(2024\)\)](https://arxiv.org/html/2609.38346#bib.bib18)by shifting the design problem from response\-level constraint to multi\-turn adaptation, and similarly suggests that response\-level evaluation criteria[Tack & Piech \(\(2022\)\)](https://arxiv.org/html/2609.38346#bib.bib42);[Maurya et al\. \(\(2025\)\)](https://arxiv.org/html/2609.38346#bib.bib30)should be complemented by measures of whether tutor behavior changes appropriately across successive failed interactions\. Our temporal analyses add a separate implication: the association between a tutor move and recovery is not fixed across an impasse\. Questioning was strongly associated with recovery when an impasse first emerged, but this association weakened as the impasse persisted, whereas addressing an incorrect response became relatively more positively associated with recovery\. This extends the assistance dilemma[Koedinger & Aleven \(\(2007\)\)](https://arxiv.org/html/2609.38346#bib.bib24)by suggesting that the appropriateness of support may depend not only on the learner’s current difficulty, but also on how long that difficulty has persisted\. It also qualifies work emphasizing the value of eliciting student reasoning[VanLehn et al\. \(\(2003\)\)](https://arxiv.org/html/2609.38346#bib.bib47);[Chi et al\. \(\(2001\)\)](https://arxiv.org/html/2609.38346#bib.bib11): questioning was associated with higher recovery early in an impasse, but this association weakened following repeated unresolved turns\. This pattern is consistent with contingent tutoring[H\. Wood & Wood \(\(1999\)\)](https://arxiv.org/html/2609.38346#bib.bib51)and graduated hinting in intelligent tutoring systems[VanLehn \(\(2011\)\)](https://arxiv.org/html/2609.38346#bib.bib46), where support becomes more explicit after failure\. For LLM tutors, impasse persistence may therefore serve as a practical signal for when to consider shifting from elicitation toward more targeted assistance\.

From a learning analytics perspective, impasses may be better represented as evolving interaction processes than as isolated incorrect turns\. In our case, recovery declined as impasses deepened, and the associations between particular tutor moves and recovery changed with depth, indicating that the sequence and history of an interaction contain information that a single response cannot capture in isolation\. This supports temporal approaches to learning analytics that treat the ordering and duration of events as analytically meaningful[Reimann \(\(2009\)\)](https://arxiv.org/html/2609.38346#bib.bib36);[Chen et al\. \(\(2018\)\)](https://arxiv.org/html/2609.38346#bib.bib10)\. Methodologically, impasse type, depth, and recovery provide relatively simple dialogue\-derived measures for representing this process over time\. For LLM\-tutor evaluation, this shifts attention from whether an individual response is pedagogically appropriate toward whether assistance remains appropriate as the interaction unfolds\. The trajectories of students who left sessions early further illustrate the value of this temporal view\. These students became concentrated in repeated concept\-elicitation loops before reaching problem execution, suggesting that stalled progress can emerge from the interaction between student difficulty and a tutor strategy that continues to elicit information the student is not providing\. This resembles wheel\-spinning[Beck & Gong \(\(2013\)\)](https://arxiv.org/html/2609.38346#bib.bib7);[Gong & Beck \(\(2015\)\)](https://arxiv.org/html/2609.38346#bib.bib16);[Kai et al\. \(\(2018\)\)](https://arxiv.org/html/2609.38346#bib.bib19), but at the level of tutoring dialogue rather than repeated skill practice\. Repeated nonprogress may therefore provide learning analytics systems with an early signal of a stalled tutoring interaction, creating an opportunity to identify when adaptation may be needed before disengagement occurs\. Taken together, these findings point toward graduated, state\-sensitive assistance as a useful design target for LLM tutors\. Rather than treating questioning or answer withholding as fixed pedagogical policies, tutors could use signals such as impasse type, persistence, and prior unsuccessful assistance to determine when more targeted support may be warranted\. Such escalation need not imply immediately providing an answer; it could instead move from elicitation toward error\-directed feedback, increasingly specific hints, or worked sub\-steps as evidence accumulates that the current strategy is not resolving the impasse\. This remains a design hypothesis rather than a demonstrated optimal policy, however, and should be tested through randomized comparisons that connect escalation strategies to both immediate recovery and subsequent learning\.

### 5\.1Limitations and Future Work

Although the current work provides insights into students’ impasse recovery process while working with an AI tutor, there are a few notable limitations\. The authentic dialogue analyses are observational, so associations between tutor moves and recovery are not causal\. Observations at greater impasse depths also represent the subset of episodes that remained unresolved at earlier turns, so depth\-dependent associations may partly reflect differences in the composition of persistent impasses rather than changes caused by persistence itself\. Recovery also captures only next\-turn correctness, not longer\-term learning\. Our impasse categories reflect observable dialogue behaviors rather than latent cognitive or affective states, and the full corpus was coded with a human\-validated LLM pipeline rather than exhaustive human annotation\. Generalizability is also limited by the use of a single guided tutor in one undergraduate chemistry setting, while the controlled comparison examines tutor response policies rather than student outcomes\. Future work should test these patterns across domains and tutor architectures, connect impasse recovery to longer\-term learning, and experimentally compare escalation strategies\.

## References

- Aleven et al\. \(\(2006\)\)Aleven, V\., McLaren, B\., Roll, I\. & Koedinger, K\.\(2006\)\.Toward meta\-cognitive tutoring: A model of help seeking with a cognitive tutor\.International Journal of Artificial Intelligence in Education 16 2 101–128\.
- Aleven et al\. \(\(2003\)\)Aleven, V\., Stahl, E\., Schworm, S\., Fischer, F\. & Wallace, R\.\(2003\)\.Help seeking and help design in interactive learning environments\.Review of Educational Research 73 3 277–320\.doi:10\.3102/00346543073003277
- Anonymous \(\(2026\)\)Anonymous\.\(2026\)\.Development of an AI tutor to scaffold problem\-solving \(details withheld for review\)\.
- Bakeman & Quera \(\(1995\)\)Bakeman, R\. & Quera, V\.\(1995\)\.Analyzing interaction: Sequential analysis with SDIS and GSEQ\.: Cambridge University Press\.
- Baker et al\. \(\(2010\)\)Baker, R\.S\.J\.d\., D’Mello, S\.K\., Rodrigo, M\.M\.T\. & Graesser, A\.C\.\(2010\)\.Better to be frustrated than bored: The incidence, persistence, and impact of learners’ cognitive–affective states during interactions with three different computer\-based learning environments\.International Journal of Human\-Computer Studies 68 4 223–241\.doi:10\.1016/j\.ijhcs\.2009\.12\.003
- Bastani et al\. \(\(2025\)\)Bastani, H\., Bastani, O\., Sungu, A\., Ge, H\., Kabakçı, Ö\. & Mariman, R\.\(2025\)\.Generative AI without guardrails can harm learning: Evidence from high school mathematics\.Proceedings of the National Academy of Sciences 122 26 e2422633122\.doi:10\.1073/pnas\.2422633122
- Beck & Gong \(\(2013\)\)Beck, J\.E\. & Gong, Y\.\(2013\)\.Wheel\-spinning: Students who fail to master a skill\.In Artificial intelligence in education \(AIED 2013\) \( 431–440\)\.
- Borchers et al\. \(\(2024\)\)Borchers, C\., Yang, K\., Lin, J\., Rummel, N\., Koedinger, K\.R\. & Aleven, V\.\(2024\)\.Combining dialog acts and skill modeling: What chat interactions enhance learning rates during AI\-supported peer tutoring?In Proceedings of EDM 2024\.
- Burkholder et al\. \(\(2020\)\)Burkholder, E\.W\., Miles, J\.K\., Layden, T\.J\., Wang, K\.D\., Fritz, A\.V\. & Wieman, C\.E\.\(2020\)\.Template for teaching and assessment of problem solving in introductory physics\.Physical Review Physics Education Research 16 1 010123\.doi:10\.1103/PhysRevPhysEducRes\.16\.010123
- Chen et al\. \(\(2018\)\)Chen, B\., Knight, S\. & Wise, A\.F\.\(2018\)\.Critical issues in designing and implementing temporal analytics\.Journal of Learning Analytics 5 1 1–9\.doi:10\.18608/jla\.2018\.53\.1
- Chi et al\. \(\(2001\)\)Chi, M\.T\.H\., Siler, S\.A\., Jeong, H\., Yamauchi, T\. & Hausmann, R\.G\.\(2001\)\.Learning from human tutoring\.Cognitive Science 25 4 471–533\.doi:10\.1207/s15516709cog2504˙1
- Cohen \(\(1960\)\)Cohen, J\.\(1960\)\.A coefficient of agreement for nominal scales\.Educational and Psychological Measurement 20 1 37–46\.doi:10\.1177/001316446002000104
- Daheim et al\. \(\(2024\)\)Daheim, N\., Macina, J\., Kapur, M\., Gurevych, I\. & Sachan, M\.\(2024\)\.Stepwise verification and remediation of student reasoning errors with large language model tutors\.In Proceedings of EMNLP 2024 \( 8386–8411\)\.
- D’Mello et al\. \(\(2014\)\)D’Mello, S\., Lehman, B\., Pekrun, R\. & Graesser, A\.\(2014\)\.Confusion can be beneficial for learning\.Learning and Instruction 29 153–170\.doi:10\.1016/j\.learninstruc\.2012\.05\.003
- Gilardi et al\. \(\(2023\)\)Gilardi, F\., Alizadeh, M\. & Kubli, M\.\(2023\)\.ChatGPT outperforms crowd workers for text\-annotation tasks\.Proceedings of the National Academy of Sciences 120 30 e2305016120\.doi:10\.1073/pnas\.2305016120
- Gong & Beck \(\(2015\)\)Gong, Y\. & Beck, J\.E\.\(2015\)\.Towards detecting wheel\-spinning: Future failure in mastery learning\.In Proceedings of L@S 2015 \( 67–74\)\.
- Graesser et al\. \(\(1995\)\)Graesser, A\.C\., Person, N\.K\. & Magliano, J\.P\.\(1995\)\.Collaborative dialogue patterns in naturalistic one\-to\-one tutoring\.Applied Cognitive Psychology 9 6 495–522\.doi:10\.1002/acp\.2350090604
- Jurenka et al\. \(\(2024\)\)Jurenka, I\., Kunesch, M\., McKee, K\.R\., Gillick, D\. et al\.\(2024\)\.Towards responsible development of generative AI for education: An evaluation\-driven approach\.arXiv preprint arXiv:2407\.12687\.
- Kai et al\. \(\(2018\)\)Kai, S\., Almeda, M\.V\., Baker, R\.S\., Heffernan, C\. & Heffernan, N\.\(2018\)\.Decision tree modeling of wheel\-spinning and productive persistence in skill builders\.Journal of Educational Data Mining 10 1 36–71\.
- Kapur \(\(2008\)\)Kapur, M\.\(2008\)\.Productive failure\.Cognition and Instruction 26 3 379–424\.doi:10\.1080/07370000802212669
- Kapur \(\(2016\)\)Kapur, M\.\(2016\)\.Examining productive failure, productive success, unproductive failure, and unproductive success in learning\.Educational Psychologist 51 2 289–299\.doi:10\.1080/00461520\.2016\.1155457
- Kestin et al\. \(\(2025\)\)Kestin, G\., Miller, K\., Klales, A\., Milbourne, T\. & Ponti, G\.\(2025\)\.AI tutoring outperforms in\-class active learning: An RCT introducing a novel research\-based design in an authentic educational setting\.Scientific Reports 15 17458\.doi:10\.1038/s41598\-025\-97652\-6
- Knight et al\. \(\(2017\)\)Knight, S\., Wise, A\.F\. & Chen, B\.\(2017\)\.Time for change: Why learning analytics needs temporal analysis\.Journal of Learning Analytics 4 3 7–17\.
- Koedinger & Aleven \(\(2007\)\)Koedinger, K\.R\. & Aleven, V\.\(2007\)\.Exploring the assistance dilemma in experiments with cognitive tutors\.Educational Psychology Review 19 3 239–264\.doi:10\.1007/s10648\-007\-9049\-0
- Landis & Koch \(\(1977\)\)Landis, J\.R\. & Koch, G\.G\.\(1977\)\.The measurement of observer agreement for categorical data\.Biometrics 33 1 159–174\.doi:10\.2307/2529310
- Lehmann et al\. \(\(2024\)\)Lehmann, M\., Cornelius, P\.B\. & Sting, F\.J\.\(2024\)\.AI meets the classroom: When do large language models harm learning?arXiv preprint arXiv:2409\.09047\.
- Liang & Zeger \(\(1986\)\)Liang, K\-Y\. & Zeger, S\.L\.\(1986\)\.Longitudinal data analysis using generalized linear models\.Biometrika 73 1 13–22\.doi:10\.1093/biomet/73\.1\.13
- Liffiton et al\. \(\(2023\)\)Liffiton, M\., Sheese, B\., Savelka, J\. & Denny, P\.\(2023\)\.CodeHelp: Using large language models with guardrails for scalable support in programming classes\.In Proceedings of koli calling 2023\.
- Macina et al\. \(\(2023\)\)Macina, J\., Daheim, N\., Chowdhury, S\., Sinha, T\., Kapur, M\., Gurevych, I\. & Sachan, M\.\(2023\)\.MathDial: A dialogue tutoring dataset with rich pedagogical properties grounded in math reasoning problems\.In Findings of EMNLP 2023 \( 5602–5621\)\.
- Maurya et al\. \(\(2025\)\)Maurya, K\.K\., Srivatsa, K\.A\., Petukhova, K\. & Kochmar, E\.\(2025\)\.Unifying AI tutor evaluation: An evaluation taxonomy for pedagogical ability assessment of LLM\-powered AI tutors\.In Proceedings of NAACL 2025\.doi:10\.18653/v1/2025\.naacl\-long\.57
- Pardos & Bhandari \(\(2024\)\)Pardos, Z\.A\. & Bhandari, S\.\(2024\)\.ChatGPT\-generated help produces learning gains equivalent to human tutor\-authored help on mathematics skills\.PLOS ONE 19 5 e0304013\.doi:10\.1371/journal\.pone\.0304013
- Pauzi et al\. \(\(2025\)\)Pauzi, Z\., Dodman, M\. & Mavrikis, M\.\(2025\)\.Automating pedagogical evaluation of LLM\-based conversational agents\.In CEUR workshop proceedings \( 4006\)\.
- Petukhova & Kochmar \(\(2025\)\)Petukhova, K\. & Kochmar, E\.\(2025\)\.Intent matters: Enhancing AI tutoring with fine\-grained pedagogical intent annotation\.arXiv preprint arXiv:2506\.07626\.
- A\.M\. Price et al\. \(\(2021\)\)Price, A\.M\., Kim, C\.J\., Burkholder, E\.W\., Fritz, A\.V\. & Wieman, C\.E\.\(2021\)\.A detailed characterization of the expert problem\-solving process in science and engineering: Guidance for teaching and assessment\.CBE—Life Sciences Education 20 3 ar43\.doi:10\.1187/cbe\.20\-12\-0276
- T\.W\. Price et al\. \(\(2017\)\)Price, T\.W\., Zhi, R\. & Barnes, T\.\(2017\)\.Hint generation under uncertainty: The effect of hint quality on help\-seeking behavior\.In Artificial intelligence in education \(AIED 2017\) \( 311–322\)\.
- Reimann \(\(2009\)\)Reimann, P\.\(2009\)\.Time is precious: Variable\- and event\-centred approaches to process analysis in CSCL research\.International Journal of Computer\-Supported Collaborative Learning 4 3 239–257\.doi:10\.1007/s11412\-009\-9070\-z
- Scarlatos, Baker & Lan \(\(2025\)\)Scarlatos, A\., Baker, R\.S\. & Lan, A\.\(2025\)\.Exploring knowledge tracing in tutor\-student dialogues using LLMs\.In Proceedings of LAK ’25\.doi:10\.1145/3706468\.3706501
- Scarlatos, Liu et al\. \(\(2025\)\)Scarlatos, A\., Liu, N\., Lee, J\., Baraniuk, R\. & Lan, A\.\(2025\)\.Training LLM\-based tutors to improve student learning outcomes in dialogues\.In Artificial intelligence in education \(AIED 2025\) \( 251–266\)\.
- Schwartz Poehlmann et al\. \(\(2024\)\)Schwartz Poehlmann, J\.K\., Nardo, J\.E\., Rojas, M\. & Salehi, S\.\(2024\)\.Introducing the problem\-solving template as a tool for equity: Addressing incoming preparation disparities\.Journal of Chemical Education 101 3 1332–1340\.doi:10\.1021/acs\.jchemed\.3c00732
- Sheese et al\. \(\(2024\)\)Sheese, B\., Liffiton, M\., Savelka, J\. & Denny, P\.\(2024\)\.Patterns of student help\-seeking when using a large language model\-powered programming assistant\.In Proceedings of ACE ’24 \( 49–57\)\.
- Singer & Willett \(\(1993\)\)Singer, J\.D\. & Willett, J\.B\.\(1993\)\.It’s about time: Using discrete\-time survival analysis to study duration and the timing of events\.Journal of Educational Statistics 18 2 155–195\.doi:10\.3102/10769986018002155
- Tack & Piech \(\(2022\)\)Tack, A\. & Piech, C\.\(2022\)\.The AI teacher test: Measuring the pedagogical ability of Blender and GPT\-3 in educational dialogues\.In Proceedings of EDM 2022 \(p\. 529\)\.
- Thomas et al\. \(\(2024\)\)Thomas, D\.R\., Lin, J\., Gatz, E\., Gurung, A\., Gupta, S\., Norberg, K\.Koedinger, K\.R\.\(2024\)\.Improving student learning with hybrid human\-AI tutoring: A three\-study quasi\-experimental investigation\.In Proceedings of LAK ’24 \( 404–415\)\.
- VanLehn \(\(1988\)\)VanLehn, K\.\(1988\)\.Toward a theory of impasse\-driven learning\.In H\. Mandl & A\. Lesgold \(Eds\.\), Learning issues for intelligent tutoring systems \( 19–41\)\.: Springer\.doi:10\.1007/978\-1\-4684\-6350\-7˙2
- VanLehn \(\(1999\)\)VanLehn, K\.\(1999\)\.Rule\-learning events in the acquisition of a complex skill: An evaluation of Cascade\.Journal of the Learning Sciences 8 1 71–125\.doi:10\.1207/s15327809jls0801˙3
- VanLehn \(\(2011\)\)VanLehn, K\.\(2011\)\.The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems\.Educational Psychologist 46 4 197–221\.doi:10\.1080/00461520\.2011\.611369
- VanLehn et al\. \(\(2003\)\)VanLehn, K\., Siler, S\., Murray, C\., Yamauchi, T\. & Baggett, W\.B\.\(2003\)\.Why do only some events cause learning during human tutoring?Cognition and Instruction 21 3 209–249\.doi:10\.1207/S1532690XCI2103˙01
- K\.D\. Wang et al\. \(\(2025\)\)Wang, K\.D\., Wu, Z\., Tufts, L\., Wieman, C\., Salehi, S\. & Haber, N\.\(2025\)\.Scaffold or crutch? Examining college students’ use and views of generative AI tools for STEM education\.In IEEE EDUCON 2025 \( 1–10\)\.
- R\.E\. Wang et al\. \(\(2024\)\)Wang, R\.E\., Ribeiro, A\.T\., Robinson, C\.D\., Loeb, S\. & Demszky, D\.\(2024\)\.Tutor CoPilot: A human\-AI approach for scaling real\-time expertise\.arXiv preprint arXiv:2410\.03017\.
- D\. Wood et al\. \(\(1976\)\)Wood, D\., Bruner, J\.S\. & Ross, G\.\(1976\)\.The role of tutoring in problem solving\.Journal of Child Psychology and Psychiatry 17 2 89–100\.doi:10\.1111/j\.1469\-7610\.1976\.tb00381\.x
- H\. Wood & Wood \(\(1999\)\)Wood, H\. & Wood, D\.\(1999\)\.Help seeking, learning and contingent tutoring\.Computers & Education 33 2–3 153–169\.doi:10\.1016/S0360\-1315\(99\)00030\-5
- Yan et al\. \(\(2024\)\)Yan, L\., Martinez\-Maldonado, R\. & Gašević, D\.\(2024\)\.Generative artificial intelligence in learning analytics: Contextualising opportunities and challenges through the learning analytics cycle\.In Proceedings of LAK ’24\.

Similar Articles

TutorMoments: Do AI tutors know when to help and when to hold back?

Hugging Face Blog

Ai2 introduces TutorMoments, a replay-based evaluation framework and dataset for measuring whether LLMs can balance when to help and when to hold back in one-on-one math tutoring. Preliminary results show models tend to over-help, and prompt engineering only partially closes the gap to human tutors.

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

Hugging Face Daily Papers

LLM-as-a-Tutor introduces a framework that extends LLM's role from judge to tutor by dynamically adjusting prompt difficulty through pairwise comparison and constraint addition, improving instruction-following performance in reinforcement learning.