I code or AI code: A comparative evaluation of AI-rated scores in classroom observations

arXiv cs.CL Papers

Summary

This study evaluates using GPT-5 to score teacher-child interactions in early childhood classrooms against human raters, finding partial alignment but limitations for full assessment.

arXiv:2609.18274v1 Announce Type: new Abstract: Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on trained observers. This study evaluated the feasibility of using a LLM (GPT-5 model) to score teacher-child interactions in early childhood classrooms, benchmarked against human raters. The study analyzed 87 video-recorded observations from 38 classrooms across 30 kindergartens in Hong Kong. Using observation transcripts, the AI model was configured to apply the full Classroom Assessment Scoring System (CLASS) framework. AI-rated scores were then compared with human ratings by examining correlations and differences in mean scores of the CLASS domains and dimensions. The results showed greater convergence between AI and raters for the Emotional Support domain and, in particular, the Quality of Feedback dimension, which captures how teachers use feedback to extend children's learning. Greater divergence emerged for interactions that were more procedural or context-dependent, particularly within the Classroom Organization and Instructional Support domains. These findings suggest that transcript-based AI scoring may capture some of the relative variation in teacher-child interactions but cannot yet reproduce calibrated human judgements consistently across the full CLASS framework. AI-assisted observation may therefore be more appropriate as a preliminary screening tool rather than as a replacement for trained observers, providing teachers with evidence for reflection rather than high-stakes evaluation. Future research should examine whether domain-specific training and incorporation of contextual and visual information can improve alignment between AI and human rated scores.
Original Article
View Cached Full Text

Cached at: 09/17/26, 09:12 AM

# I code or AI code: A comparative evaluation of AI-rated scores in classroom observations
Source: [https://arxiv.org/abs/2609.18274](https://arxiv.org/abs/2609.18274)
[View PDF](https://arxiv.org/pdf/2609.18274)

> Abstract:Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource\-intensive and dependent on trained observers\. This study evaluated the feasibility of using a LLM \(GPT\-5 model\) to score teacher\-child interactions in early childhood classrooms, benchmarked against human raters\. The study analyzed 87 video\-recorded observations from 38 classrooms across 30 kindergartens in Hong Kong\. Using observation transcripts, the AI model was configured to apply the full Classroom Assessment Scoring System \(CLASS\) framework\. AI\-rated scores were then compared with human ratings by examining correlations and differences in mean scores of the CLASS domains and dimensions\. The results showed greater convergence between AI and raters for the Emotional Support domain and, in particular, the Quality of Feedback dimension, which captures how teachers use feedback to extend children's learning\. Greater divergence emerged for interactions that were more procedural or context\-dependent, particularly within the Classroom Organization and Instructional Support domains\. These findings suggest that transcript\-based AI scoring may capture some of the relative variation in teacher\-child interactions but cannot yet reproduce calibrated human judgements consistently across the full CLASS framework\. AI\-assisted observation may therefore be more appropriate as a preliminary screening tool rather than as a replacement for trained observers, providing teachers with evidence for reflection rather than high\-stakes evaluation\. Future research should examine whether domain\-specific training and incorporation of contextual and visual information can improve alignment between AI and human rated scores\.

## Submission history

From: Yasmin Fong \[[view email](https://arxiv.org/show-email/5f0fb1b9/2609.18274)\] **\[v1\]**Wed, 16 Sep 2026 07:53:28 UTC \(558 KB\)

Similar Articles

Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation

arXiv cs.AI

This paper introduces a compete-then-collaborate framework where multiple frontier AI teachers (Claude, Codex-GPT, Grok, Gemini) are ranked by execution-based tests and then collaborate to build a verifiable curriculum. It finds that imitation (SFT) on teacher solutions degrades a competent coder student, while using the same curriculum for reinforcement learning with verifiable rewards (RLVR) improves performance, particularly on competition problems.

Are AI coding agents hitting a wall, or are we just measuring them wrong?

Reddit r/AI_Agents

This article examines the gap between hype and reality for AI coding agents, arguing that they are effective for accelerating workflow parts but still require human oversight for architecture, debugging, and review, and questioning whether current benchmarks measure the right things.

Teaching with AI

OpenAI Blog

OpenAI shares perspectives from educators on integrating AI tools like ChatGPT into teaching, including using AI for language support and teaching students to think critically about AI-generated information.