I code or AI code: A comparative evaluation of AI-rated scores in classroom observations
In plain terms
Classroom observations are essential for improving teaching quality and guiding how teachers teach, but they are expensive and need highly trained human experts. This paper investigated whether a large language model (LLM), specifically GPT-5, could automatically score teacher-child interactions in early childhood classrooms. The researchers analyzed 87 video-recorded observations from kindergartens in Hong Kong, using only the observation transcripts. The AI was configured to apply the Classroom Assessment Scoring System (CLASS), a common framework for evaluating classroom quality. They found that AI scores aligned better with human scores for aspects of "Emotional Support," especially how teachers give feedback. However, there was less agreement for more routine or context-dependent interactions in "Classroom Organization" and "Instructional Support." This suggests that while AI can capture some differences in teacher-child interactions from text, it cannot yet consistently match the nuanced judgments of trained human observers. Therefore, AI-assisted observation might be more useful as a preliminary tool for teacher reflection rather than for high-stakes evaluations.