← All research

AI for Learning

AI shows promise for scoring some aspects of teacher-child interactions in early childhood classrooms, but it can't fully replace human observers yet.

cs.CLBeginner-friendly

I code or AI code: A comparative evaluation of AI-rated scores in classroom observations

Y. Fong, J. Xiang, T. Y. D. Chan, K. Lee +1 more

In plain terms

Classroom observations are essential for improving teaching quality and guiding how teachers teach, but they are expensive and need highly trained human experts. This paper investigated whether a large language model (LLM), specifically GPT-5, could automatically score teacher-child interactions in early childhood classrooms. The researchers analyzed 87 video-recorded observations from kindergartens in Hong Kong, using only the observation transcripts. The AI was configured to apply the Classroom Assessment Scoring System (CLASS), a common framework for evaluating classroom quality. They found that AI scores aligned better with human scores for aspects of "Emotional Support," especially how teachers give feedback. However, there was less agreement for more routine or context-dependent interactions in "Classroom Organization" and "Instructional Support." This suggests that while AI can capture some differences in teacher-child interactions from text, it cannot yet consistently match the nuanced judgments of trained human observers. Therefore, AI-assisted observation might be more useful as a preliminary tool for teacher reflection rather than for high-stakes evaluations.

Why it matters · This study demonstrates a practical application of AI, specifically large language models, in a critical area of education: evaluating teaching quality. It offers insights into how AI can potentially assist in professional development for teachers while also highlighting the current limitations of AI in complex human-centric assessments.

About this work · This research contributes to the growing field of applying Artificial Intelligence, especially Large Language Models, to automate and enhance qualitative assessments within educational settings. It explores the feasibility and challenges of using AI to support educational measurement and teacher professional development.

Classroom observationsTeacher professional developmentLarge Language ModelsEducational assessmentEarly childhood education