Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
In plain terms
Large language models (LLMs) are increasingly used as educational tutors, but their ability to solve tasks doesn't automatically mean they are good at teaching. This paper addresses the problem of evaluating LLM tutors by distinguishing between their task-solving ability and their ability to support student learning (pedagogy). The researchers developed a simple diagnostic method that looks at the difference in performance on solving-oriented versus pedagogy-oriented benchmarks. They found that for eight public LLMs, these two aspects are only partially aligned, with a low correlation of 0.421, meaning models can shift ranks when evaluated for teaching quality. They also analyzed existing benchmarks, showing that good teaching behaviors like guiding questions and hints are already valued in rubrics. These findings suggest that evaluating LLM tutors should separately measure task success and learning support to better assess their true educational impact.