The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment
In plain terms
Large Language Models (LLMs) are increasingly being used to evaluate students, but it's not well understood how student demographic information, like education level or background, affects these AI tools. This paper addresses that problem, investigating whether considering demographics might improve things (e.g., tailoring feedback) or lead to unfair outcomes (e.g., biased scoring). The researchers set up controlled experiments where they tested six advanced LLMs on three educational tasks: automated essay scoring, giving helpful feedback, and answering questions about language rules. They examined two scenarios: "explicit demographic effects," where details like education level were directly stated, and "implicit effects," where demographics were hinted at through conversation history. They found that LLMs are indeed sensitive to these demographic cues in both explicit and implicit cases, altering their scoring, feedback, and answers. For example, LLMs frequently adjusted the readability of feedback when a student's education level was explicitly mentioned. However, implicit cues sometimes led to unpredictable biases, such as responses from lower-education levels receiving lower sentiment scores in question-answering tasks. These results provide clear evidence that demographics influence how LLMs perform in educational assessment tasks.