Towards Scalable Measurement of Durable Skills
In plain terms
Durable skills, such as teamwork and critical thinking, are vital for success in today's workforce but are challenging to measure, leading to them often being overlooked in educational curricula. Effective assessments for these skills need to feel like real-world interactions (ecological validity) while also being standardized and repeatable (psychometric rigor). This paper introduces a framework that uses Large Language Models (LLMs), which are AI systems capable of understanding and generating human-like text, to achieve both goals. In this setup, a person interacts with AI teammates in a conversation designed to mimic human interaction. An 'Executive LLM' strategically guides the discussion to elicit specific evidence of skill proficiency, and a separate 'AI evaluator' then measures these skills from the conversation. The research found that this method significantly increases the observable evidence of skills, and that AI-automated scoring largely aligns with assessments from human experts.