Matthias von Davier

Authors: Matthias von Davier, Ummugul Bezirhan, Ji-Yoon Jung

Abstract

This presentation details the implementation of an AI-based automated scoring (AS) pipeline to promote and compare LLM-based scoring with human scoring in PIRLS 2026, an international reading assessment conducted in almost 60 countries and at least as many language versions. As PIRLS transitions to a fully computer-based format, leveraging Large Language Models (LLMs) provides a scalable solution for scoring constructed-response (CR) items, addressing operational burdens while improving consistency across diverse linguistic contexts.

The proposed system integrates a rigorous, systematic pipeline including secure data preparation with automated PII masking, optimized prompt engineering based on content-expert validated scoring templates, and a post-processing framework designed to detect hallucinations and ensure internal consistency. To evaluate the scoring reliability of this automated approach, we introduce the Linguistic-integrated Reliability Audit (LiRA), a novel, semantic-similarity-based framework that replaces traditional human double-scoring with data-driven benchmark scores.

Results from field tests demonstrate that the AS pipeline achieves high-level alignment with human scoring, yielding 96.1% exact agreement and robust cross-country consistency. Furthermore, psychometric analyses comparing EAP estimates show near-perfect overlap between human and AI-scored conditions, validating the reliability of this modernization effort. This talk will outline the full end-to-end process, discuss the efficiency gains, and highlight the implications for the future of large-scale educational assessments.