Multilingual Automated Scoring in International Large-Scale Assessments: Scaling and Quality Control via LLM-Based Workflows
Authors: Matthias von Davier, Ummugul Bezirhan, Ji-Yoon Jung
Abstract
This presentation details the implementation of an AI-based automated scoring (AS) pipeline to promote and compare LLM-based scoring with human scoring in PIRLS 2026, an international reading assessment conducted in almost 60 countries and at least as many language versions. As PIRLS transitions to a fully computer-based format, leveraging Large Language Models (LLMs) provides a scalable solution for scoring constructed-response (CR) items, addressing operational burdens while improving consistency across diverse linguistic contexts.
The proposed system integrates a rigorous, systematic pipeline including secure data preparation with automated PII masking, optimized prompt engineering based on content-expert validated scoring templates, and a post-processing framework designed to detect hallucinations and ensure internal consistency. To evaluate the scoring reliability of this automated approach, we introduce the Linguistic-integrated Reliability Audit (LiRA), a novel, semantic-similarity-based framework that replaces traditional human double-scoring with data-driven benchmark scores.
Results from field tests demonstrate that the AS pipeline achieves high-level alignment with human scoring, yielding 96.1% exact agreement and robust cross-country consistency. Furthermore, psychometric analyses comparing EAP estimates show near-perfect overlap between human and AI-scored conditions, validating the reliability of this modernization effort. This talk will outline the full end-to-end process, discuss the efficiency gains, and highlight the implications for the future of large-scale educational assessments.