🤖 AI Summary
This study addresses the loss of student responses in multilingual assessment systems due to translation failures, which compromises reliability auditing. To mitigate this issue, the authors propose bypassing translation altogether by directly employing native-language multilingual sentence embeddings instead of translating responses into English. The approach is evaluated on 11 PIRLS constructed-response items using three state-of-the-art multilingual embedding models. Results demonstrate, for the first time, that native-language embeddings can reproduce translation-based reliability estimates without relying on machine translation, successfully recover responses previously excluded due to translation errors, and maintain comparable reliability metrics with no statistically significant degradation. This method enhances both the completeness and robustness of multilingual educational assessments.
📝 Abstract
Multilingual assessment systems commonly rely on translation for scoring and quality-control processes. We evaluate whether multilingual sentence embeddings can replace translated English input for Linguistic-Integrated Reliability Auditing (LiRA) across 11 PIRLS constructed-response items and three embedding models. Native-language embeddings reproduced translation-based reliability estimates closely while recovering responses excluded after translation failure, with no meaningful change in reliability.