🤖 AI Summary
This study addresses a critical gap in language model research, which has predominantly focused on language selection while overlooking fundamental script-level knowledge, particularly regarding generation in non-target scripts. To bridge this gap, this work shifts the focus toward script recognition and adaptation capabilities by constructing a multilingual test set spanning multiple writing systems. Two complementary experimental paradigms—input adaptation and explicit instruction following—are designed to comparatively evaluate the script-processing proficiency of models across varying scales. The findings demonstrate that these models possess substantial script knowledge, achieving over 98% fidelity in Latin scripts. Moreover, larger models significantly outperform smaller counterparts under non-standard script combinations. By systematically investigating graphical symbol recognition capabilities, this research fills a notable void in the existing literature on the orthographic competencies of large language models.
📝 Abstract
Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.