🤖 AI Summary
This study investigates how learning biases in language models shape word order typological universals, focusing specifically on crossed-serial dependencies and the mechanisms through which limited working memory influences the learning preferences of stack-based language models. Methodologically, it employs a dual-scaling strategy across both data and models, integrating artificial language simulations with cross-series dependency testing to systematically evaluate the generalization capabilities of stack-based architectures under varying memory constraints. The findings demonstrate that standard stack-based models struggle to process complex dependencies, whereas memory-constrained variants achieve superior generalization performance. These results reveal that limited working memory can serve as an inductive basis for specific word order universals, thereby providing empirical support for theoretical hypotheses positing that cognitive constraints drive linguistic structure.
📝 Abstract
Some properties of languages, e.g., subject-object-verb (SOV) word order, are more prevalent than others among the thousands of attested natural languages (NLs). Such typological commonality is often attributed to learning biases. Computational simulations, recently with language models (LMs), have facilitated the exploration of this theory. In this paper, we extend existing analyses of the relationship between LMs' learning biases and typological commonality on both data and model sides, focusing on: (i) cross-serial dependencies, the upper limit of attested syntactic complexity, and (ii) stack-based LMs (SLMs), potentially facilitating learning of hierarchical patterns. We first evaluate generalization of SLMs on cross-serial dependencies across diverse artificial languages and confirm that they struggle with such constructions. However, SLMs with limited working memory generalize better suggesting a possible basis for such inductive bias and thus the typological commonality of some word order configurations.