🤖 AI Summary
It remains unclear whether phoneme-level language models implicitly acquire word boundary representations without explicit word-segmentation supervision. Method: We frame word segmentation as a phonological probing task, evaluating child-directed phoneme-level speech language models across 31 languages. We introduce a novel unsupervised approach that localizes word onsets via peaks in prediction error, complemented by linear probing to assess implicit encoding of unseen word boundaries. Contribution/Results: Experiments demonstrate that these models spontaneously develop cross-linguistic sensitivity to word boundaries through statistical learning alone, providing empirical support for the universality of statistical learning theory in multilingual speech acquisition. The findings offer new evidence and a principled training paradigm for designing subword tokenizers—particularly those operating at the phonemic level—by revealing how distributional regularities in speech input suffice for emergent word boundary detection without lexical supervision.
📝 Abstract
Language models provide a key framework for studying linguistic theories based on prediction, but phonological analysis using large language models (LLMs) is difficult; there are few phonological benchmarks beyond English and the standard input representation used in LLMs (subwords of graphemes) is not suitable for analyzing the representation of phonemes. In this work, we demonstrate how word segmentation can be used as a phonological probing task, allowing us to study the representations learned by phoneme-based language models trained on child-directed speech across 31 languages. Following computational models of word segmentation, we present unsupervised methods for extracting word boundaries from a trained model using the observation that prediction-error peaks at the start of words. We also use linear probes to identify that these models implicitly track word boundaries, even when they do not appear in training. This cross-lingual work corroborates statistical learning theories of acquisition and empirically motivates new methods for training subword tokenizers.