How Do Language Models Represent and Use Phonological Information for Allomorph Selection?

📅 2026-09-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨语言模型如何表示和使用音韵信息进行异形词选择,通过分析英语不定冠词a/an的选择机制,并验证该规则是否适用于其他语言。
📝 Abstract
Language models are trained on tokenized text that obscures the sound structure of words, yet they reliably produce morphemes whose form is phonologically conditioned. It remains unclear whether they rely on item-specific memorization or rule-like generalization and, if the latter, how that generalization is implemented. We therefore ask whether this phonological condition is represented within language models and how it is causally used for allomorph selection. For the English indefinite article a/an, we show that the phonological condition is encoded along a single linear direction in trigger-token embeddings, that this direction causally drives article selection in token-level wug tests, and that, at the article-prediction position, the model forecasts the upcoming trigger token and uses the forecasted trigger's phonological feature to choose the article. We then ask whether this rule-like generalization extends beyond English article selection, both to allomorph selection in other languages and to explicit phonological judgment. Together, these results provide a mechanistic account of phonologically conditioned allomorph selection in language models, and dissociate this generation-time ability from explicit metalinguistic judgments.
Problem

Research questions and friction points this paper is trying to address.

phonological information
allomorph selection
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

phonological condition
linear direction in embeddings
causal drive
token-level wug tests
forecasted trigger's phonological feature
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.