Analysing the Linearity of Linguistic Relations in Language Model Embedding Spaces

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种框架,通过约束线性近似方法分析不同语言关系在语言模型嵌入空间中的线性编码强度,并发现形态和派生关系几乎完美线性编码,而词典和百科关系错误率较高。
📝 Abstract
We propose a framework to analyse how strongly different linguistic relations are linearly encoded in language model embedding spaces. We formalise linear encoding via a constrained linear approximation over related and unrelated word pairs and apply this to an extended BATS dataset covering inflectional, derivational, lexicographic, and encyclopedic relations in GloVe, RoBERTa, and ModernBERT. Our experiments show near-perfect linear encodings for inflectional and derivational relations, but substantially higher errors for lexicographic and encyclopedic relations, especially for one-to-many and many-to-many associations. We also find that RoBERTa and ModernBERT generally encode relations more linearly than GloVe. These results indicate that our framework can reveal which relational structures are most linearly accessible in embeddings, offering a compact tool for probing and comparing relational geometry across models.
Problem

Research questions and friction points this paper is trying to address.

Linguistic Relations
Linear Encoding
Embedding Spaces
BATS Dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

linear encoding
linguistic relations
embedding spaces
constrained linear approximation
relational geometry
🔎 Similar Papers
V
Vasudevan Nedumpozhimana
ADAPT Research Centre, Trinity College Dublin, Ireland
F
Fathima Thekkekara
Indian Institute of Technology Bombay, Mumbai, India
J
John Kelleher
ADAPT Research Centre, Trinity College Dublin, Ireland