Divergent large language model predictions from convergent representations in ambiguous word pairs

📅 2026-08-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the disambiguation mechanisms of large language models in contexts of lexical ambiguity, revealing a dissociation between their output behavior and internal representational similarity. Through layer-wise analysis of decoder-only Transformer models across three scales—employing inter-layer representation probing, KL divergence measurements, activation patching, and single-layer ablation experiments—the authors find that semantic differentiation is most pronounced in intermediate layers. Although late-layer representations exhibit converging cosine similarity in embedding space, they nonetheless causally support correct predictions. These findings raise critical concerns for approaches that rely on late-layer embeddings for semantic search or clustering, suggesting such methods may overlook crucial disambiguation signals present earlier in the network.
📝 Abstract
In this work we investigate how decoder-only transformers resolve lexical ambiguity through layer-by-layer analysis of three models spanning three parameter sizes (GPT-2-Small-117M, Llama-3.2-3B, Qwen2.5-32B). For both homonyms and polysemes, we find that representations become maximally distinct in middle layers, then partially reconverge in late layers, while the KL divergence between their next-token predictions reaches its maximum in the final layers. The activation patching experiment provides causal evidence that late-layer representational differences directly determine outputs despite apparent increased similarity in embedding space. Our single-layer ablation experiment indicates that models achieve equivalent disambiguation despite qualitatively different layer-wise vulnerabilities. These findings offer a mechanism for recent observations where models' internal embedding similarities show low correlation with their behavioural outputs despite strong performance. The semantic distinctions therefore remain present but become increasingly invisible to similarity measures over the embeddings, with implications for embedding-based methods such as semantic search, retrieval, and clustering that rely on late-layer cosine similarity.
Problem

Research questions and friction points this paper is trying to address.

lexical ambiguity
representation convergence
embedding similarity
semantic search
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

lexical ambiguity
layer-wise analysis
representation convergence
activation patching
embedding similarity