When Cosine Similarity Fails to Reflect Linearly Accessible Structure in Dialogue Models

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了在对话模型中,余弦相似度未能反映线性可解码的人物结构的问题,并通过对比线性探针AUC和余弦kNN的表现来说明此问题。
📝 Abstract
Cosine similarity is widely used to analyze transformer representations, implicitly assuming that similarity reflects task-relevant structure. We study when this assumption fails in dialogue-conditioned large language models. Across three 7-8B chat-tuned models, ambient cosine similarity substantially underestimates linearly decodable persona structure on the same hidden states; numerically, linear probe AUC is in the 0.73-0.97 range while cosine kNN is in the 0.56-0.77 range on a 30-class task. A low-dimensional supervised subspace recovers much of this gap, whereas a matched-rank PCA subspace does not and in some cases degrades performance. This mismatch is regime-dependent: it is absent in single-sentence sentiment classification (SST-5), and a matched-cardinality control rules out attribute cardinality as a confound. The gap does not systematically increase across dialogue turns, and the task-aligned subspace remains stable over time. However, two of three models violate a pre-registered within-subspace separability invariance criterion (|Delta AUC| <= 0.03), and one model violates a pre-registered turn-invariance criterion (|Delta L| <= 0.05). These results show that cosine similarity can fail to reflect task-aligned structure in dialogue representations even when that structure is linearly accessible.
Problem

Research questions and friction points this paper is trying to address.

Cosine Similarity
Dialogue Models
Linearly Decodable Structure
Transformer Representations
Task-Relevant Structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cosine Similarity
Linear Decodability
Dialogue Models
Supervised Subspace
PCA Subspace
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.