The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出答案盆地表示假设,通过模型产生的答案概率分布来组织线性结构,解释概念相关线性结构的形成机制。
📝 Abstract
The Linear Representation Hypothesis associates high-level concepts with directions in language models, but it remains unclear how these concept-related linear structures are organized within the model. We propose the Answer-Basin Representation Hypothesis: the probability measure induced over answers by the model's continuation distribution organizes these linear structures, with its statistics represented along linear directions shared across questions. All continuations yielding the same answer form an answer basin, whose mass is their total probability. These basin masses define the pushforward probability measure over answers. We posit that concept-related linear structure emerges from differences in the answer measure rather than being determined by changes in concept labels. Experiments across models and tasks link concept-consistent effects and their reversals in probing and steering to the alignment between concept labels and the answer measure.
Problem

Research questions and friction points this paper is trying to address.

Answer-Basin Representation Hypothesis
Linear Structure
Language Models
Concept Labels
Probability Measure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Answer-Basin Representation Hypothesis
probability measure over answers
concept-related linear structure
🔎 Similar Papers
No similar papers found.