Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing LLM activation analyses constrained by linear assumptions, which fail to capture internal nonlinear concept manifolds. To overcome this bottleneck, we propose Non-Linear Multidimensional Concept Discovery (NLMCD) to model low-dimensional manifolds and design a Concept-Based Alignment score (CBA) using the generalized Rand index to quantify geometric structural similarity. Our analysis reveals hierarchical block structures obscured by conventional metrics and identifies concept evolution trajectories from syntax-dominated to semantically mixed representations. Furthermore, we systematically quantify alignment discrepancies across languages, models, and training stages. This work establishes a novel nonlinear analytical paradigm for mechanistic interpretability in large language models.
📝 Abstract
Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer-by-layer alignment matrices reveal two block structures in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic-semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal, strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure, with strong correspondence between same-family Qwen models of different scale but weak alignment across model families; and (vi) across Tulu-3 training stages, alignment is highest between adjacent stages, with the largest shift between the base model and SFT, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO's concepts in late layers.
Problem

Research questions and friction points this paper is trying to address.

mechanistic interpretability
non-linear concept manifolds
large language models
concept alignment
internal representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Non-Linear Concept Manifolds
Mechanistic Interpretability
Concept-Based Alignment
NLMCD
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tido Specht
University of Groningen
E
Elias Benedict Krey
Carl von Ossietzky University of Oldenburg
N
Nils Neukirch
Carl von Ossietzky University of Oldenburg
Nils Strodthoff
Nils Strodthoff
Professor for eHealth/AI4Health, Oldenburg University, Germany
Machine LearningDeep LearningBiomedical Data Analysis