Contrastive-Difference CKA Reveals Concept-Specific Structural Alignment Across Language Model Architectures

📅 2026-06-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Do different large language model architectures encode high-level concepts in a structurally compatible manner? Existing methods struggle to disentangle generic representational similarity from concept-specific alignment. This work proposes Contrastive Differential Centered Kernel Alignment (CKA_Delta), a training-free approach that isolates concept-specific structural convergence signals by leveraging sample-level contrastive differences. Applying CKA_Delta reveals, for the first time, a decoupling between geometric convergence and functional transfer: across six conceptual domains, models exhibit moderate geometric convergence alongside near-perfect functional transfer. CKA_Delta substantially outperforms standard CKA in identifying cross-architecture concept alignment and successfully flags outlier models such as Gemma (AUC = 0.79), offering a novel diagnostic tool for model analysis.
📝 Abstract
Do different LLM architectures encode high-level concepts in structurally compatible ways? We systematically characterize a geometric-functional universality dissociation: across multiple concept domains and architectural families, moderate geometric convergence coexists with near-perfect functional transfer. Using contrastive-difference CKA (CKA_Delta), a training-free diagnostic that computes kernel alignment on per-sample contrastive differences, we isolate concept-specific convergence from generic similarity -- achieving significant discrimination where standard CKA cannot. The dissociation replicates across all six concept domains we test (five with p <= 0.017 geometric discrimination and safety as a converging-functional trend, p = 0.08), including two non-instruction concepts (code-vs-NL, reasoning-vs-recall) validated without system prompts; a single 70B--70B pair provides an observational note that universality may strengthen with scale, requiring replication with additional >=70B models. We position CKA_Delta as a practical regime classifier and architectural outlier detector (Gemma: d = 1.08, AUC = 0.79) rather than an absolute transfer-accuracy predictor, providing a training-free diagnostic for cross-architecture concept monitoring.
Problem

Research questions and friction points this paper is trying to address.

structural alignment
language model architectures
concept encoding
geometric-functional universality
cross-architecture comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive-Difference CKA
structural alignment
geometric-functional dissociation
cross-architecture analysis
concept-specific representation
🔎 Similar Papers