Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of ambiguous layer correspondence and misaligned feature spaces among independently trained large language models. To overcome these limitations, we propose MATCHA, a framework that transcends fixed layer-pairing constraints by jointly optimizing explicit many-to-many layer mapping matrices and a shared feature transformation module, thereby achieving end-to-end alignment across model activation spaces. Evaluated on 42 model pairs, MATCHA significantly improves reconstruction fidelity and retrieval performance while revealing the monotonic many-to-many nature of inter-model layer mappings. Furthermore, the method offers strong interpretability and effectively facilitates the cross-model transfer of intervention vectors and probing classifiers.
📝 Abstract
LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them? We study this by learning an activation alignment, a map from a source model's layerwise activations to a target's. Our method, MATCHA, factors this map into a layer map, whose output is an explicit target-by-source matrix that can be extracted and inspected, and a layer-shared feature map between hidden spaces. Most of prior work fixes the layer correspondence in advance, pairing layers at roughly the same relative depth; in contrast, we learn both factors jointly from prompts. Across 42 pairs of seven models spanning three different families, MATCHA reconstructs the target's activations more faithfully and improves retrieval-based metrics substantially, w.r.t. previous approaches. The recovered maps are broadly monotone in depth but, in contrast with most previous approaches, are consistently many-to-many: each target layer draws on a band of source layers. Our alignments also enable transfer of activation-space interventions, allowing steering vectors and probes developed for one model to transfer to another.
Problem

Research questions and friction points this paper is trying to address.

Activation Alignment
Cross-Model Mapping
Large Language Models
Layer Correspondence
Feature Transformation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation Alignment
Many-to-Many Layer Map
Cross-Model Transfer
Feature Map
MATCHA
🔎 Similar Papers
No similar papers found.