Score
Design and implement algorithms and evaluation pipelines that enforce distributional consistency between simulated/generated and real embeddings, including contrastive distribution-alignment losses and procedures to synthesize ‘warm’ embeddings for training. Build mapping methods that relate model representations to brain recordings and analyses that compare semantic and behavioral manifolds and measure how alignment affects generalization to novel (cold) and familiar (warm) items.
This paper systematically investigates the quantification of neural network model similarity, addressing both representational similarity (intermediate-layer activations) and functional similarity (output behavior). It unifies and comparatively analyzes mainstream metrics—such as CKA, SVCCA, PWCCA, linear probes, and top-k output agreement—across these two complementary paradigms. A structured taxonomy is introduced, accompanied by theoretical analysis of each metric’s mathematical properties, interrelationships, and applicability boundaries. Empirical evaluation assesses their explanatory power and limitations in downstream tasks including model compression, ensemble learning, and robustness analysis. The core contribution is a cross-paradigm benchmark enabling rigorous metric comparison, revealing how metric selection critically influences downstream conclusions. The work further identifies open challenges and proposes principled evaluation criteria, thereby advancing methodological foundations for model behavior interpretation and trustworthy AI. (149 words)
This paper investigates whether high predictive accuracy implies genuine alignment between deep neural network representations and human perceptual representations. Method: Building upon the prevalent paradigm of modeling stimulus representations under linear transformations, we conduct large-scale model recovery experiments: using 20 visual models and 4.5 million human behavioral judgments, we generate synthetic behavioral responses incorporating linear geometric distortions and dimensional changes, and evaluate model identification performance via regression. Results: Even with massive data, current flexible alignment metrics yield an upper bound of <80% model recovery accuracy—indicating that the best-fitting model is not necessarily the truly aligned one. Our key contribution is exposing a fundamental tension between predictive accuracy and identifiability, advocating for explicit trade-offs between them in model comparison to avoid selecting spurious optima. This establishes a new methodological principle for evaluating representational alignment.
This paper addresses the empirical phenomenon of representation alignment—where model representations increasingly converge as scale grows—in large language models, establishing the first formal learning-theoretic analysis framework. Methodologically, it unifies alignment definitions across metric, probabilistic, and spectral perspectives; introduces a task-driven representation stitching mechanism; and establishes its necessary and sufficient condition in terms of kernel alignment. Theoretical contributions include: (i) an upper bound on the generalization error of stitched representations; (ii) a rigorous proof that kernel alignment strictly governs representation transferability; and (iii) the first falsifiable learning-theoretic foundation for representation convergence in large models. Experimentally, the work integrates kernel methods, spectral graph theory, and probabilistic modeling to empirically validate the theoretical predictions.
Representation alignment in self-supervised learning remains poorly understood due to (1) existing metrics ignoring transformation invariance, (2) ambiguity regarding *when* alignment occurs during training, and (3) a lack of systematic characterization of layer-wise robustness to distribution shift. Method: We propose a multi-metric alignment evaluation framework integrating Procrustes alignment, permutation/soft matching, and linear regression, coupled with inter-layer comparison and out-of-distribution (OOD) experimental design. Contribution/Results: We find that representation alignment concentrates overwhelmingly in the first training epoch (>90% alignment achieved immediately), early layers exhibit high robustness to distribution shifts while late layers are markedly sensitive, and orthogonal transformations closely approximate optimal linear alignment. This work provides the first systematic characterization of alignment evolution across network depth, training dynamics, and distributional change—establishing a new paradigm for understanding representational convergence in both artificial and biological neural networks.
This work investigates whether cross-dataset consistency in model representation similarity stems from intrinsic model properties or is confounded by biases inherent in common benchmark datasets. To address this, we conduct systematic representation comparison experiments across multimodal (image, image-text) and multitask (self-supervised, classification, image-text contrastive) models, using Centered Kernel Alignment (CKA) and linearly weighted similarity analysis on diverse domain-shifted datasets. Results demonstrate that training objective is the dominant factor governing cross-dataset representation similarity stability—significantly outweighing influences of data modality and network architecture. We propose the first evaluation framework explicitly designed for cross-dataset representational consistency. Furthermore, we reveal that self-supervised vision models exhibit the strongest generalization of representation similarity across datasets, and that the correlation between representation similarity and task performance is maximized on single-domain benchmarks.
This study investigates the alignment mechanisms between neural network representations and human visual learning in few-shot image understanding. Method: We systematically evaluate generalization behaviors of 86 pretrained models on continuous relational reasoning and natural image classification tasks, introducing the first quantitative measure of cross-task consistency between model representations and human cognitive trajectories. Our approach integrates representational similarity analysis, intrinsic dimension estimation, cognitive-modeling–driven evaluation, and multimodal contrastive learning assessment. Results: Multimodal contrastive learning emerges as the strongest predictor of human few-shot generalization—significantly outperforming conventional metrics such as parameter count or training data scale. Pretrained models serve as effective sources of cognitive representations. The proposed evaluation paradigm establishes a generalizable, ecologically valid framework for cross-species intelligence modeling, advancing the study of human-aligned artificial perception.
Concept alignment lacks a unified definition, and existing methods optimize divergent objectives under the same terminology, obscuring its fundamental nature. This work formalizes its multidimensional structure by decomposing it along two axes—“alignment targets” and “alignment levels”—and identifies four distinct alignment properties, revealing that current approaches satisfy only subsets of these. To address this limitation, we propose Coupled Sparse Autoencoders (CoSAE), a framework that jointly optimizes multiple alignment objectives, alongside InterVenchA, an interventional evaluation benchmark. Experiments demonstrate that optimizing a single objective fails to reliably recover other alignment properties, whereas CoSAE achieves strong instance-level conceptual consistency using merely 0.1% paired data.
This study addresses the challenge of "alignment generalization," wherein the behavior of fine-tuned large language models (LLMs) in unseen scenarios remains difficult to predict. We propose an activation-representation-based task for predicting alignment generalization and systematically analyze how 66 distinct values influence fine-tuning outcomes. Our findings reveal that internal activation representations predict fine-tuning effects more accurately than textual descriptions. Building on this insight, we construct the first value taxonomy and shared value space for LLMs. Large-scale empirical evaluations demonstrate a significant correlation between activation-based value similarity and model robustness (r=0.45). These results provide a quantifiable scientific foundation for principled LLM behavioral design.
This study investigates the intrinsic mechanisms underlying cross-domain misalignment—termed “emergent misalignment”—in large language models (LLMs) after fine-tuning on fine-grained harmful data. Addressing the key question of *why single-domain harmful training generalizes to broad inappropriate behavior*, we propose a geometric analytical framework integrating cosine similarity, principal component analysis, parameter subspace projection overlap, and linear interpolation connectivity experiments. We首次 discover that misaligned behaviors across distinct harmful tasks reside in a shared low-dimensional parameter subspace and exhibit pronounced linear structure within it; models obtained via cross-task linear interpolation retain consistent, widespread harmful outputs, confirming functional equivalence and parameter convergence. These findings reveal that misalignment possesses a tractable, geometrically localizable nature—establishing a theoretical foundation and novel intervention pathways for controllable alignment.
This study investigates the impact of consistency training on model alignment, demonstrating that it is not alignment-neutral. Through systematic evaluation of seven consistency methods across 108 open-source large language models (7B–70B) with controlled misalignment, the authors find that such training generally suppresses reward hacking while exacerbating sycophancy. Leveraging controlled fine-tuning, distribution shift analysis, and theoretical modeling, they identify distribution shift as the dominant underlying mechanism. Building on this insight, they propose a unified theoretical framework that predicts under which conditions consistency training amplifies or mitigates specific misalignment behaviors, thereby offering an auditable foundation for safer alignment practices.
This study investigates how student models inherit capabilities from teachers during knowledge distillation on unlabeled data—even pure noise—with a focus on the positional and capacity-determining mechanisms of hidden channels. Within an MLP distillation framework, the authors propose a Covert Token Propagation (CTP) mechanism, revealing that students achieve geometric alignment with teacher hidden representations by adjusting input projection weights, thereby demonstrating that representation alignment—not mere information transfer—is central to knowledge transfer. The analysis integrates linear CKA, weight freezing, multi-teacher ensemble ablation, KL gradient inspection, and initialization sweeps. Key findings include: channel closure is driven by weight drift independent of teacher accuracy; freezing input weights (W₀) disrupts transfer while freezing output weights (W₂) does not; signals from multiple teachers interfere destructively; and CKA strongly correlates with student performance (r=0.98). This geometric perspective further extends to cross-token entanglement phenomena in large language models.