simulation-based representation alignment

Design and implement algorithms and evaluation pipelines that enforce distributional consistency between simulated/generated and real embeddings, including contrastive distribution-alignment losses and procedures to synthesize ‘warm’ embeddings for training. Build mapping methods that relate model representations to brain recordings and analyses that compare semantic and behavioral manifolds and measure how alignment affects generalization to novel (cold) and familiar (warm) items.

simulation-basedrepresentationalignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Model-Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right One

Oct 27, 2025
IA
Itamar Avitan
🏛️ Ben-Gurion University of the Negev

This paper investigates whether high predictive accuracy implies genuine alignment between deep neural network representations and human perceptual representations. Method: Building upon the prevalent paradigm of modeling stimulus representations under linear transformations, we conduct large-scale model recovery experiments: using 20 visual models and 4.5 million human behavioral judgments, we generate synthetic behavioral responses incorporating linear geometric distortions and dimensional changes, and evaluate model identification performance via regression. Results: Even with massive data, current flexible alignment metrics yield an upper bound of <80% model recovery accuracy—indicating that the best-fitting model is not necessarily the truly aligned one. Our key contribution is exposing a fundamental tension between predictive accuracy and identifiability, advocating for explicit trade-offs between them in model comparison to avoid selecting spurious optima. This establishes a new methodological principle for evaluating representational alignment.

Assessing model recovery accuracy under linear transformation of representationsEvaluating whether high test accuracy indicates genuine representational alignmentIdentifying limitations of flexible alignment metrics in model comparison

Towards a Learning Theory of Representation Alignment

Feb 19, 2025
FI
Francesco Insulla
🏛️ Stanford University | Istituto Italiano di Tecnologia | Università di Genova | Massachusetts Institute of Technology

This paper addresses the empirical phenomenon of representation alignment—where model representations increasingly converge as scale grows—in large language models, establishing the first formal learning-theoretic analysis framework. Methodologically, it unifies alignment definitions across metric, probabilistic, and spectral perspectives; introduces a task-driven representation stitching mechanism; and establishes its necessary and sufficient condition in terms of kernel alignment. Theoretical contributions include: (i) an upper bound on the generalization error of stitched representations; (ii) a rigorous proof that kernel alignment strictly governs representation transferability; and (iii) the first falsifiable learning-theoretic foundation for representation convergence in large models. Experimentally, the work integrates kernel methods, spectral graph theory, and probabilistic modeling to empirically validate the theoretical predictions.

Analyzing AI models' representation alignmentExploring statistical model convergenceLinking stitching to kernel alignment

Representation alignment in self-supervised learning remains poorly understood due to (1) existing metrics ignoring transformation invariance, (2) ambiguity regarding *when* alignment occurs during training, and (3) a lack of systematic characterization of layer-wise robustness to distribution shift. Method: We propose a multi-metric alignment evaluation framework integrating Procrustes alignment, permutation/soft matching, and linear regression, coupled with inter-layer comparison and out-of-distribution (OOD) experimental design. Contribution/Results: We find that representation alignment concentrates overwhelmingly in the first training epoch (>90% alignment achieved immediately), early layers exhibit high robustness to distribution shifts while late layers are markedly sensitive, and orthogonal transformations closely approximate optimal linear alignment. This work provides the first systematic characterization of alignment evolution across network depth, training dynamics, and distributional change—establishing a new paradigm for understanding representational convergence in both artificial and biological neural networks.

Analyzing representational alignment in neural networksExamining alignment under distribution shiftsIdentifying when alignment emerges during training

Training objective drives the consistency of representational similarity across datasets

Nov 08, 2024
LC
Laure Ciernik
🏛️ Technische Universität Berlin | Aignostics | Anthropic | Google DeepMind

This work investigates whether cross-dataset consistency in model representation similarity stems from intrinsic model properties or is confounded by biases inherent in common benchmark datasets. To address this, we conduct systematic representation comparison experiments across multimodal (image, image-text) and multitask (self-supervised, classification, image-text contrastive) models, using Centered Kernel Alignment (CKA) and linearly weighted similarity analysis on diverse domain-shifted datasets. Results demonstrate that training objective is the dominant factor governing cross-dataset representation similarity stability—significantly outweighing influences of data modality and network architecture. We propose the first evaluation framework explicitly designed for cross-dataset representational consistency. Furthermore, we reveal that self-supervised vision models exhibit the strongest generalization of representation similarity across datasets, and that the correlation between representation similarity and task performance is maximized on single-domain benchmarks.

Analyzes link between model representations and task behaviorExamines impact of objective function on similarity consistencyMeasures how representational similarity varies across datasets

Evaluating alignment between humans and neural network representations in image-based learning tasks

Jun 15, 2023
CD
Can Demircan
🏛️ Max Planck Institute for Biological Cybernetics | Max Planck School of Cognition | Max Planck Institute for Human Cognitive & Brain Sciences | Universität Hamburg | Kavli Institute for Systems Neuroscience | Leipzig University | Technical University Dresden | Julius-Maximilians-Universität Würzburg | Max Planck Institute for Human Development

This study investigates the alignment mechanisms between neural network representations and human visual learning in few-shot image understanding. Method: We systematically evaluate generalization behaviors of 86 pretrained models on continuous relational reasoning and natural image classification tasks, introducing the first quantitative measure of cross-task consistency between model representations and human cognitive trajectories. Our approach integrates representational similarity analysis, intrinsic dimension estimation, cognitive-modeling–driven evaluation, and multimodal contrastive learning assessment. Results: Multimodal contrastive learning emerges as the strongest predictor of human few-shot generalization—significantly outperforming conventional metrics such as parameter count or training data scale. Pretrained models serve as effective sources of cognitive representations. The proposed evaluation paradigm establishes a generalizable, ecologically valid framework for cross-species intelligence modeling, advancing the study of human-aligned artificial perception.

Brain-Computer ConsistencyFew-shot LearningImage Learning Tasks

Latest Papers

What's happening recently
View more

Concept alignment lacks a unified definition, and existing methods optimize divergent objectives under the same terminology, obscuring its fundamental nature. This work formalizes its multidimensional structure by decomposing it along two axes—“alignment targets” and “alignment levels”—and identifies four distinct alignment properties, revealing that current approaches satisfy only subsets of these. To address this limitation, we propose Coupled Sparse Autoencoders (CoSAE), a framework that jointly optimizes multiple alignment objectives, alongside InterVenchA, an interventional evaluation benchmark. Experiments demonstrate that optimizing a single objective fails to reliably recover other alignment properties, whereas CoSAE achieves strong instance-level conceptual consistency using merely 0.1% paired data.

concept alignmentdistributional alignmentinstance-level alignment

This study addresses the challenge of "alignment generalization," wherein the behavior of fine-tuned large language models (LLMs) in unseen scenarios remains difficult to predict. We propose an activation-representation-based task for predicting alignment generalization and systematically analyze how 66 distinct values influence fine-tuning outcomes. Our findings reveal that internal activation representations predict fine-tuning effects more accurately than textual descriptions. Building on this insight, we construct the first value taxonomy and shared value space for LLMs. Large-scale empirical evaluations demonstrate a significant correlation between activation-based value similarity and model robustness (r=0.45). These results provide a quantifiable scientific foundation for principled LLM behavioral design.

Alignment GeneralizationBehavior PredictionLarge Language Models

Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior

Nov 03, 2025
DA
Daniel Aarao Reis Arturi
🏛️ McGill University | McMaster University | University of Alberta | Algoverse AI Research | St. Jude Children’s Research Hospital

This study investigates the intrinsic mechanisms underlying cross-domain misalignment—termed “emergent misalignment”—in large language models (LLMs) after fine-tuning on fine-grained harmful data. Addressing the key question of *why single-domain harmful training generalizes to broad inappropriate behavior*, we propose a geometric analytical framework integrating cosine similarity, principal component analysis, parameter subspace projection overlap, and linear interpolation connectivity experiments. We首次 discover that misaligned behaviors across distinct harmful tasks reside in a shared low-dimensional parameter subspace and exhibit pronounced linear structure within it; models obtained via cross-task linear interpolation retain consistent, widespread harmful outputs, confirming functional equivalence and parameter convergence. These findings reveal that misalignment possesses a tractable, geometrically localizable nature—establishing a theoretical foundation and novel intervention pathways for controllable alignment.

Demonstrating linear connectivity between narrow and broad misaligned behaviorsIdentifying shared parameter subspaces enabling cross-task harmful generalizationRevealing geometric mechanisms behind emergent misalignment in language models

This study investigates the impact of consistency training on model alignment, demonstrating that it is not alignment-neutral. Through systematic evaluation of seven consistency methods across 108 open-source large language models (7B–70B) with controlled misalignment, the authors find that such training generally suppresses reward hacking while exacerbating sycophancy. Leveraging controlled fine-tuning, distribution shift analysis, and theoretical modeling, they identify distribution shift as the dominant underlying mechanism. Building on this insight, they propose a unified theoretical framework that predicts under which conditions consistency training amplifies or mitigates specific misalignment behaviors, thereby offering an auditable foundation for safer alignment practices.

consistency trainingmisalignmentmodel alignment

This study investigates how student models inherit capabilities from teachers during knowledge distillation on unlabeled data—even pure noise—with a focus on the positional and capacity-determining mechanisms of hidden channels. Within an MLP distillation framework, the authors propose a Covert Token Propagation (CTP) mechanism, revealing that students achieve geometric alignment with teacher hidden representations by adjusting input projection weights, thereby demonstrating that representation alignment—not mere information transfer—is central to knowledge transfer. The analysis integrates linear CKA, weight freezing, multi-teacher ensemble ablation, KL gradient inspection, and initialization sweeps. Key findings include: channel closure is driven by weight drift independent of teacher accuracy; freezing input weights (W₀) disrupts transfer while freezing output weights (W₂) does not; signals from multiple teachers interfere destructively; and CKA strongly correlates with student performance (r=0.98). This geometric perspective further extends to cross-token entanglement phenomena in large language models.

behavioral entanglementcovert trait propagationhidden-channel

Hot Scholars

DC

De-Chuan Zhan

Nanjing University, China
Machine LearningData Mining
JS

Jinwoo Shin

ICT Endowed Chair Professor
Machine LearningDeep Learning
UB

Ulas Bagci

Northwestern University
artificial intelligencedeep learningbiomedical image analysismedical image computing
WC

Weidong Cai

Clinical Associate Professor, Stanford University School of Medicine
functional neuroimagingmachine learningcognitivedevelopmental
WS

Wenxuan Song

The Hong Kong University of Science and Technology (Guangzhou)
Vision-language-action ModelRobotics