domain feature alignment

Designs and implements methods and modules that align intermediate feature representations between source and target domains—using feature-distribution matching, layer-wise or multi-scale alignment modules, and alignment losses—to reduce cross-domain representation mismatch. Also analyzes feature-space discrepancies and measures alignment effectiveness to improve model robustness and generalization under domain shift.

domainfeaturealignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the unclear interaction between feature alignment and target fitting in cross-modal fine-tuning, which often leads to a mismatch between feature-label structures across source and target domains, thereby degrading generalization. For the first time, this study theoretically characterizes their relationship by introducing the notion of “feature-label distortion,” and establishes a provable generalization bound on target error. Based on this analysis, a principle for joint optimization of alignment and fitting is derived. The resulting framework offers interpretable and actionable design guidelines for cross-modal fine-tuning. Extensive experiments demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmark datasets, confirming its effectiveness and broad applicability.

cross-modal fine-tuningfeature alignmentfeature-label distortion

Towards a Learning Theory of Representation Alignment

Feb 19, 2025
FI
Francesco Insulla
🏛️ Stanford University | Istituto Italiano di Tecnologia | Università di Genova | Massachusetts Institute of Technology

This paper addresses the empirical phenomenon of representation alignment—where model representations increasingly converge as scale grows—in large language models, establishing the first formal learning-theoretic analysis framework. Methodologically, it unifies alignment definitions across metric, probabilistic, and spectral perspectives; introduces a task-driven representation stitching mechanism; and establishes its necessary and sufficient condition in terms of kernel alignment. Theoretical contributions include: (i) an upper bound on the generalization error of stitched representations; (ii) a rigorous proof that kernel alignment strictly governs representation transferability; and (iii) the first falsifiable learning-theoretic foundation for representation convergence in large models. Experimentally, the work integrates kernel methods, spectral graph theory, and probabilistic modeling to empirically validate the theoretical predictions.

Analyzing AI models' representation alignmentExploring statistical model convergenceLinking stitching to kernel alignment

This work addresses three key challenges under distribution shift: performance estimation, attribution-based explanation, and model improvement. To this end, the authors propose the Entropic Projection Alignment (EPA) framework, which simultaneously matches critical moments and minimizes the KL divergence between the labeled source domain and the unlabeled target domain—without requiring explicit density ratio estimation. EPA yields closed-form importance weights and incorporates implicit variance control to enhance robustness. Comprehensive theoretical analysis and extensive experiments demonstrate that EPA significantly outperforms existing methods in estimation accuracy, interpretability, and downstream model performance, while offering computational efficiency and strong theoretical guarantees.

distribution shiftdomain adaptationfeature attribution

DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative Perception

Jul 24, 2025
CT
Chengchang Tian
🏛️ Southeast University | Washington State University

In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.

Address domain gaps from hardware diversity and deployment conditionsEnhance semantic feature quality for collaborative perception fusionMitigate temporal misalignment caused by transmission delays

Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs

Nov 11, 2024
MT
Megh Thakkar
🏛️ Chandar Research Lab | Mila - Quebec AI Institute | IBM Research | Polytechnique Montréal

To address the significant degradation in safety capabilities observed during the specialization of large language models for domain experts, this paper proposes MergeAlign—a novel method that pioneers the application of model merging to jointly optimize safety and domain utility. MergeAlign integrates domain-specific vectors with alignment vectors via an interpretable vector interpolation mechanism, enabling fine-grained trade-offs between safety and performance. Evaluated on Llama3-based fine-tuned variants for medicine and finance, the approach leverages parameter-efficient alignment, model similarity analysis, and contribution decomposition assessment. Experiments demonstrate a 32% improvement in harmful response interception rate, with zero degradation in domain task performance (±0.2% fluctuation). The core contribution lies in introducing the model merging paradigm into safety alignment, thereby unifying knowledge fidelity and robustness within a single modeling framework.

Balancing domain expertise and safety in specialized LLMsMerging domain and alignment vectors for safer LLMsPreventing harmful content generation in domain-expert models

Latest Papers

What's happening recently
View more

This work proposes a knowledge-distillation-free method for optimizing domain mixing in pretraining data to align the distribution of a base language model with that of a target model. Treating language models as points in log-likelihood space, the approach dynamically adjusts the mixing weights of data domains by minimizing the Kullback–Leibler (KL) divergence between the base and target models, thereby steering model updates toward the target distribution. Experiments based on the NanoGPT framework demonstrate that, compared to uniform sampling from the Pile dataset, the proposed method substantially reduces distributional discrepancy with the target model and yields downstream task performance markedly closer to that of the target. This study offers a novel, distribution-alignment-based perspective on data recipe design for language model pretraining.

distribution alignmentdomain mixturelanguage models

This work addresses performance degradation in domain adaptation caused by distributional shift by proposing a trust-aware domain adaptation framework. It introduces, for the first time, a sample-level trust mechanism into the joint alignment of feature and prediction spaces. The approach constructs dual trust indicators based on predictive entropy—quantifying uncertainty—and prototype similarity—measuring semantic consistency—to weight the Joint Feature-Prediction Discrepancy (JFPD) metric. This weighting strategy steers the model to focus on target samples that are both reliable and semantically consistent. Extensive experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches on standard benchmarks. Moreover, the estimated domain discrepancy exhibits a strong correlation with target-domain error, confirming the method’s effectiveness and robustness.

distribution shiftdomain adaptationfeature-prediction discrepancy

This work addresses the challenge of distributional discrepancy between source and target domains in unsupervised tensor domain adaptation by proposing a novel method that jointly optimizes alignment matrices and a shared invariant subspace. By imposing constraints on the more flexible oblique manifold—rather than the conventional Stiefel manifold—and incorporating a variance-preserving regularizer to enhance robustness, the proposed framework generalizes existing tensor alignment approaches while significantly improving both domain adaptation efficiency and classification accuracy. Extensive experiments demonstrate that the method consistently outperforms state-of-the-art techniques across multiple benchmark datasets.

cross-domaindomain adaptationinvariant subspace

Existing representation alignment methods are largely confined to pairwise model alignment, limiting their scalability to multi-model settings and their ability to establish a consistent global reference space. This work proposes Geometrically Corrected Procrustes Alignment (GCPA), which leverages generalized Procrustes analysis to construct a shared orthogonal space while incorporating a posterior directional correction mechanism. This approach enables efficient and consistent multi-way alignment without compromising the intrinsic geometric structure of individual models. Experimental results demonstrate that GCPA consistently improves performance across cross-model retrieval tasks between any pair of models and maintains a practical, unified shared reference space.

multi-way alignmentneural network representationsPlatonic Representation Hypothesis

This work addresses the limitations of conventional model fusion approaches—such as parameter averaging—which often introduce non-generalizable features when source models differ substantially and lack theoretical grounding or effective metrics for assessing fusion compatibility. The study establishes, for the first time, a theoretical connection between model fusion and ensemble learning, and introduces M-Loss, a layer- or neuron-level inconsistency metric computed using only a small amount of unlabeled data to quantify the discrepancy between parameter averaging and model ensembling. This metric effectively guides the optimization of fusion strategies and the evaluation of parameter importance, thereby enhancing pruning efficiency. Experiments demonstrate that M-Loss significantly improves alignment between fused and ensemble models, enabling efficient and accurate model integration while reducing inference cost and storage overhead.

compatibility evaluationlimited unlabeled datamerging metric

Hot Scholars

DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
JQ

Jing Qin

University of Southern Denmark
MathematicsStatistics
SH

Shengfeng He

Singapore Management University
Visual ComputingGenerative ModelsComputer VisionComputational Photography
JH

Junjun He

Shanghai Jiao Tong University
DD

Donglin Di

Li Auto Inc.
Generative ModelsEmbodied AIMedical ImageMultimedia