Score
Designs and implements methods and modules that align intermediate feature representations between source and target domains—using feature-distribution matching, layer-wise or multi-scale alignment modules, and alignment losses—to reduce cross-domain representation mismatch. Also analyzes feature-space discrepancies and measures alignment effectiveness to improve model robustness and generalization under domain shift.
This work addresses the unclear interaction between feature alignment and target fitting in cross-modal fine-tuning, which often leads to a mismatch between feature-label structures across source and target domains, thereby degrading generalization. For the first time, this study theoretically characterizes their relationship by introducing the notion of “feature-label distortion,” and establishes a provable generalization bound on target error. Based on this analysis, a principle for joint optimization of alignment and fitting is derived. The resulting framework offers interpretable and actionable design guidelines for cross-modal fine-tuning. Extensive experiments demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmark datasets, confirming its effectiveness and broad applicability.
This paper addresses the empirical phenomenon of representation alignment—where model representations increasingly converge as scale grows—in large language models, establishing the first formal learning-theoretic analysis framework. Methodologically, it unifies alignment definitions across metric, probabilistic, and spectral perspectives; introduces a task-driven representation stitching mechanism; and establishes its necessary and sufficient condition in terms of kernel alignment. Theoretical contributions include: (i) an upper bound on the generalization error of stitched representations; (ii) a rigorous proof that kernel alignment strictly governs representation transferability; and (iii) the first falsifiable learning-theoretic foundation for representation convergence in large models. Experimentally, the work integrates kernel methods, spectral graph theory, and probabilistic modeling to empirically validate the theoretical predictions.
This work addresses three key challenges under distribution shift: performance estimation, attribution-based explanation, and model improvement. To this end, the authors propose the Entropic Projection Alignment (EPA) framework, which simultaneously matches critical moments and minimizes the KL divergence between the labeled source domain and the unlabeled target domain—without requiring explicit density ratio estimation. EPA yields closed-form importance weights and incorporates implicit variance control to enhance robustness. Comprehensive theoretical analysis and extensive experiments demonstrate that EPA significantly outperforms existing methods in estimation accuracy, interpretability, and downstream model performance, while offering computational efficiency and strong theoretical guarantees.
In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.
To address the significant degradation in safety capabilities observed during the specialization of large language models for domain experts, this paper proposes MergeAlign—a novel method that pioneers the application of model merging to jointly optimize safety and domain utility. MergeAlign integrates domain-specific vectors with alignment vectors via an interpretable vector interpolation mechanism, enabling fine-grained trade-offs between safety and performance. Evaluated on Llama3-based fine-tuned variants for medicine and finance, the approach leverages parameter-efficient alignment, model similarity analysis, and contribution decomposition assessment. Experiments demonstrate a 32% improvement in harmful response interception rate, with zero degradation in domain task performance (±0.2% fluctuation). The core contribution lies in introducing the model merging paradigm into safety alignment, thereby unifying knowledge fidelity and robustness within a single modeling framework.
This work proposes a knowledge-distillation-free method for optimizing domain mixing in pretraining data to align the distribution of a base language model with that of a target model. Treating language models as points in log-likelihood space, the approach dynamically adjusts the mixing weights of data domains by minimizing the Kullback–Leibler (KL) divergence between the base and target models, thereby steering model updates toward the target distribution. Experiments based on the NanoGPT framework demonstrate that, compared to uniform sampling from the Pile dataset, the proposed method substantially reduces distributional discrepancy with the target model and yields downstream task performance markedly closer to that of the target. This study offers a novel, distribution-alignment-based perspective on data recipe design for language model pretraining.
This work addresses performance degradation in domain adaptation caused by distributional shift by proposing a trust-aware domain adaptation framework. It introduces, for the first time, a sample-level trust mechanism into the joint alignment of feature and prediction spaces. The approach constructs dual trust indicators based on predictive entropy—quantifying uncertainty—and prototype similarity—measuring semantic consistency—to weight the Joint Feature-Prediction Discrepancy (JFPD) metric. This weighting strategy steers the model to focus on target samples that are both reliable and semantically consistent. Extensive experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches on standard benchmarks. Moreover, the estimated domain discrepancy exhibits a strong correlation with target-domain error, confirming the method’s effectiveness and robustness.
This work addresses the challenge of distributional discrepancy between source and target domains in unsupervised tensor domain adaptation by proposing a novel method that jointly optimizes alignment matrices and a shared invariant subspace. By imposing constraints on the more flexible oblique manifold—rather than the conventional Stiefel manifold—and incorporating a variance-preserving regularizer to enhance robustness, the proposed framework generalizes existing tensor alignment approaches while significantly improving both domain adaptation efficiency and classification accuracy. Extensive experiments demonstrate that the method consistently outperforms state-of-the-art techniques across multiple benchmark datasets.
Existing representation alignment methods are largely confined to pairwise model alignment, limiting their scalability to multi-model settings and their ability to establish a consistent global reference space. This work proposes Geometrically Corrected Procrustes Alignment (GCPA), which leverages generalized Procrustes analysis to construct a shared orthogonal space while incorporating a posterior directional correction mechanism. This approach enables efficient and consistent multi-way alignment without compromising the intrinsic geometric structure of individual models. Experimental results demonstrate that GCPA consistently improves performance across cross-model retrieval tasks between any pair of models and maintains a practical, unified shared reference space.
This work addresses the limitations of conventional model fusion approaches—such as parameter averaging—which often introduce non-generalizable features when source models differ substantially and lack theoretical grounding or effective metrics for assessing fusion compatibility. The study establishes, for the first time, a theoretical connection between model fusion and ensemble learning, and introduces M-Loss, a layer- or neuron-level inconsistency metric computed using only a small amount of unlabeled data to quantify the discrepancy between parameter averaging and model ensembling. This metric effectively guides the optimization of fusion strategies and the evaluation of parameter importance, thereby enhancing pruning efficiency. Experiments demonstrate that M-Loss significantly improves alignment between fused and ensemble models, enabling efficient and accurate model integration while reducing inference cost and storage overhead.