Score
Designs, builds, or analyzes methods, transforms, and evaluation metrics that enforce or measure how a model’s outputs or internal representations match a target distribution or reference space across modalities, languages, coordinates, or data types. This includes creating alignment losses and scoring functions, distributional filtering and proposal-divergence measures, and latent-space/coordinate/positional transforms for embedding alignment (bilingual, multimodal, image, point-cloud, pose, etc.).
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
Embedding models trained independently on similar data capture stable semantic meanings but yield inconsistent representation spaces, hindering interoperability across models. This work addresses compatibility challenges in multimodal search and model upgrades via orthogonal transformation-based embedding alignment. Theoretically, we derive the first tight Procrustes alignment error bound, proving the existence of a near-isometric orthogonal transformation that approximately preserves pairwise inner products—establishing rigorous theoretical foundations for alignment. Methodologically, we employ efficient Procrustes analysis as a post-hoc alignment procedure, preserving the intrinsic geometric structure of each embedding space while enabling cross-model alignment. Experiments demonstrate substantial improvements in model retraining compatibility, text retrieval fusion accuracy, and cross-modal search performance; our method achieves state-of-the-art results in hybrid multimodal search.
Conventional dimensionality reduction (DR) evaluation suffers from systematic bias due to the frequent adoption of highly correlated metrics, leading to overemphasis on specific structural properties. Method: We propose an empirically grounded metric redundancy reduction framework: first computing Pearson correlation matrices across diverse datasets and DR algorithms; then applying clustering to identify functionally redundant metric groups; and finally retaining only the most representative metric per group—replacing subjective, intent-driven metric selection with objective, behavior-based clustering. Contribution/Results: Our approach significantly improves cross-dataset and cross-algorithm stability of DR evaluations, effectively mitigating structural biases inherent in traditional assessment protocols. Experimental validation demonstrates enhanced reproducibility and generalizability, establishing a principled, data-driven framework for fair and robust comparative evaluation of DR methods.
Training deep neural networks often suffers from catastrophic loss explosions, leading to costly failures. Conventional monitoring metrics—such as weight or gradient norms—are lagging and lack discriminative power for early detection. This paper proposes a spectral alignment–based early-warning mechanism: it quantifies the alignment between layer-wise input distributions and the dominant left singular vectors of weight matrices to detect incipient representational collapse. Theoretically, we show that the collapse of sign diversity in spectral alignment serves as an interpretable, pre-divergence indicator of training instability. Our method requires only lightweight SVD computation and statistical tracking, entailing minimal overhead and straightforward deployment. Empirical evaluation on language models demonstrates that our approach issues warnings significantly earlier than conventional metrics, with clearer signals and stronger generalization across architectures. This work establishes a novel paradigm for stabilizing large-model training through interpretable, spectrum-aware monitoring.
Existing approaches to measuring functional similarity between models rely on the true data distribution, making it difficult to characterize alignment of decision boundaries across the entire input space. This work proposes Rashomon Alignment (RA), a novel framework that, for the first time, evaluates functional similarity between models from a geometric perspective over the full input space without dependence on any specific data distribution. By uniformly sampling the input space and employing geometric similarity metrics, RA enables a global analysis of decision boundary alignment. Experiments across more than 90 datasets demonstrate that geometric alignment provides a complementary perspective to distribution-based alignment, and that RA effectively supports model selection, ensemble construction, and enhanced interpretability.
To address the misalignment between automatic evaluation metrics and human preferences in generative tasks, this paper proposes MetaMetrics—a calibratable meta-metric that supervisely weights and fuses existing metrics to model fine-grained human preferences across multimodal (language/vision), multilingual, and multi-domain settings. Methodologically, it introduces the first preference-dimension-aware metric calibration framework, enabling cross-modal unified evaluation and plug-and-play integration. The approach combines supervised meta-learning, multi-task joint optimization, and explicit modeling of human preference annotations. Experiments demonstrate that MetaMetrics significantly improves correlation with human judgments across multilingual text and vision generation tasks (average Kendall’s τ increase of +18.7%). Moreover, it exhibits strong generalization to unseen domains and models, maintaining robust alignment with human preferences without task-specific retraining.
In online education, precise alignment between learning resources and objectives relies on costly manual curation, hindering scalable personalized instruction. To address this, we propose the first text-embedding-based automated alignment evaluation framework, systematically benchmarking multiple models on educational alignment tasks and demonstrating that semantic similarity reliably predicts learning outcomes. Leveraging the Voyage embedding model, we achieve fine-grained semantic matching between learning objectives and resources, supported by a rigorous evidence chain comprising model evaluation, expert validation, and empirical testing. Three randomized controlled trials (N = 360) show that our framework achieves 79% alignment identification accuracy on human-curated resources and 83% on LLM-generated resources. Critically, high-alignment resources significantly improve learning performance (χ² = 15.39, p < 0.001). This work establishes a novel, low-cost, and scalable paradigm for intelligent educational resource filtering.
Existing vision-language models exhibit limited performance on medical image–text tasks and lack effective tools to quantify inter-modal information imbalance. This work proposes the Asymmetric Spectral Alignment Score (SAS), introducing for the first time a directional alignment metric that projects multimodal representations onto the principal component basis of an anchor modality and computes modality-wise correlations weighted by eigenvalues. SAS reveals an asymmetry in which medical images retain richer structural information than clinical text. Integrated into an evaluation framework encompassing 15 vision-language models and six alignment metrics, SAS demonstrates the strongest correlation with bidirectional retrieval performance under label-free conditions, offering a practical and interpretable tool for assessing medical multimodal models.
This study addresses the lack of cognitive alignment and neglect of population heterogeneity in existing face similarity metrics by proposing an interpretable, human-perception-aligned method grounded in cognitive psychology. The approach explicitly models cognitive mechanisms—including facial features, nonlinear responses, and group biases—and integrates vision-language models, gated cross-attention, concept bottlenecks, and neural generalized additive models for joint optimization. Experimental results on the FACETS dataset demonstrate that the proposed method significantly outperforms state-of-the-art metrics, effectively enhancing both alignment accuracy with human perceptual judgments across diverse populations and model interpretability. This work bridges computational modeling and cognitive science to advance more psychologically plausible face representation learning.
Existing segmentation evaluation metrics often lack transparency and modularity, making them ill-suited for diverse tasks such as transparent object, specular surface, or lesion segmentation. This work proposes a unified evaluation framework that decomposes metrics into five modular components: prediction representation, target extraction, target matching, score computation, and metric reporting. For the first time, it systematically analyzes the implicit assumptions and design limitations of mainstream binary segmentation metrics through this modular lens. The framework enables task-aware customization of evaluation protocols, reveals evolutionary trajectories among existing metrics, and is accompanied by an open-source toolkit. By offering a principled and interpretable foundation, this approach paves the way for developing more rational and adaptable segmentation evaluation methodologies.
This work addresses the critical challenge of evaluating model generalization in high-stakes scenarios with scarce labels, where existing methods lack reliable, label-free metrics for pre-deployment model selection and post-deployment performance monitoring. To bridge this gap, the study introduces, for the first time, the internal causal circuit mechanisms of Vision Transformers into generalization assessment, proposing two novel unsupervised metrics: Dependency Depth Bias and Circuit Shift Score. The former quantifies depth-wise biases in representational dependency structures, while the latter measures changes in circuit stability under distribution shifts. Extensive experiments across diverse tasks demonstrate that these metrics achieve substantially higher correlations with true generalization performance—improving by 13.4% and 34.1% on average over current approaches—thereby significantly enhancing the reliability of generalization prediction without requiring ground-truth labels.