Score
Designs and implements normalization schemes applied separately to each data modality so that embeddings or features from different modalities become comparable; this includes per-modality scaling, distributional adjustment, and procedures that preserve scale invariance. Builds methods that handle differing feature dimensionalities and counts and produce normalized representations robust to missing modalities for downstream analyses such as clustering or similarity computation.
Computing approximate vanishing ideals from scarce, uncertain data points suffers from poor numerical stability under traditional coefficient-norm-based polynomial normalization, which is scale-sensitive and lacks robustness. Method: We propose a gradient-weighted, data-driven seminorm normalization strategy—novelly integrated into border basis construction—to replace conventional normalization. This approach rigorously guarantees scale invariance and significantly enhances robustness against data perturbations. Contribution/Results: Theoretical analysis and experiments on three affine varieties demonstrate that the new method completely eliminates scale dependence, improves noise resilience substantially, and requires only minor algorithmic adjustments without increasing time complexity. Its core contribution is a geometrically meaningful and numerically stable normalization framework, establishing a new paradigm for approximate algebraic geometry modeling.
Existing multimodal recommendation models suffer from an imbalance between embedding space alignment and uniformity—overemphasizing uniform distribution on the hypersphere manifold undermines the geometric proximity of semantically similar items. Method: We propose a spherical representation learning framework with multimodal similarity calibration, featuring: (1) a calibrated uniformity loss that dynamically adjusts item density on the hypersphere according to multimodal similarity; (2) a spherical Bessel fusion mechanism enabling geometrically consistent integration of multimodal features within the manifold space; and (3) incorporation of fine-grained semantic features extracted by multimodal large language models (MLLMs). Results: Evaluated on five real-world datasets, our method achieves up to a 5.4% improvement in NDCG@20 over state-of-the-art baselines, demonstrating the effectiveness of enhanced alignment and manifold-aware multimodal fusion.
Existing visualization quality metrics—such as normalized stress and KL divergence—are highly sensitive to uniform scaling of projections, despite such transformations preserving structural fidelity and thus inducing evaluation distortion. This work is the first to systematically characterize this scale dependence and proposes a theoretically grounded, scale-invariant correction: normalizing the pairwise distance matrix prior to metric computation, ensuring that assessments reflect only relative structural relationships. Experiments across multiple standard benchmark datasets demonstrate that the corrected metrics achieve significantly improved stability and discriminative power—accurately distinguishing high- from low-quality projections in controlled evaluations while exhibiting complete robustness to arbitrary uniform scaling. Crucially, the modification preserves the original computational efficiency and interpretability of the base metrics. This advancement establishes a more perceptually aligned, fair, and reliable foundation for evaluating dimensionality reduction algorithms.
This work addresses the limited interpretability of cosine similarity in pretrained embedding spaces, where absolute similarity values are concentrated in a narrow range due to anisotropy. The authors propose a monotonic calibration method based on isotonic regression that reparameterizes similarity scores without altering the underlying embeddings or their geometric structure. The approach strictly preserves the original ranking order of similarities, thereby maintaining all ordinal-dependent structures—such as nearest neighbors, angular rankings, and threshold graphs—intact. The calibrated similarities achieve near-perfect alignment in absolute values while retaining 98% local stability under seven types of perturbations and fully preserving rank correlation with the original similarities.
Normalized stress is a widely used metric for evaluating dimensionality reduction projections; however, it exhibits sensitivity to uniform scaling of projections, introducing scale-dependent bias that violates the physical plausibility of fidelity assessment. This work proposes the first scale-invariant stress correction method: we theoretically derive a scaling-invariance constraint and integrate distance-metric analysis with empirical validation across multiple benchmark datasets. The corrected stress metric fully eliminates unphysical scale dependence, restores expected monotonicity and ranking consistency in small-scale evaluations, and significantly enhances the fairness and reliability of cross-algorithm comparisons. By establishing rigorous scale invariance—grounded in geometric principles and empirically verified—the proposed metric constitutes the first strictly scale-invariant quantitative standard for assessing high-dimensional visualizations and embeddings.
This work addresses the challenge of high-proportion missing data across arbitrary modality combinations in multimodal learning by proposing UL4M4, a task-agnostic, lightweight, and universally applicable unsupervised modality imputation framework. The method leverages modality-specific normalization and a novel partial-modality distance metric to enable fair clustering under frozen encoders, with cluster centroids guiding an iterative greedy imputation process. UL4M4 is the first approach to support decoupled imputation for any number of modalities and arbitrary missing patterns while effectively preserving cross-modal structural and scale invariance. Experimental results demonstrate that, even under extreme settings with over 50% missing modalities, UL4M4 consistently achieves F1-Micro scores above 0.7, significantly outperforming existing methods and exhibiting robustness across varying clustering scales.
本文针对受污染的不相似数据问题,提出了一种鲁棒条件多维尺度方法(rcMDS),通过使用Fair M-估计目标替代平方应力准则,并开发了一种重加权条件SMACOF算法来优化此目标。
This study addresses the issue of erroneous invariance attribution in preprocessing for spectral foundation models by proposing a decoupled evaluation paradigm that distinguishes the contributions of normalization from representation learning. Through Raman spectroscopy analysis, controlled experiments, and numerical validation, we demonstrate that the invariance observed in multimodal models primarily stems from parameter-free preprocessing normalization rather than learned representations. These findings reveal that current models fail to surpass simple normalization baselines, effectively correcting prevailing cognitive biases regarding spectral representation learning capabilities within the field. Consequently, this work establishes a new benchmark for evaluation methodologies, emphasizing the necessity of rigorously isolating preprocessing effects when assessing the true efficacy of spectral foundation models.
This work addresses the challenges of cross-modal structural distortion and preference misalignment caused by missing modalities in real-world multimodal recommendation systems. To tackle these issues, the authors propose a two-stage calibration framework: first, structured modality completion with cross-modal correspondence supervision ensures consistency in the imputed content; second, preference-guided representation calibration aligns the completed features with the recommendation ranking objective. A completion-aware item graph is further constructed by integrating the imputed modalities with collaborative signals. The approach innovatively jointly optimizes cross-modal structural consistency and task-oriented representation alignment, incorporating mechanisms such as structural regularization and pseudo-missing data augmentation. Extensive experiments demonstrate that the proposed method significantly outperforms existing approaches across multiple datasets and missing-modality settings, exhibiting both strong effectiveness and robustness.
This work investigates the generalization capabilities of tabular foundation models in cross-modal settings and introduces a unified evaluation framework accompanied by a standardized classification pipeline. The approach leverages equiangular tight frame (ETF) preprocessing, in-context learning, and probability calibration to systematically assess model performance across 95 datasets spanning seven distinct modalities. A novel validation-free ETF-based training stopping criterion is proposed, along with a lightweight baseline built upon frozen features. The method achieves performance comparable to task-specific fine-tuned models on most benchmarks while accelerating inference by 4–200× and producing well-calibrated confidence estimates, thereby substantially enhancing practical deployability.