modality-specific normalization

Designs and implements normalization schemes applied separately to each data modality so that embeddings or features from different modalities become comparable; this includes per-modality scaling, distributional adjustment, and procedures that preserve scale invariance. Builds methods that handle differing feature dimensionalities and counts and produce normalized representations robust to missing modalities for downstream analyses such as clustering or similarity computation.

modality-specificnormalization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Gradient-Weighted, Data-Driven Normalization for Approximate Border Bases -- Concept and Computation

Jun 11, 2025
HK
Hiroshi Kera
🏛️ Chiba University | Zuse Institute Berlin | Rhine-Waal University of Applied Sciences

Computing approximate vanishing ideals from scarce, uncertain data points suffers from poor numerical stability under traditional coefficient-norm-based polynomial normalization, which is scale-sensitive and lacks robustness. Method: We propose a gradient-weighted, data-driven seminorm normalization strategy—novelly integrated into border basis construction—to replace conventional normalization. This approach rigorously guarantees scale invariance and significantly enhances robustness against data perturbations. Contribution/Results: Theoretical analysis and experiments on three affine varieties demonstrate that the new method completely eliminates scale dependence, improves noise resilience substantially, and requires only minor algorithmic adjustments without increasing time complexity. Its core contribution is a geometrically meaningful and numerically stable normalization framework, establishing a new paradigm for approximate algebraic geometry modeling.

Adapting border basis concept for approximate treatment of uncertain data pointsDemonstrating superior robustness and invariance compared to coefficient normalizationProposing gradient-weighted normalization for better stability and scaling invariance

CM$^3$: Calibrating Multimodal Recommendation

Aug 02, 2025
XZ
Xin Zhou
🏛️ Nanyang Technological University

Existing multimodal recommendation models suffer from an imbalance between embedding space alignment and uniformity—overemphasizing uniform distribution on the hypersphere manifold undermines the geometric proximity of semantically similar items. Method: We propose a spherical representation learning framework with multimodal similarity calibration, featuring: (1) a calibrated uniformity loss that dynamically adjusts item density on the hypersphere according to multimodal similarity; (2) a spherical Bessel fusion mechanism enabling geometrically consistent integration of multimodal features within the manifold space; and (3) incorporation of fine-grained semantic features extracted by multimodal large language models (MLLMs). Results: Evaluated on five real-world datasets, our method achieves up to a 5.4% improvement in NDCG@20 over state-of-the-art baselines, demonstrating the effectiveness of enhanced alignment and manifold-aware multimodal fusion.

Addresses imbalance between alignment and uniformity in multimodal recommender systemsIntroduces Spherical Bézier method to fuse multimodal features effectivelyProposes calibrated uniformity loss using multimodal similarity for better embeddings

How Scale Breaks "Normalized Stress" and KL Divergence: Rethinking Quality Metrics

Oct 09, 2025
KS
Kiran Smelser
🏛️ University of Arizona | Technical University of Munich

Existing visualization quality metrics—such as normalized stress and KL divergence—are highly sensitive to uniform scaling of projections, despite such transformations preserving structural fidelity and thus inducing evaluation distortion. This work is the first to systematically characterize this scale dependence and proposes a theoretically grounded, scale-invariant correction: normalizing the pairwise distance matrix prior to metric computation, ensuring that assessments reflect only relative structural relationships. Experiments across multiple standard benchmark datasets demonstrate that the corrected metrics achieve significantly improved stability and discriminative power—accurately distinguishing high- from low-quality projections in controlled evaluations while exhibiting complete robustness to arbitrary uniform scaling. Crucially, the modification preserves the original computational efficiency and interpretability of the base metrics. This advancement establishes a more perceptually aligned, fair, and reliable foundation for evaluating dimensionality reduction algorithms.

KL divergence in t-SNE suffers from scale sensitivity issuesNormalized stress metric is sensitive to uniform scalingQuality metrics need scale-invariance for accurate projection evaluation

This work addresses the limited interpretability of cosine similarity in pretrained embedding spaces, where absolute similarity values are concentrated in a narrow range due to anisotropy. The authors propose a monotonic calibration method based on isotonic regression that reparameterizes similarity scores without altering the underlying embeddings or their geometric structure. The approach strictly preserves the original ranking order of similarities, thereby maintaining all ordinal-dependent structures—such as nearest neighbors, angular rankings, and threshold graphs—intact. The calibrated similarities achieve near-perfect alignment in absolute values while retaining 98% local stability under seven types of perturbations and fully preserving rank correlation with the original similarities.

anisotropycalibrationcosine similarity

“Normalized Stress” is Not Normalized: How to Interpret Stress Correctly

Aug 14, 2024
KS
Kiran Smelser
🏛️ University of Arizona

Normalized stress is a widely used metric for evaluating dimensionality reduction projections; however, it exhibits sensitivity to uniform scaling of projections, introducing scale-dependent bias that violates the physical plausibility of fidelity assessment. This work proposes the first scale-invariant stress correction method: we theoretically derive a scaling-invariance constraint and integrate distance-metric analysis with empirical validation across multiple benchmark datasets. The corrected stress metric fully eliminates unphysical scale dependence, restores expected monotonicity and ranking consistency in small-scale evaluations, and significantly enhances the fairness and reliability of cross-algorithm comparisons. By establishing rigorous scale invariance—grounded in geometric principles and empirically verified—the proposed metric constitutes the first strictly scale-invariant quantitative standard for assessing high-dimensional visualizations and embeddings.

Current stress metrics distort evaluation of dimension reduction techniquesNormalized stress metric incorrectly responds to projection scaling changesProposing scale-invariant method to properly assess projection quality

Latest Papers

What's happening recently
View more

This work addresses the challenge of high-proportion missing data across arbitrary modality combinations in multimodal learning by proposing UL4M4, a task-agnostic, lightweight, and universally applicable unsupervised modality imputation framework. The method leverages modality-specific normalization and a novel partial-modality distance metric to enable fair clustering under frozen encoders, with cluster centroids guiding an iterative greedy imputation process. UL4M4 is the first approach to support decoupled imputation for any number of modalities and arbitrary missing patterns while effectively preserving cross-modal structural and scale invariance. Experimental results demonstrate that, even under extreme settings with over 50% missing modalities, UL4M4 consistently achieves F1-Micro scores above 0.7, significantly outperforming existing methods and exhibiting robustness across varying clustering scales.

feature imputationincomplete observationsmissing modalities

本文针对受污染的不相似数据问题,提出了一种鲁棒条件多维尺度方法(rcMDS),通过使用Fair M-估计目标替代平方应力准则,并开发了一种重加权条件SMACOF算法来优化此目标。

conditional dimension reductioncontaminated dissimilaritydissimilarity data

This study addresses the issue of erroneous invariance attribution in preprocessing for spectral foundation models by proposing a decoupled evaluation paradigm that distinguishes the contributions of normalization from representation learning. Through Raman spectroscopy analysis, controlled experiments, and numerical validation, we demonstrate that the invariance observed in multimodal models primarily stems from parameter-free preprocessing normalization rather than learned representations. These findings reveal that current models fail to surpass simple normalization baselines, effectively correcting prevailing cognitive biases regarding spectral representation learning capabilities within the field. Consequently, this work establishes a new benchmark for evaluation methodologies, emphasizing the necessity of rigorously isolating preprocessing effects when assessing the true efficacy of spectral foundation models.

AttributionNormalizationPreprocessing Invariance

This work addresses the challenges of cross-modal structural distortion and preference misalignment caused by missing modalities in real-world multimodal recommendation systems. To tackle these issues, the authors propose a two-stage calibration framework: first, structured modality completion with cross-modal correspondence supervision ensures consistency in the imputed content; second, preference-guided representation calibration aligns the completed features with the recommendation ranking objective. A completion-aware item graph is further constructed by integrating the imputed modalities with collaborative signals. The approach innovatively jointly optimizes cross-modal structural consistency and task-oriented representation alignment, incorporating mechanisms such as structural regularization and pseudo-missing data augmentation. Extensive experiments demonstrate that the proposed method significantly outperforms existing approaches across multiple datasets and missing-modality settings, exhibiting both strong effectiveness and robustness.

cross-modal structural distortionincomplete multimodal recommendationmodality imputation

This work investigates the generalization capabilities of tabular foundation models in cross-modal settings and introduces a unified evaluation framework accompanied by a standardized classification pipeline. The approach leverages equiangular tight frame (ETF) preprocessing, in-context learning, and probability calibration to systematically assess model performance across 95 datasets spanning seven distinct modalities. A novel validation-free ETF-based training stopping criterion is proposed, along with a lightweight baseline built upon frozen features. The method achieves performance comparable to task-specific fine-tuned models on most benchmarks while accelerating inference by 4–200× and producing well-calibrated confidence estimates, thereby substantially enhancing practical deployability.

cross-modality transferfrozen feature evaluationin-context learning

Hot Scholars

JY

Junsong Yuan

State University of New York at Buffalo
computer visionvideo analyticsaction and gesture analysismultimedia
WT

Wenjie Tian

Northwest Polytechnical University
speech generation
ZZ

Zhou Zhao

Zhejiang University
Machine LearningData MiningMultimedia Computing
JM

Jiayi Ma

Wuhan University
Computer VisionImage FusionImage Matching