Similarity-Distance-Magnitude Activations

📅 2025-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing softmax lacks explicit modeling of sample similarity and distance to the training distribution, resulting in poor robustness under covariate shift and out-of-distribution (OOD) inputs, as well as limited interpretability. This paper proposes the Similarity–Distribution–Magnitude-aware (SDM) activation function, the first to jointly embed deep similarity matching, training-distribution distance estimation, and output magnitude control directly into the activation layer—enabling multi-dimensional disentangled modeling of prediction confidence. SDM supports instance-level interpretability, class-wise empirical cumulative distribution partitioning, and recall preservation, effectively mitigating low-recall issues in selective classification. Experiments demonstrate that SDM significantly outperforms softmax and state-of-the-art post-hoc calibration methods in robustness against OOD samples and covariate shift—particularly within high-confidence prediction regions—while maintaining strong discriminative performance.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationNatural Language Processing: Safety and RobustnessReasoning under Uncertainty: Relational Probabilistic Models

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
We introduce a more robust and interpretable formulation of the standard softmax activation function commonly used with neural networks by adding Similarity (i.e., correctly predicted depth-matches into training) awareness and Distance-to-training-distribution awareness to the existing output Magnitude (i.e., decision-boundary) awareness. When used as the final-layer activation with language models, the resulting Similarity-Distance-Magnitude (SDM) activation function is more robust than the softmax function to co-variate shifts and out-of-distribution inputs in high-probability regions, and provides interpretability-by-exemplar via dense matching. Complementing the prediction-conditional estimates, the SDM activation enables a partitioning of the class-wise empirical CDFs to guard against low class-wise recall among selective classifications. These properties make it preferable for selective classification, even when considering post-hoc calibration methods over the softmax.
Problem

Research questions and friction points this paper is trying to address.

Enhancing neural network robustness against distribution shifts
Improving interpretability via similarity-aware activation functions
Addressing selective classification reliability with empirical CDFs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates similarity, distance, and magnitude awareness
Enhances robustness to distribution shifts and OOD inputs
Provides interpretability via dense exemplar matching