Score
Design and implement contrastive learning objectives, positive/negative sampling and augmentation strategies, and training protocols that form positive pairs across distinct instances or variants of the same semantic entity so that encoders and composition modules produce representations invariant to instance- or variant-level differences. Analyze and evaluate loss formulations and model behavior to suppress background- or shortcut-specific features, align composed and target representations, and improve robustness of the learned representations.
Existing contrastive learning methods define similarity exclusively over semantically consistent sample pairs, overlooking latent structural similarities inherent in semantically dissimilar pairs. To address this limitation, we propose SimLAP—a novel framework that redefines positive pairs not as semantically identical samples but as learnable, discriminative subspace-aligned pairs. SimLAP jointly optimizes pairwise similarity estimation and subspace projection within an end-to-end training paradigm, incorporating both contrastive loss and explicit subspace alignment constraints. By uncovering and leveraging structural similarities among inter-class samples residing in shared latent subspaces, SimLAP breaks from conventional similarity modeling paradigms. Extensive experiments across multiple benchmarks demonstrate its effectiveness: SimLAP significantly improves few-shot transfer performance and model robustness under distribution shifts. Moreover, it offers a new perspective on unsupervised similarity learning—shifting focus from semantic identity to geometric consistency in learned representation subspaces.
In instance-discriminative self-supervised learning, conventional data augmentation may erroneously repel semantically similar samples and discard class-discriminative features. To address this, we propose a semantic-augmented contrastive learning framework that explicitly incorporates semantically similar images as additional positive pairs—a novel design first introduced in this work. Our method dynamically mines semantic positives based on feature similarity and seamlessly integrates with mainstream architectures such as MoCo and SimSiam. On ImageNet, it achieves a +4.1% improvement in linear evaluation accuracy (800 epochs) over MoCo-v2. Significant gains are also observed on STL-10, CIFAR-10, and downstream object detection tasks. Our core contributions are threefold: (1) the first explicit modeling of semantic similarity as a prior for positive pair construction; (2) enhanced semantic richness and intra-class discriminability of learned representations; and (3) strong generality and extensibility across architectures and downstream tasks.
Random cropping in contrastive learning often induces semantic inconsistency between dual views, degrading representation quality. To address this, we propose a novel contrastive learning paradigm that incorporates the original (uncropped) image: a dual-branch encoder processes one augmented view in one branch and the full-resolution original image in the other; an enhanced InfoNCE loss is designed to explicitly enforce consistency between local cropped views and global semantic context, thereby mitigating feature confusion caused by cropping distortion. This is the first work to integrate the original image into the instance discrimination framework without requiring additional annotations or computational overhead. In linear evaluation on ImageNet-1K, our method outperforms MoCo-v2 by 5.1%. It also achieves substantial gains over mainstream self-supervised baselines on transfer classification and object detection tasks.
Self-supervised learning faces two key challenges: Siamese networks suffer from representational collapse, while contrastive learning relies on negative samples, leading to poor robustness under small batch sizes. To address these issues, this paper proposes a negative-sample-free implicit contrastive learning paradigm. Its core innovation is a guided stop-gradient mechanism that dynamically blocks gradients between symmetric positive sample pairs, thereby implicitly encoding contrastive signals without requiring negative samples, prediction heads, or asymmetric encoders. The method is fully compatible with SimSiam and BYOL frameworks, needing only standard momentum updates and symmetric loss functions. Extensive experiments demonstrate significant performance gains on ImageNet, robust training with extremely small batch sizes (e.g., 8), and superior training stability and generalization compared to existing negative-sample-free approaches.
Existing vision-language pretraining (VLP) models suffer from low computational efficiency in modeling long visual sequences and suboptimal semantic alignment due to the conventional InfoNCE loss in cross-modal contrastive learning, which erroneously treats semantically similar samples as negatives. To address these issues, we propose Semantic-Aware Contrastive Learning (SACL), a mutual information maximization–inspired framework. SACL introduces, for the first time, a cross-modal similarity modulation mechanism that dynamically adjusts negative-pair contrastive strength based on semantic proximity. It further formulates a theoretically grounded, mutual information–based weighted contrastive loss and integrates it into a lightweight VLP architecture. Extensive experiments on downstream tasks—including VQA, NLVR2, and image/text retrieval—demonstrate consistent and significant performance gains. Results validate that preserving semantically informative “false negatives” enhances both model generalization and cross-modal alignment fidelity.
This work investigates under what positive sample sampling conditions contrastive learning can recover a meaningful geometric structure in the latent space. By constructing a measure-theoretic framework, the study introduces a “diversity condition” as a necessary requirement for the identifiability of latent geometry and elucidates the joint influence of sampling support and encoder inductive bias on representation identifiability. Theoretically, it is shown that under full-support sampling, the global optimum of InfoNCE recovers the latent structure up to orthogonal equivalence; however, under non-full support, non-orthogonal mappings may yield better solutions. To address this, the authors propose a support-corrected variant of InfoNCE and model representations using the von Mises–Fisher distribution, empirically validating on both synthetic and real-world data the critical role of inductive bias when sampling diversity is limited.
Concept alignment lacks a unified definition, and existing methods optimize divergent objectives under the same terminology, obscuring its fundamental nature. This work formalizes its multidimensional structure by decomposing it along two axes—“alignment targets” and “alignment levels”—and identifies four distinct alignment properties, revealing that current approaches satisfy only subsets of these. To address this limitation, we propose Coupled Sparse Autoencoders (CoSAE), a framework that jointly optimizes multiple alignment objectives, alongside InterVenchA, an interventional evaluation benchmark. Experiments demonstrate that optimizing a single objective fails to reliably recover other alignment properties, whereas CoSAE achieves strong instance-level conceptual consistency using merely 0.1% paired data.
This work reveals that the often-overlooked embedding norm in contrastive learning inherently encodes critical semantic information, such as semantic specificity. From the perspective of optimization dynamics, we theoretically demonstrate—for the first time—the mechanism by which embedding norms naturally capture semantic attributes during training under scale-invariant losses. We derive analytical relationships between the norm and established semantic metrics, including concept specificity, word frequency, and human uncertainty. Furthermore, we show that the norm serves as a calibration signal without requiring additional training, offering significant practical utility in retrieval and confidence calibration tasks.
Class imbalance significantly degrades the performance of contrastive learning, yet the underlying mechanisms by which it affects training dynamics and induces representation bias remain theoretically underexplored. This work addresses this gap by introducing a novel theoretical framework that analyzes the training process of Transformer-based contrastive learning under imbalanced data through the lens of neuron weight evolution. The analysis reveals three characteristic phases in the evolution of neuronal weights during training. Guided by these theoretical insights, the authors propose a targeted neuron pruning strategy that effectively mitigates representation bias. Experimental results demonstrate that the proposed method substantially enhances feature separability and overall representation quality in imbalanced scenarios.