contrastive learning

Designs and builds training objectives, encoders, and pretraining pipelines that learn discriminative latent representations by pulling together positive pairs and pushing apart negatives; this includes creating contrastive losses and architectures that combine autoencoder or masked-reconstruction objectives with contrastive terms, CLIP-style cross‑modal alignment, and soft‑target/soft‑label formulations. Implements and analyzes practical components such as pair/negative sampling strategies, masking strategies, reconstruction‑contrastive hybrids, and pretraining protocols to enforce desired invariances, suppress dataset shortcuts, and improve transfer of the learned embeddings.

contrastivelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$218K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Compression then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding

Nov 11, 2025
DL
Da Li
🏛️ University of Chinese Academy of Sciences | Kuaishou Technology

This work addresses the fundamental bottleneck in vision-language models (VLMs): the difficulty of simultaneously preserving semantic fidelity and ensuring downstream discriminability. To this end, we propose CoMa—a novel pretraining paradigm that decouples semantic preservation from discriminative feature learning by introducing compression learning as a warm-up stage preceding contrastive learning. CoMa achieves efficient semantic distillation and feature compression using only a small amount of data. On the MMEB benchmark, it attains state-of-the-art performance among models of comparable scale, significantly improving embedding quality for cross-modal retrieval, clustering, and classification. Its core innovation lies in the first explicit formulation of compression objectives as a semantic initialization mechanism for contrastive learning—thereby jointly optimizing training efficiency and representation capability. Crucially, CoMa delivers high-performance multimodal embeddings even under stringent low-data-budget constraints.

Decoupling comprehensive understanding from discriminative feature optimizationDeveloping efficient pre-training for multimodal embedding modelsEnhancing vision-language model performance with limited data

Latent space configuration for improved generalization in supervised autoencoder neural networks

Feb 13, 2024
NG
Nikita Gabdullin
🏛️ Joint Stock Research and Production Company Kryptonite

To address the uncontrollable topology of latent spaces (LS) in autoencoders (AEs), this paper proposes a geometric-loss-guided supervised AE co-optimization framework—the first to explicitly configure LS topology in supervised AEs. The method jointly optimizes encoder architecture and geometric constraint losses, enabling user-defined cluster positions and shapes, decoder-free label prediction, and cross-sample similarity assessment. Key innovations include zero-shot cross-dataset generalization, similarity estimation for unseen classes, and text-driven image retrieval without classifiers or language models. Experiments demonstrate 12–19% improvements in zero-shot accuracy on LIP, Market-1501, and WildTrack, and achieve 78.3% mAP in cross-modal retrieval.

Configuring latent space topology for supervised autoencodersEnabling similarity measurement and label prediction directly in latent spaceImproving generalization to unseen datasets without fine-tuning

Implicit Contrastive Representation Learning with Guided Stop-gradient

Mar 12, 2025
BL
Byeongchan Lee
🏛️ Gauss Labs | KAIST

Self-supervised learning faces two key challenges: Siamese networks suffer from representational collapse, while contrastive learning relies on negative samples, leading to poor robustness under small batch sizes. To address these issues, this paper proposes a negative-sample-free implicit contrastive learning paradigm. Its core innovation is a guided stop-gradient mechanism that dynamically blocks gradients between symmetric positive sample pairs, thereby implicitly encoding contrastive signals without requiring negative samples, prediction heads, or asymmetric encoders. The method is fully compatible with SimSiam and BYOL frameworks, needing only standard momentum updates and symmetric loss functions. Extensive experiments demonstrate significant performance gains on ImageNet, robust training with extremely small batch sizes (e.g., 8), and superior training stability and generalization compared to existing negative-sample-free approaches.

Enhance robustness with fewer negative samplesPrevent collapse in self-supervised representation learningStabilize training and improve performance with small batch sizes

Unity by Diversity: Improved Representation Learning in Multimodal VAEs

Mar 08, 2024
TM
Thomas M. Sutter
🏛️ ETH Zurich | UC Irvine

Multimodal variational autoencoders (MVAEs) suffer from overly rigid cross-modal representation coupling, making it difficult to simultaneously ensure high-quality shared representations and modality-specific fidelity. Method: We propose a soft-constrained Mixture-of-Experts (Soft-MoE) prior that replaces hard parameter sharing with learnable gating weights, enabling flexible alignment of modality-specific latent distributions under a unified posterior. This decouples modality-invariant and modality-specific representations while preserving information integrity via variational inference and a soft alignment loss. Contribution/Results: Experiments on multiple benchmarks and real-world multimodal datasets demonstrate that our approach significantly outperforms existing shared-architecture MVAEs. It achieves state-of-the-art performance in both latent representation quality—measured by disentanglement and downstream task accuracy—and missing modality imputation accuracy.

Missing Data ImputationMultimodal Data FusionVariational Autoencoder

Latest Papers

What's happening recently
View more

Existing image representation methods often struggle to simultaneously support both recognition and generation tasks. This work proposes a hypernetwork architecture based on Implicit Neural Representations (INRs), which encodes images into compact model weights that enable efficient reconstruction. By integrating knowledge distillation with pixel-level and perceptual losses, the method establishes a unified visual representation framework. It is the first approach to achieve high-accuracy recognition and high-quality image generation within a single shared embedding space, demonstrating state-of-the-art performance across diverse vision tasks while maintaining a highly compressed embedding dimensionality.

generationimage representation learningimplicit neural representation

This work addresses the limitation of existing unlearning methods for large language models (LLMs), which primarily suppress target information at the output layer but fail to disentangle forgotten and retained knowledge in the representation space. To overcome this, we propose CLReg, a contrastive representation regularization approach that, for the first time, introduces contrastive learning into LLM unlearning. CLReg explicitly separates forgotten and retained features in the latent space, achieving representational disentanglement. Theoretical analysis reveals a direct link between such disentanglement and improved unlearning efficacy, surpassing the constraints of conventional methods that operate solely in the prediction space. Experiments demonstrate that CLReg significantly reduces feature entanglement across multiple benchmarks, consistently enhances the performance of mainstream unlearning techniques, and does so without introducing additional privacy risks.

contrastive representationforget-retain interferenceLLM unlearning

Class imbalance significantly degrades the performance of contrastive learning, yet the underlying mechanisms by which it affects training dynamics and induces representation bias remain theoretically underexplored. This work addresses this gap by introducing a novel theoretical framework that analyzes the training process of Transformer-based contrastive learning under imbalanced data through the lens of neuron weight evolution. The analysis reveals three characteristic phases in the evolution of neuronal weights during training. Guided by these theoretical insights, the authors propose a targeted neuron pruning strategy that effectively mitigates representation bias. Experimental results demonstrate that the proposed method substantially enhances feature separability and overall representation quality in imbalanced scenarios.

contrastive learningimbalanced datamodel bias

This work investigates under what positive sample sampling conditions contrastive learning can recover a meaningful geometric structure in the latent space. By constructing a measure-theoretic framework, the study introduces a “diversity condition” as a necessary requirement for the identifiability of latent geometry and elucidates the joint influence of sampling support and encoder inductive bias on representation identifiability. Theoretically, it is shown that under full-support sampling, the global optimum of InfoNCE recovers the latent structure up to orthogonal equivalence; however, under non-full support, non-orthogonal mappings may yield better solutions. To address this, the authors propose a support-corrected variant of InfoNCE and model representations using the von Mises–Fisher distribution, empirically validating on both synthetic and real-world data the critical role of inductive bias when sampling diversity is limited.

contrastive learningidentifiabilityinductive bias

Existing methods jointly optimize contrastive alignment and masked reconstruction objectives, which often introduces semantic noise and causes optimization interference, thereby limiting cross-modal representation learning performance. This work proposes the TG-DP framework, which decouples reconstruction and alignment tasks along separate optimization paths for the first time. Each path employs a visibility pattern tailored to its specific objective, and a teacher model is introduced to guide the organization of visible tokens in the contrastive path, reducing interference and enhancing representation quality. The proposed method achieves significant improvements in zero-shot retrieval on AudioSet—R@1 increases from 35.2% to 37.4% (video→audio) and from 27.9% to 37.1% (audio→video)—and attains state-of-the-art linear probe performance on both AS20K and VGGSound benchmarks.

audio-visual representation learningcontrastive alignmentmasked reconstruction

Hot Scholars

TY

Tianbao Yang

Texas A&M University
machine learningstochastic optimization
TS

Tania Stathaki

Imperial College London
Object TrackingImage FusionImage RegistrationImage Processing
YM

Yuki Mitsufuji

Distinguished Engineer, Sony
Machine LearningAudioSource SeparationMusic Technology
YW

Yisen Wang

Assistant Professor, Peking University
Machine LearningSelf-Supervised LearningLarge Language ModelsSafety