Score
Designs and builds training objectives, encoders, and pretraining pipelines that learn discriminative latent representations by pulling together positive pairs and pushing apart negatives; this includes creating contrastive losses and architectures that combine autoencoder or masked-reconstruction objectives with contrastive terms, CLIP-style cross‑modal alignment, and soft‑target/soft‑label formulations. Implements and analyzes practical components such as pair/negative sampling strategies, masking strategies, reconstruction‑contrastive hybrids, and pretraining protocols to enforce desired invariances, suppress dataset shortcuts, and improve transfer of the learned embeddings.
This work addresses the fundamental bottleneck in vision-language models (VLMs): the difficulty of simultaneously preserving semantic fidelity and ensuring downstream discriminability. To this end, we propose CoMa—a novel pretraining paradigm that decouples semantic preservation from discriminative feature learning by introducing compression learning as a warm-up stage preceding contrastive learning. CoMa achieves efficient semantic distillation and feature compression using only a small amount of data. On the MMEB benchmark, it attains state-of-the-art performance among models of comparable scale, significantly improving embedding quality for cross-modal retrieval, clustering, and classification. Its core innovation lies in the first explicit formulation of compression objectives as a semantic initialization mechanism for contrastive learning—thereby jointly optimizing training efficiency and representation capability. Crucially, CoMa delivers high-performance multimodal embeddings even under stringent low-data-budget constraints.
To address the uncontrollable topology of latent spaces (LS) in autoencoders (AEs), this paper proposes a geometric-loss-guided supervised AE co-optimization framework—the first to explicitly configure LS topology in supervised AEs. The method jointly optimizes encoder architecture and geometric constraint losses, enabling user-defined cluster positions and shapes, decoder-free label prediction, and cross-sample similarity assessment. Key innovations include zero-shot cross-dataset generalization, similarity estimation for unseen classes, and text-driven image retrieval without classifiers or language models. Experiments demonstrate 12–19% improvements in zero-shot accuracy on LIP, Market-1501, and WildTrack, and achieve 78.3% mAP in cross-modal retrieval.
Self-supervised learning faces two key challenges: Siamese networks suffer from representational collapse, while contrastive learning relies on negative samples, leading to poor robustness under small batch sizes. To address these issues, this paper proposes a negative-sample-free implicit contrastive learning paradigm. Its core innovation is a guided stop-gradient mechanism that dynamically blocks gradients between symmetric positive sample pairs, thereby implicitly encoding contrastive signals without requiring negative samples, prediction heads, or asymmetric encoders. The method is fully compatible with SimSiam and BYOL frameworks, needing only standard momentum updates and symmetric loss functions. Extensive experiments demonstrate significant performance gains on ImageNet, robust training with extremely small batch sizes (e.g., 8), and superior training stability and generalization compared to existing negative-sample-free approaches.
Multimodal variational autoencoders (MVAEs) suffer from overly rigid cross-modal representation coupling, making it difficult to simultaneously ensure high-quality shared representations and modality-specific fidelity. Method: We propose a soft-constrained Mixture-of-Experts (Soft-MoE) prior that replaces hard parameter sharing with learnable gating weights, enabling flexible alignment of modality-specific latent distributions under a unified posterior. This decouples modality-invariant and modality-specific representations while preserving information integrity via variational inference and a soft alignment loss. Contribution/Results: Experiments on multiple benchmarks and real-world multimodal datasets demonstrate that our approach significantly outperforms existing shared-architecture MVAEs. It achieves state-of-the-art performance in both latent representation quality—measured by disentanglement and downstream task accuracy—and missing modality imputation accuracy.
Existing image representation methods often struggle to simultaneously support both recognition and generation tasks. This work proposes a hypernetwork architecture based on Implicit Neural Representations (INRs), which encodes images into compact model weights that enable efficient reconstruction. By integrating knowledge distillation with pixel-level and perceptual losses, the method establishes a unified visual representation framework. It is the first approach to achieve high-accuracy recognition and high-quality image generation within a single shared embedding space, demonstrating state-of-the-art performance across diverse vision tasks while maintaining a highly compressed embedding dimensionality.
This work addresses the limitation of existing unlearning methods for large language models (LLMs), which primarily suppress target information at the output layer but fail to disentangle forgotten and retained knowledge in the representation space. To overcome this, we propose CLReg, a contrastive representation regularization approach that, for the first time, introduces contrastive learning into LLM unlearning. CLReg explicitly separates forgotten and retained features in the latent space, achieving representational disentanglement. Theoretical analysis reveals a direct link between such disentanglement and improved unlearning efficacy, surpassing the constraints of conventional methods that operate solely in the prediction space. Experiments demonstrate that CLReg significantly reduces feature entanglement across multiple benchmarks, consistently enhances the performance of mainstream unlearning techniques, and does so without introducing additional privacy risks.
Class imbalance significantly degrades the performance of contrastive learning, yet the underlying mechanisms by which it affects training dynamics and induces representation bias remain theoretically underexplored. This work addresses this gap by introducing a novel theoretical framework that analyzes the training process of Transformer-based contrastive learning under imbalanced data through the lens of neuron weight evolution. The analysis reveals three characteristic phases in the evolution of neuronal weights during training. Guided by these theoretical insights, the authors propose a targeted neuron pruning strategy that effectively mitigates representation bias. Experimental results demonstrate that the proposed method substantially enhances feature separability and overall representation quality in imbalanced scenarios.
This work investigates under what positive sample sampling conditions contrastive learning can recover a meaningful geometric structure in the latent space. By constructing a measure-theoretic framework, the study introduces a “diversity condition” as a necessary requirement for the identifiability of latent geometry and elucidates the joint influence of sampling support and encoder inductive bias on representation identifiability. Theoretically, it is shown that under full-support sampling, the global optimum of InfoNCE recovers the latent structure up to orthogonal equivalence; however, under non-full support, non-orthogonal mappings may yield better solutions. To address this, the authors propose a support-corrected variant of InfoNCE and model representations using the von Mises–Fisher distribution, empirically validating on both synthetic and real-world data the critical role of inductive bias when sampling diversity is limited.
Existing methods jointly optimize contrastive alignment and masked reconstruction objectives, which often introduces semantic noise and causes optimization interference, thereby limiting cross-modal representation learning performance. This work proposes the TG-DP framework, which decouples reconstruction and alignment tasks along separate optimization paths for the first time. Each path employs a visibility pattern tailored to its specific objective, and a teacher model is introduced to guide the organization of visible tokens in the contrastive path, reducing interference and enhancing representation quality. The proposed method achieves significant improvements in zero-shot retrieval on AudioSet—R@1 increases from 35.2% to 37.4% (video→audio) and from 27.9% to 37.1% (audio→video)—and attains state-of-the-art linear probe performance on both AS20K and VGGSound benchmarks.