Score
Designs and implements loss functions and optimization procedures that explicitly account for individual instances by separating instance-specific and background semantics, penalizing incorrect assignment, and steering embedding updates toward per-instance feature representations. Builds and analyzes objective terms, sampling and weighting schemes, and gradient targets that produce instance-discriminative embeddings and optimization dynamics for tasks requiring instance-level distinction.
This work addresses the limitation of instance discrimination—a dominant self-supervised learning paradigm—in capturing high-level semantic invariances. To this end, we propose a novel dataset and evaluation framework grounded in *semantic pairs*: positive sample pairs explicitly constructed to reflect meaningful semantic relationships (e.g., fine-grained subclasses within the same category or cross-domain synonymous instances), rather than relying on stochastic data augmentations. Our method integrates such semantically informed pairs into contrastive learning to explicitly guide encoders toward learning invariant representations at the semantic level. Experiments demonstrate substantial improvements in transfer performance on downstream tasks—including ImageNet classification and PASCAL VOC detection—achieving an average +2.3% gain over strong baselines (e.g., SimCLR, MoCo). Furthermore, the publicly released dataset fills a critical gap in semantic-aware self-supervised evaluation, establishing a new benchmark that empirically supports the shift from low-level feature modeling toward semantic understanding in self-supervised representation learning.
This work addresses the challenge of overfitting and poor generalization in multiple instance learning (MIL) under label-scarce conditions by proposing a context-based, fine-tuning-free approach. The method leverages a Perceiver architecture pretrained on diverse synthetic bag-structured datasets, integrating complementary inductive biases from varied generation strategies. This enables the model to perform accurate classification on new MIL tasks through a single forward pass with only a few labeled bags, without requiring any gradient-based adaptation. Evaluated across twelve established MIL benchmarks, the proposed approach consistently outperforms supervised baselines that rely on task-specific training, demonstrating substantially improved generalization and practical utility in few-shot MIL scenarios.
This work addresses the challenge of fine-grained relational modeling between a query instance and heterogeneous, relationally ambiguous target instance bags in multi-instance verification. We propose Cross-Attention Pooling (CAP), the first framework to generate query-aware bag representations. CAP introduces two novel query-guided attention functions that dynamically aggregate discriminative instances and explicitly model inter-instance dependencies—overcoming fundamental limitations of conventional MIL and Siamese architectures in capturing complex relevance structures. Evaluated on three distinct verification tasks, CAP consistently outperforms state-of-the-art MIL variants and strong baselines, achieving simultaneous gains in classification accuracy and explanation quality. Ablation studies confirm CAP’s robust capability in identifying critical instances. Overall, CAP establishes a new paradigm for interpretable multi-instance verification by unifying representation learning, relational reasoning, and explainability within a single differentiable framework.
This work addresses the family of parametric optimization problems and proposes the first unified, data-driven framework for analyzing the generalization performance of both classical and learned optimizers. Methodologically: (1) it introduces PAC-Bayes theory to the analysis of learned optimizers, deriving verifiable, high-probability generalization upper bounds; (2) it establishes performance bounds for classical optimizers based on empirical convergence rates; and (3) it pioneers a learning paradigm that directly minimizes the PAC-Bayes bound during training. Evaluated on signal processing, control, and meta-learning tasks, the derived bounds are significantly tighter than conventional worst-case guarantees. Moreover, the theoretical generalization guarantees for learned optimizers consistently exceed the empirical performance of their non-learned baselines—thereby unifying theoretical rigor with practical efficacy.
Hard negative sampling in contrastive learning critically influences representation geometry, yet its precise roles in mitigating dimensional collapse (DC) and inducing neural collapse (NC) remain poorly understood. Method: We develop a generalized contrastive loss framework, integrating equiangular tight frame (ETF) geometric modeling and unit-sphere normalization analysis. Contribution/Results: We provide the first rigorous proof that, under both supervised and unsupervised hard contrastive learning (HSCL/HUCL), any global optimum necessarily exhibits NC—i.e., class means form an ETF and intra-class features collapse to identical points. Crucially, we extend this result to generic losses including InfoNCE without assuming class-conditional independence. Theory and experiments jointly demonstrate that NC emerges stably *only* when hard negative sampling is synergistically combined with feature normalization under Adam-based batch optimization; otherwise, DC prevails. Our code is publicly available.
This work addresses the challenge of attributing misclassifications and evaluating robustness in black-box classifiers by proposing an explainability-aware optimization framework. The approach integrates L₀ sparsity regularization (XA-L₀) with a tolerance-region confusion matrix (TOR-Confusion Matrix) to generate minimal input perturbations that induce target predictions while preserving sparsity and semantic interpretability. This unified framework simultaneously enables the generation of highly interpretable counterfactual examples and fine-grained quantification of model robustness. Empirical evaluations on both image and tabular datasets demonstrate the method’s effectiveness, significantly outperforming existing black-box analysis techniques in terms of interpretability and robustness assessment fidelity.
This work addresses the scalability challenge of pairwise loss functions in large-scale machine learning, whose computational complexity grows quadratically with data size. The authors propose a survey sampling–based estimator that directly samples pairs with carefully designed inclusion probabilities, leveraging auxiliary information to assign higher sampling weights to more informative pairs. This approach substantially reduces computational overhead while preserving optimization performance. Theoretical analysis establishes a precise trade-off bound between estimation accuracy and computational efficiency. Empirical results demonstrate that, in high-dimensional embedding tasks such as visual representation learning and graph learning, the method achieves performance comparable to full pairwise computation using only a small fraction of sampled pairs, thereby validating its effectiveness and scalability.
Standard image classifiers employing global average pooling (GAP) discard spatial information, making it difficult to localize class-discriminative evidence in multi-object scenes. This work reveals, for the first time, that the conventional architecture—comprising GAP followed by a linear classification head—inherently exhibits multi-instance learning (MIL) characteristics, naturally treating an image as a bag of spatial instances. Leveraging this insight, we propose a post-hoc method that requires no model modification and recovers localized class evidence obscured by pooling through predictive grid decomposition. Experiments demonstrate that our approach effectively reconstructs faithful foreground responses using off-the-shelf classifiers and further uncovers that classification failures often stem from the intrinsic limitations of mean aggregation inherent in GAP.
Existing self-supervised learning methods exhibit limited performance on link prediction tasks in graphs without node attributes. This work proposes the first self-supervised learning framework centered on link representations, introducing a link-level contrastive learning mechanism that integrates instance discrimination with a community structure-aware graph augmentation strategy. The proposed models, L-GRACE and L-BGRL, significantly outperform current state-of-the-art approaches under both self-supervised and supervised settings, achieving particularly strong results on attribute-free graphs. These empirical gains validate the effectiveness of link-centric representation learning and structure-aware augmentation for improving link prediction performance in the absence of node features.
This work reveals that the often-overlooked embedding norm in contrastive learning inherently encodes critical semantic information, such as semantic specificity. From the perspective of optimization dynamics, we theoretically demonstrate—for the first time—the mechanism by which embedding norms naturally capture semantic attributes during training under scale-invariant losses. We derive analytical relationships between the norm and established semantic metrics, including concept specificity, word frequency, and human uncertainty. Furthermore, we show that the norm serves as a calibration signal without requiring additional training, offering significant practical utility in retrieval and confidence calibration tasks.