Score
Designs and implements methods and pipelines that select, generate, or weight informative negative examples for training and evaluation, including hard-negative mining, cluster-aware and prototype-based sampling, multi-source (peer/teacher/external) negatives, and policies to balance positive/negative ratios and reduce duplicate or variant predictions. Builds the accompanying loss functions, thresholds and prioritization rules (e.g., dual-confidence thresholds, difficulty-aware or hard-example mining losses), clustering and prototyping procedures, and sampling/prioritization workflows for unlabeled pools or manual labeling to improve discriminator robustness and coverage while avoiding self-reinforcement.
This work addresses the limitations of conventional negative sampling strategies in two-tower model training, which often yield easy negatives that hinder discriminative learning and exacerbate popularity bias and feedback loops. To overcome these issues, the authors propose the first large language model (LLM)-based framework for real-time hard negative sampling. During training, the LLM performs semantic clustering and dynamically generates challenging yet relevant negatives from within semantically similar clusters. This approach significantly enhances the representation learning capability of two-tower models, effectively mitigates popularity bias, and breaks detrimental feedback cycles—all while maintaining low computational overhead. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art industrial negative sampling methods on both public benchmarks and a billion-scale production system, yielding substantial gains in retrieval performance.
This work addresses the limitations of traditional hard negative mining—such as insufficient corpus coverage, retriever scoring bias, and false positive interference—and the performance degradation often caused by negatives directly generated by large language models due to misalignment between generation and discrimination objectives. The authors propose CausalNeg, a novel framework that formally characterizes the generation-discrimination gap for the first time. It generates hard negatives of controllable difficulty through causal-guided counterfactual perturbations and mitigates source dependency via a query-perspective entropy maximization strategy. Integrating chain-of-thought reasoning, counterfactual data augmentation, and contrastive learning, CausalNeg enables interpretable, shortcut-free negative synthesis, significantly boosting retrieval performance across multiple benchmarks while effectively avoiding the performance drop commonly associated with generated negatives.
In knowledge graph link prediction, low-quality negative samples severely hinder the performance of embedding models. To address this, we propose EMU, a novel framework that theoretically derives a sufficient condition for the effectiveness of negative sample distributions. EMU shifts from conventional sampling to a *generative* negative sampling paradigm: guided by optimal embedding conditions, it actively constructs negative triples via embedding perturbations—including directional disturbance and norm scaling—that provably satisfy theoretical optimality. EMU is model-agnostic and seamlessly integrates with any knowledge graph embedding (KGE) model (e.g., TransE, ComplEx, RotatE) and existing sampling strategies, supporting end-to-end joint training. Extensive experiments on standard benchmarks (FB15k-237, WN18RR) demonstrate significant improvements in MRR and Hits@1—equivalent to increasing embedding dimensionality fivefold. The code is publicly available.
This work addresses the challenge of retrieving scientific literature sources from multilingual social media short texts, where performance is often degraded by semantically similar distractors. To mitigate this issue, the authors propose a clustering-aware, staged hard negative mining approach that integrates dense retrieval, multilingual cross-encoder reranking, and large language model (LLM)-based evidence selection. By leveraging the semantic cluster structure of the candidate pool, the method distinguishes between local cluster negatives and global semantic negatives to construct stage-aware hard negatives. Additionally, constrained prompting strategies—such as constrained classification—are designed to enhance the reliability of LLM-based evidence selection. Evaluated on the CheckThat! 2026 shared task, the proposed method achieved 6th place out of 37 participating teams, demonstrating significant improvements in cross-lingual retrieval and reranking performance.
Recommender systems face persistent challenges including filter bubbles, sparse user-item interactions, cold-start problems, and feedback loops. Existing approaches predominantly leverage positive behavioral signals while underutilizing the critical role of negative feedback in preference modeling. This survey establishes, for the first time, a taxonomy of negative sampling methodologies—categorizing over one hundred works into five paradigmatic classes. It further introduces a scenario-aware adaptation framework that elucidates the pivotal roles of negative sampling in modeling dynamic preferences, mitigating filter bubbles, and alleviating feedback bias. Finally, it identifies three emerging research frontiers: enhancing interpretability, integrating causal inference, and synergizing with large language models. By rigorously delineating the theoretical boundaries and practical implementation pathways of negative sampling, this work elevates it from an empirical engineering heuristic to a foundational paradigm in recommender system design.
Current out-of-distribution (OOD) detection methods based on vision-language models suffer from performance limitations due to false negatives introduced during negative sample mining. This work proposes a debiased negative sampling mechanism that, for the first time, formulates the correction of negative label distribution bias as a tractable Monte Carlo sampling process. By leveraging labeled in-distribution data and unlabeled wild corpora to approximate the true negative label distribution, the approach effectively mitigates the false negative problem. Extensive experiments demonstrate that the proposed method consistently outperforms existing techniques across diverse OOD detection settings, establishing new state-of-the-art performance.
This work addresses the performance limitations of implicit-feedback recommender systems caused by mislabeling unobserved interactions as negative samples—so-called false negatives. To mitigate this issue, the authors propose a novel Corrected and Weighted (CW) loss function that theoretically derives the true distribution of negative samples to correct for sampling bias, while simultaneously incorporating a dynamic reweighting mechanism based on the model’s prediction confidence. This approach effectively alleviates the false-negative problem without altering the sampling procedure or incurring significant computational overhead. Extensive experiments on four large-scale sparse datasets demonstrate that the CW loss consistently outperforms state-of-the-art loss functions across multiple ranking metrics, yielding substantial and robust improvements in recommendation quality.
This work addresses the challenge of learning from highly imbalanced data, where positive instances are not only scarce but also resemble negative examples, making them difficult to identify. To tackle this issue, the authors propose a focused empirical risk estimator that, for the first time within the positive-unlabeled (PU) learning framework, effectively handles both extreme class imbalance and hard-to-discriminate positive samples. The method is compatible with both SCAR (Selected Completely At Random) and SAR (Selected At Random) labeling assumptions and jointly models positive and unlabeled data to significantly enhance generalization under sparse annotation scenarios. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance across multiple imbalanced benchmarks and successfully applies to real-world tasks such as financial misstatement detection.
Existing graph contrastive learning methods rely on static negative sampling, which struggles to dynamically balance informativeness and computational overhead. This work proposes AdNGCL, a novel framework that introduces, for the first time, a budget-aware, loss-sensitive Hardness-Aware Negative Scheduler (HANS). HANS formulates negative sample selection as a dynamic process governed by loss gating and computational budget constraints, adaptively adjusting sampling strides across hard, medium, and easy negatives while periodically refreshing the pool to preserve diversity. Evaluated on nine benchmark graph datasets, AdNGCL achieves state-of-the-art performance on seven and runner-up results on two, significantly improving accuracy while enabling explicit control over computational cost.
This study challenges the foundational assumption in mainstream machine learning that objective ground-truth labels exist, an assumption often violated in real-world scenarios and leading to inaccurate evaluation and learning. Adopting a negative ontological stance—asserting that no single true label exists—the work introduces, for the first time, this philosophical perspective into machine learning through a democratic supervision framework. It proposes representing each instance with multiple imprecise yet authentic truth labels (MIATTs), accompanied by a logic-driven mechanism for generating and evaluating MIATTs, as well as a truth-learning strategy that operates without relying on a predefined ground truth. The resulting EL-MIATTs methodology is validated in real educational settings, demonstrating not only a departure from the conventional single-ground-truth paradigm but also practical utility in supporting personalized education and career development.