Score
Designs, implements, and evaluates machine‑learning models and end‑to‑end pipelines that estimate click‑through rate (CTR) and conversion rate (CVR) from user, item, and context features, including feature engineering, model selection, calibration, and serving. Works on related tasks such as CTR/CVR score estimation, bias correction, optimization for ranking or auction objectives, and offline/online evaluation and monitoring.
CVR estimation suffers from sample selection bias (SSB) due to training exclusively on clicked samples, causing inconsistency between training and inference spaces. Existing methods fail to distinguish ambiguous negatives (impressions without clicks) from factual negatives (clicked but non-converting impressions), undermining model robustness. This paper proposes a full-space unbiased CVR modeling framework that, for the first time, explicitly disentangles the semantics of these two negative classes within the entire impression space. Its core innovation is “choral supervision”—a unified mechanism integrating multi-granularity counterfactual objectives, contrastive learning, multi-task collaborative distillation, and consistency regularization to enable discriminative and robust modeling of non-clicked samples. Evaluated on industrial datasets, the method achieves a 1.23% AUC gain and a 2.1% online GMV lift, significantly mitigating SSB while enhancing generalization and deployment stability.
This paper addresses the limitation in click-through rate (CTR) prediction accuracy for search advertising, which stems from insufficient modeling of user behavior sequences. To tackle this, we propose a generative-discriminative collaborative framework. Methodologically, we innovatively integrate generative pretraining—specifically, category-conditioned next-item prediction—into a discriminative CTR model, establishing a two-stage training paradigm: first, large-scale user behavior sequences are leveraged for generative pretraining to capture high-order behavioral patterns; second, the model is fine-tuned end-to-end for the discriminative CTR task. Evaluated on a proprietary industrial-scale dataset and online A/B tests, our approach achieves significant improvements in CTR prediction performance (AUC +0.82%, LogLoss −1.35%). It has been successfully deployed on a leading global e-commerce platform, demonstrating that generative modeling effectively enhances discriminative recommendation tasks.
The CTR prediction field has long suffered from a lack of standardized benchmarks and unified evaluation protocols, leading to irreproducible experiments and incomparable results. Method: We introduce the first open-source, reproducible CTR benchmark platform, systematically re-evaluating 24 state-of-the-art models across five public datasets under consistent preprocessing and evaluation protocols—conducting over 7,000 experiments (>12,000 GPU hours). Contribution/Results: We propose a standardized CTR evaluation paradigm that reveals widespread overestimation of deep model performance differences; we demonstrate that rigorous hyperparameter optimization and fair experimental design are critical for valid comparisons. After thorough tuning, performance gaps among most models significantly narrow. We fully open-source all code, configurations, and results—including training scripts, data pipelines, and evaluation metrics—to advance reproducible research and trustworthy SOTA assessment.
This paper investigates the choice of training objectives in e-commerce recommendation systems: click-through rate (CTR) versus order submission rate (OSR). Using large-scale online A/B tests on an industrial multi-objective recommendation model, we jointly analyze click, add-to-cart, and purchase behavioral data to systematically evaluate the impact of CTR- versus OSR-oriented optimization on GMV, new-item exposure, and user exploration diversity. Results demonstrate that optimizing for OSR increases GMV by over fivefold compared to CTR, without compromising new-item visibility or user behavioral diversity. Feature importance analysis further reveals a fundamental shift in model attention toward conversion-critical signals under OSR optimization. The study provides empirical evidence and methodological guidance for objective design in e-commerce recommender systems, highlighting the superiority of downstream conversion metrics over proximal engagement proxies in driving business outcomes.
To address weak generalization caused by uniform sample training and limited representation capacity due to shared single supervision across multiple encoders in CTR prediction, this paper proposes TF4CTR—a dual-focus framework. Methodologically, it introduces (1) a Sample Selection Embedding Module (SSEM) that enables difficulty-aware sample selection and dynamic encoder assignment; (2) a Dual-Focus Loss (TF Loss) providing hierarchical, sample-level differentiated supervision; and (3) a Dynamic Fusion Module (DFM) enhancing multi-granularity feature interaction modeling. The framework is plug-and-play compatible with existing architectures. Extensive experiments on five real-world datasets demonstrate consistent and significant performance gains over mainstream models—including Wide&Deep, DeepFM, and AutoInt—validating its strong compatibility and superior generalization capability. Code and experimental logs are publicly available.
This work addresses a critical limitation in conventional conversion rate (CVR) prediction models, which treat all clicks as homogeneous events and thereby overlook the heterogeneity of user click intent—leading to underestimation of high-intent clicks and overestimation of low-intent ones. To remedy this, the study introduces a novel click-intent disentanglement framework that leverages interface interaction signals (e.g., click type) as proxy labels for intent, enabling the construction of multiple intent-specific CVR sub-models. The final prediction is dynamically fused based on the estimated intent distribution. The authors further propose an end-to-end consistency constraint and a first-click–last-impression credit assignment mechanism to resolve attribution ambiguity in multi-impression, multi-click scenarios. Online experiments demonstrate near-perfect calibration accuracy across intent segments, a 2.80% uplift in per-click conversion, and a cumulative 0.98% improvement in core business metrics.
This work addresses the challenge that click-through rate (CTR) prediction models often yield unreliable estimates when encountering sparse feature combinations and struggle to dynamically handle instance-level uncertainty during inference. To this end, the authors propose the UTTSI framework, which introduces test-time computation scaling into CTR prediction for the first time. UTTSI employs a dual-signal mechanism—combining model confidence and data frequency priors—to disentangle epistemic from aleatoric uncertainty, enabling adaptive decisions: high-uncertainty instances trigger stochastic feature-path exploration and consistency-based ensemble weighting to enhance predictions, while confident instances bypass additional computation. Notably, UTTSI requires no retraining and is model-agnostic. Evaluated across four datasets and three backbone models, it consistently outperforms training-phase baselines, achieving a statistically significant 5.3% relative CTR gain (p < 0.01) in online A/B tests with only a 2.8× average computational overhead.
Traditional Transformer-based click-through rate (CTR) prediction models suffer from excessive computational and memory costs due to parameter scale expansion, hindering their industrial deployment. This work proposes LoopCTR, which introduces a novel recurrent scaling paradigm: during training, it decouples computation from parameter growth by recursively reusing shared layers, enhanced with a sandwich architecture, hyper-connected residuals, Mixture-of-Experts (MoE), and intermediate supervision; during inference, it achieves high performance without requiring recurrence. Evaluated on three public benchmarks and one industrial dataset, LoopCTR attains state-of-the-art results. Oracle analysis reveals a remaining performance margin of 0.02–0.04 AUC, and models trained with fewer recurrence steps demonstrate an even higher performance ceiling.
This work addresses the challenge of conversion rate (CVR) prediction, which is hindered by extreme data sparsity in user behavior. To overcome this, the authors propose a novel multi-task knowledge transfer architecture featuring a router module that dynamically allocates cross-task knowledge, a receiver module that selectively captures and transforms relevant information, and an enhancement module designed to ensure transferred knowledge positively contributes to the target task. This framework enables efficient, controllable, and beneficial knowledge sharing across tasks. Extensive experiments on multiple public benchmarks demonstrate significant performance gains over state-of-the-art methods. Furthermore, online A/B tests show a 3.93% increase in eCPM, and the approach has been successfully deployed in two major industrial recommendation scenarios.
This work addresses the selection bias and high variance inherent in causal estimation of post-click conversion rates (CVR), which arise from reliance solely on clicked samples. Drawing on semiparametric theory, the authors propose the first doubly robust estimator tailored to chain-structured outcomes and incorporate outcome regularization to enhance numerical stability. The method remains consistent even when nuisance parameters are misspecified and achieves a faster convergence rate than existing approaches. Theoretical analysis further reveals fundamental limitations in naively debiasing composite losses or directly applying standard causal estimators. Empirical results demonstrate that the proposed estimator significantly outperforms current baselines on both synthetic and real-world datasets, exhibiting strong effectiveness and robustness.