Score
Designs, trains, and evaluates predictive models that estimate the probability a user will click a presented item or link (click-through rate/CTR) from observed features and interaction history; this includes pointwise CTR estimation, click-through prediction for ranking, and analysis of clickstream sequences for feature engineering and behavioral patterns. The work covers model selection and validation, calibration or bias-correction of propensity estimates, and producing scores or forecasts usable in online or batch decision systems.
This paper addresses the limitation in click-through rate (CTR) prediction accuracy for search advertising, which stems from insufficient modeling of user behavior sequences. To tackle this, we propose a generative-discriminative collaborative framework. Methodologically, we innovatively integrate generative pretraining—specifically, category-conditioned next-item prediction—into a discriminative CTR model, establishing a two-stage training paradigm: first, large-scale user behavior sequences are leveraged for generative pretraining to capture high-order behavioral patterns; second, the model is fine-tuned end-to-end for the discriminative CTR task. Evaluated on a proprietary industrial-scale dataset and online A/B tests, our approach achieves significant improvements in CTR prediction performance (AUC +0.82%, LogLoss −1.35%). It has been successfully deployed on a leading global e-commerce platform, demonstrating that generative modeling effectively enhances discriminative recommendation tasks.
To address data sparsity and selection bias in conversion rate (CVR) prediction—arising from exclusive reliance on clicked samples—this paper proposes a counterfactual CVR prediction method grounded in structural causal models (SCMs). We construct a user-behavior causal graph and introduce a hypothetical intervention mechanism to generate credible counterfactual conversion labels for non-clicked samples, enabling full-space modeling. Crucially, we replace heuristic rules with principled causal inference, eliminating ad-hoc assumptions. Furthermore, we integrate multi-task learning to enhance label reliability. Extensive experiments on multiple public benchmarks demonstrate significant improvements over state-of-the-art methods. Online A/B tests show substantial gains in the joint CTR+CVR metric. Notably, our approach exhibits superior robustness and generalization in implicit conversion scenarios, where conversion signals are weak or unobserved.
Existing click model taxonomies are outdated, hindering unified evaluation of probabilistic graphical models (PGMs) and neural networks (NNs), and failing to accommodate modern interfaces such as carousels—thereby constraining model innovation. To address this, we reconstruct the foundational theory of click model design and propose three mathematically grounded core design choices. Based on these, we introduce the first unified taxonomy that jointly supports both PGMs and NNs across single-list, grid, and carousel interfaces. This taxonomy overcomes traditional classification limitations by enabling systematic, cross-model-type and cross-interface categorization. It integrates probabilistic modeling, deep learning, and statistical behavioral analysis to support interpretable and scalable click behavior modeling. Furthermore, leveraging this framework, we derive a novel model design paradigm specifically tailored for carousel interfaces—providing both theoretical foundations and methodological pathways for click modeling in emerging interactive interfaces.
Existing CTR prediction models predominantly rely on explicit feature interactions over ID embeddings, often leading to embedding dimension collapse and information redundancy. To address this, we propose the Supervised Feature Generation (SFG) framework—the first to shift CTR modeling from discriminative *feature interaction* to generative *feature generation*. SFG employs an encoder-decoder architecture to implicitly capture high-order feature relationships within the ID embedding latent space, using click labels as supervision signals to guide feature reconstruction. A novel supervised reconstruction loss is introduced to significantly enhance feature discriminability. The framework is plug-and-play and seamlessly integrates with mainstream models—including DeepFM, DCN, and xDeepFM—without architectural modification. Extensive experiments on benchmark datasets (Criteo, Ali-CCP) demonstrate consistent AUC improvements of 0.5–1.2%. The implementation is publicly available.
The CTR prediction field has long suffered from a lack of standardized benchmarks and unified evaluation protocols, leading to irreproducible experiments and incomparable results. Method: We introduce the first open-source, reproducible CTR benchmark platform, systematically re-evaluating 24 state-of-the-art models across five public datasets under consistent preprocessing and evaluation protocols—conducting over 7,000 experiments (>12,000 GPU hours). Contribution/Results: We propose a standardized CTR evaluation paradigm that reveals widespread overestimation of deep model performance differences; we demonstrate that rigorous hyperparameter optimization and fair experimental design are critical for valid comparisons. After thorough tuning, performance gaps among most models significantly narrow. We fully open-source all code, configurations, and results—including training scripts, data pipelines, and evaluation metrics—to advance reproducible research and trustworthy SOTA assessment.
This work addresses the limitations of traditional click-through rate (CTR) prediction models, which often lose fine-grained information when aggregating user behavior sequences and struggle to effectively capture intricate interactions between behaviors and contextual features. To overcome these challenges, we propose CDNet, a novel architecture featuring a dual-perspective collaboration mechanism. CDNet simultaneously identifies the most relevant historical behaviors through a core behavior selection module for fine-grained feature interaction and models the global interest distribution to provide coarse-grained compensation. This design balances local behavioral details with overall user interests while maintaining computational efficiency, thereby significantly enhancing the interaction modeling between sequential behaviors and contextual features. Extensive experiments on multiple real-world datasets demonstrate that CDNet consistently outperforms state-of-the-art methods, achieving substantial improvements in CTR prediction accuracy.
This work addresses the limitations of existing CTR prediction methods, which are prone to overfitting dominant features in user history, struggle to capture rapidly evolving immediate intent, and employ pointwise ranking that neglects the global context of the retrieved candidate set—often causing long-term preferences to overshadow short-term interests. To overcome these issues, the authors propose GenCI, a novel framework that introduces a generative interest grouping mechanism to explicitly model candidate-agnostic immediate intent through a next-item prediction (NTP) objective. Furthermore, a hierarchical candidate-aware network is designed, leveraging cross-attention to inject group-level semantic context into the ranking stage. This enables end-to-end joint optimization of intent generation and ranking, transcending conventional discriminative pointwise paradigms. Extensive experiments on three mainstream datasets demonstrate significant CTR improvements, validating the framework’s effectiveness in dynamic interest modeling and contextual utilization of retrieval candidates.
This study addresses the diminishing returns observed when merely increasing model interaction capacity for click-through rate (CTR) prediction. To overcome this limitation, we introduce the concept of estimator scaling and propose RECAP, a novel framework featuring a parameter-efficient recursive estimator scaling mechanism. By integrating recurrent neural networks with weight-sharing techniques, RECAP efficiently consolidates multi-source diverse features across three dimensions: knowledge distillation, exponential moving average, and inference pathways, thereby transcending the capacity bottleneck of single models. Extensive experiments demonstrate that RECAP establishes new state-of-the-art performance across multiple benchmarks, achieving an exceptional trade-off between predictive accuracy and parameter efficiency.
Existing click models are primarily designed for single-column ranked lists and struggle to capture the complex browsing and clicking behaviors of users in multi-column, horizontally scrollable carousel interfaces. This work proposes three novel position-dependent click models, among which OEPBM is the first to directly leverage eye-tracking data to construct observable exposure signals without relying on latent variables. Through a systematic comparison of optimization approaches—including gradient-based methods, EM, and maximum likelihood estimation—experiments demonstrate that OEPBM achieves superior click prediction performance and its inferred exposure patterns align most closely with actual user behavior. These findings underscore a fundamental limitation of conventional click-based models: they cannot accurately reflect users’ true examination processes when eye-tracking evidence is absent.
This work addresses a critical limitation in conventional conversion rate (CVR) prediction models, which treat all clicks as homogeneous events and thereby overlook the heterogeneity of user click intent—leading to underestimation of high-intent clicks and overestimation of low-intent ones. To remedy this, the study introduces a novel click-intent disentanglement framework that leverages interface interaction signals (e.g., click type) as proxy labels for intent, enabling the construction of multiple intent-specific CVR sub-models. The final prediction is dynamically fused based on the estimated intent distribution. The authors further propose an end-to-end consistency constraint and a first-click–last-impression credit assignment mechanism to resolve attribution ambiguity in multi-impression, multi-click scenarios. Online experiments demonstrate near-perfect calibration accuracy across intent segments, a 2.80% uplift in per-click conversion, and a cumulative 0.98% improvement in core business metrics.