Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the diminishing returns observed when merely increasing model interaction capacity for click-through rate (CTR) prediction. To overcome this limitation, we introduce the concept of estimator scaling and propose RECAP, a novel framework featuring a parameter-efficient recursive estimator scaling mechanism. By integrating recurrent neural networks with weight-sharing techniques, RECAP efficiently consolidates multi-source diverse features across three dimensions: knowledge distillation, exponential moving average, and inference pathways, thereby transcending the capacity bottleneck of single models. Extensive experiments demonstrate that RECAP establishes new state-of-the-art performance across multiple benchmarks, achieving an exceptional trade-off between predictive accuracy and parameter efficiency.
📝 Abstract
Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction capacity remains the most effective way to improve predictive performance, and find that its benefits quickly exhibit diminishing returns even as capacity continues to grow. This motivates a complementary scaling direction that we call estimator scaling, where additional resources are used to incorporate multiple related estimators rather than only enlarging a single predictor. Through theoretical analysis, we show that the gains from estimator scaling are governed by the amount of non-shared predictive variation available across estimators. However, exploiting this variation naively can be expensive: independently trained models provide substantial estimator diversity but require deployment cost to grow with ensemble size. This motivates a parameter-efficient realization of estimator scaling that can incorporate diversity from multiple estimator sources without maintaining multiple full models. Building on this view, we introduce RECursive Averaged Predictor (RECAP), a parameter-efficient recursive CTR model that operationalizes estimator scaling at three levels: distillation across independently trained models, exponential moving averaging over training trajectories, and aggregation over inference-time routes within a weight-shared recursive backbone. Experiments across multiple benchmarks establish new state-of-the-art predictive performance on standard benchmarks, while placing the RECAP on a favorable performance-parameter Pareto frontier.
Problem

Research questions and friction points this paper is trying to address.

Click-Through Rate prediction
interaction capacity
estimator scaling
parameter efficiency
diminishing returns
Innovation

Methods, ideas, or system contributions that make the work stand out.

Estimator Scaling
CTR Prediction
Parameter-Efficient Recursive Model
RECAP
Knowledge Distillation
🔎 Similar Papers
No similar papers found.