🤖 AI Summary
This work addresses the discrepancy between high offline metrics and poor online performance of lead-ranking models in CRM systems by proposing SalesLoop, a closed-loop reinforcement learning framework. SalesLoop introduces a performance-aware reward mechanism and a novel Discriminative Group Relative Policy Optimization (Discriminative GRPO) method, which for the first time adapts group relative policy optimization to discriminative ranking models. This enables listwise objective optimization and dynamic policy adaptation under temporal distribution shifts. Experimental results demonstrate that the approach improves NDCG@K and P@K by 7.9% and 15.8%, respectively. A 160-day A/B test shows a significant 4.7%–8.7% increase in cumulative conversion rate, achieves a 44.1% recall rate within the top-10% ranked leads, and enhances conversion rates for high-intent leads by 2.3×.
📝 Abstract
Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often underperform in production. We identify three fundamental gaps responsible for this disconnect: offline-online metric mismatch, pointwise-listwise objective misalignment, and temporal distribution drift. To address these gaps, we propose SalesLoop, a reinforcement learning framework that establishes a closed feedback loop between model predictions and real-world business outcomes. Our approach introduces (1) a performance-aware reward that encodes conversion outcomes weighted by ranking position and conversion velocity, and (2) Discriminative GRPO, a listwise optimization objective that adapts Group Relative Policy Optimization to discriminative ranking models.
SalesLoop improves NDCG@K by +7.9\% and P@K by +15.8\% over the strongest static baseline. A 160-day production A/B test at a New Energy Vehicle manufacturer, spanning 16.5M leads and 280 sales specialists across two provincial markets, validates statistically significant cumulative lift of +4.7\% ($p=0.047$) and +8.7\% ($p=0.002$). In production, the ranking backbone achieves Top-10\% recall of 44.1\% and surfaces high-intent leads at $2.3\times$ the conversion rate of specialist baselines.