SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking

📅 2026-07-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the discrepancy between high offline metrics and poor online performance of lead-ranking models in CRM systems by proposing SalesLoop, a closed-loop reinforcement learning framework. SalesLoop introduces a performance-aware reward mechanism and a novel Discriminative Group Relative Policy Optimization (Discriminative GRPO) method, which for the first time adapts group relative policy optimization to discriminative ranking models. This enables listwise objective optimization and dynamic policy adaptation under temporal distribution shifts. Experimental results demonstrate that the approach improves NDCG@K and P@K by 7.9% and 15.8%, respectively. A 160-day A/B test shows a significant 4.7%–8.7% increase in cumulative conversion rate, achieves a 44.1% recall rate within the top-10% ranked leads, and enhances conversion rates for high-intent leads by 2.3×.
📝 Abstract
Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often underperform in production. We identify three fundamental gaps responsible for this disconnect: offline-online metric mismatch, pointwise-listwise objective misalignment, and temporal distribution drift. To address these gaps, we propose SalesLoop, a reinforcement learning framework that establishes a closed feedback loop between model predictions and real-world business outcomes. Our approach introduces (1) a performance-aware reward that encodes conversion outcomes weighted by ranking position and conversion velocity, and (2) Discriminative GRPO, a listwise optimization objective that adapts Group Relative Policy Optimization to discriminative ranking models. SalesLoop improves NDCG@K by +7.9\% and P@K by +15.8\% over the strongest static baseline. A 160-day production A/B test at a New Energy Vehicle manufacturer, spanning 16.5M leads and 280 sales specialists across two provincial markets, validates statistically significant cumulative lift of +4.7\% ($p=0.047$) and +8.7\% ($p=0.002$). In production, the ranking backbone achieves Top-10\% recall of 44.1\% and surfaces high-intent leads at $2.3\times$ the conversion rate of specialist baselines.
Problem

Research questions and friction points this paper is trying to address.

lead ranking
offline-online mismatch
temporal distribution drift
CRM systems
ranking performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

reinforcement learning
lead ranking
listwise optimization
performance-aware reward
Discriminative GRPO
🔎 Similar Papers
No similar papers found.
C
Chenyu Zhang
Li Auto Inc.