DEGR: Dual Exploration-Driven Generative Re-Ranking for Adaptive Cross-Request Context Bridging

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing reranking methods, which are constrained by fixed upstream candidate sets and struggle to improve performance when faced with low-quality inputs, while also lacking explicit modeling of exploration value. To overcome these challenges, we propose a dual-exploration-driven generative reranking framework that integrates supervised learning with reinforcement learning. Our approach innovatively combines diversity-aware exploration constraints with an ORPO optimization strategy featuring adaptive reward weighting, and introduces an exploratory reward model capable of dynamically balancing immediate gains against long-term browsing potential. Evaluated on JD.com’s e-commerce recommendation system, the proposed method significantly outperforms current state-of-the-art approaches, with offline and online experiments demonstrating consistent improvements—yielding a 1.22% increase in UCTR and a 0.20% gain in PV.
📝 Abstract
In industrial recommendation systems, the re-ranking stage balances business objectives and diversity for sequence-level optimization while modeling contextual information. However, constrained by fixed upstream supply, existing methods fail to deliver further effectiveness gains, especially under low-quality supply. To overcome this, re-ranking can actively balance immediate and exploratory value, for instance, by prioritizing exploratory exposure under low-quality supply to preserve browsing potential and facilitate serendipitous conversions. Therefore, we propose a Dual Exploration-Driven Generative Re-Ranking (DEGR) method. DEGR adopts a hybrid supervised-reinforcement exploration and optimization paradigm, guided by an exploratory reward model that adaptively balances immediate and exploratory value. The hybrid optimization paradigm integrates three key components: supervised learning, exploration diversity constraint, and adaptive reward-weighted ORPO for preference optimization. Through this dual exploration, the generator ultimately acts as an adaptive cross-request contextual bridge. Offline and online experiments indicate that DEGR outperforms SOTA methods, achieving improvements of up to 1.22% UCTR and 0.20% PV in the JD E-commerce recommendation system.
Problem

Research questions and friction points this paper is trying to address.

re-ranking
exploratory value
low-quality supply
contextual bridging
recommendation systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Re-Ranking
Dual Exploration
Adaptive Reward
Cross-Request Context Bridging
ORPO
🔎 Similar Papers
No similar papers found.