ASARL: Autonomous Social-Aware Relevance Learning for QQ Search

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of general-purpose large language models in social search, where queries and titles often employ informal, community-specific language, leading to performance degradation due to contextual mismatch, data scarcity, and dynamic user behavior. To overcome these challenges, the authors propose ASARL, an end-to-end optimization framework that uniquely integrates multi-agent automated annotation with preference-aligned training. The approach leverages a collaborative triad of agents—ReasonAgent, CriticAgent, and GenAgent—to generate and validate socially aware relevance labels, followed by a three-stage strategy comprising Social Context Training (SCT), Preference-Guided Optimization (PGO), and Social Distillation (SD). Evaluated on the QQ search platform, ASARL significantly improves offline relevance metrics and online user engagement while dramatically enhancing annotation efficiency, thereby enabling interpretable model training and effective data governance in long-tail scenarios.
📝 Abstract
The rapid growth of online social platforms has transformed communication and information retrieval, giving rise to social search, where queries-titles are typically expressed in informal, community-specific language. While large language models provide strong general-purpose semantic understanding, their effectiveness in social search is constrained by contextual discrepancy, data scarcity, and behavior-driven dynamics. To address these challenges, we propose the Autonomous Social-Aware Relevance Learning (ASARL), a fully automated framework that integrates multi-agent data curation with staged model training. ASARL leverages a collaborative agent system: ReasonAgent generates interpretable relevance labels grounded in social attributes, CriticAgent validates and ensures logical consistency, and GenAgent augments long-tail data through synthetic query-title pairs. Building on the curated dataset, ASARL employs three-stage training: Social Context Training (SCT) to capture social language patterns, Preference-Guided Optimization (PGO) to align model predictions with behavioral signals, and Social Distillation (SD) to transfer these improvements into compact models for efficient deployment. Extensive offline and online experiments on the QQ search platform demonstrate significant improvements in both offline relevance metrics and online user engagement indicators, along with enhanced annotation efficiency. These results validate the effectiveness of combining autonomous, socially grounded data governance with preference-aligned training in practical search systems.
Problem

Research questions and friction points this paper is trying to address.

social search
contextual discrepancy
data scarcity
behavior-driven dynamics
informal language
Innovation

Methods, ideas, or system contributions that make the work stand out.

social search
multi-agent system
autonomous data curation
preference-guided optimization
social distillation
🔎 Similar Papers
No similar papers found.