Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of static supervision in capturing interactions between query understanding components and search engines, which impedes multi-task optimization. To this end, it proposes a search-aware reinforcement learning framework that pioneers a "distillation-then-RL" paradigm. Specifically, the policy model is initialized via teacher-student supervised fine-tuning, followed by reinforcement learning with real-time search feedback to enable fine-grained reward allocation based on component operational roles, thereby overcoming the limitations of conventional end-to-end outcome-oriented approaches. Experimental results demonstrate that the proposed method significantly enhances both component utility and downstream search quality, achieving an 8.9-point improvement in NDCG@20 over the SFT baseline and a 3.5-point gain compared to single-reward methods.
📝 Abstract
Query understanding (QU) plays a critical role in production search systems, translating raw user queries into search execution plans that drive downstream retrieval and ranking. While large language models (LLMs) have enabled QU to be framed as a structured multi-task generation problem (e.g., intent classification, query expansion), optimizing such models to produce search-engine-coupled outputs remains challenging: static, label-based supervision fails to capture how each component actually interacts with the underlying search pipeline to affect downstream performance. We present a search-aware reinforcement learning (RL) framework for QU based on a distill-then-RL paradigm. Teacher-student supervised fine-tuning (SFT) first yields a well-formed, schema-compliant policy initialization. The RL stage then optimizes each QU component with rewards derived from live interaction with the search engine, tailored to that component's operational role, rather than a single reward tied to the final search outcome. Experiments on Roblox search show that this component-specific optimization improves both per-component utility and downstream search quality, raising NDCG@20 by 8.9 points over the SFT policy and by 3.5 points over training with a single end-to-end reward.
Problem

Research questions and friction points this paper is trying to address.

Query Understanding
Reinforcement Learning
Large Language Models
Search Systems
Multi-Component Optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Search-Aware Reinforcement Learning
Query Understanding
Distill-then-RL
Component-Specific Optimization
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Nayoung Choi
Nayoung Choi
PhD Student @ Emory CS
Natural Language ProcessingInformation Retrieval
S
Shengjian Chen
Roblox Corporation
X
Xiaokai Wei
Roblox Corporation
Wenzheng Zhang
Wenzheng Zhang
Rutgers University
Natural Language ProcessingDeep Learning
D
Daiyao Yi
Roblox Corporation
R
Rachit Pareek
Roblox Corporation
V
Vincent Su
Roblox Corporation
M
Michelle Gong
Roblox Corporation
Jinho D. Choi
Jinho D. Choi
Associate Professor, Emory University
Natural Language ProcessingComputational LinguisticsConversational AI