ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high inference cost and scalability bottlenecks in large-scale recommendation systems, which persist despite advances in prediction accuracy. The authors propose a request-oriented computation sharing paradigm that defers interaction between user requests and candidate items and executes request-specific computations only once, substantially reducing inference overhead. Key innovations include Generalized Layer Masking (GLM), Deep Cross-Attention (DCA), and In-Kernel Broadcast Optimization (IKBO), which collectively enable isolation and efficient sharing of candidate representations during feature interaction and sequential modeling. Evaluated in a real-world production system, the approach achieves up to a 3× increase in queries per second (QPS) and a 50% throughput gain in typical scenarios, all while maintaining or even improving model quality—evidenced by a 0.5% reduction in LogLoss for short-video ranking.
📝 Abstract
Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this work, we propose Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm that exploits a unique property of recommendation inference: each user request is evaluated against many candidates, while request-side features are shared across candidates. ROCS defers request-candidate interactions as late as possible, isolates candidate-dependent representations, and evaluates substantial portions of the model once per request rather than once per candidate, significantly improving inference efficiency while maintaining or improving prediction quality. To realize this paradigm, we develop Generalized Layer Masking (GLM) to enforce candidate isolation in feature-interaction architectures, and Deep Cross Attention (DCA) to extend request-oriented sharing to sequence architectures. To support efficient GPU deployment, we co-design In-Kernel Broadcast Optimization (IKBO) that significantly accelerates ROCS model execution. Experiments on public benchmarks show that ROCS consistently improves the quality-efficiency tradeoff across recommendation backbones. On production-scale workloads, ROCS achieves up to a 3x QPS improvement on retrieval models without quality degradation and a 0.5% relative LogLoss improvement with a 50% QPS gain on a short-form video ranking model. ROCS has been deployed across large-scale recommendation systems spanning ads and organic surfaces, retrieval and ranking stages, and more than two orders of magnitude in inference complexity, delivering significant online gains at reduced infrastructure cost.
Problem

Research questions and friction points this paper is trying to address.

recommendation systems
inference efficiency
large-scale
compute sharing
quality-efficiency tradeoff
Innovation

Methods, ideas, or system contributions that make the work stand out.

Request-Oriented Compute Sharing
Generalized Layer Masking
Deep Cross Attention
In-Kernel Broadcast Optimization
Efficient Recommendation Inference
🔎 Similar Papers
Yuxin Chen
Yuxin Chen
Meta
Liang Luo
Liang Luo
University of Washington
Systems for Machine LearningComputer SystemsComputer ArchitectureMachine Learning for Systems
B
Buyun Zhang
Meta AI, Menlo Park, California, USA
J
Jian Jiao
Meta AI, Menlo Park, California, USA
Boda Li
Boda Li
ABB Corporate Research Center
Power systemcyber physical system
H
Haoyu Wang
Meta AI, Menlo Park, California, USA
T
Tongyi Tang
Meta AI, Menlo Park, California, USA
A
Ao Cai
Meta AI, Menlo Park, California, USA
Z
Zijian Shen
Meta AI, Menlo Park, California, USA
Z
Zhengkai Zhang
Meta AI, Menlo Park, California, USA
W
Wenyi Xie
Meta AI, Menlo Park, California, USA
R
Ryan Dick
Meta AI, Menlo Park, California, USA
Han Liu
Han Liu
AI Research Scientist at Meta
Security & PrivacyResponsible AIComputer Vision
Neng Shi
Neng Shi
Meta AI
Computer GraphicsMachine LearningData Visualization
B
Bin Yu
Meta AI, Menlo Park, California, USA
J
Jianbo Xiao
Meta AI, Menlo Park, California, USA
S
Shuyao Bi
Meta AI, Menlo Park, California, USA
H
Hongtao Yu
Meta AI, Menlo Park, California, USA
Yuanwei Fang
Yuanwei Fang
Meta
Computer System
Zhuoran Zhao
Zhuoran Zhao
The Hong Kong University of Science and Technology (GZ)
3D Hand Pose EstimationMultimodal Generation
S
Sijia Chen
Meta AI, Menlo Park, California, USA
Y
Yang Chen
Meta AI, Menlo Park, California, USA
S
Shuqi Yang
Meta AI, Menlo Park, California, USA
Qianru Li
Qianru Li
University of California Los Angeles
Mobile systems5G/LTE
Z
Zikun Liu
Meta AI, Menlo Park, California, USA