🤖 AI Summary
This study addresses the inefficiency of Transformers in in-database reasoning caused by the neglect of relational predicates and metadata. To overcome this limitation, we propose a query-context-aware Transformer slicing framework. This approach introduces a pioneering query-granularity routing mechanism that leverages predicate-based preselection of feed-forward network (FFN) experts to enable sparse inference. Furthermore, it incorporates system-level optimizations through an asynchronous CPU-GPU pipeline and routing-aware batching. Evaluations on BERT and Qwen benchmarks demonstrate that the proposed framework achieves up to a 4.42× reduction in latency without compromising predictive accuracy.
📝 Abstract
In-database predictive query processing increasingly applies Transformer-based models within relational pipelines. However, existing in-database inference typically exposes only tuple-level model inputs to the inference runtime, leaving relational predicates and metadata statistics invisible to neural execution planning. In this paper, we propose QCATS, a query context-aware transformer slicing framework that enables efficient sparse inference inside database systems. QCATS executes at query granularity: instead of routing individual tokens or tuples during inference, it uses query predicates and metadata statistics to pre-select context-aligned FFN slices before model execution. The framework comprises offline expert construction and lightweight query-level routing that dynamically selects experts during execution. QCATS further introduces system optimizations, including asynchronous CPU-GPU pipelines and routing-aware batching. Experiments on four predictive-query workloads with BERT-base and Qwen-0.6B show that QCATS achieves up to 4.42x latency reduction while preserving prediction accuracy comparable to dense baselines.