🤖 AI Summary
This study addresses the limitations of numerical embeddings in click-through rate (CTR) prediction, where scalar position and semantics are conflated, leading to training-inference asynchrony. To this end, we propose ScalarLens, a framework that reformulates numerical embedding as a measurement problem. Specifically, it constructs strictly invariant coordinates via monotonic local grids to decouple intrinsic values, and dynamically generates context-aware responses through bounded low-rank dynamics, thereby achieving thorough disentanglement of feature representations. Extensive evaluations comprising 1,539 experiments across multiple backbone networks demonstrate the efficacy of the proposed approach, which achieves state-of-the-art performance in 25 out of 27 experimental settings and significantly outperforms baseline models such as DEER.
📝 Abstract
Numerical embeddings for click-through rate (CTR) prediction are built on a convenient but restrictive premise: a scalar has one representation. This premise conflates where a value lies with what it means for the current sample. On the Criteo validation split, the same numerical interval carries residual click evidence with opposite signs across categorical and numerical contexts, even after additive main effects are removed. Production pipelines compound this mismatch because externally normalized features require transformations and statistics to remain synchronized between training and serving. We introduce ScalarLens, a numerical embedding that preserves what a value is while adapting how it should be interpreted. A monotone local mesh constructs a stable coordinate from the focal scalar alone; bounded low-rank dynamics then produce a contextual response without moving that coordinate or replacing categorical tokens and the CTR backbone. In a 1,539-run primary evaluation covering 19 representations, three datasets, nine backbones, and three seeds, ScalarLens ranks first in 25 of 27 settings on original numerical scales and second in the remaining two. Matched ablations show that scale correction, additional local capacity, and generic conditioning do not reproduce the gain. A controlled study further recovers categorical, numerical, and mixed response mechanisms under context shift while the focal coordinate remains exactly invariant. A complete rerun under shared standardization retains significant advantages over DEER, DAES, and NaryDis, showing that the result is not explained by tolerance to raw scales alone. ScalarLens therefore recasts numerical embedding as a measurement problem: coordinates belong to values, while predictive responses belong to values in context.