🤖 AI Summary
This study addresses the challenges of quantifying subjective judgments and accommodating significant individual perceptual differences in high-traffic applications by proposing SubJudge, a personalized reasoning framework. Methodologically, it constructs a multidimensional comparison scale that integrates a sparse Elo algorithm with multi-judge voting to optimize offline ranking. By learning relative ordinal relationships through batch preference optimization, the framework reduces comparison complexity to O(NK). During the online phase, it achieves personalized reasoning with O(1) complexity. Experimental results demonstrate that a 9B-parameter model attains performance comparable to frontier large language models while accelerating inference latency by up to 261 times. These findings confirm that SubJudge effectively enables efficient and precise modeling of subjective preferences, offering a scalable solution for real-time personalized AI systems.
📝 Abstract
Subjective judgments are central to many high-traffic applications, but subjective intensity is difficult to quantify and perceptions vary substantially across individuals. To address these challenges, we propose a pairwise comparative framework for multi-dimensional scale construction. By comparing case-person pairs along case and profile dimensions, the framework constructs relative scales that capture both fine-grained intensity and individual variation. To support practical high-traffic deployment, we optimize both offline scale construction and online inference. For scale construction, we combine sparse Elo comparisons with multi-judge voting, cutting the comparison cost from $O(N^2)$ to $O(NK)$ for $N$ objects and a budget of $K$ opponents per object, while limiting reliance on any single judge. For inference, we propose SubJudge, a System One model for personalized scoring with Batchwise Preference Optimization (BPO). Using Bradley-Terry comparisons, BPO trains the model to learn relative orderings, and SubJudge reads a continuous score from digit-token probabilities at the first response position, requiring only one forward pass per criterion and reducing the inference complexity to $O(1)$. Experiments on PluriHarms and iNews show that our 9B models match or surpass the evaluated frontier LLMs on multiple metrics. On the H100 GPU, SubJudge achieves an approximately $1.29\times$ to $261\times$ speedup in mean inference latency over Qwen3.5-9B with different thinking budgets. The code is available at https://github.com/Longchentong/SubJudge.