🤖 AI Summary
This study addresses the challenges of unified prompt standard interference and high post-training update costs in industrial search evaluation by proposing a dynamic routing framework that externalizes evaluation criteria into an evolvable skill library. Methodologically, it constructs a gating-based skill repository capable of assimilating new standards via a replay mechanism without model retraining. Furthermore, a two-stage training pipeline coupled with a listwise evaluation architecture is designed to enable page-level precise judgment and fine-grained attribution. The proposed approach significantly enhances both the accuracy and diagnostic capability of short-video search evaluation. It has been deployed at scale within Kuaishou, effectively optimizing the quality and efficiency of online evaluation systems.
📝 Abstract
Search quality evaluation provides essential supervision and diagnostic signals for the development and iteration of industrial search systems. Although large language models (LLMs) offer a scalable alternative to manual assessment, reliable automatic evaluation remains challenging: users experience search results at the page level, while the applicable evaluation criteria are multi-dimensional and continuously evolving. Packing all evaluation criteria into a unified prompt introduces irrelevant context and potential criterion interference, whereas internalizing them through post-training tightly couples rule updates with costly model retraining cycles.
To address these issues, we propose Skill-routed Evaluation with Evolvable Knowledge (SEEK). Specifically, SEEK externalizes specific search evaluation criteria into a skill bank, dynamically routes relevant skills for each query-result list pair, and employs a task-adapted listwise evaluator to produce page-level judgments and failure mode attribution. A two-stage training pipeline teaches the evaluator to align evaluation criteria with human preferences, while a replay-gated skill bank allows recurring evaluation knowledge gaps to be incorporated without model retraining. Experiments on industrial short-video search show that SEEK improves listwise quality evaluation accuracy and achieves significant progress in attribution diagnosis. SEEK has been deployed at Kuaishou, a short-video platform with over 400 million daily active users, significantly improving the scale and quality of online search evaluation.