🤖 AI Summary
This study addresses the high latency and unreliable confidence estimation of generative models in search decision-making by proposing SearchJev, a framework grounded in a dual-system architecture. The method employs non-autoregressive scoring to rapidly filter valid options, delegating complex reasoning to a System 2 module, while introducing soft-label learning for calibration decisions (SLCD) to achieve effective decision calibration. Furthermore, we construct SearchDecision-Bench, a unified benchmark encompassing six task categories. Experimental results demonstrate that the proposed framework accelerates decision-making by 5.3× and search efficiency by 4.7×, reduces calibration error by 74%, and achieves an accuracy of 54%.
📝 Abstract
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.