SearchJev: A Fast and Calibrated System-1 Model for Search Agents

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high latency and unreliable confidence estimation of generative models in search decision-making by proposing SearchJev, a framework grounded in a dual-system architecture. The method employs non-autoregressive scoring to rapidly filter valid options, delegating complex reasoning to a System 2 module, while introducing soft-label learning for calibration decisions (SLCD) to achieve effective decision calibration. Furthermore, we construct SearchDecision-Bench, a unified benchmark encompassing six task categories. Experimental results demonstrate that the proposed framework accelerates decision-making by 5.3× and search efficiency by 4.7×, reduces calibration error by 74%, and achieves an accuracy of 54%.
📝 Abstract
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.
Problem

Research questions and friction points this paper is trying to address.

Search Agents
Latency
Confidence Calibration
Decision Making
Generative Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

System-1 Model
Soft-Label Learning
Calibrated Decisions
Dual-system Search Agent
SearchDecision-Bench