From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how search scale influences the research performance of autonomous agents in quantitative factor mining. Across 50 financial factor tasks, we employ large language model agents, parallel search algorithms, and model grafting techniques to systematically analyze the effects of model capability, search depth, and organizational structure on end-to-end research loop quality. Our findings reveal that parallel search significantly outperforms sequential search, with early-stage states largely determining final outcomes. Moreover, deeper search effectively bridges gaps in model capability, while stronger models demonstrate superior proficiency in diagnosing failures and correcting trajectories. Building on these insights, we propose an adaptive test-time compute strategy and empirically validate a scaling law for search.
📝 Abstract
Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as *search scaling*. Although prior work has characterized the mechanisms, scaling behavior, and performance limits of LLM inference scaling, much less is known about these questions in autonomous research. Therefore, we investigate how search scaling affects research performance and what mechanisms drive these gains using 50 quantitative factor-mining tasks grounded in financial research reports. Each task requires an agent to carry out an end-to-end research loop, from interpreting a hypothesis and implementing it in code to evaluating and iteratively refining the resulting factor. Across nine models, we examine how model capability, search depth, and search organization shape factor quality by tracing performance across varying budgets, transferring intermediate research states between models, and comparing different search strategies. We find that (1) initial performance is more strongly associated with model capability, while deeper search can narrow cross-model gaps; (2) model grafting shows that the early research state materially shapes final performance; and (3) parallel search outperforms sequential search under the same iteration budget, consistent with benefits from broader coverage of the search space. Further trajectory analysis shows that higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates. These findings suggest that future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process.
Problem

Research questions and friction points this paper is trying to address.

Search Scaling
Autonomous LLM Agents
Quantitative Factor Mining
Research Performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Search Scaling
Autonomous Agents
Quantitative Factor Mining
Model Grafting
Parallel Search
💼 Related Jobs
No related jobs found.
K
Kangcheng Deng
StepFun
H
Hui Cai
StepFun
J
Jiacheng Lu
StepFun, Shanghai Jiao Tong University
C
Chester Zhongshu Qian
StepFun, University of California, Los Angeles
R
Rui Sun
StepFun
B
Beidi Luan
StepFun
J
Jing Li
StepFun
Daxin Jiang
Daxin Jiang
Co-Founder & CEO, StepFun Corporation
Deep LearningFoundation Models
Z
Zuo Bai
StepFun, FinStep