🤖 AI Summary
This study addresses the challenge that the effectiveness of search mechanisms during the iterative optimization of large language models (LLMs) is highly state-dependent and difficult to coordinate dynamically. To overcome this, we propose a state-conditioned composition paradigm grounded in opportunity loss theory. This method constructs a unified configuration space and employs a bilevel adaptive controller to achieve both the instantaneous dynamic composition of mechanisms and the long-term optimization of strategic preferences. Evaluated across 32 benchmarks, the proposed framework attains an average rank of first place and achieves the highest scores on 23 tasks. These results demonstrate that our approach significantly enhances the robustness and generality of test-time scaling for LLMs, thereby validating its superiority and broad applicability.
📝 Abstract
Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Targeted empirical diagnostics reveal that mechanism effectiveness is highly state-dependent. Motivated by this, we analyze how individual decisions drive final outcomes, decomposing the expected terminal improvement under a shared budget into cumulative decision opportunities minus cumulative selection losses. Guided by this opportunity-loss theoretical foundation, we propose OptiCom, a unified framework that represents LLM-driven optimizers within a shared configuration space: C=(A,Q,O,E,M,S), corresponding to artifact, query, operator, evaluation, memory, and strategy. Operating within this space, a fast LLM-based Optimization Controller dynamically composes immediate mechanisms through structured Action Packages, while a slower Strategy Adapter refines long-term selection preferences, operator weights, and templates based on accumulated trajectory feedback. Comprehensive evaluations across 32 benchmark groups demonstrate the superiority of framework: OptiCom achieves an average Max-score rank of 1.72 among 14 evaluated configurations, securing the top score in 23 groups. Ultimately, these results highlight the broad applicability and high extensibility of OptiCom as a general-purpose paradigm for robust LLM test-time scaling.