Score
Design, build, or analyze policies that perform dual-informed dives (vertical expansions) in tree search: selecting best-bound frontier nodes to start dives, following promising children by leveraging dual-bound information and parent–child locality, pruning branches using incumbent-gap criteria during dives, and deciding when to re-anchor to the best-bound between dives.
This work addresses a key limitation of traditional Conflict-Based Search (CBS)—its fixed best-first node selection strategy, which often leads to frontier explosion, premature interruption of deep search, and an inability to deliver feasible solutions before termination. The paper introduces Dual-Informed Vertical Expansion (DIVE), the first approach that treats node selection as a core design component of CBS. Operating within a branch-and-bound framework, DIVE performs depth-first dives from the current best boundary node, integrating feasible-solution pruning with periodic boundary re-anchoring to balance exploration efficiency and solution quality. Experimental results demonstrate that DIVE substantially reduces dive interruptions, yields feasible solutions with certified optimality gaps earlier than baseline methods, and maintains significantly lower queue memory consumption—advantages that are especially pronounced in dense or memory-constrained scenarios.
This paper investigates the formal relationship between two tree abstraction problems—hard-constrained and soft-constrained—within the information bottleneck (IB) framework. Method: Leveraging Lagrangian duality theory, we derive the necessary and sufficient conditions for their equivalence and uncover an intrinsic connection between the Q-function and the dual function of the hard constraint in tree structures. Integrating insights from tree phase transitions and total unimodularity of constraint matrices, we propose an efficient variable selection method. Through Lagrangian relaxation and strong duality analysis of integer/linear programming formulations, we establish equivalence criteria and characterize the underlying phase-transition mechanism. Results: Empirical validation confirms the effectiveness of dual-variable optimization. Our work provides both theoretical foundations and algorithmic tools for hierarchical information compression and interpretable tree-structured representation learning.
This paper addresses the problem of efficiently locating a target node in an implicit binary tree of height $n$ containing $k$ nodes with two children (all other internal nodes having exactly one child), where node values satisfy inorder traversal ordering. Starting from the root, a player may traverse edges or query an oracle to test whether the current node is the target. We propose *bifurcated exploration*, a novel divide-and-conquer algorithm that jointly exploits structural sparsity and inorder ordering. We prove it requires only $O(sqrt{k} + log n)$ oracle queries—matching the information-theoretic lower bound—and runs in $O(nsqrt{k})$ time, improving upon prior state-of-the-art algorithms achieving $O(sqrt{k}log n)$ queries. Our key innovation lies in leveraging the inorder constraint to enable aggressive path pruning and optimized backtracking, thereby decoupling oracle complexity from the logarithmic dependence on tree height—a first for this model.
This work addresses the high computational overhead and scalability limitations of subgoal-based policy tree search in complex deterministic single-agent tasks, which stem from explicit subgoal generation. To overcome these challenges, the authors propose a learned “rerooter” mechanism integrated with the √LTS algorithm to enable implicit soft subtask decomposition, thereby eliminating the need for explicit subgoal construction and inference and allowing more efficient allocation of search resources. Three rerooter variants are introduced: one leveraging global state structure via clustering, another fusing learned heuristics with cost-to-go estimates, and a hybrid combining both strategies—collectively enabling scalable tree search without handcrafted rerooters for the first time. Experiments demonstrate that the approach significantly outperforms conventional subgoal-based tree search across multiple complex environments, achieving state-of-the-art online training efficiency and successfully scaling to problem sizes previously intractable for existing methods.
This paper addresses complex pure-exploration objectives beyond best-arm identification—such as threshold testing and ε-optimal arm identification—by establishing the first duality-based minimax-optimal sampling allocation framework. It provides the first necessary and sufficient conditions for optimal sampling allocation in pure exploration; generalizes the top-two paradigm to arbitrary pure-exploration problems; and proposes a hyperparameter-free, information-directed selection rule driven by KL divergence and entropy. The rule is rigorously proven to achieve asymptotic optimality in Gaussian settings and resolves the long-standing open problem of asymptotic optimality for top-two Thompson sampling. Experiments demonstrate substantial improvements in sampling efficiency across Gaussian best-arm identification, threshold-bandwidth testing, and ε-optimal arm identification, consistently outperforming state-of-the-art methods.
本文针对预算限制下的智能体搜索问题,提出ExTS树搜索策略,通过奖励塑形、随机虚拟子节点和质量条件分支来优化预算分配。
This work addresses the challenge of effectively navigating multiple uncertain yet plausible reasoning paths in deep search scenarios involving multi-step retrieval and inference. To this end, the authors propose TreeSeeker, a framework that adopts a tree-based branching-and-backtracking search paradigm. TreeSeeker dynamically evaluates the value, uncertainty, and risk of each branch using a textualized UCB (Upper Confidence Bound) signal and incorporates a TreeMem mechanism to store branch-level evidence and failure cues, thereby enabling controlled exploration, exploitation, and pruning. Experimental results demonstrate that TreeSeeker significantly outperforms existing open-source baselines on the XBench-DeepSearch, BrowseComp, and BrowseComp-ZH benchmarks, confirming the efficacy of explicit branch-level control in enhancing deep search performance.
This study addresses the gradient bias in finite-sample policy gradient methods caused by overlooking rare high-reward trajectories. To mitigate this, we propose online parallel tree search and tree trajectory optimization. Leveraging a branch aggregation lemma, our approach enhances trajectory coverage while controlling bias under a fixed computational budget. Furthermore, we introduce an online tree sampling mechanism that eliminates the need for action distribution correction, and formally prove that the expected return increases monotonically with the budget under deterministic dynamics. Empirical evaluations demonstrate that our method improves tail returns by 28.6% over PPO in MuJoCo environments and significantly outperforms strong baselines on both Atari games and large language model reasoning tasks.
This study addresses the statistical asymmetry and selection history bias arising from adaptive expansion in tree-structured reinforcement learning. To mitigate these issues, we propose Selective Reasoning Policy Optimization, a method incorporating scale-invariant branching criteria, exchangeable sampling, and order-statistic correction mechanisms. These components effectively eliminate selection bias in credit estimation, ensuring that leaf-node budget allocation remains consistent with the optimization objective while enhancing fairness in tree search. Extensive experiments conducted on the Qwen model series across seven question-answering benchmarks demonstrate that our approach achieves superior average performance, yielding significant improvements in both single-hop and multi-hop question-answering accuracy.
该研究提出了一种概率焦点搜索(PFS)方法,通过平衡启发式引导与下界推进来加速有界次优搜索,在多个问题上减少了节点扩展数量。