perform long-horizon planning

Designs, implements, and evaluates planners and training regimes that produce and analyze multi-step action sequences over extended temporal horizons, including algorithms for adaptive horizon selection, look-ahead / prediction-horizon tuning, entropy-gated MCTS node expansion, dual- or cross-horizon feedback, and mechanisms for deliberate periodic or cross-episode goal evolution. Builds procedures and benchmarks to stabilize long-term rollouts and temporal consistency (e.g., ensuring physical plausibility across steps), to optimize horizon and expansion policies, and to measure and diagnose multi-step tool invocation, retrieval-limited planning failures, and other performance trade-offs under different horizon settings.

performlong-horizonplanning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$230K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the critical yet poorly understood role of long-horizon multi-turn planning in foundation model agents, which is hindered by the uncontrolled nature of internet-scale pretraining data. The authors construct a unified and controllable multi-turn environment to systematically investigate how the format, distribution, and quality of pretraining data influence planning capabilities. During post-training, they introduce GRPO, Online Policy Distillation (OPD), and Multi-teacher Online Policy Distillation (MOPD) to shape and integrate planning skills. Their findings reveal the essential role of explicit world models in enabling long-horizon generalization, demonstrate OPD’s superiority over GRPO under low-quality long-horizon data, and present the first successful fusion and transfer of cross-environment planning abilities via MOPD. Experiments further validate the efficacy of limited high-quality data and elucidate MOPD’s robust generalization, continual learning, and interference resilience across compatible, partially shared, or conflicting planning paradigms.

foundation model agentslong-horizon planningmulti-turn planning

Improving planning and MBRL with temporally-extended actions

May 21, 2025
PC
Palash Chatterjee
🏛️ Indiana University

Discrete-time modeling of continuous-time systems in model-based reinforcement learning (MBRL) and trajectory optimization leads to excessive planning steps, high computational overhead, and cumulative model prediction errors. Method: We propose the Temporal-Expansion Action (TEA) framework, which treats action duration as an optimizable variable to explicitly control the decision-making timescale. TEA is the first approach to jointly optimize action duration within MBRL and trajectory optimization; it employs a multi-armed bandit to adaptively select duration ranges and decouples deep primitive-action horizons from shallow planning depths. Contribution/Results: Experiments demonstrate that TEA significantly accelerates planning, improves solution quality, resolves convergence failures of standard methods on multiple benchmark tasks, reduces model training time, and mitigates error accumulation—thereby enhancing both efficiency and robustness of continuous-control planning.

Addressing computational demands in discrete-time planning with continuous systemsImproving MBRL performance by minimizing compounding model errorsReducing planning horizon by optimizing action durations directly

Existing reinforcement learning agents often overfit to idiosyncratic patterns in closed environments and lack verifiable behavioral generalization. This work proposes the first cross-domain, long-horizon, multi-tool post-training framework, built upon the open-source MoE model Qwen3.5-122B-A10B and combining two-stage supervised fine-tuning (SFT) with reinforcement learning (RL). Training is conducted on 363 tasks across 27 categories within the MCP benchmark, strictly isolating external evaluation tasks and reward signals. Experimental results demonstrate that the proposed approach substantially enhances out-of-distribution transfer performance, achieving consistent gains across five external benchmarks—including Toolathlon (+9.6 percentage points) and τ²-Bench (+5.3 pp)—and even improves performance on SWE-Bench Pro and Terminal-Bench 2 despite the absence of software engineering tasks in training. The study further uncovers four consistent cross-scenario behavioral divergence patterns.

behavioral evaluationcross-benchmark generalizationlong-horizon agents

This study addresses the challenge that large language models often fail in multi-hour long-horizon tasks due to a “horizon gap,” leading to forgetting early decisions, premature termination, or goal drift. Through a systematic review of 1,547 papers, it introduces the first cross-classification framework based on task lifecycle phases—planning, memory, execution, training, and evaluation—and the locus of information representation. This framework disentangles the commonly conflated notions of task length, context window size, and long-term memory, while highlighting the critical role of process signals. By integrating systematic data collection, leakage filtering, and cross-dimensional categorization with process rewards, credit assignment, and trajectory diagnostics, the work demonstrates that reliance solely on outcome signals inevitably fails as task horizons extend. It concludes by identifying three key open problems: capability decoupling, bias management in process signals, and the development of reliability theory for long-horizon tasks.

horizon gaplarge language modelslong-horizon tasks

Closed-Loop Long-Horizon Robotic Planning via Equilibrium Sequence Modeling

Oct 02, 2024
JL
Jinghan Li
🏛️ Peking University | China Tower

Long-horizon task planning for autonomous robots faces challenges in reliably translating high-level instructions into executable action sequences, while existing language-model-based agents suffer from limited foresight and error-proneness. Method: This paper proposes a closed-loop self-correcting planning framework grounded in balanced sequence modeling. Contribution/Results: Its core innovations are (1) the first end-to-end differentiable self-correction mechanism—requiring no external verifiers or reward models—and (2) a nested balanced sequence modeling architecture that integrates environmental feedback for efficient closed-loop iterative optimization. Trained via supervised end-to-end learning, the framework achieves significant improvements in long-horizon planning accuracy and reasoning scalability on the VirtualHome-Env benchmark, outperforming all state-of-the-art language-model agents across key metrics.

Addressing long-horizon robotic task planning errorsEnhancing autonomous robot action sequence planningImproving closed-loop planning with environmental feedback

Latest Papers

What's happening recently
View more

This work addresses the challenge in long-horizon planning where the exponential growth of the action sequence search space with planning horizon leads to poor-quality candidates, limiting the performance of latent-variable world models. To overcome this, the authors propose a multi-scale subgoal-conditioned planning method that decomposes long-term goals into achievable latent subgoals at varying temporal granularities, thereby guiding structured action generation. Leveraging a frozen world model for sequence evaluation and refinement, the approach replaces random initialization with semantic subgoal priors. Evaluated on PushT and OGBench Cube tasks under a 150-unit goal displacement, the method significantly improves success rates—from 12.7% to 64.7% and from 26.7% to 67.3%, respectively—demonstrating a balanced capability in both precise local control and effective long-horizon goal progression.

action proposallatent world modelslong-horizon planning

This work addresses the limitation of conventional reinforcement learning methods that employ a fixed discount factor, which often struggle to balance short-term and long-term rewards effectively in dynamic environments. To overcome this, the paper proposes an adaptive multi-horizon reinforcement learning approach that dynamically selects and fuses discounting strategies across multiple time scales. This enables the agent to automatically adapt to shifts in reward structures and task transitions without requiring manual hyperparameter tuning. Evaluated in continuous-task MiniGrid environments, the method demonstrates significantly improved parameter efficiency and environmental adaptability. It achieves this by online identification and composition of optimal discount factors, thereby enhancing agent performance in continual learning scenarios.

adaptive decision-makingcontinual learningmulti-horizon

This work addresses a critical limitation in existing automated scientific discovery systems, which rely on myopic information-gain strategies and fail to evaluate the long-term value of constructive actions—such as developing new instruments—in problems requiring chains of capabilities to achieve a goal. The authors formalize goal-directed scientific discovery as a stochastic shortest path problem in belief space, where constructive experiments dynamically expand the action space. They propose CG-Plan, an incremental replanning algorithm that integrates a capability-aware heuristic combining capability acquisition (h_cap) and experimental progress (h_exp). Theoretical analysis introduces “capability gating” as a novel dimension of problem hardness, proving that any fixed-horizon myopic planner suffers either unbounded approximation ratios or incompleteness in such settings. Experiments demonstrate that CG-Plan substantially outperforms myopic baselines in capability-gated scenarios, with consistent performance advantages across all fixed planning horizons.

capability-gated planningconstructive actionsmyopic experiment selection

Existing latent world models rely on single-step prediction, leading to error accumulation during recursive rollout in long-horizon planning and suffering from a mismatch between their training objective and the actual planning task. This work proposes the Variable-Length World Model (VLWM), which introduces, for the first time, a mechanism for predicting future latent states based on variable-length action sequences. VLWM employs a curriculum learning strategy that progressively optimizes the model from short- to long-horizon predictions and is accompanied by a tailored latent-space planning algorithm. Evaluated across multiple long-horizon control tasks, VLWM outperforms the current state-of-the-art method, LeWM, by an average of 13%, with particularly pronounced gains in tasks requiring extended planning horizons.

action-conditioned predictioncompounding errorslatent world models

This work addresses a critical gap in existing robotic benchmarks, which overlook the challenge of long-horizon tasks requiring agents to continuously track state evolution driven jointly by their own exploration and environmental dynamics. To tackle this, the paper introduces “Task State Horizon” (TSH) as a novel dimension for quantifying task difficulty and presents RoboGraph, a compiler that automatically translates state-transition dependencies into executable symbolic task graphs grounded in spatiotemporal causal relationships—including failures and interventions. The framework enables structured evaluation of agents’ state-tracking capabilities and is accompanied by a benchmark dataset comprising 84 scenarios and 588 episodes. Experiments reveal that 15 state-of-the-art agents exhibit significant performance degradation on high-TSH tasks, exposing fundamental limitations in their ability to maintain, explore, and update task states.

embodied agentslong-horizon tasksrobotic benchmarking

Hot Scholars

AS

Aditya Shirwatkar

PhD Student, IISc Bangalore
RoboticsRobot LearningLegged Locomotion
SL

Sergey Levine

UC Berkeley, Physical Intelligence
Machine LearningRoboticsReinforcement Learning
SZ

Shiyi Zhang

Tsinghua University
Video GenerationVideo Understanding
CF

Chelsea Finn

Stanford University, Physical Intelligence
machine learningroboticsreinforcement learning