Consistent Plan-Act for Long-Horizon Agentic Tasks

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the coordination failures between planning and execution agents in long-horizon tasks caused by inconsistent state cognition. To mitigate this issue, we propose ConPAct, a framework that localizes state conflicts through structured state assertions and programmatic contradiction detection. Furthermore, it integrates consistency-interactive fine-tuning to enable inference-time correction and collaborative training. Experimental results demonstrate that ConPAct improves task success rates on MiniGrid from 38.6% to 54.4% while significantly enhancing cross-environment generalization capabilities. By effectively aligning state representations across heterogeneous agents, this work establishes a novel paradigm for efficient multi-agent collaboration in complex sequential decision-making environments.
📝 Abstract
Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles. To investigate coordination failures in these tasks, we prompt both agents for structured state assertions and compare their reports programmatically to detect explicit contradictions. Our analyses reveal systematic disagreement about the same task-relevant state facts, a phenomenon we term planner-actor state mismatch. We further find that providing agents with task-relevant state information reduces mismatch and improves coordination and task performance. Based on the systematic analysis of the state mismatch, we propose Consistent Plan-Act (ConPAct), which feeds detected contradictions back to both agents to form consistent state interpretations and fine-tunes them on curated consistent interactions for better coordination. ConPAct improves performance across various environments and model configurations, e.g., increasing MiniGrid success rate from 38.6% to 54.4% with GPT-5.6-sol/terra as planner and actor respectively, demonstrating that state consistency can guide both inference-time correction and coordination training.
Problem

Research questions and friction points this paper is trying to address.

long-horizon agentic tasks
planner-actor coordination
state mismatch
multi-agent systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Planner-Actor State Mismatch
Consistent Plan-Act (ConPAct)
Long-Horizon Agentic Tasks
State Assertions
Coordination Fine-tuning
H
Heng-Zhuang Li
School of Artificial Intelligence, Nanjing University; National Key Laboratory for Novel Software Technology, Nanjing University; LongCat Team, Meituan
Yi-Kai Zhang
Yi-Kai Zhang
Nanjing University
Model RecommendationMultimodal Large Language Model
Yu Wang
Yu Wang
University of Science and Technology of China
LLM ReasoningLLM AgentReinforcement Learning
Y
Yueqing Sun
LongCat Team, Meituan
Jiayuan Zhang
Jiayuan Zhang
Beihang University
Federated Learning
Q
Qi Gu
LongCat Team, Meituan
Han-Jia Ye
Han-Jia Ye
Nanjing University
Machine LearningData MiningMetric LearningMeta-Learning