PoAct: Policy and Action Dual-Control Agent for Generalized Applications

📅 2025-01-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language model (LLM) agents commonly suffer from “planning–action mismatch” in complex tasks—i.e., a lack of dynamic alignment between high-level reasoning strategies and low-level tool invocations—leading to poor code-action quality and divergent reasoning paths. Method: This paper proposes a dual-control mechanism jointly governing strategy and action. It introduces a hierarchical coordination architecture: an upper-layer strategy routing module that dynamically selects and adapts reasoning paradigms (e.g., ReAct, Chain-of-Thought), and a lower-layer action-space remapping module that enables differentiable reconstruction of tool calls, augmented by an environment-feedback-driven online switching algorithm. Contribution/Results: The approach decouples the rigid coupling between strategy and action inherent in conventional frameworks. Evaluated on LegalAgentBench, it achieves a 20% accuracy gain, substantially reduces token consumption, and demonstrates strong generalization and cross-model scalability across GPT-4o and GLM-4 series models.

Technology Category

Multiagent Systems: Mechanism DesignPlanning, Routing, and Scheduling: Planning with Language ModelsCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Based on their superior comprehension and reasoning capabilities, Large Language Model (LLM) driven agent frameworks have achieved significant success in numerous complex reasoning tasks. ReAct-like agents can solve various intricate problems step-by-step through progressive planning and tool calls, iteratively optimizing new steps based on environmental feedback. However, as the planning capabilities of LLMs improve, the actions invoked by tool calls in ReAct-like frameworks often misalign with complex planning and challenging data organization. Code Action addresses these issues while also introducing the challenges of a more complex action space and more difficult action organization. To leverage Code Action and tackle the challenges of its complexity, this paper proposes Policy and Action Dual-Control Agent (PoAct) for generalized applications. The aim is to achieve higher-quality code actions and more accurate reasoning paths by dynamically switching reasoning policies and modifying the action space. Experimental results on the Agent Benchmark for both legal and generic scenarios demonstrate the superior reasoning capabilities and reduced token consumption of our approach in complex tasks. On the LegalAgentBench, our method shows a 20 percent improvement over the baseline while requiring fewer tokens. We conducted experiments and analyses on the GPT-4o and GLM-4 series models, demonstrating the significant potential and scalability of our approach to solve complex problems.
Problem

Research questions and friction points this paper is trying to address.

Complex Tasks
Data Organization
Tool Execution Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

PoAct
Code Action Optimization
Dynamic Strategy Adjustment
G
Guozhi Yuan
Central South University
Y
Youfeng Liu
Zhipu AI
J
Jingli Yang
Zhipu AI
W
Wei Jia
Zhipu AI
K
Kai Lin
Amarcredit
Y
Yansong Gao
Tsinghua University
S
Shan He
Beihang University
Z
Zilin Ding
Amarcredit
H
Haitao Li
Tsinghua University