Action Chunking Proximal Policy Optimization with Feedback Correction

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of value function learning caused by action chunking and the lack of feedback during open-loop execution in high-dimensional robotic control. To this end, we propose the ACPPO algorithm. This method retains a standard state-value critic to circumvent the learning bottlenecks associated with chunked Q-functions while introducing a step-wise feedback correction mechanism that adjusts actions online during chunked execution, thereby achieving closed-loop control. Built upon the PPO framework, ACPPO integrates chunked planning with local feedback correction, significantly enhancing reactivity in contact-rich tasks. Comprehensive evaluations across 25 simulation tasks demonstrate that ACPPO achieves superior overall performance, validating the effectiveness of moderate chunk lengths and correction regularization.
📝 Abstract
Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as action dimensionality and chunk length grow. Second, executing chunks open-loop removes within-chunk feedback, limiting reactivity in contact-rich tasks. We present Action Chunking PPO (ACPPO), a PPO extension that uses a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online within each chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregate performance among evaluated methods and performs best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that corrector regularization is important for balancing chunk-level planning with local feedback. These results suggest that action chunking can be effective in online PPO when chunk-level planning is paired with closed-loop correction. The code is available at: https://github.com/hshhahn/ACPPO.
Problem

Research questions and friction points this paper is trying to address.

Action Chunking
Reinforcement Learning
High-dimensional Robotic Control
Value Function Estimation
Open-loop Execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Action Chunking
Proximal Policy Optimization
Feedback Correction
Reinforcement Learning
Robotic Control
🔎 Similar Papers
No similar papers found.