🤖 AI Summary
This study addresses the inherent mismatch between the linear structure of video demonstrations and agent contexts, which often causes a disconnect between global planning and local execution details. To overcome this limitation, we propose RPent, a training-free agent framework that recursively organizes linear videos into navigable hierarchical sub-event structures. Through a read-only tool interface, the agent dynamically loads relevant sub-event segments on demand to perform recursive in-context learning from videos, effectively balancing macro-level planning with micro-level manipulation requirements. This approach establishes an efficient closed loop for task planning and execution. Extensive evaluations on the LIBERO benchmark suite demonstrate that our method yields substantial improvements in success rates, achieving up to 96.5%.
📝 Abstract
LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.