Score
Designs and implements algorithms, models, and pipelines that generate ordered sequences of manipulation actions to accomplish multi-step tasks, producing symbolic plans or executable motion/command sequences for manipulators or embodied agents. This work includes methods for discovering or imposing action chunking—grouping primitive actions into reusable higher-level units—sequencing them across task phases, and ensuring temporal and causal coherence for reliable execution.
This paper presents a systematic review of movement primitive approaches in robot control, with a focus on learning from human demonstrations to generate complex action sequences. Integrating chronological and systematic perspectives, it comprehensively traces the theoretical evolution of movement primitives, key technical advances—including spring-damper modeling, probabilistic coupling of multiple demonstration trajectories, and neural network applications in high-dimensional systems—and their empirical effectiveness in tasks such as grasping and throwing. The study offers an in-depth comparative analysis of prevailing frameworks, establishes for the first time a structured developmental trajectory of the field, and clearly identifies current open challenges and practical limitations, thereby providing both theoretical guidance and a practical roadmap for research in robotic motor skill learning.
This work addresses the challenge of action boundary inconsistency in asynchronous robotic execution, where delays in generating action chunks disrupt the continuity of real-time control. To resolve this issue without retraining, gradient computation, or policy modification, the authors propose PAINT—a training-agnostic method that reframes the problem as an initial noise selection task. By leveraging backward Euler inversion and a redrawing rule, PAINT enables an unmodified flow ODE to directly synthesize future action chunks consistent with the already executed action prefix. Evaluated across 12 simulation benchmarks and 6 real-world manipulation tasks, the approach significantly improves action consistency and task success rates, demonstrating broad applicability on single-arm, dual-arm, and humanoid robotic platforms.
Existing approaches lack scalable and verifiable methods for programmatically generating multi-step, high-contact robotic manipulation tasks. Method: This paper proposes a symbolic-physical co-verification framework that enables users to define atomic actions, objects, and spatial predicates; it enforces three complementary constraints—logical consistency checking, object-predicate compatibility analysis, and simulation-based solvability verification—to ensure semantic and physical feasibility of generated tasks. Contribution/Results: The framework generates structured task sequences up to 15 steps long, producing dense semantic task sets annotated with subgoal rewards and state labels—directly usable for reinforcement learning training or high-quality benchmark construction. Experiments yield over one million unique, verifiably solvable tasks, supporting semantic similarity computation and efficient downstream policy learning. The approach significantly improves controllability, formal verifiability, and practical utility in robotic task generation.
This work addresses the “execution gap” between high-level semantic tasks and executable robot motions by introducing Motion Statecharts—a symbolic, executable motion representation that supports concurrency and hierarchical nesting. Coupled with a unified differentiable kinematic world model, this framework enables end-to-end mapping from semantic task specifications to low-level motion control. Smooth and dynamically feasible trajectories are generated through a linear model predictive control (lMPC)-driven task-function approach incorporating snap (jerk derivative) constraints. The proposed system has been successfully deployed across eight heterogeneous robotic platforms, demonstrating strong cross-platform generalization and real-world efficacy. The accompanying software framework, Giskard, has been publicly released.
Traditional hierarchical robotic planning simplifies task-level actions into open-loop kinematic skills, hindering seamless integration of pre-trained closed-loop motor controllers. Method: We propose Composable Interaction Primitives (CIPs), a framework enabling plug-and-play composition of heterogeneous, non-composable pre-trained skills within task-and-motion planning. Building upon CIPs, we introduce Task-and-Skill Planning (TASP), a unified architecture that jointly models symbolic task planning, geometric motion planning, and learned closed-loop control. Contribution/Results: TASP transcends reliance on motion-centric skills by elevating task semantics to the perception–action closed-loop level. Evaluated on a real mobile manipulator, it achieves end-to-end autonomous execution of multi-step complex tasks—including dynamic collaborative transport and tool manipulation—demonstrating significantly improved skill reusability and environmental adaptability.
This paper addresses the fragmentation between task planning and motion control in humanoid robot loco-manipulation. We propose a unified task and motion planning (TAMP) framework that employs contact modes as high-level symbolic representations—enabling, for the first time, fully acyclic, dynamics-driven integrated TAMP. Our method jointly incorporates whole-body dynamics, robot–object–environment contact constraints, and object interaction models, combining high-order trajectory optimization with combinatorial search. Unlike conventional hierarchical approaches, our framework supports physically consistent, long-horizon, multimodal behavior generation. Experimental validation on a real humanoid robot demonstrates autonomous execution of complex, logic-intensive loco-manipulation tasks—including stepping-and-grasping and push-pull transport—over extended durations. The framework significantly improves task adaptability and behavioral consistency while ensuring dynamic feasibility and contact-aware coordination.
This work addresses the limited precision and controllability of existing vision–language–action (VLA) models in interpreting and executing language instructions that involve dense kinematic attributes—such as direction, trajectory, orientation, and relative displacement. To this end, we propose KineVLA, a novel framework that formally defines the kinematically rich VLA task and introduces a dual-layer action representation coupled with a two-level reasoning token mechanism. This design explicitly decouples task-goal invariance from kinematic variability, enabling precise responses to kinematics-sensitive instructions. We further construct a kinematics-aware VLA dataset spanning both simulated and real robotic environments, complete with a dedicated annotation protocol. Through joint vision–language–action modeling and intermediate-variable supervision for alignment, KineVLA significantly outperforms prior methods on the LIBERO benchmark and the Realman-75 robot, demonstrating superior accuracy, controllability, and generalization in kinematically dense tasks.
This work proposes a Task–Environment–Ontology (TEE) class and a Semantic Digital Twin (SDT) framework to ensure robotic actions are correct at semantic, causal, and ontological levels. It introduces the first formal axiomatization of bodily motion for task achievement, decomposing action correctness into three verifiable predicates: semantic satisfaction, causal sufficiency, and ontological feasibility. The approach enables cross-platform action verification, typed failure diagnosis, and counterfactual reasoning. Evaluated in a kitchen environment, the system successfully synthesizes, verifies, and attributes container manipulation tasks across three heterogeneous mobile manipulator platforms, demonstrating its generality and practical utility.
This work addresses the inefficiency and verbosity of object rearrangement on tabletops when limited to grasp-and-place actions by introducing non-grasping topple operations to enrich the action space. The authors formulate the task as a pebble-motion problem with aggregation operations through a directed graph abstraction, systematically incorporating topple—alongside other aggregate actions such as scooping—into the planning framework for stacked rearrangement for the first time. This abstraction generalizes naturally to additional aggregate manipulations beyond toppling. Experimental results in the IsaacSim simulation environment demonstrate that plans integrating topple actions significantly reduce execution time compared to pure grasp-and-place strategies, thereby validating the effectiveness and potential of enriched interactive action abstractions in manipulation tasks.
This work addresses the high latency and sluggish responsiveness of large language models (LLMs) in real-time embodied agent control by introducing an operating system–inspired runtime architecture. For the first time, core OS principles are integrated into embodied intelligence, enabling overlapping planning and execution through asynchronous multi-scale planning loops, typed skill kernels, preemptive scheduling, and speculative skill streaming. The architecture supports natural language programming by compiling natural language commands into interrupt handlers, substantially enhancing concurrent task handling. Experiments on the Unitree Go2 quadruped robot demonstrate a 50% reduction in per-step latency, a 73% decrease in time-to-first-action, and efficient low-overhead support for concurrent multi-task execution.
Existing theoretical frameworks struggle to fully account for the empirical success of action chunking in behavioral cloning. Through both simulated and real-world robotic experiments, this work demonstrates that its core advantages stem from non-Markovian representational capacity, mitigation of error accumulation, and an implicit ensemble effect. We find that action chunking effectively learns diverse temporal dependencies, thereby achieving performance akin to model ensembling without explicit architectural modifications. Building on this insight, we propose a novel approach that reproduces the benefits of action chunking without requiring chunked action outputs, and further introduce an explicit ensemble strategy that significantly outperforms conventional action chunking across multiple tasks.