Score
Designs and implements methods that integrate explicit 3D scene representations (e.g., volumetric maps, meshes, point clouds or implicit 3D features) into action planners so that generated plans, trajectories, or control policies respect geometric, kinematic, and collision constraints. Builds and evaluates algorithms for fusing 3D perception and representation with planning and control—covering representation conversion, uncertainty handling, real‑time updates, and performance analysis of planning in three‑dimensional environments.
Coordinating whole-body motion during long-horizon, multi-step mobile manipulation involving articulated objects remains challenging due to the coupling between navigation and manipulation in structured environments. Method: This paper introduces the Augmented Configuration Space (A-Space), which abstracts scene structure into a kinematic model unified with the robot’s own kinematics, enabling joint navigation–manipulation planning. The approach integrates symbolic task planning, optimization-based trajectory planning, and an intermediate-layer refinement mechanism within a three-tiered task–action–trajectory framework to ensure cross-scene generalization and long-term feasibility. A-Space jointly models joint reachability and supports unified constraint solving. Results: In simulation, task success rate improves by 84.6%; on a real robot, the method successfully executes up to 14 consecutive manipulation steps across 17 diverse scenes involving seven categories of rigid and articulated objects.
In grid-based motion planning with finite motion primitives, conventional A* suffers from low search efficiency due to high branching factor. Method: We propose a novel joint grid-level and primitive-level search paradigm: motion primitive sequences are embedded synchronously onto grid cells to structurally compress the action space; a sound and falsifiable pruning strategy is designed to significantly reduce the search space while preserving theoretical completeness and optimality; and the approach is integrated within the classical A* framework to ensure robustness. Contribution/Results: Experiments show a 1.5× speedup in runtime with only marginal degradation in solution quality (<2%), achieving an effective trade-off between efficiency and reliability. The core innovation lies in the first deep coupling of motion primitive sequence modeling with grid-level search—overcoming the longstanding tension between branching factor and optimality in lattice-based planning.
Achieving generalizable 3D manipulation in dynamic environments remains challenging for robotic systems. Method: This paper introduces LMM-3DP, the first framework to tightly integrate high-level semantic planning from Large Multimodal Models (LMMs) with low-level, semantics-aware 3D feature field control. Contribution/Results: Its core innovations include (1) a language-3D joint attention mechanism enabling cross-modal feature alignment within a 3D Transformer; (2) a closed-loop collaborative architecture incorporating self-feedback critique, hierarchical policy memory, and failure-driven retry; and (3) robust long-horizon kitchen manipulation under environmental perturbations. Experiments in real kitchen settings demonstrate a 1.45× improvement in low-level control success rate and ~1.5× gain in high-level planning accuracy over pure LLM-based baselines, significantly advancing embodied 3D reasoning and execution.
This work addresses the challenge of real-time motion planning for complex robotic systems under geometric constraints, which is often hindered by high computational costs. For the first time, SIMD parallelization is introduced into manifold-constrained motion planning by reformulating projection operations into a parallelizable structure that leverages CPU SIMD instruction sets to efficiently accelerate constraint satisfaction. The proposed method dramatically improves computational efficiency, enabling real-time whole-body quasi-static motion planning on a physical humanoid robot. Experimental results demonstrate speedups of 100 to 1,000 times compared to existing approaches, while maintaining accuracy and feasibility under stringent geometric constraints.
This paper presents a systematic review of learning-based 3D representations for tasks such as 3D reconstruction, novel view synthesis, and rendering, spanning from traditional explicit formulations—including meshes, point clouds, and voxels—to emerging implicit neural fields and primitive-based splatting techniques like 3D Gaussian Splatting. Emphasizing the evolutionary trajectory of 3D representations themselves, the work highlights the paradigm shift from explicit to implicit modeling, clarifying the mathematical formulations, strengths, limitations, and suitable application scenarios of each approach. Distinct from prior task-centric surveys, this study centers on representation as the core organizing principle, offering fresh insights for 3D/4D content generation and identifying key challenges and future research directions to serve as a theoretical reference for the computer graphics and vision communities.
This work addresses the challenge of reliably and scalably representing collision-free configuration spaces in high-dimensional, complex environments. The authors propose the ILD framework, which uniquely integrates invertible latent-space mapping with explicit modeling of free space as a union of convex polytopes. Specifically, an invertible neural network maps the configuration space to a latent space, where the free region is explicitly represented as a union of convex sets. Visibility-guided sampling ensures connectivity among these sets, and paths planned in the latent space are decoded back to the original space via the invertible mapping, guaranteeing strict feasibility without false positives. Experiments demonstrate that the method significantly improves coverage, connectivity, and planning success across scenarios ranging from 2D to 14 degrees of freedom, supports real-time planning, and adapts robustly to geometric changes in real-world environments.
This work addresses the limitations of existing world model–based visual navigation approaches, which typically decouple goal intent verification from trajectory generation, leading to computational redundancy and inconsistencies between actions and visual predictions. To overcome these issues, the authors propose SWAM—a task-driven, joint observation–action generation framework that, for the first time, enables end-to-end zero-shot cross-environment generalization from only monocular RGB inputs. SWAM simultaneously synthesizes intermediate RGB-D sequences and corresponding action trajectories in a single forward pass, starting solely from source and goal RGB images. The method incorporates a vision-guided action refinement module and a trajectory-scale regularization loss, while leveraging depth pseudo-labels to internalize spatial priors. Experimental results demonstrate that SWAM significantly outperforms current two-stage planners in terms of success rate, trajectory accuracy, and inference efficiency.
This work addresses the challenge of achieving safe and real-time obstacle avoidance and motion control for robots operating in unstructured dynamic environments. The authors propose a novel approach that integrates 3D Gaussian Splatting–based scene reconstruction with reactive control. By leveraging voxel filtering and dynamic Gaussian relocation, the method enables efficient reconstruction of dynamic scenes. Notably, it introduces—for the first time—a differentiable, continuous signed distance function derived from isotropic Gaussians, seamlessly bridging implicit representations with classical distance fields. This representation is further combined with Control Barrier Functions to establish a closed-loop perception-to-action pipeline. Extensive evaluations in simulation, on physical robots, and within human-robot shared workspaces demonstrate the system’s capability to perform high-fidelity dynamic scene reconstruction and ensure real-time, collision-free navigation in previously unknown environments.