Score
Designs and implements models and simulation tools that predict, synthesize, or infer geometric changes to terrain surfaces—e.g., generating deformation fields or edited terrain meshes from static geometry. These methods enforce physical plausibility and scene consistency, ensuring continuity with surrounding geometry and adherence to physical constraints during deformation.
Existing terrain generation tools prioritize artistic expression and visual realism but lack parametric control, reproducibility, and scriptability—hindering their use in intelligent robotic simulation-driven development, where controllable and explicitly defined terrains are essential. To address this, we propose TerrainGen: a highly modular Python library for procedural terrain generation that integrates rule-based modeling with multi-scale noise synthesis. It enables fine-grained parameterization of physical attributes—including slope, surface roughness, and rock density—and adopts a loosely coupled architecture compatible with Blender for automated rendering and object placement. A declarative configuration interface further simplifies terrain specification. Experimental evaluation demonstrates TerrainGen’s effectiveness in synthetic data generation and perception ground-truth annotation, significantly improving controllability, reproducibility, and deployment efficiency in environment construction. By providing a scalable, programmable infrastructure, TerrainGen advances simulation-based robotics development and machine learning training pipelines.
This study addresses the lack of systematic integration among existing models and algorithms for terrain visibility analysis, a core technique in landscape planning and spatial decision-making. It provides a comprehensive review of line-of-sight and cumulative viewshed computation models based on Triangulated Irregular Networks (TIN) and regular grids, systematically categorizing algorithmic approaches into exact and approximate solutions. Furthermore, this work innovatively establishes a multi-scale theoretical framework for terrain visibility and deeply integrates interdisciplinary application cases from archaeology, landscape management, and land conservation. By identifying critical technical bottlenecks and future research directions, this paper offers a systematic reference for efficient visibility computation in complex terrain scenarios.
This study investigates whether the motion evolution in video generation models adheres to physical laws. To this end, it proposes a paired-intervention evaluation framework that conducts controlled experiments and physics benchmarking across nine image-to-video models by manipulating geometric variables such as trajectories and obstacles. The findings reveal that current models fundamentally perform geometry-conditioned motion synthesis, with their state evolution systematically deviating from physical laws. Although these models respond to geometric features that influence velocity, they lack the capacity for continuous physical state modeling. Furthermore, text prompts predominantly determine endpoints while duration governs pacing, and non-causal editing phenomena are observed. Collectively, this work offers novel insights into the physical limitations of video generation.
Reconstructing high-quality, manifold, and feature-preserving triangular meshes from unstructured point clouds remains challenging. Method: This paper proposes an isotropic reconstruction method based on unit-sphere covering. It systematically models local point cloud geometry using spherical covering theory, rigorously derives parameter bounds to theoretically guarantee manifold output, and jointly optimizes geometric uniformity (edge-length and angle distributions) and topological correctness. The approach integrates geometric approximation analysis, adaptive neighborhood graph construction, and Delaunay-type triangulation optimization, supporting feature detection and multi-patch remeshing. Results: Experiments show that our method reduces edge-length and angle standard deviations by ~35% compared to Poisson surface reconstruction and Ball-Pivoting, while decreasing average runtime by 45% (1.8× speedup). Under reasonable sampling conditions, the resulting mesh is provably manifold.
Existing 3D generative methods suffer from slow optimization, irregular topology, noisy surfaces, and limited editability. To address these challenges, we propose the first 3D-native diffusion framework supporting both text and image conditioning, enabling second-level generation of topologically regular coarse meshes. Our method introduces multi-view joint conditional modeling and latent-space set representation to enhance geometric consistency across views. Furthermore, we pioneer a normal-field-driven geometric refinement mechanism that enables automated detail enhancement and intuitive, interactive user editing. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods in both qualitative and quantitative evaluations—producing high-fidelity, geometrically coherent meshes with rich surface details and flexible, real-time editability. The code and pretrained models are publicly released and seamlessly integrated into practical 3D modeling workflows.
This study addresses the limitations of conventional two-dimensional coverage path planning in adapting to complex three-dimensional terrains by proposing a terrain-following, full-coverage path planning algorithm for 3D environments. The method extends 2D coverage paths into 3D space for the first time, constructing a uniform elevation grid via inverse distance weighting (IDW) interpolation. It incorporates projection distance constraints and a local search strategy to generate adjacent paths spaced precisely at the width of agricultural machinery and maintained at a constant operational height above the terrain. Experimental results on real-world agricultural terrain data demonstrate that the proposed algorithm effectively produces paths satisfying both full coverage and operational consistency requirements, significantly outperforming traditional 2D approaches in complex terrains.
Existing lunar rover simulators struggle to simultaneously achieve high visual fidelity and physically accurate terrain interaction, limiting comprehensive replication of lunar surface conditions. This work proposes a real-time simulation framework that integrates a data-driven terramechanics model with high-fidelity visualization. A regression model, trained on full-vehicle and single-wheel experimental and simulation data, accurately predicts wheel slip and sinkage under both steady-state and dynamic conditions. Furthermore, the method enhances terrain deformation and rut-rendering algorithms to produce interactions that are both physically plausible and visually realistic. Validated on both flat terrain and slopes up to 20°, the approach faithfully reproduces wheel–terrain coupling behaviors and has been empirically verified to meet the demands of real-time, high-fidelity lunar rover simulation.
This study addresses the challenge of synchronously propagating road edits to vehicle dynamics, camera viewpoints, and inter-vehicle spacing in autonomous driving scenarios. It proposes the first framework that unifies the propagation of road geometry and surface condition changes to multi-vehicle dynamics, camera poses, and headway distances. The method constructs a unified road model linking scene deformation with four-wheel vehicle dynamics, generating physically consistent counterfactual driving videos from reconstructed multi-vehicle trajectory segments. Furthermore, a surrogate model is trained to enable low-cost safety screening. Validated through CarSim simulations and experiments on the Waymo dataset, the proposed approach reduces terminal spacing prediction errors by 22–40%, significantly enhancing both the efficiency and safety of counterfactual data generation.
This work addresses the challenge of enabling Transformer-based models to generate high-quality, editable CAD sketches while preserving design intent and parametric control. Inspired by classical compass-and-straightedge constructions, the authors propose modeling sketch generation as a sequence of geometric construction steps—such as offsetting, rotation, and intersection—and introduce, for the first time, a “chain-of-thought”-like sequence of geometric operations. The generation process is optimized via reinforcement learning and implemented using a Transformer architecture capable of handling floating-point parameters, thereby supporting fine-grained parametric editing. Experimental results demonstrate that the proposed method significantly outperforms existing baselines in terms of sketch quality, editability, and multiple evaluation metrics, with consistent improvements even on metrics not explicitly optimized during training.
Existing 3D generation methods struggle to meet production-grade requirements for real-time interactive applications, such as consistent topology, UV unwrapping, physically based rendering (PBR) materials, skeletal rigging, and physically plausible scene layout. To address this gap, this work proposes a two-dimensional taxonomy centered on asset production pipelines—structured by asset type and production stage—and systematically constructs a comprehensive generation framework encompassing geometry synthesis, topology optimization, UV parameterization, PBR appearance modeling, skeletal rigging, and physics-aware scene assembly. The authors further introduce a cross-dimensional evaluation protocol to rigorously assess the direct usability of generated assets in game engines and simulation platforms. Their analysis highlights critical challenges in data quality, controllable generation, and end-to-end assetization, underscoring the pivotal role of deployable 3D content as foundational infrastructure for embodied intelligence and interactive world models.