🤖 AI Summary
This study addresses the challenges of action space explosion, intermediate-state instability, limited multimodal support, and slow inference in autonomous 3D construction. To this end, we propose a lightweight multimodal autoregressive framework that formulates construction as a probabilistic next-block generation task. The method introduces auxiliary scaffold tokens to explicitly ensure intermediate structural stability and incorporates a Sequential Monte Carlo (SMC) algorithm to explore optimal assembly trajectories in parallel. Experimental results demonstrate that the proposed model reduces parameter count to one-quarter of the baseline while accelerating inference by 5 to 20 times. It maintains comparable construction quality while significantly enhancing overall structural stability. Furthermore, the effectiveness of the framework is validated through real-world robotic experiments.
📝 Abstract
Autonomously constructing physically realizable 3D structures remains a significant challenge due to combinatorial action spaces, interchangeable components, equifinal assembly sequences, and strict stability requirements during construction. State-of-the-art methods fine-tune large language models for text-based generative construction. However, these approaches do not allow for Multimodal (text, image, sketch) conditioning, overlook the practical role of scaffolding for stabilizing intermediate structures, and suffer from slow inference speeds. Therefore, we formulate construction as a probabilistic next-block generation task with multiple potential assembly actions and multiple potential task conditioning modalities. Concurrently, we explicitly consider the utility of scaffolding by introducing an auxiliary scaffold block token. We present Scaffold Multimodal Monte Carlo (ScaffoldM3C), a multimodal, lightweight, auto-regressive model for stable block-based construction, that proposes a set of next-step candidate blocks. Leveraging these candidates, we utilize Sequential Monte Carlo (SMC) to maintain a population of possible assembly sequences, allowing us to consider multiple, potentially different, assembly directions simultaneously. We train our multimodal architecture by extending the StableText2Brick dataset to contain image conditioning prompts and scaffold-stabilized build sequences. ScaffoldM3C is 4x smaller than competing baselines, yielding a 5x to 20x speedup during inference, while achieving comparable construction quality to state-of-the-art methods and higher overall stability. We demonstrate the effectiveness of our approach through simulations and real-world robot assembly demonstrations.