I2CD: Direct Image-to-Convex Decomposition for Simulation-Ready Collision Geometry

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation that 3D meshes generated from images are predominantly non-convex, necessitating complex post-processing for physical simulation and motion planning. To overcome this, we propose an end-to-end framework that directly predicts convex decomposition geometry from a single RGB image. Specifically, our method freezes a pretrained Hunyuan3D-2 diffusion transformer and trains only a lightweight cross-attention head to generate β€œconvex slot” tokens. Combined with a shape decoder and a half-space parameterization, it outputs compact collision primitives that can be directly loaded into physics engines without any post-processing. Experimental results demonstrate that our approach achieves superior volumetric IoU compared to eight baselines while accelerating inference by 6–37Γ—. Furthermore, the method exhibits compatibility across multiple physics engines, substantially reducing deployment preparation time in real-world robotic scenarios.
πŸ“ Abstract
Physics simulators and motion planners require convex collision geometry, yet image-to-3D generative models output dense, frequently non-manifold visual meshes. Bridging the two today takes a slow, brittle reconstruct-then-decompose pipeline of repair, decimation, and approximate convex decomposition. We present I2CD, which predicts a convex decomposition directly from a single RGB image. Rather than train a new image-to-3D model, I2CD freezes the pretrained Hunyuan3D-2 image-conditioned diffusion transformer and shape decoder and trains only a lightweight cross-attention head (38M parameters, under ten GPU-hours) whose learned "convex-slot" tokens emit the halfplane parameters of $K$ convex polytopes. The output is compact, convex by construction, and loads into physics engines without any post-processing, in ${\sim}0.5$s per image. On $227$ held-out OmniObject3D and Google Scanned Objects instances, I2CD attains the highest volumetric IoU among eight reconstruct-then-decompose pipelines while running $6$-$37\times$ faster end-to-end. In a cross-simulator study in MuJoCo, PyBullet, Genesis, and Isaac Sim, every engine uses I2CD geometry as delivered, whereas raw generated meshes "load" everywhere but are silently replaced by a different collision shape in most cases or need seconds to minutes of per-object preprocessing. On a physical xArm7, I2CD produces planner-ready geometry for a $20$-object cluttered scene in $11$s versus $328$s for the strongest baseline, at comparable pick-and-place execution success ($85$ vs. $90$ of $100$ trials).
Problem

Research questions and friction points this paper is trying to address.

convex decomposition
collision geometry
image-to-3D
physics simulation
motion planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Image-to-Convex Decomposition
Collision Geometry
Parameter-Efficient Fine-Tuning
Physics Simulation
Cross-Attention
πŸ”Ž Similar Papers
No similar papers found.