🤖 AI Summary
This work addresses the instability of local deformations in coarse-mesh mass-spring models caused by oversimplified axial spring mechanics. We propose a bending-aware, differentiable mass-spring model that, for the first time within this framework, incorporates surface-triplet-based bending stiffness and damping constraints. This enables physically consistent reconstruction and dynamics prediction of deformable objects from sparse-view RGB-D videos. The method substantially improves mechanical stability and reconstruction accuracy across varying sampling densities, outperforming the axial-spring-only PhysTwin baseline while preserving model simplicity—making it well-suited for physics-driven digital twin applications.
📝 Abstract
Reconstructing objects with mechanical properties from video observations enables physically consistent dynamic prediction, benefiting robotics planning and interaction. Existing spring--mass based physical driven reconstruction approaches offer efficient and differentiable physical reconstruction, but they typically rely on axial springs alone. Such formulations oversimplify the underlying structural mechanics and can become mechanically under-constrained when the physical graph is coarsened, limiting their ability to preserve stable local deformation. We present BendTwin, a bending-aware differentiable spring--mass framework for video-based reconstruction and future prediction of deformable objects. BendTwin introduces bending stiffness and damping over local surface triplets, penalizing deviations from rest angles and regularizing higher-order deformation. These bending constraints improve mechanical stability while preserving the simplicity of spring--mass system. Experiments show that BendTwin consistently outperforms the axial-only PhysTwin baseline. Ablation studies further demonstrate that the bending constraints maintain system stability across different downsampling ratios and consistently improve upon the original PhysTwin formulation. Overall, BendTwin provides an effective approach for constructing mechanically faithful digital twins from sparse-view RGB-D videos.