🤖 AI Summary
This study addresses the performance bottleneck of pure RGB skeleton detection in complex scenarios by proposing a novel depth-dominant, RGB-assisted paradigm embodied in the DDSkel model. The framework employs an asymmetric encoder for cross-modal feature fusion and incorporates a lightweight RGB branch comprising only 12% of the parameters, thereby ensuring robustness while significantly reducing computational overhead. Experimental evaluations on the SymPASCAL dataset demonstrate that the proposed method surpasses current state-of-the-art approaches while utilizing merely 36% of their trainable parameters. Consequently, this work achieves a substantial balance between accuracy and efficiency, offering a highly effective solution for resource-constrained skeleton detection tasks without compromising performance in challenging environments.
📝 Abstract
To date, all natural scene skeleton detection follows the paradigm of taking RGB images as the sole input; despite notable progress, methods under this paradigm suffer significant performance degradation on complex-content images. We observe that depth images are inherently insensitive to color and texture, and can provide clear regional contours and inter-region spatial relationships, which naturally alleviates the difficulty of skeleton detection in complex scenarios. Motivated by this observation, this paper proposes for the first time a novel skeleton detection paradigm where depth images serve as the dominant modality and RGB images act as the auxiliary, and accordingly presents a model DDSkel (short for Depth-Dominant Skeleton Detection) under this paradigm. DDSkel employs an asymmetric encoder design to fuse RGB information into depth features, with the RGB modality branch having only 12% the parameters of the depth modality branch. DDSkel has a simple structure without intricate designs. Nevertheless, with only 36% of the trainable parameters of the current best method, DDSkel outperforms all state-of-the-art approaches on SymPASCAL, the most challenging dataset with a large volume of complex images.