SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出SenseFuse方法,通过无标签融合2D图像和3D形状编码器来改进开放词汇3D实例分割,增强现有流程中的掩码标注准确性。
📝 Abstract
Open-vocabulary scene understanding is fundamental for robotics, laying the groundwork for spatial reasoning and object manipulation. While closed-vocabulary 3D instance segmentation heavily leverages 3D shape information, state-of-the-art open-vocabulary methods remain predominantly restricted to 2D image features or image-distilled representations during mask labeling. In this paper, we propose SenseFuse, a label-free fusion method that balances 2D image and 3D shape encoders for robust open-vocabulary 3D instance segmentation, refining only the mask-labeling stage of existing pipelines. We reveal that 2D image and 3D shape encoders exhibit largely disjoint failure patterns and rarely share identical wrong labels, whereas two 2D image encoders frequently repeat the same errors. This distinct behavior makes the 2D and 3D pair inherently complementary. We introduce an adaptive mechanism that selects a scene-level fusion weight to maximize a label-free sensitivity measure, estimated directly from a single scene's unlabeled proposals in milliseconds. SenseFuse improves labeling accuracy in every evaluated setting across ScanNet200, Replica, and ScanNet++, recovering 67-100% (median 93%) of the gain achievable with an oracle weight, and it raises instance AP in 21 of 22 reported settings. Code is available at https://github.com/hanes1207/SenseFuse.
Problem

Research questions and friction points this paper is trying to address.

open-vocabulary 3D instance segmentation
2D image features
3D shape information
mask labeling
label-free fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

label-free fusion
open-vocabulary 3D instance segmentation
adaptive mechanism
sensitivity measure
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
E
Euiseok Han
Korea Advanced Institute of Science and Technology (KAIST)
Tri Ton
Tri Ton
KAIST, Korea
Computer Vision
H
Hwanhee Kim
Korea Advanced Institute of Science and Technology (KAIST)
S
Seungyeon Ryu
Korea Advanced Institute of Science and Technology (KAIST)
Chang D. Yoo
Chang D. Yoo
kaist
machine learningcomputer visionsignal processingspeech enhancementspeech recognition