PartLLM: A Unified Multimodal Foundation for 3D Part Segmentation

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出PartLLM,一种统一的多模态模型,通过意图条件生成方法解决3D部件分割问题,支持文本引导、交互式及全形状语义分解。
📝 Abstract
Part segmentation is a fundamental problem in computer graphics and 3D vision. Recent works have expanded 3D part segmentation beyond fixed taxonomies, but existing approaches typically only address a specific setting, such as text-guided part segmentation or point-based interaction. In this work, we argue that these settings can be unified as an intent-conditioned generative problem, where different prompts specify the desired part decomposition. To this end, we introduce PartLLM, a unified multimodal model that formulates 3D part segmentation as autoregressive semantic decomposition. Conditioned on an input shape and a user prompt, PartLLM autoregressively generates semantic part hypotheses as queries for mask prediction and feeds them to a decomposition-aware decoder that jointly predicts coherent part masks. This unified design supports text-guided part segmentation, interactive segmentation, and full-shape semantic decomposition with controllable granularity within a single model. Extensive experiments across these task settings show that PartLLM consistently outperforms task-specific baselines, demonstrating the effectiveness of unifying 3D part segmentation under an intent-conditioned generative formulation.
Problem

Research questions and friction points this paper is trying to address.

3D part segmentation
unified approach
intent-conditioned
Innovation

Methods, ideas, or system contributions that make the work stand out.

PartLLM
unified multimodal model
intent-conditioned generative problem
autoregressive semantic decomposition
decomposition-aware decoder
🔎 Similar Papers
No similar papers found.