🤖 AI Summary
Existing 3D large language models (3D LLMs) generally lack the capacity to perceive fine-grained attributes and non-salient objects. To address this limitation, this work proposes LensDistill, a mechanism that seamlessly distills external fine-grained visual knowledge into native 3D models by integrating 3D grounding, multi-view selection, and vision-language model collaboration. Notably, this approach effectively enhances the model's comprehension of scene details without requiring architectural modifications. Experimental results demonstrate that LensDistill significantly improves fine-grained captioning performance while fully preserving the model's original capabilities in scene question answering and spatial localization. Consequently, this method establishes an efficient new paradigm for developing high-precision 3D understanding models.
📝 Abstract
Existing 3D large language models often overlook fine-grained attributes and less visually salient objects and parts, even when relevant evidence is present in scene videos. We introduce Lens3D to improve fine-grained object understanding through external visual assistance and knowledge transfer. Its LensUnd pipeline adopts 3D localization to select informative, complementary views for an external 2D vision-language model, supporting fine-grained object captioning, small-object grounding, and fine-grained object question answering. LensDistill transfers the resulting fine-grained knowledge to 3D LLMs through detailed caption supervision, enabling captioning from native inputs without external VLM calls. We also construct LensBench, a held-out evaluation set of 2,068 objects with three silver-standard reference descriptions per object. Experiments with Video-3D LLM and 3DRS demonstrate that LensDistill substantially improves fine-grained object captioning while preserving existing grounding and scene-level QA performance. These results establish the feasibility of transferring externally acquired fine-grained knowledge into native 3D LLMs.