Octrees as an Explicit 3D Language

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of spatial structure in existing 3D large models and the degradation of general language capabilities caused by fine-tuning. To this end, we propose OctLLM, which pioneers the use of sparse octrees as an explicit 3D language to precisely represent geometric structures. By employing a dual-stream shared attention mechanism and a decoupled training strategy, 3D capabilities are injected through independent parameter branches, enabling parameter-efficient fine-tuning. With minimal parameter updates, OctLLM substantially reduces the FID score for image-to-3D generation and improves captioning accuracy, establishing new state-of-the-art results across multiple benchmarks. Furthermore, it fully preserves the frozen language-vision pathways, achieving unified multimodal representation while maintaining lossless compatibility with general-purpose capabilities.
📝 Abstract
Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding. On the other front, existing methods introduce a new modality with full fine-tuning or LoRA, but full fine-tuning is costly, LoRA limits 3D capacity, and both modify the language pathway. OctLLM instead adds 3D capacity in parameters separate from the pretrained ones: mesh tokens are routed through independent trainable branches in a subset of blocks while text and image tokens retain the frozen vision-language pathway, and the two streams interact through shared self-attention. It trains far fewer parameters than full fine-tuning, yet sets a new state of the art among unified multimodal LLMs, lowering image-to-3D FID by $17.4\%$ and raising render-grounded captioning by $28.7$ points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.
Problem

Research questions and friction points this paper is trying to address.

3D Large Language Models
Spatial Structure Preservation
Catastrophic Forgetting
Multimodal Integration
Efficient Fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Octree
Sparse Octree
Multimodal LLM
Parameter-efficient
Mask Modeling
🔎 Similar Papers
No similar papers found.