BTC3D: Blended Tile Conditioning for Detail-Enhancing Image-to-3D Generation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of high-fidelity details in image-to-3D generation caused by global encoding compression. To overcome this limitation, we propose BTC3D, a training-free inference framework that first reveals the additivity of diffusion model features. By designing hybrid tile embeddings coupled with a dynamic conditioning scheduling mechanism, our method precisely extracts and fuses local high-frequency signals, enhancing detail preservation without requiring retraining. Experimental results demonstrate that BTC3D significantly improves texture quality and visual fidelity while maintaining global structural consistency. Furthermore, it can be seamlessly integrated into existing generative pipelines.
📝 Abstract
Recent diffusion-based pipelines have achieved promising progress in image-to-3D synthesis. However, generating high-fidelity details remains challenging, especially when the input image contains rich details. Existing approaches often rely on globally encoded conditioning features, which compress spatial information and limit the model to reproduce fine-grained details. This common design often leads to a phenomenon we term detail attenuation. Moreover, improving image-to-3D synthesis quality typically requires retraining or fine-tuning large diffusion models, which can be computationally expensive and impractical for complex 3D pipelines. In this work, we present Blended Tile Conditioning for image-to-3D generation (BTC3D), a training-free inference time framework that enhances fine-grained detail preservation in image-to-3D diffusion pipelines. To alleviate detail attenuation, we first examine the image feature additivity in image-to-3D models. Based on this property, we introduce a blended tile embedding that extracts local conditioning signals from split image regional patches, allowing the diffusion model to better preserve fine-grained visual details. To integrate the global and local conditioning guidance stably, we propose a dynamic conditioning schedule that gradually increases the influence of tile-level conditioning during later low-noise stages of diffusion. Our proposed method BTC3D operates entirely at inference time and can be seamlessly integrated into existing image-to-3D diffusion pipelines. Experimental results demonstrate that the proposed approach significantly improves texture quality and visual fidelity of the base model while maintaining global structural consistency in a training-free manner.
Problem

Research questions and friction points this paper is trying to address.

Image-to-3D generation
Detail attenuation
Fine-grained details
Diffusion models
Training-free
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free
Image-to-3D generation
Blended tile conditioning
Detail attenuation
Dynamic conditioning schedule
🔎 Similar Papers
No similar papers found.