🤖 AI Summary
This study addresses the persistent challenge in image-to-3D generation, where existing methods struggle to preserve fine-grained geometric details such as text and logos, frequently resulting in distortion or information loss. To overcome this limitation, this work proposes the Hi3D 3.0 system, centered on the Twinkle3D model. By extending a Diffusion Transformer (DiT) architecture to process sequences of 300,000 tokens and introducing a fine-grained cross-modal interaction mechanism, the approach enables single-stage joint modeling of global and local features. Furthermore, an O-Voxel optimization strategy is incorporated to generate high-fidelity, watertight meshes at 2048³ resolution. Experimental results demonstrate that the proposed method achieves a character recovery rate of 82.1% with 98.2% precision, significantly outperforming four commercial systems that yield only a 21.7% recall rate, thereby facilitating high-quality object-level 3D asset generation.
📝 Abstract
Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution. Twinkle3D advances high-fidelity geometry generation along four dimensions. First, while O-Voxel/FaithC offers high representational precision, it often suffers from poor surface quality and non-watertight geometry. We address both issues while retaining its $2048^{3}$-level precision. Second, we scale diffusion generation to sequences of up to 300K geometric tokens through a redesigned DiT architecture and large-scale distributed training optimizations, reducing training time per step from approximately ten minutes to ten seconds. Third, subsequent refinement cannot fully compensate for errors introduced during initial generation; we therefore strengthen both global shape and local detail in the initial generation stage, and the resulting single-stage model surpasses prior two-stage pipelines with $512^{3}$ refinement. Finally, we introduce a fine-grained image-3D cross-modal interaction mechanism that strengthens correspondence between visual evidence and geometric tokens, improving the recovery of object-specific structures. We evaluate geometric fidelity using alignment metrics derived from silhouettes and normal fields. Hi3D 3.0 outperforms four commercial systems across all reported metrics, recovering 82.1% of inscribed characters at 98.2% precision, compared with 21.7% recall for the strongest competitor.