HI3D 3.0 (Twinkle3D): Object-specific 3D Asset Generation with High Resolution

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the persistent challenge in image-to-3D generation, where existing methods struggle to preserve fine-grained geometric details such as text and logos, frequently resulting in distortion or information loss. To overcome this limitation, this work proposes the Hi3D 3.0 system, centered on the Twinkle3D model. By extending a Diffusion Transformer (DiT) architecture to process sequences of 300,000 tokens and introducing a fine-grained cross-modal interaction mechanism, the approach enables single-stage joint modeling of global and local features. Furthermore, an O-Voxel optimization strategy is incorporated to generate high-fidelity, watertight meshes at 2048³ resolution. Experimental results demonstrate that the proposed method achieves a character recovery rate of 82.1% with 98.2% precision, significantly outperforming four commercial systems that yield only a 21.7% recall rate, thereby facilitating high-quality object-level 3D asset generation.
📝 Abstract
Image-to-3D generation has become increasingly capable of producing objects that closely resemble the input image, and an outstanding challenge is to reproduce the depicted object itself, including the specific geometry that defines it. Inscriptions, brand marks, and repeated structures are frequently distorted or lost, despite being critical to object identity. We present Hi3D 3.0, an image-to-3D generation system targeting object-specific fidelity, with Twinkle3D as its geometry model for generating watertight triangle meshes at $2048^{3}$ resolution. Twinkle3D advances high-fidelity geometry generation along four dimensions. First, while O-Voxel/FaithC offers high representational precision, it often suffers from poor surface quality and non-watertight geometry. We address both issues while retaining its $2048^{3}$-level precision. Second, we scale diffusion generation to sequences of up to 300K geometric tokens through a redesigned DiT architecture and large-scale distributed training optimizations, reducing training time per step from approximately ten minutes to ten seconds. Third, subsequent refinement cannot fully compensate for errors introduced during initial generation; we therefore strengthen both global shape and local detail in the initial generation stage, and the resulting single-stage model surpasses prior two-stage pipelines with $512^{3}$ refinement. Finally, we introduce a fine-grained image-3D cross-modal interaction mechanism that strengthens correspondence between visual evidence and geometric tokens, improving the recovery of object-specific structures. We evaluate geometric fidelity using alignment metrics derived from silhouettes and normal fields. Hi3D 3.0 outperforms four commercial systems across all reported metrics, recovering 82.1% of inscribed characters at 98.2% precision, compared with 21.7% recall for the strongest competitor.
Problem

Research questions and friction points this paper is trying to address.

Image-to-3D generation
Object-specific fidelity
High-resolution geometry
Detail preservation
Watertight mesh
Innovation

Methods, ideas, or system contributions that make the work stand out.

Image-to-3D Generation
High-Fidelity Geometry
Diffusion Transformer
Cross-Modal Interaction
Object-Specific Fidelity
💼 Related Jobs
No related jobs found.
Z
Ziying Li
S
Shengchu Zhao
H
Huiang He
Y
Yiyang Chen
J
Jianwen Huang
B
Bailin Li
Changhao Li
Changhao Li
Ph.D. of Massachusetts Institute of Technology '23
Quantum computingQuantum sensingQuantum algorithmsNV center in diamond
J
Jianhui Li
J
Jie Li
R
Ruiyang Liu
Y
Yibo Luo
T
Tengjiao Sun
P
Pei Tang
S
Shiwen Wang
J
Jiaqi Wu
K
Kang Wu
K
Kaiqiao Yang
Z
Zherui Yang
H
Hu Zhang
X
Xuezhi Zhao
X
Xinhe Zheng
Y
Yukun Li
Heliang Zheng
Heliang Zheng
Unknown affiliation
Image GenerationRepresentation Learning
R
Rongfei Jia