Score
Designs and implements methods to estimate, extract, and reconstruct 2D UV texture atlases that encode the surface appearance of 3D meshes, converting image or scan data into per-mesh UV maps and aligning pixels to UV coordinates. Builds pipelines for UV map mapping, completion, and residual/decoded UV outputs so textures are consistent across frames, can be filled or inpainted where missing, and are suitable for rendering or further appearance analysis.
This paper addresses the longstanding challenge of balancing visual realism and computational efficiency in 3D mesh texturing by providing a systematic survey of recent advances in neural texture generation. It introduces a unified classification framework that encompasses methods ranging from early generative adversarial networks (GANs) to modern diffusion models, integrating core technical components such as mesh geometry processing, texture mapping, differentiable rendering, and neural synthesis. The work comprehensively traces the evolution of research in texture synthesis, transfer, and completion, while establishing a structured taxonomy of learning-based approaches. Furthermore, it clarifies prevailing supervision strategies, benchmark datasets, and evaluation metrics, and concludes with a critical discussion of current limitations and promising future directions, offering a cohesive reference for the growing field of learning-driven 3D texture generation.
本文针对网格纹理压缩中UV映射的局限性,提出了一种基于表面对齐纹理场TexF的方法,通过稀疏体素组织纹理属性,并结合3DNTC技术实现高效压缩。
AI-generated meshes often exhibit noise, non-manifold topology, and geometric degeneracies, causing existing UV unwrapping methods to produce highly fragmented charts, excessive seams, and uncontrolled distortion. To address this, we propose the first semantic-aware UV unfolding framework that jointly integrates learning-driven part decomposition with geometry-informed heuristics. Our method employs PartField for top-down recursive part segmentation, followed by parameterization optimization and multi-chart packing—minimizing chart count under user-specified distortion constraints. It robustly handles non-manifold and degenerate inputs and achieves, for the first time, part-level semantic alignment alongside low-fragmentation unfolding. Evaluated on multiple benchmarks, our approach significantly outperforms both classical tools and neural methods: chart count is reduced by 32–57%, seam length by 41–63%, distortion is better controlled, and success rate on complex meshes reaches 98.5%.
This study addresses the reliance on manual intervention for seam planning in production-level automatic UV unwrapping of quadrilateral meshes. We propose a training-free agent-based method that leverages vision-language models (VLMs) integrated with domain knowledge. By employing query-based mesh representations and a domain-specific language (DSL), our approach decouples high-level intent planning from low-level edge selection, while introducing a feedback loop mechanism to iteratively refine seams. This design ensures compatibility across backend VLMs and enables scalability to extremely large meshes. Experimental results demonstrate that the proposed method reduces the number of charts by 2.9× and shortens seam length by 1.63×, achieving an 80.9% preference rate among professional artists.
Existing UV unwrapping methods suffer from high computational cost, fragmented UV islands, semantic inconsistency, and irregular boundaries—hindering artist-driven texture editing. This paper introduces the first end-to-end generative framework for producing high-quality, artist-style UV maps in two stages: (1) a semantic-aware SeamGPT model predicts topologically sound, semantically coherent, and structurally clean seams; (2) an optimization-initialized, autoencoder-guided parameterization jointly minimizes overlap, distortion, and semantic discontinuity while maximizing UV space utilization. Experiments demonstrate significant improvements in topological fidelity and editability across multiple benchmarks. The resulting UV maps are directly compatible with professional rendering pipelines and support rapid standalone generation. Our approach establishes a new paradigm for semantics-controllable, automatic UV unfolding.
To address geometric distortion and inaccurate recovery of complex textures (e.g., text, portraits) in multi-view image-based 3D mesh reconstruction, this paper proposes a geometry-texture co-optimization framework. First, we enhance the LRM architecture with differentiable Dual Contouring to enable full-resolution geometric supervision. Second, we introduce a rendering-driven NeRF fine-tuning mechanism to improve surface detail modeling. Third, we propose a lightweight, instance-aware texture refinement module—achieving high-fidelity texture recovery in just 4 seconds while preserving feed-forward inference speed. Built upon triplane representation and differentiable rendering, our method achieves a PSNR of 29.79 on the GSO dataset—setting a new state-of-the-art—and significantly improves both 2D image fidelity and 3D geometric accuracy. Moreover, it natively supports text- or image-to-3D generation.
本文提出ExMesh++,通过自适应顶点分割与合并及UV一致性的保持,从多视图图像重建可重照明的UV-PBR网格资产,优化几何、材质和光照。
This study addresses the baked-in illumination and structural inconsistencies prevalent in single-image 3D garment texture generation, which hinder physical simulation and relighting. We propose synthesizing globally consistent textures within the 2D pattern space. Our approach leverages vision-language models for initialization and employs a diffusion Transformer conditioned on 3D positional features to refine textures in the UV domain, effectively eliminating artifacts. Furthermore, we design an automated synthetic data engine to train the conditional diffusion model, enabling standardized texture extraction free from illumination contamination. The proposed method significantly outperforms existing baselines, generating high-fidelity, physically simulatable 3D garment assets suitable for downstream applications.
Traditional UV unwrapping methods struggle to simultaneously minimize geometric distortion and satisfy artists’ stylistic preferences, such as straight seams and axis-aligned UV islands. This work formulates UV unwrapping for the first time as an end-to-end flow-matching generative problem, learning a mesh-conditioned transport process that maps noise to artist-style UV layouts, thereby producing diverse, production-ready results. To bridge the gap between geometric fidelity and artistic style, the authors introduce a boundary-aware loss and a model-in-the-loop fine-tuning mechanism. Evaluated on a large-scale professional dataset, the proposed method significantly outperforms existing approaches, generating notably straighter seams and more compact, axis-aligned UV islands while maintaining low distortion—results that achieve strong approval from professional artists.
This work addresses the challenge of generating semantic regions for 3D asset segmentation, which traditionally relies on manual intervention and struggles to integrate into interactive content creation pipelines. The authors propose a human-in-the-loop approach for producing editable semantic texture atlases by leveraging multi-view rendering and interactive 2D segmentation—combining SAM² with Label Studio—and back-projecting the results into UV space. A greedy set cover strategy is employed to select key views, enhancing computational efficiency. This method delivers the first unified, editable semantic atlas tailored for XR and game development workflows, enabling downstream tasks such as material assignment and style transfer. Experiments on eight cultural heritage objects demonstrate its effectiveness in handling complex geometries and accurately identifying fine details, cavities, and weak boundaries that require human refinement.
This study addresses the issue in 3D mesh UV texture generation where attention mechanisms are bound to 2D UV grids rather than 3D surface geometry, resulting in seam discontinuities and missing textures in occluded regions. To overcome this, we propose an image-conditioned latent UV-space diffusion framework that employs a Diffusion Transformer for denoising within a pretrained VAE latent space. The core innovation is Surface-Aware Positional Encoding (SAPE), which replaces standard 2D encodings with 3D surface coordinates to enable attention interactions based on 3D proximity, complemented by a multi-level attention head allocation strategy. This approach effectively restores consistency across seams and disconnected islands, generating sharper and globally coherent textures. It significantly outperforms baseline methods in occluded and unseen-view regions, successfully mitigating the gaps and stretching artifacts inherent in projection-based techniques.