🤖 AI Summary
This study addresses the issue in 3D mesh UV texture generation where attention mechanisms are bound to 2D UV grids rather than 3D surface geometry, resulting in seam discontinuities and missing textures in occluded regions. To overcome this, we propose an image-conditioned latent UV-space diffusion framework that employs a Diffusion Transformer for denoising within a pretrained VAE latent space. The core innovation is Surface-Aware Positional Encoding (SAPE), which replaces standard 2D encodings with 3D surface coordinates to enable attention interactions based on 3D proximity, complemented by a multi-level attention head allocation strategy. This approach effectively restores consistency across seams and disconnected islands, generating sharper and globally coherent textures. It significantly outperforms baseline methods in occluded and unseen-view regions, successfully mitigating the gaps and stretching artifacts inherent in projection-based techniques.
📝 Abstract
Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D information, leaving the attention mechanism tied to UV-grid positions rather than to the underlying surface geometry. This mismatch limits coherence across seams and disconnected UV islands.
We propose DirectUV, an image-conditioned UV texture diffusion framework that operates in the latent UV space of a pretrained image VAE, in which a Diffusion Transformer denoises the UV latent given a single input image and a coarse UV map. At its core, Surface-Aware Positional Encoding (SAPE) replaces the standard 2D-grid positional encoding with encodings derived from per-token 3D surface coordinates obtained via UV-to-surface correspondence. As positional encodings define the distance metric used by attention, SAPE enables tokens to interact according to 3D positional proximity derived from surface correspondence rather than UV-grid distance, restoring coherence across seams and disconnected islands. A multi-level extension further assigns different attention heads to progressively finer subdivisions of the same latent UV patch, allowing the model to reason about surface structure at multiple granularities.
Experiments show that DirectUV produces sharper and more globally consistent textures than other baselines, with the largest improvements in occluded and view-unseen regions where projection-based methods leave gaps or stretched textures.