MeshCarve: Artisan Mesh Generation with Flow Matching in Compact Latent Spaces

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of autoregressive 3D mesh generation and the compression quality degradation caused by joint VAE encoding. To this end, we propose MeshCarve, a framework that introduces the first full-stage spatially compressed latent generation paradigm. By decoupling vertices from edge connectivity, MeshCarve employs separate VertexVAE and EdgeVAE modules for spatial-aware compression. Furthermore, it integrates a hierarchical sparse Transformer with vertex-link encoding to achieve efficient generation via flow matching within a compact latent space. Experimental results demonstrate that MeshCarve substantially reduces token sequence length while outperforming existing state-of-the-art methods on benchmarks such as Objaverse, exhibiting strong generalization capabilities across diverse datasets.
📝 Abstract
Prior artisan mesh generation works largely predict face tokens autoregressively, which makes inference slow. Recent methods instead flow match continuous latents built by Variational AutoEncoders (VAEs), but reconstruction quality drops significantly when geometry and topology are jointly encoded, and further when the latent space is compressed. We present MeshCarve, a flow matching method that generates entirely in compact latent spaces, generating vertex positions and edge connections separately and sidestepping the difficulty of a joint compact latent. To shorten the token sequence, we propose a hierarchical sparse transformer backbone, instantiated as VertexVAE and EdgeVAE. Instead of encoding fields over the surface voxels, both VAEs anchor on discrete vertices in their latent spaces, which drastically reduces the token sequence length, and our spatial-aware compression shortens it further without costing reconstruction. VertexVAE directly encodes vertex occupancy. For connectivity, we propose vertex-link encoding, which turns arbitrary connectivity between vertices into fixed-length continuous per-vertex embeddings and recovers complex artistic topology faithfully. MeshCarve combines these VAEs with an anchor generator and flow matches on the shortened token sequences. It shows advantages over state-of-the-art autoregressive and flow matching methods on Objaverse and generalizes to Toys4K. To the best of our knowledge, it is among the first artisan mesh generation methods whose every generative stage runs in a spatially compressed latent, with a token sequence only a fraction of the most compressed previous autoregressive and flow matching works.
Problem

Research questions and friction points this paper is trying to address.

artisan mesh generation
flow matching
compact latent space
reconstruction quality
inference speed
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow Matching
Compact Latent Space
Artisan Mesh Generation
Hierarchical Sparse Transformer
Vertex-Link Encoding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiyu Wang
Nanyang Technological University
R
Ruocheng Wu
SparAI Inc.
Y
Yufei Wang
SparAI Inc.
Zhihao Li
Zhihao Li
Nanyang Technological University
ISPimage/video coding
L
Lanqing Guo
University of Texas at Austin
Bihan Wen
Bihan Wen
Associate Professor, Nanyang Technological University
Machine LearningImage ProcessingComputational ImagingComputer VisionTrustworthy AI