Meshy T2: Fast Native Mesh Generation with Flow Matching

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing autoregressive methods for 3D mesh generation suffer from slow inference and error accumulation, making them ill-suited for real-time creation of high-quality meshes with artist-friendly topology. This work proposes the first native mesh generation framework based on flow matching: it employs a vertex-set mesh VAE to encode meshes into continuous latent representations and introduces a two-stage cascaded flow matching model that enables single-pass, parallel decoding from image to complete mesh. The approach eliminates the need for vertex quantization and welding, supports interactive generation (median time of 6 seconds), explicit control over face count, and multi-part asset modeling. It significantly outperforms current autoregressive baselines in both geometric fidelity and generation efficiency, achieving speedups of over an order of magnitude.
📝 Abstract
Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes with artist-style topology is essential for film, gaming, and interactive 3D applications. Mainstream approaches serialize a mesh into a token sequence and decode it autoregressively, which is slow at inference and sensitive to error accumulation, making them impractical for interactive asset creation. We present Meshy T2, a fast native mesh generation framework built on flow matching. At its core is a vertex-set mesh VAE that encodes a mesh into one continuous latent token per vertex and decodes vertices, edge connectivity, and face winding order in a single pass, preserving high-precision geometry and artist-authored topology without vertex quantization or welding. Generation proceeds as a coarse-to-fine cascade of two flow-matching models: an image-conditioned voxel flow first sketches the overall shape as a coarse occupancy scaffold, and a mesh flow then populates the scaffold with per-vertex latent tokens, conditioned on the image, the scaffold, and a requested vertex budget. This design delivers three practical capabilities: interactive generation speed through parallel flow-based synthesis; effective face-count control through the requested vertex budget; and native support for multi-part assets, whose components emerge directly from the generated connectivity. In our experiments, Meshy T2 achieves state-of-the-art geometric fidelity and completes end-to-end image-to-mesh generation within a median of 6 seconds, over an order of magnitude faster than autoregressive baselines. Code and weights will be available at https://github.com/meshy-dev/meshy-t2.
Problem

Research questions and friction points this paper is trying to address.

mesh generation
autoregressive decoding
interactive 3D
artist-style topology
inference speed
Innovation

Methods, ideas, or system contributions that make the work stand out.

flow matching
native mesh generation
vertex-set VAE
coarse-to-fine cascade
interactive 3D generation
🔎 Similar Papers
No similar papers found.