Position Forcing: Self-Conditioning 3D Generation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the quality limitations of single-stage VecSet 3D generation models caused by the absence of explicit positional guidance. To overcome this, we propose a position enforcement mechanism that leverages latent space correspondences to recover and quantify positional information during denoising, feeding it back into the Transformer. This enables coarse-to-fine self-conditioned generation, providing progressive spatial guidance without requiring an independent positional stage. Built upon a diffusion Transformer with adaptive-resolution positional encoding, our method substantially improves generation quality, outperforming both mainstream multi-stage pipelines and comparable single-stage approaches.
📝 Abstract
Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation quality. We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences. Building on this observation, we propose Position Forcing, a position-based self-conditioning framework. During denoising, Position Forcing recovers token positions from the current clean latent estimate, quantizes them at progressively finer resolutions according to the denoising stage, and feeds the resulting positional encodings back into the diffusion Transformer. This progressively refined positional feedback provides spatial guidance at a granularity appropriate to each denoising stage, guiding shape generation along a coarse-to-fine trajectory and substantially improving generation quality without a separate position generation stage. Experiments demonstrate that Position Forcing achieves strong performance among single-stage 3D generative methods and outperforms several competitive multi-stage approaches.
Problem

Research questions and friction points this paper is trying to address.

3D generation
VecSet representation
single-stage model
positional guidance
denoising
Innovation

Methods, ideas, or system contributions that make the work stand out.

Position Forcing
Self-Conditioning
3D Generation
VecSet Representation
Diffusion Transformer
🔎 Similar Papers
No similar papers found.