🤖 AI Summary
This study addresses the challenge of generating temporally evolving dynamic 3D textures while preserving geometric consistency, a limitation of existing methods. To this end, this work presents the first approach to propagate video dynamics onto fixed geometric meshes. By encoding the geometry once to ensure constancy, it integrates a conditional video diffusion model with an image-to-3D generator to achieve dynamic texture evolution. Furthermore, a temporal window blending strategy and Low-Rank Adaptation (LoRA) fine-tuning are introduced to recover high-frequency details and eliminate flickering artifacts. The proposed method significantly outperforms existing video-to-4D and texture generation baselines, demonstrating generalization capability to unseen shapes and achieving high-quality, flicker-free dynamic texture synthesis.
📝 Abstract
We present DynaMesh, a dynamic texture generation method for 3D meshes. Given a textureless shape and a text prompt describing an effect, our method produces an appearance that evolves while the object's geometry remains unchanged. Previous works on dynamic 3D content generation have focused on motion, where an object's geometry and position change while keeping its appearance the same. Methods on texture generation sit on the other side of the problem, painting appearance onto a shape as a fixed surface property and not as an evolving process. Neither addresses a visual effect that propagates on a 3D object. A natural route consists of two generators: a video model that shows the effect from a single view, and an image-to-3D generator that lifts each frame to 3D. However, the latter has no notion of time, so running it per video frame produces a sequence that flickers, loses effect details, and yields a different mesh at every video frame. Our method addresses these failures by conditioning a video model on a render of the mesh and the prompt to obtain a reference video, then running a frozen image-to-3D generator on the video with two changes. The conditioning of each frame is blended over a temporal window, and low-rank adapters are fit per shape to restore the lost details. The mesh is encoded once for the whole sequence, so geometry is constant by construction, and the output is a single mesh with a texture per frame. Applied to various objects and effects, DynaMesh substantially improves over recent video-to-4D and texturing methods, and can generalize its temporal effect to different shapes never seen during training. Our project page is at https://threedle.github.io/dynamesh/.