CoaG: Cylinders on a Grid: Coarse 3D Layout Control for Video Generation

๐Ÿ“… 2026-09-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถ้€š่ฟ‡็”จๆˆท็ป˜ๅˆถๅœฐ้ข็ฝ‘ๆ ผๅ’Œไบบ็‰ฉๅœ†ๆŸฑไฝ“ๆฅๆŽงๅˆถ็”Ÿๆˆ่ง†้ข‘ไธญ็š„ไบบ็‰ฉไฝ็ฝฎไธŽ็›ธๆœบ็งปๅŠจ๏ผŒๅˆฉ็”จๆ–‡ๆœฌๆ็คบๅ’Œ่ƒŒๆ™ฏๅ‚่€ƒๅ›พ็”Ÿๆˆ้€ผ็œŸ่ง†้ข‘ใ€‚
๐Ÿ“ Abstract
We ask how little geometry a person has to draw to control both where people stand and where the camera moves in a generated video. Our answer is a ground plane and one cylinder per person. A user draws a grid on the ground, places one cylinder where each person should stand, moves the cylinders and the camera over 81 frames, and the model renders a photoreal video in which the people occupy the cylinders' positions, move as the cylinders move, and are seen from the drawn camera. Appearance comes from a text prompt and a background reference image; layout and motion come from the geometry. Because no dataset pairs such a signal with video, we build the pairs ourselves: an automatic engine writes 2000 captions from a combinatorial seed, generates a clip for each with a text-to-video model, and lifts every clip back to its geometry with person tracking, background inpainting, an agentic ground-mask loop, feed-forward multi-view reconstruction and a plane fit, with no real footage and no manual labels. A LoRA on Wan2.2-Fun-Control trained on 1935 such tuples follows drawn layouts and camera paths on hold-out clips: the generated people match the cylinders' count, order, position and height, the text changes who they are, the reference image changes where they are, and dolly-in, orbit, pan and crane paths are followed, dolly-out only weakly.
Problem

Research questions and friction points this paper is trying to address.

3D Layout Control
Video Generation
Geometry
Camera Path
Person Tracking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cylinders on a Grid
3D Layout Control
Video Generation
Text-to-Video Model
Automatic Data Pairing
๐Ÿ”Ž Similar Papers
No similar papers found.