๐ค AI Summary
็ ็ฉถ้่ฟ็จๆท็ปๅถๅฐ้ข็ฝๆ ผๅไบบ็ฉๅๆฑไฝๆฅๆงๅถ็ๆ่ง้ขไธญ็ไบบ็ฉไฝ็ฝฎไธ็ธๆบ็งปๅจ๏ผๅฉ็จๆๆฌๆ็คบๅ่ๆฏๅ่ๅพ็ๆ้ผ็่ง้ขใ
๐ Abstract
We ask how little geometry a person has to draw to control both where people stand and where the camera moves in a generated video. Our answer is a ground plane and one cylinder per person. A user draws a grid on the ground, places one cylinder where each person should stand, moves the cylinders and the camera over 81 frames, and the model renders a photoreal video in which the people occupy the cylinders' positions, move as the cylinders move, and are seen from the drawn camera. Appearance comes from a text prompt and a background reference image; layout and motion come from the geometry. Because no dataset pairs such a signal with video, we build the pairs ourselves: an automatic engine writes 2000 captions from a combinatorial seed, generates a clip for each with a text-to-video model, and lifts every clip back to its geometry with person tracking, background inpainting, an agentic ground-mask loop, feed-forward multi-view reconstruction and a plane fit, with no real footage and no manual labels. A LoRA on Wan2.2-Fun-Control trained on 1935 such tuples follows drawn layouts and camera paths on hold-out clips: the generated people match the cylinders' count, order, position and height, the text changes who they are, the reference image changes where they are, and dolly-in, orbit, pan and crane paths are followed, dolly-out only weakly.