arXiv · 2609.24208
CoaG: Cylinders on a Grid: Coarse 3D Layout Control for Video Generation
Abstract
We ask how little geometry a person has to draw to control both where people stand and where the camera moves in a generated video. Our answer is a ground plane and one cylinder per person. A user draws a grid on the ground, places one cylinder where each person should stand, moves the cylinders and the camera over 81 frames, and the model renders a photoreal video in which the people occupy the cylinders' positions, move as the cylinders move, and are seen from the drawn camera. Appearance comes from a text prompt and a background reference image; layout and motion come from the geometry. Because no dataset pairs such a signal with video, we build the pairs ourselves: an automatic engine writes 2000 captions from a combinatorial seed, generates a clip for each with a text-to-video model, and lifts every clip back to its geometry with person tracking, background inpainting, an agentic ground-mask loop, feed-forward multi-view reconstruction and a plane fit, with no real footage and no manual labels. A LoRA on Wan2.2-Fun-Control trained on 1935 such tuples follows drawn layouts and camera paths on hold-out clips: the generated people match the cylinders' count, order, position and height, the text changes who they are, the reference image changes where they are, and dolly-in, orbit, pan and crane paths are followed, dolly-out only weakly.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhangsihao Yang, Mengyi Shan. 2026-09-21. CoaG: Cylinders on a Grid: Coarse 3D Layout Control for Video Generation. https://arxiv.org/abs/2609.24208
Cite the original work for its findings. Save a collection to share your selection of sources.