CPU3DAI tools for 3D, video and audio

Generative Cinematographer Teaches a Video Model to Control Camera and Objects in 3D

On arXiv, on October 1, 2026, a paper was published on Generative Cinematographer (GenCine) — a system that lifts a single image into an editable 3D scene and allows joint specification of camera and foreground motion. The artist lays out a camera path and moves selected foreground regions with local 3D handles that can independently move different parts of an object. The control signals are projected into guidance maps: each handle gets a fixed color, and its 3D positions are encoded in the world coordinate system of the background, which describes the object's motion relative to the scene even with a moving camera. The authors train a lightweight guidance branch and LoRA adapters on the pretrained Wan model. The authors did not release the code. More details in the paper on arXiv.

What it means

GenCine is built on top of Wan — Alibaba's open-source video generator with a text-to-video task. Wan's code is distributed under the Apache-2.0 license, and the weights are also Apache-2.0. The largest model in the repository takes up 10.6 GB in a single file, and all weight files total 64.3 GB. The Wan authors note that the T2V-1.3B model requires 8.19 GB of VRAM and generates a five-second 480P clip on an RTX 4090 in about four minutes without optimizations such as quantization. The GenCine paper does not specify VRAM requirements for training the guidance branch and LoRA adapters, does not state the size of the fine-tuned model, and does not describe the conditions under which the work could be used. The code has not been published, so the system cannot be reproduced from the paper. For the reader of the reference, this means that there is no practical tool based on GenCine yet, and comparison with Wan's current capabilities remains at the level of the method description. The section on video generators and other tools running on your own hardware is collected on the AI Video page.
Generative Cinematographer (GenCine) One image — editable 3D scene — video with controllable camera One image input frame lifted into 3D 3D scene editable world coordinate system 3D handles camera path foreground motion Guidance maps handle color = signal 3D positions in world coordinates Wan + LoRA pretrained model guidance branch LoRA adapters Video controllable camera object motion Code not released · no practical tool · comparison with Wan is at the level of the method description
How the method works. The diagram was drawn based on this news note.

See also