A Survey of Post-Training and Alignment Methods for Video Generators on arXiv
A survey has been published on arXiv: Video Generation Models: A Survey of Post-Training and Alignment. The authors systematize approaches to post-training video models that do not require retraining from scratch: supervised fine-tuning, self-training and distillation, preference- and reward-based methods, as well as inference-time methods. Datasets, benchmarks and evaluation practices are covered separately. The code is released as open source.
What it means
The survey concerns the entire topic of video generation, which we collect in the AI Video section. Of the tools in our reference, three open-source generators are directly affected.
Maestro is a local studio built on the WanGP pipeline; the MiniMax H3 weights take up 464.1 GB, the largest file is 9.7 GB. The author states a VRAM requirement of 12—24 GB and the CUDA platform. The WanGP code license permits only non-commercial use of the program itself; the MiniMax H3 weights allow commercial use up to 20 million dollars in annual revenue, but do not apply in the EU, the UK, Korea and the US.
LTX-Video from Lightricks is stated to require only 1 GB of VRAM and supports MPS on macOS. The weights take up 236.4 GB, the largest file is 26.6 GB. The weights license permits commercial use, but companies with revenue from $10 million need a separate paid license.
CogVideoX from Zhipu AI solves the text-to-video task; the 5b model weights take up 20 GB, the largest file is 9.2 GB. Commercial use requires registration, free up to 1 million visits per month.
The survey is not tied to specific implementations and does not report which of the listed methods are already used in these tools. The practical value for users of local generators lies in systematizing approaches that may influence the next versions of models and post-processing pipelines.How the method works. The diagram was drawn for this news note.