CogVideoX: Video Generator — What It Does and What You Need to Run It

Here we have CogVideoX — an open-source text-to-video generator. The developer, Zhipu AI, positions it as a tool for turning text prompts into video clips. The model is also mentioned in connection with image-to-video generation, although the main task in the fact sheet is listed as text-to-video.
What it does
The authors describe CogVideoX as a system for generating video from text and images. In the repository, the project is tagged with text-to-video and image-to-video, which points to two operating modes. Among the keywords are also llm and sora — this makes it clear that language models are at the core, and the technology itself belongs to the class of solutions similar to Sora. There are no details in the fact sheet about the maximum clip length or supported resolutions.
What you need to run it
The code is written in Python and distributed under the Apache-2.0 license. The model weights carry The CogVideoX License: commercial use after registration, up to 1 million visits per month. Running it requires an NVIDIA GPU with CUDA support. Verbatim from the authors' description: devices with NVIDIA H100 and above, and you need to install the torch and torchao packages from source. The recommended CUDA version is 12.4. The largest weights file is 9.2 GB, and the total size of all files is 20 GB. The model is integrated into the diffusers library. Local launch is possible, but the hardware requirements are quite specific.
Who it suits
The tool is aimed at developers and researchers who have access to powerful server accelerators. The open source code and the presence of a GitHub repository make it possible to study the architecture and modify the model for your own tasks. For everyday use on consumer GPUs, judging by the stated requirements, CogVideoX is not intended. Those looking for a cloud service without having to set up an environment should also consider other solutions — here you will need to manually build dependencies and work with Python code.
The model continues to evolve: the last code change is dated November 2025, and the latest release is labeled v1.0. This suggests that the project is in an active support phase, but has not yet accumulated a large number of intermediate releases.
Fact sheet
Repository · Model on HuggingFace
| Task | text-to-video source |
|---|---|
| Code license | Apache-2.0 source |
| Weights license | The CogVideoX License: commercial use after registration, up to 1M visits per month source |
| Platform | CUDA per the author’s description source |
| Largest weights file | 9.2 GB source |
| All weights files | 20 GB source |
| Library | diffusers source |
| Language | Python source |
| Last code change | 2025-11-04 source |
| Repository created | 2022-05-29 source |
| Latest release | v1.0 source |
|---|---|
| Release date | 2024-11-08 source |
| GitHub stars | 12996 source |
| Forks | 1339 source |
| Downloads per month | 17161 source |
| Model updated | 2024-11-23 source |
Values are collected automatically from official sources and were checked on 2026-09-14. Each one links to its source, and values taken from the developer’s pages also carry a verbatim quote — hover over the note. Pricing and versions are shown as of the check date and change most often; verify on the vendor’s site before buying.



