CPU3DAI tools for 3D, video and audio

CogVideoX: Video Generator — What It Does and What You Need to Run It

Here we have CogVideoX — an open-source text-to-video generator. The developer, Zhipu AI, positions it as a tool for turning text prompts into video clips. The model is also mentioned in connection with image-to-video generation, although the main task in the fact sheet is listed as text-to-video.

What it does

The authors describe CogVideoX as a system for generating video from text and images. In the repository, the project is tagged with text-to-video and image-to-video, which points to two operating modes. Among the keywords are also llm and sora — this makes it clear that language models are at the core, and the technology itself belongs to the class of solutions similar to Sora. There are no details in the fact sheet about the maximum clip length or supported resolutions.

What you need to run it

The code is written in Python and distributed under the Apache-2.0 license. The model weights carry The CogVideoX License: commercial use after registration, up to 1 million visits per month. Running it requires an NVIDIA GPU with CUDA support. Verbatim from the authors' description: devices with NVIDIA H100 and above, and you need to install the torch and torchao packages from source. The recommended CUDA version is 12.4. The largest weights file is 9.2 GB, and the total size of all files is 20 GB. The model is integrated into the diffusers library. Local launch is possible, but the hardware requirements are quite specific.

Who it suits

The tool is aimed at developers and researchers who have access to powerful server accelerators. The open source code and the presence of a GitHub repository make it possible to study the architecture and modify the model for your own tasks. For everyday use on consumer GPUs, judging by the stated requirements, CogVideoX is not intended. Those looking for a cloud service without having to set up an environment should also consider other solutions — here you will need to manually build dependencies and work with Python code.

The model continues to evolve: the last code change is dated November 2025, and the latest release is labeled v1.0. This suggests that the project is in an active support phase, but has not yet accumulated a large number of intermediate releases.

CogVideoX pipeline Text- and image-to-video generation INPUT Text prompt Image Prompt and reference frame ENCODING Tokenization Embeddings Text and frame to vectors GENERATION CogVideoX-5b Diffusion Denoising, frames DECODING VAE decoder Frame assembly Latent code to video RESULT Video file Format: MP4 Generated video Modes: Text-to-Video Image-to-Video Platform: CUDA 12.4 Library: diffusers License: Apache-2.0 Weights: 20 GB
How the CogVideoX pipeline works. The diagram is drawn from the tool’s fact sheet.

Fact sheet

Tasktext-to-video source
Code licenseApache-2.0 source
Weights licenseThe CogVideoX License: commercial use after registration, up to 1M visits per month source
PlatformCUDA per the author’s description source
Largest weights file9.2 GB source
All weights files20 GB source
Librarydiffusers source
LanguagePython source
Last code change2025-11-04 source
Repository created2022-05-29 source
Changes often — as of 2026-09-14
Latest releasev1.0 source
Release date2024-11-08 source
GitHub stars12996 source
Forks1339 source
Downloads per month17161 source
Model updated2024-11-23 source

Values are collected automatically from official sources and were checked on 2026-09-14. Each one links to its source, and values taken from the developer’s pages also carry a verbatim quote — hover over the note. Pricing and versions are shown as of the check date and change most often; verify on the vendor’s site before buying.

See also