CPU3DAI tools for 3D, video and audio

SparkDiffusion: video generator — what it does and what you need to run it

SparkDiffusion is an open-source text-to-video generator from Alibaba Research. The project tackles the text-to-video task and, as described by the authors, speeds up video generation based on diffusion transformers by 265x through the combined use of sparsity, distillation and quantization.

What it does

The tool takes a text description and creates video. Its technical foundation is the SparkWan2.1-T2V-1.3B-480P-0.90Sparsity model, published on Hugging Face. The authors describe the approach as accelerating video generation through a combination of three techniques: sparsity, distillation and quantization. The description mentions a sparsity level of 0.90 and the w8a8 quantization format.

The code is written in Python. Among the repository tags are crossdistill, diffusion-models, distillation, high-sparsity, inference-acceleration, rola, triton and video-generation.

What you need to run it

As described by the author, running it requires Linux with a GPU that supports CUDA. The Python version is 3.10 or newer. The total size of the model weights is 2.6 GB, and the largest file weighs the same.

The code is distributed under the Apache-2.0 license. The model weights are also published under the Apache-2.0 license. The license permits use, modification and distribution of both the source code and the trained model.

Who it suits

The project will be useful to those working with text-to-video generation who are looking for an open-source implementation with an emphasis on inference acceleration. The publicly available weights let you run the model without training it yourself. The Apache-2.0 license on the code and weights is suitable for commercial and research tasks, including embedding into your own products.

The tool is not aimed at users without experience with the command line and a Python environment. The requirement for a CUDA-compatible GPU on Linux limits its use on machines without a discrete graphics card or on other operating systems. The authors do not specify a minimum amount of VRAM, so it is hard to assess compatibility with a specific graphics card in advance.

The repository was created on September 18, 2026, and the last code change was on October 5, 2026. The project is young, and at the time of this description activity is concentrated in a short time window.

SparkDiffusion is an open-source text-to-video implementation with claimed inference acceleration and available weights. Running it requires a Linux machine with a CUDA-compatible GPU and Python 3.10 or newer. The project is aimed at developers and researchers who need an open-source codebase and a model under the Apache-2.0 license.

SparkDiffusion — video generation pipeline Text prompt scene description in Russian Text encoder conversion to embeddings Diffusion model frame generation from noise Inference acceleration sparsity quantization Video file finished clip Pipeline: text prompt → embeddings → diffusion generation → acceleration → video file Acceleration technologies • Sparsity • Distillation • Quantization Model parameters • Model: SparkWan2.1-T2V • Weight size: 2.6 GB • Platform: CUDA
How the SparkDiffusion pipeline works. The diagram is drawn from the tool’s fact sheet.

Fact sheet

Tasktext-to-video source
Code licenseApache-2.0 source
Weights licenseapache-2.0 source
PlatformCUDA per the author’s description source
Largest weights file2.6 GB source
All weights files2.6 GB source
Python3.10 per the author’s description source
LanguagePython source
Last code change2026-10-05 source
Repository created2026-09-18 source
Changes often — as of 2026-10-08
GitHub stars520 source
Forks66 source
Downloads per month0 source
Model updated2026-09-24 source

Values are collected automatically from official sources and were checked on 2026-10-08. Each one links to its source, and values taken from the developer’s pages also carry a verbatim quote — hover over the note. Pricing and versions are shown as of the check date and change most often; verify on the vendor’s site before buying.

See also