SparkDiffusion: video generator — what it does and what you need to run it

SparkDiffusion is an open-source text-to-video generator from Alibaba Research. The project tackles the text-to-video task and, as described by the authors, speeds up video generation based on diffusion transformers by 265x through the combined use of sparsity, distillation and quantization.
What it does
The tool takes a text description and creates video. Its technical foundation is the SparkWan2.1-T2V-1.3B-480P-0.90Sparsity model, published on Hugging Face. The authors describe the approach as accelerating video generation through a combination of three techniques: sparsity, distillation and quantization. The description mentions a sparsity level of 0.90 and the w8a8 quantization format.
The code is written in Python. Among the repository tags are crossdistill, diffusion-models, distillation, high-sparsity, inference-acceleration, rola, triton and video-generation.
What you need to run it
As described by the author, running it requires Linux with a GPU that supports CUDA. The Python version is 3.10 or newer. The total size of the model weights is 2.6 GB, and the largest file weighs the same.
The code is distributed under the Apache-2.0 license. The model weights are also published under the Apache-2.0 license. The license permits use, modification and distribution of both the source code and the trained model.
Who it suits
The project will be useful to those working with text-to-video generation who are looking for an open-source implementation with an emphasis on inference acceleration. The publicly available weights let you run the model without training it yourself. The Apache-2.0 license on the code and weights is suitable for commercial and research tasks, including embedding into your own products.
The tool is not aimed at users without experience with the command line and a Python environment. The requirement for a CUDA-compatible GPU on Linux limits its use on machines without a discrete graphics card or on other operating systems. The authors do not specify a minimum amount of VRAM, so it is hard to assess compatibility with a specific graphics card in advance.
The repository was created on September 18, 2026, and the last code change was on October 5, 2026. The project is young, and at the time of this description activity is concentrated in a short time window.
SparkDiffusion is an open-source text-to-video implementation with claimed inference acceleration and available weights. Running it requires a Linux machine with a CUDA-compatible GPU and Python 3.10 or newer. The project is aimed at developers and researchers who need an open-source codebase and a model under the Apache-2.0 license.
Fact sheet
Repository · Model on HuggingFace
| Task | text-to-video source |
|---|---|
| Code license | Apache-2.0 source |
| Weights license | apache-2.0 source |
| Platform | CUDA per the author’s description source |
| Largest weights file | 2.6 GB source |
| All weights files | 2.6 GB source |
| Python | 3.10 per the author’s description source |
| Language | Python source |
| Last code change | 2026-10-05 source |
| Repository created | 2026-09-18 source |
| GitHub stars | 520 source |
|---|---|
| Forks | 66 source |
| Downloads per month | 0 source |
| Model updated | 2026-09-24 source |
Values are collected automatically from official sources and were checked on 2026-10-08. Each one links to its source, and values taken from the developer’s pages also carry a verbatim quote — hover over the note. Pricing and versions are shown as of the check date and change most often; verify on the vendor’s site before buying.



