Stable Video Diffusion: video generator — what it does and what you need to run it

Stable Video Diffusion is an open source generative model from Stability AI that solves the task of turning a static image into a short video (image-to-video). The tool is aimed at researchers and developers who are ready to work with the code directly and run generation on their own hardware.
What it does
- Takes a single image as input and generates a video sequence based on it.
- Implemented as the stable-video-diffusion-img2vid-xt model, available on Hugging Face.
- The code is distributed under the MIT license, which allows free use and modification.
- Integrates with the diffusers library.
- The authors position the repository as a set of generative models rather than as a finished application with a graphical interface.
What you need to run it
- Platform: local launch. No cloud version is listed in the fact sheet.
- Memory and disk: the largest model weights file takes up 8.9 GB, and the full set of files is 30.4 GB. Loading the model into video memory will require a corresponding amount of VRAM.
- Runtime: Python. As described by the author, the code is tested on version 3.10. Verbatim from the description: «NOTE: This is tested under python3.10. For other python versions, you might encounter version conflicts.»
- License: code — MIT, model weights — Stability AI Community License (commercial use free up to $1 million in annual revenue; commercial use requires registration on the Stability AI website; creating or improving other generative models based on this model or its outputs is prohibited).
- Code: available in the repository github.com/Stability-AI/generative-models. The last change is dated December 16, 2025. The latest release is version 0.1.0.
Who it suits
The tool is designed for technically proficient users: researchers, developers and enthusiasts familiar with Python and the diffusers ecosystem. It will suit those who need image-to-video generation with the ability to fine-tune and modify the code.
It will not suit those looking for a ready-made application with an interface, a cloud service without the need for local installation, or a solution that runs on low-power machines without a discrete graphics card with a large amount of memory. Also keep in mind that the Python version is hard-pinned to 3.10 — other versions may have dependency conflicts.
Stable Video Diffusion is a low-level tool in the form of code and model weights. Working with it will require setting up the runtime environment yourself, downloading tens of gigabytes of data and figuring out the specifics of launching it. The authors provide no guarantees of compatibility outside the tested environment.
Fact sheet
Repository · Model on HuggingFace
| Task | image-to-video source |
|---|---|
| Code license | MIT source |
| Weights license | Stability AI Community License: commercial use free up to $1M in annual income source |
| Largest weights file | 8.9 GB source |
| All weights files | 30.4 GB source |
| Python | 3.10 per the author’s description source |
| Library | diffusers source |
| Language | Python source |
| Last code change | 2025-12-16 source |
| Repository created | 2023-06-22 source |
| Latest release | 0.1.0 source |
|---|---|
| Release date | 2023-07-27 source |
| GitHub stars | 27280 source |
| Forks | 3107 source |
| Downloads per month | 212012 source |
| Model updated | 2024-07-10 source |
Values are collected automatically from official sources and were checked on 2026-09-14. Each one links to its source, and values taken from the developer’s pages also carry a verbatim quote — hover over the note. Pricing and versions are shown as of the check date and change most often; verify on the vendor’s site before buying.



