Stable Audio Open: Audio Generator — What It Does and What You Need to Run It

Stable Audio Open is an open source audio generator from Stability AI. It handles the text-to-audio task: it turns a text description into an audio file. The tool will suit those who need text-based audio generation and are ready to work with code and model weights directly.
What it does
The authors describe the project as generative models for conditional audio generation. The task the model handles is text-to-audio: text goes in, audio comes out. No other modes of operation, such as speech generation or music generation from notes, are stated in the fact sheet.
The code is written in Python. The last change in the repository was on August 9, 2026; the repository itself was created on May 23, 2023.
What you need to run it
It requires the stable-audio-tools library. The largest weights file is 4.5 GB, and all weights files together are 14.6 GB. This is worth keeping in mind when planning disk space.
As described by the author, Python 3.10 is used for development. The repository description states: "Requires PyTorch 2.5 or later for Flash Attention and Flex Attention support. Development for the repo is done in Python 3.10". That is, PyTorch 2.5 or newer is required for Flash Attention and Flex Attention support.
The code is distributed under the MIT license.
Who it suits
The tool is aimed at those who know how to work with Python and aren't afraid to run models through code. The open source and available weights let you understand how the model works and, if needed, modify it for your own tasks.
For those looking for a ready-made application with an interface, this option may not be suitable: the fact sheet has no information about a graphical interface or a web service. Also worth considering is the size of the weights — 14.6 GB in total, which can be noticeable on weak machines.
Stable Audio Open is a tool for developers and researchers who need text-based audio generation and are ready to work with code and weights directly. The code is open, the model is available, but running it requires technical preparation and free disk space.
Fact sheet
Repository · Model on HuggingFace
| Task | text-to-audio source |
|---|---|
| Code license | MIT source |
| Largest weights file | 4.5 GB source |
| All weights files | 14.6 GB source |
| Python | 3.10 per the author’s description source |
| Library | stable-audio-tools source |
| Language | Python source |
| Last code change | 2026-09-18 source |
| Repository created | 2023-05-23 source |
| GitHub stars | 3872 source |
|---|---|
| Forks | 487 source |
| Downloads per month | 20884 source |
| Model updated | 2025-06-19 source |
Values are collected automatically from official sources and were checked on 2026-10-03. Each one links to its source, and values taken from the developer’s pages also carry a verbatim quote — hover over the note. Pricing and versions are shown as of the check date and change most often; verify on the vendor’s site before buying.



