Fish Speech: Audio Generator — What It Does and What You Need to Run It

Fish Speech is an open-source speech generator from Fish Audio that handles text-to-speech: it turns text into voiceover. The project publishes both the code and ready-made model weights, so you can run it yourself or study how a modern speech synthesis system is built.
What it does
The authors describe Fish Speech as SOTA Open Source TTS, that is, they position it as a cutting-edge open speech synthesis system. It is built on the Dual-AR architecture, which, as described by the developer, is structurally isomorphic to standard large language models. Thanks to this, the model supports SGLang inference acceleration mechanisms: Continuous Batching, Paged KV Cache, CUDA Graph and RadixAttention-based Prefix Caching.
The repository lists the tags llama, transformer, tts, valle, vits, vqgan and vqvae — this gives an idea of the technology stack: transformer architecture, vector quantization and approaches typical of neural speech synthesis.
What you need to run it
The code is written in Python. The repository was created in October 2023, the last code change was on August 3, 2026. The latest release is v1.5.1.
The model weights are published on Hugging Face. The largest weights file is 1.2 GB, and all weights files together are 1.4 GB. The weights license is cc-by-nc-sa-4.0: you may use them with attribution, for non-commercial purposes and keeping the same license for derivative works. The code license is the Fish Audio Research License: free use, reproduction, distribution and creation of derivative works are permitted for research and non-commercial purposes, while commercial use requires a separate written agreement with Fish Audio.
The platform, as described by the author, is CUDA. This is the developer's claim, not the result of independent testing: the description states that the model supports CUDA acceleration through SGLang mechanisms.
Who it suits
Fish Speech will suit those looking for an open speech synthesis system with public weights and the option of running on your own hardware. The project will be of interest to developers who want to understand the architecture of modern TTS models or integrate speech generation into their pipeline, provided the non-commercial weights license is respected.
The project is unlikely to suit those who need commercial voiceover without restrictions: the cc-by-nc-sa-4.0 weights license prohibits commercial use. It is also worth noting that, as described by the author, running it requires CUDA — this limits the range of suitable hardware. The Fish Audio Research License for the code permits free use only for research and non-commercial purposes, while commercial use requires a separate written agreement with Fish Audio.
Fish Speech is an open speech synthesis system with public weights and a transparent architecture. The project has been in development since 2023, has a current release v1.5.1 and relies on approaches typical of large language models. The cc-by-nc-sa-4.0 weights license limits commercial use, while the Fish Audio Research License for the code permits free use only for research and non-commercial purposes, requiring a separate written agreement for commercial use.
Fact sheet
Repository · Model on HuggingFace · Developer’s site
| Task | text-to-speech source |
|---|---|
| Code license | Fish Audio Research License: commercial use requires a separate license source |
| Weights license | cc-by-nc-sa-4.0 source |
| Platform | CUDA per the author’s description source |
| Largest weights file | 1.2 GB source |
| All weights files | 1.4 GB source |
| Language | Python source |
| Last code change | 2026-09-16 source |
| Repository created | 2023-10-10 source |
| Latest release | v1.5.1 source |
|---|---|
| Release date | 2025-05-31 source |
| GitHub stars | 32925 source |
| Forks | 2845 source |
| Downloads per month | 3519 source |
| Model updated | 2025-03-25 source |
Values are collected automatically from official sources and were checked on 2026-10-03. Each one links to its source, and values taken from the developer’s pages also carry a verbatim quote — hover over the note. Pricing and versions are shown as of the check date and change most often; verify on the vendor’s site before buying.



