CPU3DAI tools for 3D, video and audio

Fish Speech: Audio Generator — What It Does and What You Need to Run It

Fish Speech is an open-source speech generator from Fish Audio that handles text-to-speech: it turns text into voiceover. The project publishes both the code and ready-made model weights, so you can run it yourself or study how a modern speech synthesis system is built.

What it does

The authors describe Fish Speech as SOTA Open Source TTS, that is, they position it as a cutting-edge open speech synthesis system. It is built on the Dual-AR architecture, which, as described by the developer, is structurally isomorphic to standard large language models. Thanks to this, the model supports SGLang inference acceleration mechanisms: Continuous Batching, Paged KV Cache, CUDA Graph and RadixAttention-based Prefix Caching.

The repository lists the tags llama, transformer, tts, valle, vits, vqgan and vqvae — this gives an idea of the technology stack: transformer architecture, vector quantization and approaches typical of neural speech synthesis.

What you need to run it

The code is written in Python. The repository was created in October 2023, the last code change was on August 3, 2026. The latest release is v1.5.1.

The model weights are published on Hugging Face. The largest weights file is 1.2 GB, and all weights files together are 1.4 GB. The weights license is cc-by-nc-sa-4.0: you may use them with attribution, for non-commercial purposes and keeping the same license for derivative works. The code license is the Fish Audio Research License: free use, reproduction, distribution and creation of derivative works are permitted for research and non-commercial purposes, while commercial use requires a separate written agreement with Fish Audio.

The platform, as described by the author, is CUDA. This is the developer's claim, not the result of independent testing: the description states that the model supports CUDA acceleration through SGLang mechanisms.

Who it suits

Fish Speech will suit those looking for an open speech synthesis system with public weights and the option of running on your own hardware. The project will be of interest to developers who want to understand the architecture of modern TTS models or integrate speech generation into their pipeline, provided the non-commercial weights license is respected.

The project is unlikely to suit those who need commercial voiceover without restrictions: the cc-by-nc-sa-4.0 weights license prohibits commercial use. It is also worth noting that, as described by the author, running it requires CUDA — this limits the range of suitable hardware. The Fish Audio Research License for the code permits free use only for research and non-commercial purposes, while commercial use requires a separate written agreement with Fish Audio.

Fish Speech is an open speech synthesis system with public weights and a transparent architecture. The project has been in development since 2023, has a current release v1.5.1 and relies on approaches typical of large language models. The cc-by-nc-sa-4.0 weights license limits commercial use, while the Fish Audio Research License for the code permits free use only for research and non-commercial purposes, requiring a separate written agreement for commercial use.

Fish Speech — audio and speech generator pipeline Text input data Tokenization splitting into tokens Dual-AR feature generation Vocoder audio synthesis Result — audio file synthesized speech export formats: not specified in fact sheet Input Processing Generation Synthesis
How the Fish Speech pipeline works. The diagram is drawn from the tool’s fact sheet.

Fact sheet

Tasktext-to-speech source
Code licenseFish Audio Research License: commercial use requires a separate license source
Weights licensecc-by-nc-sa-4.0 source
PlatformCUDA per the author’s description source
Largest weights file1.2 GB source
All weights files1.4 GB source
LanguagePython source
Last code change2026-09-16 source
Repository created2023-10-10 source
Changes often — as of 2026-10-03
Latest releasev1.5.1 source
Release date2025-05-31 source
GitHub stars32925 source
Forks2845 source
Downloads per month3519 source
Model updated2025-03-25 source

Values are collected automatically from official sources and were checked on 2026-10-03. Each one links to its source, and values taken from the developer’s pages also carry a verbatim quote — hover over the note. Pricing and versions are shown as of the check date and change most often; verify on the vendor’s site before buying.

See also