CPU3DAI tools for 3D, video and audio

XTTS: Audio Generator — What It Does and What You Need to Run It

XTTS is an open source speech generator from Coqui. It handles text-to-speech: it turns text into spoken audio. The project is developed in a GitHub repository and is distributed as a Python library.

What it does

The authors describe the project as a set of deep learning tools for speech synthesis that, according to them, has been validated in research and in production environments. The repository tags list multi-speaker-tts, voice-cloning and voice-conversion — that is, work with multiple voices, voice cloning and voice conversion. The architectures glow-tts, hifigan, melgan, tacotron and the speaker-encoder and vocoder components are also mentioned.

The XTTS-v2 model is hosted on Hugging Face. The total size of the model weights files is 1.9 GB, with the largest file taking up 1.7 GB. The weights are distributed under the Coqui Public Model License 1.0.0, which permits non-commercial use only.

What you need to run it

The code is written in Python. The code license is MPL-2.0. The latest release is v0.22.0, and the last code change was on August 16, 2024. The repository was created on May 20, 2020.

As described by the author, the platform is cuda. The description references a guide for installing with CUDA on Windows written by user GuyPaddock. The authors do not specify memory requirements or specific CUDA versions.

Who it suits

The project will suit those looking for an open tool for speech synthesis in Python who are ready to figure out installation and running it themselves. The presence of the voice-cloning and voice-conversion tags may be of interest to those working with multiple voices or voice transfer tasks.

It will not suit those expecting a ready-made application with an interface: this is a library and a model, not a boxed product. It is also worth noting that the Coqui Public Model License 1.0.0 for the weights permits non-commercial use only. The authors do not specify memory requirements, so it is hard to estimate the minimum configuration in advance.

XTTS is an open research tool for speech synthesis that has been in development since 2020 and continues to receive updates. It provides access to the model and the code, but requires technical expertise and attention to the terms of the weights license.

XTTS pipeline — audio and speech generator Coqui · open source · text-to-speech Input text text input arbitrary string Text analysis tokenization normalization Acoustic model neural network pytorch Vocoder waveform synthesis hifigan Sound speech audio Voice cloning speaker-encoder voice-cloning Model: XTTS-v2 · Weights: 1.7 GB · Platform: cuda · Code license: MPL-2.0 Library: coqui · Language: Python · Latest release: v0.22.0
How the XTTS pipeline works. The diagram is drawn from the tool’s fact sheet.

Fact sheet

Tasktext-to-speech source
Code licenseMPL-2.0 source
Weights licenseCoqui Public Model License 1.0.0: non-commercial use only source
Platformcuda per the author’s description source
Largest weights file1.7 GB source
All weights files1.9 GB source
Librarycoqui source
LanguagePython source
Last code change2024-08-16 source
Repository created2020-05-20 source
Changes often — as of 2026-10-03
Latest releasev0.22.0 source
Release date2023-12-12 source
GitHub stars46100 source
Forks6167 source
Downloads per month6656524 source
Model updated2023-12-11 source

Values are collected automatically from official sources and were checked on 2026-10-03. Each one links to its source, and values taken from the developer’s pages also carry a verbatim quote — hover over the note. Pricing and versions are shown as of the check date and change most often; verify on the vendor’s site before buying.

See also