AI Audio and Speech Generators: Open Models
Here you will find models that generate audio: music from text descriptions, sound effects, speech from text, and voice cloning. Each one has a fact sheet: license, memory requirements, platform, and formats. Everything that can be verified programmatically is taken from official sources and marked with a verification date.
This section only includes open-source models — those you can install on your own machine. Russian-language reviews usually stay silent about running them on different hardware, but for audio this is the key question: many models run fine even without a high-end GPU.
Articles
15 articles
YuE: audio generator — what it does and what you need to run itYuE is an open source audio generator from the M-A-P team, which the authors describe as a model for creating full songs. According to the developers…
XTTS: audio generator — what it does and what you need to run itXTTS is an open source speech generator from Coqui. It tackles the text-to-speech task: converting text into spoken audio. The project is developed in…
Stable Audio Open: audio generator — what it does and what you need to run itStable Audio Open is an open source audio generator from Stability AI. It handles the text-to-audio task: it turns a text description into an audio fi…
OpenVoice: audio generator — what it does and what you need to run itOpenVoice is an open source speech generator from the MyShell team. It handles the text-to-speech task and, as described by the authors, belongs to th…
MusicGen: audio generator — what it does and what you need to run itMusicGen is an audio generator from Meta that turns a text description into audio. It is part of the Audiocraft library and handles the text-to-audio…
MOSS-SoundEffect without CUDA (Mac and AMD): audio generator — what it does and what you need to run itMOSS-SoundEffect without CUDA is a desktop application for generating sound effects from text descriptions. The build is aimed at Mac computers with A…
MOSS-SoundEffect: audio generator — what it does and what you need to run itMOSS-SoundEffect is an open-source audio generator that turns a text description into audio. The model's task is text-to-audio: the user describes the…
MMAudio: audio generator — what it does and what you need to run itMMAudio is an open source audio generator that synthesizes an audio track from video or a text description. The developers position it as a tool for h…
Fish Speech: audio generator — what it does and what you need to run itFish Speech is an open source speech generator from Fish Audio that handles text-to-speech: it turns text into voiceover. The project publishes both t…
F5-TTS: audio generator — what it does and what you need to run itF5-TTS is an open source speech generator that turns text into sound. The project was created by developer SWivid and is distributed as code and ready…
DiffRhythm: audio generator — what it does and what you need to run itDiffRhythm is an audio generator that, as described by the authors, is designed to create complete songs end-to-end, from text to finished audio, usin…
CosyVoice: audio generator — what it does and what you need to run itCosyVoice is an open source speech generator from Alibaba. It tackles the text-to-speech task: converting text into spoken audio. The authors describe…
ChatTTS: audio generator — what it does and what you need to run itChatTTS is an open source audio generator from the 2noise team. It tackles the text-to-audio task: converting text into speech geared toward everyday…
Bark: audio generator — what it does and what you need to run itBark is an open source audio generator from Suno that tackles the text-to-speech task: it converts text into speech and other sounds based on a text d…
ACE-Step: audio generator — what it does and what you need to run itACE-Step is an open source audio generator that turns a text description into audio. The developers position it as a step toward a base model for musi…