Audio generators under 24 GB of VRAM
VRAM thresholds, backends and licenses come from the tool fact sheets: every value is collected by code from an official source and stamped with the date it was checked. Requirements marked with an asterisk * are the authors’ own description, not our measurement. Details are on each tool’s page.
Showing 5 of 68
| Tool | Type | Runs | VRAM | Backends | Formats |
|---|---|---|---|---|---|
| ACE-Steptext-to-audio | Audio | open source | from 8 GB * | — | — |
| Barktext-to-speech | Audio | open source | from 12 GB * | CUDA | — |
| ChatTTStext-to-audio | Audio | open source | from 4 GB * | — | — |
| MMAudio | Audio | open source | from 6 GB * | CUDA | — |
| MOSS-SoundEffect without CUDA (Mac and AMD) | Audio | open source | from 8 GB (8–10 GB) * | Apple MPSAMD ROCm | — |
Requirements not stated by the authors
The official sources for these tools have no data for the selected filter — that means “unknown”, not “won’t work”.
| CosyVoicetext-to-speech | Audio | open source | not stated | — | — |
| DiffRhythm | Audio | open source | not stated | — | — |
| F5-TTStext-to-speech | Audio | open source | not stated | AMD ROCm | — |
| Fish Speechtext-to-speech | Audio | open source | not stated | CUDA | — |
| MOSS-SoundEffecttext-to-audio | Audio | open source | not stated | — | — |
| MusicGentext-to-audio | Audio | open source | not stated | — | — |
| OpenVoicetext-to-speech | Audio | open source | not stated | — | — |
| Stable Audio Opentext-to-audio | Audio | open source | not stated | — | — |
| XTTStext-to-speech | Audio | open source | not stated | CUDA | — |
| YuEtext-generation | Audio | open source | not stated | — | — |
Fact sheets are re-checked monthly; every value on a tool page has a source link and a check date. The asterisk * means “as described by the author”: a developer’s statement, not an independent measurement.