CPU3DAI tools for 3D, video and audio

Audio generators on CUDA

VRAM thresholds, backends and licenses come from the tool fact sheets: every value is collected by code from an official source and stamped with the date it was checked. Requirements marked with an asterisk * are the authors’ own description, not our measurement. Details are on each tool’s page.

Showing 4 of 68

ToolTypeRunsVRAMBackendsFormats
Barktext-to-speechAudioopen sourcefrom 12 GB *CUDA
Fish Speechtext-to-speechAudioopen sourcenot statedCUDA
MMAudioAudioopen sourcefrom 6 GB *CUDA
XTTStext-to-speechAudioopen sourcenot statedCUDA

Requirements not stated by the authors

The official sources for these tools have no data for the selected filter — that means “unknown”, not “won’t work”.

ACE-Steptext-to-audioAudioopen sourcefrom 8 GB *
ChatTTStext-to-audioAudioopen sourcefrom 4 GB *
CosyVoicetext-to-speechAudioopen sourcenot stated
DiffRhythmAudioopen sourcenot stated
MOSS-SoundEffecttext-to-audioAudioopen sourcenot stated
MusicGentext-to-audioAudioopen sourcenot stated
OpenVoicetext-to-speechAudioopen sourcenot stated
Stable Audio Opentext-to-audioAudioopen sourcenot stated
YuEtext-generationAudioopen sourcenot stated

Fact sheets are re-checked monthly; every value on a tool page has a source link and a check date. The asterisk * means “as described by the author”: a developer’s statement, not an independent measurement.