CPU3DAI tools for 3D, video and audio

AI tools with open source on CUDA under 12 GB of VRAM

VRAM thresholds, backends and licenses come from the tool fact sheets: every value is collected by code from an official source and stamped with the date it was checked. Requirements marked with an asterisk * are the authors’ own description, not our measurement. Details are on each tool’s page.

Showing 9 of 96

ToolTypeRunsVRAMBackendsFormats
MeshAnything3Dopen sourcefrom 7 GB *CUDA—
Stable Fast 3Dimage-to-3d3Dopen sourcefrom 6 GB *CUDAApple MPSglb
threestudio3Dopen sourcefrom 6 GB *CUDAobj
TripoSRimage-to-3d3Dopen sourcefrom 6 GB *CUDA—
Barktext-to-speechAudioopen sourcefrom 12 GB *CUDA—
MMAudioAudioopen sourcefrom 6 GB *CUDA—
Allegrotext-to-videoVideoopen sourcefrom 9.3 GB (9.3–27.5 GB) *CUDA—
FramePackVideoopen sourcefrom 6 GB *CUDA—
Maestroimage-text-to-videoVideoopen sourcefrom 12 GB (12–24 GB) *CUDA—

Requirements not stated by the authors

The official sources for these tools have no data for the selected filter — that means “unknown”, not “won’t work”.

3DTopia3Dopen sourcenot statedCUDAglb
CRM3Dopen sourcenot stated—obj
DreamGaussian3Dopen sourcenot stated——
Era3D3Dopen sourcenot stated——
Hunyuan3D 2.1image-to-3d3Dopen sourcefrom 10 GB (10–29 GB) *——
InstantMeshimage-to-3d3Dopen sourcenot statedCUDAobj
LGMtext-to-3d3Dopen sourcenot stated——
LichtFeld-Studio3Dopen sourcenot statedCUDAply
Michelangelo3Dopen sourcenot stated——
MVDream3Dopen sourcenot stated——
Nerfstudio3Dopen sourcenot statedCUDA—
One-2-3-453Dopen sourcenot stated——
OpenLRMimage-to-3d3Dopen sourcenot statedCUDAobj
PhysX-Omniimage-to-3d3Dopen sourcenot stated——
Point-E3Dopen sourcenot stated——
Scal3Rimage-to-3d3Dopen sourcenot statedCUDA—
Shap-E3Dopen sourcenot stated——
TRELLISimage-to-3d3Dopen sourcenot statedCUDAglb
Unique3D3Dopen sourcenot statedCUDA—
video_to_world3Dopen sourcenot statedCUDAply
Wonder3D3Dopen sourcenot stated——
Zero123++3Dopen sourcefrom 5 GB *——
ACE-Steptext-to-audioAudioopen sourcefrom 8 GB *——
ChatTTStext-to-audioAudioopen sourcefrom 4 GB *——
CosyVoicetext-to-speechAudioopen sourcenot stated——
DiffRhythmAudioopen sourcefrom 8 GB *——
Fish Speechtext-to-speechAudioopen sourcenot statedCUDA—
MOSS-SoundEffecttext-to-audioAudioopen sourcenot stated——
MusicGentext-to-audioAudioopen sourcenot stated——
OpenVoicetext-to-speechAudioopen sourcenot stated——
Stable Audio Opentext-to-audioAudioopen sourcenot stated——
XTTStext-to-speechAudioopen sourcenot statedCUDA—
YuEtext-generationAudioopen sourcefrom 24 GB *——
AI Comic BuilderVideoopen sourcenot stated——
AlayaRendererimage-to-imageVideoopen sourcenot stated——
AnimateDiffVideoopen sourcefrom 13 GB *——
CogVideoXtext-to-videoVideoopen sourcenot statedCUDA—
Cosmos3-Super-Image2Videoimage-to-videoVideoopen sourcenot statedCUDA—
forge-filmVideoopen sourcenot statedCUDA—
GAE-GeometricAutoEncoderimage-to-videoVideoopen sourcenot stated—ply
Infinite WorldVideoopen sourcenot stated——
LatteVideoopen sourcenot statedCUDA—
LingBot-VideoVideoopen sourcenot stated——
Mochi 1text-to-videoVideoopen sourcefrom 80 GB *——
nano-world-modelVideoopen sourcenot statedCUDA—
object-permanencevideo-to-videoVideoopen sourcenot stated——
Open-SoraVideoopen sourcenot stated——
Open-Sora-PlanVideoopen sourcefrom 24 GB *——
RefAlignimage-text-to-videoVideoopen sourcenot stated——
SparkDiffusiontext-to-videoVideoopen sourcenot statedCUDA—
Stable Video Diffusionimage-to-videoVideoopen sourcenot stated——
TE-Speed-MiniMaxH3-OSSVideoopen sourcenot stated——
video-generator-clientVideoopen sourcenot stated——
VideoCrafterVideoopen sourcenot stated——
ViMaxVideoopen sourcenot stated——
Wantext-to-videoVideoopen sourcefrom 8.19 GB *——
wind-comicVideoopen sourcenot stated——
WorldCrafterimage-to-videoVideoopen sourcenot statedCUDA—

Fact sheets are re-checked monthly; every value on a tool page has a source link and a check date. The asterisk * means “as described by the author”: a developer’s statement, not an independent measurement.