CPU3DAI tools for 3D, video and audio

AI tools on CUDA

VRAM thresholds, backends and licenses come from the tool fact sheets: every value is collected by code from an official source and stamped with the date it was checked. Requirements marked with an asterisk * are the authors’ own description, not our measurement. Details are on each tool’s page.

Showing 35 of 96

ToolTypeRunsVRAMBackendsFormats
3DTopia3Dopen sourcenot statedCUDAglb
Gaussian Splatting (3DGS)3Dopen sourcefrom 24 GB *CUDA—
habitat-gs3Dopen sourcefrom 80 GB *CUDAply
InstantMeshimage-to-3d3Dopen sourcenot statedCUDAobj
LichtFeld-Studio3Dopen sourcenot statedCUDAply
MeshAnything3Dopen sourcefrom 7 GB *CUDA—
Nerfstudio3Dopen sourcenot statedCUDA—
OpenLRMimage-to-3d3Dopen sourcenot statedCUDAobj
Scal3Rimage-to-3d3Dopen sourcenot statedCUDA—
Stable Fast 3Dimage-to-3d3Dopen sourcefrom 6 GB *CUDAApple MPSglb
threestudio3Dopen sourcefrom 6 GB *CUDAobj
TRELLISimage-to-3d3Dopen sourcenot statedCUDAglb
TripoSRimage-to-3d3Dopen sourcefrom 6 GB *CUDA—
Unique3D3Dopen sourcenot statedCUDA—
video_to_world3Dopen sourcenot statedCUDAply
Barktext-to-speechAudioopen sourcefrom 12 GB *CUDA—
Fish Speechtext-to-speechAudioopen sourcenot statedCUDA—
MMAudioAudioopen sourcefrom 6 GB *CUDA—
XTTStext-to-speechAudioopen sourcenot statedCUDA—
4DAnyoneVideoopen sourcefrom 24 GB *CUDA—
Allegrotext-to-videoVideoopen sourcefrom 9.3 GB (9.3–27.5 GB) *CUDA—
CogVideoXtext-to-videoVideoopen sourcenot statedCUDA—
Cosmos3-Super-Image2Videoimage-to-videoVideoopen sourcenot statedCUDA—
forge-filmVideoopen sourcenot statedCUDA—
FramePackVideoopen sourcefrom 6 GB *CUDA—
HelixWorldVideoopen sourcefrom 80 GB *CUDA—
HunyuanVideotext-to-videoVideoopen sourcefrom 45 GB (45–60 GB) *CUDA—
Lanceany-to-anyVideoopen sourcefrom 40 GB *CUDAobj
LatteVideoopen sourcenot statedCUDA—
Maestroimage-text-to-videoVideoopen sourcefrom 12 GB (12–24 GB) *CUDA—
nano-world-modelVideoopen sourcenot statedCUDA—
SkyReelsVideoopen sourcefrom 18.5 GB *CUDA—
SparkDiffusiontext-to-videoVideoopen sourcenot statedCUDA—
Wan-Dancer-14Bimage-to-videoVideoopen sourcefrom 80 GB *CUDA—
WorldCrafterimage-to-videoVideoopen sourcenot statedCUDA—

Cloud services — no VRAM needed

They compute on their own servers, so a hardware filter does not apply to them.

KaedimTurning sketches, references, product photos and briefs into 3D assets3Dcloudcloudnot needed—
Luma Genievideo, image, audio and text generation3Dcloudcloudnot neededJPG
Masterpiece XAI editor for creating 3D scenes and assets3Dcloudcloudnot needed—
Meshy3D model generation from text and images3Dcloudcloudnot neededFBX, OBJ, STL, GLTF, GLB, USDZ, …
Polycam3D scanning, photogrammetry, 3D model and Gaussian splat generation3Dcloudcloudnot neededGLTF, OBJ, FBX, DAE, STL, USDZ, …
Rodin3D model generation from text or images3Dcloudcloudnot neededSTL, FBX, OBJ, GLB, GLTF, USDZ
Tripo AItext-to-3D, image-to-3D3Dcloudcloudnot neededGLB, FBX, OBJ, USD, STL, 3MF
Dream MachineMultimodal creative agent for creating and editing images, video, audio and textVideocloudcloudnot neededMP4
HailuoVideo and image generationVideocloudcloudnot needed—
KlingImage and video generationVideocloudcloudnot needed—
PikaVideo, image and audio generationVideocloudcloudnot needed—
RunwayVideo, image, speech, sound effect, avatar generation and other tasksVideocloudcloudnot neededMP4, MOV, ZIP, PNG, WAV, EXR
SoraVideo generation from text, images or videoVideocloudcloudnot neededMP4
VeoVideo generationVideocloudcloudnot needed—
ViduVideo generation: Reference2Video, Image2Video, Start-End2Video, Text2Video, Upscale, TemplateVideocloudcloudnot neededpng, jpeg, jpg, webp

Requirements not stated by the authors

The official sources for these tools have no data for the selected filter — that means “unknown”, not “won’t work”.

CRM3Dopen sourcenot stated—obj
DreamGaussian3Dopen sourcenot stated——
Era3D3Dopen sourcenot stated——
Hunyuan3D 2.1image-to-3d3Dopen sourcefrom 10 GB (10–29 GB) *——
LGMtext-to-3d3Dopen sourcenot stated——
Michelangelo3Dopen sourcenot stated——
MVDream3Dopen sourcenot stated——
One-2-3-453Dopen sourcenot stated——
PhysX-Omniimage-to-3d3Dopen sourcenot stated——
Point-E3Dopen sourcenot stated——
Shap-E3Dopen sourcenot stated——
Wonder3D3Dopen sourcenot stated——
Zero123++3Dopen sourcefrom 5 GB *——
ACE-Steptext-to-audioAudioopen sourcefrom 8 GB *——
ChatTTStext-to-audioAudioopen sourcefrom 4 GB *——
CosyVoicetext-to-speechAudioopen sourcenot stated——
DiffRhythmAudioopen sourcefrom 8 GB *——
MOSS-SoundEffecttext-to-audioAudioopen sourcenot stated——
MusicGentext-to-audioAudioopen sourcenot stated——
OpenVoicetext-to-speechAudioopen sourcenot stated——
Stable Audio Opentext-to-audioAudioopen sourcenot stated——
YuEtext-generationAudioopen sourcefrom 24 GB *——
AI Comic BuilderVideoopen sourcenot stated——
AlayaRendererimage-to-imageVideoopen sourcenot stated——
AnimateDiffVideoopen sourcefrom 13 GB *——
GAE-GeometricAutoEncoderimage-to-videoVideoopen sourcenot stated—ply
Infinite WorldVideoopen sourcenot stated——
LingBot-VideoVideoopen sourcenot stated——
Mochi 1text-to-videoVideoopen sourcefrom 80 GB *——
object-permanencevideo-to-videoVideoopen sourcenot stated——
Open-SoraVideoopen sourcenot stated——
Open-Sora-PlanVideoopen sourcefrom 24 GB *——
RefAlignimage-text-to-videoVideoopen sourcenot stated——
Stable Video Diffusionimage-to-videoVideoopen sourcenot stated——
TE-Speed-MiniMaxH3-OSSVideoopen sourcenot stated——
video-generator-clientVideoopen sourcenot stated——
VideoCrafterVideoopen sourcenot stated——
ViMaxVideoopen sourcenot stated——
Wantext-to-videoVideoopen sourcefrom 8.19 GB *——
wind-comicVideoopen sourcenot stated——

Fact sheets are re-checked monthly; every value on a tool page has a source link and a check date. The asterisk * means “as described by the author”: a developer’s statement, not an independent measurement.