CPU3DAI tools for 3D, video and audio

AI tools under 16 GB of VRAM

VRAM thresholds, backends and licenses come from the tool fact sheets: every value is collected by code from an official source and stamped with the date it was checked. Requirements marked with an asterisk * are the authors’ own description, not our measurement. Details are on each tool’s page.

Showing 19 of 96

ToolTypeRunsVRAMBackendsFormats
Hunyuan3D 2.1image-to-3d3Dopen sourcefrom 10 GB (10–29 GB) *——
MeshAnything3Dopen sourcefrom 7 GB *CUDA—
Stable Fast 3Dimage-to-3d3Dopen sourcefrom 6 GB *CUDAApple MPSglb
threestudio3Dopen sourcefrom 6 GB *CUDAobj
TripoSRimage-to-3d3Dopen sourcefrom 6 GB *CUDA—
Zero123++3Dopen sourcefrom 5 GB *——
ACE-Steptext-to-audioAudioopen sourcefrom 8 GB *——
Barktext-to-speechAudioopen sourcefrom 12 GB *CUDA—
ChatTTStext-to-audioAudioopen sourcefrom 4 GB *——
DiffRhythmAudioopen sourcefrom 8 GB *——
MMAudioAudioopen sourcefrom 6 GB *CUDA—
MOSS-SoundEffect without CUDA (Mac and AMD)Audioopen sourcefrom 8 GB (8–10 GB) *Apple MPSAMD ROCm—
Allegrotext-to-videoVideoopen sourcefrom 9.3 GB (9.3–27.5 GB) *CUDA—
AnimateDiffVideoopen sourcefrom 13 GB *——
FramePackVideoopen sourcefrom 6 GB *CUDA—
LTX-Videoimage-to-videoVideoopen sourcefrom 1 GB *Apple MPS—
Maestroimage-text-to-videoVideoopen sourcefrom 12 GB (12–24 GB) *CUDA—
Pyramid Flowtext-to-videoVideoopen sourcefrom 8 GB *Apple MPS—
Wantext-to-videoVideoopen sourcefrom 8.19 GB *——

Cloud services — no VRAM needed

They compute on their own servers, so a hardware filter does not apply to them.

KaedimTurning sketches, references, product photos and briefs into 3D assets3Dcloudcloudnot needed—
Luma Genievideo, image, audio and text generation3Dcloudcloudnot neededJPG
Masterpiece XAI editor for creating 3D scenes and assets3Dcloudcloudnot needed—
Meshy3D model generation from text and images3Dcloudcloudnot neededFBX, OBJ, STL, GLTF, GLB, USDZ, …
Polycam3D scanning, photogrammetry, 3D model and Gaussian splat generation3Dcloudcloudnot neededGLTF, OBJ, FBX, DAE, STL, USDZ, …
Rodin3D model generation from text or images3Dcloudcloudnot neededSTL, FBX, OBJ, GLB, GLTF, USDZ
Tripo AItext-to-3D, image-to-3D3Dcloudcloudnot neededGLB, FBX, OBJ, USD, STL, 3MF
Dream MachineMultimodal creative agent for creating and editing images, video, audio and textVideocloudcloudnot neededMP4
HailuoVideo and image generationVideocloudcloudnot needed—
KlingImage and video generationVideocloudcloudnot needed—
PikaVideo, image and audio generationVideocloudcloudnot needed—
RunwayVideo, image, speech, sound effect, avatar generation and other tasksVideocloudcloudnot neededMP4, MOV, ZIP, PNG, WAV, EXR
SoraVideo generation from text, images or videoVideocloudcloudnot neededMP4
VeoVideo generationVideocloudcloudnot needed—
ViduVideo generation: Reference2Video, Image2Video, Start-End2Video, Text2Video, Upscale, TemplateVideocloudcloudnot neededpng, jpeg, jpg, webp

Requirements not stated by the authors

The official sources for these tools have no data for the selected filter — that means “unknown”, not “won’t work”.

3DTopia3Dopen sourcenot statedCUDAglb
CRM3Dopen sourcenot stated—obj
DreamGaussian3Dopen sourcenot stated——
Era3D3Dopen sourcenot stated——
Hunyuan3D 2.1 without CUDA (Mac and AMD)3Dopen sourcenot statedApple MPSAMD ROCmobj, glb, ply, stl, fbx, dae, 3mf
InstantMeshimage-to-3d3Dopen sourcenot statedCUDAobj
LGMtext-to-3d3Dopen sourcenot stated——
LichtFeld-Studio3Dopen sourcenot statedCUDAply
Michelangelo3Dopen sourcenot stated——
MVDream3Dopen sourcenot stated——
Nerfstudio3Dopen sourcenot statedCUDA—
One-2-3-453Dopen sourcenot stated——
OOOSplat3Dopen sourcenot statedApple MPSply
OpenLRMimage-to-3d3Dopen sourcenot statedCUDAobj
PhysX-Omniimage-to-3d3Dopen sourcenot stated——
Point-E3Dopen sourcenot stated——
Scal3Rimage-to-3d3Dopen sourcenot statedCUDA—
Shap-E3Dopen sourcenot stated——
TRELLISimage-to-3d3Dopen sourcenot statedCUDAglb
Unique3D3Dopen sourcenot statedCUDA—
video_to_world3Dopen sourcenot statedCUDAply
Wonder3D3Dopen sourcenot stated——
CosyVoicetext-to-speechAudioopen sourcenot stated——
F5-TTStext-to-speechAudioopen sourcenot statedAMD ROCm—
Fish Speechtext-to-speechAudioopen sourcenot statedCUDA—
MOSS-SoundEffecttext-to-audioAudioopen sourcenot stated——
MusicGentext-to-audioAudioopen sourcenot stated——
OpenVoicetext-to-speechAudioopen sourcenot stated——
Stable Audio Opentext-to-audioAudioopen sourcenot stated——
XTTStext-to-speechAudioopen sourcenot statedCUDA—
AI Comic BuilderVideoopen sourcenot stated——
AlayaRendererimage-to-imageVideoopen sourcenot stated——
CogVideoXtext-to-videoVideoopen sourcenot statedCUDA—
Cosmos3-Super-Image2Videoimage-to-videoVideoopen sourcenot statedCUDA—
forge-filmVideoopen sourcenot statedCUDA—
GAE-GeometricAutoEncoderimage-to-videoVideoopen sourcenot stated—ply
Infinite WorldVideoopen sourcenot stated——
LatteVideoopen sourcenot statedCUDA—
LingBot-VideoVideoopen sourcenot stated——
nano-world-modelVideoopen sourcenot statedCUDA—
object-permanencevideo-to-videoVideoopen sourcenot stated——
Open-SoraVideoopen sourcenot stated——
RefAlignimage-text-to-videoVideoopen sourcenot stated——
SparkDiffusiontext-to-videoVideoopen sourcenot statedCUDA—
Stable Video Diffusionimage-to-videoVideoopen sourcenot stated——
TE-Speed-MiniMaxH3-OSSVideoopen sourcenot stated——
video-generator-clientVideoopen sourcenot stated——
VideoCrafterVideoopen sourcenot stated——
ViMaxVideoopen sourcenot stated——
wind-comicVideoopen sourcenot stated——
WorldCrafterimage-to-videoVideoopen sourcenot statedCUDA—

Fact sheets are re-checked monthly; every value on a tool page has a source link and a check date. The asterisk * means “as described by the author”: a developer’s statement, not an independent measurement.