ManifoldSplat edits 3D head shape with text in 90 seconds
Researchers presented ManifoldSplat — a method for language-based shape editing of animatable 3D heads reconstructed from monocular video. Instead of directly optimizing the Gaussian point cloud, edits are performed inside the structured FLAME manifold, which preserves identity and animation. The authors claim reconstruction and editing of an avatar in about 90 seconds on a consumer GPU and rendering at around 800 FPS. The code had not been released at the time of publication.
What it means
ManifoldSplat addresses a task that isn't covered in our AI for 3D section: it's not generation of a model from an image, but semantic editing of an existing avatar via text commands. The closest related tools — Hunyuan3D 2.1, Stable Fast 3D and OpenLRM — work as image-to-3d generators: they create a static model from scratch, but can't change its shape based on a text description and aren't tied to an animation rig.
According to the fact sheets for our tools: Hunyuan3D 2.1 requires 10 to 29 GB of VRAM depending on the stage, Stable Fast 3D — about 6 GB, OpenLRM doesn't state exact requirements but mentions CUDA OOM. The authors position ManifoldSplat as running on a consumer GPU, but don't specify the exact amount of VRAM. The authors also didn't state the license or weights, and the code is not yet available.
As long as there's neither a repository nor weights, the claimed 90 seconds and 800 FPS can't be verified on your own hardware. If the code appears, it will be the first tool in our field for text-based shape editing of animatable heads, rather than for generating static models.How the method works. The diagram was drawn based on this news note.