CPU3DAI tools for 3D, video and audio

Iris-3B: New text-to-image model generates pixels directly, without a VAE

The model speridlabs/iris-3b has been published on HuggingFace — a text-to-image generator with 3 billion parameters. What sets it apart from most generators is that Iris-3B operates in pixel space: the network outputs every pixel of a 1024x1024 image directly, bypassing the VAE and latent space. The author also reports that after fine-tuning, the same model handles monocular depth estimation, image restoration and upscaling.

What it means

The model is distributed under the apache-2.0 license, with weights posted in safetensors format. The largest file — depth/model.safetensors — weighs 11.1 GB, and the repository totals 33.5 GB. It runs via pytorch. To generate a single image, the author recommends CFG 3 and 100 steps. The developer does not specify VRAM requirements, and there are no quantized builds in the repository. Full weights at 33.5 GB are a noticeable load for running on your own hardware, and without information on the minimum configuration it is hard to assess whether the model will run on a specific card. You can follow the arrival of new AI tools for 3D and video in the CPU3D news feed.
Iris-3B: text-to-image without VAE Direct pixel generation 1024x1024 Text prompt scene description Iris-3B 3 billion parameters CFG 3, 100 steps Pixels 1024x1024 directly Image no VAE, no latent space Fine-tuning depth, restoration, upscale 33.5 GB safetensors apache-2.0, pytorch VAE not used
How the method works. The diagram was drawn based on this news note.

See also