Iris-3B: New text-to-image model generates pixels directly, without a VAE
The model speridlabs/iris-3b has been published on HuggingFace — a text-to-image generator with 3 billion parameters. What sets it apart from most generators is that Iris-3B operates in pixel space: the network outputs every pixel of a 1024x1024 image directly, bypassing the VAE and latent space. The author also reports that after fine-tuning, the same model handles monocular depth estimation, image restoration and upscaling.
What it means
The model is distributed under the apache-2.0 license, with weights posted in safetensors format. The largest file — depth/model.safetensors — weighs 11.1 GB, and the repository totals 33.5 GB. It runs via pytorch. To generate a single image, the author recommends CFG 3 and 100 steps.
The developer does not specify VRAM requirements, and there are no quantized builds in the repository. Full weights at 33.5 GB are a noticeable load for running on your own hardware, and without information on the minimum configuration it is hard to assess whether the model will run on a specific card.
You can follow the arrival of new AI tools for 3D and video in the CPU3D news feed.How the method works. The diagram was drawn based on this news note.