Iris-3B

A 3-billion-parameter text-to-image model that paints pixels directly: no VAE, no latent space.

Generative priors are a promising foundation for downstream vision tasks. In this project we explore pixel-space generative models as an alternative to vision foundation models such as DINOv2.

Most image generators work in a compressed "latent" space and rely on a separate decoder to turn that into an image. Iris-3B skips that step: the network itself outputs every pixel. You type a prompt, and you get a 1024-pixel image.

Project page ยท Code

Examples

Generated at native aspect ratios of about one megapixel with the default settings (CFG 3, 100 steps).


Studio portrait of a Maasai elder wearing vibrant beaded jewelry, deep red cloth, dark backdrop, Rembrandt lighting, ultra detailed skin texture

Close-up portrait of a young woman with silver glitter freckles and iridescent makeup, soft pastel background, high fashion beauty photography

A glass sculpture of a heart filled with flowers, caustics and reflections, 3D render

Black and white portrait of a fisherman with a thick grey beard and deep wrinkles, piercing eyes, overcast light, fine grain film photograph

A Byzantine-style mosaic of a peacock made of tiny gold and turquoise tiles, shimmering texture

A bronze sculpture of a horse in motion, patina, dramatic museum spotlight

The ancient city of Petra with the Treasury carved into pink sandstone, morning light

A tree with lightbulbs instead of fruits glowing at dusk, surreal concept art

Thousands of sky lanterns rising into the night sky at Yi Peng festival in Chiang Mai

The aurora borealis swirling green and violet over a snowy Lofoten fishing village with red wooden cabins, reflections in a calm fjord, night photograph

Volcanic eruption at night in Iceland, rivers of glowing lava flowing across black fields, plumes of steam lit orange, long exposure

Get started

1. Install the code

git clone https://github.com/speridlabs/iris-3b.git
cd iris-3b
pip install -e .          # Python 3.11+, PyTorch 2.7.1+

2. Download the weights (about 12 GB)

hf download speridlabs/iris-3b --local-dir iris-3b

3. Generate an image

python scripts/sample.py --checkpoint iris-3b \
    --prompt "a red fox sleeping in fresh snow, golden hour"

The text encoder (Qwen3-VL-4B-Instruct) downloads automatically on the first run. You need an NVIDIA GPU with CUDA.

Tips

Setting Default What it does
--cfg-scale 3 How strictly the image follows the prompt. Higher values follow the prompt more literally but can look harsher.
--steps 100 Number of denoising steps. Fewer steps are faster but lose some detail.
--seed โ€” Fix it to get the same image again.
--negative-prompt โ€” Things you don't want in the image.
--txt-file โ€” A text file with one prompt per line, for batches.

Default output is 1024ร—1024. Write prompts as plain descriptive English sentences.

What's in this repo

File Contents
model.safetensors Model weights (EMA, float32; run with bfloat16 autocast)
config.yaml Architecture and sampling settings read by scripts/sample.py

About the model

Size 3B parameters (plus a frozen 4B text encoder)
Architecture Diffusion transformer: 8 dual-stream + 16 single-stream blocks, then a 4-block pixel head that turns each 16ร—16 patch back into pixels
Text encoder Qwen3-VL-4B-Instruct (frozen)
Training Rectified flow, trained from scratch at 256 โ†’ 512 โ†’ 1024 px, then supervised fine-tuning at 1024 px (665K steps total)

The full training recipe is in the code repository.

Limitations

  • Like every image generator, Iris-3B can produce inaccurate, biased or unsafe content. Review outputs before using them.
  • Text rendering inside images and exact object counts are not always reliable.
  • Images are generated at about one megapixel; for larger prints, use an upscaler.

License

Released under the Apache License 2.0. The text encoder, Qwen3-VL-4B-Instruct, is downloaded from its publisher and stays under its own license (Apache 2.0).

Citation

@techreport{licai2026iris,
  title       = {Iris-3B: Going Beyond the Latent with Pixel-Space
                 Diffusion Training, Conversion and Fine-Tuning},
  author      = {Li Cai, Hanqiu and Garabito, Chema},
  institution = {Speridlabs},
  year        = {2026}
}

Made with โค๏ธ by speridlabs.com

Downloads last month
16
Safetensors
Model size
3B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using speridlabs/iris-3b 1