Announced 7 Oct 2026 · Sources checked
What did Sperid Labs release, and when?
The dated research object is the 7 October 2026 arXiv preprint, identifier 2610.09450. Authors listed are Hanqiu Li Cai, marked as project lead, and Chema Garabito, both of Sperid Labs. The abstract says they release Iris-3B’s weights and training code. We opened that abstract page on 9 October.
The inspectable software landing is later than the preprint stamp. The Hugging Face card speridlabs/iris-3b was live when we opened it on 9 October. It describes a 3-billion-parameter model that “paints every pixel directly — no VAE, no latent space,” plus fine-tunes for monocular depth and for image restoration and upscaling. The GitHub repository speridlabs/iris-3b carries an Apache 2.0 badge and points at the same paper, project page and Hub repo. The project page at speridlabs.com/research/iris is dated “Oct 2026” without a day.
This is a research weight drop, not a hosted image API and not a claim that pixel space now beats every latent generator. For the current commercial image-model path on this site, the nearer product story remains Google’s Nano Banana 2.1, which is a Gemini image endpoint rather than a downloadable pixel-space backbone.
How does a 3B model write pixels instead of latents?
Most popular text-to-image systems compress the picture with a variational autoencoder, run diffusion or flow matching in that smaller grid, then decode back to pixels. Iris-3B skips the codec. The paper and project page describe a flow-matching transformer on raw RGB: the image is cut into 16×16 pixel patches, so a 1024² frame becomes a 64×64 token grid.
The trunk follows a FLUX-style hybrid: 8 dual-stream blocks with separate text and image weights, then 16 single-stream blocks, at width 2560. A four-block pixel-transformer (PiT) head, taken from PixelDiT, turns each patch token back into pixels. Text comes from a frozen Qwen3-VL-4B-Instruct encoder. 2D axial RoPE is used so one set of weights can run at different resolutions and aspect ratios up to about one megapixel.
The card says the model was trained from scratch on a 256 → 512 → 1024 curriculum, then supervised-fine-tuned at 1024, for 665,000 steps in total. That is a training recipe, not a user setting. Readers who want the usual encoder-decoder picture of vision-language models still have our multimodal AI explainer; Iris is a generator that happens to use a frozen multimodal encoder for text, not a general image-understanding chat model.
Did skipping the VAE help depth and restoration?
The paper’s reason for building a pixel-space prior is transfer, not prettier samples. The authors expected a generator that never discards detail to a codec to fine-tune better on monocular depth and on image restoration. They test that idea two ways: pretrain Iris-3B from scratch, and convert a latent FLUX.2 Klein base 4B model to pixel space.
The result, in the authors’ words, is “no significant improvement from using a pixel-space generative prior.” Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind. On 4× DIV2K restoration, neither pixel model beats a latent FLUX.2 Klein fine-tune; the converted one trails it slightly. The abstract says they document recipes, failure modes and remaining confounds rather than claim the question is closed.
The Hub card still ships those fine-tunes. It reports depth AbsRel 0.071 and δ1 0.946 as a mean over NYUv2, KITTI, ETH3D, ScanNet and DIODE under a Marigold V2 protocol, and restoration LPIPS 0.292 / PSNR 20.68 on DIV2K validation at 4×. Depth is affine-invariant, not metric. Those figures are Sperid’s. They describe the released heads; they do not cancel the paper’s comparison with the latent twin.
How do you run the weights?
The card’s getting-started path is `git clone https://github.com/speridlabs/iris-3b.git`, `pip install -e .` on Python 3.11+ with PyTorch 2.7.1+, then `hf download speridlabs/iris-3b` and `python scripts/sample.py`. Default sampling is CFG 3 and 100 steps at 1024×1024. The text encoder downloads on first run. The card says you need an NVIDIA GPU with CUDA. We did not install that stack.
Text-to-image weights are about 12 GB. The `depth/` and `upscaler/` folders add about 12 GB each. Depth and restore run a single forward pass with an empty prompt, so the text encoder is not loaded for those scripts. The card lists `scripts/depth.py` and `scripts/upscale.py`. License on the Iris weights is Apache 2.0; Qwen3-VL-4B-Instruct stays under its own Apache-2.0 terms.
What is not established?
We did not sample Iris-3B, rerun GenEval or OneIG, or fine-tune the depth head. We did not verify the GitHub star count or the Hub download counter as a quality signal. A pixel-space prior is also not the same object as a contrastive image encoder; our CLIP explainer is about embedding images and text in one space, not about writing pixels.
The project page’s “fraction of their training compute” claim is not a public FLOP table on the pages we opened. The negative downstream result is the authors’ comparison under one matched recipe; they say confounds remain. The 7 October preprint is the date for the paper. The 9 October Hub and GitHub pages are the date we confirmed the files were public.
This is an evidence review of pages opened on 9 October. It is not a first-hand generation test and not a ranking of Iris against Qwen-Image, FLUX or Nano Banana on a shared prompt set.
Common questions
Is Iris-3B open source?
The Hub card and GitHub README mark the release Apache 2.0, and the paper says weights and training code are released. The frozen Qwen3-VL-4B text encoder is a separate Apache-2.0 download from its publisher. We did not audit training data rights.
Did the authors claim pixel space is better for depth?
No. They expected an advantage and report that it did not show up under their matched fine-tunes. Iris-3B was level with latent FLUX.2 Klein on depth; the converted pixel model fell behind.
Can I call this as a hosted API?
Not on the pages we opened. The release is a GitHub repository and a Hugging Face repo, plus a Spaces demo named on the README. There is no price list or SLA on the project page.
What to remember
Use the 7 October preprint for the architecture and the negative transfer result. Use the 9 October Hub card for file sizes, sampling flags and the Apache-2.0 grant. Treat OneIG 0.540 and the depth/restore tables as Sperid’s until an independent run exists.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





