Image-to-Image
PEFT
Safetensors

Stable Layers

Decomposing Images into Editable RGBA Layers

Stable Layers Teaser

Paper    Project Page    Model (Hugging Face)    Code

About

Stable Layers decomposes a single RGB image into a small stack of back-to-front RGBA layers (background plus object layers) suitable for compositing and editing.

This repository hosts the LoRA adapter over Qwen/Qwen-Image-Layered. This model was trained using Qwen 3.5 9b as the VLM teacher.

Model Artifacts

This repository is the LoRA adapter (adapter_config.json, adapter_model.safetensors). The base model is pulled from Hugging Face automatically (Qwen/Qwen-Image-Layered). See the code repository for the inference script.

Installation

pip install torch diffusers transformers peft pillow numpy

Tested with torch 2.11, diffusers 0.37, transformers 5.5, peft 0.18. Requires one GPU — the base model is ~40 GB in bf16, so an 80 GB-class card (A100-80 / H100 / H200) is comfortable.

Recommended Inference Settings

Heun sampler · 50 steps · CFG 1.0 · 640 px · 4 layers

Setting Value Flag
Sampler Heun (2nd order) always used — not configurable
Steps 50 --steps 50
CFG / guidance 1.0 (off) --guidance-scale 1.0
Resolution 640 px (max dim) --size 640
Layers 4 --num-layers 4

Note: using a higher resolution, lower steps, or a non-Heun sampler will garble the results.

Quickstart

# single image
python decompose.py --input photo.png --output results/

# a directory of images
python decompose.py --input images/ --output results/

# RGBA layers with real alpha (for compositing / editors)
python decompose.py --input images/ --output results/ --transparent

# explicit LoRA location
python decompose.py --input photo.png --output results/ --lora ./model

Outputs

results/<image_name>/
  source.png       # input, resized
  composite.png    # layers recomposited — compare against source as a sanity check
  layer_0.png      # background (inpainted behind the removed objects)
  layer_1.png      # object layers, back-to-front
  layer_2.png
  layer_3.png

Layers are ordered back-to-front: layer_0 is the background, higher indices sit on top. Not every image needs all 4 — unused layers come out blank, which is normal.

By default layers are composited onto white (easy to eyeball). Pass --transparent to get RGBA with real alpha, which is what you want when importing into an editor or compositing them yourself.

Notes:

  • Reproducible: noise is seeded per image as seed + image_index (--seed 42 by default), so the same inputs give the same outputs.
  • Prompt: decomposition is driven by the source image; the text prompt only nudges guidance. The default ("a clean, well composed image") is fine — override with --prompt if you want.
  • Aspect ratio is preserved; the longest side is scaled to --size and both dimensions are rounded to multiples of 16.

License

This code and model usage are subject to Stability AI Community License terms.

For individuals or organizations generating annual revenue of USD 1,000,000 (or local currency equivalent) or more, commercial usage requires an enterprise license from Stability AI.

Citation

@article{stablelayers2026,
  author  = {},
  title   = {Stable Layers: Decomposing Images into Editable RGBA Layers},
  journal = {arXiv preprint arXiv:2605.30257},
  year    = {2026}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for StabilityLabs/Stable-Layers

Base model

Qwen/Qwen-Image
Adapter
(3)
this model

Paper for StabilityLabs/Stable-Layers