This is a modular diffusion pipeline built with 🧨 Diffusers' modular pipeline framework.

Pipeline Type: QwenImage21AutoBlocks

Description: Auto Modular pipeline for text-to-image and image-conditioned generation using Qwen-Image 2.1.

  • for text-to-image generation, all you need to provide is prompt
  • for image-conditioned generation, you need to provide prompt and image (one image or a list)

This pipeline uses a 4-block architecture that can be customized and extended.

Example Usage

[TODO]

Pipeline Architecture

This modular pipeline is composed of the following blocks:

  1. text_encoder (QwenImage21AutoTextEncoderStep)
    • Text encoder step that encodes the prompt, together with the condition images when there are any.
  2. vae_encoder (QwenImage21AutoVaeEncoderStep)
    • VAE encoder step that encodes the condition images into their latent representations.
  3. denoise (QwenImage21AutoCoreDenoiseStep)
    • Auto core denoise step that performs the denoising process.
  4. decode (QwenImage21DecodeStep)
    • Step that decodes the latents to RGBA images and postprocesses them

Model Components

  1. image_processor (VaeImageProcessor)
  2. text_encoder (Qwen3VLForConditionalGeneration)
  3. processor (Qwen3VLProcessor)
  4. guider (ClassifierFreeGuidance)
  5. vae (AutoencoderKLQwenImage21)
  6. scheduler (FlowMatchEulerDiscreteScheduler)
  7. transformer (QwenImage21Transformer2DModel)

Configuration Parameters

sample_sigmas (default: None): Default sampling grid of the checkpoint, used when sigmas is not passed.

Workflow Input Specification

text2image
  • prompt (str): The prompt or prompts to guide image generation.
image_conditioned
  • image (Image | list): Reference image(s) for denoising. Can be a single image or list of images.
  • prompt (str): The prompt or prompts to guide image generation.

Input/Output Specification

Inputs:

  • image (Image | list, optional): Reference image(s) for denoising. Can be a single image or list of images.
  • output_resolution (int, optional, defaults to 1024): Target side length used to derive the output size and to resize condition images.
  • prompt (str): The prompt or prompts to guide image generation.
  • negative_prompt (str, optional): The prompt or prompts not to guide the image generation.
  • generator (Generator, optional): Torch generator for deterministic generation.
  • num_images_per_prompt (int, optional, defaults to 1): The number of images to generate per prompt.
  • height (int, optional): The height in pixels of the generated image.
  • width (int, optional): The width in pixels of the generated image.
  • image_latents (list, optional): Normalized latents of each condition image. Can be generated from vae_encoder step.
  • latents (Tensor): Pre-generated noisy latents for image generation.
  • num_inference_steps (int): The number of denoising steps.
  • sigmas (list, optional): Custom sigmas for the denoising process.
  • use_kv_cache (bool, optional, defaults to True): Cache the text and condition-image keys and values after the first step. Valid because causal_condition modulates those tokens from t = 0, making their activations step-independent. Toggling it does not reproduce the same image bit-for-bit in reduced precision.
  • attention_kwargs (dict, optional): Additional kwargs for attention processors.
  • **denoiser_input_fields (None, optional): conditional model inputs for the denoiser: e.g. prompt_embeds, negative_prompt_embeds, etc.
  • output_type (str, optional, defaults to pil): Output format: 'pil', 'np', 'pt'.

Outputs:

  • images (list): Generated images.
Downloads last month
111
Safetensors
Model size
36.9k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support