Instructions to use akshan-main/tiny-qwenimage21-modular-pipe with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use akshan-main/tiny-qwenimage21-modular-pipe with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("akshan-main/tiny-qwenimage21-modular-pipe", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
This is a modular diffusion pipeline built with 🧨 Diffusers' modular pipeline framework.
Pipeline Type: QwenImage21AutoBlocks
Description: Auto Modular pipeline for text-to-image and image-conditioned generation using Qwen-Image 2.1.
- for text-to-image generation, all you need to provide is
prompt - for image-conditioned generation, you need to provide
promptandimage(one image or a list)
This pipeline uses a 4-block architecture that can be customized and extended.
Example Usage
[TODO]
Pipeline Architecture
This modular pipeline is composed of the following blocks:
- text_encoder (
QwenImage21AutoTextEncoderStep)- Text encoder step that encodes the prompt, together with the condition images when there are any.
- vae_encoder (
QwenImage21AutoVaeEncoderStep)- VAE encoder step that encodes the condition images into their latent representations.
- denoise (
QwenImage21AutoCoreDenoiseStep)- Auto core denoise step that performs the denoising process.
- decode (
QwenImage21DecodeStep)- Step that decodes the latents to RGBA images and postprocesses them
Model Components
- image_processor (
VaeImageProcessor) - text_encoder (
Qwen3VLForConditionalGeneration) - processor (
Qwen3VLProcessor) - guider (
ClassifierFreeGuidance) - vae (
AutoencoderKLQwenImage21) - scheduler (
FlowMatchEulerDiscreteScheduler) - transformer (
QwenImage21Transformer2DModel)
Configuration Parameters
sample_sigmas (default: None): Default sampling grid of the checkpoint, used when sigmas is not passed.
Workflow Input Specification
text2image
prompt(str): The prompt or prompts to guide image generation.
image_conditioned
image(Image | list): Reference image(s) for denoising. Can be a single image or list of images.prompt(str): The prompt or prompts to guide image generation.
Input/Output Specification
Inputs:
image(Image | list, optional): Reference image(s) for denoising. Can be a single image or list of images.output_resolution(int, optional, defaults to1024): Target side length used to derive the output size and to resize condition images.prompt(str): The prompt or prompts to guide image generation.negative_prompt(str, optional): The prompt or prompts not to guide the image generation.generator(Generator, optional): Torch generator for deterministic generation.num_images_per_prompt(int, optional, defaults to1): The number of images to generate per prompt.height(int, optional): The height in pixels of the generated image.width(int, optional): The width in pixels of the generated image.image_latents(list, optional): Normalized latents of each condition image. Can be generated from vae_encoder step.latents(Tensor): Pre-generated noisy latents for image generation.num_inference_steps(int): The number of denoising steps.sigmas(list, optional): Custom sigmas for the denoising process.use_kv_cache(bool, optional, defaults toTrue): Cache the text and condition-image keys and values after the first step. Valid becausecausal_conditionmodulates those tokens fromt = 0, making their activations step-independent. Toggling it does not reproduce the same image bit-for-bit in reduced precision.attention_kwargs(dict, optional): Additional kwargs for attention processors.**denoiser_input_fields(None, optional): conditional model inputs for the denoiser: e.g. prompt_embeds, negative_prompt_embeds, etc.output_type(str, optional, defaults topil): Output format: 'pil', 'np', 'pt'.
Outputs:
images(list): Generated images.
- Downloads last month
- 111