WikiQwen-Illustrator-7B

A style LoRA for Qwen-Image-2.1 that draws how-to step illustrations in a flat, clean house style from a one-sentence caption. It is the illustrator of the WikiQwen app, drawing a picture for every step of an article.

Sample with true classifier-free guidance. Qwen-Image-2.1 is meant to be sampled without guidance, but after this much fine-tuning the LoRA needs it: without guidance the style is weaker and lettering falls apart. The settings below (20 steps, true_cfg_scale=4, empty negative prompt) were chosen on held-out captions.

import torch
from diffusers import QwenImage21Pipeline   # diffusers main (the pipeline is not in a release yet)

pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16).to("cuda")
pipe.load_lora_weights("devon7y/WikiQwen-Illustrator-7B", weight_name="pytorch_lora_weights.safetensors")
image = pipe(prompt="A hand waters a potted basil plant on a sunny windowsill.", width=512, height=384,
             num_inference_steps=20, true_cfg_scale=4.0, negative_prompt="").images[0]
image.convert("RGB").save("step.png")   # the pipeline returns RGBA

Training

  • Data: 587,815 captioned Creative Commons step illustrations by wikiHow contributors (Kiwix March 2023 archive, CC BY-NC-SA 3.0), at 512x384; captions used as-is.
  • LoRA rank 128 / alpha 128 on every linear layer inside the 32 transformer blocks (attention q/k/v/out and the SwiGLU MLP), 335.5M parameters; AdamW, lr 5e-5 cosine with 500 warmup steps, effective batch 32, 18,000 steps (1 epoch), fp32 LoRA weights under bf16, one H100 (18 h).

Non-commercial research demo. Not affiliated with or endorsed by the source site.

Downloads last month
36
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for devon7y/WikiQwen-Illustrator-7B

Adapter
(91)
this model

Space using devon7y/WikiQwen-Illustrator-7B 1

Collection including devon7y/WikiQwen-Illustrator-7B