Text-to-Image
lora
simpletuner
training-assistant
qwen-image
experimental

SimpleTuner Training Assistant v2 for Qwen Image 2.1

Built with Qwen. This training assistant LoRA was trained from scratch for 1,000 updates, using corrected adamw_bf16, learning rate 1e-4 and global gradient norm clipping at 1.0. This standalone repository contains the same assistant-v2 step-1000 weights used in the Qwen Image 2.1 LoRA experiments.

Keep this adapter frozen and active during downstream LoRA training, then disable it when generating images with the downstream adapter. Its purpose is to help preserve coherence and image quality during fine-tuning. Unlike v1, v2 mixes model-generated images with real CC12M and e621 images during assistant training.

Use in SimpleTuner

Use a SimpleTuner checkout with Qwen Image 2.1 and assistant-LoRA support. Add these fields to a complete concept-training configuration:

{
  "model_family": "qwen_image",
  "model_flavour": "v2.1",
  "model_type": "lora",
  "lora_type": "standard",
  "assistant_lora_path": "SimpleTuner/Qwen-Image-2.1-training-assistant-v2",
  "disable_assistant_lora": false,
  "assistant_lora_strength": 1.0,
  "assistant_lora_inference_strength": 0.0
}

SimpleTuner freezes the assistant alongside the trainable concept adapter. Use ordinary concept training; distillation_method: assistant_lora creates an assistant using online teacher generation. At inference, load the base model and your resulting concept LoRA with this assistant disabled. This assistant has no trigger word. It is not a character adapter or a few-step generation adapter. Older Qwen Image flavours have not been tested with these weights.

The assistant is intentionally trained to absorb changes that would otherwise affect downstream tuning. Images made with the assistant enabled by itself do not represent the intended final inference setup. The relevant evaluation is downstream training with the assistant enabled and frozen, followed by inference with it disabled.

Training

Training mixed three equally weighted source groups:

The actual 1,000 batches were 338 synthetic, 331 CC12M and 331 e621. Synthetic backends were capped at 128 samples per aspect bucket; each real-data resolution backend was capped at 256 samples. These were bounded subsets, not passes over the complete source datasets, and images can be reused across resolution backends. No Domokun dataset was configured. Captions were checked for Domokun-name/trigger matches; this does not establish that no unlabelled character appears in the real images.

Setting Value
GPU NVIDIA H100
Updates / batch / accumulation 1,000 / 1 / 1
Precision BF16
LoRA rank / alpha 32 / 32
Attention projections to_q, to_k, to_v, to_out.0
Optimizer Corrected adamw_bf16
Learning rate 1e-4, constant after 25 warmup steps
Gradient clipping Global norm, maximum 1.0
Gradient checkpointing Enabled, interval 2
Sampling Equal probability per source group, distributed across its resolution backends
VAE Original Qwen Image 2.1 VAE; full-frame encoding, batch 1
Trainable parameters 33,554,432

Training covered base resolutions 512, 1024, 1536 and 2048. Synthetic data included square, portrait and landscape buckets; real-image backends used aspect-preserving area buckets. The nominal resolution in the weights metadata is not the complete list of training sizes. See training_details.json, training_config.json and dataloader.json for the recipe and provenance.

Evaluation and limitations

The experiment collection compares baseline training, the assistants themselves, and downstream Domokun adapters. The focus is preservation of coherence and quality while still learning the new concept. This adapter is not intended to eliminate every instance of concept leakage or dataset bias.

The extended photo-aesthetics experiment includes 50,000 downstream updates at 512px and 10,000 at 1024px, with the frozen v2 assistant active during training and disabled for validation. Photo-aesthetics was not a configured assistant-training source; no claim of zero image overlap with CC12M or e621 is made. Together with the earlier controls, these comparisons show preservation of coherence and quality in the tested training setup. They do not establish that every learning rate, dataset or training duration will behave the same way.

More recent assisted, regularised Domokun tests completed 2,000 single-resolution updates at 512px and 1024px. A 4,000-update probabilistic multi-scale run produced recognizable Domokun in inspected beach samples at both output resolutions, while an inspected held-out portrait remained coherent. These are limited qualitative observations. REPA and flow-shift comparisons are additional downstream experiments; neither was used to train this assistant.

The published comparison samples use the original Qwen VAE. Newer SimpleTuner versions can use Ollin's texture-fixed decoder, which affects decoded texture independently of the assistant.

Artifact and license

pytorch_lora_weights.safetensors contains 256 finite BF16 tensors and matches the completed v2 checkpoint at update 1,000. It is byte-identical to the assistant-v2 checkpoint in the experiment repository.

SHA-256: 374b0d936a42f75f4af64d219f1a6b022dc3d6584f7e7e09b1d94a999800b7d2.

Modification notice: this separately trained LoRA changes Qwen Image 2.1 transformer attention projections when loaded. Distribution and use are subject to the Qwen Research License Agreement. See Notice for upstream attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SimpleTuner/Qwen-Image-2.1-training-assistant-v2

Adapter
(55)
this model

Datasets used to train SimpleTuner/Qwen-Image-2.1-training-assistant-v2