Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation
Model weights for PVD class-conditioned image generation (C2I) and text-to-image generation with SD3.5 Medium, FLUX.1-dev and Qwen-Image.
Model Details
| Variant | Weight directory | Released components | Configuration |
|---|---|---|---|
| C2I | weights/pvd_c2i/ |
model.safetensors; VAE and latent statistics in cache/ |
ImageNet classes, 256×256 |
| SD3.5 Medium | weights/pvd_sd35m/ |
part1.safetensors, part2.safetensors, model_config.json |
12 layers per transformer |
| SD3.5 Medium (Unsplash) | weights/pvd_sd35m_unsplash/ |
part1.safetensors, part2.safetensors, model_config.json |
Further trained on the Unsplash dataset |
| FLUX.1-dev | weights/pvd_flux/ |
backbone.safetensors, part1/ and part2/ LoRA adapters |
9 dual-stream + 19 single-stream layers; rank 64 |
| FLUX.1-dev (Unsplash) | weights/pvd_flux_unsplash/ |
part1/ and part2/ LoRA adapters; reuses weights/pvd_flux/backbone.safetensors |
Further trained on the Unsplash dataset |
| Qwen-Image | weights/pvd_qwenimage/ |
backbone.safetensors, part1/ and part2/ LoRA adapters |
30 layers; rank 32 |
| Qwen-Image (Unsplash) | weights/pvd_qwen_unsplash/ |
part1/ and part2/ LoRA adapters; reuses weights/pvd_qwenimage/backbone.safetensors |
Further trained on the Unsplash dataset |
Each adapter folder contains adapter_model.safetensors and adapter_config.json. The C2I decoder and latent statistics are cache/vavae-imagenet256-f16d32-dinov2.pt and cache/latents_stats.pt.
Usage
Place this repository's weights/ and cache/ folders in the PVD code repository root. Keep the folder names as listed above. See the code repository's README for base model components and inference commands.