Qwen-Image-2.1 7B β INT4 (W4A8) ConvRot for ComfyUI
INT4 quantized weights of Qwen-Image-2.1 for fast, low-VRAM inference in ComfyUI.
This is a modified (quantized) version of the Qwen-Image-2.1 model. It is not an official Qwen release and is not endorsed by the Qwen team.
About Qwen-Image-2.1
A unified text-to-image generation and image editing model in the Qwen family. With just 7B parameters in its visual generation component (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility.
Highlights
- Efficient Image Generation β combines strong visual performance with fast inference and a compact design, making high-quality image creation accessible across a wide range of creative workflows.
- Flexible Creative Control β supports diverse inputs, outputs, and localized edits, giving creators the flexibility to explore ideas and refine details within a unified workflow.
Key improvements in 2.1
- Compact and Efficient β lightweight architecture with mixed-granularity attention and prefix KV cache reuse delivers strong image quality at low computational cost.
- Native Transparency, Unified Creation and Editing β generate regular or transparent (RGBA) images from text, edit transparent layers, and extract subjects from photographs β all in one model.
- Versatile Editing β support up to 10 reference images, specify local edits via circles, painted annotations, or separate masks, and preserve identity for people and products.
- Realistic Textures and Refined Aesthetics β improved typography, portrait lighting, and fine details for more visually compelling results.
Model tree for tsolful/Qwen2.1_INT4W4A8
Base model
Qwen/Qwen-Image-2.1