Kandinsky 6 — ComfyUI INT8 ConvRot

English | 简体中文

Three single-file INT8 ConvRot DiTs for joint video/audio generation with ComfyUI Kandinsky T8. The complete video, audio and internal text branches remain in each DiT. Shared text encoders and decoders are separate, unmodified files.

Downloads

Place the selected model in ComfyUI/models/diffusion_models:

Model Download Size Workflow
Lite k6_lite_int8_convrot.safetensors 4.87 GB Base: 50 steps, CFG 5
Lite distill k6_lite_distill_int8_convrot.safetensors 3.77 GB PiFlow: 10 steps, CFG 1
Pro distill k6_pro_distill_int8_convrot.safetensors 34.97 GB PiFlow: 10 steps, CFG 1

For example, download Lite distill and the independent components directly into your ComfyUI models directory:

hf download t8star/Kandinsky-Comfy diffusion_models/k6_lite_distill_int8_convrot.safetensors \
  text_encoders/qwen_2.5_vl_7b.safetensors text_encoders/clip_l.safetensors \
  vae/hunyuan_video_vae_bf16.safetensors audio_vae/v1-44.pth \
  audio_vae/bigvgan_vocoder/config.json audio_vae/bigvgan_vocoder/bigvgan_generator.pt \
  --local-dir ComfyUI/models

The repository folders match ComfyUI's model folders. Keep audio_vae/bigvgan_vocoder/config.json beside bigvgan_generator.pt. Import a workflow, select the model in UNETLoader, and use weight_dtype=default.

PiFlow requires CFG=1, denoise=1 and one global, full-range conditioning; masks are limited to the Image workflow's clean reference tail.

Download the six ready GUI workflows, unzip and drag a root-level JSON onto the ComfyUI canvas. Each selects the corresponding model. Use Kandinsky T8 0.1.4+ and restart ComfyUI; the image example is installed as ComfyUI/input/kandinsky6_i2va_portrait.png. The pack includes this image and actual run verification records.

Requirements and validation

Requires ComfyUI 0.39.0+, comfy-kitchen 0.2.37+, the T8 nodes and an NVIDIA CUDA GPU with INT8 Tensor Core support. Default output is 864×480, 121 frames, 24 fps, with audio. Pro distill was tested using dynamic VRAM/offloading on RTX 5090 Laptop 24 GB / 64 GB RAM.

All three variants completed full text and image workflows; all six outputs passed audio/video decoding checks. See validation. Results cover the listed hardware and example prompts.

ConvRot uses normalized grouped Hadamard rotation with per-output-channel INT8 weights and dynamically quantized activations. Lite has 920 quantized Linear layers; Pro distill has 1728. Sensitive projections, modulation, norms, biases and embeddings retain their original tensors. SHA256SUMS, conversion manifests and MODEL_INDEX.json record checksums and pinned sources.

Credits and licenses

The converted Kandinsky DiTs are derived from Kandinsky Lab and retain the MIT license. INT8 conversion by T8star. Repository-level MIT metadata applies to the Kandinsky DiTs; independent components retain their own licenses:

Component Source License
Qwen 2.5 VL 7B Comfy-Org / Qwen Apache 2.0
CLIP-L Comfy-Org / OpenAI CLIP MIT
HunyuanVideo VAE Comfy-Org / Tencent Tencent Hunyuan Community
TOD audio decoder v1-44.pth MMAudio / Sony Research CC BY-NC 4.0
BigVGAN 44 kHz NVIDIA MIT

Shared components are redistributed unchanged. License texts and references are retained in licenses and NOTICE. TOD weights are licensed for noncommercial use; HunyuanVideo VAE retains its upstream territory and usage restrictions.

T8star

Bilibili · YouTube · API · Free gallery

Online AI apps · ComfyUI distribution · Hugging Face

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for t8star/Kandinsky-Comfy

Quantized
(2)
this model