Kroma INT8 Quants

Experimental INT8 quantizations of Kroma by lodestones, primarily targeting lower-VRAM and older consumer GPUs.

Current variants

Mixed V3

A mixed-precision quantization using a combination of INT8 and W4A4.

This version currently uses standard INT8 rather than ConvRot INT8, as standard INT8 has been faster in my testing on a GTX 1660 Super.

These are experimental community quantizations and are not official Krea/Kroma releases.

Benchmarks

Steps: 8, CFG: 1, Sampler: Euler, Scheduler: Simple, Resolution: 1024x1024 Duration: ~108s ~11.30 s/it

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PotatoForge/Kroma-INT8-Quants

Base model

lodestones/Kroma
Quantized
(5)
this model