mlx-community/Ming-Image-0.1-Design-Layer-8bit

Pre-quantized MLX tier of inclusionAI/Ming-Image-0.1-Design-Layer (MIT) for Apple Silicon, loaded by the Swift/MLX port ming-image-swift. 8-bit tier of the layer-decomposition model. It matches bf16 on every job we gate, and its memory fits a 48 GB Mac.

A Mac never has to hold the bf16 weights to use this tier. Total size: 28.3 GB; the bf16 repo is 49.8 GB. It is part of the Ming-Image (MLX) collection, next to mlx-community/Ming-Image-0.1-Design-Layer-bf16.

What is quantized

Component This tier
mllm/: MoE MLLM (attention, dense and shared-expert MLPs, the 256 routed experts) 8-bit
connector/: Qwen2-1.5B connector 8-bit
transformer/: DiT attention and feed-forward 8-bit
Kept at full precision: the MoE routers, embeddings, norms, the Qwen2.5 ViT, the f32 projection heads (mlp/), the VAE, and the DiT's conditioning layers (adaLN, embedders, final layer) bf16 / f32

Weight-only affine quantization, group size 64. The DiT never goes below 8 bits: weight-only int4 on a DiT is real quality damage that buys no speed.

The layout is the bf16 repo's upstream tree, with MLX .scales / .biases stored beside each quantized weight and a quantization block in each quantized component's config.json. The Swift loader reads it as published. Loading this repo gives parameters bit-identical to quantizing the bf16 snapshot at load time (verified for every parameter).

Quality

Gate: five real signage stills at 512, plus one at 1024 (same noise as bf16) bf16 8-bit
Layers recomposited vs the input 24.4–33.7 dB within 0.1 dB of bf16 on every job
Layer coverage — identical
Stray text outside the text layer (the hardest still) 26 chars 30 chars

Layer decomposition is driven mostly by the input image, so it is robust to quantizing the conditioning.

Full tables are in GATE-RESULTS §8 in the port's oracle (https://github.com/xocialize/ming-image-swift).

Memory and speed

Measured on an M5 Max as process phys_footprint, with MLX's buffer cache capped at 2 GB (MLXEngine's default).

  • Post-load resident: 7.6 GB. Peak process footprint: 28.1–28.4 GB for 4 to 12 layers at the 1024 bucket. The peak is the conditioning stage.
  • MLXEngine declares 7.9 GB resident plus 25.1 GB activation, which is admitted on a 48 GB Mac. Requests are capped at 12 layers, the measured envelope.
  • Speed (M5 Max, 12 steps, CFG 2.0): 137 s for 4 layers at the 512 bucket. bf16 takes 131 s.

Use (Swift / MLXEngine)

import MLXMingImage
import MLXToolKit

let package = MingImageLayerPackage(configuration: MingImageLayerConfiguration(quant: .int8))
try await package.load()
let response = try await package.run(LayerDecomposeRequest(
    image: Image(format: .png, data: try Data(contentsOf: designURL)), spec: spec, resolution: 1024)) as! LayerDecomposeResponse
// response.layers[0] is the front-most layer; response.composite is the model's recomposition

Code: https://github.com/xocialize/ming-image-swift

License

MIT, as the upstream weights. The upstream LICENSE is included.

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Ming-Image-0.1-Design-Layer-8bit

Finetuned
(4)
this model

Collection including mlx-community/Ming-Image-0.1-Design-Layer-8bit