Instructions to use mlx-community/Ming-Image-0.1-Design-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/Ming-Image-0.1-Design-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Ming-Image-0.1-Design-4bit mlx-community/Ming-Image-0.1-Design-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
mlx-community/Ming-Image-0.1-Design-4bit
Pre-quantized MLX tier of inclusionAI/Ming-Image-0.1-Design (MIT) for Apple Silicon, loaded by the Swift/MLX port ming-image-swift. 4-bit tier, for memory-bound Macs. It fits a 36 GB Mac, at a small, measured quality cost. Keep OCR-verify and re-roll in the loop for text.
A Mac never has to hold the bf16 weights to use this tier. Total size: 20.5 GB; the bf16 repo is 49.8 GB. It is part of the Ming-Image (MLX) collection, next to mlx-community/Ming-Image-0.1-Design-bf16.
What is quantized
| Component | This tier |
|---|---|
mllm/: MoE MLLM (attention, dense and shared-expert MLPs, the 256 routed experts) |
4-bit |
connector/: Qwen2-1.5B connector |
8-bit |
transformer/: DiT attention and feed-forward |
8-bit |
Kept at full precision: the MoE routers, embeddings, norms, the Qwen2.5 ViT, the f32 projection heads (mlp/), the VAE, and the DiT's conditioning layers (adaLN, embedders, final layer) |
bf16 / f32 |
Weight-only affine quantization, group size 64. The DiT never goes below 8 bits: weight-only int4 on a DiT is real quality damage that buys no speed.
The layout is the bf16 repo's upstream tree, with MLX
.scales / .biases stored beside each quantized weight and a quantization block in each quantized component's
config.json. The Swift loader reads it as published. Loading this repo gives parameters bit-identical to
quantizing the bf16 snapshot at load time (verified for every parameter).
Quality
| Gate (same noise and scorers as bf16) | bf16 | 8-bit | 4-bit |
|---|---|---|---|
| Render vs the fp32-truth render (1024², same noise) | 41.9 dB | 36.5 dB | 29.0 dB |
| Text boards: exact strings · fully correct boards | 117/122 · 12/16 | 118/122 · 12/16 | 116/122 · 10/16 |
| Zones reserved for live text: clean · text-free | 19/20 · 20/20 | 19/20 · 20/20 | 19/20 · 20/20 |
| Native alpha: the recipe's usable subjects | 5/6 | 5/6 | 4/6 |
| Conditioning drift vs fp32 truth (learnable relL2, short / json) | 0.122 / 0.140 | 0.118 / 0.155 | 0.290 / 0.291 |
The misses are single-character typos on otherwise-correct boards ("Kexmote", "6 frt 8 in"). Composition and zones hold.
Native alpha is the weakest point of this tier. On the gate's six subjects, the alpha recipe loses one (a potted plant). In a live run at a different seed, a coffee cup that comes back transparent on the first try at bf16 and 8-bit came back opaque on both phrases. For transparent assets, use the 8-bit tier or bf16, and check alpha on import either way.
Full tables are in GATE-RESULTS §8 in the port's oracle (https://github.com/xocialize/ming-image-swift).
Memory and speed
Measured on an M5 Max as process phys_footprint, with MLX's buffer cache capped at 2 GB (MLXEngine's default).
- Post-load resident: 7.6 GB. Peak process footprint: 19.7 GB at 1024², 2048² and 2560×1440 alike. The peak is the conditioning stage, when the 4-bit MLLM loads, conditions and is released.
- The 2048²-class VAE decode is bounded (chunked mid-attention plus a tiled up path, exact).
- MLXEngine declares 7.9 GB resident plus 14.6 GB activation, which is admitted on a 36 GB Mac.
- Speed at 12 steps (M5 Max): 42 s at 1024² and 243 s at 2048². bf16 takes 39 s and 245 s.
Use (Swift / MLXEngine)
import MLXMingImage
import MLXToolKit
let package = MingImageT2IPackage(configuration: MingImageConfiguration(quant: .int4))
try await package.load()
let response = try await package.run(T2IRequest(prompt: "…", width: 1024, height: 1024, seed: 42)) as! T2IResponse
Code: https://github.com/xocialize/ming-image-swift
License
MIT, as the upstream weights. The upstream LICENSE is included.
4-bit
Model tree for mlx-community/Ming-Image-0.1-Design-4bit
Base model
inclusionAI/Ming-Image-0.1-Design