Muse-Glimmer-30B-FP8

fp8 quantization of meta-models/Muse-Glimmer-30B, produced with compressed-tensors by streaming the checkpoint tensor-by-tensor (the model is never fully instantiated).

8-bit float weights (float8_e4m3fn), per output channel, with dynamic per-token activation quantization. Highest fidelity of the set and the largest.

All quantizations of this model

Variant Format Size vs BF16 Mean rel. error Linears quantized Left BF16
Muse-Glimmer-30B-FP8 ← this one float-quantized 34.40 GB 58% 0.0266 416 0
Muse-Glimmer-30B-NVFP4 nvfp4-pack-quantized 23.38 GB 39% 0.0947 416 0
Muse-Glimmer-30B-INT4-W4A16 pack-quantized 22.20 GB 37% 0.1156 416 0

Mean relative error is ||dequant(W) - W|| / ||W||, averaged over a sample of quantized Linear layers, measured against the original BF16 weights. Lower is better.

This variant

Format float-quantized
Weight bits 8
Group size per-channel (no grouping)
Strategy channel
Linears quantized 416
Left in BF16 0
Shards 8
On disk 34.40 GB
Mean relative error 0.0266
Shape/dtype conformance failures 0

Use with vLLM

vllm serve dudeman2512/Muse-Glimmer-30B-FP8

How this was made

Every produced tensor is checked for shape/dtype conformance against what the server expects, then reconstruction error is measured against the source BF16 weights, before anything is published. The numbers in the table above are those measurements — not estimates.

Downloads last month
19
Safetensors
Model size
30B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dudeman2512/Muse-Glimmer-30B-FP8

Quantized
(150)
this model