Hunmin-397B-A17B-CUA-FP8

Block-wise FP8 (e4m3) quantization of mncai/Hunmin-397B-A17B-CUA: 128×128 weight blocks in a mixed FP8 + BF16 layout (following the official Qwen3.5-397B-A17B-FP8 scheme — precision-sensitive tensors kept in BF16). ~373 GB (about half the BF16 release), for cheaper multi-GPU serving.

For the model description, benchmarks, and evaluation methodology, see the main model card.

Quickstart (vLLM)

vllm serve <path>/Hunmin-397B-A17B-CUA-FP8 \
  --tensor-parallel-size 8 \
  --max-model-len 131072 \
  --gpu-memory-utilization 0.90 \
  --enable-prefix-caching
Downloads last month
102
Safetensors
Model size
397B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mncai/Hunmin-397B-A17B-CUA-FP8

Quantized
(1)
this model