Qwen3-30B-A3B-12L-MXFP8-test

This is a 12-layer test checkpoint derived from a Qwen3-30B-A3B AutoRound MXFP8 checkpoint. It is intended for vLLM loading and inference tests, including MXFP8 linear and fused MoE coverage. It is not intended for quality evaluation or production use.

Configuration

  • Architecture: Qwen3MoeForCausalLM
  • Transformer layers: 12 (the first 12 layers of the source checkpoint)
  • Weight format: MXFP8, group size 32
  • Activation format: dynamic MXFP8, group size 32
  • Packing format: auto_round:llm_compressor
  • AutoRound version: 0.14.2

The tokenizer, embedding, final normalization, and language model head are retained. Quantization metadata is trimmed to the retained layers.

vLLM test

This checkpoint is prepared for the following vLLM test model identifier:

INCModel/Qwen3-30B-A3B-12L-MXFP8-test

The test uses eager execution and generates eight tokens from the prompt The capital of France is.

Downloads last month
19
Safetensors
Model size
8B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for INCModel/Qwen3-30B-A3B-12L-MXFP8-test

Quantized
(144)
this model