Model Overview

  • Model Architecture: Qwen3_5MoeForConditionalGeneration
    • Input: Text, Image, Video
    • Output: Text
  • Supported Hardware Microarchitecture: AMD MI350/MI355
  • ROCm: 7.0.0
  • PyTorch: 2.9.1
  • Transformers: 5.3.0
  • Operating System(s): Linux
  • Inference Engine: SGLang/vLLM
  • Model Optimizer: AMD-Quark (v0.12)
    • Quantized layers: Router Experts and MTP Router Experts
    • Weight quantization: OCP MXFP4, Static
    • Activation quantization: OCP MXFP4, Dynamic

Model Quantization

The model was quantized from Qwen/Qwen3.5-397B-A17B-FP8 using AMD-Quark. The weights are quantized to MXFP4 and activations are quantized to MXFP4.

Quantization scripts:

import os
from quark.torch import LLMTemplate, ModelQuantizer


# Configuration
ckpt_path = "Qwen/Qwen3.5-397B-A17B-FP8"
output_dir = "amd/Qwen3.5-397B-A17B-Quark-MXFP4"
quant_scheme = "mxfp4"
exclude_layers = ["lm_head", "model.visual.*", "*mlp.gate", "*shared_expert_gate*", "*.linear_attn.*", "*.self_attn.*", "*.shared_expert.*", "mtp.fc"]

# Get quant config from template
template = LLMTemplate.get("qwen3_5_moe")
quant_config = template.get_config(scheme=quant_scheme, exclude_layers=exclude_layers)

# Quantize with File-to-file mode
quantizer = ModelQuantizer(quant_config)
quantizer.direct_quantize_checkpoint(
    pretrained_model_path=ckpt_path,
    save_path=output_dir,
)

For further details or issues, please refer to the AMD-Quark documentation or contact the respective developers.

Evaluation

The model was evaluated on the GSM8K benchmark using the SGLang framework.

Accuracy

Benchmark Qwen/Qwen3.5-397B-A17B-FP8 amd/Qwen3.5-397B-A17B-Quark-MXFP4(this model) Recovery
gsm8k 97.3 97.3 100.0%

Reproduction

The GSM8K results were obtained with SGLang running inside the Docker image rocm/sgl-dev:v0.5.17-rocm720-mi35x-20260820. For this MTP-quantized model, SGLang must include PR #38870.

Launching server

MODEL="amd/Qwen3.5-397B-A17B-Quark-MXFP4"
AITER_FLYDSL_FORCE=1 HIP_VISIBLE_DEVICES=2,3 \
SGLANG_USE_AITER_UNIFIED_ATTN=1 SGLANG_USE_AITER=1 \
nohup python3 -m sglang.launch_server \
  --model-path "${MODEL}" --tp 2 \
  --attention-backend aiter --trust-remote-code \
  --chunked-prefill-size 32768 \
  --model-loader-extra-config '{"enable_multithread_load": true}' \
  --watchdog-timeout 1200 --mem-fraction-static 0.9 \
  --host 0.0.0.0 --port 8002 --disable-radix-cache \
  --enable-aiter-allreduce-fusion --max-running-requests 128 \
  --page-size 16 \
  --speculative-algorithm NEXTN \
  --speculative-draft-model-path "${MODEL}" \
  --speculative-num-steps 1 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 2 > "tmp_qwen35_mi35x_accuracy_$(date +%Y%m%d_%H%M%S).log" 2>&1 &

Evaluating model in a new terminal

python3 -m sglang.test.run_eval \
  --port 8002 \
  --model amd/Qwen3.5-397B-A17B-Quark-MXFP4 \
  --eval-name gsm8k \
  --num-examples 1319 \
  --num-threads 512 \
  --max-tokens 2048 \
  --chat-template-kwargs '{"enable_thinking": false}' \
  2>&1 | tee /tmp/gsm8k_run_eval.log

License

Modifications Copyright(c) 2026 Advanced Micro Devices, Inc. All rights reserved.

Downloads last month
19
Safetensors
Model size
207B params
Tensor type
U8
·
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for amd/Qwen3.5-397B-A17B-Quark-MXFP4

Quantized
(7)
this model