Working 2x DGX Spark (GB10) multi-node vLLM config - 20-23 tok/s with MTP + NVFP4 KV

#52
by H-K-B - opened

Sharing a fully documented working config for running this model tensor-parallel across two NVIDIA DGX Spark (GB10) nodes with vLLM, since almost no numbers exist for this hardware:

https://huggingface.co/datasets/H-K-B/glm-5.3-flash-dgx-spark-vllm

Headline results (NVFP4 modelopt quant, 64K ctx):

  • ~20-23 tok/s single-stream, ~49 tok/s aggregate at 4 concurrent
  • Stack: RoCEv2 RDMA NCCL + NVFP4 KV cache (nvfp4_ds_mla, 384 B/tok) + CUDA graphs + MTP-3 speculative decoding
  • Measured MTP acceptance: mean length 2.3-2.6, per-position ~0.70/0.45/0.26

The writeup includes the exact launch flags, four small vLLM patches the stack needs (incl. one that passes hf_overrides to the MTP draft config - without it the draft rebuilds with index_topk=2048 and dies on an unsupported NVFP4 decode shape), a measured perf ladder for each optimization step, and every failure mode we hit (each cost a ~25-min model load).

Two things worth knowing even if you're not on Sparks:

  1. PP > 1 is architecturally unavailable for Glm5NextForCausalLM (no make_empty_intermediate_tensors), and TP is head-locked to {2,4,8} - so a 3-node cluster is exactly the size this model cannot use.
  2. Validate configs with load_format="dummy" + create_engine_config() before real launches - catches nearly all init failures in ~30s instead of a full weight load.

Happy to answer questions about the setup.

Sign up or log in to comment