DeepSeek-V4-Flash-0731 DSpark Graph-8 on two GB10 nodes

A reproducible deployment and benchmark package for the official deepseek-ai/DeepSeek-V4-Flash-0731 checkpoint on two NVIDIA GB10-class nodes using vLLM tensor parallelism, DSpark speculative decoding, NVFP4 DS-MLA KV cache, FlashInfer B12X MoE, RoCE, and ordinary CUDA Graph capture through batch shape 8.

Independent community recipe. This repository is not an official DeepSeek release and contains no model weights. Download the official checkpoint from DeepSeek-AI. DSpark accelerates target-verified generation; it does not change the target model's knowledge or semantics.

Headline result

On a two-node production pair, three warm single-request repetitions per shape produced:

Workload Median decode Median TTFT DSpark acceptance
Count 1–300, exact bounded output 91.89 tok/s 0.186 s 98.81%
Twenty Python dataclasses 84.43 tok/s 0.249 s 98.66%
Low-predictability technical prose 40.08 tok/s 0.235 s 32.10%

The replaced GGUF/llama.cpp lane delivered about 9.3 tok/s on the same two-node class, so the measured uplift was approximately 9.9Γ—, 9.1Γ—, and 4.3Γ— respectively. These are workload-specific measurements, not a universal throughput floor.

Graph-8 improved matched decode over Graph-4 by:

Workload Graph-4 Graph-8 Gain
Count 1–300 74.28 93.90 tok/s +26.42%
Dataclasses 75.22 86.33 tok/s +14.78%
Technical prose 32.12 40.87 tok/s +27.24%

The K5 proposer verifies six-token shapes. Capturing [1,2,4,8] rather than stopping at 4 covers that shape.

Qualification gates

  • Official checkpoint revision: 9e165c30e2704aec5d9d593cce3eebd58bbef1cb
  • 48/48 safetensor shards verified independently on both ranks
  • 96/96 draft parameters loaded; no skipped tensors
  • Configured context: 1,048,576 tokens
  • KV capacity observed at startup: 1,878,944 tokens (1.79Γ— full-context concurrency)
  • Production exact uncached retrieval: 159,979 prompt tokens, 97.50 s, cached tokens 0
  • Equivalent pre-promotion qualification pair: 949,896 prompt tokens, exact key retrieval, 1,057.98 s TTFT, cached tokens 0
  • Native OpenAI-format tool call: passed
  • Tool-result continuation: passed
  • Closed-world strict JSON schema: passed
  • Production service restart and post-restart smoke: passed

One unconstrained prose parity sample differed by one output token between speculative and non-speculative runs (256 vs. 255 tokens). Bounded outputs, retrieval keys, tool names/arguments, and strict JSON remained exact. Do not interpret this package as a claim of byte-identical unconstrained prose.

Canonical runtime profile

Model: DeepSeek-V4-Flash-0731
Tensor parallel: 2 nodes
DSpark draft tokens: 5 (K5)
KV cache: nvfp4_ds_mla
MoE backend: flashinfer_b12x
VLLM_USE_BREAKABLE_CUDAGRAPH=0
CUDA graph capture set observed: [1, 2, 4, 8]
--max-cudagraph-capture-size 8
--max-model-len 1048576
--block-size 256
RoCE v2 / NCCL IB

Pinned runtime image:

ghcr.io/anemll/dspark-vllm-gx10:0.1.1
registry digest: sha256:a83948492cf13df455170fb42885f5ef4db54fefe0feff0f841ecbff464ac9d8
validated local image ID: sha256:3430d6614a8e2925f34d059af6caf05aff42387326db4d05639a60f10f2654d8

Quick start

1. Prepare both nodes

Each node needs:

  • one GB10-class CUDA GPU;
  • Docker with Compose and NVIDIA Container Toolkit;
  • RoCE-capable connectivity visible under /dev/infiniband;
  • the complete official 0731 checkpoint at the same local path;
  • passwordless SSH only if you add your own coordinator wrapper.

Verify all 48 model shards against a trusted manifest before launch.

2. Configure node-specific environments

Copy runtime/.env.example to .env.worker and .env.head. Set a different VLLM_HOST_IP and NCCL_IB_GID_INDEX on each node. MASTER_ADDR must be the head node's fabric address on both ranks.

Do not reuse one GID index across nodes without checking sysfs: RoCEv2 GID indices can differ and can move after reboot or link changes.

3. Start worker, then head

Worker rank:

COMPOSE_DISABLE_ENV_FILE=1 NODE_RANK=1 HEADLESS=1 \
  docker compose --env-file runtime/.env.worker \
  -f runtime/docker-compose.yml -p deepseek-v4-flash-0731 up -d

Head rank:

COMPOSE_DISABLE_ENV_FILE=1 NODE_RANK=0 HEADLESS= \
  docker compose --env-file runtime/.env.head \
  -f runtime/docker-compose.yml -p deepseek-v4-flash-0731 up -d

Keep the API loopback-only unless you deliberately provide authentication and firewalling. If remote clients need access, put an authenticated reverse proxy in front of vLLM.

4. Verify

BASE_URL=http://127.0.0.1:8000 MODEL=deepseek-v4-flash-0731 \
  python3 scripts/production_proof.py --context-tokens 160000

BASE_URL=http://127.0.0.1:8000 MODEL=deepseek-v4-flash-0731 \
  python3 scripts/benchmark.py --repetitions 3 --output benchmark.json

Repository contents

  • runtime/docker-compose.yml β€” two-node vLLM/DSpark Graph-8 profile
  • runtime/.env.example β€” node-specific configuration template
  • scripts/production_proof.py β€” model, tools, strict JSON, and long-context gates
  • scripts/benchmark.py β€” streaming TTFT/decode/acceptance harness
  • benchmarks/production-summary.json β€” sanitized measured results
  • docs/METHODOLOGY.md β€” definitions, test conditions, caveats, and parity disclosure
  • NOTICE β€” upstream projects and pinned revisions

Attribution

This package builds on the official DeepSeek model and community GB10 work from Anemll, MiaAI-Lab, Tony Deangelo, and Weschera. Exact repositories and commits are recorded in NOTICE.

License

Code and documentation in this package are MIT-licensed. The official model and weights are separately provided by DeepSeek-AI under MIT; obtain them from the upstream repository. Preserve all upstream notices when redistributing upstream code.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for bertholomus/DeepSeek-V4-Flash-0731-DSpark-Graph8

Finetuned
(8)
this model