You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
This is a preview research build. Access is gated. By requesting access you acknowledge it is a GB10/sm_121a-specific serving build that requires a custom vLLM plugin, is not a stock-vLLM checkpoint, and inherits all terms of the base model tencent/Hy3-295B-A21B.
Log in or Sign Up to review the conditions and access this model content.
Hy3-295B-A21B — GB10 / sm_121a single-node build (R2fc)
Status: preview, gated. This card is intentionally minimal. Quantization recipe and allocation details are withheld pending an official-release decision. Full benchmark numbers and the sm_121a build details will be added when the model is officially released. Not optimized for public download.
What this is
A single-node serving build of Hy3-295B-A21B (base: tencent/Hy3-295B-A21B) that
fits and serves on one NVIDIA DGX Spark (GB10, sm_121a, 121 GB unified memory) under
vLLM. Mixture-of-Experts, 80 transformer layers + 1 MTP head, 8-of-192 routed experts.
- Resident weights: ~86.9 GiB
- Context: up to ~137k tokens on a single Spark with nothing else resident (model natively supports 262,144)
- Target hardware: NVIDIA GB10 (
sm_121a). This build's serving kernels are compiled forcompute_121aand are not portable to other GPUs as-is.
License
Inherits and preserves the license of the base model tencent/Hy3-295B-A21B. All base
license terms and use restrictions apply unchanged. See the base model's license.
Serving
Serving requires a custom vLLM plugin (codebook-quantized MoE) plus a config overlay —
this is not a stock-vLLM-loadable checkpoint. The corrected config.json is already
baked into this repo. In brief:
- Image:
ghcr.io/spark-arena/dgx-vllm-eugr-nightly(sm_121a build) - Plugin: the
gridbookcodebook-MoE plugin + the load-chain patches (X3/X4) - Launch:
--quantization gridbook,--max-num-batched-tokens 16384, fp8 KV cache
Intended use / limitations
Research and evaluation on GB10-class hardware. Chinese-trained MoE base; response-language behavior follows the base model. Not evaluated for production safety-critical use. Benchmark numbers are withheld from this preview card by design.
Provenance
Built 2026-07 for the DGX Spark single-node target. Quantization method, calibration set, and per-layer allocation are recorded internally and withheld from this preview.
- Downloads last month
- 4