Xor NVFP4

Xor NVFP4 (xor-nvfp4) is a mixed-precision NVFP4 build of Xor for typed decision tasks. It pairs NVIDIA's 4-bit NVFP4 quantization of Qwen/Qwen3.6-35B-A3B with a new post-trained adapter blend, and keeps the parts of the network that decision quality depends on in BF16. It is served through the same TypeSafe-compatible /v1/systemone API as juspay/xor.

On the public JEVBench tiers and the kev transfer-v4 development suite it matches or exceeds Xor 1.1 (BF16) measured on the same hardware, at 39 GB instead of 66 GB and with lower latency.

Differences from Xor 1.1

  • Weights: NVFP4 routed experts in 26 of 40 layers; BF16 everywhere else that matters for the readout (see below)
  • Adapter: a blend of the F12 adapter (75%) and the Xor 1.1 adapter F10 (25%), instead of F10 alone
  • Size: approximately 39 GB instead of 66 GB
  • Runtime: needs two extra SGLang flags (--moe-runner-backend flashinfer_cutlass, --kv-cache-dtype bf16)
  • Validated on 2 x NVIDIA RTX PRO 6000 Blackwell (96 GB), data parallelism 2

Composition

The checkpoint starts from nvidia/Qwen3.6-35B-A3B-NVFP4 (ModelOpt MIXED_PRECISION) and replaces selected modules with BF16 weights:

Part Precision Source
310 adapter tables: linear-attention in_proj_qkv/z/a/b, out_proj (30 layers); attention q/k/v/o_proj (10 layers); shared-expert gate/up/down_proj (40 layers) BF16 adapter blend, below
lm_head BF16 Qwen3.6-35B-A3B base
Routed experts, layers 20-29 and 36-39 (14 of 40) BF16 Qwen3.6-35B-A3B base
Routed experts, layers 0-19 and 30-35 (26 of 40) NVFP4 (W4A16, group 16) NVIDIA ModelOpt
Embeddings, norms, router, vision tower, MTP as in the NVIDIA checkpoint NVIDIA

Every module moved to BF16 is removed from quantized_layers in both hf_quant_config.json and config.json, so it loads as an unquantized layer. BF16 experts are stored per expert (experts.{e}.gate_proj/up_proj/down_proj.weight), the layout of the NVFP4 checkpoint.

Adapter blend

Each of the 310 adapter tables is merged in FP32 and stored in BF16:

W = bf16( W_base + 1.125 * (B @ A)_F12 + 0.25 * (W_xor1.1 - W_base) )
  • W_base: Qwen/Qwen3.6-35B-A3B at revision 995ad96eacd98c81ed38be0c5b274b04031597b0
  • (B @ A)_F12: the F12 LoRA adapter (rank 16). Its release scale is alpha 24 / rank 16 = 1.5; 1.125 is 75% of it.
  • W_xor1.1 - W_base: the complete F10 update, recovered from juspay/xor v1.1. Xor 1.1 differs from the base in exactly these 310 tables.
  • 0.25: 25% of the F10 update.

The two adapters come from the same training lineage and change the same tables in closely aligned directions (per-table cosine similarity 0.77-0.94). The 75/25 blend kept F12's gains on kev and recovered Xor 1.1's JEVBench result. recipe/ contains the merge and build scripts; RELEASE_PROVENANCE.json records all input hashes.

Interface

Identical to Xor 1.1:

  • noul: binary probability
  • choice: categorical decision and full probability distribution (2 to 255 candidates)
  • score: expected ordinal score and full probability distribution

Requests may include an images array with up to eight image data URLs, or one video data URL. The complete request body must not exceed 8 MB.

The serving layer (candidate readout, forward and reverse option-order evaluation, per-type calibration, schema conversion) is the Xor 1.1 serving bundle and must be used for reproducible results.

Quick start

Download the release, verify and extract the serving bundle, and start Xor NVFP4 on two GPUs:

hf download juspay/xor-nvfp4 --local-dir xor-nvfp4

(cd xor-nvfp4/serving && sha256sum -c xor-nvfp4-serving.tar.gz.sha256)

mkdir -p xor-nvfp4-runtime
tar -xzf xor-nvfp4/serving/xor-nvfp4-serving.tar.gz -C xor-nvfp4-runtime --strip-components=1

cd xor-nvfp4-runtime
cp .env.example .env
sed -i "s|^MODEL_DIR=.*|MODEL_DIR=$(cd ../xor-nvfp4 && pwd)|" .env
./run.sh

run.sh verifies every model file against checksums.sha256 before starting. When the smoke test succeeds, the API is available at http://127.0.0.1:30002/v1/systemone.

curl -sS -X POST http://127.0.0.1:30002/v1/systemone \
  -H 'Content-Type: application/json' \
  --data @examples/request.json

curl -sS -X POST http://127.0.0.1:30002/v1/systemone \
  -H 'Content-Type: application/json' \
  --data @examples/image-request.json

The bundle is the Xor 1.1 serving bundle with the same compatibility server, adapted for this checkpoint (two SGLang flags, this model's checksums, model name xor-nvfp4). .env.example uses GPUs 0,1 with one replica each; for one GPU set CUDA_VISIBLE_DEVICES=0 and DP_SIZE=1. The setup requires Linux x86-64, the Hugging Face CLI, Docker Engine with Docker Compose v2, the NVIDIA Container Toolkit, and approximately 60 GB of free disk space.

Manual start (without Docker Compose)

Start SGLang with the two extra flags this checkpoint needs:

docker run -d --name xor-nvfp4-sglang --gpus '"device=0,1"' --network host --ipc host -e HF_HUB_OFFLINE=1 \
  -v "$PWD/xor-nvfp4:/models/xor:ro" \
  prakhar1611/xor-sglang@sha256:94c48d2a6cc98dc456cf93f723707ea7dd81dddfe1061e823b348d68bbe8158f \
  python3 -m sglang.launch_server --model-path /models/xor --trust-remote-code \
  --tp-size 1 --dp-size 2 --port 30000 --host 127.0.0.1 \
  --max-prefill-tokens 250000 --mem-fraction-static 0.85 \
  --moe-runner-backend flashinfer_cutlass --kv-cache-dtype bf16
  • --moe-runner-backend flashinfer_cutlass is required: the automatic MoE backend selects flashinfer_trtllm, which does not support NVFP4 MoE on these GPUs and fails at load.
  • --kv-cache-dtype bf16 overrides the FP8 KV-cache setting inherited from the NVIDIA checkpoint; all results below use a BF16 KV cache.
  • For one GPU, use --gpus '"device=0"' --dp-size 1 (not separately benchmarked).

Then start the compatibility server from the extracted bundle against it:

cd xor-nvfp4-runtime
OPENJEV_SGLANG_URL=http://127.0.0.1:30000 \
OPENJEV_TEMP_JSON='{"choice":1.1,"noul":1.4,"score":1.0}' \
OPENJEV_CACHE=0 OPENJEV_IMAGES=1 \
OPENJEV_MODEL_ID=xor-nvfp4 OPENJEV_MODEL_ALIAS=xor-nvfp4 \
  python3 server.py 30002

curl -sS -X POST http://127.0.0.1:30002/v1/systemone \
  -H 'Content-Type: application/json' --data @examples/request.json

The bundle's smoke_test.py checks for the model name xor-nvfp4.

Evaluation

Xor NVFP4 and Xor 1.1 were evaluated with identical setups on the same machine: 2 x NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), SGLang image above, data parallelism 2, the Xor 1.1 serving bundle server.py unmodified with its per-type temperatures, request caching disabled, one request at a time, and the KV cache flushed before every run. Each figure is the mean of three runs.

Public JEVBench self-run

JEVBench public tiers, typesafe adapter.

Tier Attempted Valid Xor NVFP4 correct Xor 1.1 correct Xor NVFP4 p50 / p95 Xor 1.1 p50 / p95 Xor NVFP4 ECE Xor 1.1 ECE
Easy 48 48 48 48 0.038 / 0.040 s 0.044 / 0.047 s 0.015 0.017
Original 72 72 70 70 0.038 / 0.040 s 0.045 / 0.047 s 0.109 0.107
Hard public 111 111 90 90 0.053 / 0.125 s 0.068 / 0.135 s 0.093 0.088
All public 231 231 208 208

Xor NVFP4 answered 208 of 231 correctly in each of the three runs.

kev transfer-v4 (development split)

Metric Xor NVFP4 Xor 1.1
Accuracy 84.96% 84.76%
Objective (higher is better) -0.3380 -0.3434
NLL 0.3774 0.3818
Brier 0.2110 0.2122
ECE 0.0253 0.0192
Confident-error rate 1.42% 1.37%
Latency p50 / p95 37 / 42 ms 45 / 53 ms

These are self-run results, not an official JEVBench rank. Latency is hardware-specific and was measured locally without network overhead. Calibration is slightly behind Xor 1.1 on JEVBench-hard and kev ECE.

Operational notes

  • Model files occupy approximately 39 GB; each data-parallel worker loads a complete replica.
  • Alternative hardware and parallelism settings must be validated independently before publishing performance results.
  • checksums.sha256 covers every model file; RELEASE_PROVENANCE.json records base, NVFP4 and Xor 1.1 revisions, the F12 adapter hashes, the blend, and the runtime.
  • The compatibility server has no authentication. Remote deployments must add authentication, TLS, rate limits, and request-size limits at the ingress layer.

License

Apache License 2.0. Derived from Qwen3.6-35B-A3B (Apache 2.0), NVIDIA Qwen3.6-35B-A3B-NVFP4 (Apache 2.0), and Xor 1.1 (Apache 2.0). See THIRD_PARTY_NOTICES.md.

Downloads last month
145
Safetensors
Model size
19B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for juspay/xor-nvfp4

Quantized
(851)
this model