Instructions to use juspay/xor-nvfp4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use juspay/xor-nvfp4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="juspay/xor-nvfp4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("juspay/xor-nvfp4") model = AutoModelForMultimodalLM.from_pretrained("juspay/xor-nvfp4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use juspay/xor-nvfp4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "juspay/xor-nvfp4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juspay/xor-nvfp4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/juspay/xor-nvfp4
- SGLang
How to use juspay/xor-nvfp4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "juspay/xor-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juspay/xor-nvfp4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "juspay/xor-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "juspay/xor-nvfp4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use juspay/xor-nvfp4 with Docker Model Runner:
docker model run hf.co/juspay/xor-nvfp4
Xor NVFP4
Xor NVFP4 (xor-nvfp4) is a mixed-precision NVFP4 build of Xor for typed decision tasks. It pairs NVIDIA's 4-bit NVFP4 quantization of Qwen/Qwen3.6-35B-A3B with a new post-trained adapter blend, and keeps the parts of the network that decision quality depends on in BF16. It is served through the same TypeSafe-compatible /v1/systemone API as juspay/xor.
On the public JEVBench tiers and the kev transfer-v4 development suite it matches or exceeds Xor 1.1 (BF16) measured on the same hardware, at 39 GB instead of 66 GB and with lower latency.
Differences from Xor 1.1
- Weights: NVFP4 routed experts in 26 of 40 layers; BF16 everywhere else that matters for the readout (see below)
- Adapter: a blend of the F12 adapter (75%) and the Xor 1.1 adapter F10 (25%), instead of F10 alone
- Size: approximately 39 GB instead of 66 GB
- Runtime: needs two extra SGLang flags (
--moe-runner-backend flashinfer_cutlass,--kv-cache-dtype bf16) - Validated on 2 x NVIDIA RTX PRO 6000 Blackwell (96 GB), data parallelism 2
Composition
The checkpoint starts from nvidia/Qwen3.6-35B-A3B-NVFP4 (ModelOpt MIXED_PRECISION) and replaces selected modules with BF16 weights:
| Part | Precision | Source |
|---|---|---|
310 adapter tables: linear-attention in_proj_qkv/z/a/b, out_proj (30 layers); attention q/k/v/o_proj (10 layers); shared-expert gate/up/down_proj (40 layers) |
BF16 | adapter blend, below |
lm_head |
BF16 | Qwen3.6-35B-A3B base |
| Routed experts, layers 20-29 and 36-39 (14 of 40) | BF16 | Qwen3.6-35B-A3B base |
| Routed experts, layers 0-19 and 30-35 (26 of 40) | NVFP4 (W4A16, group 16) | NVIDIA ModelOpt |
| Embeddings, norms, router, vision tower, MTP | as in the NVIDIA checkpoint | NVIDIA |
Every module moved to BF16 is removed from quantized_layers in both hf_quant_config.json and config.json, so it loads as an unquantized layer. BF16 experts are stored per expert (experts.{e}.gate_proj/up_proj/down_proj.weight), the layout of the NVFP4 checkpoint.
Adapter blend
Each of the 310 adapter tables is merged in FP32 and stored in BF16:
W = bf16( W_base + 1.125 * (B @ A)_F12 + 0.25 * (W_xor1.1 - W_base) )
W_base:Qwen/Qwen3.6-35B-A3Bat revision995ad96eacd98c81ed38be0c5b274b04031597b0(B @ A)_F12: the F12 LoRA adapter (rank 16). Its release scale is alpha 24 / rank 16 = 1.5; 1.125 is 75% of it.W_xor1.1 - W_base: the complete F10 update, recovered from juspay/xorv1.1. Xor 1.1 differs from the base in exactly these 310 tables.- 0.25: 25% of the F10 update.
The two adapters come from the same training lineage and change the same tables in closely aligned directions (per-table cosine similarity 0.77-0.94). The 75/25 blend kept F12's gains on kev and recovered Xor 1.1's JEVBench result. recipe/ contains the merge and build scripts; RELEASE_PROVENANCE.json records all input hashes.
Interface
Identical to Xor 1.1:
noul: binary probabilitychoice: categorical decision and full probability distribution (2 to 255 candidates)score: expected ordinal score and full probability distribution
Requests may include an images array with up to eight image data URLs, or one video data URL. The complete request body must not exceed 8 MB.
The serving layer (candidate readout, forward and reverse option-order evaluation, per-type calibration, schema conversion) is the Xor 1.1 serving bundle and must be used for reproducible results.
Quick start
Download the release, verify and extract the serving bundle, and start Xor NVFP4 on two GPUs:
hf download juspay/xor-nvfp4 --local-dir xor-nvfp4
(cd xor-nvfp4/serving && sha256sum -c xor-nvfp4-serving.tar.gz.sha256)
mkdir -p xor-nvfp4-runtime
tar -xzf xor-nvfp4/serving/xor-nvfp4-serving.tar.gz -C xor-nvfp4-runtime --strip-components=1
cd xor-nvfp4-runtime
cp .env.example .env
sed -i "s|^MODEL_DIR=.*|MODEL_DIR=$(cd ../xor-nvfp4 && pwd)|" .env
./run.sh
run.sh verifies every model file against checksums.sha256 before starting. When the smoke test succeeds, the API is available at http://127.0.0.1:30002/v1/systemone.
curl -sS -X POST http://127.0.0.1:30002/v1/systemone \
-H 'Content-Type: application/json' \
--data @examples/request.json
curl -sS -X POST http://127.0.0.1:30002/v1/systemone \
-H 'Content-Type: application/json' \
--data @examples/image-request.json
The bundle is the Xor 1.1 serving bundle with the same compatibility server, adapted for this checkpoint (two SGLang flags, this model's checksums, model name xor-nvfp4). .env.example uses GPUs 0,1 with one replica each; for one GPU set CUDA_VISIBLE_DEVICES=0 and DP_SIZE=1. The setup requires Linux x86-64, the Hugging Face CLI, Docker Engine with Docker Compose v2, the NVIDIA Container Toolkit, and approximately 60 GB of free disk space.
Manual start (without Docker Compose)
Start SGLang with the two extra flags this checkpoint needs:
docker run -d --name xor-nvfp4-sglang --gpus '"device=0,1"' --network host --ipc host -e HF_HUB_OFFLINE=1 \
-v "$PWD/xor-nvfp4:/models/xor:ro" \
prakhar1611/xor-sglang@sha256:94c48d2a6cc98dc456cf93f723707ea7dd81dddfe1061e823b348d68bbe8158f \
python3 -m sglang.launch_server --model-path /models/xor --trust-remote-code \
--tp-size 1 --dp-size 2 --port 30000 --host 127.0.0.1 \
--max-prefill-tokens 250000 --mem-fraction-static 0.85 \
--moe-runner-backend flashinfer_cutlass --kv-cache-dtype bf16
--moe-runner-backend flashinfer_cutlassis required: the automatic MoE backend selectsflashinfer_trtllm, which does not support NVFP4 MoE on these GPUs and fails at load.--kv-cache-dtype bf16overrides the FP8 KV-cache setting inherited from the NVIDIA checkpoint; all results below use a BF16 KV cache.- For one GPU, use
--gpus '"device=0"' --dp-size 1(not separately benchmarked).
Then start the compatibility server from the extracted bundle against it:
cd xor-nvfp4-runtime
OPENJEV_SGLANG_URL=http://127.0.0.1:30000 \
OPENJEV_TEMP_JSON='{"choice":1.1,"noul":1.4,"score":1.0}' \
OPENJEV_CACHE=0 OPENJEV_IMAGES=1 \
OPENJEV_MODEL_ID=xor-nvfp4 OPENJEV_MODEL_ALIAS=xor-nvfp4 \
python3 server.py 30002
curl -sS -X POST http://127.0.0.1:30002/v1/systemone \
-H 'Content-Type: application/json' --data @examples/request.json
The bundle's smoke_test.py checks for the model name xor-nvfp4.
Evaluation
Xor NVFP4 and Xor 1.1 were evaluated with identical setups on the same machine: 2 x NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB), SGLang image above, data parallelism 2, the Xor 1.1 serving bundle server.py unmodified with its per-type temperatures, request caching disabled, one request at a time, and the KV cache flushed before every run. Each figure is the mean of three runs.
Public JEVBench self-run
JEVBench public tiers, typesafe adapter.
| Tier | Attempted | Valid | Xor NVFP4 correct | Xor 1.1 correct | Xor NVFP4 p50 / p95 | Xor 1.1 p50 / p95 | Xor NVFP4 ECE | Xor 1.1 ECE |
|---|---|---|---|---|---|---|---|---|
| Easy | 48 | 48 | 48 | 48 | 0.038 / 0.040 s | 0.044 / 0.047 s | 0.015 | 0.017 |
| Original | 72 | 72 | 70 | 70 | 0.038 / 0.040 s | 0.045 / 0.047 s | 0.109 | 0.107 |
| Hard public | 111 | 111 | 90 | 90 | 0.053 / 0.125 s | 0.068 / 0.135 s | 0.093 | 0.088 |
| All public | 231 | 231 | 208 | 208 |
Xor NVFP4 answered 208 of 231 correctly in each of the three runs.
kev transfer-v4 (development split)
| Metric | Xor NVFP4 | Xor 1.1 |
|---|---|---|
| Accuracy | 84.96% | 84.76% |
| Objective (higher is better) | -0.3380 | -0.3434 |
| NLL | 0.3774 | 0.3818 |
| Brier | 0.2110 | 0.2122 |
| ECE | 0.0253 | 0.0192 |
| Confident-error rate | 1.42% | 1.37% |
| Latency p50 / p95 | 37 / 42 ms | 45 / 53 ms |
These are self-run results, not an official JEVBench rank. Latency is hardware-specific and was measured locally without network overhead. Calibration is slightly behind Xor 1.1 on JEVBench-hard and kev ECE.
Operational notes
- Model files occupy approximately 39 GB; each data-parallel worker loads a complete replica.
- Alternative hardware and parallelism settings must be validated independently before publishing performance results.
checksums.sha256covers every model file;RELEASE_PROVENANCE.jsonrecords base, NVFP4 and Xor 1.1 revisions, the F12 adapter hashes, the blend, and the runtime.- The compatibility server has no authentication. Remote deployments must add authentication, TLS, rate limits, and request-size limits at the ingress layer.
License
Apache License 2.0. Derived from Qwen3.6-35B-A3B (Apache 2.0), NVIDIA Qwen3.6-35B-A3B-NVFP4 (Apache 2.0), and Xor 1.1 (Apache 2.0). See THIRD_PARTY_NOTICES.md.
- Downloads last month
- 145
Model tree for juspay/xor-nvfp4
Base model
Qwen/Qwen3.6-35B-A3B