Model Overview

Description:

This model is a quantized version of DeepSeek AI's DeepSeek R1 model, produced with the official NVIDIA TensorRT Model Optimizer DeepSeek recipe and calibrated with the MLPerf Inference deepseek-r1 calibration dataset. It is quantized with FP4 weights and activations (the same quantization scheme as nvidia/DeepSeek-R1-FP4-v2, including the wo module in attention layers) and is ready for inference with TensorRT-LLM. This model is ready for commercial/non-commercial use.

Third-Party Community Consideration

This model is not owned or developed by NVIDIA or DeepSeek AI. It has been produced by a third party from the open DeepSeek R1 checkpoint for the MLPerf Inference deepseek-r1 benchmark. See the original DeepSeek R1 Model Card.

License/Terms of Use:

MIT

Model Architecture:

Architecture Type: Transformers

Network Architecture: DeepSeek R1

Input:

Input Type(s): Text

Input Format(s): String

Input Parameters: 1D (One Dimensional): Sequences

Other Properties Related to Input: Context length up to 128K. Per DeepSeek's usage recommendations: temperature 0.5–0.7 (0.6 recommended), no system prompt, and for math problems ask for step-by-step reasoning with the final answer in \boxed{}.

Output:

Output Type(s): Text

Output Format: String

Output Parameters: 1D (One Dimensional): Sequences

Software Integration:

Supported Runtime Engine(s):

  • TensorRT-LLM

Supported Hardware Microarchitecture Compatibility:

  • NVIDIA Blackwell

Preferred Operating System(s):

  • Linux

Model Version(s):

Quantized with NVIDIA TensorRT Model Optimizer, main branch @ 089c06e4 (July 2026), using the examples/deepseek recipe (deepseek_v3/ptq.py for calibration, deepseek_v3/quantize_fp8_to_nvfp4.sh for weight conversion).

Datasets:

Calibration Dataset: mlperf_deepseek_r1_calibration_dataset_500_fp8_eval — the official 500-sample MLPerf Inference deepseek-r1 calibration set. See the MLPerf deepseek-r1 reference README for dataset details and download instructions.

  • Data collection method: Hybrid: Human, Automated
  • Labeling method: Hybrid: Human, Automated

Inference:

Engine: TensorRT-LLM

Test Hardware: NVIDIA GB200 NVL72

Post Training Quantization

This model was obtained by quantizing the weights and activations of DeepSeek R1 to FP4 data type, ready for inference with TensorRT-LLM. Only the weights and activations of the linear operators within the transformer blocks are quantized (fine-grained FP4 block scaling with group size 16; the KV cache is quantized to FP8; MLA attention projections remain in higher precision). This optimization reduces the number of bits per parameter from 8 to 4, reducing the disk size and GPU memory requirements by approximately 1.6x.

Calibration runs the official recipe unmodified except for the calibration dataloader, which reads the MLPerf deepseek-r1 calibration set (500 samples, max sequence length 2048, batch size 4).

Usage

Deploy with TensorRT-LLM

To deploy the quantized checkpoint with TensorRT-LLM LLM API, use the following sample code (requires 8 GPUs of Blackwell generation or later, and TensorRT-LLM built from source with the latest main branch):

from tensorrt_llm import SamplingParams
from tensorrt_llm._torch import LLM

def main():

    prompts = [
        "Hello, my name is",
        "The president of the United States is",
        "The capital of France is",
        "The future of AI is",
    ]
    sampling_params = SamplingParams(max_tokens=32)

    llm = LLM(model="centml/DeepSeek-R1-NVFP4-v2-mlpinf", tensor_parallel_size=8, enable_attention_dp=True)

    outputs = llm.generate(prompts, sampling_params)

    # Print the outputs.
    for output in outputs:
        prompt = output.prompt
        generated_text = output.outputs[0].text
        print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")


# The entry point of the program need to be protected for spawning processes.
if __name__ == '__main__':
    main()
Downloads last month
525
Safetensors
Model size
394B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for centml/DeepSeek-R1-NVFP4-v2-mlpinf

Finetuned
(323)
this model