AxionML Step-3.7-Flash-NVFP4

Mirrored by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.

Quantized by StepFun. The weights in this repository are an unmodified copy of stepfun-ai/Step-3.7-Flash-NVFP4 (revision 4275532ffd9a9496ff36b7a2dc4a9db1048da438). All credit for the quantization belongs to StepFun.

This is an NVFP4-quantized version of stepfun-ai/Step-3.7-Flash (198B total parameters, ~11B activated), quantized with NVIDIA Model Optimizer.

About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity while higher-precision FP32 accumulation protects dot-product accuracy.

Ready for commercial and non-commercial use under Apache 2.0.

Model Summary

Architecture Sparse MoE vision-language model
Total Parameters 198B (196B language backbone + 1.8B vision encoder)
Activated Parameters ~11B
Experts 288 routed
Reasoning Selectable levels: low / medium / high
Context Length 256K tokens
Checkpoint Size ~129 GB

Evaluation Results

Step 3.7 Flash benchmark results

Benchmark Step 3.7 Flash
SimpleVQA (Search) 79.2
V* (Python) 95.3
ClawEval-1.1 67.1
Toolathlon 49.5
HLE (w/ tool) 48.1
SWE-Bench Pro 56.3
Terminal-Bench 2.1 59.5
GDPVal-AA 45.8

Scores are from the Step-3.7-Flash model card (full-precision baseline).

Quantization Details

  • Quantization format: NVFP4 (W4A4, group size 16) on the MoE linear layers; attention, router and vision encoder kept in higher precision
  • KV cache: FP8
  • Tool: NVIDIA Model Optimizer v0.45.0

Usage

Deploy with SGLang

sglang serve \
    --model-path AxionML/Step-3.7-Flash-NVFP4 \
    --tp 4 --ep 4 \
    --moe-runner-backend flashinfer_trtllm \
    --kv-cache-dtype fp8_e4m3 \
    --quantization modelopt_fp4 \
    --trust-remote-code \
    --reasoning-parser step3p5 \
    --tool-call-parser step3p5 \
    --attention-backend trtllm_mha

Deploy with vLLM

vllm serve AxionML/Step-3.7-Flash-NVFP4 \
    --tensor-parallel-size 4 \
    --gpu-memory-utilization 0.9 \
    --enable-expert-parallel \
    --trust-remote-code \
    --quantization modelopt \
    --kv-cache-dtype fp8 \
    --reasoning-parser step3p5 \
    --enable-auto-tool-choice \
    --tool-call-parser step3p5 \
    --async-scheduling

StepFun publishes prebuilt images: lmsysorg/sglang:dev-step-3.7-flash and vllm/vllm-openai:stepfun37. The NVFP4 build requires ModelOpt quantization and an FP8 KV cache.

Limitations

The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the original model card and the upstream quantized model card for full details.

Credits

Downloads last month
-
Safetensors
Model size
104B params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AxionML/Step-3.7-Flash-NVFP4

Quantized
(43)
this model

Space using AxionML/Step-3.7-Flash-NVFP4 1