AxionML Inkling-Small-NVFP4

Mirrored by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.

Quantized by Thinking Machines. The weights in this repository are an unmodified copy of thinkingmachines/Inkling-Small-NVFP4 (revision b6a99534467840620d411e4cd4ad5819b2610d9c). All credit for the quantization belongs to Thinking Machines.

This is an NVFP4-quantized version of thinkingmachines/Inkling-Small (276B total parameters, 12B activated; text, image and audio input), quantized with NVIDIA Model Optimizer.

About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to 卤6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity while higher-precision FP32 accumulation protects dot-product accuracy.

Ready for commercial and non-commercial use under Apache 2.0, subject to Thinking Machines' Acceptable Use Policy.

Model Summary

Architecture 42-layer decoder-only multimodal MoE; hybrid local/global attention
Total Parameters 276B
Activated Parameters 12B
Experts 256 routed (6 active) + 2 shared
Input Text, image, audio (16 kHz WAV)
Checkpoint Size ~171 GB

Evaluation Results

Benchmark Inkling-Small
SWE-bench Verified 80.2
SWE-bench Pro (Public) 55.9
Terminal-Bench 2.1 (best harness) 64.7
SciCode 48.7
GDPval-AA v2 1269
MCP Atlas (public) 79.6
BrowseComp (w/ context) 77.4
Toolathlon Verified 54.4
GPQA Diamond 89.5
HLE (text only) 31.6
HLE (with tools) 47.8

Scores are from the Inkling-Small model card (full-precision baseline).

Quantization Details

  • Quantization format: NVFP4 (group size 16), Thinking Machines' official NVFP4 release; attention, routers, vision/audio encoders and other excluded modules stay in BF16
  • KV cache: not quantized

Usage

Deploy with SGLang

python3 -m sglang.launch_server \
    --model-path AxionML/Inkling-Small-NVFP4 \
    --tp 4 \
    --trust-remote-code

Deploy with vLLM

vllm serve AxionML/Inkling-Small-NVFP4 \
    --tensor-parallel-size 4 \
    --trust-remote-code

Minimal commands; for tuned serving flags follow the official recipes: SGLang, vLLM, TokenSpeed.

Limitations

The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the original model card and the upstream quantized model card for full details.

Credits

Downloads last month
336
Safetensors
Model size
156B params
Tensor type
I64
路
F32
路
BF16
路
F8_E4M3
路
U8
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for AxionML/Inkling-Small-NVFP4

Quantized
(54)
this model