Instructions to use AxionML/Inkling-Small-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AxionML/Inkling-Small-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="AxionML/Inkling-Small-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("AxionML/Inkling-Small-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("AxionML/Inkling-Small-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AxionML/Inkling-Small-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AxionML/Inkling-Small-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxionML/Inkling-Small-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/AxionML/Inkling-Small-NVFP4
- SGLang
How to use AxionML/Inkling-Small-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AxionML/Inkling-Small-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxionML/Inkling-Small-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AxionML/Inkling-Small-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxionML/Inkling-Small-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use AxionML/Inkling-Small-NVFP4 with Docker Model Runner:
docker model run hf.co/AxionML/Inkling-Small-NVFP4
AxionML Inkling-Small-NVFP4
Mirrored by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.
Quantized by Thinking Machines. The weights in this repository are an unmodified copy of thinkingmachines/Inkling-Small-NVFP4 (revision
b6a99534467840620d411e4cd4ad5819b2610d9c). All credit for the quantization belongs to Thinking Machines.
This is an NVFP4-quantized version of thinkingmachines/Inkling-Small (276B total parameters, 12B activated; text, image and audio input), quantized with NVIDIA Model Optimizer.
About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to 卤6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity while higher-precision FP32 accumulation protects dot-product accuracy.
Ready for commercial and non-commercial use under Apache 2.0, subject to Thinking Machines' Acceptable Use Policy.
Model Summary
| Architecture | 42-layer decoder-only multimodal MoE; hybrid local/global attention |
| Total Parameters | 276B |
| Activated Parameters | 12B |
| Experts | 256 routed (6 active) + 2 shared |
| Input | Text, image, audio (16 kHz WAV) |
| Checkpoint Size | ~171 GB |
Evaluation Results
| Benchmark | Inkling-Small |
|---|---|
| SWE-bench Verified | 80.2 |
| SWE-bench Pro (Public) | 55.9 |
| Terminal-Bench 2.1 (best harness) | 64.7 |
| SciCode | 48.7 |
| GDPval-AA v2 | 1269 |
| MCP Atlas (public) | 79.6 |
| BrowseComp (w/ context) | 77.4 |
| Toolathlon Verified | 54.4 |
| GPQA Diamond | 89.5 |
| HLE (text only) | 31.6 |
| HLE (with tools) | 47.8 |
Scores are from the Inkling-Small model card (full-precision baseline).
Quantization Details
- Quantization format: NVFP4 (group size 16), Thinking Machines' official NVFP4 release; attention, routers, vision/audio encoders and other excluded modules stay in BF16
- KV cache: not quantized
Usage
Deploy with SGLang
python3 -m sglang.launch_server \
--model-path AxionML/Inkling-Small-NVFP4 \
--tp 4 \
--trust-remote-code
Deploy with vLLM
vllm serve AxionML/Inkling-Small-NVFP4 \
--tensor-parallel-size 4 \
--trust-remote-code
Minimal commands; for tuned serving flags follow the official recipes: SGLang, vLLM, TokenSpeed.
Limitations
The base model was trained on data that may contain toxic language and societal biases. The quantized model inherits these limitations. It may generate inaccurate, biased, or offensive content. Please refer to the original model card and the upstream quantized model card for full details.
Credits
- Base model: thinkingmachines/Inkling-Small
- Quantization: thinkingmachines/Inkling-Small-NVFP4 by Thinking Machines
- Mirror: AxionML
- Downloads last month
- 336
Model tree for AxionML/Inkling-Small-NVFP4
Base model
thinkingmachines/Inkling-Small