Instructions to use LiquidAI/d1-3B-w8a8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LiquidAI/d1-3B-w8a8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="LiquidAI/d1-3B-w8a8", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("LiquidAI/d1-3B-w8a8", trust_remote_code=True) model = AutoModelForMultimodalLM.from_pretrained("LiquidAI/d1-3B-w8a8", trust_remote_code=True, device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use LiquidAI/d1-3B-w8a8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LiquidAI/d1-3B-w8a8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LiquidAI/d1-3B-w8a8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/LiquidAI/d1-3B-w8a8
- SGLang
How to use LiquidAI/d1-3B-w8a8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "LiquidAI/d1-3B-w8a8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LiquidAI/d1-3B-w8a8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "LiquidAI/d1-3B-w8a8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LiquidAI/d1-3B-w8a8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use LiquidAI/d1-3B-w8a8 with Docker Model Runner:
docker model run hf.co/LiquidAI/d1-3B-w8a8
d1-3B-w8a8
d1-3B-w8a8 is the INT8 quantized version of d1-3B, a 3B parameter
decision model, intended for devices with INT8 matrix multiplication (torch._int_mm) on INT8 tensor
cores (e.g. NVIDIA Jetson devices, and NVIDIA GPUs from Ampere on). You give it a state (text, JSON, images, or
a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with
zero output tokens.
📖 For the model's description and benchmark results, see the base model card: LiquidAI/d1-3B.
- W8A8: INT8 weights and INT8 activations on the language model, so its matrix products run as INT8
matrix multiplications (
torch._int_mm) on the GPU's INT8 tensor cores. - Quantized ahead of time with torchao: the checkpoint stores the INT8 weights, and loads as INT8 with no calibration or conversion at load time.
- Smaller: 3.8 GB instead of 6.2 GB for the bf16 checkpoint, which leaves more memory to the rest of the application (on a Jetson, the GPU shares it with the CPU).
Find more information about open d1 in our blog post.
On other devices (CPUs, Apple silicon, AMD GPUs, NVIDIA GPUs before Ampere), use d1-3B.
🗒️ Model Details
| Model | Parameters | Description |
|---|---|---|
| LFM2.5-VL-3B | 3.1B | General-purpose vision-language model (base) |
| d1-3B | 3.1B | Post-trained for single-pass, calibrated decisions |
| d1-3B-w8a8 | 3.1B | d1-3B in INT8 (W8A8), for devices with INT8 tensor cores (e.g. Jetson) |
Quantization
| Method | torchao Int8DynamicActivationInt8WeightConfig |
| Weights | INT8, symmetric, one scale per output channel, round to nearest |
| Activations | INT8, symmetric, one scale per token, computed at run time |
| Quantized layers | the 166 linear layers of the language model (attention, short convolution and MLP projections) |
| Kept in bf16 | the vision encoder, the multimodal projector, the embeddings and the output head |
We recommend d1-3B-w8a8 wherever a pipeline on such a device needs a yes/no, a pick from named options, or a rating: routing and triage, moderation, intent and topic classification, extraction checks, reranking, agent guardrails, and visual inspection. It is not a chat model and does not write text.
🏃 How to use
The model needs a GPU where torch._int_mm runs on INT8 tensor cores (NVIDIA, from Ampere on, e.g. Jetson
Orin), PyTorch, and torchao. On a Jetson with JetPack 6, install PyTorch from the Jetson AI Lab index, then the
rest:
pip install --index-url https://pypi.jetson-ai-lab.io/jp6/cu126 torch torchvision triton
pip install "transformers>=5.19" "torchao>=0.18" pillow
On another NVIDIA GPU, pip install torch torchvision "transformers>=5.19" "torchao>=0.18" pillow.
The model ships its own code, so load it with trust_remote_code=True. Call model.compile(): torchao's INT8
kernels are only fast compiled.
from transformers import AutoModel
from transformers.image_utils import load_image
model = AutoModel.from_pretrained("LiquidAI/d1-3B-w8a8", trust_remote_code=True).to("cuda")
model.compile(mode="reduce-overhead")
# Text: several named questions over one state, answered in one pass
questions = {
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?",
},
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "App or site faults",
"fraud": "Suspected unauthorised use",
},
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"],
},
}
print(model.system_one("I was charged twice this month, please refund one of them.", questions))
# Image: the photo is the whole state
image = load_image("http://images.cocodataset.org/val2017/000000039769.jpg") # two cats on a sofa
cats = {
"type": "choice",
"instructions": "How many cats are there?",
"criteria": {"one": "One", "two": "Two", "more": "Three or more"},
}
print(model.system_one(None, {"cats": cats}, images=[image]))
# Batch: many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
print(model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets]))
The first call with a new shape compiles it, which takes a while on a Jetson, so warm up the shapes you serve before timing or serving.
| call | |
|---|---|
system_one(state, questions, images=None) |
Named questions over one state, in one pass. The state and its images are read once for all questions. |
system_one_batch([(state, questions[, images]), ...]) |
Many requests, packed with no padding. |
A state is a string, any JSON value, or None when the images are the whole state.
Questions and answers
Questions follow the Decision Index schema: type, instructions, and criteria.
type |
criteria |
answer fields |
|---|---|---|
noul: yes or no |
optional: {"true": "...", "false": "..."} to define each side |
noul: P(yes) |
choice: one of named options |
{name: description} |
choice, confidence, probabilities |
score: 2 to 10 ordered levels |
a list of level descriptions, lowest first | score (the expected level), confidence, probabilities, legend |
Each call returns {"answers": {name: answer}, "usage": {"input_tokens": n, "output_tokens": 0}}.
⚡ Speed
Warm calls, one request at a time: a single question, three questions over one state, a 3.4k-token state and a
384 px image. The last column is throughput with 64 states packed into one pass. We measure on an NVIDIA Jetson
AGX Orin 64 GB (MAXN, JetPack 6.2, PyTorch 2.11, torchao 0.18), fastest of 5 runs. Both checkpoints run the same
code, compiled with model.compile(mode="reduce-overhead"):
| one question | 3 questions, one pass | 3.4k-token state | 384 px image | 64 states, packed | |
|---|---|---|---|---|---|
| d1-3B (bf16) | 37 ms | 65 ms | 831 ms | 137 ms | 68 / s |
| d1-3B-w8a8 | 31 ms | 45 ms | 562 ms | 99 ms | 107 / s |
Peak memory of the process (on a Jetson, it includes the GPU's) drops from 12.4 GB to 7.9 GB.
📬 Contact
- Got questions or want to connect? Join our Discord community
- If you are interested in custom solutions with edge deployment, please contact our sales team.
Citation
@article{liquidAI2026opend1,
author = {Liquid AI},
title = {Open d1: Edge decision models for text, vision, and audio},
journal = {Liquid AI Blog},
year = {2026},
note = {https://www.liquid.ai/blog/open-d1},
}
@article{liquidai2025lfm2,
title = {LFM2 Technical Report},
author = {Liquid AI},
journal = {arXiv preprint arXiv:2511.23404},
year = {2025}
}
- Downloads last month
- 36