Instructions to use EliovpAI/clef-MXFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use EliovpAI/clef-MXFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="EliovpAI/clef-MXFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("EliovpAI/clef-MXFP4") model = AutoModelForMultimodalLM.from_pretrained("EliovpAI/clef-MXFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use EliovpAI/clef-MXFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "EliovpAI/clef-MXFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EliovpAI/clef-MXFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/EliovpAI/clef-MXFP4
- SGLang
How to use EliovpAI/clef-MXFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "EliovpAI/clef-MXFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EliovpAI/clef-MXFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "EliovpAI/clef-MXFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "EliovpAI/clef-MXFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use EliovpAI/clef-MXFP4 with Docker Model Runner:
docker model run hf.co/EliovpAI/clef-MXFP4
Clef MXFP4
This is an MXFP4 quantization of Cloudflare/clef, a 27B multimodal decision model post-trained from Qwen3.8-27B. It was quantized with AMD Quark 0.13.
The checkpoint is a standard Quark hf_format export. It loads in vLLM, and in Transformers with amd-quark
installed. It is not tied to one GPU family: it runs natively on hardware with MXFP4 matrix units (for example
MI350/MI355X) and emulated elsewhere.
| BF16 original | This repo | |
|---|---|---|
| Backbone weights on disk | 54.7 GB | 18.9 GB |
| Joint schema head | BF16 | BF16, byte-identical |
| vLLM weight memory | 50.2 GiB | 17.0 GiB |
Quantization
| Format | OCP MXFP4: E2M1 elements with one E8M0 scale per 32 values along the input dimension |
| Weights | MXFP4, static, even scale rounding |
| Activations | MXFP4, dynamic per 32-value block |
| Algorithm | AWQ (Quark's qwen3_5 template) on the MLP projections |
| Calibration | 128 samples × 512 tokens from pileval (mit-han-lab/pile-val-backup) |
| Quantized | All language-model linear layers: full attention, Gated DeltaNet projections and MLP, in all 64 layers |
| Kept in BF16 | Vision encoder and merger, lm_head, embeddings, norms, the conv1d layers, and the joint schema head |
lm_head stays in BF16 deliberately. Clef's joint schema head reads the output-embedding matrix directly to embed answer
options.
The recipe is the same as AMD's own amd/Qwen3.8-27B-Quark-AWQ-MXFP4 for the base model:
python3 quantize_quark.py \
--model_dir Cloudflare/clef \
--output_dir ./clef-MXFP4 \
--quant_scheme mxfp4 \
--quant_algo awq \
--num_calib_data 128 \
--seq_len 512 \
--model_export hf_format \
--data_type auto \
--device cuda \
--skip_evaluation
quantize_quark.py is the LLM PTQ example from the Quark v0.13 repository.
After export, the backbone was resharded into five files with a safetensors index (bit-identical tensors), and the
joint head files were copied unchanged. algo_config is set to null in config.json because AWQ is already folded
into the weights and vLLM does not parse that field. The exported original is kept as config.json.orig_with_algo_config.
Usage
Clef decisions (Transformers)
Use the original Clef code unchanged; only the repository name changes. Loading needs amd-quark:
pip install amd-quark transformers pillow
import sys
import torch
from huggingface_hub import snapshot_download
path = snapshot_download("EliovpAI/clef-MXFP4")
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone
model, processor = load_release_model(path, device="cuda")
response = systemone(model, processor, {
"model": "clef",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])
See the Clef model card for the input format, image and video inputs, and batching. In Transformers, Quark simulates the MXFP4 arithmetic. This is the reference path for accuracy, not for speed.
Backbone serving (vLLM)
vLLM loads the quantized backbone and runs native MXFP4 GEMMs where the hardware supports them. The Clef joint head is not part of vLLM; use the Transformers path above for decisions.
vllm serve EliovpAI/clef-MXFP4 --max-model-len 16384
Evaluation
All runs were on one AMD Instinct MI355X (gfx950) with ROCm 7.2.3. BF16 is Cloudflare/clef at revision 2f3de3d.
Both models used the same code and inputs.
Clef decisions: release code (Transformers 5.8.1, Quark 0.13)
| Task | Records | BF16 accuracy | MXFP4 accuracy | Top-1 agreement | Mean KL (BF16 ‖ MXFP4) |
|---|---|---|---|---|---|
| BANKING77 test, 77-way choice | 300 | 94.0 % | 93.3 % | 98.0 % | 0.022 |
| CLINC150 plus test, 151-way choice incl. out-of-scope | 300 | 96.7 % | 96.7 % | 98.7 % | 0.018 |
| Receipt image, 2 × noul + 3-way choice | 3 questions | – | – | 100 % | 0.0008 |
The records are a fixed random sample (seed 0). Each label is a choice option whose description is the humanized label name. These are our own prompts, not the Decision Index protocol, so the absolute numbers are not comparable to Cloudflare's published results. Use the BF16/MXFP4 difference.
The SystemOne example from the Clef model card gives the same answers. technical is chosen at 0.918
(BF16 0.917), urgency peaks at "Today" with 0.920 (BF16 0.862), and outage is 0.841 (BF16 0.895).
Backbone language modelling (vLLM 0.21.0)
| BF16 | MXFP4 | |
|---|---|---|
| WikiText-2 perplexity (64 × 1,024 tokens) | 9.974 | 10.505 (+5.3 %) |
| Greedy chat answers (4 prompts) | coherent | coherent, same content |
Files
| File | Purpose |
|---|---|
model-0000X-of-00005.safetensors, model.safetensors.index.json |
Quantized backbone (MXFP4 language model, BF16 vision encoder) |
config.json |
Model and Quark quantization config |
config.json.orig_with_algo_config |
Config as exported by Quark, including the AWQ settings |
joint_head.safetensors, joint_head_config.json, joint_schema_model.py |
Clef joint schema head and code, unchanged from the original |
tokenizer*, chat_template.jinja, processor_config.json, preprocessor_config.json, generation_config.json |
Tokenizer and processors |
LICENSE |
Apache-2.0, from the original repository |
SHA256SUMS |
Checksums of every file above |
License and attribution
Apache-2.0, the same as Cloudflare/clef, which is post-trained from Qwen/Qwen3.8-27B. All credit for the model goes to Cloudflare and the Qwen team. This repository changes only the weight format of the backbone, as described above. It is not affiliated with or endorsed by Cloudflare, Qwen or AMD.
- Downloads last month
- 24