Clef MXFP4

This is an MXFP4 quantization of Cloudflare/clef, a 27B multimodal decision model post-trained from Qwen3.8-27B. It was quantized with AMD Quark 0.13.

The checkpoint is a standard Quark hf_format export. It loads in vLLM, and in Transformers with amd-quark installed. It is not tied to one GPU family: it runs natively on hardware with MXFP4 matrix units (for example MI350/MI355X) and emulated elsewhere.

BF16 original This repo
Backbone weights on disk 54.7 GB 18.9 GB
Joint schema head BF16 BF16, byte-identical
vLLM weight memory 50.2 GiB 17.0 GiB

Quantization

Format OCP MXFP4: E2M1 elements with one E8M0 scale per 32 values along the input dimension
Weights MXFP4, static, even scale rounding
Activations MXFP4, dynamic per 32-value block
Algorithm AWQ (Quark's qwen3_5 template) on the MLP projections
Calibration 128 samples × 512 tokens from pileval (mit-han-lab/pile-val-backup)
Quantized All language-model linear layers: full attention, Gated DeltaNet projections and MLP, in all 64 layers
Kept in BF16 Vision encoder and merger, lm_head, embeddings, norms, the conv1d layers, and the joint schema head

lm_head stays in BF16 deliberately. Clef's joint schema head reads the output-embedding matrix directly to embed answer options.

The recipe is the same as AMD's own amd/Qwen3.8-27B-Quark-AWQ-MXFP4 for the base model:

python3 quantize_quark.py \
  --model_dir Cloudflare/clef \
  --output_dir ./clef-MXFP4 \
  --quant_scheme mxfp4 \
  --quant_algo awq \
  --num_calib_data 128 \
  --seq_len 512 \
  --model_export hf_format \
  --data_type auto \
  --device cuda \
  --skip_evaluation

quantize_quark.py is the LLM PTQ example from the Quark v0.13 repository. After export, the backbone was resharded into five files with a safetensors index (bit-identical tensors), and the joint head files were copied unchanged. algo_config is set to null in config.json because AWQ is already folded into the weights and vLLM does not parse that field. The exported original is kept as config.json.orig_with_algo_config.

Usage

Clef decisions (Transformers)

Use the original Clef code unchanged; only the repository name changes. Loading needs amd-quark:

pip install amd-quark transformers pillow
import sys

import torch
from huggingface_hub import snapshot_download

path = snapshot_download("EliovpAI/clef-MXFP4")
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone

model, processor = load_release_model(path, device="cuda")

response = systemone(model, processor, {
    "model": "clef",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])

See the Clef model card for the input format, image and video inputs, and batching. In Transformers, Quark simulates the MXFP4 arithmetic. This is the reference path for accuracy, not for speed.

Backbone serving (vLLM)

vLLM loads the quantized backbone and runs native MXFP4 GEMMs where the hardware supports them. The Clef joint head is not part of vLLM; use the Transformers path above for decisions.

vllm serve EliovpAI/clef-MXFP4 --max-model-len 16384

Evaluation

All runs were on one AMD Instinct MI355X (gfx950) with ROCm 7.2.3. BF16 is Cloudflare/clef at revision 2f3de3d. Both models used the same code and inputs.

Clef decisions: release code (Transformers 5.8.1, Quark 0.13)

Task Records BF16 accuracy MXFP4 accuracy Top-1 agreement Mean KL (BF16 ‖ MXFP4)
BANKING77 test, 77-way choice 300 94.0 % 93.3 % 98.0 % 0.022
CLINC150 plus test, 151-way choice incl. out-of-scope 300 96.7 % 96.7 % 98.7 % 0.018
Receipt image, 2 × noul + 3-way choice 3 questions – – 100 % 0.0008

The records are a fixed random sample (seed 0). Each label is a choice option whose description is the humanized label name. These are our own prompts, not the Decision Index protocol, so the absolute numbers are not comparable to Cloudflare's published results. Use the BF16/MXFP4 difference.

The SystemOne example from the Clef model card gives the same answers. technical is chosen at 0.918 (BF16 0.917), urgency peaks at "Today" with 0.920 (BF16 0.862), and outage is 0.841 (BF16 0.895).

Backbone language modelling (vLLM 0.21.0)

BF16 MXFP4
WikiText-2 perplexity (64 × 1,024 tokens) 9.974 10.505 (+5.3 %)
Greedy chat answers (4 prompts) coherent coherent, same content

Files

File Purpose
model-0000X-of-00005.safetensors, model.safetensors.index.json Quantized backbone (MXFP4 language model, BF16 vision encoder)
config.json Model and Quark quantization config
config.json.orig_with_algo_config Config as exported by Quark, including the AWQ settings
joint_head.safetensors, joint_head_config.json, joint_schema_model.py Clef joint schema head and code, unchanged from the original
tokenizer*, chat_template.jinja, processor_config.json, preprocessor_config.json, generation_config.json Tokenizer and processors
LICENSE Apache-2.0, from the original repository
SHA256SUMS Checksums of every file above

License and attribution

Apache-2.0, the same as Cloudflare/clef, which is post-trained from Qwen/Qwen3.8-27B. All credit for the model goes to Cloudflare and the Qwen team. This repository changes only the weight format of the backbone, as described above. It is not affiliated with or endorsed by Cloudflare, Qwen or AMD.

Downloads last month
24
Safetensors
Model size
15B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EliovpAI/clef-MXFP4

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(18)
this model

Dataset used to train EliovpAI/clef-MXFP4