Clef-Flash-W4A16-AutoRound-LLM-Compressor

Quantized W4A16 (4-bit weights, 16-bit activations) release of Cloudflare/clef-flash using Intel AutoRound.

Clef-Flash is a 9B multimodal decision model post-trained from Qwen3.5 that evaluates structured schemas (text, JSON, image, video) and returns calibrated probability distributions across typed questions in a single forward pass without autoregressive token generation.


Quantization Details

  • Method: Intel AutoRound W4A16 (sym=True, group_size=32)
  • Vision Tower: Preserved in native BF16 (ensures zero OCR/visual degradation)
  • Joint Schema Head (joint_head.safetensors): Preserved in native BF16
  • Linear Convolutions: Preserved in native BF16 to prevent drift in Qwen3.5 linear attention layers
  • Accuracy: Zero loss observed across text, JSON state, and multimodal test suites compared to the base checkpoint.

Quickstart

Installation

pip install torch transformers huggingface_hub pillow

Usage (systemone API)

import sys
from huggingface_hub import snapshot_download

# Download repo and load custom joint schema model
repo_id = "Vishva007/clef-flash-W4A16-AutoRound"  # or auto-gptq / llm-compressor variant
model_path = snapshot_download(repo_id)
sys.path.insert(0, model_path)

from joint_schema_model import load_release_model, systemone

model, processor = load_release_model(model_path, device="cuda")

# 1. Text / JSON Decision Example
response = systemone(model, processor, {
    "model": "clef-flash",
    "state": "Prod database latency spiked to 4,000ms. Checkout failing with 504 Gateway Timeouts.",
    "questions": {
        "severity": {
            "type": "choice",
            "instructions": "Determine incident severity level",
            "criteria": {
                "SEV_1": "Critical revenue outage",
                "SEV_2": "Major feature degradation",
                "SEV_3": "Minor issue"
            }
        },
        "urgency": {
            "type": "score",
            "instructions": "Urgency rating",
            "criteria": ["Low", "Medium", "Immediate page"]
        },
        "rollback": {
            "type": "noul",
            "instructions": "Should a rollback be initiated?"
        }
    }
})

print(response["answers"])

Multimodal (Image) Evaluation

from PIL import Image

response = systemone(model, processor, {
    "model": "clef-flash",
    "state": "Review the uploaded invoice receipt.",
    "images": [Image.open("receipt.png")],
    "questions": {
        "legible": {"type": "noul", "instructions": "Is the receipt text clear and legible?"},
        "amount_exceeds_1000": {"type": "noul", "instructions": "Is total > $1000 USD?"}
    }
})

print(response["answers"])

Formats Available

Repository Format Engine Target
Vishva007/clef-flash-W4A16-AutoRound AutoRound / AutoGPTQ Transformers / Native Python
Vishva007/clef-flash-W4A16-AutoRound-GPTQ AutoGPTQ Standard Transformers / ExLlama / AutoGPTQ
Vishva007/clef-flash-W4A16-AutoRound-LLM-Compressor Compressed-Tensors vLLM / SGLang

🚀 Deploy on RunPod

One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.

🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.

PyTorch 2.14

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.14 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.14-runpod d7lxsa4w9m Deploy to RunPod
PyTorch 2.14 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.14-runpod yk0y6j6rpg Deploy to RunPod
PyTorch 2.14 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.14-runpod gsp4gwx0nw Deploy to RunPod

PyTorch 2.13

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.13 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.13-runpod gmlupxnxfk Deploy to RunPod
PyTorch 2.13 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.13-runpod y3j8xvk4f4 Deploy to RunPod
PyTorch 2.13 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.13-runpod vigpissn5w Deploy to RunPod

PyTorch 2.12

Template CUDA Version Docker Image Template ID Deploy
PyTorch 2.12 (CUDA 12.6) 12.6 vishva123/cuda-12.6-pytorch-2.12-runpod ctmz86zmf0 Deploy to RunPod
PyTorch 2.12 (CUDA 13.0) 13.0 vishva123/cuda-13.0-pytorch-2.12-runpod qjko5yiwzi Deploy to RunPod
PyTorch 2.12 (CUDA 13.2) 13.2 vishva123/cuda-13.2-pytorch-2.12-runpod ifg6xmye0f Deploy to RunPod

Acknowledgements

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vishva007/clef-flash-W4A16-AutoRound-LLM-Compressor

Finetuned
Qwen/Qwen3.5-9B
Quantized
(26)
this model

Collection including Vishva007/clef-flash-W4A16-AutoRound-LLM-Compressor