You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Clef 27B, NF4 (bitsandbytes)

This is Cloudflare's Clef 27B decision model with its backbone quantised to 4-bit NF4, so it fits on a single 24 GB GPU. It's about 18 GB on disk instead of 55 GB, and it loads with Cloudflare's own load_release_model, unchanged.

I made it to run Clef on one RTX 4090 for a video-classification pipeline. Before sharing it, I checked it against the full bf16 model on that task. The results are below, along with what I didn't test.

What changed, and what didn't

  • Changed: the backbone's linear layers are quantised to NF4 with this config. config.json and the weight shards are the only files that differ from Cloudflare's.

    BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_use_double_quant=True,
                       bnb_4bit_compute_dtype=torch.bfloat16, llm_int8_skip_modules=["lm_head"])
    
  • Unchanged: the joint schema head (joint_head.safetensors, joint_head_config.json), the loader (joint_schema_model.py), the tokenizer and processor files, and the license are Cloudflare's, copied byte for byte.

  • lm_head stays in bf16 on purpose. The joint schema head reads the output-embedding weights directly, so quantising lm_head would hand it packed 4-bit data. If you quantise Clef yourself, skip lm_head.

Usage

Tested with torch 2.11.0 (CUDA 13.0), transformers 5.10.2, bitsandbytes 0.50.2, torchvision 0.26.0 and accelerate, on an RTX 4090.

For about a third more speed, install the fast kernels for Qwen 3.5's linear-attention layers too. Without them transformers falls back to plain PyTorch, which works but is slower.

pip install flash-linear-attention   # pure Triton, nothing to compile
# causal-conv1d compiles a CUDA extension. There's no prebuilt wheel for torch 2.11 + CUDA 13,
# so it builds from source and needs the CUDA 13 toolkit's nvcc (12.x won't do: torch refuses
# to build across a major-version mismatch).
CUDA_HOME=/usr/local/cuda CAUSAL_CONV1D_FORCE_BUILD=TRUE pip install --no-build-isolation causal-conv1d
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("jlancaster/clef-27b-nf4")
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone

# The quantisation settings come from config.json; no extra arguments needed.
model, processor = load_release_model(path, device="cuda")

answer = systemone(model, processor, {
    "model": "clef",
    "state": {"title": "...", "transcript_excerpt": "..."},
    "questions": {"video_type": {
        "type": "choice",
        "instructions": "Classify this video.",
        "criteria": {"sermon": "A complete sermon.", "worship_music": "Music is the main content."},
    }},
})
print(answer["answers"]["video_type"])  # choice, confidence, probabilities

The request and response shapes are Cloudflare's System One format; see their model card for noul, choice and score questions.

How close it is to the bf16 model

I ran the same 152 requests through the full bf16 model and through this quantisation, and compared their answers request by request. Each request was one 13-option choice question about a church video (title, description, transcript excerpts and metadata), about 4–5k tokens.

Compared with bf16 Clef 27B Result
Same top-level decision (the 7 options we keep vs the 6 we hide) 151 of 152
Same top choice 149 of 152
Difference in the summed "hide" probability mean 0.011, max 0.092

The one flipped decision was a near tie that moved from 0.51 to 0.49. Against our own blind labels, the two scored within a point of each other: 97% vs 98% keep/hide agreement and Brier 0.034 vs 0.033.

Resources

  • GPU memory: 18.2 GB after loading. It ran our 4–5k-token requests one at a time within 24 GB. I didn't test batching or longer inputs.

  • Speed on an RTX 4090, the same 152 requests (4–5k tokens each), one at a time:

    Kernels installed Median 90th percentile
    flash-linear-attention 0.5.2 + causal-conv1d 1.7.0 1.70 s 1.79 s
    flash-linear-attention only 1.78 s 1.88 s
    Neither (PyTorch fallback) 2.51 s 2.66 s

    The kernels don't change the decisions: with both installed, all 152 kept the same keep/hide call as the fallback run, and the summed "hide" probability moved by at most 0.015.

What I didn't test

  • Other tasks: these numbers come from one classification task and 152 requests. They're evidence for that task, not a general quality claim.
  • Images and video inputs: text-only requests only.
  • Other GPUs and library versions: only the setup listed above.
  • score and noul questions: only choice.

License

Apache 2.0, the same as Clef. This is a derivative of Cloudflare/clef; the only change is the quantisation described above.

Downloads last month
-
Safetensors
Model size
27B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jlancaster/clef-27b-nf4

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(23)
this model