Instructions to use jlancaster/clef-27b-nf4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jlancaster/clef-27b-nf4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="jlancaster/clef-27b-nf4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("jlancaster/clef-27b-nf4") model = AutoModelForMultimodalLM.from_pretrained("jlancaster/clef-27b-nf4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jlancaster/clef-27b-nf4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jlancaster/clef-27b-nf4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jlancaster/clef-27b-nf4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/jlancaster/clef-27b-nf4
- SGLang
How to use jlancaster/clef-27b-nf4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jlancaster/clef-27b-nf4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jlancaster/clef-27b-nf4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jlancaster/clef-27b-nf4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jlancaster/clef-27b-nf4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use jlancaster/clef-27b-nf4 with Docker Model Runner:
docker model run hf.co/jlancaster/clef-27b-nf4
Clef 27B, NF4 (bitsandbytes)
This is Cloudflare's Clef 27B decision model with its backbone quantised to 4-bit NF4, so it fits on a single 24 GB GPU. It's about 18 GB on disk instead of 55 GB, and it loads with Cloudflare's own load_release_model, unchanged.
I made it to run Clef on one RTX 4090 for a video-classification pipeline. Before sharing it, I checked it against the full bf16 model on that task. The results are below, along with what I didn't test.
What changed, and what didn't
Changed: the backbone's linear layers are quantised to NF4 with this config.
config.jsonand the weight shards are the only files that differ from Cloudflare's.BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16, llm_int8_skip_modules=["lm_head"])Unchanged: the joint schema head (
joint_head.safetensors,joint_head_config.json), the loader (joint_schema_model.py), the tokenizer and processor files, and the license are Cloudflare's, copied byte for byte.lm_headstays in bf16 on purpose. The joint schema head reads the output-embedding weights directly, so quantisinglm_headwould hand it packed 4-bit data. If you quantise Clef yourself, skiplm_head.
Usage
Tested with torch 2.11.0 (CUDA 13.0), transformers 5.10.2, bitsandbytes 0.50.2, torchvision 0.26.0 and accelerate, on an RTX 4090.
For about a third more speed, install the fast kernels for Qwen 3.5's linear-attention layers too. Without them transformers falls back to plain PyTorch, which works but is slower.
pip install flash-linear-attention # pure Triton, nothing to compile
# causal-conv1d compiles a CUDA extension. There's no prebuilt wheel for torch 2.11 + CUDA 13,
# so it builds from source and needs the CUDA 13 toolkit's nvcc (12.x won't do: torch refuses
# to build across a major-version mismatch).
CUDA_HOME=/usr/local/cuda CAUSAL_CONV1D_FORCE_BUILD=TRUE pip install --no-build-isolation causal-conv1d
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("jlancaster/clef-27b-nf4")
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone
# The quantisation settings come from config.json; no extra arguments needed.
model, processor = load_release_model(path, device="cuda")
answer = systemone(model, processor, {
"model": "clef",
"state": {"title": "...", "transcript_excerpt": "..."},
"questions": {"video_type": {
"type": "choice",
"instructions": "Classify this video.",
"criteria": {"sermon": "A complete sermon.", "worship_music": "Music is the main content."},
}},
})
print(answer["answers"]["video_type"]) # choice, confidence, probabilities
The request and response shapes are Cloudflare's System One format; see their model card for noul, choice and score questions.
How close it is to the bf16 model
I ran the same 152 requests through the full bf16 model and through this quantisation, and compared their answers request by request. Each request was one 13-option choice question about a church video (title, description, transcript excerpts and metadata), about 4–5k tokens.
| Compared with bf16 Clef 27B | Result |
|---|---|
| Same top-level decision (the 7 options we keep vs the 6 we hide) | 151 of 152 |
| Same top choice | 149 of 152 |
| Difference in the summed "hide" probability | mean 0.011, max 0.092 |
The one flipped decision was a near tie that moved from 0.51 to 0.49. Against our own blind labels, the two scored within a point of each other: 97% vs 98% keep/hide agreement and Brier 0.034 vs 0.033.
Resources
GPU memory: 18.2 GB after loading. It ran our 4–5k-token requests one at a time within 24 GB. I didn't test batching or longer inputs.
Speed on an RTX 4090, the same 152 requests (4–5k tokens each), one at a time:
Kernels installed Median 90th percentile flash-linear-attention0.5.2 +causal-conv1d1.7.01.70 s 1.79 s flash-linear-attentiononly1.78 s 1.88 s Neither (PyTorch fallback) 2.51 s 2.66 s The kernels don't change the decisions: with both installed, all 152 kept the same keep/hide call as the fallback run, and the summed "hide" probability moved by at most 0.015.
What I didn't test
- Other tasks: these numbers come from one classification task and 152 requests. They're evidence for that task, not a general quality claim.
- Images and video inputs: text-only requests only.
- Other GPUs and library versions: only the setup listed above.
scoreandnoulquestions: onlychoice.
License
Apache 2.0, the same as Clef. This is a derivative of Cloudflare/clef; the only change is the quantisation described above.
- Downloads last month
- -