Instructions to use salyamq/clewen-flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use salyamq/clewen-flash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="salyamq/clewen-flash", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("salyamq/clewen-flash", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use salyamq/clewen-flash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "salyamq/clewen-flash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "salyamq/clewen-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/salyamq/clewen-flash
- SGLang
How to use salyamq/clewen-flash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "salyamq/clewen-flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "salyamq/clewen-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "salyamq/clewen-flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "salyamq/clewen-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use salyamq/clewen-flash with Docker Model Runner:
docker model run hf.co/salyamq/clewen-flash
Clewen-Flash
Clewen-Flash combines Qwen3.5-9B text and image generation with Cloudflare Clef-Flash structured decisions, sharing one backbone. The decision mode uses recovered switchable adapters and the original Clef joint schema head. No additional training was performed. Recovered adapters approximate the learned changes; they are not the original LoRA factors.
Usage
Both modes support text and images. Decision inputs can also contain JSON. Modes are selected explicitly:
- Qwen: text generation and image understanding.
- Clef: structured answers with probabilities for
choice,noul(yes/no), andscorequestions.
Install on Linux with Python 3.11 or 3.12 and a CUDA GPU:
pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu128
pip install "clewen[cuda] @ git+https://github.com/salyamq/clewer.git"
from PIL import Image
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
"salyamq/clewen-flash", # or "salyamq/clewen"
trust_remote_code=True,
device_map="cuda",
dtype="bfloat16",
)
# Text generation
reply = model.text(
[{"role": "user", "content": "Explain gradient descent briefly."}],
max_new_tokens=128,
)
print(reply["answer"])
# Image understanding
image = Image.open("example.jpg").convert("RGB")
reply = model.text(
[{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": "Describe this image briefly."},
],
}],
max_new_tokens=128,
)
print(reply["answer"])
# Structured decision from an image
result = model.decision({
"state": "Inspect the attached image.",
"images": [image],
"questions": {
"cat_visible": {
"type": "noul",
"instructions": "Is a cat visible?",
},
},
})
print(result["answers"])
For text or JSON decisions, omit images and put your information in state. Both modes use the same loaded model instance.
Results
Decision-mode scores from complete evaluation runs. Scores are percentages; higher is better. Workflow exact-action scores require the complete action set to match the reference labels.
| Benchmark / metric | Clewen | Clewen-Flash | Clef | Clef-Flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|---|---|---|---|---|---|---|---|
| BFCL — case exact accuracy | 98.41 | 98.88 | 98.5 | 98.8 | 95.8 | 96.5 | 94.5 | 38.1 |
| API-Bank — accuracy | 91.73 | 93.11 | 91.9 | 93.1 | 88.2 | 83.7 | 56.3 | 11.5 |
| BANKING77 — macro-F1 | 94.08 | 90.80 | 94.2 | 90.9 | 79.7 | 74.3 | 84.8 | 14.3 |
| RAGTruth — hallucination F1 | 79.30 | 35.74 | 79.4 | 35.6 | 76.5 | 70.4 | 46.2 | 48.8 |
| When2Call MCQ — accuracy | 72.48 | 65.36 | 72.4 | 65.6 | 81.0 | 75.4 | 49.6 | 11.9 |
| Invoice processing — exact actions | 64.67 | 57.11 | 64.7 | 57.1 | 61.8 | — | — | — |
| Invoice processing — primary action | 86.44 | 74.22 | 86.2 | 73.3 | 83.1 | — | — | — |
| Customer service — exact actions | 76.31 | 76.96 | 76.3 | 77.0 | 76.0 | — | — | — |
| Security incidents — exact actions | 63.33 | 61.67 | 62.9 | 61.7 | 61.7 | — | — | — |
| Agent trace observability — primary action | 68.47 | 69.82 | 68.5 | 69.8 | 71.6 | — | — | — |
| Decision latency — median, ms ↓ | 142.30 | 78.06 | 209.3 | 38.8 | 524.1 | 84.4 | 51.4 | 5.8 |
| Decision latency — p95, ms ↓ | 274.15 | 111.26 | 238.6 | 122.4 | 536.0 | 211.2 | 187.9 | 222.5 |
Latency was measured on one H200 in BF16 across the five Decision Index benchmarks. It includes input encoding and answer decoding, excluding model loading and warmup. These evaluations measure structured decisions, not text-generation quality.
License
Apache-2.0. Qwen components are credited to the Qwen authors; Clef components are credited to Cloudflare. See LICENSE, LICENSE-QWEN, LICENSE-CLEF, and NOTICE for attribution and license details.
- Downloads last month
- -