Clewen-Flash

Clewen-Flash combines Qwen3.5-9B text and image generation with Cloudflare Clef-Flash structured decisions, sharing one backbone. The decision mode uses recovered switchable adapters and the original Clef joint schema head. No additional training was performed. Recovered adapters approximate the learned changes; they are not the original LoRA factors.

Usage

Both modes support text and images. Decision inputs can also contain JSON. Modes are selected explicitly:

  • Qwen: text generation and image understanding.
  • Clef: structured answers with probabilities for choice, noul (yes/no), and score questions.

Install on Linux with Python 3.11 or 3.12 and a CUDA GPU:

pip install torch==2.10.0 torchvision==0.25.0 --index-url https://download.pytorch.org/whl/cu128
pip install "clewen[cuda] @ git+https://github.com/salyamq/clewer.git"
from PIL import Image
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "salyamq/clewen-flash",  # or "salyamq/clewen"
    trust_remote_code=True,
    device_map="cuda",
    dtype="bfloat16",
)

# Text generation
reply = model.text(
    [{"role": "user", "content": "Explain gradient descent briefly."}],
    max_new_tokens=128,
)
print(reply["answer"])

# Image understanding
image = Image.open("example.jpg").convert("RGB")
reply = model.text(
    [{
        "role": "user",
        "content": [
            {"type": "image", "image": image},
            {"type": "text", "text": "Describe this image briefly."},
        ],
    }],
    max_new_tokens=128,
)
print(reply["answer"])

# Structured decision from an image
result = model.decision({
    "state": "Inspect the attached image.",
    "images": [image],
    "questions": {
        "cat_visible": {
            "type": "noul",
            "instructions": "Is a cat visible?",
        },
    },
})
print(result["answers"])

For text or JSON decisions, omit images and put your information in state. Both modes use the same loaded model instance.

Results

Decision-mode scores from complete evaluation runs. Scores are percentages; higher is better. Workflow exact-action scores require the complete action set to match the reference labels.

Benchmark / metric Clewen Clewen-Flash Clef Clef-Flash Jev DiffusionGemma Jev Kev 9B Laya
BFCL — case exact accuracy 98.41 98.88 98.5 98.8 95.8 96.5 94.5 38.1
API-Bank — accuracy 91.73 93.11 91.9 93.1 88.2 83.7 56.3 11.5
BANKING77 — macro-F1 94.08 90.80 94.2 90.9 79.7 74.3 84.8 14.3
RAGTruth — hallucination F1 79.30 35.74 79.4 35.6 76.5 70.4 46.2 48.8
When2Call MCQ — accuracy 72.48 65.36 72.4 65.6 81.0 75.4 49.6 11.9
Invoice processing — exact actions 64.67 57.11 64.7 57.1 61.8 — — —
Invoice processing — primary action 86.44 74.22 86.2 73.3 83.1 — — —
Customer service — exact actions 76.31 76.96 76.3 77.0 76.0 — — —
Security incidents — exact actions 63.33 61.67 62.9 61.7 61.7 — — —
Agent trace observability — primary action 68.47 69.82 68.5 69.8 71.6 — — —
Decision latency — median, ms ↓ 142.30 78.06 209.3 38.8 524.1 84.4 51.4 5.8
Decision latency — p95, ms ↓ 274.15 111.26 238.6 122.4 536.0 211.2 187.9 222.5

Latency was measured on one H200 in BF16 across the five Decision Index benchmarks. It includes input encoding and answer decoding, excluding model loading and warmup. These evaluations measure structured decisions, not text-generation quality.

License

Apache-2.0. Qwen components are credited to the Qwen authors; Clef components are credited to Cloudflare. See LICENSE, LICENSE-QWEN, LICENSE-CLEF, and NOTICE for attribution and license details.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for salyamq/clewen-flash

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Finetuned
(5)
this model

Collection including salyamq/clewen-flash