ClefSift 27B

ClefSift is an experimental English text classifier fine-tuned from Cloudflare/Clef. It returns scores for four document-level labels in one forward pass:

Label Training interpretation
human A text with an upstream human-authorship label
ai-assisted An AI copy edit of a provisional human source
mixed Human and AI passages combined in one text
ai A text with an upstream AI-generation label

These are provisional, process-based labels. The model sees writing style, not the author's actual workflow. Its output does not prove authorship. Do not use it as the sole basis for academic, employment, or moderation decisions.

This release uses the selected v2 step-128 checkpoint. The additional rank-64 LoRA weights are merged into the BF16 backbone. The trained joint decision head is stored separately in FP32. This is a decision model, not a chat model.

Model details

Item Value
Immediate base Cloudflare/clef
Base revision 2f3de3dd85f379784083b0814d997ab627200f0c
Retained backbone architecture Qwen3_5ForConditionalGeneration, model type qwen3_5
Parameters Approximately 27 billion, plus the joint decision head
Task English, document-level, four-class text classification
Training toolkit MersivMedia/clef-finetune, version 0.1.0
Training date 2026-10-06 UTC
Selected checkpoint clef-27b-llmtrace-en-fourway-v2, step 128
Full merged release size Approximately 55.25 GB

The retained config and loader use the Qwen3.5 architecture identifiers above. Pin the Clef revision when you reproduce this work. A current upstream model card is not a substitute for the retained config.

The release retains the base vision encoder, but all ClefSift training and evaluation used text. Fine-tuned image, video, and general decision performance are untested.

Files and inference

The merged release needs both the backbone and the joint decision head. A standard Transformers text-classification pipeline does not load the complete Clef decision model. Use the supplied joint_schema_model.py instead. Review that code before you import it.

Files Purpose
model-*.safetensors, model.safetensors.index.json, config.json Merged BF16 backbone
joint_head.safetensors, joint_head_config.json Trained FP32 joint decision head
joint_schema_model.py Unmodified loader and SystemOne request/response functions from the pinned Clef release
Tokenizer, processor, chat template, and generation config Retained base input assets
ClefSift-Q8_0.gguf Q8_0 export for a Clef-compatible llama.cpp server
release.json, weights-manifest.json Portable provenance and SHA-256 hashes without local machine paths
LICENSE, NOTICE Retained Apache-2.0 license and upstream attribution

Reference evaluation used Python 3.11, torch 2.14.1+cu130, transformers 5.19.0, and safetensors on one 96 GB RTX PRO 6000 Blackwell GPU. The full release requires about 55 GB for weights, plus memory for inference. Transformers must provide Qwen3_5ForConditionalGeneration.

The example below uses the exact training question and v2 text normalization. It uses the upstream loader's default BF16 inference, including its BF16 cast of the stored FP32 head. Reference scores below use the training toolkit's FP32-head path, so this example is not an exact reproduction of those scores.

import html
import sys

from huggingface_hub import snapshot_download

path = snapshot_download(
    "spmurrayzzz/clefsift",
    ignore_patterns=["*.gguf"],
)
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone

model, processor = load_release_model(path, device="cuda")
text = "Replace this text with the document to check."
text = " ".join(
    html.unescape(text)
    .replace("\\n", " ")
    .replace("\\r", " ")
    .replace("\\t", " ")
    .split()
)
request = {
    "model": "clefsift",
    "state": text,
    "questions": {
        "classification": {
            "type": "choice",
            "instructions": "Judge only the writing style, not the topic or quality. Human writing tends to show irregular grammar, typos, uneven rhythm, slang, humor, and lived personal detail. AI-generated writing tends to show flawless grammar, uniformly balanced sentences, generic abstraction, hedging, and formulaic transitions.",
            "criteria": {
                "human": "Irregular grammar or typos, colloquial voice, slang or humor, uneven sentence rhythm, concrete personal detail from lived experience",
                "ai-assisted": "Mostly uniform, polished, machine-like prose, but with some personal specifics or human irregularities breaking through",
                "mixed": "Contains some clearly human-written passages mixed with some clearly machine-generated passages",
                "ai": "Flawless uniform prose, balanced clause structures, generic abstract content, hedging language, formulaic transitions, no personal irregularities",
            },
        },
    },
}
response = systemone(model, processor, request, max_length=1536)
print(response["answers"]["classification"])

Use a fixed Hub commit as the download revision for reproducible deployment. Complete training and validation requests fit 1,536 tokens without truncation. Longer documents and changed question wording are outside this evaluation.

The answer contains a winning choice, its confidence, and all four probabilities. These probabilities are not a verified probability of AI use. The ClefSift extension shows UNCERTAIN when the top probability is below 0.35 or its lead over the next option is below 0.10. Those thresholds are display rules, not independently calibrated reliability guarantees.

GGUF inference with llama.cpp

The Q8 export needs a llama.cpp build with Clef's joint decision head and /v1/systemone support. The verified build is 11450 at commit a46709b683aba9274d8ab29f5b42f7e551d1dff6. A generic GGUF chat loader is not an equivalent decision backend.

Download the Q8 file and the original Clef projector:

hf download spmurrayzzz/clefsift ClefSift-Q8_0.gguf --local-dir ./clefsift
hf download ggml-org/Clef-GGUF mmproj-Clef-BF16.gguf --local-dir ./clefsift

Start the server on an unused port:

llama-server --host 127.0.0.1 --port 8000 --alias clefsift \
  --model ./clefsift/ClefSift-Q8_0.gguf \
  --mmproj ./clefsift/mmproj-Clef-BF16.gguf \
  -b 32768 -ub 32768

The batch settings are required in the verified build. Its default physical batch size of 512 rejects requests above that size. The retained projector is not a fine-tuned vision model. Our Mac serving setup uses roughly 30 GB of RAM including that projector.

Send the same state and questions from the Python example to POST http://127.0.0.1:8000/v1/systemone, with model set to clefsift. The response uses the same four-label SystemOne contract.

Training data and provenance

The source corpus is iitolstykh/LLMTrace_detection, revision 332053804f8798177ff2ecb1148e1839ed429ba3. Only English examples entered these runs. Source domains include articles, questions, reviews, and short-form text.

The v2 training set has 298 source groups and 1,192 examples, with one example per class in each group. Official source splits remain intact across all variants. Known duplicate groups and malformed candidates were excluded. The audit does not establish author-level independence.

AI-assisted examples are whole-text copy edits from gpt-5.6-luna through Pi's openai-codex provider. Training mixed examples include 124 Luna 5.6 insertions and 174 upstream Gemini 2.5 Flash examples. Structural checks preserve source substrings for insertions, but they do not verify meaning in whole-text edits. Manual samples found some changes in meaning or viewpoint.

Matched validation uses the same editing model version as training. Held-out validation uses gpt-6-luna for all edits and insertions. This tests another version within the same family, not an unrelated family. Both panels share their 47 parent groups and human/fully-AI controls.

For v2, every class decodes HTML entities and literal escaped whitespace, then collapses whitespace. The original v1 regression requests remain unchanged. No synthetic bootstrap rows or Pangram outputs supplied training labels.

Authorship and source rights remain unresolved. The upstream dataset declares Apache-2.0, but that declaration does not establish rights in every source text. The exploratory runs used provisional labels without independent source clearance. This model release does not distribute the training texts or generation transcripts.

Training procedure

v2 starts from the pinned Clef reference, not from the earlier v1 adapter. It trains the LoRA adapters and the joint head together. The vision tower and language-model output head are excluded from adapters.

Setting Value
Backbone / trainable head dtype BF16 / FP32
LoRA rank / alpha / dropout 64 / 128 / 0.05
Seed / optimization steps 0 / 128
Batch size / gradient accumulation 1 / 4
Example exposures 512, less than one full epoch
Backbone / head learning rate 1e-4 / 5e-5
Warmup 12 steps
Loss Cross-entropy with 0.05 label smoothing, plus four-option Brier loss with weight 1.0
Gradient checkpointing / schema augmentation Enabled / disabled

Checkpoint selection uses matched fixed-four macro F1, then lower Brier and earlier step for ties. Step 128 was selected before held-out scoring. Training took about 861 seconds. The run, including evaluations, used about 24.11 GPU-minutes on one RTX PRO 6000 Blackwell.

Validation results

These scores describe small internal validation panels with provisional labels. They are not a final-test benchmark or an estimate of LinkedIn accuracy. Reference scores use the selected unmerged adapter with an FP32 head. Q8 scores use its merged, converted export through llama.cpp.

Panel Requests Reference accuracy Q8 accuracy Reference macro F1 Q8 macro F1 Reference Brier Q8 Brier
Matched v2 188 89.9% (169/188) 89.4% (168/188) 0.8991 0.8936 0.1829 0.2737
Held-out Luna version 188 86.2% (162/188) 86.2% (162/188) 0.8602 0.8602 0.2421 0.3115
Original v1 validation 150 85.3% (128/150) 85.3% (128/150) 0.8766 0.8766 0.2596 0.3374

The first two rows use four-class macro F1. The last uses the three target classes in original v1 validation, with all four prediction options retained. Brier is the sum of squared errors over all four probabilities. Lower is better. Accuracy and F1 use the winning class before the display uncertainty rule.

Human false positives are 5/47 in each v2 panel and 5/50 in original validation, for both the reference and Q8. A false positive means any non-human winning class on a provisional human target. These rates do not establish a low false-positive rate for real users.

AI-assisted recall is 45/47 matched and 33/47 held-out, for both versions. Original mixed recall fell from v1's 45/50 to v2's 39/50 (90% to 78%). That regression remains unresolved.

Q8 export comparison

The retained ClefSift-Q8_0.gguf export is 28,732,214,816 bytes (26.76 GiB). Its SHA-256 is 1fbdedec7989395973e5364a47dcd4a01a5f419a425b26ba476faed5556cb7fa. Conversion used llama.cpp commit a46709b683aba9274d8ab29f5b42f7e551d1dff6, build 11450, through an F16 intermediate.

Q8 preserves 525/526 reference class predictions across the saved requests. One matched AI prediction becomes mixed. Human false positives and AI-assisted recall do not change. Four display verdicts become uncertain at unchanged thresholds.

Q8 probabilities are softer and Brier scores are worse. This comparison includes merging, conversion, quantization, and backend differences. It does not isolate quantization alone. No intermediate F16 inference evaluation was completed. The validation panels share controls and groups, so the 526 requests are not 526 independent observations.

Limitations and intended use

Use ClefSift for exploratory text review and further research with human review. Evaluate it on your own domain before you rely on its scores.

  • No independent professional-post evaluation with verified authorship and rights has run.
  • No final-test scoring, unrelated-generator-family evaluation, or multiple-seed experiment has run.
  • Non-English text, long documents, paraphrase attacks, and fine-tuned vision behavior are untested.
  • The style rubric can penalize polished human prose or miss deliberately irregular AI prose.
  • Training labels and copy-edit fidelity remain provisional.
  • Confidence calibration changes across inference formats and lacks independent validation.
  • The held-out AI-assisted recall gap and original mixed-recall regression remain open.

The final test was used only for structural and duplicate auditing. Its text did not enter generation, prompt development, training, checkpoint selection, hard-negative mining, or model scoring.

License and acknowledgments

The release uses Apache-2.0 and retains the base LICENSE file. Base weights and the unmodified loader come from Cloudflare's Clef release. The retained license includes Alibaba Cloud attribution for the Qwen backbone. Training used the Apache-2.0 clef-finetune toolkit. ClefSift is an independent fine-tune and is not affiliated with Cloudflare or Alibaba.

Modified artifacts are the LoRA-merged backbone weights and trained joint head. The tokenizer, processor assets, and joint_schema_model.py retain their upstream contents. The model license does not resolve the source-rights uncertainty stated above.

Downloads last month
52
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for spmurrayzzz/clefsift

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Finetuned
(7)
this model

Dataset used to train spmurrayzzz/clefsift