clef-flash-NVFP4

This is an NVFP4 (W4A4) quantization of Cloudflare/clef-flash at revision 17f0b0ad.

  • Quantized: the 96 text MLP projections (gate/up/down_proj).
  • Kept in BF16: the gated-deltanet and full-attention layers, lm_head (the joint head reads its rows), the vision tower and the joint head.
  • Size: 12.1 GB, versus 19.1 GB for BF16 (decimal GB).

Method

  • Tool: llmcompressor 0.13.0 oneshot with QuantizationModifier(scheme="NVFP4"). Weights use FP4 with group size 16 and FP8 scales.
  • Calibration: per-tensor input-activation global scales, taken from 248 records (3.4M tokens, up to 87K tokens long) built from legal, news, science, books, tools, routing and tax sources.
  • Decontamination: the calibration data was checked against the evaluation sets.
  • Stack: torch 2.13 (cu132), transformers 5.14.1, compressed-tensors 0.18.0, on an RTX PRO 6000 (SM120).

Parity vs BF16

The held-out set has 256 records, one question each: LongBench v2 (128, 4 options), banking77 (64, 77 options) and LEDGAR (64, 100 options). Inputs run from 1.7K to 237K tokens. Each model gets one forward pass per record through the joint head.

BF16 NVFP4 (this repo) NVFP4 incl. attention + deltanet (not released)
Accuracy 62.1% 63.7% 60.9%
Δ accuracy, 95% CI (paired bootstrap) +1.6 pt [−1.2, +4.3] −1.2 pt [−5.1, +2.7]
Top-1 agreement with BF16 0.926 [0.887, 0.952] 0.883 [0.838, 0.917]
Agreement where BF16 margin ≥ 0.5 (n=156) 1.000 0.994
Mean KL(BF16 ‖ quant) 0.020 0.048
Size (GB) 19.1 12.1 9.1

Release gates: the accuracy delta's 95% lower bound must be ≥ −2 pt, and agreement on confident decisions must be ≥ 0.98. This model passes both. The variant that also quantizes attention and deltanet fails the accuracy gate, losing mostly at 8K–32K tokens (−6.1 pt).

Every disagreement with BF16 is on a question where BF16's own margin is below 0.5.

These gates were set after the first full report. Under the originally proposed gates (overall agreement ≥ 0.95, every length band ≥ 0.90) this model fails at 0.926 overall and 0.878 on the 8K–32K band. Those gates were dropped because band slices of n=26–53 are too small to gate on.

Per-band, per-cohort and per-margin breakdowns are in metrics/.

Usage

The repo uses the same code as the base model (joint_schema_model.py). Load the backbone decompressed: the joint head calls model.language_model directly, so the packed weights have to be expanded at load time.

import json, torch, joint_schema_model as jsm
from safetensors.torch import load_file
from transformers import Qwen3_5ForConditionalGeneration, CompressedTensorsConfig

d = "<local path to this repo>"
backbone = Qwen3_5ForConditionalGeneration.from_pretrained(
    d, dtype=torch.bfloat16, device_map={"": "cuda:0"},
    quantization_config=CompressedTensorsConfig(run_compressed=False))
head = jsm.JointSchemaHead(**json.load(open(f"{d}/joint_head_config.json")))
head.load_state_dict(load_file(f"{d}/joint_head.safetensors"))
model = jsm.ClefModel(backbone, head.to("cuda:0", torch.bfloat16)).eval()

The metrics above were measured with this path, which simulates W4A4 numerically in eager PyTorch. Native FP4 kernel serving (e.g. vLLM on Blackwell) has not been evaluated.

License

Apache-2.0, inherited from the base model. This is a quantized derivative of Cloudflare/clef-flash. The weights were not otherwise modified.

Downloads last month
22
Safetensors
Model size
9B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Code4me2/clef-flash-NVFP4

Finetuned
Qwen/Qwen3.5-9B
Quantized
(29)
this model