clef-p300

A tt-model v6 thin bundle that serves Cloudflare/clef (revision 2f3de3dd85f379784083b0814d997ab627200f0c) on 2 Tenstorrent blackhole chips (mesh P150x2). The tt-orchard harness staged it from a run of that model.

License

The weights this bundle points to are licensed Apache 2.0 (https://www.apache.org/licenses/LICENSE-2.0). The bundle ships no weights. tt-model downloads them from Cloudflare/clef with your own Hugging Face account.

What it runs

  • Weights: Cloudflare/clef at revision 2f3de3dd85f379784083b0814d997ab627200f0c.
  • Model code: the Tenstorrent implementation of Qwen/Qwen3.8-27B, taken from the bundle qwen3.8-27b-dflash2-p300. Cloudflare/clef has the same text architecture, so only the weights differ.
  • At each start the launcher builds model-dir/ from the configuration files of Qwen/Qwen3.8-27B (shipped in base_config/) and the tokenizer and weights of Cloudflare/clef. It sets MODEL_WEIGHTS_DIR and HF_MODEL to that directory, so the chips load the weights of Cloudflare/clef.
  • Context length 262144 tokens; up to 4 sequences at a time.
  • Speculative decoding is off. The drafter of Qwen/Qwen3.8-27B needs the MTP tensors of the model, and Cloudflare/clef has none.
  • Sampling runs on the host, because on-device sampling is not used with the drafter off on this mesh.

Not served

Cloudflare/clef ships files this bundle does not load: joint_head.safetensors. The bundle serves the language-model backbone only, so the sidecar head is not served. The run checked the head on the host against the CPU reference; that check does not make it part of this bundle.

Intended use

Text generation with Cloudflare/clef through the OpenAI-compatible server that tt-model serve starts. Out of scope: the sidecar head, which this bundle does not serve.

Expected performance

Every number is labelled. A measured number names its evidence files; the first file holds the value. The files are in this repository under evidence/, at the paths shown, with the run's directory, the home directory and the host name replaced by <RUN_DIR>, <HOME> and <HOST>. TODO means not measured.

Number Value Label Evidence
top1 agreement with the CPU reference, stage 2 (2 chips, bundle qwen3.8-27b-dflash2-p300) 0.96875 fraction measured stages/2/result.json, stages/2/evidence/swap-check.json, stages/2/evidence/server.log, stages/2/evidence/sidecar-parity.json
server ready after start, stage 2 (empty tensor cache) 236.1 s measured stages/2/result.json, stages/2/evidence/swap-check.json, stages/2/evidence/server.log, stages/2/evidence/sidecar-parity.json
top1 agreement with the CPU reference, this package (stage 7) 0.96875 fraction measured stages/7/verify/evidence/verify.json, stages/7/verify/evidence/server.log
server ready after start, this package (stage 7, fresh install, empty tensor cache) 1484.3 s measured stages/7/verify/evidence/verify.json, stages/7/verify/evidence/server.log

Limitations

  • Tested on 2 chips (mesh P150x2) only; other configurations are not claimed by this bundle.
  • Accuracy was compared with the CPU reference on one short fixed prompt only.
  • Speculative decoding is off, so decode speed is that of plain decoding.
  • Sampling runs on the host.
  • joint_head.safetensors is not served.

Risks and safety considerations

The check above does not cover long outputs, tool calling or safety behaviour. Output can differ from Cloudflare/clef run on a CPU or GPU in ways it does not show.

Not measured

  • a download of the weights through tt-model pull, and a boot of the package from the Hub

Boot check

Stage 7 of the run installed this bundle, served it on a leased board and compared its tokens with the CPU reference. The Expected performance table has the result.

How to serve

tt-model serve episod/clef-p300
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for episod/clef-p300

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Finetuned
(7)
this model