skim projection artifacts

Per-model projection files for skim, a drop-in sparse-KV decode backend for vLLM 0.26. Each file is a static, low-rank linear projection fitted offline from that checkpoint's own activations, with the fitting method withheld. It contains no model weights and no text, is useless without the upstream checkpoint, and is refused by skim for any other model id.

What is here

file calibrated against upstream revision upstream licence size sha256
skim_proj_mistral24b_16k_d48_v1.pt mistralai/Mistral-Small-24B-Base-2501 b0a2e4ed093c26997495ae625528f81ea04b749f Apache-2.0 42 MB 49482a4ce2b43253c3eac09a9bc47b2bbcb74928ef18dbee2c666fe330f85388
skim_proj_qwen25-32b_32k_fresh_v1.pt Qwen/Qwen2.5-32B 1818d35814b8319459f4bd55ed1ac8709630f003 Apache-2.0 67 MB 9c4610491e1c0d9b036a83e50fba5ff6c2dee124e2b88f80813aaaba71346441

Both models are certified for skim at their shipped op-points (Mistral-Small: d′=48 int4, k=2048, N=30720; Qwen2.5-32B: d′=32 int8, k=1024, N=30720) by paired non-inferiority to dense fp8 on RULER within 5 accuracy points. A certification is valid only at its own context length, tasks, anchor and tensor-parallel size; see the skim docs.

Terms

These artifacts are distributed under the strictest reading of the upstream licence: each is treated as if it were subject to the licence of the checkpoint it was calibrated against, whether or not it is legally a derivative work of that checkpoint. Both upstream checkpoints here are Apache-2.0, so each artifact is redistributed under Apache-2.0 with the upstream attribution in NOTICE. skim's own code is separately Apache-2.0.

You must hold a valid licence to the upstream weights; possession of a .pt conveys no right to the model. Do not represent an artifact as the upstream model or as endorsed by its authors. Keep the provenance recorded in the file and in NOTICE.

Use

pip install "skimattn[vllm]"
skim pull Qwen/Qwen2.5-32B                       # fetches the file from this repo and verifies its sha256
skim pull mistralai/Mistral-Small-24B-Base-2501
skim serve Qwen/Qwen2.5-32B --max-model-len 30720

skim pins the revision of this repository it fetches from and refuses a file whose sha256 does not match its shipped config, so a change here never reaches a user silently.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support