skim projection artifacts
Per-model projection files for skim, a drop-in sparse-KV decode backend for vLLM 0.26. Each file is a static, low-rank linear projection fitted offline from that checkpoint's own activations, with the fitting method withheld. It contains no model weights and no text, is useless without the upstream checkpoint, and is refused by skim for any other model id.
What is here
| file | calibrated against | upstream revision | upstream licence | size | sha256 |
|---|---|---|---|---|---|
skim_proj_mistral24b_16k_d48_v1.pt |
mistralai/Mistral-Small-24B-Base-2501 |
b0a2e4ed093c26997495ae625528f81ea04b749f |
Apache-2.0 | 42 MB | 49482a4ce2b43253c3eac09a9bc47b2bbcb74928ef18dbee2c666fe330f85388 |
skim_proj_qwen25-32b_32k_fresh_v1.pt |
Qwen/Qwen2.5-32B |
1818d35814b8319459f4bd55ed1ac8709630f003 |
Apache-2.0 | 67 MB | 9c4610491e1c0d9b036a83e50fba5ff6c2dee124e2b88f80813aaaba71346441 |
Both models are certified for skim at their shipped op-points (Mistral-Small: d′=48 int4, k=2048, N=30720; Qwen2.5-32B: d′=32 int8, k=1024, N=30720) by paired non-inferiority to dense fp8 on RULER within 5 accuracy points. A certification is valid only at its own context length, tasks, anchor and tensor-parallel size; see the skim docs.
Terms
These artifacts are distributed under the strictest reading of the upstream licence: each is treated as if it
were subject to the licence of the checkpoint it was calibrated against, whether or not it is legally a derivative
work of that checkpoint. Both upstream checkpoints here are Apache-2.0, so each artifact is redistributed under
Apache-2.0 with the upstream attribution in NOTICE. skim's own code is separately Apache-2.0.
You must hold a valid licence to the upstream weights; possession of a .pt conveys no right to the model. Do not
represent an artifact as the upstream model or as endorsed by its authors. Keep the provenance recorded in the file
and in NOTICE.
Use
pip install "skimattn[vllm]"
skim pull Qwen/Qwen2.5-32B # fetches the file from this repo and verifies its sha256
skim pull mistralai/Mistral-Small-24B-Base-2501
skim serve Qwen/Qwen2.5-32B --max-model-len 30720
skim pins the revision of this repository it fetches from and refuses a file whose sha256 does not match its shipped config, so a change here never reaches a user silently.