context-cache-lab — Stage 2b memory extractor (research artifact)
⚠️ This is a negative-result research artifact, not a useful model.
It is published so that a failed experiment is reproducible. Its own preregistered gates returned INCONCLUSIVE (Stage 2b) and FAIL (Stage 3). Do not use it for anything in production. Do not read a capability claim into it.
A 566M-parameter C²KV-style sidecar that compiles chunks of text into compressed,
pre-RoPE KV "pages" for a frozen Qwen/Qwen3-4B. It does not generate text and
cannot be used standalone — it only produces key/value state that the frozen target
consumes.
Code, protocol, preregistration and all results: https://github.com/johnathonkillaly/context-cache-lab
What was measured
| Stage | Verdict | Result |
|---|---|---|
| 2b — does learned compressed state beat equal-budget raw text? | 🟡 INCONCLUSIVE | Beat BUDGET at every ratio, but missed 2 of 4 frozen criteria by ~1% |
| 3 — do independently compiled pages support two-page reasoning? | ❌ FAIL | 0/13 valid items vs NATIVE 13/13; all 5 gates missed |
Stage 2b (single-document QA, 4096 tokens, clean_hit):
| ratio | this extractor | equal-budget raw text | NATIVE |
|---|---|---|---|
| 2× | 0.569 | 0.472 | 0.958 |
| 4× | 0.389 | 0.222 | 0.958 |
| 8× | 0.347 | 0.125 | 0.958 |
| 16× | 0.278 | 0.069 | 0.958 |
Semantic facts survived compression (0.625 at 4×) while exact strings did not — hashes collapsed to 0.000 and identifiers to 0.250, where raw text at the same budget scored 0.750. Warm reuse was 23–32× faster than native prefill with breakeven at 2 queries, but compilation costs more than a single prefill, so the first query is slower.
Stage 3 then evaluated this exact frozen checkpoint on a preregistered cross-page compositional task. It scored zero clean accuracy on every native-valid item, with a mean rank margin (−2.47) worse than no context at all (−1.10). Compiling the document jointly instead of per-page also scored zero, so this is not an independence tax.
What that failure does not establish: Stage 3 used a different corpus (~94–100 token pages against ~256-token training chunks), so it measures insufficient transfer of this carrier to that task — not the impossibility of independently compiled memory in general. It also cannot fully separate loss of within-page fidelity from failure to combine intact relations. See the repository's Stage 3 report for the full limitations.
Files
| File | Purpose |
|---|---|
stage2b_extractor.safetensors |
Use this. 109 fp32 tensors, safe to load. Verified to round-trip exactly against the original. |
stage2b_extractor.pt |
Provenance artifact only. Its SHA-256 is pinned in the repo's results/stage3/freeze.json, so the Stage 3 integrity audit reproduces exactly. It is a pickle — the repo's scripts load it with weights_only=False. Only use it if you need the hash to verify. |
config.json |
layer_share, n_sink, training step, target model revision, and the original .pt SHA-256. |
original .pt sha256: b8384dd2eeaa87890a6b65e22f791183399764b93dd9a7f4e8a6e4ce13fca425
target model revision: 1cfa9a7208912126459214e8b04321603b3df60c
training step: 2200
Usage
The extractor is only meaningful alongside the repository's code:
git clone https://github.com/johnathonkillaly/context-cache-lab
cd context-cache-lab
uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e ".[dev]"
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from ccl.compressor import MemoryExtractor
from ccl.target import TargetModel
target = TargetModel("Qwen/Qwen3-4B") # frozen
extractor = MemoryExtractor(target, layer_share=1, n_sink=8)
extractor.load_state_dict(load_file(hf_hub_download(
"jkillay/context-cache-lab-stage2b-extractor",
"stage2b_extractor.safetensors")))
extractor.eval()
pages = extractor.compile_chunks([target.encode(c) for c in chunks], ratio=4.0)
Requires transformers>=5.0 (the legacy tuple KV-cache format is gone) and roughly
10 GB of memory for the frozen target plus the sidecar. Developed on Apple Silicon
(MPS); no CUDA is assumed.
Training
2,200 steps, ~108 minutes on an Apple M4 Max, compression ratio sampled per example
(C²KV's -Dyn variant). The target model was frozen throughout — only the shared
memory-token embedding and per-layer Q/K/V projection heads received gradients.
Supervision was on answer tokens only, applied after page concatenation, so the
extractor is pushed toward states that compose rather than states that are merely
individually informative.
Training data was a synthetic corpus generated deterministically from integer seeds, with value pools provably disjoint from the held-out evaluation draw. No scraped text and no personal data.
License and attribution
Apache-2.0. This is a derivative of Qwen/Qwen3-4B
(Apache-2.0) in the sense that its projections were initialized from that model's
weights; no Qwen weights are redistributed here — only the trained sidecar.
The mechanism is a reimplementation of the published design of C²KV (Du et al., KDD 2026). No C²KV source code was copied. Please cite that paper alongside this artifact.
Citation
See CITATION.cff in the
GitHub repository.
- Downloads last month
- -