context-cache-lab — Stage 2b memory extractor (research artifact)

⚠️ This is a negative-result research artifact, not a useful model.

It is published so that a failed experiment is reproducible. Its own preregistered gates returned INCONCLUSIVE (Stage 2b) and FAIL (Stage 3). Do not use it for anything in production. Do not read a capability claim into it.

A 566M-parameter C²KV-style sidecar that compiles chunks of text into compressed, pre-RoPE KV "pages" for a frozen Qwen/Qwen3-4B. It does not generate text and cannot be used standalone — it only produces key/value state that the frozen target consumes.

Code, protocol, preregistration and all results: https://github.com/johnathonkillaly/context-cache-lab

What was measured

Stage Verdict Result
2b — does learned compressed state beat equal-budget raw text? 🟡 INCONCLUSIVE Beat BUDGET at every ratio, but missed 2 of 4 frozen criteria by ~1%
3 — do independently compiled pages support two-page reasoning? FAIL 0/13 valid items vs NATIVE 13/13; all 5 gates missed

Stage 2b (single-document QA, 4096 tokens, clean_hit):

ratio this extractor equal-budget raw text NATIVE
0.569 0.472 0.958
0.389 0.222 0.958
0.347 0.125 0.958
16× 0.278 0.069 0.958

Semantic facts survived compression (0.625 at 4×) while exact strings did not — hashes collapsed to 0.000 and identifiers to 0.250, where raw text at the same budget scored 0.750. Warm reuse was 23–32× faster than native prefill with breakeven at 2 queries, but compilation costs more than a single prefill, so the first query is slower.

Stage 3 then evaluated this exact frozen checkpoint on a preregistered cross-page compositional task. It scored zero clean accuracy on every native-valid item, with a mean rank margin (−2.47) worse than no context at all (−1.10). Compiling the document jointly instead of per-page also scored zero, so this is not an independence tax.

What that failure does not establish: Stage 3 used a different corpus (~94–100 token pages against ~256-token training chunks), so it measures insufficient transfer of this carrier to that task — not the impossibility of independently compiled memory in general. It also cannot fully separate loss of within-page fidelity from failure to combine intact relations. See the repository's Stage 3 report for the full limitations.

Files

File Purpose
stage2b_extractor.safetensors Use this. 109 fp32 tensors, safe to load. Verified to round-trip exactly against the original.
stage2b_extractor.pt Provenance artifact only. Its SHA-256 is pinned in the repo's results/stage3/freeze.json, so the Stage 3 integrity audit reproduces exactly. It is a pickle — the repo's scripts load it with weights_only=False. Only use it if you need the hash to verify.
config.json layer_share, n_sink, training step, target model revision, and the original .pt SHA-256.
original .pt sha256: b8384dd2eeaa87890a6b65e22f791183399764b93dd9a7f4e8a6e4ce13fca425
target model revision: 1cfa9a7208912126459214e8b04321603b3df60c
training step: 2200

Usage

The extractor is only meaningful alongside the repository's code:

git clone https://github.com/johnathonkillaly/context-cache-lab
cd context-cache-lab
uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e ".[dev]"
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from ccl.compressor import MemoryExtractor
from ccl.target import TargetModel

target = TargetModel("Qwen/Qwen3-4B")          # frozen
extractor = MemoryExtractor(target, layer_share=1, n_sink=8)
extractor.load_state_dict(load_file(hf_hub_download(
    "jkillay/context-cache-lab-stage2b-extractor",
    "stage2b_extractor.safetensors")))
extractor.eval()

pages = extractor.compile_chunks([target.encode(c) for c in chunks], ratio=4.0)

Requires transformers>=5.0 (the legacy tuple KV-cache format is gone) and roughly 10 GB of memory for the frozen target plus the sidecar. Developed on Apple Silicon (MPS); no CUDA is assumed.

Training

2,200 steps, ~108 minutes on an Apple M4 Max, compression ratio sampled per example (C²KV's -Dyn variant). The target model was frozen throughout — only the shared memory-token embedding and per-layer Q/K/V projection heads received gradients. Supervision was on answer tokens only, applied after page concatenation, so the extractor is pushed toward states that compose rather than states that are merely individually informative.

Training data was a synthetic corpus generated deterministically from integer seeds, with value pools provably disjoint from the held-out evaluation draw. No scraped text and no personal data.

License and attribution

Apache-2.0. This is a derivative of Qwen/Qwen3-4B (Apache-2.0) in the sense that its projections were initialized from that model's weights; no Qwen weights are redistributed here — only the trained sidecar.

The mechanism is a reimplementation of the published design of C²KV (Du et al., KDD 2026). No C²KV source code was copied. Please cite that paper alongside this artifact.

Citation

See CITATION.cff in the GitHub repository.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jkillay/context-cache-lab-stage2b-extractor

Finetuned
Qwen/Qwen3-4B
Finetuned
(973)
this model

Paper for jkillay/context-cache-lab-stage2b-extractor