ctxprune-small: context compression for AI agents

Live demo: ctxprune vs LLMLingua-2 on logs, JSON and tracebacks

By @imdariotoo · code: riceharvest/ctxprune

A 140M-parameter token classifier that shortens text before it reaches an LLM by deleting the words that matter least. It is a successor to Microsoft's LLMLingua-2, trained for what agents actually read: tool outputs (JSON, logs, stack traces, diffs, directory listings), code, documents and chat, in many languages.

  • Keeps identifiers intact. Text is rebuilt from whole source spans, so IDs, paths, URLs, IPs and versions are never split or re-spaced (LLMLingua-2 turns ord_8f3a2c91 into ord _ 8f3a2c91 and 0.21 into 0. 21). Optional force_protected=True guarantees no identifier is dropped.
  • Better answers from compressed text. At the same share of LLM tokens kept, QA accuracy from the compressed text is +15 points higher than with LLMLingua-2-large at 50% kept, and +21 points at 33%. The gain comes from agent tool outputs, agent text and code; on multilingual web prose LLMLingua-2-large is still better (see per-source table).
  • Small and fast. ~102 ms per 500-token chunk on CPU (ONNX fp32), ~93 ms int8, vs ~600 ms for LLMLingua-2-large. ONNX files included; no PyTorch needed.
  • Apache-2.0, trained only on permissively licensed data (LLMLingua-2's training labels are CC BY-NC-SA).

Usage

pip install ctxprune   # CPU, no torch
pip install "ctxprune[torch]"  # PyTorch (GPU/CPU)
from ctxprune import Compressor

c = Compressor("darioooooo0o/ctxprune-small", backend="onnx")    # or backend="torch"
out = c.compress(tool_output, rate=0.5)       # keep ~50% of the text
print(out["text"])

c.compress(tool_output, rate=0.33, force_protected=True)  # never drop IDs, numbers, paths

With LLMLingua (drop-in, needs the tokenizer-detection fix proposed upstream in microsoft/LLMLingua#261):

from llmlingua import PromptCompressor
pc = PromptCompressor(model_name="darioooooo0o/ctxprune-small", use_llmlingua2=True)
pc.compress_prompt(text, rate=0.5, force_tokens=["\n", "?"])

Raw transformers / ONNX: label 1 = keep (same convention as LLMLingua-2). Score each token with softmax(logits)[..., 1], average per word, and keep the highest-scoring words. onnx/model.onnx is exact fp32. onnx/model_quantized.onnx is int8 (143 MB): its keep decisions match fp32 on ~96% of tokens, and QA accuracy at 50% kept is 68.2 vs 68.8 for fp32 (same eval as below).

Results

Held-out evaluation: 378 questions over 196 chunks from repos, MCP servers and sites never seen in training. Each method compresses the chunk to the same share of the reader's tokens; the reader then answers from the compressed text and is scored by exact match on the gold span.

QA exact match, independent reader (DeepSeek V4 Flash):

Method 50% of tokens kept 33% of tokens kept
Original (no compression) 91.8
ctxprune-small 68.8 53.4
ctxprune-small + force_protected 63.5 52.9
LLMLingua-2 large 53.4 32.5
LLMLingua-2 base 41.0 24.9

Per source, at 50%:

Source fineweb2 fineweb edu github code oasst2 swe rebench assistant swe rebench tool swe smith tool toucan mcp
ctxprune-small @50% 56 73 72 75 70 60 84 59
LLMLingua-2 large @50% 83 73 62 62 49 48 29 43

Teacher as reader. An earlier checkpoint (conservative labels only, ~2k chunks) was also scored with Qwen3.8-27B, the labeling teacher, as reader: 65.9 / 51.1 vs LLMLingua-2 large 55.0 / 32.5. That is the same ranking as with the independent reader, so the gain is not an artifact of the teacher grading itself.

Identifiers lost (IDs, numbers, paths, versions that no longer appear intact), at 50%:

Method IDs lost LLM-token ratio
ctxprune-small 41% 0.49
ctxprune-small + force_protected 0% 0.54
LLMLingua-2 large 34% 0.48
LLMLingua-2 base 47% 0.59

Reasoning on vs off for the reader made no difference on a 40-question check, so readers ran without it.

Training

  • Data: 5,249 training chunks (≤500 tokens) from 7 permissive sources: coding-agent tool outputs (SWE-smith, MIT; SWE-rebench OpenHands, CC BY 4.0), MCP tool outputs (Toucan-1.5M, Apache-2.0), source code (github-code-clean, MIT/Apache/BSD/ISC files only), English and multilingual web text (FineWeb-Edu and FineWeb-2, ODC-BY), assistant chat (oasst2, Apache-2.0). Train/val/test are split per repository, MCP server, website or conversation.
  • Labels: Qwen3.8-27B (Apache-2.0) compresses each chunk by deletion only. Its output is aligned to the source with an LCS over atoms. Two prompts are used: a conservative one, and an aggressive one with an explicit word budget. Training on both gives graded targets: words kept even under the aggressive prompt score highest. This teaches ranking inside the large "keep" set, which is what a 33–50% rate needs.
  • Model: mmBERT-small (MIT) with a 2-class token head, 5 epochs, best epoch by ROC-AUC against teacher labels. Training takes 5 minutes on one Intel Arc Pro B70.
  • What mattered (same eval, DeepSeek reader, 50% / 33% kept): conservative labels only, 65.6 / 45.8; aggressive labels only, 65.6 / 50.8; both together (graded), 68.8 / 53.4. Doubling the labels from ~1k to ~2k added 4 points; mmBERT-base instead of small added ~1 point at 2.3x the latency.

Code, data pipeline and evals: https://github.com/riceharvest/ctxprune.

Limitations

  • Multilingual web prose is the weakest domain: LLMLingua-2-large answers more questions there (small sample: 18 questions). The teacher often rewrote non-English text instead of deleting from it, so fewer of those labels survived filtering.
  • force_protected=True guarantees no identifier is dropped, but at 50% it spends budget on IDs the question did not need and scores a few points lower overall. Use it when exact IDs matter more than everything else (e.g. you will act on them).
  • Compression always loses information. At 50% kept, QA accuracy drops from ~92% (uncompressed) to 69%. Use it where the token savings are worth that.
  • Labels come from one teacher LLM, and the eval questions were written by that same model.
  • Rates are measured in the target LLM's tokens in our evals. With the default character budget, CJK text compresses slightly less than the requested rate.

Attribution

Training data includes content under CC BY 4.0 (nebius/SWE-rebench-openhands-trajectories) and ODC-BY (FineWeb-Edu, FineWeb-2). The deletion-only teacher prompt adapts LLMLingua-2's (MIT, Microsoft). Base model: jhu-clsp/mmBERT-small (MIT).

Downloads last month
30
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for darioooooo0o/ctxprune-small

Quantized
(282)
this model

Datasets used to train darioooooo0o/ctxprune-small

Space using darioooooo0o/ctxprune-small 1