SLIMDER Qwen3.8 REAM-288 Depth-32 Agentic — N-gram 50% Compact
This is the promoted BF16 compact checkpoint derived from
sjakek/slimder-qwen38-ream288-depth32-agentic at revision
7134cf0f0b9db5450e17f7db5c090daeb58bca7b.
The PLE n-gram table was compacted to 50% capacity while the original source checkpoint remains unchanged. The selected map uses the validated promoted hybrid policy: activation-aware selections for bigram heads 0–7 and the frequency baseline for trigram heads 8–15.
Size
- Parameters: 74,615,655,680
- Serialized tensor bytes: 150,511,416,232 (150.51 GB / 140.17 GiB)
- Original checkpoint tensor bytes: 200,431,688,376
- Trainable parameters removed: approximately 25.60 billion
- Compact PLE rows: 160,000,768
- Original PLE rows represented by the remap: 320,001,446
Provenance and validation
- Selection manifest SHA-256:
24a0cf8360699362b6a99da498a4138b6a390918bf9cfc9df5b15bcdc3397983 - Compact PLE SHA-256:
8e3bbd1425e4bae6c49a789f12e4ec708890662ad6bff308fe3da1533f76e391 - Global remap SHA-256:
d7eb33bfa9212c7ee2ce82c42feaab782c05051f0c78899256d8d274688140a8 - Model index SHA-256:
79e5cf5c268369d0be19afe61a4b67beb53110fc294fbc08a74cfb85328bab46 - Two-GPU structural forward and deterministic generation smoke: passed
Selection and holdout evidence is published in
sjakek/slimder-qwen38-ngram-activation-refine-20260901.
Runtime note
The compact PLE uses a remap table and requires the included
transformers_compact_runtime.py compatibility installer with the pinned
Qwen4 experimental Transformers revision. This is a BF16 research checkpoint,
not the later NVFP4 smoke artifact.
Limitations
The structural and n-gram coverage tests passed, but this checkpoint still requires task-level evaluation and substantive post-pruning fine-tuning before production deployment.
Qwen3.8 Perian project lineage
This repository is retained in the Qwen3.8 Perian checkpoints collection. Its exact position in the lineage is: BF16 compact checkpoint after depth, expert-width, and PLE n-gram reductions, before the final rank-32 QLoRA.
The final Qwen3.8 Perian GGUF release combines three reductions and one post-training stage:
- depth: 48 to 32 transformer layers;
- routed-expert width: 384 to 288 experts per layer;
- PLE n-gram capacity: 320,001,446 to 160,000,768 rows (50%, about 25.60B parameters removed), using activation-aware bigram and frequency-ranked trigram selections validated on a document-disjoint 5M-token holdout;
- rank-32 QLoRA on 12,558 normalized traces spanning math/STEM reasoning, coding/debugging, agentic tool use, retrieval, and general multi-step reasoning. The trace mixture draws from several frontier-model families, including Fable 5, GLM 5.2, Kimi K3, Claude Opus 4.7, Qwen3.8-Max, and GPT-5.6-Sol. The final merged milestone was trained through 9,336,692 supervised assistant tokens.
Earlier checkpoints in this collection do not inherit later stages merely by being listed beside them; the stage statement above is authoritative for this artifact.
- Downloads last month
- -
Model tree for jakeatx/Qwen3.8-Perian-BF16
Base model
jakeatx/slimder-qwen38-reap384-s0