SLIMDER Qwen3.8 REAM-288 Depth-32 Agentic — N-gram 50% Compact

This is the promoted BF16 compact checkpoint derived from sjakek/slimder-qwen38-ream288-depth32-agentic at revision 7134cf0f0b9db5450e17f7db5c090daeb58bca7b.

The PLE n-gram table was compacted to 50% capacity while the original source checkpoint remains unchanged. The selected map uses the validated promoted hybrid policy: activation-aware selections for bigram heads 0–7 and the frequency baseline for trigram heads 8–15.

Size

  • Parameters: 74,615,655,680
  • Serialized tensor bytes: 150,511,416,232 (150.51 GB / 140.17 GiB)
  • Original checkpoint tensor bytes: 200,431,688,376
  • Trainable parameters removed: approximately 25.60 billion
  • Compact PLE rows: 160,000,768
  • Original PLE rows represented by the remap: 320,001,446

Provenance and validation

  • Selection manifest SHA-256: 24a0cf8360699362b6a99da498a4138b6a390918bf9cfc9df5b15bcdc3397983
  • Compact PLE SHA-256: 8e3bbd1425e4bae6c49a789f12e4ec708890662ad6bff308fe3da1533f76e391
  • Global remap SHA-256: d7eb33bfa9212c7ee2ce82c42feaab782c05051f0c78899256d8d274688140a8
  • Model index SHA-256: 79e5cf5c268369d0be19afe61a4b67beb53110fc294fbc08a74cfb85328bab46
  • Two-GPU structural forward and deterministic generation smoke: passed

Selection and holdout evidence is published in sjakek/slimder-qwen38-ngram-activation-refine-20260901.

Runtime note

The compact PLE uses a remap table and requires the included transformers_compact_runtime.py compatibility installer with the pinned Qwen4 experimental Transformers revision. This is a BF16 research checkpoint, not the later NVFP4 smoke artifact.

Limitations

The structural and n-gram coverage tests passed, but this checkpoint still requires task-level evaluation and substantive post-pruning fine-tuning before production deployment.

Qwen3.8 Perian project lineage

This repository is retained in the Qwen3.8 Perian checkpoints collection. Its exact position in the lineage is: BF16 compact checkpoint after depth, expert-width, and PLE n-gram reductions, before the final rank-32 QLoRA.

The final Qwen3.8 Perian GGUF release combines three reductions and one post-training stage:

  • depth: 48 to 32 transformer layers;
  • routed-expert width: 384 to 288 experts per layer;
  • PLE n-gram capacity: 320,001,446 to 160,000,768 rows (50%, about 25.60B parameters removed), using activation-aware bigram and frequency-ranked trigram selections validated on a document-disjoint 5M-token holdout;
  • rank-32 QLoRA on 12,558 normalized traces spanning math/STEM reasoning, coding/debugging, agentic tool use, retrieval, and general multi-step reasoning. The trace mixture draws from several frontier-model families, including Fable 5, GLM 5.2, Kimi K3, Claude Opus 4.7, Qwen3.8-Max, and GPT-5.6-Sol. The final merged milestone was trained through 9,336,692 supervised assistant tokens.

Earlier checkpoints in this collection do not inherit later stages merely by being listed beside them; the stage statement above is authoritative for this artifact.

Downloads last month
-
Safetensors
Model size
75B params
Tensor type
BF16
·
I64
·
I32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jakeatx/Qwen3.8-Perian-BF16