Marco-Mini-Instruct-REAP20

Marco-Mini-Instruct (17.3B total, ~0.86B active, 256 experts, 8 active per token) with 20% of its experts removed by REAP: 205 of 256 experts kept. These are the full-precision (bf16) safetensors weights, for fine-tuning, re-quantising or running with transformers.

For ready-to-run quantised files at every pruning ratio, see kueizen/Marco-Mini-Instruct-REAP-GGUF.

Load it

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "kueizen/Marco-Mini-Instruct-REAP20"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")

How this was made

  • Pruning: REAP (Router-weighted Expert Activation Pruning, Cerebras Research), using the reference implementation. REAP scores each expert by its router weight times the size of its output over a calibration set, then removes the lowest-scoring experts whole. It is one-shot: no retraining. Seed 42.
  • Calibration set: theblackcat102/evol-codealpaca-v1 (train split, shuffled, seed 42), 64 samples per category, batch size 1, max sequence length 2048 tokens.
  • Experts removed: 20% (205 of 256 kept).

We have not run task benchmarks on these checkpoints (MMLU, coding, multilingual). Test them on your own workload before relying on them.

License and credits

Derived from ATH-MaaS/Marco-Mini-Instruct and released under the same Apache 2.0 licence. What we changed: removed experts with REAP. Nothing else was modified or retrained. REAP is by Cerebras Research: paper, code. All credit for the base model goes to its authors.

Published by Kueizen.

Downloads last month
-
Safetensors
Model size
14B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kueizen/Marco-Mini-Instruct-REAP20

Finetuned
(7)
this model

Paper for kueizen/Marco-Mini-Instruct-REAP20