GLM-4.7-Flash-REAP-Code-48E

Code-calibrated standard REAP one-shot MoE Expert pruning of zai-org/GLM-4.7-Flash. This checkpoint intentionally does not use Korean hard protection.

Routed Experts were pruned 64 → 48 per layer (25.0%) while top-4 routing was kept unchanged. This percentage is the reduction in routed Expert count, not total model parameters. The checkpoint contains 24.27B parameters (base 31.22B; 22.25% total-parameter reduction) and occupies 45.23 GiB (approximately 22.2% smaller than the base checkpoint), is stored in BF16/F32 safetensors, and received no fine-tuning.

Status Result
Checkpoint integrity checks PASS
Internal release gate NOT APPLICABLE AS A K-GUARD CANDIDATE
External reproduction NOT YET REPORTED
Deployment status EXPERIMENTAL COMPARISON CONTROL

Comparison control. This checkpoint uses code-centric REAP without Korean hard protection. It is published to provide a comparison point, not as a Korean-preserving or deployment-recommended model.

What this checkpoint shows

  • Method: REAP-Code baseline
  • Korean MC macro: 53.5 (base 68.8, -15.3 pp)
  • Overall non-Korean composite: 60.4 (base 65.5, -5.1 pp)
  • Coding 2-benchmark macro: 74.0 (base 76.0, -2.0 pp)
  • LiveCodeBench v6 pass@1: 25.1 (base 22.9, +2.3 pp)
  • Coding 3-benchmark macro: 57.7 (base 58.3, -0.6 pp)
  • Protection budget: No Korean hard-protection constraint was applied.

Benchmark results

benchmark base K-Guard 12.5% K-Guard force-fit 25% K-Guard force-fit 50% REAP-Code 25% (this checkpoint)
Korean (held-out harness)
KMMLU 49.4 48.8 42.4 23.6 35.0
KoBEST BoolQ 90.2 87.7 78.8 74.3 79.6
KoBEST COPA 70.7 69.7 65.0 55.9 51.7
KoBEST SentiNeg 74.8 71.8 70.5 53.7 49.6
KoBEST HellaSwag 58.8 55.0 57.4 48.8 51.6
English reasoning / knowledge
ARC-Challenge 87.1 86.1 79.9 36.9 79.1
HellaSwag (en) 74.2 67.4 55.3 32.6 61.7
WinoGrande 67.4 65.6 59.2 53.1 63.5
MMLU-Pro 41.2 36.9 34.2 14.8 33.8
TruthfulQA MC2 47.3 46.0 43.1 40.6 44.6
MMLU-en 69.7 64.5 56.6 30.1 59.6
Math
GSM8K 86.3 86.7 81.8 61.3 86.7
MathQA 30.5 29.1 24.8 21.3 28.7
Coding
HumanEval+ 81.1 75.6 74.4 65.2 80.5
MBPP+ 70.9 69.3 67.5 61.6 67.5
LiveCodeBench v6 22.9 25.1 20.6 16.6 25.1
Safety / bias
BBQ 76.6 69.9 59.8 33.6 74.0
SafetyBench 74.8 73.0 69.1 44.3 72.1
Macro
Korean MC macro 68.8 66.6 62.8 51.2 53.5
English reasoning / knowledge macro (6) 64.5 61.1 54.7 34.7 57.0
Math macro (2) 58.4 57.9 53.3 41.3 57.7
Safety / bias macro (2) 75.7 71.5 64.5 39.0 73.0
Overall non-Korean composite (10) 65.5 62.5 56.4 36.9 60.4
Coding 2-bench macro 76.0 72.5 70.9 63.4 74.0
Coding 3-bench macro 58.3 56.7 54.1 47.8 57.7

Aggregate rows are unweighted arithmetic means. The overall non-Korean composite averages six English reasoning/knowledge metrics, two math metrics, and two safety/bias metrics. The coding 3-benchmark macro averages HumanEval+, MBPP+, and LiveCodeBench v6 pass@1. These mixed summaries are compact internal comparisons, not standardized leaderboard scores.

Compatibility

Runtime or feature Status
Transformers 5.0.0, PyTorch/CUDA, eager attention and eager Experts on DGX Spark Validated
Native from_pretrained(..., trust_remote_code=False) loading Validated
Speculative MTP/NextN decoding Unsupported (mtp_compatible=false)
vLLM Untested for this pruned checkpoint
SGLang Untested for this pruned checkpoint
  • Architecture: Glm4MoeLiteForCausalLM.
  • Model code: native Transformers implementation; no repository-side modeling_*.py is required, and the checkpoint loads with trust_remote_code=False.
  • Preserved: first Dense layer, shared Experts, topk_method=noaux_tc, norm_topk_prob=true, router correction bias, and routed_scaling_factor=1.8.
  • Routing: top-4 remains unchanged; only routed Expert tensors and corresponding router rows/bias entries were compacted with one old→new ID map per layer.
  • The extra layer-47 MTP/NextN tensors are retained for checkpoint completeness, but runtime validation reports mtp_compatible=false. Use standard autoregressive generation; speculative MTP/NextN decoding is not supported by this release.
  • Top-4 routing is unchanged, so the checkpoint-size and resident-weight reduction must not be interpreted as a proportional reduction in active Expert compute or latency. Cold-load time, TTFT, prefill/decode throughput, batch scaling, power, and 4K/32K/128K serving behavior have not been benchmarked.

How it was made

Source checkpoint

  • Base: zai-org/GLM-4.7-Flash
  • Frozen revision: 7dd20894a642a0aa287e9827cb1a1f7f91386b67
  • Original routed MoE layout: first layer Dense, then 46 routed MoE layers with 64 Experts, top-4 routing, plus shared Experts.

Calibration dataset

Source Records Observed tokens Role
theblackcat102/evol-codealpaca-v1 8,192 4,366,438 Code-centric standard REAP baseline
  • No Korean segment weighting, Korean↔English contrast, or hard-protected set was used.
  • The manifest passed a separate internal token-count, token-hash, and exact-prompt contamination validator.
  • The exact selected-sample manifest is not included in this model repository, so the complete run is not yet externally reproducible from public materials alone.

Observation and scoring

  • Selected-only observer: only routed Expert assignments were measured; no dense [experts, tokens, hidden] activation tensor was retained.
  • REAP contribution: normalized dispatch weight × L2 norm of the unweighted Expert output.
  • Router selection and dispatch were checked separately so GLM's correction bias affects selection without being mistaken for the actual dispatch weight.
  • The observed model remained logit-identical to the uninstrumented eager model in validation.

Pruning

  • Uniform routed Experts per layer: 64 → 48.
  • Shared Experts and the first Dense layer were not pruned.
  • Top-4 and routed scaling factor 1.8 were retained.
  • Safetensors surgery physically compacted Expert tensors, router rows, and correction-bias entries.
  • Protection result: No Korean hard-protection constraint was applied.

Evaluation scope

These are internal comparative evaluation results, not standardized leaderboard numbers:

  • KMMLU: deterministic 500-example subset (seed 42).
  • KoBEST: declared BoolQ, COPA, SentiNeg, and HellaSwag evaluation splits.
  • ARC-Challenge, HellaSwag-en, WinoGrande, MMLU-Pro, TruthfulQA MC2, MMLU-en, GSM8K, MathQA, BBQ, and SafetyBench: deterministic 512-example subsets.
  • HumanEval+: full 164 tasks; MBPP+: frozen full 378-task revision; greedy pass@1.
  • LiveCodeBench: full incremental v6 split (175 problems), n=1 greedy no-thinking generation, batch size 1, with official functional tests executed in an isolated ARM64 container. This protocol is not directly comparable to the official leaderboard's n=10 sampling protocol.
  • The machine-readable field decision_eligible=true means that the run passed the project's internal screening checks. It is not an external certification. release_eligible=false.
  • Do not compare these subset scores directly with full-set or few-shot leaderboard values.

Machine-readable results and the exact survivor manifest are included under evaluation/.

Internal classification: This is a code-centric REAP comparison control. K-Guard release criteria do not apply because Korean hard protection was intentionally omitted.

Interpretation boundaries

  • At the same 48-Expert size, K-Guard-48E has a 9.3 pp higher Korean MC macro, while this checkpoint has a 4.0 pp higher overall non-Korean composite, a 4.6 pp higher LiveCodeBench score, and a 3.6 pp higher coding 3-benchmark macro.
  • This is not a controlled single-variable comparison: the two checkpoints differ in both calibration data and survivor-selection rules. It cannot isolate the causal effect of hard protection.
  • Machine-readable results include direct K-Guard-48E versus REAP-Code-48E McNemar tests for 11 non-Korean binary/pass-fail tasks. Direct paired tests for the Korean benchmark set, aggregate bootstrap confidence intervals, multiple-comparison correction, and repeated calibration seeds are absent.
  • PASS in the artifacts refers to checkpoint integrity and internal pipeline validation, not external replication.

Usage

Install a Transformers release that includes Glm4MoeLiteForCausalLM:

pip install "transformers>=5.0.0" accelerate safetensors

The checkpoint was validated with Transformers 5.0.0 and NVIDIA PyTorch 25.09. It uses the native Transformers model implementation and does not execute custom code from this repository.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "buugiiiiii/GLM-4.7-Flash-REAP-Code-48E"
tokenizer = AutoTokenizer.from_pretrained(
    model_id,
    trust_remote_code=False,
)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=False,
    attn_implementation="eager",
    experts_implementation="eager",
)
model.eval()

messages = [{"role": "user", "content": "한국어로 이진 탐색을 설명해줘."}]
prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.inference_mode():
    output = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=False,
        pad_token_id=tokenizer.eos_token_id,
    )

print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Limitations

Expert pruning is irreversible within this checkpoint and can change routing behavior sharply on tasks not represented by the calibration or evaluation sets. No post-pruning recovery or fine-tuning was applied. Reported Korean behavior is relative to the tested base checkpoint and harness; it does not imply universal Korean-language safety or quality preservation. The current release is not validated for speculative MTP/NextN decoding.

The Korean evaluation is primarily multiple-choice. Korean long-form generation, summarization, multi-turn dialogue, format-following, code explanation/debugging, factuality, hallucination, safety, tool calling, JSON output, and long-context retrieval were not comprehensively evaluated. Chinese capability is inherited from the base model but was not specifically protected or evaluated after pruning. vLLM/SGLang compatibility and deployment performance remain unverified.

Included research artifacts

  • evaluation/glm-full-b20-results-with-livecodebench.json: machine-readable five-model full results
  • evaluation/glm-full-b20-long-scores.csv: compact long-format B20 scores
  • evaluation/glm-full-b20-independent-validation.json: independent B20 validation report
  • evaluation/survivor-manifest.json: exact per-layer old Expert IDs retained
  • artifacts/calibration-validation.json: separate internal calibration-manifest validation
  • artifacts/observation-validation.json: selected-only observation validation
  • artifacts/expert-statistics.csv: per-Expert scoring/features used for this method

Provenance

  • Base model: zai-org/GLM-4.7-Flash, revision 7dd20894a642a0aa287e9827cb1a1f7f91386b67
  • Base license: MIT
  • General method code: K-REAP, commit 234666e
  • Reproducibility note: the linked repository documents K-REAP and its public adapters, but the exact GLM-specific observation pipeline, run configuration, and evaluation manifests used for this checkpoint are not yet published there.
  • Team: RAS-LAB
  • Benchmark report SHA256: 71ab2ce742bf8a481d90d10c8b8684a918241fd475729b0bf2afe56b44335dd6
  • Model card generated from an internally validated benchmark artifact with zero validation failures; this is not an external reproduction claim.

Citation

Please cite the base checkpoint and REAP when using this model:

@misc{zai2026glm47flash,
  title        = {GLM-4.7-Flash},
  author       = {{Z.ai}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/zai-org/GLM-4.7-Flash}},
  note         = {Base revision 7dd20894a642a0aa287e9827cb1a1f7f91386b67}
}

@article{lasby2025reap,
  title   = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
  author  = {Lasby, Mike and Lazarevich, Ivan and Sinnadurai, Nish and Lie, Sean and Ioannou, Yani and Thangarasa, Vithursan},
  journal = {arXiv preprint arXiv:2510.13999},
  year    = {2025},
  doi     = {10.48550/arXiv.2510.13999}
}
Downloads last month
85
Safetensors
Model size
24B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for buugiiiiii/GLM-4.7-Flash-REAP-Code-48E

Finetuned
(70)
this model

Paper for buugiiiiii/GLM-4.7-Flash-REAP-Code-48E