DeepSeek V3 Pruned 8 Experts With MTP

Yeahh... so I pruned this like an idiot before. I totally removed the MTP module.

We now have num_nextn_predict_layers = 1 and the appended model.layers.61.* MTP module, including enorm, hnorm, eh_proj, embed_tokens, and shared_head tensors.

The routed expert banks have been pruned to experts 0..7, and layer 61 uses the same 8 routed experts as the main MoE layers.

Downloads last month
44
Safetensors
Model size
40B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ckoh04/deepseek-v3-pruned-8experts-mtp

Quantized
(17)
this model