YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Pruned Qwen1.5 MoE (thin wrapper)
This repository is a thin wrapper that loads base weights from Qwen/Qwen1.5-MoE-A2.7B and masks selected experts per layer during inference routing.
Usage
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
"ffront/qwen1_5_moe_pruned",
trust_remote_code=True,
pruned_experts_by_layer={
"*": [0],
5: [1, 2, 3],
},
)
Behavior
- Pruned experts are masked before softmax/top-k in sparse MoE routing.
- Masked experts do not occupy top-k slots.
- Base model weights are loaded from Qwen/Qwen1.5-MoE-A2.7B at load time.
- This wrapper does not zero expert weights.
Publish to Hugging Face Hub
- Create repository: https://huggingface.co/new?name=qwen1_5_moe_pruned&organization=Ffr0nt
- Add files from this folder to repo root.
- Push with git or upload in web UI.
- Validate loading using the usage snippet above.
- Downloads last month
- 7
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support