YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Pruned Qwen1.5 MoE (thin wrapper)

This repository is a thin wrapper that loads base weights from Qwen/Qwen1.5-MoE-A2.7B and masks selected experts per layer during inference routing.

Usage

from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "ffront/qwen1_5_moe_pruned",
    trust_remote_code=True,
    pruned_experts_by_layer={
        "*": [0],
        5: [1, 2, 3],
    },
)

Behavior

  • Pruned experts are masked before softmax/top-k in sparse MoE routing.
  • Masked experts do not occupy top-k slots.
  • Base model weights are loaded from Qwen/Qwen1.5-MoE-A2.7B at load time.
  • This wrapper does not zero expert weights.

Publish to Hugging Face Hub

  1. Create repository: https://huggingface.co/new?name=qwen1_5_moe_pruned&organization=Ffr0nt
  2. Add files from this folder to repo root.
  3. Push with git or upload in web UI.
  4. Validate loading using the usage snippet above.
Downloads last month
7
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support