Instructions to use mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse
- SGLang
How to use mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse with Docker Model Runner:
docker model run hf.co/mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse
qwen2.5_0.5b-PATCH-45Sparse
This checkpoint (Qwen-2.5 0.5B, PATCH-Joint, 45% sparsity): 40.29% average zero-shot accuracy, 14.57 WikiText2 perplexity.
This repository hosts a mask only release for the paper PATCH: Learnable Tile-level Hybrid Sparsity for LLMs. PATCH (Pruning with a Learnable Tile-level Configuration for Hybrid Sparsity) learns a structured mask on frozen pretrained weights, assigning each tile as dense (0% sparsity) or 2:4 sparse (50% sparsity) to hit a flexible global sparsity target while staying hardware-friendly.
Because PATCH/MaskLLM keep the base weights frozen, we distribute only the
binary keep/prune mask (bit-packed in mask.npz) - no weight values. You
recover the sparse model by downloading the original base model and applying the
mask.
- Base model:
Qwen/Qwen2.5-0.5B - Method: PATCH-Joint | Target sparsity: 45% | Pattern: Dense / 2:4 tiles
- Measured mask sparsity: 44.99%.
Results (Qwen-2.5 0.5B)
| Sparsity | Method | Pattern | Avg Acc (% ↑) | WikiText2 PPL (↓) |
|---|---|---|---|---|
| 0% | Dense | - | 46.00 | 12.08 |
| 50% | Magnitude | 2:4 | 30.16 | 6734.97 |
| 50% | Wanda | 2:4 | 32.97 | 72.48 |
| 50% | SparseGPT | 2:4 | 34.81 | 36.59 |
| 50% | Thanos | 2:4 | 34.85 | 37.32 |
| 50% | ProxSparse | 2:4 | 32.05 | 111.05 |
| 50% | MaskLLM | 2:4 | 39.33 | 15.22 |
| 45% | PATCH-Joint ⭐ | Dense/2:4 | 40.29 | 14.57 |
| 35% | PATCH-Joint | Dense/2:4 | 41.15 | 13.84 |
| 25% | PATCH-Joint | Dense/2:4 | 42.39 | 13.47 |
Per-task zero-shot accuracy (%) for this checkpoint:
| MMLU | PIQA | ARC-E | ARC-C | WinoG. | OBQA | RACE | HellaS. | Average |
|---|---|---|---|---|---|---|---|---|
| 27.39 | 68.44 | 59.13 | 25.77 | 53.67 | 19.80 | 32.15 | 35.99 | 40.29 |
All numbers are from the PATCH paper (arXiv:2509.23410); accuracy is the average over MMLU, PIQA, ARC-Easy, ARC-Challenge, Winogrande, OpenBookQA, RACE and HellaSwag, evaluated with the LM-Evaluation-Harness. PPL is WikiText2.
Training hyper-parameters
| Hyper-parameter | Value |
|---|---|
| Fine-tuning dataset | SlimPajama (2B tokens) |
| Training steps | 2000 |
| Global batch size | 256 |
| Sequence length | 4096 |
| Mask tile size | 128 x 128 (hardware tiles: 128x128 / 128x64 / 64x128 / 64x64) |
| Logits init. | N(0, 0.014) |
| Tile-logit prior | SparseGPT (strength 3) |
| Regularization scope | Global (single target density) |
| Evaluation | LM-Eval-Harness (8 zero-shot tasks) + WikiText2 PPL @ seqlen 4096 |
| Hardware | 1 node x 4 GPUs, data parallel (HuggingFace Trainer) |
| Optimizer | Adam |
| Learning rate | 1e-3 |
| Gumbel scaling (kappa) | 25 -> 350 |
| Gumbel temp (tau) | 4 -> 0.05 |
| Sparsity reg. (lambda1) | 7 |
| Weight reg. (lambda2) | 10 |
How to use
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM
import torch
from load_patch_mask import apply_patch_mask # shipped in this repo
npz = hf_hub_download(repo_id="mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse", filename="mask.npz")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B", torch_dtype=torch.bfloat16)
apply_patch_mask(model, npz) # zeroes the pruned weights in place
Or from the command line:
python load_patch_mask.py --base_model Qwen/Qwen2.5-0.5B --mask_repo mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse
Speedup on real hardware requires a 2:4-aware / hybrid sparse kernel; see the GitHub repository and STOICC.
License
The released mask is a derivative of the base model and is distributed under the
base model's license (apache-2.0). You must comply with that license and
obtain access to the base model separately.
The mask-generation code is released under the MIT license (see the PATCH repository).
Citation
@article{hourri2025patch,
title = {PATCH: Learnable Tile-level Hybrid Sparsity for LLMs},
author = {Hourri, Younes and Mozaffari, Mohammad and Mehri Dehnavi, Maryam},
year = 2025,
journal = {arXiv preprint arXiv:2509.23410}
}
Model tree for mohammad-mozaffari/qwen2.5_0.5b-PATCH-45Sparse
Base model
Qwen/Qwen2.5-0.5B