Instructions to use kamyarkazemi1373/Qwen3-4B-W8A8-RK3588 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RKLLM
How to use kamyarkazemi1373/Qwen3-4B-W8A8-RK3588 with RKLLM:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Qwen3-4B W8A8 for RK3588 (Orange Pi 5)
Quantized Qwen/Qwen3-4B models for Rockchip RK3588 NPU using RKLLM.
Files
| File | hybrid_rate | Description | Size |
|---|---|---|---|
Qwen3-4B-w8a8-npu.rkllm |
0.0 | All layers on NPU โ fastest throughput | 4.51 GB |
Qwen3-4B-w8a8-hybrid.rkllm |
0.5 | 50% CPU A76 + 50% NPU โ lower NPU memory pressure | 4.54 GB |
Quantization Details
- Source model: Qwen/Qwen3-4B (HuggingFace)
- Toolkit: rkllm-toolkit 1.2.1b1
- dtype: W8A8 (8-bit weights + 8-bit activations)
- Algorithm: normal
- Platform: rk3588, num_npu_core=3
- Max context: 4096 tokens
- Calibration: 20 representative prompts (reasoning, math, code, multilingual)
Note: W4A16 is not supported by rkllm-toolkit 1.2.1b1 for RK3588.
Usage
See the full deployment stack at: https://github.com/kamyarkazremi/orangepi5-rkllm
Includes:
rkllm_enhancedbinary with n_keep=4 sliding-window KV cache- Patched API server with ChatML support and zombie recovery
- Systemd service configuration
- Downloads last month
- 22
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support