Qwen3-4B (LiteRT-LM)
Qwen3-4B by Alibaba — Quantized as a LiteRT-LM model (.litertlm) for on-device inference.
Model Details
- Base model: Qwen3-4B by Alibaba
- Format: LiteRT-LM (.litertlm)
- Quantized: INT8 channelwise with float32 KV cache
- Size: ~5.4 GB
- Use case: Chat on mobile devices (AI Chat task only)
- Runtime: LiteRT-LM on-device (no cloud)
Usage
Deploy the model using any compatible LiteRT runtime.
Or via Hugging Face Hub:
huggingface-cli download Xenna/qwen3.6-4b --local-dir ./models/qwen3.6-4b
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support