Qwen3-4B (LiteRT-LM)

Qwen3-4B by Alibaba — Quantized as a LiteRT-LM model (.litertlm) for on-device inference.

Model Details

  • Base model: Qwen3-4B by Alibaba
  • Format: LiteRT-LM (.litertlm)
  • Quantized: INT8 channelwise with float32 KV cache
  • Size: ~5.4 GB
  • Use case: Chat on mobile devices (AI Chat task only)
  • Runtime: LiteRT-LM on-device (no cloud)

Usage

Deploy the model using any compatible LiteRT runtime.

Or via Hugging Face Hub:

huggingface-cli download Xenna/qwen3.6-4b --local-dir ./models/qwen3.6-4b
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support