Qwen3 0.6B Dynamic INT8 for Qualcomm SM8650

This is a Qualcomm SM8650 AOT compilation of litert-community/Qwen3-0.6B file Qwen3-0.6B.litertlm. The original Qwen metadata and tokenizer are packaged with the compiled prefill/decode graph.

Artifact

  • File: qwen3-0.6b-sm8650-dynamic-int8.litertlm
  • Size: 658,959,760 bytes
  • SHA-256: dbb105f85adb0d88c028ea7a4a4fbd27439df72e15e002e9f2201a5fa22b8d5c
  • Context limit recorded by the source package: 4096 tokens
  • Target: Qualcomm SM8650
  • Build completed: 2026-09-11T21:49:59+00:00

The source bundle uses dynamic INT8 weights and a floating-point KV cache. Qualcomm AOT compilation and host-side LiteRT-LM container inspection passed. No physical SM8650 inference test has been performed, so runtime compatibility, quality, latency, memory use, and accelerator fallback remain unverified.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for macunaima/Qwen3-0.6B-SM8650-LiteRT-LM-Dynamic-INT8

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1249)
this model