Instructions to use h2loop-ai/qwen3-0.6b-hexagon with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use h2loop-ai/qwen3-0.6b-hexagon with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf h2loop-ai/qwen3-0.6b-hexagon # Run inference directly in the terminal: llama cli -hf h2loop-ai/qwen3-0.6b-hexagon
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf h2loop-ai/qwen3-0.6b-hexagon # Run inference directly in the terminal: llama cli -hf h2loop-ai/qwen3-0.6b-hexagon
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf h2loop-ai/qwen3-0.6b-hexagon # Run inference directly in the terminal: ./llama-cli -hf h2loop-ai/qwen3-0.6b-hexagon
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf h2loop-ai/qwen3-0.6b-hexagon # Run inference directly in the terminal: ./build/bin/llama-cli -hf h2loop-ai/qwen3-0.6b-hexagon
Use Docker
docker model run hf.co/h2loop-ai/qwen3-0.6b-hexagon
- LM Studio
- Jan
- vLLM
How to use h2loop-ai/qwen3-0.6b-hexagon with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "h2loop-ai/qwen3-0.6b-hexagon" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "h2loop-ai/qwen3-0.6b-hexagon", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/h2loop-ai/qwen3-0.6b-hexagon
- Ollama
How to use h2loop-ai/qwen3-0.6b-hexagon with Ollama:
ollama run hf.co/h2loop-ai/qwen3-0.6b-hexagon
- Unsloth Studio
How to use h2loop-ai/qwen3-0.6b-hexagon with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for h2loop-ai/qwen3-0.6b-hexagon to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for h2loop-ai/qwen3-0.6b-hexagon to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for h2loop-ai/qwen3-0.6b-hexagon to start chatting
- Pi
How to use h2loop-ai/qwen3-0.6b-hexagon with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf h2loop-ai/qwen3-0.6b-hexagon
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "h2loop-ai/qwen3-0.6b-hexagon" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use h2loop-ai/qwen3-0.6b-hexagon with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf h2loop-ai/qwen3-0.6b-hexagon
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "h2loop-ai/qwen3-0.6b-hexagon" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use h2loop-ai/qwen3-0.6b-hexagon with Docker Model Runner:
docker model run hf.co/h2loop-ai/qwen3-0.6b-hexagon
- Lemonade
How to use h2loop-ai/qwen3-0.6b-hexagon with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull h2loop-ai/qwen3-0.6b-hexagon
Run and chat with the model
lemonade run user.qwen3-0.6b-hexagon-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use h2loop-ai/qwen3-0.6b-hexagon with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf h2loop-ai/qwen3-0.6b-hexagon
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default h2loop-ai/qwen3-0.6b-hexagon
Run Hermes
hermes
- Atomic Chat
Qwen3-0.6B for the Qualcomm Hexagon NPU
Every v79 binary here is verified end to end on Hexagon v79 (SM8750). Each was pushed to a physical Snapdragon 8 Elite, loaded, and run to completion โ coherence checked across multiple prompts at temperature 0, with KPIs measured from those same runs.
The v81 binary is profiled on real v81 silicon but not executed end to end, because no Snapdragon 8 Elite Gen 5 device is available to us for interactive runs. It is built from the same checkpoint as the verified v79 build, with byte-identical KV quantization. What that does and does not license is spelled out in its section below.
Qwen3-0.6B quantized for Snapdragon Hexagon NPUs, in two deployment paths:
- 4-bit โ 4-bit weights / 16-bit activations, run through llama.cpp's Hexagon HTP backend. Fully measured on real silicon, coherence verified across multiple prompts.
- A16W8 โ int8 weights / int16 activations / 16-bit KV cache, compiled to QAIRT Genie context binaries. Statistically indistinguishable from fp32 on MMLU, with much faster prefill and init. Published here for Hexagon v79, coherence verified on device.
Everything on this page was measured on a Snapdragon 8 Elite (SM8750, Hexagon v79) on Qualcomm Device Cloud. Nothing is estimated from per-graph profiling unless explicitly labelled.
4-bit โ on-device KPIs
Context 4096, all 87 layers resident on the NPU, no CPU or GPU fallback.
| KPI | value |
|---|---|
| init (cold load) | ~900 ms |
| TTFT (17-token prompt) | 38 ms |
| decode | 91.7 tok/s |
| prefill | 3618 tok/s |
| peak RAM | 466 MB host + 448 MB NPU-side KV |
| MMLU vs fp32 | 93.9% |
Decode is quoted at an empty cache โ the figure comparable to a short chat turn. It is not constant with context:
| KV depth | decode | prefill |
|---|---|---|
| 0 | 91.7 tok/s | 3618 tok/s |
| 128 | 90.1 | 3278 |
| 256 | 80.3 | 3293 |
| 512 | 71.0 | 3046 |
| 768 | 60.1 | 2970 |
| 1024 | 55.5 | 2888 |
| 2048 | 43.1 | 2198 |
Peak RAM is two separate real numbers. Host resident memory is 466 MB; the KV cache is a further 448 MB allocated NPU-side, which does not appear in host RSS. Total device footprint is about 914 MB at context 4096. The cache scales linearly with context, so at context 1024 it needs about 112 MB instead of 448 MB. Reporting host RSS alone undercounts by the size of the cache.
One environment variable is worth 67% of decode*
Without it the backend waits on interrupts for batch completion; with it the backend polls.
| configuration | decode @ empty cache |
|---|---|
| polling enabled | 91.7 tok/s |
Measured at an empty cache; 55.1 tok/s without it, hence the 67%.
* Polling is done on a CPU thread, so one CPU core will sit at 100%. Generally not recommended when building an app.
export LD_LIBRARY_PATH=lib
export ADSP_LIBRARY_PATH=lib # FastRPC finds the v79 DSP skel here
export GGML_HEXAGON_OPPOLL=1 # +67% decode*
export GGML_HEXAGON_HOSTBUF=0 # +8% prefill
./llama-bench --device HTP0 -ngl 99 -fa 1 --mmap 0 \
-t 6 --cpu-mask 0xfc --cpu-strict 1 --ubatch-size 2048 \
-m qwen3-0.6b-a16w4-hexagon.gguf -p 512 -n 128
Three further results from the same sweep:
- Keep operator fusion on โ disabling it costs 8% decode.
- Do not quantize the KV cache. An 8-bit cache collapses prefill by up to 8.3x on these kernels (3448 to 416 tok/s) and is slower at every realistic depth, because the HTP kernels have no optimized quantized-KV path. It costs no measurable accuracy; the throughput penalty is what makes it a bad trade.
- Operator batch and queue depth are already tuned โ raising either changed nothing measurable.
4-bit accuracy
Full MMLU, all 14,042 questions โ the complete test set, not a subsample.
Evaluation protocol. All MMLU figures on this page are 0-shot, scored by loglikelihood (the log-probability of each of "A"/"B"/"C"/"D" as a continuation, highest wins โ nothing is generated), over the complete 14,042-question test set. Quantized and baseline runs use the identical protocol, so the difference between them isolates the quantization.
For reference, the published 0-shot MMLU for the unquantized Qwen3-0.6B is 0.4039, against the 0.4017 we measure for the same model in our own harness โ a 0.0022 difference, well inside one standard error. MMLU scores are only comparable when the prompting protocol matches; figures quoted elsewhere under a different protocol are not comparable to these.
| config | MMLU | ยฑSE | above-chance capability retained* |
|---|---|---|---|
| fp32 parent (0-shot) | 0.4017 | 0.0041 | 100.0% |
| 4-bit (this build) | 0.3924 | 0.0041 | 93.9% |
* (acc - 0.25) / (fp32 - 0.25). MMLU is 4-way multiple choice, so a non-functional model still
scores about 0.25. This column is more informative than a raw percentage of fp32, which flatters
every model by roughly 62 points before it has learned anything.
The 4-bit QAIRT Genie build we previously published retained 50.8% on the same measure. This build retains 93.9%.
KV-cache precision does not affect accuracy here โ 16-bit scores 0.3924 and an 8-bit block-wise cache scores 0.3915, a 0.0009 spread against a standard error of 0.0041. Keep the cache at 16-bit for throughput, not accuracy.
Task dependence, stated plainly. A second 4-bit build scores higher on generative reasoning and lower on MMLU: 0.376 vs 0.334 on 5-shot chain-of-thought GSM8K (500 questions), against 0.3551 vs 0.3924 on MMLU. There is no single "best 4-bit" answer โ pick by workload. Note also that agreement with the parent's next-token distribution does not predict task accuracy here: the build with better top-1 agreement and lower KL divergence is the one that scores worse on MMLU.
4-bit coherence
Verified across multiple prompts at temperature 0.
prompt: What is the capital of France?
output: [thinking] ... France is a country in Europe, and the capital is Paris ...
The capital of France is **Paris**.
Correct, terminates cleanly, no repetition and no language drift. We test more than one prompt deliberately โ a single well-behaved sample is not evidence of coherence, and we have measured builds of this model that answer one prompt cleanly and then degenerate on the next.
A16W8 โ int8 weights, 16-bit KV cache
The higher-accuracy option: int8 weights, 16-bit activations, and a 16-bit KV cache.
Accuracy
Full MMLU, all 14,042 questions, with the complete deployed pipeline modelled โ int8 weights, int16 activations, and the 16-bit KV cache. This is not a weight-only figure.
Same protocol as above: 0-shot, loglikelihood-scored, complete test set.
| config | MMLU | ยฑSE | above-chance capability retained* |
|---|---|---|---|
| fp32 parent (0-shot) | 0.4017 | 0.0041 | 100.0% |
| A16W8, 16-bit KV | 0.4036 | 0.0041 | 101.2% |
* (acc - 0.25) / (fp32 - 0.25).
The difference from fp32 is +0.0019, about a third of one standard error, so it is statistically indistinguishable from zero. The honest phrasing is not "lossless" but "any loss is below what the test set can resolve". The cache dtype has no measurable effect among 16-bit formats: 16-bit integer scores 0.4036, and the two 16-bit float formats score 0.4042 and 0.4044 โ a 0.0013 spread against a standard error of 0.0041.
On-device KPIs
Context 1024, same device and methodology as the 4-bit numbers above.
| KPI | value |
|---|---|
init (GenieDialog_create) |
358 ms (335โ378) |
| prefill | 6428 tok/s |
| decode | 72.0 tok/s (71.3โ72.3) |
| TTFT | ~36 ms (17-token prompt) |
| peak RAM | 754 MiB host resident (753โ755) |
Init, decode and peak RAM are measured directly: six profiled runs of a 17-token prompt producing 88 generated tokens, with resident memory sampled during each run. The decode figure is insensitive to how prefill time is apportioned โ it lands at 71โ74 tok/s across any plausible prefill value.
Prefill is not reported directly by the runtime โ its profile emits only create/query/free โ so it
is solved by regression over four prompt lengths (17, 126, 369, 639 tokens):
query_ms = 15.6 + 0.1556ยทn_prompt + 13.824ยทn_gen. The regression predicts 72.3 tok/s decode
against the 72.0 tok/s measured here, which is the cross-check for the prefill figure resting on it.
A16W8 decode is flat with context depth (13.79โ13.83 ms per token from depth 17 to 782), because its graph is compiled for a fixed context and always pays full padded-attention cost. The 4-bit build attends only over live KV, so it is faster on short contexts and degrades as the cache fills. The two cross at roughly 475 tokens, and over a full 1024-token generation they land within 4% of each other.
A16W8's decisive advantages are prefill (1.75x) and init (2.5x), not decode.
Why the KV cache is 16-bit
An 8-bit KV cache with a single per-tensor scale does not work for this model. It costs most of the model's above-chance capability and makes generation degenerate on device. The cause is scale granularity, not bit width: one static scale cannot cover the K/V dynamic range regardless of how the range is chosen, and clipping to the 99.9th percentile recovers almost nothing. A block-wise 8-bit cache would be fine, but QNN has no block-wise equivalent, so 16-bit is the correct choice here.
Measured cost of the 16-bit cache on device: decode ~72 vs 77.0 tok/s (โ6%) and +56 MB of buffers. Cheap for the accuracy it buys.
Binaries
Published for Hexagon v79 (SM8750), built from the 16-bit-KV checkpoint and verified on device.
| file | size |
|---|---|
qwen3_part1_a16w8_int16kv_v79.bin |
311 MB โ token-embedding graph |
qwen3_part2_a16w8_int16kv_v79.bin |
620 MB โ 28 transformer layers + output projection |
v79/genie_config_a16w8_int16kv.json |
context 1024 |
v79/htp_backend_ext_config.json |
dsp_arch: v79, burst |
v79/metadata.json |
graph I/O shapes and quantization parameters |
SHA256SUMS_v79_int16kv |
checksums |
Each part carries both graph families: prompt_ar128_cl1024_* (prefill, 128-token chunks) and
token_ar1_cl1024_* (autoregressive decode).
The A16W8 binaries that previously sat here carried an 8-bit KV cache and have been removed.
Host libraries and DSP skels must live in separate directories โ the hexagon-v79/unsigned tree
ships a 32-bit libQnnSystem.so that overwrites the aarch64 host copy if both are pushed into one
directory, after which startup fails with a 32-bit/64-bit error.
adb push qwen3_part1_a16w8_int16kv_v79.bin qwen3_part2_a16w8_int16kv_v79.bin /data/local/tmp/qwen/
adb push v79/. host-model/tokenizer.json /data/local/tmp/qwen/
adb push <QAIRT>/lib/aarch64-android/. /data/local/tmp/genie_rt/
adb push <QAIRT>/bin/aarch64-android/genie-t2t-run /data/local/tmp/genie_rt/
adb push <QAIRT>/lib/hexagon-v79/unsigned/. /data/local/tmp/genie_dsp/
adb shell 'cd /data/local/tmp/qwen && \
LD_LIBRARY_PATH=/data/local/tmp/genie_rt ADSP_LIBRARY_PATH=/data/local/tmp/genie_dsp \
/data/local/tmp/genie_rt/genie-t2t-run -c genie_config_a16w8_int16kv.json --prompt_file prompt.txt'
Rename the two .bin files to match the ctx-bins entries in the config, or edit the config to
match. Verify what you downloaded with sha256sum -c SHA256SUMS.
Prompt format (required)
Prefill the opening <think> tag as part of the assistant turn:
<|im_start|>user
{QUESTION}<|im_end|>
<|im_start|>assistant
<think>
This matches Qwen3's own chat template in thinking mode. It matters for this build: if the model is
left to emit <think> itself as the first generated token, a single spurious token can follow it and
derail the response, sometimes switching the reasoning into another language. With <think>
prefilled, generation is coherent across prompts, closes </think> correctly and terminates on EOS.
Prompts that do not use the chat template at all are unaffected.
Verified at temperature 0 on four prompts (capital of France, haiku, arithmetic, gravity):
prompt: What is the capital of France?
output: Okay, the user is asking for the capital of France. ... Just state the capital.
</think>
The capital of France is Paris.
A16W8 for Hexagon v81 (SM8850, 8 Elite Gen 5)
Same 16-bit-KV checkpoint as the verified v79 build, compiled for v81 with the same QAIRT 2.45.
The KV quantization is byte-identical (past_key_0_in scale 0.01514680590480566 in both), so this
is the same numerics on a different Hexagon version.
Measured on v81 silicon (per-graph)
| graph | latency | peak memory |
|---|---|---|
part2 token_ar1_cl1024 (decode) |
15.326 ms | 77.9 โ 88.3 MB |
part2 prompt_ar128_cl1024 (prefill) |
22.531 ms | 64.0 โ 74.9 MB |
part1 token_ar1_cl1024 (embedding) |
0.040 ms | ~10 MB |
Which gives decode ~65.1 tok/s and prefill ~5,681 tok/s (128 tokens per 22.531 ms chunk).
These are per-graph figures and are not comparable like-for-like with the end-to-end v79 numbers above. Running the same per-graph method on v79 yields 65.8 tok/s, against 72.0 tok/s measured end-to-end โ so this method understates by roughly 9%. Applying that correction suggests v81 would land near 71 tok/s in practice, but that is an inference, not a measurement.
The headline result is that v81 is not faster than v79 for this model: 15.326 ms vs 15.2 ms per decode step, within a percent. Decode is DRAM-bandwidth-bound, and a newer Hexagon version does not move that wall.
What is and is not established
Carries over from the v79 verification, because both are properties of the weights and the
quantization scheme: MMLU 0.4036, and the <think> prefill requirement documented above.
Not established: that this binary loads and runs on a real Gen 5 device, its init time, TTFT, host peak RAM, and thermal behaviour. A successful compile is not evidence of a working binary โ a 2.47-compiled build of this same checkpoint compiled cleanly and then failed to initialise on device.
| file | size |
|---|---|
qwen3_part1_a16w8_int16kv_v81.bin |
311 MB |
qwen3_part2_a16w8_int16kv_v81.bin |
622 MB |
v81/genie_config_a16w8_int16kv.json |
context 1024, tuned to match v79 |
v81/htp_backend_ext_config.json |
dsp_arch: v81, burst, rpc_control_latency: 10 |
v81/metadata.json |
graph I/O shapes and quantization parameters |
results_v81_profile.json |
the raw profile behind the table above |
Files
qwen3-0.6b-a16w4-hexagon.gguf 382 MB 4-bit weights / 16-bit activations
qwen3_part1_a16w8_int16kv_v79.bin 311 MB A16W8 embedding graph (v79)
qwen3_part2_a16w8_int16kv_v79.bin 620 MB A16W8 layers + output projection (v79)
qwen3_part1_a16w8_int16kv_v81.bin 311 MB A16W8 embedding graph (v81)
qwen3_part2_a16w8_int16kv_v81.bin 622 MB A16W8 layers + output projection (v81)
v79/ Genie + HTP configs and graph metadata (v79)
v81/ Genie + HTP configs and graph metadata (v81)
results_v81_profile.json raw v81 per-graph profile
host-model/tokenizer.json tokenizer (referenced by the Genie config)
host-model/tokenizer_config.json
llama.cpp/bin/llama-bench prebuilt arm64-android, Hexagon HTP backend
llama.cpp/bin/llama-cli
llama.cpp/lib/libggml-htp-v79.so Hexagon v79 DSP skel
llama.cpp/lib/*.so backend + runtime libraries
runtime/run_qdc_kpis.sh device KPI sweep
runtime/cli_kpis.sh single-model init / TTFT / TPOT / peak-RAM run
runtime/RUN_ON_DEVICE.sh staging helper
SHA256SUMS checksums for every shipped artifact
requirements.txt
The runtime is prebuilt and shipped, so the 4-bit path needs no llama.cpp build. Verify what you
downloaded with sha256sum -c SHA256SUMS.
Deployment
Root is not required; /data/local/tmp is writable and executable, which is sufficient.
adb push qwen3-0.6b-a16w4-hexagon-v79.zip /data/local/tmp/
adb shell 'cd /data/local/tmp && unzip -o qwen3-0.6b-a16w4-hexagon-v79.zip -d qwen-a16w4'
adb shell 'cd /data/local/tmp/qwen-a16w4 && chmod +x bin/* run_kpis.sh && sh run_kpis.sh'
Confirming it is actually running on the NPU
Do not infer NPU execution from the fact that it ran. Check the load log for HTP0:
load_tensors: HTP0 model buffer size = ... MiB
llama_kv_cache: HTP0 KV buffer size = 448.00 MiB
If those read CPU, the model is on the CPU. A 0.6B 4-bit model prefills in the thousands of tokens
per second on the HTP and roughly an order of magnitude slower on the CPU, which is a second,
independent check.
Known gotchas
GGML_HEXAGON_NHVX=8hangs. v79 has six HVX units; leave the variable unset.llama-cliblocks on stdin even with-no-cnv. Pass-stand redirect stdin from/dev/null, or it appears to hang forever and any memory sampled reads as a few MB from a process that never loaded the model.llama-benchis the reliable timing tool.llama-cliandllama-benchdo not take the same flags:--mmap 0and-fa 1work forllama-bench, whilellama-clineeds--no-mmapand-fa on.llama-cliprints no load log or buffer sizes without--verbose, which makes memory look unmeasurable when it is not.- VTCM on v79 is 8 MB, so no model of this size is VTCM-resident. Weight-placement tuning does not help decode on this backend; the polling flag does.
- For the Genie path, host libraries and DSP skels must live in separate directories.
hexagon-v79/unsigned/ships a 32-bit DSPlibQnnSystem.sothat silently overwrites the aarch64 host copy if both are pushed into one directory, and startup then fails with "is 32-bit instead of 64-bit".
Measurement methodology
Every figure was measured on physical silicon. Prefill and decode for the 4-bit build are
llama-bench with 3-5 repetitions and reported standard deviations. TTFT and init come from the
runtime's own timing output. Peak host memory is a resident high-water mark sampled during the run;
NPU-side memory is the runtime's per-backend buffer accounting. Accuracy is the full 14,042-question
MMLU test set, 0-shot and loglikelihood-scored. A16W8 init, decode and peak RAM are measured directly from profiled runs; only A16W8
prefill is derived by regression, because the Genie profile does not separate the phases.
Accuracy is a property of the weights, so it does not change with Hexagon architecture version. Only speed and memory do.
- Downloads last month
- -
We're not able to determine the quantization variants.