qwen3.5-9b-gguf (Enclave model volume)

Curated model volume for Enclave confidential inference: Qwen3.5-9B at unsloth Dynamic Q4_K_XL (~6 GB), a single-file GGUF bundled with the official tokenizer.json so the volume is self-contained.

Upstream unsloth/Qwen3.5-9B-GGUF ships 22 quantizations plus a BF16 export and vision mmproj-* files (~160 GB total) and no tokenizer.json. This volume carries ONLY the one served quant plus the tokenizer (same rationale as qwen2.5-0.5b-gguf / qwen3.5-122b-gguf): a multi-gguf volume makes the host's preload pick ambiguous, and wrapping the whole repo wastes ~154 GB of attested volume for files that never serve. The mmproj files are omitted deliberately โ€” llm-chat's ggml path is text-only and preloads a single LLM gguf.

Provenance

File Upstream Revision sha256
Qwen3.5-9B-UD-Q4_K_XL.gguf unsloth/Qwen3.5-9B-GGUF 3885219b6810b007914f3a7950a8d1b469d598a5 6f5d30666c2d8ae16a306e616d95341dcf3cc46810df84d7e6f5a7d1e4c1b293
tokenizer.json Qwen/Qwen3.5-9B c202236235762e1c871ad0ccb60c8ee5ba337b9a 5f9e4d4901a92b997e463c1f46055088b6cca5ca61a6522d1b9f64c4bb81cb42

The tokenizer is byte-identical to the one in qwen3.5-122b-gguf (the Qwen3.5 family shares it โ€” same sha256, same vocab 248320, <|im_end|> 248046 / <|endoftext|> 248044).

LICENSE is the Apache-2.0 text; both upstreams are Apache-2.0.

Usage on Enclave

Wrapped as a Tinfoil Modelwrap volume (dm-verity; the root hash is part of the enclave measurement). Deployments attach it by name and the guest reads it at /models/<name>; the host preloads the GGUF as the wasi-nn ggml graph. As the volume's single gguf it needs no MODEL_VOLUMES third field and no model_file override, and with the tokenizer bundled, llm-chat's default tokenizer.json lookup works with no cross-volume configuration.

Downloads last month
89
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for EnclaveHost/qwen3.5-9b-gguf

Finetuned
Qwen/Qwen3.5-9B
Quantized
(409)
this model