⚠️ STOCK llama.cpp WILL NOT LOAD THIS MODEL

8.59 GiB · 21.08 tok/s on a Ryzen AI MAX+ 395.

ZAYA1-8B — ROCmFPX 8-bit GGUF

An 8-bit ROCmFPX quantization for AMD gfx1151 (Ryzen AI MAX+ 395 / Strix Halo), quantized from BF16 GGUF — a lossless source, not a requantization of a lower-bit build.

File ZAYA1-8B-Q8_0_ROCMFPX.gguf
Size 8.59 GiB
BPW 8.32
ftype Q8_0_ROCMFPX (111)

tie_word_embeddings is TRUE, so output.weight does not exist — --output-tensor-type is a silent no-op here and --token-embedding-type is the flag that lands (262K vocab).

⛔ Requires a llama.cpp with the ROCmFPX quant types

Q8_0_ROCMFPX (ftype 111) and Q8_0_ROCMFPX_AGENT (ftype 115) exist only in charlie12345/ROCmFPX, not upstream llama.cpp. Stock llama.cpp reports invalid ggml type 103. Ignore the auto-generated "Use this model" commands above.


All quant variants

Three builds of this model, all measured in one session on one box with one binary (Ryzen AI MAX+ 395, gfx1151, ROCm 7.2.4, ROCmFPX-2809dc5) — so these rows are directly comparable. Median of 3, warm-up discarded, otherwise-idle box.

variant ftype size bpw decode (median) range repo
4-bit COHERENT 102 4.86 GiB 4.71 23.04 22.83 – 23.70 ZAYA1-8B-ROCmFP4-GGUF
8-bit AGENT 115 8.72 GiB 8.45 21.02 20.95 – 21.47 ZAYA1-8B-ROCmFPX-Q8_0-AGENT-GGUF
8-bit plain 111 8.59 GiB 8.32 21.08 20.99 – 21.20 ZAYA1-8B-ROCmFPX-Q8_0-GGUF

⚠️ Decode is ~88% weight-independent on this architecture (the CCA grouped conv is ~55% of decode). All three builds land within ~10% of each other; the 4-bit is smallest and marginally fastest. No 8-bit or 4-bit format will make this model meaningfully faster.

What AGENT actually changes: it keeps far more tensors at true Q8_0 instead of the packed 8-bit type — measured in these files, 154 tensors vs 1 tensor. On models with an MTP draft head that raises draft acceptance and wins ~6%; these two models have no MTP head, and here the two 8-bit builds are within noise of each other.

Correctness: All three builds answer correctly. On some prompts content is empty with finish_reason=length and the correct answer sits in reasoning_content — this model is verbose, give it ≥1024 tokens.

Per-tensor types (audited in this finished file)

token_embd Q8_0 · 400 packed TYPE_103 · 842 F32 · 40 BF16 (CCA conv kept at BF16) · 1 Q8_0


What was NOT measured

  • No perplexity run, and no quality A/B against the source. The checks above are memorized-fact prompts — necessary but not sufficient; a damaged model can pass them.
  • No long-context testing. · No tool-calling evaluation.

Base model licence inherited; credit for the model goes to its authors.

Downloads last month
115
GGUF
Model size
9B params
Architecture
zaya
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kingjones777/ZAYA1-8B-ROCmFPX-Q8_0-GGUF

Quantized
(2)
this model