โš ๏ธ STOCK llama.cpp WILL NOT LOAD THIS MODEL

7.72 GiB ยท 88.82 tok/s on a Ryzen AI MAX+ 395.

Ling-3.0-tiny โ€” ROCmFPX 8-bit AGENT GGUF

An 8-bit ROCmFPX quantization for AMD gfx1151 (Ryzen AI MAX+ 395 / Strix Halo), quantized from BF16 GGUF โ€” a lossless source, not a requantization of a lower-bit build.

File Ling-3.0-tiny-Q8_0_ROCMFPX_AGENT.gguf
Size 7.72 GiB
BPW 8.40
ftype Q8_0_ROCMFPX_AGENT (115)

Requires the Q-LoRA bailingmoe3 arch port (q_lora_rank=256). Our port is in patches/; without it no GGUF of this model loads at all.

โ›” Requires a llama.cpp with the ROCmFPX quant types

Q8_0_ROCMFPX (ftype 111) and Q8_0_ROCMFPX_AGENT (ftype 115) exist only in charlie12345/ROCmFPX, not upstream llama.cpp. Stock llama.cpp reports invalid ggml type 103. Ignore the auto-generated "Use this model" commands above.


All quant variants

Three builds of this model, all measured in one session on one box with one binary (Ryzen AI MAX+ 395, gfx1151, ROCm 7.2.4, ROCmFPX-2809dc5) โ€” so these rows are directly comparable. Median of 3, warm-up discarded, otherwise-idle box.

variant ftype size bpw decode (median) range repo
4-bit COHERENT 102 4.30 GiB 4.67 104.04 104.00 โ€“ 104.24 Ling-3.0-tiny-ROCmFP4-GGUF
8-bit AGENT 115 7.72 GiB 8.40 88.82 88.80 โ€“ 88.83 Ling-3.0-tiny-ROCmFPX-Q8_0-AGENT-GGUF
8-bit plain 111 7.62 GiB 8.28 89.51 89.48 โ€“ 89.51 Ling-3.0-tiny-ROCmFPX-Q8_0-GGUF

โš ๏ธ The 4-bit build is faster (104.04 vs ~89 tok/s) and 44% smaller. These 8-bit builds exist for accuracy headroom, not speed โ€” pick them only if you need the extra precision.

What AGENT actually changes: it keeps far more tensors at true Q8_0 instead of the packed 8-bit type โ€” measured in these files, 135 tensors vs 2 tensors. On models with an MTP draft head that raises draft acceptance and wins ~6%; these two models have no MTP head, and here the two 8-bit builds are within noise of each other.

Correctness: 3/3 clean on all three builds (391 ยท Tokyo ยท 366).

Per-tensor types (audited in this finished file)

output.weight Q8_0 ยท token_embd Q8_0 ยท 176 packed TYPE_103 ยท 215 F32 ยท 135 Q8_0 ยท attn_v 14 of 24 at Q8_0


What was NOT measured

  • No perplexity run, and no quality A/B against the source. The checks above are memorized-fact prompts โ€” necessary but not sufficient; a damaged model can pass them.
  • No long-context testing. ยท No tool-calling evaluation.

Base model licence inherited; credit for the model goes to its authors.

Downloads last month
93
GGUF
Model size
8B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for kingjones777/Ling-3.0-tiny-ROCmFPX-Q8_0-AGENT-GGUF

Quantized
(14)
this model