⛔ THIS BUILD DOES NOT FIT ON A 128 GB STRIX HALO

llama.cpp reports 113.03 GiB addressable on a Ryzen AI MAX+ 395. These 8-bit builds are 114.38 GiB and 116.15 GiB. Attempting -ngl 999 hard-wedges the machine — we did it twice: a KFD SVM D-state livelock (svm_range_cpu_invalidate_pagetables) that survives a GPU reset and needs a power cycle. -fit off does not save you; it only stops llama.cpp from shrinking the model, so it allocates until the driver dies.

On a single 128 GB Strix Halo, use the 4-bit build (63.07 GiB, 37.86 tok/s). These 8-bit builds are for machines with more memory, or for CPU / partial-offload inference.

Mistral-Small-4-119B-A6.5B — ROCmFPX 8-bit GGUF

An 8-bit ROCmFPX quantization built from BF16 (222 GiB) — a lossless source, not a requantization of a lower-bit build. 119B total / 6.5B active MoE.

File Mistral-Small-4-119B-2603-Q8_0_ROCMFPX.gguf
Size 114.38 GiB
BPW 8.26
ftype Q8_0_ROCMFPX (111)
Tensors 579

Built with --output-tensor-type q8_0 --token-embedding-type q8_0 --tensor-type shexp=q8_0. The shexp override matched 108 shared-expert tensors (confirmed in the dry-run receipt — a --tensor-type pattern that matches nothing is a silent no-op, so we check the count).

⛔ Requires a llama.cpp with the ROCmFPX quant types

Q8_0_ROCMFPX (111) / Q8_0_ROCMFPX_AGENT (115) exist only in charlie12345/ROCmFPX. Stock llama.cpp reports invalid ggml type 103. Ignore the auto-generated "Use this model" commands above.


All quant variants

variant ftype size bpw GPU on 128 GB Strix Halo decode
4-bit COHERENT 102 63.07 GiB 4.55 ✅ fits 37.86 tok/s
8-bit AGENT 115 116.15 GiB 8.39 ⛔ does not fit not measurable on this box
8-bit plain 111 114.38 GiB 8.26 ⛔ does not fit not measurable on this box

Repos: 4-bit · 8-bit AGENT · 8-bit plain

Verified

  • Structural: GGUF v3, 579 tensors, 59 KV pairs — matches the source topology.
  • Correctness (CPU-only, -ngl 0): 17×23 ⇒ ✅ 391 · capital of Japan ⇒ ✅ Tokyo. Loaded in 256 s from disk.

What was NOT measured

  • No GPU decode benchmark. The build does not fit in addressable GPU memory on our hardware, so we have no tok/s figure for it and do not quote one.
  • No perplexity run, no quality A/B against the source, no long-context or tool-calling tests.

Base model licence inherited; credit for the model goes to its authors.

Downloads last month
74
GGUF
Model size
119B params
Architecture
mistral4
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kingjones777/Mistral-Small-4-119B-ROCmFPX-Q8_0-GGUF

Quantized
(34)
this model