YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Qwen3.8-27B - MXFP4 (imatrix)

4.25-bit MXFP4 quantization of ggml-org/Qwen3.8-27B, quantized with imatrix.

Files:

  • 27b-mxfp4-imx.gguf - MXFP4 weights (4.25-bit, block-scale e8m0), OCP scale search with imatrix weighting
  • 27b.imatrix - the imatrix file (wiki train data) used for the weighted scale search

Quantization recipe (llama.cpp):

llama-imatrix -m base.gguf -f train.txt -o 27b.imatrix
llama-quantize --imatrix 27b.imatrix base.gguf 27b-mxfp4-imx.gguf mx

Scale settings: weights use the OCP scale search (target 4.0); the MXFP4 KV cache path uses UOS scaling (e = ceil(log2(amax/7.25))+127, MXAttention) as default.

Related: https://github.com/timlikesai/llama.cpp/pull/14

Downloads last month
187
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support