mAPEX Quantized Models

Source Data

The source data for quantization was taken from deepgrove/maple-preview.

Original source file: model-00001-of-00009.safetensors

Quantization Method

All quants were produced using the modified APEX quantization scheme.

mAPEX (modified Automated Precision EXpert allocation) assigns per-layer, per-tensor precision for MoE models using llama.cpp's --tensor-type-file.

For more information, see: https://github.com/DrMoriarty/apex-quant/

Quantized Models

Note: TierN-s quants are typically slightly larger than their regular counterparts, but provide roughly 10โ€“20% faster inference (highly dependent on the model and GPU).

Name Size (GB) Comments
maple-preview-Tier7.gguf 11.59
maple-preview-Tier10.gguf 9.13
maple-preview-Tier10-s.gguf 9.13
Downloads last month
709
GGUF
Model size
20B params
Architecture
maple
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for DrMoriarty0/maple-preview

Quantized
(15)
this model