Ling 3.0 Tiny APEX Quant (and a TQ1_0 version)
A compressed mixed-precision MoE quantization of Ling 3.0 Tiny.
Credits & Sources
- Base model: inclusionAI/Ling-3.0-tiny
- Importance matrix: SC117/Ling-3.0-tiny-abliterated-APEX-GGUF
- Quantization: localai-org/apex-quant
- Quantization backend: llama.cpp
The quantization uses a mixed-precision approach rather than applying a single quantization format to every tensor. Different parts of the MoE are assigned different precisions to preserve important model capabilities while aggressively reducing the size of the large expert weights. (Thanks APEX)
It is ~2.9GB so it will fit in most small GPUs and genorate at a moderately good speed.
Ling-3.0-tiny-TQ1_0.gguf was made using normal llama.cpp tooling and is very dumb and not recomended. It is 1.8GB.
This is an unofficial community quantization and is not affiliated with or endorsed by InclusionAI.
- Downloads last month
- 70
Hardware compatibility
Log In to add your hardware
1-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support