GGUF
llama.cpp
quantization
imatrix
conversational

Apertus-v1.1-4B-Instruct-GGUF

This repository contains unofficial llama.cpp GGUF quantizations of swiss-ai/Apertus-v1.1-4B-Instruct, a distilled multilingual language model.

Quantized Files

File Name Description
Q4_K_M.gguf Balanced 4-bit quantization offering a strong trade-off between performance and memory usage.
Q5_K_M.gguf Higher-fidelity 5-bit quantization for improved output quality.
Q8_0.gguf Near-lossless 8-bit quantization.
imatrix.dat Importance matrix calibration data file generated during the quantization process.

Calibration & Multilingual Preservation

These quantizations were built using an importance matrix (imatrix) generated from the eaddario/imatrix-calibration dataset (combined_all_medium). This dataset was specifically selected to safeguard and maintain the model's robust multilingual capabilities across lower-bit quants.

License

This project is licensed under the Apache 2.0 License, matching the licensing terms of the original model.

Downloads last month
551
GGUF
Model size
4B params
Architecture
apertus
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for agentlans/Apertus-v1.1-4B-Instruct-GGUF

Quantized
(15)
this model

Dataset used to train agentlans/Apertus-v1.1-4B-Instruct-GGUF