Nex-N2.5-mini-INT4-W4A16
This model is a quantized version of nex-agi/Nex-N2.5-mini.
Quantization details
- Method: GPTQ
- Scheme: W4A16
- Targets: Linear
- Group size: 128
- Calibration dataset: mlabonne/open-perfectblend
- Calibration samples: 256
- Maximum sequence length: 2048
The following modules were excluded from quantization:
lm_head- Embeddings
- Vision modules
- Linear attention modules
- MLP gates
- Shared expert gates
Sampling parameters
The default sampling parameters are set to the ones recommended in Nex-AGI's original model card:
temperature: 0.7top_p: 0.95top_k: 40
- Downloads last month
- 57
Model tree for Arahide/Nex-N2.5-mini-INT4-W4A16
Base model
nex-agi/Nex-N2.5-mini