Nex-N2.5-mini-INT4-W4A16

This model is a quantized version of nex-agi/Nex-N2.5-mini.

Quantization details

  • Method: GPTQ
  • Scheme: W4A16
  • Targets: Linear
  • Group size: 128
  • Calibration dataset: mlabonne/open-perfectblend
  • Calibration samples: 256
  • Maximum sequence length: 2048

The following modules were excluded from quantization:

  • lm_head
  • Embeddings
  • Vision modules
  • Linear attention modules
  • MLP gates
  • Shared expert gates

Sampling parameters

The default sampling parameters are set to the ones recommended in Nex-AGI's original model card:

  • temperature: 0.7
  • top_p: 0.95
  • top_k: 40
Downloads last month
57
Safetensors
Model size
7B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Arahide/Nex-N2.5-mini-INT4-W4A16

Quantized
(39)
this model