apex-flash-1-heretic-FP8

FP8 (e4m3, 128x128 block scales) quantization of pqhaz/apex-flash-1-heretic, in the same tensor layout and quantization_config as the official zai-org/GLM-5.3-Flash release. Produced with convert.py (included), the same script as pqhaz/apex-flash-1-abliterated-FP8.

With reasoning enabled the model still refuses about four in ten requests of the evaluation set (source: 99/100). See the BF16 repository for what was changed and how it was measured; ablation.json holds the exact parameter set. The numbers there were measured on the BF16 weights, not on this quantization.

Intended for authorized security research in environments the researcher owns or has permission to test.

Downloads last month
4
Safetensors
Model size
321B params
Tensor type
BF16
·
F8_E4M3
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pqhaz/apex-flash-1-heretic-FP8