Laya, 8-bit ONNX

An 8-bit weight-only quantization of convaiinnovations/laya (English checkpoint, Apache-2.0), made from the fp32 ONNX export at receptron/laya-onnx. Same inputs and outputs, so @receptron/laya loads it with Laya.load({ modelDir }).

fp32 export this bundle
laya.onnx size 1,685 MB 633 MB
RAM (Linux, loaded) ~2.8 GB ~0.9 GB

Method: ONNX Runtime MatMulNBits: 8-bit, block size 32, symmetric, accuracy_level=4 (int8 compute). The word-embedding table stays fp32.

Check: on 22 hand-labeled brand descriptions from an ad-objective recommender (one Choice with 3-8 options plus a Score and a Noul per request), answers matched the fp32 model on 21 of 22, and no probability moved by more than 0.014. The one change was a 28% vs 27% near-tie. Dynamic INT8 quantization (quantize_dynamic) was tried first and rejected: it changed 6 of 22 answers.

laya.onnx sha256: 42b043dfd9ad0ceacae1629a7f030cae6a2d2c3c235e5e1ba87b4ee84a854e9a

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bgoldfarb93/laya-int8

Quantized
(56)
this model