multilingual-e5-large-instruct-fp8

FP8_DYNAMIC quantization of intfloat/multilingual-e5-large-instruct, produced with llm-compressor 0.13.0 and measured end to end on NVIDIA H100 80 GB with vLLM 0.26.0.

Only the Linear layers of the encoder are quantized; the vocabulary, position and token-type embeddings, the pooler and any classification head stay in bf16.

Measured

variant size SciFact nDCG@10 STS-B de ρ STS17 en-de ρ texts/s (bs=256)
bf16 source 1068 MiB 0.7204 0.8347 0.8525 1488
FP8 802 MiB 0.7164 0.8349 0.8526 1511
NVFP4 675 MiB 0.7036 0.8371 0.8440

Retrieval and STS are unchanged against bf16. Size is −25%. Throughput is unchanged — the gain is memory only. If you want the smaller checkpoint, the NVFP4 sibling is 675 MiB for 1.7 pp nDCG@10.

SciFact: 300 queries over 5183 documents, nDCG@10. STS-B de: mteb/stsb_multi_mt (1379 pairs). STS17 en-de: mteb/sts17-crosslingual-sts (250 pairs). Reranking: 1 positive + 19 sampled negatives per query. Throughput over 4096 multilingual texts through the vLLM OpenAI endpoint.

Serving

vllm serve mbehr90/multilingual-e5-large-instruct-fp8 --runner pooling
Downloads last month
-
Safetensors
Model size
0.6B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mbehr90/multilingual-e5-large-instruct-fp8

Quantized
(263)
this model