ms-marco-MiniLM-L6-v2, fused fp16 CUDA graph for CodeSage
These are ONNX graphs of cross-encoder/ms-marco-MiniLM-L6-v2 at revision c5ee24cb16019beea0893ab7796b1df96625c6b8, packaged for the CodeSage reranker. The weights are unchanged apart from precision. All credit for the model goes to its authors; the license is the upstream Apache-2.0.
Files
| File | Contents |
|---|---|
onnx/model.onnx |
upstream graph, copied unchanged |
onnx/model_cuda_fp16.onnx |
ONNX Runtime BERT optimizer fusions (EmbedLayerNormalization, 6 × Attention, SkipLayerNormalization, BiasGelu), fp16 weights, int64 inputs, fp32 logits |
tokenizer.json |
upstream, copied unchanged |
model_cuda_fp16.onnx targets the ONNX Runtime CUDA execution provider. It also loads on the CPU provider.
Measured effect
Measured with ONNX Runtime 1.24.4 on an RTX 4080 Laptop GPU, scoring 50 (query, code chunk) pairs per query in batches of 32, over 40 queries.
| Graph | Median latency per query | Max logit change | Top-10 overlap | Same top-1 |
|---|---|---|---|---|
| upstream fp32 | 108.0 ms | |||
model_cuda_fp16.onnx |
20.6 ms | 0.024 | 0.9975 | 40/40 |
Reproduce
The script scripts/derive-reranker-onnx.py in the CodeSage repository downloads the pinned upstream files, verifies their sha256, and writes these files byte-for-byte. It was run with onnx 1.21.0 and onnxruntime 1.24.4.
tokenizer.json d241a60d5e8f04cc1b2b3e9ef7a4921b27bf526d9f6050ab90f9267a1f9e5c66
onnx/model.onnx 5d3e70fd0c9ff14b9b5169a51e957b7a9c74897afd0a35ce4bd318150c1d4d4a
onnx/model_cuda_fp16.onnx 32eef6d63e978aba96b6c4134b93b252eab2eb6ae3f8b1b5d15c96b4473f7eb2
Model tree for IA0x00/ms-marco-MiniLM-L6-v2-codesage
Base model
microsoft/MiniLM-L12-H384-uncased