jina-embeddings-v2-base-code, inference-optimized ONNX graphs for CodeSage
These are derived ONNX graphs of jinaai/jina-embeddings-v2-base-code at revision 516f4baf13dec4ddddda8631e019b5737c8bc250, built for the CodeSage indexer. The weights are unchanged apart from precision. All credit for the model goes to Jina AI; the license is the upstream Apache-2.0.
Files
| File | Precision | Attention | Output |
|---|---|---|---|
onnx/model.onnx |
fp32 | upstream (decomposed) | sentence_embedding [batch, 768], mean-pooled and L2-normalized |
onnx/model_cuda_fp16.onnx |
fp16 weights, int64/fp32 I/O | com.microsoft MultiHeadAttention with ALiBi as attention_bias |
same |
tokenizer.json |
copied unchanged from upstream |
Inputs for both graphs are input_ids and attention_mask (int64, [batch, seq]). Mean pooling over the attention mask and L2 normalization are part of the graph, so you don't pool the output yourself.
model_cuda_fp16.onnx targets the ONNX Runtime CUDA execution provider. Pad sequences to a multiple of 8 tokens so the memory-efficient attention kernel accepts the bias.
Measured effect
Measured with ONNX Runtime 1.24.4 on an RTX 4080 Laptop GPU, using 2,048 real chunks of up to 512 tokens. Cosine and neighbour overlap are against the upstream fp32 graph with host-side mean pooling.
| Corpus | Upstream fp32 | model.onnx |
model_cuda_fp16.onnx |
min cosine (fp16) | top-10 neighbour overlap (fp16) |
|---|---|---|---|---|---|
| php-src (C) | 59.7 chunks/s | 89 | 249โ257 | 0.99999 | 0.994 |
| Laravel app (PHP) | 67.2 | 113 | 301โ317 | 0.999996 | 0.996 |
| React app (TS/JS) | 69.8 | 109 | 296โ317 | 0.999995 | 0.997 |
model.onnx matches the upstream vectors to within float rounding (min cosine 0.999999).
Reproduce
The script scripts/derive-jina-onnx.py in the CodeSage repository downloads the pinned upstream files, verifies their sha256, and writes these files byte-for-byte. It was run with onnx 1.21.0 and onnxruntime 1.24.4.
tokenizer.json b01c78a902aa4facb2f47f95449f48e2f7bbfea5d2472ee2f6ce92323c6f86e5
onnx/model.onnx 98ca3fc0def59e9fea861ca8888f805fd3cd10d429695ff53bd03ab053e8aee4
onnx/model_cuda_fp16.onnx 5d34291e0d44a01924d6709c631b141a875a7e9add7f118dde87050644cfc2ab
Model tree for IA0x00/jina-embeddings-v2-base-code-codesage
Base model
jinaai/jina-embeddings-v2-base-code