jina-embeddings-v2-base-code, inference-optimized ONNX graphs for CodeSage

These are derived ONNX graphs of jinaai/jina-embeddings-v2-base-code at revision 516f4baf13dec4ddddda8631e019b5737c8bc250, built for the CodeSage indexer. The weights are unchanged apart from precision. All credit for the model goes to Jina AI; the license is the upstream Apache-2.0.

Files

File Precision Attention Output
onnx/model.onnx fp32 upstream (decomposed) sentence_embedding [batch, 768], mean-pooled and L2-normalized
onnx/model_cuda_fp16.onnx fp16 weights, int64/fp32 I/O com.microsoft MultiHeadAttention with ALiBi as attention_bias same
tokenizer.json copied unchanged from upstream

Inputs for both graphs are input_ids and attention_mask (int64, [batch, seq]). Mean pooling over the attention mask and L2 normalization are part of the graph, so you don't pool the output yourself.

model_cuda_fp16.onnx targets the ONNX Runtime CUDA execution provider. Pad sequences to a multiple of 8 tokens so the memory-efficient attention kernel accepts the bias.

Measured effect

Measured with ONNX Runtime 1.24.4 on an RTX 4080 Laptop GPU, using 2,048 real chunks of up to 512 tokens. Cosine and neighbour overlap are against the upstream fp32 graph with host-side mean pooling.

Corpus Upstream fp32 model.onnx model_cuda_fp16.onnx min cosine (fp16) top-10 neighbour overlap (fp16)
php-src (C) 59.7 chunks/s 89 249โ€“257 0.99999 0.994
Laravel app (PHP) 67.2 113 301โ€“317 0.999996 0.996
React app (TS/JS) 69.8 109 296โ€“317 0.999995 0.997

model.onnx matches the upstream vectors to within float rounding (min cosine 0.999999).

Reproduce

The script scripts/derive-jina-onnx.py in the CodeSage repository downloads the pinned upstream files, verifies their sha256, and writes these files byte-for-byte. It was run with onnx 1.21.0 and onnxruntime 1.24.4.

tokenizer.json             b01c78a902aa4facb2f47f95449f48e2f7bbfea5d2472ee2f6ce92323c6f86e5
onnx/model.onnx            98ca3fc0def59e9fea861ca8888f805fd3cd10d429695ff53bd03ab053e8aee4
onnx/model_cuda_fp16.onnx  5d34291e0d44a01924d6709c631b141a875a7e9add7f118dde87050644cfc2ab
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for IA0x00/jina-embeddings-v2-base-code-codesage

Quantized
(11)
this model