Queryn adapter β€” fastembed-bge-small β†’ qwen3-emb-8b

Translates an embedding produced by fastembed-bge-small into the embedding space of qwen3-emb-8b, so a corpus already embedded with fastembed-bge-small can be served against a qwen3-emb-8b index without re-embedding it. Part of the Queryn embedding-translation engine.

Specs

Source model fastembed-bge-small (384-d)
Target model qwen3-emb-8b (4096-d)
Architecture linear (plain linear projection)
Parameters ~1.6M
Best test cosine similarity 0.7007 (epoch 15)
ONNX opset 17

Architecture ablation (best test cosine): linear 0.7007 ← saved, deep 0.6972.

Input / output contract

  • Input source_embedding β€” float32, shape [batch, 384]. Raw fastembed-bge-small embeddings; the graph L2-normalizes them itself, so pre-normalization is neither required nor harmful.
  • Output target_embedding β€” float32, shape [batch, 4096], unit-normalized, in qwen3-emb-8b space.
  • Batch axis is dynamic.

Usage

import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download

path = hf_hub_download("QuerynAi/queryn-adapter-fastembed-bge-small_to_qwen3-emb-8b", "model.onnx")
sess = ort.InferenceSession(path, providers=["CPUExecutionProvider"])

src = np.random.rand(4, 384).astype(np.float32)   # your fastembed-bge-small embeddings
tgt = sess.run(["target_embedding"], {"source_embedding": src})[0]
assert tgt.shape == (4, 4096)                     # unit vectors in qwen3-emb-8b space

Training

Trained on paired embeddings over a unified multi-domain corpus β€” arXiv abstracts, Australian case law, SQuAD passages, PubMed abstracts, and crypto/markets news (~350k rows spanning science, legal, QA, medical, and finance). Loss: 1 - mean cosine similarity, Adam, ReduceLROnPlateau, best-epoch checkpoint. Both a linear baseline and the MLP are trained for every pair; the higher-scoring one is published (ties go to linear).

Plots

fastembed-bge-small

fastembed-bge-small β†’ all targets: learning curves and best scores (this pair included).

architecture_ablation

Linear vs. deep for every pair (black ring = saved architecture).

Full adapter set: Queryn Embedding Adapters

Provenance

  • Source checkpoint: models/v1/fastembed-bge-small_to_qwen3-emb-8b.pt (sha256 ed50d7edc0209b90…)
  • Converted: 2026-08-30T20:45:22+00:00 Β· torch 2.13.0 Β· ptConverter.py

License

Released under the MIT license.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including QuerynAi/queryn-adapter-fastembed-bge-small_to_qwen3-emb-8b