motherlode-code-small-en-v0.1

This is a 33.4M-parameter English embedding model for code and text. It has the same architecture, tokenizer and 384-dimensional CLS output as bge-small-en-v1.5, so it drops in where bge-small runs today. It is a parameter-wise mean of two fine-tunes of bge-small:

  • one trained to find the code a GitHub issue needs changed;
  • one trained to reproduce the geometry of gte-modernbert-base, a model 4.5 times its size.
finding the file an issue needs, SWE-bench Verified, recall@1, the 411 instances the contamination check leaves
  gate 1's pipeline                       178  against bge-small's 145   61 wins, 28 losses, sign p 0.00061
  the Black Window engine's published     230  against bge-small's 192   67 wins, 29 losses, sign p 0.00013
  path, each model with its shipped mean
code search, CoIR nDCG@10, a probe        0.5758  against bge-small's 0.4673, on 144,461 queries

Each figure is measured against bge-small on the same items, paired, and each has its population, its artifact and a row in our claims ledger. The per-instance files are in RiverRider/motherlode-code-small-en-v0.1-evidence.

Use

sentence-transformers, 5.1.0 and 6.0.1 tested:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("RiverRider/motherlode-code-small-en-v0.1")
docs = model.encode(passages, normalize_embeddings=True)
query = model.encode(["where is the retry limit read from settings?"], normalize_embeddings=True)
scores = docs @ query[0]

No query or passage prompt is used, as with bge-small in these measurements.

transformers.js 3.7.1, the call the Black Window engine makes:

import { pipeline } from "@huggingface/transformers";
const embed = await pipeline("feature-extraction", "RiverRider/motherlode-code-small-en-v0.1", { dtype: "fp32" });
const out = await embed(texts, { pooling: "cls", normalize: true });

transformers.js 3.7.1 truncates a text longer than 512 tokens after adding the special tokens, so the final [SEP] is lost. Python keeps it. On the 13 of 200 test texts that long, the two libraries' vectors agree at a smallest cosine of 0.99640. bge-small-en-v1.5 behaves the same way through the same library, at 0.99715. The Black Window engine therefore tokenizes without special tokens, cuts at 510, adds [CLS] and [SEP] itself and normalises the CLS row (embedKeepSep in its index.html). That call matches PyTorch at a smallest cosine of 0.99999988 on all 200 texts, for this model and for bge-small.

llama.cpp. gguf/motherlode-code-small-en-v0.1-f16.gguf was converted by the converter in llama.cpp commit 465e49b9c. Through that llama.cpp, with CLS pooling, normalised, at 256 tokens with [SEP] kept, it matches PyTorch float32 at a smallest cosine of 0.9999957 on 2,000 texts, against 0.9999966 for CompendiumLabs/bge-small-en-v1.5-gguf at f16 against bge-small the same way.

Centring. Three 384-float means ship under centring/, each as {"mu": [...]}:

  • mean_shipped.json, the one the Black Window engine centres a store on: the average of the code and caption means below and a prose mean over the first 20,000 paragraphs of at least 60 characters in wikitext-103's train split;
  • mean_code_pilot64k.json, fitted on 64,000 passages from 86 Python repositories;
  • mean_captions_coco20k.json, fitted on 20,000 COCO train2017 captions.

Subtract the mean from both sides before taking a cosine when a score's magnitude matters: thresholds, gates, or a comparison across queries. The mean pairwise raw cosine over 20,000 code passages is 0.2192 for this model and 0.6233 for bge-small. After the shipped mean it is 0.1314 over the code passages, 0.2462 over scifact's abstracts, 0.1345 over COCO captions and 0.0565 over the wikitext paragraphs. A mean fitted on one domain under-corrects another: over 5,000 COCO captions the mean pairwise cosine is 0.2865 after the code mean and 0.0286 after the caption mean. For ranking, centring is not a gain here:

  • on SWE-bench inside the engine with the code mean, ranking the head by raw cosine instead of the centred score reads 274 against 269 of 500 (24 wins, 19 losses, sign p 0.54);
  • on scifact, centring with the shipped mean costs 0.0057 nDCG@10 (0.7150 raw, 0.7093 centred).

Results

Every comparison is paired by item. Sign tests are exact and two-sided, over the items only one model gets right.

Finding the file. On SWE-bench Verified, the task is to rank the repository's files for an issue, and a hit is the gold file first.

setting population this model bge-small-en-v1.5 paired
gate 1's base pipeline, float32, 512 tokens the 411 instances the contamination check leaves 178 145 61 wins, 28 losses, +0.0803, sign p 0.00061
the same all 500 219 181 73 wins, 35 losses, sign p 0.00033
a replay of the Black Window engine's published path, int4 rows, each model with the mean the engine ships for it the 411 230 192 67 wins, 29 losses, +0.0925, sign p 0.00013
the same all 500 278 230 82 wins, 34 losses, sign p 0.00001
the same path, each model with its own code mean the 411 221 180 70 wins, 29 losses, +0.0998, sign p 0.00005
the engine at 8234be7, whose lexical bonus lowers both models, each with its shipped mean the 411 207 158 69 wins, 20 losses, +0.1192, sign p below 0.00001

The published path is engine 94a8e7b with Sunstone 4b1bfe9, replayed with the engine's own chunkers. The same pipeline as gate 1 on all 500 gives these peers:

  • granite-embedding-small-english-r2: 178;
  • gte-modernbert-base: 181;
  • the issue-trained parent, a2: 216.

The replay reproduces the engine's published run at rank 1 on 495 of 500 instances (230 against 229 for bge-small). A permuted-query floor places 5 of 500 for this model and 4 for bge-small.

Searching code, as a probe. On CoIR, nDCG@10 is taken at 512 tokens in float16 as the mean of ten dataset scores. One parent trained on CoIR's training partitions, so these are probe figures and not benchmark entries.

population this model bge-small-en-v1.5 granite-embedding-small-english-r2 the distilled parent, d1
144,461 test queries neither parent's contamination check flagged 0.5758 0.4673 0.5301 0.6160
all 162,216 test queries 0.5612 0.4531 0.5172 0.6032

Prose. These are scored as the engine scores: int4 rows around each model's mean, ranked by the centred dot. The comparison is paired, with 2,000 bootstrap resamples.

task population this model, shipped mean bge-small, shipped mean gap, 95% interval
BeIR scifact, nDCG@10 300 test queries, 5,183 abstracts 0.7093 0.6766 +0.0327, +0.0173 to +0.0490
COCO val2017, caption to caption, recall@1 25,014 captions 0.3523 0.3556 -0.0033, -0.0062 to -0.0002

This model trails on captions by 0.0033, an amount the test resolves. bge-small's shipped mean was fitted on the first 20,000 of these same val2017 captions (research/bench/mu_probe.py in Black Window), averaged with a Wikipedia mean. Refitted as the average of a mean over 20,000 train2017 captions and one over 20,000 wikitext-103 paragraphs, it reads 0.3531, so 0.0025 of its lead is a mean fitted on the test captions. With its code mean in place of the shipped one this model reads 0.6891 on scifact and 0.3470 on COCO. The shuffled floors read at most 0.0066 on scifact and 0.0002 on COCO.

Heads fitted into this model's space. Black Window reads pictures through linear heads that map an image model's features into the text reader's space. Each was refitted for this model by the protocol that fitted bge-small's, and bge-small's refit reproduced its published figure in each case. Text-to-image recall@1, soup minus bge-small, paired by caption with 2,000 bootstrap resamples:

head image side population this model bge-small gap, 95% interval
heads/text_head_soup_v3gallery.safetensors the gallery's shipped image rows 5,001 captions over 123,287 images, 3 seeds 0.1122 0.1142 -0.0019, -0.0048 to +0.0009
heads/vision_qwen38_L52_soup.head Qwen3.8-27B layer 52 25,000 COCO val2017 captions 0.4148 0.4197 -0.0049, -0.0076 to -0.0020
tower/mcs0_soup.head MobileCLIP-S0 25,000 COCO val2017 captions 0.3160 0.3150 +0.0010, -0.0018 to +0.0038

tower/coco_caps_soup.bin is the browser tower's 24,000-caption word bank, the same texts as bge-small's re-embedded.

The exports. onnx/model.onnx matches PyTorch float32 in ONNX Runtime on the CPU on 2,000 texts, 1,000 code passages and 1,000 abstracts: a smallest cosine of 0.99999976 and a largest component difference of 3.6e-7. onnx/model_fp16.onnx, converted from it by ONNX Runtime's float16 converter with float32 inputs and outputs, is half the bytes and matches PyTorch float32 at a smallest cosine of 0.9999985 on the same 2,000 texts. Black Window's page loads it on WebGPU, where it reaches 0.99990 on 200 texts and embeds them 2.1 times as fast as the float32 file, and the prose results above move by 0.0009 each through it. The GGUF is above under Use. Black Window's wasm reader, the runtime's candle BertReader, runs model.safetensors and matches PyTorch fed the same token ids at a smallest cosine of 0.99999988 on the same 2,000 texts, both at 64 tokens and at the 256 it now reads.

How it was made

Base. BAAI/bge-small-en-v1.5: BERT with 12 layers, hidden size 384, 12 heads, intermediate size 1536, 512 positions and a 30,522-token vocabulary, CLS pooling, normalised.

a2, the issue-trained parent. Its data came from 86 permissively licensed Python repositories:

  • up to 80,000 commit pairs, each a cleaned commit message against the passage holding its first changed line;
  • up to 40,000 docstring pairs;
  • queries written for 64,000 seed passages, each as a GitHub issue, a search query or a question.

Three judges picked each query's best passage among eight candidates. The teachers and judges were gemma-4-31B-it, Mistral-Small-3.1-24B-Instruct-2503 and Qwen2.5-Coder-32B-Instruct. Training ran 2,500 steps of 128 examples at learning rate 2e-5, with InfoNCE at scale 20 over in-batch positives and one hard negative, 512 tokens, seed 0. None of the 12 SWE-bench Verified repositories is in the corpus.

d1, the distilled parent. bge-small was trained to reproduce gte-modernbert-base's geometry: 402,791 pairs from the training partitions of 20 CoIR subtasks, up to 25,000 queries each, with test texts removed. The loss was KL divergence between teacher and student softmaxes over each text's cosines to the other 255 in its batch, at temperature 0.05. It ran for 6,278 steps over two epochs, with AdamW at 2e-5, weight decay 0.01, a 5% warm-up, bf16 autocast, 512 tokens and seed 0.

The soup. It is the parameter-wise mean of a2 and d1 at weight 0.5 over bge-small's 199 tensors, with no further training and no other weight tried. On its own, a2 finds files as well as the soup (176 against 178 on the 411, sign p 0.90) and scores 0.4552 on CoIR. d1 leads CoIR at 0.6160 and finds fewer files than bge-small (133 against 145 on the 411). The mean keeps both skills.

Contamination

  • SWE-bench Verified. 89 of the 500 instances are flagged because a gold file at its base commit equals a text d1 trained on, or shares two or more two-line windows of 20-character lines with one. Headline figures use the other 411, with all 500 beside. a2's corpus excludes the 12 SWE-bench repositories and any directory named for their packages.
  • CoIR. d1 trained on CoIR's training partitions. 17,755 of the 162,216 test queries are flagged by one parent's check or the other's, and every CoIR figure here is a probe figure.
  • COCO. COCO captions are in the training data of some sentence encoders, and whether they are in bge-small's is not established.

Limits

  • It is English only, reads 512 tokens, and gives one 384-dimensional vector per text.
  • a2's code is Python only. d1's CoIR data spans the six CodeSearchNet languages (Go, Java, JavaScript, Ruby, Python and PHP), text-to-SQL, and the mixed languages of the Stack Overflow, CodeFeedback and CodeTransOcean sets.
  • It trails bge-small on caption-to-caption retrieval by 0.0033 (above), and its heads trail bge-small's by up to 0.0049.
  • It is Black Window's default memory reader: on the page since 2026-09-28, in Sunstone from 0.1.4 and in the app from blackwidow 1788bf9. bge-small stays available as ?mem=bge-small, sunstone.memoryReader set to bge-small, and BW_READER=bge. A store saved under bge-small is re-embedded once when opened under this model. In Sunstone, a store of 1,171 vectors written by 0.1.3 is re-read on 0.1.4's first open and read as saved on the next. In the app, a 1,063-passage store of Reddit posts is carried with its keys and texts, and on it this model places 38 of 40 planted passages first on the same 40 questions as bge-small, which places 38.
  • Black Window's wasm reader, used where a browser has no WebGPU, reads the first 256 tokens of a text from blackwidow 738d428 on, and cut at 64 before it. On that path, replayed over the same 411 SWE-bench instances, this model places the gold file first for 233 against bge-small's 190 at 256 tokens, and for 172 against 156 at 64, a gap the test could not resolve (sign p 0.072). The 256-token cut embeds at 0.306 of the 64-token rate on one CPU thread. Against the 512-token embedding, 256 tokens leave this model's rows at a median cosine of 1.000 and a 10th percentile of 0.969, where 64 left 0.850 and 0.741. The page served at blackwindow.xyz since 2026-09-28 carries the 256-token build.
  • Video enters Black Window only in the Mac app: each kept frame becomes a head row through heads/vision_qwen38_L52_soup.head and a caption row, and the frames' descriptions and any transcript become the attachment's passages. Sound enters as its whisper transcript, text rows like any other. On MSR-VTT's 1k-A test split, 1,000 videos and 20,000 captions, with every video's rows built as the app builds them, a caption's first row belongs to its video for 0.2231 of captions with this model's store against 0.2129 with bge-small's, +0.0102 with a 95% interval of +0.0052 to +0.0151 resampled by video. None of those videos carries an audio stream, so the transcript path is not measured. An earlier version of this card said Black Window's sound and video heads were fitted into bge-small's space and not refitted; no product has either head.

Files

file bytes sha256
centring/mean_captions_coco20k.json 8,699 28a5fe632aa0675f3617b2b48786056be87a7439fca342ac76f077896930ecc7
centring/mean_code_pilot64k.json 8,780 f6180ab7e11e394761409705b56897cd96182232b9c66e7679279da0dcd4e494
centring/mean_shipped.json 9,920 3e5e47695f63410fd9e388968b2d8601d4a57dfae1231ade03d6190fe0acd898
engine/gates.json 1,625 9add8f65da91bdee8262b8787e5119f7b8f7e445efb5b36485ddda506a4a4bb9
gguf/motherlode-code-small-en-v0.1-f16.gguf 67,583,456 66b77656814918815519004aa21cc6c0b5d2908f71248ea8c5050e3d3d7151c9
heads/text_head_soup_v3gallery.safetensors 789,736 3b63d8400e8d03486b686eabfc43d3a2827c2ccebbb72e052140bb21c2f95e66
heads/vision_qwen38_L52_soup.head 7,867,400 7068eaeb127d0a2f1eeae091e75050bedffc60d8e76a1bc9cd204199c8907226
model.safetensors 133,462,096 637a02871f785c69aa54ea6ae7e543692b1e6ce3596347739400ba43fa32b8e5
onnx/model.onnx 133,029,886 21b8550df9c72321f492a8d6045bd4074a785aa424238abdeecc6676735868ef
onnx/model_fp16.onnx 66,676,588 d4b036e96363c1b5515de862d87286ed4808f91a080ac206a9f5000502344a52
tower/coco_caps_soup.bin 18,432,000 b36d366d4d430bf84cfb324a542f9a721e7cc875805bdab41c3e733efae4a2c4
tower/mcs0_soup.head 789,504 789928e75e79def613d6692c1837c39f14a96c1ad264ae053fcda1cf83102771

Every file's hash, including the tokenizer and configuration, is in SHA256SUMS.

engine/gates.json holds the raw-cosine floors Black Window's chat path uses with this model under its shipped mean, mapped from bge-small's by matching quantiles over 4,193,047 scifact training pairs and checked on 1,554,900 test pairs, where each floor passes a non-gold pair within 0.00276 of bge-small's rate:

  • 0.3322, 0.3584, 0.4440, 0.4752 and 0.5078, for bge-small's 0.60, 0.62, 0.68, 0.70 and 0.72;
  • a band of 0.1375, for bge-small's 0.08.

Licence and attribution

The weights and the files fitted to them are licensed under the Business Source License 1.1, in LICENSE, with a Change Date of 15 September 2030, when they become Apache-2.0. Under its terms anyone may copy, modify, fine-tune, convert and redistribute them for non-production use, and evaluate, benchmark and publish measurements of them. Its Additional Use Grant adds production use for any internal purpose, including commercial development on your own or your employer's code, with no limit on seats. What needs a separate licence is offering the model to others as a hosted embedding, search or retrieval service, or supplying it to others inside a product or service.

This model is derived from BAAI/bge-small-en-v1.5, which is MIT-licensed. That licence's text and copyright notice are reproduced in NOTICE, and bge-small's configuration and tokenizer files remain under it. The training sources:

source used for licence
86 Python repositories, pilot/repos.txt in the evidence dataset a2's passages and commit pairs MIT, Apache-2.0, BSD-3-Clause, BSD-2-Clause, MIT-CMU, each read from its own licence file
gemma-4-31B-it, Mistral-Small-3.1-24B-Instruct-2503, Qwen2.5-Coder-32B-Instruct a2's queries and labels Apache-2.0
gte-modernbert-base d1's teacher Apache-2.0
CoIR's 20 training partitions (CoIR-Retrieval/{sub}-queries-corpus, -qrels) d1's pairs none declared: the Hub cards of all 40 datasets were read on 2026-09-27 and none states a licence. CoIR's evaluation code is Apache-2.0. Its sources include Stack Overflow question-answer pairs, which Stack Overflow publishes under CC BY-SA

The upstream models and datasets: bge-small-en-v1.5 (BAAI), gte-modernbert-base (Alibaba-NLP), CoIR (Li et al., arXiv:2407.02883), and SWE-bench Verified (princeton-nlp/SWE-bench_Verified). Weight averaging of fine-tunes follows Wortsman et al., "Model soups", ICML 2022.

Downloads last month
89
Safetensors
Model size
33.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RiverRider/motherlode-code-small-en-v0.1

Finetuned
(411)
this model

Paper for RiverRider/motherlode-code-small-en-v0.1