Instructions to use AbstractPhil/captionbert-8192-v2-b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AbstractPhil/captionbert-8192-v2-b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="AbstractPhil/captionbert-8192-v2-b", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AbstractPhil/captionbert-8192-v2-b", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
captionbert-8192-b
A 58.3M standalone sentence encoder distilled from the geometric consensus of five BERT-family teachers. No expert models at inference: tokenizer + this model, 768-d L2-normalized output.
12 layers, 512-d, 8 heads, FFN 2048, 8192 position capacity. 0.53x bert-base.
This is the complete-corpus build: all 66 CC12M chunks, 31.9M rows. Its
sibling captionbert-8192-v2
trained on 54 chunks because ModernBERT was missing from 10 of them; those were
repaired and gate-verified before this run.
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("AbstractPhil/captionbert-8192-v2-B", trust_remote_code=True)
tok = AutoTokenizer.from_pretrained("google-bert/bert-base-uncased")
emb = model.encode(["a cat on a windowsill", "a feline by the window"]) # (2, 768)
(emb[0] @ emb[1]).item()
model.attach_amoe() # this repo's NATIVE arms -- see the warning below
emb = model.encode(["a cat on a windowsill"])
Benchmark
| model | params | STS-B | SICK-R | STS12 | STS13 | STS14 | STS15 | STS16 | BIOSSES | mean |
|---|---|---|---|---|---|---|---|---|---|---|
| bert-base | 109.5M | 0.4729 | 0.5865 | 0.3087 | 0.5988 | 0.4773 | 0.6029 | 0.6373 | 0.5469 | 0.5289 |
| ModernBERT-base | 149.0M | 0.4215 | 0.5479 | 0.3527 | 0.4247 | 0.3795 | 0.5349 | 0.4174 | 0.5630 | 0.4552 |
| roberta-base | 124.6M | 0.5436 | 0.6296 | 0.3211 | 0.5631 | 0.4522 | 0.6134 | 0.6198 | 0.5777 | 0.5401 |
| albert-base-v2 | 11.7M | 0.4784 | 0.5364 | 0.3101 | 0.4831 | 0.3809 | 0.5542 | 0.5491 | 0.4863 | 0.4723 |
| distilbert | 66.4M | 0.5717 | 0.6424 | 0.4344 | 0.6490 | 0.5410 | 0.6663 | 0.6854 | 0.5162 | 0.5883 |
| captionbert-8192-b | 58.3M | 0.5752 | 0.6548 | 0.5012 | 0.6037 | 0.5470 | 0.7146 | 0.6782 | 0.5500 | 0.6031 |
| captionbert-8192-b + arms | 63.2M | 0.7675 | 0.7374 | 0.6706 | 0.7381 | 0.6945 | 0.8109 | 0.7695 | 0.6472 | 0.7295 |
| captionbert-8192-v2 | 58.3M | 0.5747 | 0.6526 | 0.5051 | 0.5995 | 0.5452 | 0.7136 | 0.6776 | 0.5933 | 0.6077 |
| all-MiniLM-L6-v2 | 22.7M | 0.8203 | 0.7758 | 0.7237 | 0.8058 | 0.7559 | 0.8539 | 0.7899 | 0.8144 | 0.7925 |
All ten models measured in one harness, same eight tasks, mean-pooled and
L2-normalized, no task tuning. Spearman correlation; mean is the unweighted
average over the eight.
all-MiniLM-L6-v2 was contrastively trained on 1B+ curated sentence pairs. It is
listed for scale, not as a peer -- nothing here saw a similarity label.
The trunk beats every teacher it was distilled from, and the best of them (distilbert, .5883) by +.0194 -- at 13% of their combined 461M parameters, having never seen a similarity label. The margin comes mostly from STS12, where every teacher collapses to .31-.43 and the trunk holds .50.
With arms it clears the best teacher by +.14 and closes to within .063 of a model trained on a billion curated pairs.
Mean-pooled BERT-family encoders are known-weak sentence encoders -- that is the reason Sentence-BERT exists -- so beating them is an efficiency result rather than a state-of-the-art one. The MiniLM row is in the table to keep that honest.
Geometry
| model | self_cos | erank |
|---|---|---|
| bert-base | +0.6071 | 32.8 |
| ModernBERT-base | +0.9001 | 26.1 |
| roberta-base | +0.9594 | 19.8 |
| albert-base-v2 | +0.7473 | 20.9 |
| distilbert | +0.6920 | 31.1 |
| captionbert-8192-b | +0.1411 | 36.1 |
| captionbert-8192-b + arms | +0.0984 | 55.5 |
| captionbert-8192-v2 | +0.1396 | 36.6 |
| all-MiniLM-L6-v2 | +0.0251 | 86.7 |
self_cos is the isotropy gauge: the mean cosine between unrelated sentences.
Mean-pooled BERT-family embeddings sit in a narrow cone (+.61 to +.96), where
cosine cannot discriminate. erank is the participation ratio -- how many of the
768 directions carry variance.
Both track capability almost perfectly across all ten models, and isotropy is the mechanism: no isotropy objective appears anywhere in the training stack. The arms then lift erank 36.6 -> 57.6, the first evidence in this line that adaptation adds usable directions rather than only rotating them.
More data bought nothing (and that is the finding)
-b trained on 19% more rows for 19% more steps than -v2. Head to head:
| v2 (54ch, 26.9M) | b (66ch, 31.9M) | delta | |
|---|---|---|---|
| 8-task mean, bare | .6077 | .6031 | -.0046 |
| 7 tasks excluding BIOSSES | -- | -- | +.0009 |
| erank (STS-B) | 36.6 | 36.1 | -0.5 |
| self_cos (STS-B) | +.1396 | +.1411 | +.0015 |
| 8-task mean, native arms | .7287 | .7295 | +.0008 |
The entire -.0046 comes from BIOSSES, which is 100 rows -- a 0.4-sigma move. Everything else is a dead heat.
The ceiling is TEACHER AGREEMENT, not corpus size. The consensus target uses 28.7 of 768 directions: five BERT-family encoders only agree on ~29, and no amount of the same distribution raises that. The trunk reaches erank ~103 in domain but ~36 out of it -- the structure it builds on captions does not transfer. The next lever is heterogeneous teachers, measurable at the consensus stage before a single training step.
AMOE arms are TRUNK-BOUND -- use this repo's
Three 1.6M-parameter anchors on the frozen trunk, under a trained dispatch. Anchors toggle bit-exact, so one artifact serves both the unsupervised baseline and the adapted model.
| mask | STS-B | SICK-R | mean (8 tasks) |
|---|---|---|---|
| OFF (bare trunk) | .5752 | .6548 | .6031 |
equiv only |
.7219 | .7200 | .6842 |
simplify only |
.5995 | .6603 | .6254 |
paraphrase only |
.6137 | .6612 | .6295 |
| all three | .7675 | .7374 | .7295 |
An -only row is that arm as damped by the dispatch -- masking never
renormalizes, so it reads lower than the same anchor trained alone.
Do not attach captionbert-8192-v2's arms to this trunk. Measured:
| configuration | mean |
|---|---|
| v2 arms on v2 | .7287 |
| v2 arms on -b | .6863 |
| + re-aligned routing keys | .6987 |
| -b native anchors | .7295 |
Transferring the arms costs 31% of their gain. Re-training only the 1,536 routing keys recovers 29% of that; retraining the anchors recovers all of it. 71% of the loss is in the anchors themselves.
These two trunks are indistinguishable on eight STS tasks and on geometry, yet 1.6M adapter parameters tell them apart -- adapters read the residual stream and the task gauges read the pooled output, and the stream carries trunk identity the output does not. Budget one anchor set per trunk (~18 min).
attach_amoe() resolves this repo's own arms by default. Files are under
amoe/b-collective/. See amoe-lora.
How it was built
- Five teachers embedded 33M CC12M llava-next captions (mean-pooled, 768-d).
- One global whitened Procrustes map per teacher into
bert-base's frame, fit on a stratified random sample and reported out-of-sample (worst arm retains 95% of its in-sample R@1 at 1,833x chance). - Consensus = normalized centroid of the aligned teachers, per chunk.
- Student trained from scratch: InfoNCE(T=0.07) + per-sample MSE against the consensus. Pure Adam, no weight decay. 31.9M rows, 62,312 steps at batch 2048, ~6.4 h on one RTX 6000 Pro.
The alignment maps in maps/ are the same maps v2 used -- refitting them
would put the consensus targets in a different frame with no signal in the loss.
Known limits
- Consensus rank ~28.7 of 768. The model's ceiling, and a property of teacher agreement rather than of this model.
- Alignment quality varies by teacher. Out-of-sample cosine into the bert frame: distil .625, roberta .372, albert .331, modern .327 -- the ordering tracks architectural distance from bert-base.
- Single seed. The AMOE results carry a measured seed spread of .003-.005; the trunk does not have one.
- Trained on image captions; expect caption-like text to be its strongest domain.
- BIOSSES is 100 rows. Treat any single-task delta there as noise.
Files
model.safetensors the trunk, HF format
config.json AutoModel config (auto_map -> modeling_captionbert)
modeling_captionbert.py CaptionBertV2Model + attach_amoe/detach_amoe
checkpoints/ training checkpoints (final_model.pt is the ship)
maps/ alignment maps -- SHARED with v2, do not refit
amoe/b-collective/ native anchors + dispatch + metrics
Output convention
| field | shape | |
|---|---|---|
last_hidden_state |
(B, L, 512) | token states |
pooler_output |
(B, 768) | the embedding, L2-normalized |
embedding |
(B, 768) | alias |
geolip-captionbert-8192 (v1) returned the pooled embedding as
last_hidden_state. If porting v1 code, use pooler_output.
Citation
@misc{abstractphil2026captionbertb,
title = {captionbert-8192-b: consensus distillation on the complete CC12M corpus},
author = {AbstractPhil},
year = {2026},
url = {https://huggingface.co/AbstractPhil/captionbert-8192-v2-B}
}
MIT.
- Downloads last month
- -
Model tree for AbstractPhil/captionbert-8192-v2-b
Base model
FacebookAI/roberta-base