Instructions to use jadermcs/lexsplade-bert-large-silver-ldv with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use jadermcs/lexsplade-bert-large-silver-ldv with sentence-transformers:
from sentence_transformers import SparseEncoder model = SparseEncoder("jadermcs/lexsplade-bert-large-silver-ldv") queries = ["Which planet is known as the Red Planet?"] documents = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", ] query_embeddings = model.encode_query(queries) document_embeddings = model.encode_document(documents) similarities = model.similarity(query_embeddings, document_embeddings) print(similarities) - Notebooks
- Google Colab
- Kaggle
LexSPLADE (bert-large-uncased, silver, Longman Defining Vocabulary)
A sparse contextual word encoder whose output vocabulary is restricted to the Longman Defining Vocabulary. Given a sentence with one word marked, it returns a sparse vector describing that word in that context โ but only 1,897 of the 30,524 vocabulary dimensions (6.2%) can ever be non-zero, and each is a basic-English defining word.
The effect is that a vector reads as a definition:
"He sat on the <t> bank </t> of the river" -> bank shore bed edge water wall river face
"She withdrew the money from her <t> bank </t>" -> bank money account pay check payment trust debt
"The soldiers began their <t> charge </t>" -> charge attack march offensive advance defense battle
"There is no <t> charge </t> for delivery" -> charge cost tax payment pay rate price ticket
The restriction is trained, not applied afterwards: the mask is in place during training, so the gradients, the regularizers and the similarity all live in the restricted subspace.
The target word must be wrapped in <t> โฆ </t>. Pooling runs over the tokens
between those markers only.
Usage
from sentence_transformers import SparseEncoder
import torch
model = SparseEncoder("jadermcs/lexsplade-bert-large-silver-ldv", trust_remote_code=True)
sents = [
"He sat on the <t> bank </t> of the river and watched the water.",
"She withdrew the money from her <t> bank </t> on Monday morning.",
]
v = model.encode(sents, convert_to_tensor=True).to_dense().float()
print(torch.nn.functional.cosine_similarity(v[0:1], v[1:2], dim=1)) # ~0.17
top = v[0].topk(8).indices
print([model.tokenizer.convert_ids_to_tokens(int(i)) for i in top])
# ['bank', 'shore', 'bed', 'edge', 'water', 'wall', 'river', 'face']
trust_remote_code=True is required. The restriction lives in the custom
pooling module (modeling_lexsplade.py, ids in
1_TargetWordSpladePooling/config.json). Loading the weights as a plain
AutoModelForMaskedLM and pooling yourself will produce unrestricted vectors,
which is not what this model was trained to emit.
Keep the marked word inside max_seq_length; if the markers are truncated away
the pooling falls back to the whole sentence.
The vocabulary
Built from the Longman Dictionary of Contemporary English defining vocabulary โ the ~2,000 basic words in which every LDOCE definition is written. Three filters:
| step | dropped | remaining |
|---|---|---|
| published list, single-word entries | โ | 2,188 |
must be one whole-word token in the BERT vocabulary (advertise โ advert ##ise has no dimension of its own) |
63 | 2,125 |
| no stopwords or punctuation | 228 | 1,897 |
Shipped in this repo as longman_defining_vocabulary.txt (the word list) and
ldv_vocabulary.json (words plus the token ids the mask uses). The word list is
reproduced from the published LDOCE defining vocabulary for research use;
the dictionary and its defining vocabulary are the work of Longman/Pearson.
Training
| base model | google-bert/bert-large-uncased |
| training data | 31,914 same-lemma word-in-context pairs (LLM-verified silver + MCL-WiC train) |
| objective | AnglE loss on the sparse target-word vectors, with a SPLADE FLOPS regularizer and a Barlow Twins term on the dense span representation |
| output space | 1,897 LDV dimensions, masked throughout training |
| seed | 42 |
Evaluation
MCL-WiC, English. AP is average precision; accuracy is at the best threshold.
| split | accuracy | AP |
|---|---|---|
| test | 90.70 | 96.16 |
| dev | 90.90 | 95.36 |
Mean nonzero dimensions per usage: 118 of 30,524 โ roughly 60% fewer than the unrestricted model, at the same MCL-WiC AP (96.16 vs 96.30).
Variants
jadermcs/lexsplade-bert-large-silver-ldv(this model) โ 1,897 LDV dimensions.jadermcs/lexsplade-bert-large-silverโ identical training, full 30,524-dim output.
- Downloads last month
- 11
Model tree for jadermcs/lexsplade-bert-large-silver-ldv
Base model
google-bert/bert-large-uncased