Sentence Similarity
sentence-transformers
Safetensors
PEFT
English
text-embeddings
retrieval
web-search
news
Instructions to use desearch/Desearch-Embedding-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use desearch/Desearch-Embedding-4B with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("desearch/Desearch-Embedding-4B") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - PEFT
How to use desearch/Desearch-Embedding-4B with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Desearch Embedding 4B
Desearch Embedding 4B is a text embedding model for web and news search, developed by Desearch. It is fine-tuned from Qwen/Qwen3-Embedding-4B and maps search queries and web documents into a shared vector space for first-stage retrieval.
Highlights
- Built for news and the open web. Fine-tuned on recent news coverage, reference pages and encyclopedic articles, the content a web search engine has to rank every day.
- Real queries in every form. Short keyword searches, natural-language questions, noisy queries, questions about dated news events, and multi-hop questions that combine facts from linked pages.
- One model for chunks and whole pages. Documents are seen as paragraph chunks, page openings and full pages, the units a search index stores.
- Binary first-stage ready. Trained with a binary objective on the leading 256 dimensions, for search stacks that shortlist with compact binary vectors before rescoring with full vectors.
Model details
| Base model | Qwen/Qwen3-Embedding-4B |
| Parameters | 4B |
| Embedding dimension | 2560 |
| Max sequence length | 32K tokens |
| Pooling | Last token, L2-normalized |
| Query instruction | Built-in query prompt |
| Language | English |
| Training method | LoRA fine-tuning |
| License | Apache 2.0 |
Usage
The weights are a LoRA adapter on Qwen/Qwen3-Embedding-4B; the base model downloads automatically.
pip install -U sentence-transformers peft
Using Sentence Transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("desearch/Desearch-Embedding-4B")
queries = [
"Who became Apple’s chief executive and will lead the Sept. 9, 2026 “Surprise and Shine” product event?",
"How much does the base Mac Mini M6 cost in US dollars as of its August 2026 announcement?",
]
documents = [
"John Ternus, Apple's new chief executive, takes the stage at the September 9, 2026 “Surprise and Shine” event to unveil the latest iPhones and Apple Watches.",
"Apple surprised buyers in late August 2026 with the Mac Mini M6 at $899 and the Mac Mini M5 Pro at $1,699, both shipping on September 22.",
]
query_embeddings = model.encode(queries, prompt_name="query")
document_embeddings = model.encode(documents)
print(model.similarity(query_embeddings, document_embeddings))
Queries use the built-in query prompt; documents are encoded as they are.
Using Transformers
import torch
import torch.nn.functional as F
from peft import PeftModel
from transformers import AutoModel, AutoTokenizer
task = "Given a web search query, retrieve relevant passages that answer the query"
questions = [
"Who became Apple’s chief executive and will lead the Sept. 9, 2026 “Surprise and Shine” product event?",
"How much does the base Mac Mini M6 cost in US dollars as of its August 2026 announcement?",
]
queries = [f"Instruct: {task}\nQuery:{q}" for q in questions]
documents = [
"John Ternus, Apple's new chief executive, takes the stage at the September 9, 2026 “Surprise and Shine” event to unveil the latest iPhones and Apple Watches.",
"Apple surprised buyers in late August 2026 with the Mac Mini M6 at $899 and the Mac Mini M5 Pro at $1,699, both shipping on September 22.",
]
tokenizer = AutoTokenizer.from_pretrained("desearch/Desearch-Embedding-4B", padding_side="left")
model = AutoModel.from_pretrained("Qwen/Qwen3-Embedding-4B", dtype=torch.bfloat16)
model = PeftModel.from_pretrained(model, "desearch/Desearch-Embedding-4B").eval()
def encode(texts):
batch = tokenizer(texts, padding=True, truncation=True, max_length=8192, return_tensors="pt")
with torch.no_grad():
hidden = model(**batch).last_hidden_state
return F.normalize(hidden[:, -1], p=2, dim=1)
print(encode(queries) @ encode(documents).T)
Recommended use cases
- First-stage retrieval for web search, news search and retrieval-augmented generation
- Semantic search over news archives, documentation and reference content
- Candidate generation ahead of a reranker
Limitations
- Tuned on English text; other languages are not a focus of this release.
- Queries need the
queryprompt; documents are encoded without one. - Trained on documents up to 2,048 tokens; split longer pages into passages.
- Embeddings differ from Qwen/Qwen3-Embedding-4B; re-embed an existing corpus when switching.
License
This model is licensed under the Apache License 2.0. It is derived from Qwen/Qwen3-Embedding-4B, which is also licensed under the Apache License 2.0.
Citation
@misc{desearch2026embedding,
title = {Desearch Embedding 4B},
author = {Desearch},
year = {2026},
url = {https://huggingface.co/desearch/Desearch-Embedding-4B}
}