Instructions to use Hcompany/NeoMME-260M-Retriever-ST-dense with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use Hcompany/NeoMME-260M-Retriever-ST-dense with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("Hcompany/NeoMME-260M-Retriever-ST-dense") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
NeoMME-Retriever (260M): Single-Tower Multimodal-Native Multilingual Foundation Encoder 🔎
NeoMME-Retriever (260M) variants:
- Default (
transformers): Returns dense and multi-vector embeddings together with a single forward pass. Recommended for most use cases and inference.- ST dense [current]: Supports independent dense fine-tuning with Sentence Transformers.
- ST late-interaction: Supports independent multi-vector fine-tuning with Sentence Transformers.
NeoMME-260M-Retriever-ST-dense is a model for multimodal document retrieval. Fine-tuned from NeoMME-260M, it encodes text queries and documents (text or page screenshots) using one shared bidirectional Transformer encoder.
This model can be used with Sentence Transformers, but can only generate dense embeddings.
| Specification | Value |
|---|---|
| Parameters | 263M |
| Vocabulary | 131,072 tokens |
| Context length | 16,384 tokens |
| Hidden size | 1,024 |
| Image patches | 32 × 32 pixels, up to 2,048 pixels on the longest side (default) |
| Dense embeddings | 1,024 dimensions (Matryoshka: [128, 256, 512, 1,024]) |
| Dense pooling strategy | Mean |
Dense embeddings are L2-normalized and use cosine similarity. They match NeoMMEForRetrieval.dense_embeddings.
Performance
All scores use the metric shown at the full trained dimensions. Higher is better. ViDoRe v3, v2, and v1 measure visual document retrieval, while BEIR-15 measures text retrieval.
| Benchmark | Metric | NeoMME-260M | NeoMME-800M | ||
|---|---|---|---|---|---|
| Late interaction | Dense [current] | Late interaction | Dense | ||
| ViDoRe v3 | nDCG@10 | 0.5226 | 0.3907 | 0.5560 | 0.4391 |
| ViDoRe v2 | nDCG@5 | 0.5218 | 0.4075 | 0.5591 | 0.4475 |
| ViDoRe v1 | nDCG@5 | 0.8598 | 0.7552 | 0.8744 | 0.7993 |
| BEIR-15 | nDCG@10 | 0.4881 | 0.3055 | 0.5126 | 0.3686 |
Usage
pip install -U "sentence-transformers[image]"
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("Hcompany/NeoMME-260M-Retriever-ST-dense")
queries = [
"Quelle partie de la production pétrolière du Kazakhstan provient de champs en mer ?",
"Which hour of the day had the highest overall electricity generation in 2019?",
]
documents = [
"https://github.com/tonywu71/colpali-cookbooks/blob/main/examples/data/shift_kazakhstan.jpg?raw=true",
"https://github.com/tonywu71/colpali-cookbooks/blob/main/examples/data/energy_electricity_generation.jpg?raw=true",
]
query_embeddings = model.encode_query(queries, convert_to_tensor=True)
document_embeddings = model.encode_document(documents, convert_to_tensor=True)
scores = model.similarity(query_embeddings, document_embeddings)
# Expected: scores[0, 0] > scores[0, 1] and scores[1, 1] > scores[1, 0].
print(scores)
The score tensor has shape (num_queries, num_documents) and scores[i, j] is the score between query i and document j. A larger value indicates a closer match.
Training
NeoMME-260M-Retriever was fine-tuned from NeoMME-260M on text retrieval and document-page images. Training uses a joint late-interaction and Matryoshka dense contrastive objective.
The NeoMME technical report describes the full fine-tuning recipe (will be released soon).
Limitations
With Sentence Transformers, only one of the two retrieval heads can be used at a time.
License
Model weights are released under the Apache 2.0 license.
Citation
@misc{lac2026neommesingletowermultimodalnativemultilingual,
title={NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference},
author={Aurélien Lac and Tony Wu},
year={2026},
eprint={2609.01657},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2609.01657},
}
- Downloads last month
- 3