ipipan
/

silver-retriever-base-v1

@@ -12,13 +12,19 @@ datasets:
 - ipipan/maupqa
 ---
-# HerBERT-base Retrieval (v2)
-HerBERT Retrieval model encodes the Polish sentences or paragraphs into a 768-dimensional dense vector space and can be used for tasks like document retrieval or semantic search.
-It was initialized from the [HerBERT-base](https://huggingface.co/allegro/herbert-base-cased) model and fine-tuned on the [PolQA](https://huggingface.co/ipipan/polqa) and [MAUPQA](https://huggingface.co/ipipan/maupqa) datasets for 40,000 steps with a batch size of 256.
-The model was trained on question-passage pairs and works best on similar tasks. The training passages consisted of `title` and `text` concatenated with the special token `</s>`. Even if your passages don't have a `title`, it is still beneficial to prefix a passage `text` with the `</s>` token.
 ## Usage (Sentence-Transformers)
 Using this model becomes easy when you have [sentence-transformers](https://www.SBERT.net) installed:
@@ -32,11 +38,11 @@ Then you can use the model like this:
 ```python
 from sentence_transformers import SentenceTransformer
 sentences = [
-    "W jakim mieście urodził się Zbigniew Herbert?",
     "Zbigniew Herbert</s>Zbigniew Bolesław Ryszard Herbert (ur. 29 października 1924 we Lwowie, zm. 28 lipca 1998 w Warszawie) – polski poeta, eseista i dramaturg.",
 ]
-model = SentenceTransformer('ipipan/herbert-base-retrieval-v2')
 embeddings = model.encode(sentences)
 print(embeddings)
 ```
@@ -55,12 +61,12 @@ def cls_pooling(model_output, attention_mask):
 # Sentences we want sentence embeddings for
 sentences = [
-    "W jakim mieście urodził się Zbigniew Herbert?",
     "Zbigniew Herbert</s>Zbigniew Bolesław Ryszard Herbert (ur. 29 października 1924 we Lwowie, zm. 28 lipca 1998 w Warszawie) – polski poeta, eseista i dramaturg.",
 ]
 # Load model from HuggingFace Hub
-tokenizer = AutoTokenizer.from_pretrained('ipipan/herbert-base-retrieval-v2')
-model = AutoModel.from_pretrained('ipipan/herbert-base-retrieval-v2')
 # Tokenize sentences
 encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')

 - ipipan/maupqa
 ---
+# Silver Retriever Base (v1)
+Silver Retriever model encodes the Polish sentences or paragraphs into a 768-dimensional dense vector space and can be used for tasks like document retrieval or semantic search.
+It was initialized from the [HerBERT-base](https://huggingface.co/allegro/herbert-base-cased) model and fine-tuned on the [PolQA](https://huggingface.co/ipipan/polqa) and [MAUPQA](https://huggingface.co/ipipan/maupqa) datasets for 15,000 steps with a batch size of 1,024.
+## Preparing inputs
+The model was trained on question-passage pairs and works best when the input is the same format as that used during training:
+- We added the phrase `Pytanie:' to the beginning of the question.
+- The training passages consisted of `title` and `text` concatenated with the special token `</s>`. Even if your passages don't have a `title`, it is still beneficial to prefix a passage with the `</s>` token.
+- Although we used the dot product during training, the model usually works better with the cosine distance.
 ## Usage (Sentence-Transformers)
 Using this model becomes easy when you have [sentence-transformers](https://www.SBERT.net) installed:
 ```python
 from sentence_transformers import SentenceTransformer
 sentences = [
+    "Pytanie: W jakim mieście urodził się Zbigniew Herbert?",
     "Zbigniew Herbert</s>Zbigniew Bolesław Ryszard Herbert (ur. 29 października 1924 we Lwowie, zm. 28 lipca 1998 w Warszawie) – polski poeta, eseista i dramaturg.",
 ]
+model = SentenceTransformer('ipipan/silver-retriever-base-v1')
 embeddings = model.encode(sentences)
 print(embeddings)
 ```
 # Sentences we want sentence embeddings for
 sentences = [
+    "Pytanie: W jakim mieście urodził się Zbigniew Herbert?",
     "Zbigniew Herbert</s>Zbigniew Bolesław Ryszard Herbert (ur. 29 października 1924 we Lwowie, zm. 28 lipca 1998 w Warszawie) – polski poeta, eseista i dramaturg.",
 ]
 # Load model from HuggingFace Hub
+tokenizer = AutoTokenizer.from_pretrained('ipipan/silver-retriever-base-v1')
+model = AutoModel.from_pretrained('ipipan/silver-retriever-base-v1')
 # Tokenize sentences
 encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')