Integrate with Sentence Transformers (#2)

- Integrate with Sentence Transformers (071e1bb3c6657116ca95005ca54faa2d8ca39798)
- Reduce to Sentence Transformers 2.3.1 (e5f16c718bef58eb279aa75c89541c56103bc4a8)

Files changed (5) hide show

1_Pooling/config.json ADDED Viewed

+{
+  "word_embedding_dimension": 384,
+  "pooling_mode_cls_token": false,
+  "pooling_mode_mean_tokens": true,
+  "pooling_mode_max_tokens": false,
+  "pooling_mode_mean_sqrt_len_tokens": false,
+  "pooling_mode_weightedmean_tokens": false,
+  "pooling_mode_lasttoken": false
+}

README.md CHANGED Viewed

@@ -7,6 +7,8 @@ tags:
 - String Matching
 - Fuzzy Join
 - Entity Retrieval
 ---
 ## PEARL-small
 [Learning High-Quality and General-Purpose Phrase Representations](https://arxiv.org/pdf/2401.10407.pdf). <br>
@@ -46,7 +48,25 @@ The FastText model here is `crawl-300d-2M-subword.bin`.
 ## Usage
-Below is an example of entity retrieval, and we reuse the code from E5.
 ```python
 import torch.nn.functional as F

 - String Matching
 - Fuzzy Join
 - Entity Retrieval
+- transformers
+- sentence-transformers
 ---
 ## PEARL-small
 [Learning High-Quality and General-Purpose Phrase Representations](https://arxiv.org/pdf/2401.10407.pdf). <br>
 ## Usage
+### Sentence Transformers
+PEARL is integrated with the Sentence Transformers library, and can be used like so:
+```python
+from sentence_transformers import SentenceTransformer, util
+query_texts = ["The New York Times"]
+doc_texts = [ "NYTimes", "New York Post", "New York"]
+input_texts = query_texts + doc_texts
+model = SentenceTransformer("Lihuchen/pearl_small")
+embeddings = model.encode(input_texts)
+scores = util.cos_sim(embeddings[0], embeddings[1:]) * 100
+print(scores.tolist())
+# [[90.56318664550781, 79.65763854980469, 75.52056121826172]]
+```
+### Transformers
+You can also use `transformers` to use PEARL. Below is an example of entity retrieval, and we reuse the code from E5.
 ```python
 import torch.nn.functional as F

config_sentence_transformers.json ADDED Viewed

+{
+  "__version__": {
+    "sentence_transformers": "2.3.1",
+    "transformers": "4.37.0",
+    "pytorch": "2.1.0+cu121"
+  }
+}

modules.json ADDED Viewed

+[
+  {
+    "idx": 0,
+    "name": "0",
+    "path": "",
+    "type": "sentence_transformers.models.Transformer"
+  },
+  {
+    "idx": 1,
+    "name": "1",
+    "path": "1_Pooling",
+    "type": "sentence_transformers.models.Pooling"
+  },
+  {
+    "idx": 2,
+    "name": "2",
+    "path": "2_Normalize",
+    "type": "sentence_transformers.models.Normalize"
+  }
+]

sentence_bert_config.json ADDED Viewed

+{
+  "max_seq_length": 512,
+  "do_lower_case": false
+}