tahaenesaslanturk commited on
Commit
e2d8782
1 Parent(s): d17aa24

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +28 -3
README.md CHANGED
@@ -1,3 +1,28 @@
1
- ---
2
- license: mit
3
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ ---
4
+ # TS-Corpus BPE Tokenizer (256k, Cased)
5
+
6
+ ## Overview
7
+ This repository hosts a Byte Pair Encoding (BPE) tokenizer with a vocabulary size of 256,000, trained cased using several datasets from the TS Corpus website. The BPE method is particularly effective for languages like Turkish, providing a balance between word-level and character-level tokenization.
8
+
9
+ ## Dataset Sources
10
+ The tokenizer was trained on a variety of text sources from TS Corpus, ensuring a broad linguistic coverage. These sources include:
11
+ - [TS Corpus V2](https://tscorpus.com/corpora/ts-corpus-v2/)
12
+ - [TS Wikipedia Corpus](https://tscorpus.com/corpora/ts-wikipedia-corpus/)
13
+ - [TS Abstract Corpus](https://tscorpus.com/corpora/ts-abstract-corpus/)
14
+ - [TS Idioms and Proverbs Corpus](https://tscorpus.com/corpora/ts-idioms-and-proverbs-corpus/)
15
+ - [Syllable Corpus](https://tscorpus.com/corpora/syllable-corpus/)
16
+ - [Turkish Constitution Corpus](https://tscorpus.com/corpora/turkish-constitution-corpus/)
17
+
18
+ The inclusion of idiomatic expressions, proverbs, and legal terminology provides a comprehensive toolkit for processing Turkish text across different domains.
19
+
20
+ ## Tokenizer Model
21
+ Utilizing the Byte Pair Encoding (BPE) method, this tokenizer excels in efficiently managing subword units without the need for an extensive vocabulary. BPE is especially suitable for handling the agglutinative nature of Turkish, where words can have multiple suffixes.
22
+
23
+ ## Usage
24
+ To use this tokenizer in your projects, load it with the Hugging Face `transformers` library:
25
+ ```python
26
+ from transformers import AutoTokenizer
27
+ tokenizer = AutoTokenizer.from_pretrained("tahaenesaslanturk/ts-corpus-bpe-256k-cased")
28
+ ```