nilq commited on
Commit
119d28a
1 Parent(s): e169e8c

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +20 -0
README.md CHANGED
@@ -1,3 +1,23 @@
1
  ---
2
  license: mit
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: mit
3
+ language:
4
+ - en
5
  ---
6
+
7
+ ## Baby Tokenizer
8
+
9
+ Compact sentencepiece tokenizer for sample-efficient English language modeling.
10
+
11
+ ### Data
12
+
13
+ This tokeniser is derived from the BabyLM 100M dataset of mixed domain data, consisting of the following sources:
14
+ - CHILDES (child-directed speech)
15
+ - Subtitles (speech), BNC (speech)
16
+ - TED talks (speech)
17
+ - children's books (simple written language).
18
+
19
+ ### Specifications
20
+
21
+ - Vocabulary size: 20k
22
+ - Alphabet limit: 150
23
+ - Minimum token frequency: 5