webbigdata-jp
commited on
Commit
•
cd9f0b8
1
Parent(s):
ef01472
version 2 upload
Browse files- README.md +406 -15
- adapter_config.json +8 -6
- adapter_model.safetensors +1 -1
- tokenizer.json +2 -2
- tokenizer_config.json +1692 -6
README.md
CHANGED
@@ -14,6 +14,18 @@ tags:
|
|
14 |
|
15 |
![image/png](https://cdn-uploads.huggingface.co/production/uploads/630469550907b9a115c91e62/m11e35NrZMi7ZpBQ7C6KV.png)
|
16 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
17 |
# モデルカード(Model Card for Model ID)
|
18 |
|
19 |
C3TR-AdapterはGoogleが発表したLLMであるgemma-7bの日英・英日翻訳性能を向上させるQLoRA Adapterです。
|
@@ -50,9 +62,9 @@ If you want to run it on your own local computer, you will need at least approxi
|
|
50 |
|
51 |
# Gemmaは最新のライブラリでなくては動かないので、以下のVersionに更新してください
|
52 |
# Gemma will not work without the latest library, so please update to the following version
|
53 |
-
pip install transformers==4.
|
54 |
-
pip install peft==0.
|
55 |
-
pip install bitsandbytes==0.
|
56 |
```
|
57 |
|
58 |
サンプルスクリプト(sample script)
|
@@ -69,6 +81,7 @@ peft_model_id = "webbigdata/C3TR-Adapter"
|
|
69 |
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
|
70 |
model = PeftModel.from_pretrained(model = model, model_id = peft_model_id)
|
71 |
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
|
|
72 |
|
73 |
def trans(my_str):
|
74 |
input_ids = tokenizer(my_str, return_tensors="pt",
|
@@ -79,33 +92,411 @@ def trans(my_str):
|
|
79 |
max_new_tokens=800, use_cache=True
|
80 |
)
|
81 |
full_outputs = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)
|
82 |
-
return full_outputs[0].split("###
|
83 |
|
84 |
ret = trans("""
|
85 |
-
|
|
|
|
|
86 |
Translate Japanese to English.
|
|
|
87 |
### Input:
|
88 |
-
|
|
|
89 |
|
90 |
-
|
91 |
-
### Answer:
|
92 |
""")
|
93 |
print(ret)
|
94 |
```
|
95 |
|
96 |
-
|
97 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
98 |
|
99 |
Instructionsは"Translate Japanese to English."(日英翻訳)と"Translate English to Japanese."(英日翻訳)の2種類です。
|
100 |
There are two types of instructions: "Translate Japanese to English." and "Translate English to Japanese.".
|
101 |
|
102 |
-
|
103 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
104 |
|
105 |
-
eg:
|
106 |
-
Translate English to Japanese within the context of subculture.
|
107 |
|
108 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
109 |
|
110 |
|
111 |
|
|
|
14 |
|
15 |
![image/png](https://cdn-uploads.huggingface.co/production/uploads/630469550907b9a115c91e62/m11e35NrZMi7ZpBQ7C6KV.png)
|
16 |
|
17 |
+
# News
|
18 |
+
|
19 |
+
2024.05.17
|
20 |
+
C3TR-AdapterのVersion2を公開しました。
|
21 |
+
Version 2 of C3TR-Adapter has been released.
|
22 |
+
|
23 |
+
Version2では主にカジュアルな会話に関する翻訳能力が大幅に向上しています。
|
24 |
+
Version 2 has greatly improved the ability to translate casual conversations.
|
25 |
+
|
26 |
+
その反面、フォーマルな文章の翻訳能力が少し落ちてしまっています。フォーマルな文章を対象にする場合、Version1を引き続きお使いください
|
27 |
+
On the other hand, translation capabilities for formal texts have declined slightly. If you are targeting formal texts, please continue to use Version 1.
|
28 |
+
|
29 |
# モデルカード(Model Card for Model ID)
|
30 |
|
31 |
C3TR-AdapterはGoogleが発表したLLMであるgemma-7bの日英・英日翻訳性能を向上させるQLoRA Adapterです。
|
|
|
62 |
|
63 |
# Gemmaは最新のライブラリでなくては動かないので、以下のVersionに更新してください
|
64 |
# Gemma will not work without the latest library, so please update to the following version
|
65 |
+
pip install transformers==4.40.0.dev0
|
66 |
+
pip install peft==0.10.0
|
67 |
+
pip install bitsandbytes==0.43.0
|
68 |
```
|
69 |
|
70 |
サンプルスクリプト(sample script)
|
|
|
81 |
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
|
82 |
model = PeftModel.from_pretrained(model = model, model_id = peft_model_id)
|
83 |
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
84 |
+
tokenizer.pad_token = tokenizer.unk_token
|
85 |
|
86 |
def trans(my_str):
|
87 |
input_ids = tokenizer(my_str, return_tensors="pt",
|
|
|
92 |
max_new_tokens=800, use_cache=True
|
93 |
)
|
94 |
full_outputs = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)
|
95 |
+
return full_outputs[0].split("### Response:\n")[-1].strip()
|
96 |
|
97 |
ret = trans("""
|
98 |
+
You are a highly skilled professional Japanese-English and English-Japanese translator. Translate the given text accurately, taking into account the context and specific instructions provided. Steps may include hints enclosed in square brackets [] with the key and value separated by a colon:. Only when the subject is specified in the Japanese sentence, the subject will be added when translating into English. If no additional instructions or context are provided, use your expertise to consider what the most appropriate context is and provide a natural translation that aligns with that context. When translating, strive to faithfully reflect the meaning and tone of the original text, pay attention to cultural nuances and differences in language usage, and ensure that the translation is grammatically correct and easy to read. After completing the translation, review it once more to check for errors or unnatural expressions. For technical terms and proper nouns, either leave them in the original language or use appropriate translations as necessary. Take a deep breath, calm down, and start translating.
|
99 |
+
|
100 |
+
### Instruction:
|
101 |
Translate Japanese to English.
|
102 |
+
|
103 |
### Input:
|
104 |
+
あら?また夜食を食べてるの?
|
105 |
+
こんにゃくは太りません
|
106 |
|
107 |
+
### Response:
|
|
|
108 |
""")
|
109 |
print(ret)
|
110 |
```
|
111 |
|
112 |
+
### プロンプトフォーマット prompt format
|
113 |
+
|
114 |
+
プロンプトフォーマットは独自です。
|
115 |
+
The prompt format is original.
|
116 |
+
|
117 |
+
Version1とVersion2ではプロンプトフォーマットも変わっており システムプロンプトも追加されています。
|
118 |
+
The prompt format has changed between Version 1 and Version 2, and system prompts have also been added.
|
119 |
+
|
120 |
+
余分な空白や改行はモデルの誤動作に繋がるのでテンプレートにミスがないようにしてください
|
121 |
+
Make sure there are no mistakes in the template, as extra spaces or line breaks can lead to malfunctioning of the model.
|
122 |
|
123 |
Instructionsは"Translate Japanese to English."(日英翻訳)と"Translate English to Japanese."(英日翻訳)の2種類です。
|
124 |
There are two types of instructions: "Translate Japanese to English." and "Translate English to Japanese.".
|
125 |
|
126 |
+
Version2からは実験的な試みとして、翻訳時に4つのタイプのヒントを与える事が出来るようになっています。
|
127 |
+
From Version 2, as an experimental trial, it is possible to give four types of hints when translating.
|
128 |
+
|
129 |
+
|
130 |
+
### (1)文体(writing style)
|
131 |
+
|
132 |
+
仕事場などではbusinessを使います
|
133 |
+
In the workplace, we use business.
|
134 |
+
|
135 |
+
|
136 |
+
```
|
137 |
+
You are a highly skilled professional Japanese-English and English-Japanese translator. Translate the given text accurately, taking into account the context and specific instructions provided. Steps may include hints enclosed in square brackets [] with the key and value separated by a colon:. Only when the subject is specified in the Japanese sentence, the subject will be added when translating into English. If no additional instructions or context are provided, use your expertise to consider what the most appropriate context is and provide a natural translation that aligns with that context. When translating, strive to faithfully reflect the meaning and tone of the original text, pay attention to cultural nuances and differences in language usage, and ensure that the translation is grammatically correct and easy to read. After completing the translation, review it once more to check for errors or unnatural expressions. For technical terms and proper nouns, either leave them in the original language or use appropriate translations as necessary. Take a deep breath, calm down, and start translating.
|
138 |
+
|
139 |
+
|
140 |
+
### Instruction:
|
141 |
+
Translate Japanese to English.
|
142 |
+
When translating, please use the following hints:
|
143 |
+
[writing_style: business]
|
144 |
+
|
145 |
+
### Input:
|
146 |
+
お疲れ様です、本日の資料を送ります。
|
147 |
+
|
148 |
+
### Response:
|
149 |
+
Thank you for your hard work today. I am sending today's materials.
|
150 |
+
```
|
151 |
+
|
152 |
+
|
153 |
+
以下の例ではsystem promptを省略しています。
|
154 |
+
In the following example, the system prompt is omitted.
|
155 |
+
|
156 |
+
コピペなどではslangやcasualを使います
|
157 |
+
Use slang or casual language when meme.
|
158 |
+
|
159 |
+
|
160 |
+
```
|
161 |
+
### Instruction:
|
162 |
+
Translate Japanese to English.
|
163 |
+
When translating, please use the following hints:
|
164 |
+
[writing_style: slang]
|
165 |
+
[牛鮭定食: Beef salmon set meal]
|
166 |
+
|
167 |
+
### Input:
|
168 |
+
そんな事より >>1 よ、ちょいと聞いてくれよ。スレとあんま関係ないけどさ。
|
169 |
+
このあいだ、近所の吉野家行ったんです。吉野家。
|
170 |
+
そしたらなんか人がめちゃくちゃいっぱいで座れないんです。
|
171 |
+
で、よく見たらなんか垂れ幕下がってて、150円引き、とか書いてあるんです。
|
172 |
+
もうね、アホかと。馬鹿かと。
|
173 |
+
お前らな、150円引き如きで普段来てない吉野家に来てんじゃねーよ、ボケが。
|
174 |
+
150円だよ、150円。
|
175 |
+
なんか親子連れとかもいるし。一家4人で吉野家か。おめでてーな。
|
176 |
+
よーしパパ特盛頼んじゃうぞー、とか言ってるの。もう見てらんない。
|
177 |
+
お前らな、150円やるからその席空けろと。
|
178 |
+
吉野家ってのはな、もっと殺伐としてるべきなんだよ。
|
179 |
+
Uの字テーブルの向かいに座った奴といつ喧嘩が始まってもおかしくない、
|
180 |
+
刺すか刺されるか、そんな雰囲気がいいんじゃねーか。女子供は、すっこんでろ。
|
181 |
+
で、やっと座れたかと思ったら、隣の奴が、大盛つゆだくで、とか言ってるんです。
|
182 |
+
そこでまたぶち切れですよ。
|
183 |
+
あのな、つゆだくなんてきょうび流行んねーんだよ。ボケが。
|
184 |
+
得意げな顔して何が、つゆだくで、だ。
|
185 |
+
お前は本当につゆだくを食いたいのかと問いたい。問い詰めたい。小1時間問い詰めたい。
|
186 |
+
お前、つゆだくって言いたいだけちゃうんかと。
|
187 |
+
吉野家通の俺から言わせてもらえば今、吉野家通の間での最新流行はやっぱり、
|
188 |
+
ねぎだく、これだね。
|
189 |
+
大盛りねぎだくギョク。これが通の頼み方。
|
190 |
+
ねぎだくってのはねぎが多めに入ってる。そん代わり肉が少なめ。これ。
|
191 |
+
で、それに大盛りギョク(玉子)。これ最強。
|
192 |
+
しかしこれを頼むと次から店員にマークされるという危険も伴う、諸刃の剣。
|
193 |
+
素人にはお薦め出来ない。
|
194 |
+
まあお前、>>1は、牛鮭定食でも食ってなさいってこった
|
195 |
+
|
196 |
+
### Response:
|
197 |
+
>1 Hey, listen up. This has nothing to do with the thread.
|
198 |
+
I went to Yoshinoya the other day. Yoshinoya.
|
199 |
+
And there were so many people that there was no place to sit.
|
200 |
+
So I looked around, and there was a sign hanging down, 150 yen off or something.
|
201 |
+
I mean, what a joke. What an idiot.
|
202 |
+
You guys, don't go to Yoshinoya just because it's 150 yen off, you morons.
|
203 |
+
It's 150 yen, 150 yen.
|
204 |
+
There were some families too. A family of four at Yoshinoya. Congratulations.
|
205 |
+
Alright, papa will order a special. I can't even look at that.
|
206 |
+
You guys, give up that seat if you're willing to pay 150 yen.
|
207 |
+
Yoshinoya should be more hardcore.
|
208 |
+
You could start a fight with anyone sitting across the U-shaped table at any moment,
|
209 |
+
stabbing or getting stabbed, that's the atmosphere I like. Women and children should shut up.
|
210 |
+
And then you think you've finally got a seat, and the guy next to you says, "I'll have a large tsuyu-daku."
|
211 |
+
And then there's another fight.
|
212 |
+
You know, tsuyu-daku isn't popular these days. Moron.
|
213 |
+
What's so proud about having tsuyu-daku.
|
214 |
+
I want to ask you if you really want to eat tsuyu-daku. I want to interrogate you. I want to interrogate you for an hour.
|
215 |
+
Do you just want to say tsuyu-daku?
|
216 |
+
From a Yoshinoya regular like me, the latest trend among Yoshinoya regulars is negi-daku.
|
217 |
+
Negi-daku gyoku. That's the way to order.
|
218 |
+
Negi-daku means more onions and less meat. That's it.
|
219 |
+
And then there's negi-daku gyoku (tamago). That's the best.
|
220 |
+
But there's also the risk of being marked by the staff if you order this. A double-edged sword.
|
221 |
+
I wouldn't recommend it to beginners.
|
222 |
+
Well, you, >>1, just eat a beef salmon set meal.
|
223 |
+
```
|
224 |
+
|
225 |
+
現在は11のwriteing styleがあります。
|
226 |
+
There are 11 writeing style at that time.
|
227 |
+
casual, formal, technical, journalistic, web-fiction, business, nsfw, educational-casual, academic-presentation, slang, sns-casual
|
228 |
+
|
229 |
+
|
230 |
+
#### (2)固有名詞の読み方 How to read proper nouns
|
231 |
+
|
232 |
+
[英語名称: 日本語訳] またはその逆。
|
233 |
+
[English name: Japanese translation name] or vice versa
|
234 |
+
|
235 |
+
eg.
|
236 |
+
|
237 |
+
```
|
238 |
+
### Instruction:
|
239 |
+
Translate Japanese to English.
|
240 |
+
When translating, please use the following hints:
|
241 |
+
[writing_style: formal]
|
242 |
+
[羽生結弦: Yuzuru Hanyu]
|
243 |
+
[羽生善治: Yoshiharu Habu]
|
244 |
+
|
245 |
+
### Input:
|
246 |
+
フィギュアスケートの羽生結弦さんが将棋棋士の羽生善治さんと対談した
|
247 |
+
|
248 |
+
### Response:
|
249 |
+
Figure skater Yuzuru Hanyu had a conversation with shogi player Yoshiharu Habu.
|
250 |
+
```
|
251 |
+
|
252 |
+
(3)キャラクタースタイル character_style
|
253 |
+
|
254 |
+
キャラクタースタイルの中で性別や個性を指定する事ができます
|
255 |
+
You can specify gender and personality in the character style.
|
256 |
+
|
257 |
+
|
258 |
+
男性指定 Male designated
|
259 |
+
```
|
260 |
+
### Instruction:
|
261 |
+
Translate Japanese to English.
|
262 |
+
When translating, please use the following hints:
|
263 |
+
[writing_style: formal]
|
264 |
+
[青山樹_character_style: male]
|
265 |
+
[青山樹: AOYAMA Itsuki]
|
266 |
+
|
267 |
+
### Input:
|
268 |
+
青山樹は週末に友達とキャンプに行って、自然を楽しんだ。そして時計を紛失した。
|
269 |
+
|
270 |
+
### Response:
|
271 |
+
Aoyama Itsuki went camping with his friends on the weekend and enjoyed nature. However, he lost his watch.
|
272 |
+
```
|
273 |
+
|
274 |
+
女性指定 Feale designated
|
275 |
+
```
|
276 |
+
### Instruction:
|
277 |
+
Translate Japanese to English.
|
278 |
+
When translating, please use the following hints:
|
279 |
+
[writing_style: formal]
|
280 |
+
[青山樹_character_style: female]
|
281 |
+
[青山樹: Itsuki Aoyama]
|
282 |
+
|
283 |
+
### Input:
|
284 |
+
青山樹は週末に友達とキャンプに行って、自然を楽しんだ。そして時計を紛失した。
|
285 |
+
|
286 |
+
### Response:
|
287 |
+
Itsuki Aoyama went camping with friends on the weekend and enjoyed nature. However, she lost her watch.
|
288 |
+
```
|
289 |
+
|
290 |
+
|
291 |
+
ノンバイナリー指定 nonbinary designated
|
292 |
+
|
293 |
+
```
|
294 |
+
### Instruction:
|
295 |
+
Translate Japanese to English.
|
296 |
+
When translating, please use the following hints:
|
297 |
+
[writing_style: formal]
|
298 |
+
[青山樹_character_style: nonbinary]
|
299 |
+
[青山樹: Tatsuki Aoyama]
|
300 |
+
|
301 |
+
### Input:
|
302 |
+
青山樹は週末に友達とキャンプに行って、自然を楽しんだ。そして時計を紛失した。
|
303 |
+
|
304 |
+
### Response:
|
305 |
+
Tatsuki Aoyama went camping with their friends on the weekend and enjoyed nature. They lost their watch.
|
306 |
+
```
|
307 |
+
|
308 |
+
残念ながら現時点では性別の指定は本文の内容が優先されるため、例えば以下の文章では性別指定が有効になりません。
|
309 |
+
以下の例では本文内の「俺は男だよ!」を消せば性別指定が有効になります。
|
310 |
+
|
311 |
+
Unfortunately, at present, the content of the text takes priority when designating gender, so for example, the gender designation will not be effective in the following sentence.
|
312 |
+
In the example below, if you delete "俺は男だよ!(I'm a guy!)" from the text, the gender specification will be effective.
|
313 |
|
|
|
|
|
314 |
|
315 |
+
```
|
316 |
+
### Instruction:
|
317 |
+
Translate Japanese to English.
|
318 |
+
When translating, please use the following hints:
|
319 |
+
[writing_style: web-fiction]
|
320 |
+
[カミーユ: kamille]
|
321 |
+
[kamille_character_style: female, rough, teenager]
|
322 |
+
[ジュリド: Jerid]
|
323 |
+
[ティターンズ: Titans]
|
324 |
+
[エマ: Emma]
|
325 |
+
[エゥーゴ: A.E.U.G.]
|
326 |
+
|
327 |
+
### Input:
|
328 |
+
ジェリド「カ��ーユ?女の名前なのに・・・何だ、男か。」
|
329 |
+
カミーユ「なめるな!!」
|
330 |
+
ジェリド「うわ!」
|
331 |
+
エマ「やめなさい!」
|
332 |
+
ジェリド「オレ達をティターンズと知ってちょっかいを出してきたのか?」
|
333 |
+
カミーユ「カミーユが男の名前で何で悪いんだ!!!俺は男だよ!」
|
334 |
+
|
335 |
+
こうして地球連邦のエリート部隊・ティターンズを殴った罪で拘束された後、母を失い、反地球連邦組織『エゥーゴ』に参加しました。
|
336 |
+
|
337 |
+
### Response:
|
338 |
+
Jerid: "Kamille? A girl's name, yet... What's with you, a guy?"
|
339 |
+
Kamille: "Don't you dare!"
|
340 |
+
Jerid: "Yikes!"
|
341 |
+
Emma: "Stop it!"
|
342 |
+
Jerid: "Are you provoking us by knowing we're Titans?"
|
343 |
+
Kamille: "What's wrong with Kamille being a guy's name!!! I'm a man!"
|
344 |
+
|
345 |
+
Thus, after being detained for assaulting the Earth Federation's elite force, the Titans, Kamille lost his mother and joined the anti-Earth Federation organization, the A.E.U.G.
|
346 |
+
```
|
347 |
+
|
348 |
+
|
349 |
+
character_styleとwriting_styleを組み合わせる
|
350 |
+
Combining character_style and writing_style
|
351 |
+
|
352 |
+
以下の例では段々と丁寧な言い回しに変化しています
|
353 |
+
In the following example, the phrase gradually changes to a more polite one.
|
354 |
+
|
355 |
+
|
356 |
+
```
|
357 |
+
### Instruction:
|
358 |
+
Translate Japanese to English.
|
359 |
+
When translating, please use the following hints:
|
360 |
+
[writing_style: slang]
|
361 |
+
[speaker_character_style: vulgar]
|
362 |
+
|
363 |
+
### Input:
|
364 |
+
今日の会議は非常に重要ですので、時間通りに来てください。
|
365 |
+
|
366 |
+
### Response:
|
367 |
+
Today’s meeting is very important, so please come on time.
|
368 |
+
```
|
369 |
+
|
370 |
+
|
371 |
+
```
|
372 |
+
### Instruction:
|
373 |
+
Translate Japanese to English.
|
374 |
+
When translating, please use the following hints:
|
375 |
+
[writing_style: casual]
|
376 |
+
[speaker_character_style: rough]
|
377 |
+
|
378 |
+
### Input:
|
379 |
+
今日の会議は非常に重要ですので、時間通りに来てください。
|
380 |
+
|
381 |
+
### Response:
|
382 |
+
The meeting today is very important, so please come on time.
|
383 |
+
```
|
384 |
+
|
385 |
+
```
|
386 |
+
### Instruction:
|
387 |
+
Translate Japanese to English.
|
388 |
+
When translating, please use the following hints:
|
389 |
+
[writing_style: formal]
|
390 |
+
[speaker_character_style: noble]
|
391 |
+
|
392 |
+
### Input:
|
393 |
+
今日の会議は非常に重要ですので、時間通りに来てください。
|
394 |
+
|
395 |
+
### Response:
|
396 |
+
As today's meeting is of utmost importance, please arrive on time.
|
397 |
+
```
|
398 |
+
|
399 |
+
|
400 |
+
日本語でも同様に丁寧になっていっています
|
401 |
+
The Japanese language is also becoming more polite.
|
402 |
+
|
403 |
+
|
404 |
+
```
|
405 |
+
### Instruction:
|
406 |
+
Translate English to Japanese.
|
407 |
+
When translating, please use the following hints:
|
408 |
+
[writing_style: slang]\[speaker_character_style: vulgar]
|
409 |
+
|
410 |
+
### Input:
|
411 |
+
Since today's meeting is very important, please arrive on time.
|
412 |
+
|
413 |
+
### Response:
|
414 |
+
今日の会議はとても重要なので、時間通りに来るように。
|
415 |
+
```
|
416 |
+
|
417 |
+
```
|
418 |
+
### Instruction:
|
419 |
+
Translate English to Japanese.
|
420 |
+
When translating, please use the following hints:
|
421 |
+
[writing_style: casual]
|
422 |
+
[speaker_character_style: rough]
|
423 |
+
|
424 |
+
### Input:
|
425 |
+
Since today's meeting is very important, please arrive on time.
|
426 |
+
|
427 |
+
### Response:
|
428 |
+
今日の会議はとても重要なので、時間通りに来るようにしてください。
|
429 |
+
```
|
430 |
+
|
431 |
+
|
432 |
+
```
|
433 |
+
### Instruction:
|
434 |
+
Translate English to Japanese.
|
435 |
+
When translating, please use the following hints:
|
436 |
+
[writing_style: formal]
|
437 |
+
[speaker_character_style: noble]
|
438 |
+
|
439 |
+
### Input:
|
440 |
+
Since today's meeting is very important, please arrive on time.
|
441 |
+
|
442 |
+
### Response:
|
443 |
+
今日の会議は非常に重要ですので、時間どおりにお越しください。
|
444 |
+
```
|
445 |
+
|
446 |
+
|
447 |
+
(4)一人称と語尾(First person and ending)
|
448 |
+
キャラクターの一人称と語尾を指定する事ができます。
|
449 |
+
この機能はまだ非常に実験的な機能であり、現時点では不完全です。
|
450 |
+
|
451 |
+
You can specify the first person and ending for your character.
|
452 |
+
This feature is still highly experimental and incomplete at this time.
|
453 |
+
|
454 |
+
|
455 |
+
```
|
456 |
+
### Instruction:
|
457 |
+
Translate English to Japanese.
|
458 |
+
[writing_style: casual, anime]
|
459 |
+
[squid girl: イカ娘]
|
460 |
+
[squid girl_first_person_and_ending: 私, イカ, ゲソ]
|
461 |
+
[marisa: 魔理沙]
|
462 |
+
[marisa_first_person_and_ending: 俺, ぜ, だぜ]
|
463 |
+
[reimu: 霊夢]
|
464 |
+
[reimu_first_person_and_ending: 私, かしら, だわ]
|
465 |
+
|
466 |
+
### Input:
|
467 |
+
Marisa: Hey Reimu, check out this cool squid girl I just met!
|
468 |
+
Reimu: A squid girl? How unusual. What brings you to Gensokyo?
|
469 |
+
Squid Girl: I'm on a mission to invade the surface world, starting with this land! ika!
|
470 |
+
Marisa: Haha, good luck with that! Reimu and I have dealt with way worse than some little squid.
|
471 |
+
Squid Girl: I'm not just some squid! I'll show you the power of the sea! ...Eventually! geso
|
472 |
+
|
473 |
+
### Response:
|
474 |
+
魔理沙: 霊夢、このイカ娘に会ったぜ!
|
475 |
+
霊夢: イカ娘?珍しいな。何で幻想郷に来たのかしら?
|
476 |
+
イカ娘: 地上世界を征服する任務を帯びて、まずはこの土地から始めるイカゲソ!
|
477 |
+
魔理沙: 笑うぜ、イカ娘。霊夢と俺には、イカなんかよりひどい奴らと戦ったことがあるんだぜ。
|
478 |
+
イカ娘: 私、ただのイカじゃないゲソ!海の力を見せつけるゲソ!...いつか��ソ!
|
479 |
+
```
|
480 |
+
|
481 |
+
|
482 |
+
```
|
483 |
+
### Instruction:
|
484 |
+
Translate Japanese to English.
|
485 |
+
When translating, please use the following hints:
|
486 |
+
[writing_style: casual, game]
|
487 |
+
[初春: hatsuharu]
|
488 |
+
|
489 |
+
[時雨: shigure]
|
490 |
+
[夕立: yuudachi]
|
491 |
+
[yuudachi_first_person_and_ending: poi]
|
492 |
+
[白露: shiratsuyu]
|
493 |
+
|
494 |
+
### Input:
|
495 |
+
時雨「雨は、いつか止むさ」 初春「わらわは、北方部隊に所属。戦雲渦巻くアッツやキスカなどの北方海域で活躍したぞ。」 夕立「白露型駆逐艦・夕立!今日も頑張るっぽい!ぽ~いっ!」
|
496 |
+
|
497 |
+
### Response:
|
498 |
+
Shigure: "The rain will eventually stop." Hatsuharu: "I belong to the Northern Fleet. I was active in the northern seas around Attu and Kiska, where the clouds of war swirl." Yuudachi: "Shiratsuyu-class destroyer, Yuudachi! I'll do my best today too! Po~i!"
|
499 |
+
```
|
500 |
|
501 |
|
502 |
|
adapter_config.json
CHANGED
@@ -5,12 +5,13 @@
|
|
5 |
"bias": "none",
|
6 |
"fan_in_fan_out": false,
|
7 |
"inference_mode": true,
|
8 |
-
"init_lora_weights":
|
|
|
9 |
"layers_pattern": null,
|
10 |
"layers_to_transform": null,
|
11 |
"loftq_config": {},
|
12 |
"lora_alpha": 32,
|
13 |
-
"lora_dropout": 0,
|
14 |
"megatron_config": null,
|
15 |
"megatron_core": "megatron.core",
|
16 |
"modules_to_save": null,
|
@@ -19,14 +20,15 @@
|
|
19 |
"rank_pattern": {},
|
20 |
"revision": "unsloth",
|
21 |
"target_modules": [
|
|
|
22 |
"k_proj",
|
|
|
|
|
23 |
"up_proj",
|
24 |
-
"v_proj",
|
25 |
"gate_proj",
|
26 |
-
"
|
27 |
-
"q_proj",
|
28 |
-
"down_proj"
|
29 |
],
|
30 |
"task_type": "CAUSAL_LM",
|
|
|
31 |
"use_rslora": false
|
32 |
}
|
|
|
5 |
"bias": "none",
|
6 |
"fan_in_fan_out": false,
|
7 |
"inference_mode": true,
|
8 |
+
"init_lora_weights": "gaussian",
|
9 |
+
"layer_replication": null,
|
10 |
"layers_pattern": null,
|
11 |
"layers_to_transform": null,
|
12 |
"loftq_config": {},
|
13 |
"lora_alpha": 32,
|
14 |
+
"lora_dropout": 0.05,
|
15 |
"megatron_config": null,
|
16 |
"megatron_core": "megatron.core",
|
17 |
"modules_to_save": null,
|
|
|
20 |
"rank_pattern": {},
|
21 |
"revision": "unsloth",
|
22 |
"target_modules": [
|
23 |
+
"o_proj",
|
24 |
"k_proj",
|
25 |
+
"q_proj",
|
26 |
+
"down_proj",
|
27 |
"up_proj",
|
|
|
28 |
"gate_proj",
|
29 |
+
"v_proj"
|
|
|
|
|
30 |
],
|
31 |
"task_type": "CAUSAL_LM",
|
32 |
+
"use_dora": false,
|
33 |
"use_rslora": false
|
34 |
}
|
adapter_model.safetensors
CHANGED
@@ -1,3 +1,3 @@
|
|
1 |
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:
|
3 |
size 800116456
|
|
|
1 |
version https://git-lfs.github.com/spec/v1
|
2 |
+
oid sha256:63574dfbad9845da5aaa22aa3344b88855f6e9ab3c33b90b2782bf8cd56bc1a7
|
3 |
size 800116456
|
tokenizer.json
CHANGED
@@ -1,3 +1,3 @@
|
|
1 |
version https://git-lfs.github.com/spec/v1
|
2 |
-
oid sha256:
|
3 |
-
size
|
|
|
1 |
version https://git-lfs.github.com/spec/v1
|
2 |
+
oid sha256:f30f819ff5b0f4cef2c8a6aafbeb20a13e7dd14409ece0f1e11e1d84bcfd281b
|
3 |
+
size 17518937
|
tokenizer_config.json
CHANGED
@@ -34,21 +34,1709 @@
|
|
34 |
"single_word": false,
|
35 |
"special": true
|
36 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
37 |
"106": {
|
38 |
"content": "<start_of_turn>",
|
39 |
"lstrip": false,
|
40 |
"normalized": false,
|
41 |
"rstrip": false,
|
42 |
"single_word": false,
|
43 |
-
"special": true
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
44 |
},
|
45 |
-
"
|
46 |
-
"content": "
|
47 |
"lstrip": false,
|
48 |
"normalized": false,
|
49 |
"rstrip": false,
|
50 |
"single_word": false,
|
51 |
-
"special":
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
52 |
}
|
53 |
},
|
54 |
"additional_special_tokens": [
|
@@ -56,10 +1744,8 @@
|
|
56 |
"<end_of_turn>"
|
57 |
],
|
58 |
"bos_token": "<bos>",
|
59 |
-
"chat_template": "{% if messages[0]['role'] == 'system' %}{{ raise_exception('System role not supported') }}{% endif %}{% for message in messages %}{% if (message['role'] == 'user') != (loop.index0 % 2 == 0) %}{{ raise_exception('Conversation roles must alternate user/assistant/user/assistant/...') }}{% endif %}{% if (message['role'] == 'assistant') %}{% set role = 'model' %}{% else %}{% set role = message['role'] %}{% endif %}{{ '<start_of_turn>' + role + '\n' + message['content'] | trim + '<end_of_turn>\n' }}{% endfor %}{% if add_generation_prompt %}{{'<start_of_turn>model\n'}}{% endif %}",
|
60 |
"clean_up_tokenization_spaces": false,
|
61 |
"eos_token": "<eos>",
|
62 |
-
"legacy": null,
|
63 |
"model_max_length": 8192,
|
64 |
"pad_token": "<unk>",
|
65 |
"padding_side": "right",
|
|
|
34 |
"single_word": false,
|
35 |
"special": true
|
36 |
},
|
37 |
+
"4": {
|
38 |
+
"content": "<mask>",
|
39 |
+
"lstrip": false,
|
40 |
+
"normalized": false,
|
41 |
+
"rstrip": false,
|
42 |
+
"single_word": false,
|
43 |
+
"special": false
|
44 |
+
},
|
45 |
+
"5": {
|
46 |
+
"content": "<2mass>",
|
47 |
+
"lstrip": false,
|
48 |
+
"normalized": false,
|
49 |
+
"rstrip": false,
|
50 |
+
"single_word": false,
|
51 |
+
"special": false
|
52 |
+
},
|
53 |
+
"6": {
|
54 |
+
"content": "[@BOS@]",
|
55 |
+
"lstrip": false,
|
56 |
+
"normalized": false,
|
57 |
+
"rstrip": false,
|
58 |
+
"single_word": false,
|
59 |
+
"special": false
|
60 |
+
},
|
61 |
+
"7": {
|
62 |
+
"content": "<unused0>",
|
63 |
+
"lstrip": false,
|
64 |
+
"normalized": false,
|
65 |
+
"rstrip": false,
|
66 |
+
"single_word": false,
|
67 |
+
"special": false
|
68 |
+
},
|
69 |
+
"8": {
|
70 |
+
"content": "<unused1>",
|
71 |
+
"lstrip": false,
|
72 |
+
"normalized": false,
|
73 |
+
"rstrip": false,
|
74 |
+
"single_word": false,
|
75 |
+
"special": false
|
76 |
+
},
|
77 |
+
"9": {
|
78 |
+
"content": "<unused2>",
|
79 |
+
"lstrip": false,
|
80 |
+
"normalized": false,
|
81 |
+
"rstrip": false,
|
82 |
+
"single_word": false,
|
83 |
+
"special": false
|
84 |
+
},
|
85 |
+
"10": {
|
86 |
+
"content": "<unused3>",
|
87 |
+
"lstrip": false,
|
88 |
+
"normalized": false,
|
89 |
+
"rstrip": false,
|
90 |
+
"single_word": false,
|
91 |
+
"special": false
|
92 |
+
},
|
93 |
+
"11": {
|
94 |
+
"content": "<unused4>",
|
95 |
+
"lstrip": false,
|
96 |
+
"normalized": false,
|
97 |
+
"rstrip": false,
|
98 |
+
"single_word": false,
|
99 |
+
"special": false
|
100 |
+
},
|
101 |
+
"12": {
|
102 |
+
"content": "<unused5>",
|
103 |
+
"lstrip": false,
|
104 |
+
"normalized": false,
|
105 |
+
"rstrip": false,
|
106 |
+
"single_word": false,
|
107 |
+
"special": false
|
108 |
+
},
|
109 |
+
"13": {
|
110 |
+
"content": "<unused6>",
|
111 |
+
"lstrip": false,
|
112 |
+
"normalized": false,
|
113 |
+
"rstrip": false,
|
114 |
+
"single_word": false,
|
115 |
+
"special": false
|
116 |
+
},
|
117 |
+
"14": {
|
118 |
+
"content": "<unused7>",
|
119 |
+
"lstrip": false,
|
120 |
+
"normalized": false,
|
121 |
+
"rstrip": false,
|
122 |
+
"single_word": false,
|
123 |
+
"special": false
|
124 |
+
},
|
125 |
+
"15": {
|
126 |
+
"content": "<unused8>",
|
127 |
+
"lstrip": false,
|
128 |
+
"normalized": false,
|
129 |
+
"rstrip": false,
|
130 |
+
"single_word": false,
|
131 |
+
"special": false
|
132 |
+
},
|
133 |
+
"16": {
|
134 |
+
"content": "<unused9>",
|
135 |
+
"lstrip": false,
|
136 |
+
"normalized": false,
|
137 |
+
"rstrip": false,
|
138 |
+
"single_word": false,
|
139 |
+
"special": false
|
140 |
+
},
|
141 |
+
"17": {
|
142 |
+
"content": "<unused10>",
|
143 |
+
"lstrip": false,
|
144 |
+
"normalized": false,
|
145 |
+
"rstrip": false,
|
146 |
+
"single_word": false,
|
147 |
+
"special": false
|
148 |
+
},
|
149 |
+
"18": {
|
150 |
+
"content": "<unused11>",
|
151 |
+
"lstrip": false,
|
152 |
+
"normalized": false,
|
153 |
+
"rstrip": false,
|
154 |
+
"single_word": false,
|
155 |
+
"special": false
|
156 |
+
},
|
157 |
+
"19": {
|
158 |
+
"content": "<unused12>",
|
159 |
+
"lstrip": false,
|
160 |
+
"normalized": false,
|
161 |
+
"rstrip": false,
|
162 |
+
"single_word": false,
|
163 |
+
"special": false
|
164 |
+
},
|
165 |
+
"20": {
|
166 |
+
"content": "<unused13>",
|
167 |
+
"lstrip": false,
|
168 |
+
"normalized": false,
|
169 |
+
"rstrip": false,
|
170 |
+
"single_word": false,
|
171 |
+
"special": false
|
172 |
+
},
|
173 |
+
"21": {
|
174 |
+
"content": "<unused14>",
|
175 |
+
"lstrip": false,
|
176 |
+
"normalized": false,
|
177 |
+
"rstrip": false,
|
178 |
+
"single_word": false,
|
179 |
+
"special": false
|
180 |
+
},
|
181 |
+
"22": {
|
182 |
+
"content": "<unused15>",
|
183 |
+
"lstrip": false,
|
184 |
+
"normalized": false,
|
185 |
+
"rstrip": false,
|
186 |
+
"single_word": false,
|
187 |
+
"special": false
|
188 |
+
},
|
189 |
+
"23": {
|
190 |
+
"content": "<unused16>",
|
191 |
+
"lstrip": false,
|
192 |
+
"normalized": false,
|
193 |
+
"rstrip": false,
|
194 |
+
"single_word": false,
|
195 |
+
"special": false
|
196 |
+
},
|
197 |
+
"24": {
|
198 |
+
"content": "<unused17>",
|
199 |
+
"lstrip": false,
|
200 |
+
"normalized": false,
|
201 |
+
"rstrip": false,
|
202 |
+
"single_word": false,
|
203 |
+
"special": false
|
204 |
+
},
|
205 |
+
"25": {
|
206 |
+
"content": "<unused18>",
|
207 |
+
"lstrip": false,
|
208 |
+
"normalized": false,
|
209 |
+
"rstrip": false,
|
210 |
+
"single_word": false,
|
211 |
+
"special": false
|
212 |
+
},
|
213 |
+
"26": {
|
214 |
+
"content": "<unused19>",
|
215 |
+
"lstrip": false,
|
216 |
+
"normalized": false,
|
217 |
+
"rstrip": false,
|
218 |
+
"single_word": false,
|
219 |
+
"special": false
|
220 |
+
},
|
221 |
+
"27": {
|
222 |
+
"content": "<unused20>",
|
223 |
+
"lstrip": false,
|
224 |
+
"normalized": false,
|
225 |
+
"rstrip": false,
|
226 |
+
"single_word": false,
|
227 |
+
"special": false
|
228 |
+
},
|
229 |
+
"28": {
|
230 |
+
"content": "<unused21>",
|
231 |
+
"lstrip": false,
|
232 |
+
"normalized": false,
|
233 |
+
"rstrip": false,
|
234 |
+
"single_word": false,
|
235 |
+
"special": false
|
236 |
+
},
|
237 |
+
"29": {
|
238 |
+
"content": "<unused22>",
|
239 |
+
"lstrip": false,
|
240 |
+
"normalized": false,
|
241 |
+
"rstrip": false,
|
242 |
+
"single_word": false,
|
243 |
+
"special": false
|
244 |
+
},
|
245 |
+
"30": {
|
246 |
+
"content": "<unused23>",
|
247 |
+
"lstrip": false,
|
248 |
+
"normalized": false,
|
249 |
+
"rstrip": false,
|
250 |
+
"single_word": false,
|
251 |
+
"special": false
|
252 |
+
},
|
253 |
+
"31": {
|
254 |
+
"content": "<unused24>",
|
255 |
+
"lstrip": false,
|
256 |
+
"normalized": false,
|
257 |
+
"rstrip": false,
|
258 |
+
"single_word": false,
|
259 |
+
"special": false
|
260 |
+
},
|
261 |
+
"32": {
|
262 |
+
"content": "<unused25>",
|
263 |
+
"lstrip": false,
|
264 |
+
"normalized": false,
|
265 |
+
"rstrip": false,
|
266 |
+
"single_word": false,
|
267 |
+
"special": false
|
268 |
+
},
|
269 |
+
"33": {
|
270 |
+
"content": "<unused26>",
|
271 |
+
"lstrip": false,
|
272 |
+
"normalized": false,
|
273 |
+
"rstrip": false,
|
274 |
+
"single_word": false,
|
275 |
+
"special": false
|
276 |
+
},
|
277 |
+
"34": {
|
278 |
+
"content": "<unused27>",
|
279 |
+
"lstrip": false,
|
280 |
+
"normalized": false,
|
281 |
+
"rstrip": false,
|
282 |
+
"single_word": false,
|
283 |
+
"special": false
|
284 |
+
},
|
285 |
+
"35": {
|
286 |
+
"content": "<unused28>",
|
287 |
+
"lstrip": false,
|
288 |
+
"normalized": false,
|
289 |
+
"rstrip": false,
|
290 |
+
"single_word": false,
|
291 |
+
"special": false
|
292 |
+
},
|
293 |
+
"36": {
|
294 |
+
"content": "<unused29>",
|
295 |
+
"lstrip": false,
|
296 |
+
"normalized": false,
|
297 |
+
"rstrip": false,
|
298 |
+
"single_word": false,
|
299 |
+
"special": false
|
300 |
+
},
|
301 |
+
"37": {
|
302 |
+
"content": "<unused30>",
|
303 |
+
"lstrip": false,
|
304 |
+
"normalized": false,
|
305 |
+
"rstrip": false,
|
306 |
+
"single_word": false,
|
307 |
+
"special": false
|
308 |
+
},
|
309 |
+
"38": {
|
310 |
+
"content": "<unused31>",
|
311 |
+
"lstrip": false,
|
312 |
+
"normalized": false,
|
313 |
+
"rstrip": false,
|
314 |
+
"single_word": false,
|
315 |
+
"special": false
|
316 |
+
},
|
317 |
+
"39": {
|
318 |
+
"content": "<unused32>",
|
319 |
+
"lstrip": false,
|
320 |
+
"normalized": false,
|
321 |
+
"rstrip": false,
|
322 |
+
"single_word": false,
|
323 |
+
"special": false
|
324 |
+
},
|
325 |
+
"40": {
|
326 |
+
"content": "<unused33>",
|
327 |
+
"lstrip": false,
|
328 |
+
"normalized": false,
|
329 |
+
"rstrip": false,
|
330 |
+
"single_word": false,
|
331 |
+
"special": false
|
332 |
+
},
|
333 |
+
"41": {
|
334 |
+
"content": "<unused34>",
|
335 |
+
"lstrip": false,
|
336 |
+
"normalized": false,
|
337 |
+
"rstrip": false,
|
338 |
+
"single_word": false,
|
339 |
+
"special": false
|
340 |
+
},
|
341 |
+
"42": {
|
342 |
+
"content": "<unused35>",
|
343 |
+
"lstrip": false,
|
344 |
+
"normalized": false,
|
345 |
+
"rstrip": false,
|
346 |
+
"single_word": false,
|
347 |
+
"special": false
|
348 |
+
},
|
349 |
+
"43": {
|
350 |
+
"content": "<unused36>",
|
351 |
+
"lstrip": false,
|
352 |
+
"normalized": false,
|
353 |
+
"rstrip": false,
|
354 |
+
"single_word": false,
|
355 |
+
"special": false
|
356 |
+
},
|
357 |
+
"44": {
|
358 |
+
"content": "<unused37>",
|
359 |
+
"lstrip": false,
|
360 |
+
"normalized": false,
|
361 |
+
"rstrip": false,
|
362 |
+
"single_word": false,
|
363 |
+
"special": false
|
364 |
+
},
|
365 |
+
"45": {
|
366 |
+
"content": "<unused38>",
|
367 |
+
"lstrip": false,
|
368 |
+
"normalized": false,
|
369 |
+
"rstrip": false,
|
370 |
+
"single_word": false,
|
371 |
+
"special": false
|
372 |
+
},
|
373 |
+
"46": {
|
374 |
+
"content": "<unused39>",
|
375 |
+
"lstrip": false,
|
376 |
+
"normalized": false,
|
377 |
+
"rstrip": false,
|
378 |
+
"single_word": false,
|
379 |
+
"special": false
|
380 |
+
},
|
381 |
+
"47": {
|
382 |
+
"content": "<unused40>",
|
383 |
+
"lstrip": false,
|
384 |
+
"normalized": false,
|
385 |
+
"rstrip": false,
|
386 |
+
"single_word": false,
|
387 |
+
"special": false
|
388 |
+
},
|
389 |
+
"48": {
|
390 |
+
"content": "<unused41>",
|
391 |
+
"lstrip": false,
|
392 |
+
"normalized": false,
|
393 |
+
"rstrip": false,
|
394 |
+
"single_word": false,
|
395 |
+
"special": false
|
396 |
+
},
|
397 |
+
"49": {
|
398 |
+
"content": "<unused42>",
|
399 |
+
"lstrip": false,
|
400 |
+
"normalized": false,
|
401 |
+
"rstrip": false,
|
402 |
+
"single_word": false,
|
403 |
+
"special": false
|
404 |
+
},
|
405 |
+
"50": {
|
406 |
+
"content": "<unused43>",
|
407 |
+
"lstrip": false,
|
408 |
+
"normalized": false,
|
409 |
+
"rstrip": false,
|
410 |
+
"single_word": false,
|
411 |
+
"special": false
|
412 |
+
},
|
413 |
+
"51": {
|
414 |
+
"content": "<unused44>",
|
415 |
+
"lstrip": false,
|
416 |
+
"normalized": false,
|
417 |
+
"rstrip": false,
|
418 |
+
"single_word": false,
|
419 |
+
"special": false
|
420 |
+
},
|
421 |
+
"52": {
|
422 |
+
"content": "<unused45>",
|
423 |
+
"lstrip": false,
|
424 |
+
"normalized": false,
|
425 |
+
"rstrip": false,
|
426 |
+
"single_word": false,
|
427 |
+
"special": false
|
428 |
+
},
|
429 |
+
"53": {
|
430 |
+
"content": "<unused46>",
|
431 |
+
"lstrip": false,
|
432 |
+
"normalized": false,
|
433 |
+
"rstrip": false,
|
434 |
+
"single_word": false,
|
435 |
+
"special": false
|
436 |
+
},
|
437 |
+
"54": {
|
438 |
+
"content": "<unused47>",
|
439 |
+
"lstrip": false,
|
440 |
+
"normalized": false,
|
441 |
+
"rstrip": false,
|
442 |
+
"single_word": false,
|
443 |
+
"special": false
|
444 |
+
},
|
445 |
+
"55": {
|
446 |
+
"content": "<unused48>",
|
447 |
+
"lstrip": false,
|
448 |
+
"normalized": false,
|
449 |
+
"rstrip": false,
|
450 |
+
"single_word": false,
|
451 |
+
"special": false
|
452 |
+
},
|
453 |
+
"56": {
|
454 |
+
"content": "<unused49>",
|
455 |
+
"lstrip": false,
|
456 |
+
"normalized": false,
|
457 |
+
"rstrip": false,
|
458 |
+
"single_word": false,
|
459 |
+
"special": false
|
460 |
+
},
|
461 |
+
"57": {
|
462 |
+
"content": "<unused50>",
|
463 |
+
"lstrip": false,
|
464 |
+
"normalized": false,
|
465 |
+
"rstrip": false,
|
466 |
+
"single_word": false,
|
467 |
+
"special": false
|
468 |
+
},
|
469 |
+
"58": {
|
470 |
+
"content": "<unused51>",
|
471 |
+
"lstrip": false,
|
472 |
+
"normalized": false,
|
473 |
+
"rstrip": false,
|
474 |
+
"single_word": false,
|
475 |
+
"special": false
|
476 |
+
},
|
477 |
+
"59": {
|
478 |
+
"content": "<unused52>",
|
479 |
+
"lstrip": false,
|
480 |
+
"normalized": false,
|
481 |
+
"rstrip": false,
|
482 |
+
"single_word": false,
|
483 |
+
"special": false
|
484 |
+
},
|
485 |
+
"60": {
|
486 |
+
"content": "<unused53>",
|
487 |
+
"lstrip": false,
|
488 |
+
"normalized": false,
|
489 |
+
"rstrip": false,
|
490 |
+
"single_word": false,
|
491 |
+
"special": false
|
492 |
+
},
|
493 |
+
"61": {
|
494 |
+
"content": "<unused54>",
|
495 |
+
"lstrip": false,
|
496 |
+
"normalized": false,
|
497 |
+
"rstrip": false,
|
498 |
+
"single_word": false,
|
499 |
+
"special": false
|
500 |
+
},
|
501 |
+
"62": {
|
502 |
+
"content": "<unused55>",
|
503 |
+
"lstrip": false,
|
504 |
+
"normalized": false,
|
505 |
+
"rstrip": false,
|
506 |
+
"single_word": false,
|
507 |
+
"special": false
|
508 |
+
},
|
509 |
+
"63": {
|
510 |
+
"content": "<unused56>",
|
511 |
+
"lstrip": false,
|
512 |
+
"normalized": false,
|
513 |
+
"rstrip": false,
|
514 |
+
"single_word": false,
|
515 |
+
"special": false
|
516 |
+
},
|
517 |
+
"64": {
|
518 |
+
"content": "<unused57>",
|
519 |
+
"lstrip": false,
|
520 |
+
"normalized": false,
|
521 |
+
"rstrip": false,
|
522 |
+
"single_word": false,
|
523 |
+
"special": false
|
524 |
+
},
|
525 |
+
"65": {
|
526 |
+
"content": "<unused58>",
|
527 |
+
"lstrip": false,
|
528 |
+
"normalized": false,
|
529 |
+
"rstrip": false,
|
530 |
+
"single_word": false,
|
531 |
+
"special": false
|
532 |
+
},
|
533 |
+
"66": {
|
534 |
+
"content": "<unused59>",
|
535 |
+
"lstrip": false,
|
536 |
+
"normalized": false,
|
537 |
+
"rstrip": false,
|
538 |
+
"single_word": false,
|
539 |
+
"special": false
|
540 |
+
},
|
541 |
+
"67": {
|
542 |
+
"content": "<unused60>",
|
543 |
+
"lstrip": false,
|
544 |
+
"normalized": false,
|
545 |
+
"rstrip": false,
|
546 |
+
"single_word": false,
|
547 |
+
"special": false
|
548 |
+
},
|
549 |
+
"68": {
|
550 |
+
"content": "<unused61>",
|
551 |
+
"lstrip": false,
|
552 |
+
"normalized": false,
|
553 |
+
"rstrip": false,
|
554 |
+
"single_word": false,
|
555 |
+
"special": false
|
556 |
+
},
|
557 |
+
"69": {
|
558 |
+
"content": "<unused62>",
|
559 |
+
"lstrip": false,
|
560 |
+
"normalized": false,
|
561 |
+
"rstrip": false,
|
562 |
+
"single_word": false,
|
563 |
+
"special": false
|
564 |
+
},
|
565 |
+
"70": {
|
566 |
+
"content": "<unused63>",
|
567 |
+
"lstrip": false,
|
568 |
+
"normalized": false,
|
569 |
+
"rstrip": false,
|
570 |
+
"single_word": false,
|
571 |
+
"special": false
|
572 |
+
},
|
573 |
+
"71": {
|
574 |
+
"content": "<unused64>",
|
575 |
+
"lstrip": false,
|
576 |
+
"normalized": false,
|
577 |
+
"rstrip": false,
|
578 |
+
"single_word": false,
|
579 |
+
"special": false
|
580 |
+
},
|
581 |
+
"72": {
|
582 |
+
"content": "<unused65>",
|
583 |
+
"lstrip": false,
|
584 |
+
"normalized": false,
|
585 |
+
"rstrip": false,
|
586 |
+
"single_word": false,
|
587 |
+
"special": false
|
588 |
+
},
|
589 |
+
"73": {
|
590 |
+
"content": "<unused66>",
|
591 |
+
"lstrip": false,
|
592 |
+
"normalized": false,
|
593 |
+
"rstrip": false,
|
594 |
+
"single_word": false,
|
595 |
+
"special": false
|
596 |
+
},
|
597 |
+
"74": {
|
598 |
+
"content": "<unused67>",
|
599 |
+
"lstrip": false,
|
600 |
+
"normalized": false,
|
601 |
+
"rstrip": false,
|
602 |
+
"single_word": false,
|
603 |
+
"special": false
|
604 |
+
},
|
605 |
+
"75": {
|
606 |
+
"content": "<unused68>",
|
607 |
+
"lstrip": false,
|
608 |
+
"normalized": false,
|
609 |
+
"rstrip": false,
|
610 |
+
"single_word": false,
|
611 |
+
"special": false
|
612 |
+
},
|
613 |
+
"76": {
|
614 |
+
"content": "<unused69>",
|
615 |
+
"lstrip": false,
|
616 |
+
"normalized": false,
|
617 |
+
"rstrip": false,
|
618 |
+
"single_word": false,
|
619 |
+
"special": false
|
620 |
+
},
|
621 |
+
"77": {
|
622 |
+
"content": "<unused70>",
|
623 |
+
"lstrip": false,
|
624 |
+
"normalized": false,
|
625 |
+
"rstrip": false,
|
626 |
+
"single_word": false,
|
627 |
+
"special": false
|
628 |
+
},
|
629 |
+
"78": {
|
630 |
+
"content": "<unused71>",
|
631 |
+
"lstrip": false,
|
632 |
+
"normalized": false,
|
633 |
+
"rstrip": false,
|
634 |
+
"single_word": false,
|
635 |
+
"special": false
|
636 |
+
},
|
637 |
+
"79": {
|
638 |
+
"content": "<unused72>",
|
639 |
+
"lstrip": false,
|
640 |
+
"normalized": false,
|
641 |
+
"rstrip": false,
|
642 |
+
"single_word": false,
|
643 |
+
"special": false
|
644 |
+
},
|
645 |
+
"80": {
|
646 |
+
"content": "<unused73>",
|
647 |
+
"lstrip": false,
|
648 |
+
"normalized": false,
|
649 |
+
"rstrip": false,
|
650 |
+
"single_word": false,
|
651 |
+
"special": false
|
652 |
+
},
|
653 |
+
"81": {
|
654 |
+
"content": "<unused74>",
|
655 |
+
"lstrip": false,
|
656 |
+
"normalized": false,
|
657 |
+
"rstrip": false,
|
658 |
+
"single_word": false,
|
659 |
+
"special": false
|
660 |
+
},
|
661 |
+
"82": {
|
662 |
+
"content": "<unused75>",
|
663 |
+
"lstrip": false,
|
664 |
+
"normalized": false,
|
665 |
+
"rstrip": false,
|
666 |
+
"single_word": false,
|
667 |
+
"special": false
|
668 |
+
},
|
669 |
+
"83": {
|
670 |
+
"content": "<unused76>",
|
671 |
+
"lstrip": false,
|
672 |
+
"normalized": false,
|
673 |
+
"rstrip": false,
|
674 |
+
"single_word": false,
|
675 |
+
"special": false
|
676 |
+
},
|
677 |
+
"84": {
|
678 |
+
"content": "<unused77>",
|
679 |
+
"lstrip": false,
|
680 |
+
"normalized": false,
|
681 |
+
"rstrip": false,
|
682 |
+
"single_word": false,
|
683 |
+
"special": false
|
684 |
+
},
|
685 |
+
"85": {
|
686 |
+
"content": "<unused78>",
|
687 |
+
"lstrip": false,
|
688 |
+
"normalized": false,
|
689 |
+
"rstrip": false,
|
690 |
+
"single_word": false,
|
691 |
+
"special": false
|
692 |
+
},
|
693 |
+
"86": {
|
694 |
+
"content": "<unused79>",
|
695 |
+
"lstrip": false,
|
696 |
+
"normalized": false,
|
697 |
+
"rstrip": false,
|
698 |
+
"single_word": false,
|
699 |
+
"special": false
|
700 |
+
},
|
701 |
+
"87": {
|
702 |
+
"content": "<unused80>",
|
703 |
+
"lstrip": false,
|
704 |
+
"normalized": false,
|
705 |
+
"rstrip": false,
|
706 |
+
"single_word": false,
|
707 |
+
"special": false
|
708 |
+
},
|
709 |
+
"88": {
|
710 |
+
"content": "<unused81>",
|
711 |
+
"lstrip": false,
|
712 |
+
"normalized": false,
|
713 |
+
"rstrip": false,
|
714 |
+
"single_word": false,
|
715 |
+
"special": false
|
716 |
+
},
|
717 |
+
"89": {
|
718 |
+
"content": "<unused82>",
|
719 |
+
"lstrip": false,
|
720 |
+
"normalized": false,
|
721 |
+
"rstrip": false,
|
722 |
+
"single_word": false,
|
723 |
+
"special": false
|
724 |
+
},
|
725 |
+
"90": {
|
726 |
+
"content": "<unused83>",
|
727 |
+
"lstrip": false,
|
728 |
+
"normalized": false,
|
729 |
+
"rstrip": false,
|
730 |
+
"single_word": false,
|
731 |
+
"special": false
|
732 |
+
},
|
733 |
+
"91": {
|
734 |
+
"content": "<unused84>",
|
735 |
+
"lstrip": false,
|
736 |
+
"normalized": false,
|
737 |
+
"rstrip": false,
|
738 |
+
"single_word": false,
|
739 |
+
"special": false
|
740 |
+
},
|
741 |
+
"92": {
|
742 |
+
"content": "<unused85>",
|
743 |
+
"lstrip": false,
|
744 |
+
"normalized": false,
|
745 |
+
"rstrip": false,
|
746 |
+
"single_word": false,
|
747 |
+
"special": false
|
748 |
+
},
|
749 |
+
"93": {
|
750 |
+
"content": "<unused86>",
|
751 |
+
"lstrip": false,
|
752 |
+
"normalized": false,
|
753 |
+
"rstrip": false,
|
754 |
+
"single_word": false,
|
755 |
+
"special": false
|
756 |
+
},
|
757 |
+
"94": {
|
758 |
+
"content": "<unused87>",
|
759 |
+
"lstrip": false,
|
760 |
+
"normalized": false,
|
761 |
+
"rstrip": false,
|
762 |
+
"single_word": false,
|
763 |
+
"special": false
|
764 |
+
},
|
765 |
+
"95": {
|
766 |
+
"content": "<unused88>",
|
767 |
+
"lstrip": false,
|
768 |
+
"normalized": false,
|
769 |
+
"rstrip": false,
|
770 |
+
"single_word": false,
|
771 |
+
"special": false
|
772 |
+
},
|
773 |
+
"96": {
|
774 |
+
"content": "<unused89>",
|
775 |
+
"lstrip": false,
|
776 |
+
"normalized": false,
|
777 |
+
"rstrip": false,
|
778 |
+
"single_word": false,
|
779 |
+
"special": false
|
780 |
+
},
|
781 |
+
"97": {
|
782 |
+
"content": "<unused90>",
|
783 |
+
"lstrip": false,
|
784 |
+
"normalized": false,
|
785 |
+
"rstrip": false,
|
786 |
+
"single_word": false,
|
787 |
+
"special": false
|
788 |
+
},
|
789 |
+
"98": {
|
790 |
+
"content": "<unused91>",
|
791 |
+
"lstrip": false,
|
792 |
+
"normalized": false,
|
793 |
+
"rstrip": false,
|
794 |
+
"single_word": false,
|
795 |
+
"special": false
|
796 |
+
},
|
797 |
+
"99": {
|
798 |
+
"content": "<unused92>",
|
799 |
+
"lstrip": false,
|
800 |
+
"normalized": false,
|
801 |
+
"rstrip": false,
|
802 |
+
"single_word": false,
|
803 |
+
"special": false
|
804 |
+
},
|
805 |
+
"100": {
|
806 |
+
"content": "<unused93>",
|
807 |
+
"lstrip": false,
|
808 |
+
"normalized": false,
|
809 |
+
"rstrip": false,
|
810 |
+
"single_word": false,
|
811 |
+
"special": false
|
812 |
+
},
|
813 |
+
"101": {
|
814 |
+
"content": "<unused94>",
|
815 |
+
"lstrip": false,
|
816 |
+
"normalized": false,
|
817 |
+
"rstrip": false,
|
818 |
+
"single_word": false,
|
819 |
+
"special": false
|
820 |
+
},
|
821 |
+
"102": {
|
822 |
+
"content": "<unused95>",
|
823 |
+
"lstrip": false,
|
824 |
+
"normalized": false,
|
825 |
+
"rstrip": false,
|
826 |
+
"single_word": false,
|
827 |
+
"special": false
|
828 |
+
},
|
829 |
+
"103": {
|
830 |
+
"content": "<unused96>",
|
831 |
+
"lstrip": false,
|
832 |
+
"normalized": false,
|
833 |
+
"rstrip": false,
|
834 |
+
"single_word": false,
|
835 |
+
"special": false
|
836 |
+
},
|
837 |
+
"104": {
|
838 |
+
"content": "<unused97>",
|
839 |
+
"lstrip": false,
|
840 |
+
"normalized": false,
|
841 |
+
"rstrip": false,
|
842 |
+
"single_word": false,
|
843 |
+
"special": false
|
844 |
+
},
|
845 |
+
"105": {
|
846 |
+
"content": "<unused98>",
|
847 |
+
"lstrip": false,
|
848 |
+
"normalized": false,
|
849 |
+
"rstrip": false,
|
850 |
+
"single_word": false,
|
851 |
+
"special": false
|
852 |
+
},
|
853 |
"106": {
|
854 |
"content": "<start_of_turn>",
|
855 |
"lstrip": false,
|
856 |
"normalized": false,
|
857 |
"rstrip": false,
|
858 |
"single_word": false,
|
859 |
+
"special": true
|
860 |
+
},
|
861 |
+
"107": {
|
862 |
+
"content": "<end_of_turn>",
|
863 |
+
"lstrip": false,
|
864 |
+
"normalized": false,
|
865 |
+
"rstrip": false,
|
866 |
+
"single_word": false,
|
867 |
+
"special": true
|
868 |
+
},
|
869 |
+
"108": {
|
870 |
+
"content": "\n",
|
871 |
+
"lstrip": false,
|
872 |
+
"normalized": false,
|
873 |
+
"rstrip": false,
|
874 |
+
"single_word": false,
|
875 |
+
"special": false
|
876 |
+
},
|
877 |
+
"109": {
|
878 |
+
"content": "\n\n",
|
879 |
+
"lstrip": false,
|
880 |
+
"normalized": false,
|
881 |
+
"rstrip": false,
|
882 |
+
"single_word": false,
|
883 |
+
"special": false
|
884 |
+
},
|
885 |
+
"110": {
|
886 |
+
"content": "\n\n\n",
|
887 |
+
"lstrip": false,
|
888 |
+
"normalized": false,
|
889 |
+
"rstrip": false,
|
890 |
+
"single_word": false,
|
891 |
+
"special": false
|
892 |
+
},
|
893 |
+
"111": {
|
894 |
+
"content": "\n\n\n\n",
|
895 |
+
"lstrip": false,
|
896 |
+
"normalized": false,
|
897 |
+
"rstrip": false,
|
898 |
+
"single_word": false,
|
899 |
+
"special": false
|
900 |
+
},
|
901 |
+
"112": {
|
902 |
+
"content": "\n\n\n\n\n",
|
903 |
+
"lstrip": false,
|
904 |
+
"normalized": false,
|
905 |
+
"rstrip": false,
|
906 |
+
"single_word": false,
|
907 |
+
"special": false
|
908 |
+
},
|
909 |
+
"113": {
|
910 |
+
"content": "\n\n\n\n\n\n",
|
911 |
+
"lstrip": false,
|
912 |
+
"normalized": false,
|
913 |
+
"rstrip": false,
|
914 |
+
"single_word": false,
|
915 |
+
"special": false
|
916 |
+
},
|
917 |
+
"114": {
|
918 |
+
"content": "\n\n\n\n\n\n\n",
|
919 |
+
"lstrip": false,
|
920 |
+
"normalized": false,
|
921 |
+
"rstrip": false,
|
922 |
+
"single_word": false,
|
923 |
+
"special": false
|
924 |
+
},
|
925 |
+
"115": {
|
926 |
+
"content": "\n\n\n\n\n\n\n\n",
|
927 |
+
"lstrip": false,
|
928 |
+
"normalized": false,
|
929 |
+
"rstrip": false,
|
930 |
+
"single_word": false,
|
931 |
+
"special": false
|
932 |
+
},
|
933 |
+
"116": {
|
934 |
+
"content": "\n\n\n\n\n\n\n\n\n",
|
935 |
+
"lstrip": false,
|
936 |
+
"normalized": false,
|
937 |
+
"rstrip": false,
|
938 |
+
"single_word": false,
|
939 |
+
"special": false
|
940 |
+
},
|
941 |
+
"117": {
|
942 |
+
"content": "\n\n\n\n\n\n\n\n\n\n",
|
943 |
+
"lstrip": false,
|
944 |
+
"normalized": false,
|
945 |
+
"rstrip": false,
|
946 |
+
"single_word": false,
|
947 |
+
"special": false
|
948 |
+
},
|
949 |
+
"118": {
|
950 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n",
|
951 |
+
"lstrip": false,
|
952 |
+
"normalized": false,
|
953 |
+
"rstrip": false,
|
954 |
+
"single_word": false,
|
955 |
+
"special": false
|
956 |
+
},
|
957 |
+
"119": {
|
958 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n",
|
959 |
+
"lstrip": false,
|
960 |
+
"normalized": false,
|
961 |
+
"rstrip": false,
|
962 |
+
"single_word": false,
|
963 |
+
"special": false
|
964 |
+
},
|
965 |
+
"120": {
|
966 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
967 |
+
"lstrip": false,
|
968 |
+
"normalized": false,
|
969 |
+
"rstrip": false,
|
970 |
+
"single_word": false,
|
971 |
+
"special": false
|
972 |
+
},
|
973 |
+
"121": {
|
974 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
975 |
+
"lstrip": false,
|
976 |
+
"normalized": false,
|
977 |
+
"rstrip": false,
|
978 |
+
"single_word": false,
|
979 |
+
"special": false
|
980 |
+
},
|
981 |
+
"122": {
|
982 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
983 |
+
"lstrip": false,
|
984 |
+
"normalized": false,
|
985 |
+
"rstrip": false,
|
986 |
+
"single_word": false,
|
987 |
+
"special": false
|
988 |
+
},
|
989 |
+
"123": {
|
990 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
991 |
+
"lstrip": false,
|
992 |
+
"normalized": false,
|
993 |
+
"rstrip": false,
|
994 |
+
"single_word": false,
|
995 |
+
"special": false
|
996 |
+
},
|
997 |
+
"124": {
|
998 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
999 |
+
"lstrip": false,
|
1000 |
+
"normalized": false,
|
1001 |
+
"rstrip": false,
|
1002 |
+
"single_word": false,
|
1003 |
+
"special": false
|
1004 |
+
},
|
1005 |
+
"125": {
|
1006 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1007 |
+
"lstrip": false,
|
1008 |
+
"normalized": false,
|
1009 |
+
"rstrip": false,
|
1010 |
+
"single_word": false,
|
1011 |
+
"special": false
|
1012 |
+
},
|
1013 |
+
"126": {
|
1014 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1015 |
+
"lstrip": false,
|
1016 |
+
"normalized": false,
|
1017 |
+
"rstrip": false,
|
1018 |
+
"single_word": false,
|
1019 |
+
"special": false
|
1020 |
+
},
|
1021 |
+
"127": {
|
1022 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1023 |
+
"lstrip": false,
|
1024 |
+
"normalized": false,
|
1025 |
+
"rstrip": false,
|
1026 |
+
"single_word": false,
|
1027 |
+
"special": false
|
1028 |
+
},
|
1029 |
+
"128": {
|
1030 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1031 |
+
"lstrip": false,
|
1032 |
+
"normalized": false,
|
1033 |
+
"rstrip": false,
|
1034 |
+
"single_word": false,
|
1035 |
+
"special": false
|
1036 |
+
},
|
1037 |
+
"129": {
|
1038 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1039 |
+
"lstrip": false,
|
1040 |
+
"normalized": false,
|
1041 |
+
"rstrip": false,
|
1042 |
+
"single_word": false,
|
1043 |
+
"special": false
|
1044 |
+
},
|
1045 |
+
"130": {
|
1046 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1047 |
+
"lstrip": false,
|
1048 |
+
"normalized": false,
|
1049 |
+
"rstrip": false,
|
1050 |
+
"single_word": false,
|
1051 |
+
"special": false
|
1052 |
+
},
|
1053 |
+
"131": {
|
1054 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1055 |
+
"lstrip": false,
|
1056 |
+
"normalized": false,
|
1057 |
+
"rstrip": false,
|
1058 |
+
"single_word": false,
|
1059 |
+
"special": false
|
1060 |
+
},
|
1061 |
+
"132": {
|
1062 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1063 |
+
"lstrip": false,
|
1064 |
+
"normalized": false,
|
1065 |
+
"rstrip": false,
|
1066 |
+
"single_word": false,
|
1067 |
+
"special": false
|
1068 |
+
},
|
1069 |
+
"133": {
|
1070 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1071 |
+
"lstrip": false,
|
1072 |
+
"normalized": false,
|
1073 |
+
"rstrip": false,
|
1074 |
+
"single_word": false,
|
1075 |
+
"special": false
|
1076 |
+
},
|
1077 |
+
"134": {
|
1078 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1079 |
+
"lstrip": false,
|
1080 |
+
"normalized": false,
|
1081 |
+
"rstrip": false,
|
1082 |
+
"single_word": false,
|
1083 |
+
"special": false
|
1084 |
+
},
|
1085 |
+
"135": {
|
1086 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1087 |
+
"lstrip": false,
|
1088 |
+
"normalized": false,
|
1089 |
+
"rstrip": false,
|
1090 |
+
"single_word": false,
|
1091 |
+
"special": false
|
1092 |
+
},
|
1093 |
+
"136": {
|
1094 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1095 |
+
"lstrip": false,
|
1096 |
+
"normalized": false,
|
1097 |
+
"rstrip": false,
|
1098 |
+
"single_word": false,
|
1099 |
+
"special": false
|
1100 |
+
},
|
1101 |
+
"137": {
|
1102 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1103 |
+
"lstrip": false,
|
1104 |
+
"normalized": false,
|
1105 |
+
"rstrip": false,
|
1106 |
+
"single_word": false,
|
1107 |
+
"special": false
|
1108 |
+
},
|
1109 |
+
"138": {
|
1110 |
+
"content": "\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n",
|
1111 |
+
"lstrip": false,
|
1112 |
+
"normalized": false,
|
1113 |
+
"rstrip": false,
|
1114 |
+
"single_word": false,
|
1115 |
+
"special": false
|
1116 |
+
},
|
1117 |
+
"139": {
|
1118 |
+
"content": "▁▁",
|
1119 |
+
"lstrip": false,
|
1120 |
+
"normalized": false,
|
1121 |
+
"rstrip": false,
|
1122 |
+
"single_word": false,
|
1123 |
+
"special": false
|
1124 |
+
},
|
1125 |
+
"140": {
|
1126 |
+
"content": "▁▁▁",
|
1127 |
+
"lstrip": false,
|
1128 |
+
"normalized": false,
|
1129 |
+
"rstrip": false,
|
1130 |
+
"single_word": false,
|
1131 |
+
"special": false
|
1132 |
+
},
|
1133 |
+
"141": {
|
1134 |
+
"content": "▁▁▁▁",
|
1135 |
+
"lstrip": false,
|
1136 |
+
"normalized": false,
|
1137 |
+
"rstrip": false,
|
1138 |
+
"single_word": false,
|
1139 |
+
"special": false
|
1140 |
+
},
|
1141 |
+
"142": {
|
1142 |
+
"content": "▁▁▁▁▁",
|
1143 |
+
"lstrip": false,
|
1144 |
+
"normalized": false,
|
1145 |
+
"rstrip": false,
|
1146 |
+
"single_word": false,
|
1147 |
+
"special": false
|
1148 |
+
},
|
1149 |
+
"143": {
|
1150 |
+
"content": "▁▁▁▁▁▁",
|
1151 |
+
"lstrip": false,
|
1152 |
+
"normalized": false,
|
1153 |
+
"rstrip": false,
|
1154 |
+
"single_word": false,
|
1155 |
+
"special": false
|
1156 |
+
},
|
1157 |
+
"144": {
|
1158 |
+
"content": "▁▁▁▁▁▁▁",
|
1159 |
+
"lstrip": false,
|
1160 |
+
"normalized": false,
|
1161 |
+
"rstrip": false,
|
1162 |
+
"single_word": false,
|
1163 |
+
"special": false
|
1164 |
+
},
|
1165 |
+
"145": {
|
1166 |
+
"content": "▁▁▁▁▁▁▁▁",
|
1167 |
+
"lstrip": false,
|
1168 |
+
"normalized": false,
|
1169 |
+
"rstrip": false,
|
1170 |
+
"single_word": false,
|
1171 |
+
"special": false
|
1172 |
+
},
|
1173 |
+
"146": {
|
1174 |
+
"content": "▁▁▁▁▁▁▁▁▁",
|
1175 |
+
"lstrip": false,
|
1176 |
+
"normalized": false,
|
1177 |
+
"rstrip": false,
|
1178 |
+
"single_word": false,
|
1179 |
+
"special": false
|
1180 |
+
},
|
1181 |
+
"147": {
|
1182 |
+
"content": "▁▁▁▁▁▁▁▁▁▁",
|
1183 |
+
"lstrip": false,
|
1184 |
+
"normalized": false,
|
1185 |
+
"rstrip": false,
|
1186 |
+
"single_word": false,
|
1187 |
+
"special": false
|
1188 |
+
},
|
1189 |
+
"148": {
|
1190 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁",
|
1191 |
+
"lstrip": false,
|
1192 |
+
"normalized": false,
|
1193 |
+
"rstrip": false,
|
1194 |
+
"single_word": false,
|
1195 |
+
"special": false
|
1196 |
+
},
|
1197 |
+
"149": {
|
1198 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁",
|
1199 |
+
"lstrip": false,
|
1200 |
+
"normalized": false,
|
1201 |
+
"rstrip": false,
|
1202 |
+
"single_word": false,
|
1203 |
+
"special": false
|
1204 |
+
},
|
1205 |
+
"150": {
|
1206 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1207 |
+
"lstrip": false,
|
1208 |
+
"normalized": false,
|
1209 |
+
"rstrip": false,
|
1210 |
+
"single_word": false,
|
1211 |
+
"special": false
|
1212 |
+
},
|
1213 |
+
"151": {
|
1214 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1215 |
+
"lstrip": false,
|
1216 |
+
"normalized": false,
|
1217 |
+
"rstrip": false,
|
1218 |
+
"single_word": false,
|
1219 |
+
"special": false
|
1220 |
+
},
|
1221 |
+
"152": {
|
1222 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1223 |
+
"lstrip": false,
|
1224 |
+
"normalized": false,
|
1225 |
+
"rstrip": false,
|
1226 |
+
"single_word": false,
|
1227 |
+
"special": false
|
1228 |
+
},
|
1229 |
+
"153": {
|
1230 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1231 |
+
"lstrip": false,
|
1232 |
+
"normalized": false,
|
1233 |
+
"rstrip": false,
|
1234 |
+
"single_word": false,
|
1235 |
+
"special": false
|
1236 |
+
},
|
1237 |
+
"154": {
|
1238 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1239 |
+
"lstrip": false,
|
1240 |
+
"normalized": false,
|
1241 |
+
"rstrip": false,
|
1242 |
+
"single_word": false,
|
1243 |
+
"special": false
|
1244 |
+
},
|
1245 |
+
"155": {
|
1246 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1247 |
+
"lstrip": false,
|
1248 |
+
"normalized": false,
|
1249 |
+
"rstrip": false,
|
1250 |
+
"single_word": false,
|
1251 |
+
"special": false
|
1252 |
+
},
|
1253 |
+
"156": {
|
1254 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1255 |
+
"lstrip": false,
|
1256 |
+
"normalized": false,
|
1257 |
+
"rstrip": false,
|
1258 |
+
"single_word": false,
|
1259 |
+
"special": false
|
1260 |
+
},
|
1261 |
+
"157": {
|
1262 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1263 |
+
"lstrip": false,
|
1264 |
+
"normalized": false,
|
1265 |
+
"rstrip": false,
|
1266 |
+
"single_word": false,
|
1267 |
+
"special": false
|
1268 |
+
},
|
1269 |
+
"158": {
|
1270 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1271 |
+
"lstrip": false,
|
1272 |
+
"normalized": false,
|
1273 |
+
"rstrip": false,
|
1274 |
+
"single_word": false,
|
1275 |
+
"special": false
|
1276 |
+
},
|
1277 |
+
"159": {
|
1278 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1279 |
+
"lstrip": false,
|
1280 |
+
"normalized": false,
|
1281 |
+
"rstrip": false,
|
1282 |
+
"single_word": false,
|
1283 |
+
"special": false
|
1284 |
+
},
|
1285 |
+
"160": {
|
1286 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1287 |
+
"lstrip": false,
|
1288 |
+
"normalized": false,
|
1289 |
+
"rstrip": false,
|
1290 |
+
"single_word": false,
|
1291 |
+
"special": false
|
1292 |
+
},
|
1293 |
+
"161": {
|
1294 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1295 |
+
"lstrip": false,
|
1296 |
+
"normalized": false,
|
1297 |
+
"rstrip": false,
|
1298 |
+
"single_word": false,
|
1299 |
+
"special": false
|
1300 |
+
},
|
1301 |
+
"162": {
|
1302 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1303 |
+
"lstrip": false,
|
1304 |
+
"normalized": false,
|
1305 |
+
"rstrip": false,
|
1306 |
+
"single_word": false,
|
1307 |
+
"special": false
|
1308 |
+
},
|
1309 |
+
"163": {
|
1310 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1311 |
+
"lstrip": false,
|
1312 |
+
"normalized": false,
|
1313 |
+
"rstrip": false,
|
1314 |
+
"single_word": false,
|
1315 |
+
"special": false
|
1316 |
+
},
|
1317 |
+
"164": {
|
1318 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1319 |
+
"lstrip": false,
|
1320 |
+
"normalized": false,
|
1321 |
+
"rstrip": false,
|
1322 |
+
"single_word": false,
|
1323 |
+
"special": false
|
1324 |
+
},
|
1325 |
+
"165": {
|
1326 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1327 |
+
"lstrip": false,
|
1328 |
+
"normalized": false,
|
1329 |
+
"rstrip": false,
|
1330 |
+
"single_word": false,
|
1331 |
+
"special": false
|
1332 |
},
|
1333 |
+
"166": {
|
1334 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1335 |
"lstrip": false,
|
1336 |
"normalized": false,
|
1337 |
"rstrip": false,
|
1338 |
"single_word": false,
|
1339 |
+
"special": false
|
1340 |
+
},
|
1341 |
+
"167": {
|
1342 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1343 |
+
"lstrip": false,
|
1344 |
+
"normalized": false,
|
1345 |
+
"rstrip": false,
|
1346 |
+
"single_word": false,
|
1347 |
+
"special": false
|
1348 |
+
},
|
1349 |
+
"168": {
|
1350 |
+
"content": "▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁▁",
|
1351 |
+
"lstrip": false,
|
1352 |
+
"normalized": false,
|
1353 |
+
"rstrip": false,
|
1354 |
+
"single_word": false,
|
1355 |
+
"special": false
|
1356 |
+
},
|
1357 |
+
"169": {
|
1358 |
+
"content": "<table>",
|
1359 |
+
"lstrip": false,
|
1360 |
+
"normalized": false,
|
1361 |
+
"rstrip": false,
|
1362 |
+
"single_word": false,
|
1363 |
+
"special": false
|
1364 |
+
},
|
1365 |
+
"170": {
|
1366 |
+
"content": "<caption>",
|
1367 |
+
"lstrip": false,
|
1368 |
+
"normalized": false,
|
1369 |
+
"rstrip": false,
|
1370 |
+
"single_word": false,
|
1371 |
+
"special": false
|
1372 |
+
},
|
1373 |
+
"171": {
|
1374 |
+
"content": "<thead>",
|
1375 |
+
"lstrip": false,
|
1376 |
+
"normalized": false,
|
1377 |
+
"rstrip": false,
|
1378 |
+
"single_word": false,
|
1379 |
+
"special": false
|
1380 |
+
},
|
1381 |
+
"172": {
|
1382 |
+
"content": "<tbody>",
|
1383 |
+
"lstrip": false,
|
1384 |
+
"normalized": false,
|
1385 |
+
"rstrip": false,
|
1386 |
+
"single_word": false,
|
1387 |
+
"special": false
|
1388 |
+
},
|
1389 |
+
"173": {
|
1390 |
+
"content": "<tfoot>",
|
1391 |
+
"lstrip": false,
|
1392 |
+
"normalized": false,
|
1393 |
+
"rstrip": false,
|
1394 |
+
"single_word": false,
|
1395 |
+
"special": false
|
1396 |
+
},
|
1397 |
+
"174": {
|
1398 |
+
"content": "<tr>",
|
1399 |
+
"lstrip": false,
|
1400 |
+
"normalized": false,
|
1401 |
+
"rstrip": false,
|
1402 |
+
"single_word": false,
|
1403 |
+
"special": false
|
1404 |
+
},
|
1405 |
+
"175": {
|
1406 |
+
"content": "<th>",
|
1407 |
+
"lstrip": false,
|
1408 |
+
"normalized": false,
|
1409 |
+
"rstrip": false,
|
1410 |
+
"single_word": false,
|
1411 |
+
"special": false
|
1412 |
+
},
|
1413 |
+
"176": {
|
1414 |
+
"content": "<td>",
|
1415 |
+
"lstrip": false,
|
1416 |
+
"normalized": false,
|
1417 |
+
"rstrip": false,
|
1418 |
+
"single_word": false,
|
1419 |
+
"special": false
|
1420 |
+
},
|
1421 |
+
"177": {
|
1422 |
+
"content": "</table>",
|
1423 |
+
"lstrip": false,
|
1424 |
+
"normalized": false,
|
1425 |
+
"rstrip": false,
|
1426 |
+
"single_word": false,
|
1427 |
+
"special": false
|
1428 |
+
},
|
1429 |
+
"178": {
|
1430 |
+
"content": "</caption>",
|
1431 |
+
"lstrip": false,
|
1432 |
+
"normalized": false,
|
1433 |
+
"rstrip": false,
|
1434 |
+
"single_word": false,
|
1435 |
+
"special": false
|
1436 |
+
},
|
1437 |
+
"179": {
|
1438 |
+
"content": "</thead>",
|
1439 |
+
"lstrip": false,
|
1440 |
+
"normalized": false,
|
1441 |
+
"rstrip": false,
|
1442 |
+
"single_word": false,
|
1443 |
+
"special": false
|
1444 |
+
},
|
1445 |
+
"180": {
|
1446 |
+
"content": "</tbody>",
|
1447 |
+
"lstrip": false,
|
1448 |
+
"normalized": false,
|
1449 |
+
"rstrip": false,
|
1450 |
+
"single_word": false,
|
1451 |
+
"special": false
|
1452 |
+
},
|
1453 |
+
"181": {
|
1454 |
+
"content": "</tfoot>",
|
1455 |
+
"lstrip": false,
|
1456 |
+
"normalized": false,
|
1457 |
+
"rstrip": false,
|
1458 |
+
"single_word": false,
|
1459 |
+
"special": false
|
1460 |
+
},
|
1461 |
+
"182": {
|
1462 |
+
"content": "</tr>",
|
1463 |
+
"lstrip": false,
|
1464 |
+
"normalized": false,
|
1465 |
+
"rstrip": false,
|
1466 |
+
"single_word": false,
|
1467 |
+
"special": false
|
1468 |
+
},
|
1469 |
+
"183": {
|
1470 |
+
"content": "</th>",
|
1471 |
+
"lstrip": false,
|
1472 |
+
"normalized": false,
|
1473 |
+
"rstrip": false,
|
1474 |
+
"single_word": false,
|
1475 |
+
"special": false
|
1476 |
+
},
|
1477 |
+
"184": {
|
1478 |
+
"content": "</td>",
|
1479 |
+
"lstrip": false,
|
1480 |
+
"normalized": false,
|
1481 |
+
"rstrip": false,
|
1482 |
+
"single_word": false,
|
1483 |
+
"special": false
|
1484 |
+
},
|
1485 |
+
"185": {
|
1486 |
+
"content": "<h1>",
|
1487 |
+
"lstrip": false,
|
1488 |
+
"normalized": false,
|
1489 |
+
"rstrip": false,
|
1490 |
+
"single_word": false,
|
1491 |
+
"special": false
|
1492 |
+
},
|
1493 |
+
"186": {
|
1494 |
+
"content": "<h2>",
|
1495 |
+
"lstrip": false,
|
1496 |
+
"normalized": false,
|
1497 |
+
"rstrip": false,
|
1498 |
+
"single_word": false,
|
1499 |
+
"special": false
|
1500 |
+
},
|
1501 |
+
"187": {
|
1502 |
+
"content": "<h3>",
|
1503 |
+
"lstrip": false,
|
1504 |
+
"normalized": false,
|
1505 |
+
"rstrip": false,
|
1506 |
+
"single_word": false,
|
1507 |
+
"special": false
|
1508 |
+
},
|
1509 |
+
"188": {
|
1510 |
+
"content": "<h4>",
|
1511 |
+
"lstrip": false,
|
1512 |
+
"normalized": false,
|
1513 |
+
"rstrip": false,
|
1514 |
+
"single_word": false,
|
1515 |
+
"special": false
|
1516 |
+
},
|
1517 |
+
"189": {
|
1518 |
+
"content": "<h5>",
|
1519 |
+
"lstrip": false,
|
1520 |
+
"normalized": false,
|
1521 |
+
"rstrip": false,
|
1522 |
+
"single_word": false,
|
1523 |
+
"special": false
|
1524 |
+
},
|
1525 |
+
"190": {
|
1526 |
+
"content": "<h6>",
|
1527 |
+
"lstrip": false,
|
1528 |
+
"normalized": false,
|
1529 |
+
"rstrip": false,
|
1530 |
+
"single_word": false,
|
1531 |
+
"special": false
|
1532 |
+
},
|
1533 |
+
"191": {
|
1534 |
+
"content": "<blockquote>",
|
1535 |
+
"lstrip": false,
|
1536 |
+
"normalized": false,
|
1537 |
+
"rstrip": false,
|
1538 |
+
"single_word": false,
|
1539 |
+
"special": false
|
1540 |
+
},
|
1541 |
+
"192": {
|
1542 |
+
"content": "</h1>",
|
1543 |
+
"lstrip": false,
|
1544 |
+
"normalized": false,
|
1545 |
+
"rstrip": false,
|
1546 |
+
"single_word": false,
|
1547 |
+
"special": false
|
1548 |
+
},
|
1549 |
+
"193": {
|
1550 |
+
"content": "</h2>",
|
1551 |
+
"lstrip": false,
|
1552 |
+
"normalized": false,
|
1553 |
+
"rstrip": false,
|
1554 |
+
"single_word": false,
|
1555 |
+
"special": false
|
1556 |
+
},
|
1557 |
+
"194": {
|
1558 |
+
"content": "</h3>",
|
1559 |
+
"lstrip": false,
|
1560 |
+
"normalized": false,
|
1561 |
+
"rstrip": false,
|
1562 |
+
"single_word": false,
|
1563 |
+
"special": false
|
1564 |
+
},
|
1565 |
+
"195": {
|
1566 |
+
"content": "</h4>",
|
1567 |
+
"lstrip": false,
|
1568 |
+
"normalized": false,
|
1569 |
+
"rstrip": false,
|
1570 |
+
"single_word": false,
|
1571 |
+
"special": false
|
1572 |
+
},
|
1573 |
+
"196": {
|
1574 |
+
"content": "</h5>",
|
1575 |
+
"lstrip": false,
|
1576 |
+
"normalized": false,
|
1577 |
+
"rstrip": false,
|
1578 |
+
"single_word": false,
|
1579 |
+
"special": false
|
1580 |
+
},
|
1581 |
+
"197": {
|
1582 |
+
"content": "</h6>",
|
1583 |
+
"lstrip": false,
|
1584 |
+
"normalized": false,
|
1585 |
+
"rstrip": false,
|
1586 |
+
"single_word": false,
|
1587 |
+
"special": false
|
1588 |
+
},
|
1589 |
+
"198": {
|
1590 |
+
"content": "</blockquote>",
|
1591 |
+
"lstrip": false,
|
1592 |
+
"normalized": false,
|
1593 |
+
"rstrip": false,
|
1594 |
+
"single_word": false,
|
1595 |
+
"special": false
|
1596 |
+
},
|
1597 |
+
"199": {
|
1598 |
+
"content": "<strong>",
|
1599 |
+
"lstrip": false,
|
1600 |
+
"normalized": false,
|
1601 |
+
"rstrip": false,
|
1602 |
+
"single_word": false,
|
1603 |
+
"special": false
|
1604 |
+
},
|
1605 |
+
"200": {
|
1606 |
+
"content": "<em>",
|
1607 |
+
"lstrip": false,
|
1608 |
+
"normalized": false,
|
1609 |
+
"rstrip": false,
|
1610 |
+
"single_word": false,
|
1611 |
+
"special": false
|
1612 |
+
},
|
1613 |
+
"201": {
|
1614 |
+
"content": "<b>",
|
1615 |
+
"lstrip": false,
|
1616 |
+
"normalized": false,
|
1617 |
+
"rstrip": false,
|
1618 |
+
"single_word": false,
|
1619 |
+
"special": false
|
1620 |
+
},
|
1621 |
+
"202": {
|
1622 |
+
"content": "<i>",
|
1623 |
+
"lstrip": false,
|
1624 |
+
"normalized": false,
|
1625 |
+
"rstrip": false,
|
1626 |
+
"single_word": false,
|
1627 |
+
"special": false
|
1628 |
+
},
|
1629 |
+
"203": {
|
1630 |
+
"content": "<u>",
|
1631 |
+
"lstrip": false,
|
1632 |
+
"normalized": false,
|
1633 |
+
"rstrip": false,
|
1634 |
+
"single_word": false,
|
1635 |
+
"special": false
|
1636 |
+
},
|
1637 |
+
"204": {
|
1638 |
+
"content": "<s>",
|
1639 |
+
"lstrip": false,
|
1640 |
+
"normalized": false,
|
1641 |
+
"rstrip": false,
|
1642 |
+
"single_word": false,
|
1643 |
+
"special": false
|
1644 |
+
},
|
1645 |
+
"205": {
|
1646 |
+
"content": "<sub>",
|
1647 |
+
"lstrip": false,
|
1648 |
+
"normalized": false,
|
1649 |
+
"rstrip": false,
|
1650 |
+
"single_word": false,
|
1651 |
+
"special": false
|
1652 |
+
},
|
1653 |
+
"206": {
|
1654 |
+
"content": "<sup>",
|
1655 |
+
"lstrip": false,
|
1656 |
+
"normalized": false,
|
1657 |
+
"rstrip": false,
|
1658 |
+
"single_word": false,
|
1659 |
+
"special": false
|
1660 |
+
},
|
1661 |
+
"207": {
|
1662 |
+
"content": "<code>",
|
1663 |
+
"lstrip": false,
|
1664 |
+
"normalized": false,
|
1665 |
+
"rstrip": false,
|
1666 |
+
"single_word": false,
|
1667 |
+
"special": false
|
1668 |
+
},
|
1669 |
+
"208": {
|
1670 |
+
"content": "</strong>",
|
1671 |
+
"lstrip": false,
|
1672 |
+
"normalized": false,
|
1673 |
+
"rstrip": false,
|
1674 |
+
"single_word": false,
|
1675 |
+
"special": false
|
1676 |
+
},
|
1677 |
+
"209": {
|
1678 |
+
"content": "</em>",
|
1679 |
+
"lstrip": false,
|
1680 |
+
"normalized": false,
|
1681 |
+
"rstrip": false,
|
1682 |
+
"single_word": false,
|
1683 |
+
"special": false
|
1684 |
+
},
|
1685 |
+
"210": {
|
1686 |
+
"content": "</b>",
|
1687 |
+
"lstrip": false,
|
1688 |
+
"normalized": false,
|
1689 |
+
"rstrip": false,
|
1690 |
+
"single_word": false,
|
1691 |
+
"special": false
|
1692 |
+
},
|
1693 |
+
"211": {
|
1694 |
+
"content": "</i>",
|
1695 |
+
"lstrip": false,
|
1696 |
+
"normalized": false,
|
1697 |
+
"rstrip": false,
|
1698 |
+
"single_word": false,
|
1699 |
+
"special": false
|
1700 |
+
},
|
1701 |
+
"212": {
|
1702 |
+
"content": "</u>",
|
1703 |
+
"lstrip": false,
|
1704 |
+
"normalized": false,
|
1705 |
+
"rstrip": false,
|
1706 |
+
"single_word": false,
|
1707 |
+
"special": false
|
1708 |
+
},
|
1709 |
+
"213": {
|
1710 |
+
"content": "</s>",
|
1711 |
+
"lstrip": false,
|
1712 |
+
"normalized": false,
|
1713 |
+
"rstrip": false,
|
1714 |
+
"single_word": false,
|
1715 |
+
"special": false
|
1716 |
+
},
|
1717 |
+
"214": {
|
1718 |
+
"content": "</sub>",
|
1719 |
+
"lstrip": false,
|
1720 |
+
"normalized": false,
|
1721 |
+
"rstrip": false,
|
1722 |
+
"single_word": false,
|
1723 |
+
"special": false
|
1724 |
+
},
|
1725 |
+
"215": {
|
1726 |
+
"content": "</sup>",
|
1727 |
+
"lstrip": false,
|
1728 |
+
"normalized": false,
|
1729 |
+
"rstrip": false,
|
1730 |
+
"single_word": false,
|
1731 |
+
"special": false
|
1732 |
+
},
|
1733 |
+
"216": {
|
1734 |
+
"content": "</code>",
|
1735 |
+
"lstrip": false,
|
1736 |
+
"normalized": false,
|
1737 |
+
"rstrip": false,
|
1738 |
+
"single_word": false,
|
1739 |
+
"special": false
|
1740 |
}
|
1741 |
},
|
1742 |
"additional_special_tokens": [
|
|
|
1744 |
"<end_of_turn>"
|
1745 |
],
|
1746 |
"bos_token": "<bos>",
|
|
|
1747 |
"clean_up_tokenization_spaces": false,
|
1748 |
"eos_token": "<eos>",
|
|
|
1749 |
"model_max_length": 8192,
|
1750 |
"pad_token": "<unk>",
|
1751 |
"padding_side": "right",
|