initial release

Files changed (8) hide show

README.md ADDED Viewed

+---
+language:
+- "ko"
+tags:
+- "korean"
+- "token-classification"
+- "pos"
+- "dependency-parsing"
+datasets:
+- "universal_dependencies"
+license: "cc-by-sa-4.0"
+pipeline_tag: "token-classification"
+widget:
+- text: "홍시 맛이 나서 홍시라 생각한다."
+---
+# roberta-base-korean-morph-upos
+## Model Description
+This is a RoBERTa model for POS-tagging and dependency-parsing, derived from [klue/roberta-base](https://huggingface.co/klue/roberta-base) and [morphUD-korean](https://github.com/jungyeul/morphUD-korean). Every morpheme is tagged by [UPOS](https://universaldependencies.org/u/pos/)(Universal Part-Of-Speech).
+## How to Use
+```py
+from transformers import AutoTokenizer,AutoModelForTokenClassification,TokenClassificationPipeline
+tokenizer=AutoTokenizer.from_pretrained("KoichiYasuoka/roberta-base-korean-morph-upos")
+model=AutoModelForTokenClassification.from_pretrained("KoichiYasuoka/roberta-base-korean-morph-upos")
+pipeline=TokenClassificationPipeline(tokenizer=tokenizer,model=model,aggregation_strategy="simple")
+nlp=lambda x:[(x[t["start"]:t["end"]],t["entity_group"]) for t in pipeline(x)]
+print(nlp("홍시 맛이 나서 홍시라 생각한다."))
+```
+or
+```py
+import esupar
+nlp=esupar.load("KoichiYasuoka/roberta-base-korean-morph-upos")
+print(nlp("홍시 맛이 나서 홍시라 생각한다."))
+```
+## See Also
+[esupar](https://github.com/KoichiYasuoka/esupar): Tokenizer POS-tagger and Dependency-parser with BERT/RoBERTa/DeBERTa models

config.json ADDED Viewed

The diff for this file is too large to render. See raw diff

pytorch_model.bin ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:f7d1fc5576bab4b9c9c47791db5edc1b9c0f977645882469dd5bb599b816d656
+size 442924913

special_tokens_map.json ADDED Viewed

+{
+  "bos_token": "[CLS]",
+  "cls_token": "[CLS]",
+  "eos_token": "[SEP]",
+  "mask_token": "[MASK]",
+  "pad_token": "[PAD]",
+  "sep_token": "[SEP]",
+  "unk_token": "[UNK]"
+}

supar.model ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:4c0bf237ca850fab143c201dbc0e710887cb554efc04bc487304df682cf101ee
+size 490658725

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

+{
+  "bos_token": "[CLS]",
+  "cls_token": "[CLS]",
+  "do_basic_tokenize": true,
+  "do_lower_case": false,
+  "eos_token": "[SEP]",
+  "mask_token": "[MASK]",
+  "model_max_length": 512,
+  "never_split": null,
+  "pad_token": "[PAD]",
+  "sep_token": "[SEP]",
+  "strip_accents": null,
+  "tokenize_chinese_chars": true,
+  "tokenizer_class": "BertTokenizerFast",
+  "unk_token": "[UNK]"
+}

vocab.txt ADDED Viewed

The diff for this file is too large to render. See raw diff