Add model

Browse files

Files changed (6) hide show

README.md +92 -0
config.json +83 -0
fairseq/model.pt +3 -0
preprocessor_config.json +9 -0
pytorch_model.bin +3 -0
rinna.png +0 -0

README.md ADDED Viewed

	@@ -0,0 +1,92 @@

+---
+thumbnail: https://github.com/rinnakk/japanese-pretrained-models/blob/master/rinna.png
+language: ja
+license: apache-2.0
+datasets: reazon-research/reazonspeech
+inference: false
+tags:
+  - hubert
+  - speech
+---
+# `rinna/japanese-hubert-large`
+![rinna-icon](./rinna.png)
+# Overview
+This is a Japanese HuBERT Large model trained by [rinna Co., Ltd.](https://rinna.co.jp/)
+* **Model summary**
+  The model architecture is the same as the [original HuBERT Large model](https://huggingface.co/facebook/hubert-large-ll60k), which contains 24 transformer layers with 16 attention heads.
+  The model was trained using code from the [official repository](https://github.com/facebookresearch/fairseq/tree/main/examples/hubert), and the detailed training configuration can be found in the same repository and the [original paper](https://ieeexplore.ieee.org/document/9585401).
+* **Training**
+  The model was trained on approximately 19,000 hours of following Japanese speech corpus ReazonSpeech v1.
+  - [ReazonSpeech](https://huggingface.co/datasets/reazon-research/reazonspeech)
+* **Contributors**
+  - [Yukiya Hono](https://huggingface.co/yky-h)
+  - [Kentaro Mitsui](https://huggingface.co/Kentaro321)
+  - [Kei Sawada](https://huggingface.co/keisawada)
+---
+# How to use the model
+```python
+import soundfile as sf
+from transformers import AutoFeatureExtractor, AutoModel
+model_name = "rinna/japanese-hubert-large"
+feature_extractor = AutoFeatureExtractor.from_pretrained(model_name)
+model = AutoModel.from_pretrained(model_name)
+model.eval()
+raw_speech_16kHz, sr = sf.read(audio_file)
+inputs = feature_extractor(
+    raw_speech_16kHz,
+    return_tensors="pt",
+    sampling_rate=sr,
+)
+outputs = model(**inputs)
+print(f"Input:  {inputs.input_values.size()}")  # [1, #samples]
+print(f"Output: {outputs.last_hidden_state.size()}")  # [1, #frames, 1024]
+```
+A fairseq checkpoint file can also be available [here](https://huggingface.co/rinna/japanese-hubert-large/tree/main/fairseq).
+---
+# How to cite
+```bibtex
+@misc{rinna-japanese-hubert-large,
+  title={rinna/japanese-hubert-large},
+  author={Hono, Yukiya and Mitsui, Kentaro and Sawada, Kei},
+  url={https://huggingface.co/rinna/japanese-hubert-large}
+}
+```
+---
+# Citations
+```bibtex
+@article{hsu2021hubert,
+  author={Hsu, Wei-Ning and Bolte, Benjamin and Tsai, Yao-Hung Hubert and Lakhotia, Kushal and Salakhutdinov, Ruslan and Mohamed, Abdelrahman},
+  journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing},
+  title={HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units},
+  year={2021},
+  volume={29},
+  number={},
+  pages={3451-3460},
+  doi={10.1109/TASLP.2021.3122291}
+}
+```
+---
+# License
+[The Apache 2.0 license](https://www.apache.org/licenses/LICENSE-2.0)

config.json ADDED Viewed

	@@ -0,0 +1,83 @@

+{
+  "_name_or_path": "rinna/japanese-hubert-large",
+  "activation_dropout": 0.0,
+  "apply_spec_augment": true,
+  "architectures": [
+    "HubertModel"
+  ],
+  "attention_dropout": 0.1,
+  "bos_token_id": 1,
+  "classifier_proj_size": 256,
+  "conv_bias": true,
+  "conv_dim": [
+    512,
+    512,
+    512,
+    512,
+    512,
+    512,
+    512
+  ],
+  "conv_kernel": [
+    10,
+    3,
+    3,
+    3,
+    3,
+    2,
+    2
+  ],
+  "conv_stride": [
+    5,
+    2,
+    2,
+    2,
+    2,
+    2,
+    2
+  ],
+  "ctc_loss_reduction": "sum",
+  "ctc_zero_infinity": false,
+  "do_stable_layer_norm": true,
+  "eos_token_id": 2,
+  "feat_extract_activation": "gelu",
+  "feat_extract_dropout": 0.0,
+  "feat_extract_norm": "layer",
+  "feat_proj_dropout": 0.1,
+  "feat_proj_layer_norm": true,
+  "final_dropout": 0.0,
+  "gradient_checkpointing": false,
+  "hidden_act": "gelu",
+  "hidden_dropout": 0.1,
+  "hidden_size": 1024,
+  "initializer_range": 0.02,
+  "intermediate_size": 4096,
+  "layer_norm_eps": 1e-05,
+  "layerdrop": 0.1,
+  "mask_channel_length": 10,
+  "mask_channel_min_space": 1,
+  "mask_channel_other": 0.0,
+  "mask_channel_prob": 0.0,
+  "mask_channel_selection": "static",
+  "mask_feature_length": 10,
+  "mask_feature_min_masks": 0,
+  "mask_feature_prob": 0.0,
+  "mask_time_length": 10,
+  "mask_time_min_masks": 2,
+  "mask_time_min_space": 1,
+  "mask_time_other": 0.0,
+  "mask_time_prob": 0.075,
+  "mask_time_selection": "static",
+  "model_type": "hubert",
+  "num_attention_heads": 16,
+  "num_conv_pos_embedding_groups": 16,
+  "num_conv_pos_embeddings": 128,
+  "num_feat_extract_layers": 7,
+  "num_hidden_layers": 24,
+  "pad_token_id": 0,
+  "tokenizer_class": "Wav2Vec2CTCTokenizer",
+  "torch_dtype": "float32",
+  "transformers_version": "4.28.1",
+  "use_weighted_layer_sum": false,
+  "vocab_size": 32
+}

fairseq/model.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:9f1046daff2169846e024d0dab6214a67e768d0c3116e4be94cabd7bfb645889
+size 1266606909

preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,9 @@

+{
+  "do_normalize": true,
+  "feature_extractor_type": "Wav2Vec2FeatureExtractor",
+  "feature_size": 1,
+  "padding_side": "right",
+  "padding_value": 0.0,
+  "return_attention_mask": true,
+  "sampling_rate": 16000
+}

pytorch_model.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6319cee367d17923b8dee987b9813bc8c9c70c7572bf828c59b4d628f177aaee
+size 1261891557

rinna.png ADDED Viewed