first commit

Browse files

Files changed (5) hide show

README.md +62 -0
config.json +71 -0
fairseq/model.pt +3 -0
pytorch_model.bin +3 -0
rinna.png +0 -0

README.md ADDED Viewed

	@@ -0,0 +1,62 @@

+---
+language: ja
+datasets:
+  - reazon-research/reazonspeech
+tags:
+  - hubert
+  - speech
+license: apache-2.0
+---
+# japanese-hubert-base
+![rinna-icon](./rinna.png)
+This is a Japanese HuBERT (Hidden Unit Bidirectional Encoder Representations from Transformers) model trained by [rinna Co., Ltd.](https://rinna.co.jp/)
+This model was traind using a large-scale Japanese audio dataset, [ReazonSpeech](https://huggingface.co/datasets/reazon-research/reazonspeech) corpus.
+## How to use the model
+```python
+import torch
+from transformers import HubertModel
+model = HubertModel.from_pretrained("rinna/japanese-hubert-base")
+model.eval()
+wav_input_16khz = torch.randn(1, 10000)
+outputs = model(wav_input_16khz)
+print(f"Input:   {wav_input_16khz.size()}")  # [1, 10000]
+print(f"Output:  {outputs.last_hidden_state.size()}")  # [1, 31, 768]
+```
+## Model summary
+The model architecture is the same as the [original HuBERT base model](https://huggingface.co/facebook/hubert-base-ls960), which contains 12 transformer layers with 8 attention heads.
+The model was trained using code from the [official repository](https://github.com/facebookresearch/fairseq/tree/main/examples/hubert), and the detailed training configuration can be found in the same repository and the [original paper](https://ieeexplore.ieee.org/document/9585401).
+A fairseq checkpoint file can also be available [here](https://huggingface.co/rinna/japanese-hubert-base/tree/main/fairseq).
+## Training
+The model was trained on approximately 19,000 hours of [ReazonSpeech](https://huggingface.co/datasets/reazon-research/reazonspeech) corpus.
+## License
+[The Apache 2.0 license](https://www.apache.org/licenses/LICENSE-2.0)
+## Citation
+```bibtex
+@article{hubert2021hsu,
+  author={Hsu, Wei-Ning and Bolte, Benjamin and Tsai, Yao-Hung Hubert and Lakhotia, Kushal and Salakhutdinov, Ruslan and Mohamed, Abdelrahman},
+  journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing},
+  title={HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units},
+  year={2021},
+  volume={29},
+  number={},
+  pages={3451-3460},
+  doi={10.1109/TASLP.2021.3122291}
+}
+```

config.json ADDED Viewed

	@@ -0,0 +1,71 @@

+{
+  "activation_dropout": 0.1,
+  "apply_spec_augment": true,
+  "architectures": [
+    "HubertModel"
+  ],
+  "attention_dropout": 0.1,
+  "bos_token_id": 1,
+  "classifier_proj_size": 256,
+  "conv_bias": false,
+  "conv_dim": [
+    512,
+    512,
+    512,
+    512,
+    512,
+    512,
+    512
+  ],
+  "conv_kernel": [
+    10,
+    3,
+    3,
+    3,
+    3,
+    2,
+    2
+  ],
+  "conv_stride": [
+    5,
+    2,
+    2,
+    2,
+    2,
+    2,
+    2
+  ],
+  "ctc_loss_reduction": "sum",
+  "ctc_zero_infinity": false,
+  "do_stable_layer_norm": false,
+  "eos_token_id": 2,
+  "feat_extract_activation": "gelu",
+  "feat_extract_norm": "group",
+  "feat_proj_dropout": 0.0,
+  "feat_proj_layer_norm": true,
+  "final_dropout": 0.1,
+  "hidden_act": "gelu",
+  "hidden_dropout": 0.1,
+  "hidden_size": 768,
+  "initializer_range": 0.02,
+  "intermediate_size": 3072,
+  "layer_norm_eps": 1e-05,
+  "layerdrop": 0.1,
+  "mask_feature_length": 10,
+  "mask_feature_min_masks": 0,
+  "mask_feature_prob": 0.0,
+  "mask_time_length": 10,
+  "mask_time_min_masks": 2,
+  "mask_time_prob": 0.05,
+  "model_type": "hubert",
+  "num_attention_heads": 12,
+  "num_conv_pos_embedding_groups": 16,
+  "num_conv_pos_embeddings": 128,
+  "num_feat_extract_layers": 7,
+  "num_hidden_layers": 12,
+  "pad_token_id": 0,
+  "torch_dtype": "float32",
+  "transformers_version": "4.28.1",
+  "use_weighted_layer_sum": false,
+  "vocab_size": 32
+}

fairseq/model.pt ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:dade3cf824ae0d214f7de8b73e70bae7c101e81f12d93577c4760bf516db4063
+size 378888853

pytorch_model.bin ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6c023ccb71e4c2b5a324c94fc5ebe12403d3081c5f370df229892419996fd113
+size 377554841

rinna.png ADDED Viewed