facebook
/

seamless-m4t-unity-small

fairseq2

SeamlessM4T

Model card Files Files and versions Community

sanchit-gandhi commited on Aug 24, 2023

Commit

dbc578f

1 Parent(s): a27233a

Update weights and README.md

Browse files

Files changed (2) hide show

.gitattributes +1 -0
README.md +11 -12

.gitattributes CHANGED Viewed

@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+*.ptl filter=lfs diff=lfs merge=lfs -text

README.md CHANGED Viewed

@@ -15,7 +15,8 @@ SeamlessM4T covers:
 - 🗣️ 35 languages for speech output.
 Apart from [SeamlessM4T-LARGE (2.3B)](https://huggingface.co/facebook/seamless-m4t-large) and [SeamlessM4T-MEDIUM (1.2B)](https://huggingface.co/facebook/seamless-m4t-medium) models, we are also developing a small model (281M) targeting for on-device inference.
-[This folder](https://huggingface.co/facebook/seamless-m4t-unity-small) contains an example to run an exported small model covering most tasks (ASR/S2TT/S2ST). The model could be executed on popular mobile devices with Pytorch Mobile (https://pytorch.org/mobile/home/).
 ## Overview
 | Model   | Checkpoint | Num Params | Disk Size | Supported Tasks         | Supported Languages|
@@ -25,30 +26,28 @@ Apart from [SeamlessM4T-LARGE (2.3B)](https://huggingface.co/facebook/seamless-m
 UnitY-Small-S2T is a pruned version of UnitY-Small without 2nd pass unit decoding.
-Note: If using pytorch runtime in python, only **pytorch<=1.11.0** is supported for **UnitY-Small(281M)**. We tested UnitY-Small-S2T(235M), it works with later versions.
 ## Inference
 To use exported model, users don't need seamless_communication or fairseq2 dependency.
 ```python
 import torchaudio
 import torch
-audio_input, _ = torchaudio.load(TEST_AUDIO_PATH) # Load waveform using torchaudio
-s2t_model = torch.jit.load("unity_on_device_s2t.ptl") # Load exported S2T model
-text = s2t_model(audio_input, tgt_lang=TGT_LANG) # Forward call with tgt_lang specified for ASR or S2TT
-print(f"{lang}:{text}")
 s2st_model = torch.jit.load("unity_on_device.ptl")
-text, units, waveform = s2st_model(audio_input, tgt_lang=TGT_LANG) # S2ST model also returns waveform
-print(f"{lang}:{text}")
-torchaudio.save(f"{OUTPUT_FOLDER}/{lang}.wav", waveform.unsqueeze(0), sample_rate=16000) # Save output waveform to local file
 ```
 Also running the exported model doesn't need python runtime. For example, you could load this model in C++ following [this tutorial](https://pytorch.org/tutorials/advanced/cpp_export.html), or building your own on-device applications similar to [this example](https://github.com/pytorch/ios-demo-app/tree/master/SpeechRecognition)
 # Citation
-If you use SeamlessM4T in your work or any models/datasets/artifacts published in SeamlessM4T, please cite :
 ```bibtex
 @article{seamlessm4t2023,
@@ -60,4 +59,4 @@ If you use SeamlessM4T in your work or any models/datasets/artifacts published i
 ```
 # License
-seamless_communication is CC-BY-NC 4.0 licensed

 - 🗣️ 35 languages for speech output.
 Apart from [SeamlessM4T-LARGE (2.3B)](https://huggingface.co/facebook/seamless-m4t-large) and [SeamlessM4T-MEDIUM (1.2B)](https://huggingface.co/facebook/seamless-m4t-medium) models, we are also developing a small model (281M) targeting for on-device inference.
+This README contains an example to run an exported small model covering most tasks (ASR/S2TT/S2ST). The model could be executed on popular mobile devices with Pytorch Mobile (https://pytorch.org/mobile/home/).
 ## Overview
 | Model   | Checkpoint | Num Params | Disk Size | Supported Tasks         | Supported Languages|
 UnitY-Small-S2T is a pruned version of UnitY-Small without 2nd pass unit decoding.
 ## Inference
 To use exported model, users don't need seamless_communication or fairseq2 dependency.
 ```python
 import torchaudio
 import torch
+audio_input, _ = torchaudio.load(TEST_AUDIO_PATH) # Load waveform using torchaudio
 s2st_model = torch.jit.load("unity_on_device.ptl")
+with torch.no_grad():
+    text, units, waveform = s2st_model(audio_input, tgt_lang=TGT_LANG) # S2ST model also returns waveform
+print(text)
+torchaudio.save(f"{OUTPUT_FOLDER}/result.wav", waveform.unsqueeze(0), sample_rate=16000) # Save output waveform to local file
 ```
 Also running the exported model doesn't need python runtime. For example, you could load this model in C++ following [this tutorial](https://pytorch.org/tutorials/advanced/cpp_export.html), or building your own on-device applications similar to [this example](https://github.com/pytorch/ios-demo-app/tree/master/SpeechRecognition)
 # Citation
+If you use SeamlessM4T in your work or any models/datasets/artifacts published in SeamlessM4T, please cite:
 ```bibtex
 @article{seamlessm4t2023,
 ```
 # License
+seamless_communication is CC-BY-NC 4.0 licensed