Opus-MT tc-big for Scribe SV (CTranslate2)

CTranslate2 conversions of two Helsinki-NLP Opus-MT tc-big models. They are used by Scribe SV, a Windows dictation and translation utility, and are published here so the application can download them on demand instead of shipping them inside the installer.

Nothing was retrained or fine-tuned: these are format conversions of the original models, quantized to int8_float16 with ctranslate2-converter.

Contents

folder direction source model
opus-mt-tc-big-zle-en-ct2 East Slavic (ru, uk, be) to English Helsinki-NLP/opus-mt-tc-big-zle-en
opus-mt-tc-big-en-zle-ct2 English to East Slavic (ru, uk, be) Helsinki-NLP/opus-mt-tc-big-en-zle

Each folder holds a complete CTranslate2 model: model.bin, config.json, shared_vocabulary.json, the SentencePiece models (source.spm, target.spm) and the tokenizer files.

Language token

opus-mt-tc-big-en-zle is a multi-target model. When translating into East Slavic, the source sentence must start with a target-language token, e.g. >>rus<< for Russian:

>>rus<< The build failed.

The other direction needs no token.

Attribution and license

Original models by the Helsinki-NLP group (Language Technology Research Group at the University of Helsinki), released under CC-BY-4.0. These conversions keep the same license, and attribution to Helsinki-NLP is required when using them.

  • Tiedemann, J., Thottingal, S. OPUS-MT — Building open translation services for the World. EAMT 2020.
  • Tiedemann, J. The Tatoeba Translation Challenge — Realistic Data Sets for Low Resource and Multilingual MT. WMT 2020.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for crash-sv/scribe-translate-ct2

Finetuned
(2)
this model